Speech synthesis system and control method thereof, and training method of the speech synthesis system
Patent Information
- Application Number
- KR1020240118325
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2026-09-23
- Estimated Expiration
- 2044-09-02
Smart Images

Figure 112024095846440-PAT00046_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a speech synthesis system, a method for controlling the same, and a method for learning a speech synthesis system. More specifically, the present invention relates to a method for learning a vocoder of a speech synthesis system. Background Technology
[0002] With the advancement of artificial intelligence, the utilization of speech synthesis technology is steadily increasing, and speech synthesis systems are being effectively utilized in Text-to-Speech (TTS) systems. TTS is a technology that synthesizes text into natural and easy-to-understand speech, and it is becoming increasingly important across various fields. TTS technology significantly improves accessibility by providing various text materials, such as web pages, documents, and e-books, in audio format to the visually impaired or those with reading difficulties. Furthermore, in various applications such as customer service, navigation systems, and voice assistants, TTS technology can enhance efficiency and productivity by synthesizing text information into speech and delivering it to users without human intervention.
[0003] Meanwhile, vocoders in speech synthesis systems used in TTS systems and the like are evolving into Neural Vocoders and Universal Vocoders through the application of deep learning technology (Neural Vocoders and Universal Vocoders are deep learning-based vocoders, so they will all be referred to as “Neural Vocoders”). These neural vocoders are trained to accurately estimate target speech when receiving acoustic features extracted from various target speech signals. However, when actually used in fields where speech synthesis systems are applied, such as TTS systems, they are configured to estimate speech signals by receiving acoustic features generated by an acoustic model. Nevertheless, since the characteristics of the data used for training differ from those used for inference, the quality of speech synthesis may degrade. Conventionally, to address this, a process of additional training (fine-tuning) the neural vocoder using features generated by the acoustic model was performed. However, this method is limited to fine-tuning only on the features generated by the acoustic model; consequently, if multiple acoustic models exist, fine-tuning must be performed for each model to create multiple neural vocoders, which presents a limitation. Consequently, since multiple neural vocoders must be trained and applied to the service, significant time and resources are required. Therefore, efficient training of neural vocoders must be considered. The problem to be solved
[0004] The present invention aims to provide a speech synthesis system capable of flexibly responding to various acoustic features, a method for controlling the same, and a method for learning the speech synthesis system.
[0005] More specifically, the present invention is intended to provide a neural vocoder that can be universally utilized for various acoustic characteristics.
[0006] Furthermore, the present invention aims to provide a training method for a neural vocoder that solves the problem of sound quality degradation and enables the generation of high-quality speech. means of solving the problem
[0007] To solve the problem described above, a learning method for a speech synthesis system according to the present invention may include the steps of: extracting first type acoustic feature data from a target speech waveform corresponding to correct answer data; processing the first type acoustic feature data obtained by the extraction as an input to a smoothing filter; obtaining smoothed second type acoustic feature data corresponding to the first type acoustic feature data from the smoothing filter; and learning a vocoder of the speech synthesis system using a generated speech waveform created using at least a portion of the first type acoustic feature data and the second type acoustic feature data, and the target speech waveform.
[0008] Furthermore, in a speech synthesis system according to the present invention comprising a storage unit and a vocoder, the vocoder includes a smoothing filter, and the speech synthesis system extracts first type acoustic feature data from a target speech waveform corresponding to correct answer data and stores it in the storage unit, processes the first type acoustic feature data as an input to the smoothing filter, obtains smoothed second type acoustic feature data corresponding to the first type acoustic feature data from the smoothing filter and stores it in the storage unit, and trains the vocoder of the speech synthesis system using a generated speech waveform created using at least a portion of the first type acoustic feature data and the second type acoustic feature data and the target speech waveform.
[0009] Furthermore, a program that is executed by one or more processes in an electronic device according to the present invention and stored in a computer-readable medium may include instructions for performing the steps of: extracting first type acoustic feature data from a target voice waveform corresponding to correct answer data; processing the first type acoustic feature data obtained by the extraction as an input to a smoothing filter; obtaining smoothed second type acoustic feature data corresponding to the first type acoustic feature data from the smoothing filter; and training a vocoder of the voice synthesis system using a generated voice waveform created using at least a portion of the first type acoustic feature data and the second type acoustic feature data, and the target voice waveform.
[0010] Furthermore, according to the control method of a speech synthesis system according to the present invention, the method may include the steps of receiving text to be synthesized from a user terminal, processing the text as input to an acoustic model, acquiring acoustic feature data as output to the acoustic model, processing the acoustic feature data as input to a vocoder trained using the trained acoustic feature data and smoothed feature data obtained by smoothing the trained acoustic feature data, and acquiring a speech waveform corresponding to the text from the trained vocoder.
[0011] Furthermore, the speech synthesis system according to the present invention comprises an input unit that receives text to be synthesized, an acoustic model that receives the text as input and generates acoustic feature data, and a vocoder trained using the learned acoustic feature data and smoothed feature data obtained by smoothing the learned acoustic feature data, wherein the trained vocoder receives the acoustic feature data as input and generates a speech waveform corresponding to the text.
[0012] Furthermore, a program stored in a computer-readable medium, which is executed by one or more processes in an electronic device according to the present invention, may include instructions for performing the steps of: receiving text to be synthesized from a user terminal; processing said text as input to an acoustic model; acquiring said acoustic feature data as output to the acoustic model; processing said acoustic feature data as input to a vocoder trained using said acoustic feature data and said smoothed acoustic feature data; and acquiring a speech waveform corresponding to said text from said vocoder. Effects of the invention
[0013] As seen above, according to the speech synthesis system and the control method thereof and the learning method of the speech synthesis system of the present invention, by providing a neural vocoder that has learned various acoustic features, a single neural vocoder can be utilized for various speech synthesis and conversion.
[0014] Furthermore, according to the speech synthesis system, the control method thereof, and the learning method of the speech synthesis system of the present invention, learning of various learning data can be provided by augmenting various acoustic features using a smoothing filter. Through this, the neural vocoder learns various forms of input data and can maintain high-quality speech synthesis performance even with various inputs during the actual usage phase.
[0015] Furthermore, according to the speech synthesis system, the control method thereof, and the learning method of the speech synthesis system of the present invention, acoustic features of various smoothing levels can be learned to solve the problem of sound quality degradation in actual usage environments. That is, according to the present invention, the ability of a neural vocoder to process smoothed acoustic features can be enhanced, thereby improving the naturalness and quality of the final speech output. Brief explanation of the drawing
[0016] FIG. 1 is a block diagram illustrating a speech synthesis system according to the present invention. FIGS. 2 and FIGS. 3 are conceptual diagrams for explaining a vocoder according to the present invention. FIG. 4 is a flowchart illustrating a vocoder according to the present invention. FIGS. 5, FIGS. 6 and FIGS. 7 are conceptual diagrams for explaining a learning method according to the present invention. FIGS. 8, FIGS. 9, and FIGS. 10 are formulas for explaining an algorithm related to the learning of a neural vocoder according to the present invention. FIG. 10 is a flowchart illustrating a speech synthesis method of a speech synthesis system according to the present invention. Specific details for implementing the invention
[0017] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components are assigned the same reference number regardless of the drawing symbols, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles. Furthermore, in describing embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the spirit and technical scope of the present invention.
[0018] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0019] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0020] A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0021] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0022] The present invention relates to a speech synthesis system, a method for controlling the same, and a method for learning the speech synthesis system. The speech synthesis system according to the present invention may also be referred to as a speech synthesis system. The speech synthesis system according to the present invention may be a Text-to-Speech (TTS) system that generates speech from text. Furthermore, the speech synthesis system according to the present invention may be a system that converts or synthesizes speech from speech.
[0023] The speech synthesis system according to the present invention includes a vocoder, and the present invention aims to provide a vocoder capable of generating high-quality speech (speech waveform).
[0024] Hereinafter, the details will be examined in more detail with reference to the attached drawings. FIG. 1 is a block diagram for explaining a speech synthesis system according to the present invention, and FIGS. 2 and 3 are conceptual diagrams for explaining a vocoder according to the present invention. Furthermore, FIG. 4 is a flowchart for explaining a vocoder according to the present invention, and FIGS. 5, 6, and 7 are conceptual diagrams for explaining a learning method according to the present invention. Furthermore, FIGS. 8, 9, and 10 are formulas for explaining an algorithm related to the learning of a neural vocoder according to the present invention, and FIG. 10 is a flowchart for explaining a speech synthesis method of a speech synthesis system according to the present invention.
[0025] Hereinafter, the speech synthesis system (1000) according to the present invention will be described assuming it is a TTS system. Meanwhile, as previously discussed, if the speech synthesis system (1000) according to the present invention is a system that synthesizes or generates speech from speech, it is possible to exclude some of the components described in FIG. 1.
[0026] First, as illustrated in FIG. 1, a speech synthesis system (1000) according to the present invention may include an input unit (100), a storage unit (200), an acoustic model (300), and a vocoder (400).
[0027] Although not illustrated, the speech synthesis system (1000) according to the present invention may include one or more processors, and such processors may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), neural network processing units (NPUs), application integrated circuits, application semiconductors (ASICs), etc.). One or more processors may be configured to execute instructions contained in the storage unit (200), computer-readable instructions and / or other instructions described herein.
[0028] Meanwhile, the input unit (100) can be configured in various types as a means of inputting data. For example, when a voice synthesis system (1000) that has completed learning is utilized, the input unit (100) can be configured to receive user input. The input unit (100) can be configured to receive user input from a user terminal. The input unit (100) can also be referred to as a user interface module. The input unit (100) may include a touch screen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, or other similar devices. In the present invention, there is no limitation on the type of input unit (100).
[0029] Here, the user input may be text (2000) itself and may correspond to speech. In this case, the speech synthesis system may further include a module that synthesizes speech into text.
[0030] Next, the storage unit (200) serves to store various data and may include one or more non-transient computer-readable storage media that can be read and / or accessed by at least one of one or more processors.
[0031] One or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the storage unit (200) may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage device), whereas in other examples, the storage unit (200) may be implemented using two or more physical devices.
[0032] The storage unit (200) may include computer-readable instructions and additional data. The storage unit may include a storage necessary to perform at least some of the methods, scenarios, and techniques described herein and / or at least some of the functions of the device and network. Furthermore, at least some of the storage unit (200) may be a cloud storage or a cloud server. At least some of the data and training data corresponding to user input received from the input unit (100) may be stored in the storage unit (200).
[0033] Next, the acoustic model (300) may be configured to generate acoustic feature (or feature) data (or referred to as acoustic feature, acoustic feature data, acoustic feature, etc.) from the text (2000). The acoustic model (300) may also be referred to as an “acoustic model.” The acoustic model (300) may be configured to convert (or synthesize) the text input into acoustic feature data. That is, the acoustic model (300) is configured to perform the process of converting the text into acoustic features of speech.
[0034] The acoustic model (300) can generate acoustic feature data based on text by performing text encoding and acoustic encoding. In text encoding, the acoustic model (300) converts (or synthesizes) the text into a format that a machine can understand, and in acoustic encoding, it can generate acoustic feature data based on the encoded text. In the present invention, various types of acoustic models (300) can be utilized, for example, acoustic models such as Tacotron, Tacotron2, Deep Voice, FastSpeech, Transformer TTS, and ClariNet can be utilized.
[0035] Meanwhile, as illustrated in FIG. 5, acoustic feature data can be configured to include various features of speech. For example, acoustic feature data may be two-dimensional data (or two-dimensional array data) composed of a time axis and a frequency axis. Acoustic feature data may be configured to include energy, which can be expressed as a change in color, as shown in the illustration.
[0036] Two-dimensional data consists of a time axis and a frequency axis, and can represent temporal changes of the speech signal based on the time axis and frequency components of the speech signal based on the frequency axis.
[0037] Acoustic feature data may be a Mel-spectrogram, as illustrated in FIG. 5, as an example. A Mel-spectrogram consists of two-dimensional array data representing acoustic characteristics in the time-frequency domain based on the time axis and the frequency axis.
[0038] Next, the vocoder (400) performs the role of converting (or synthesizing) acoustic feature data (or acoustic feature, acoustic feature) into a speech waveform (3000). The vocoder (400) can be utilized for speech synthesis, text-to-speech (TTS) systems, and speech modulation. The vocoder receives acoustic feature data as input and can output a speech waveform (3000) as output. The speech waveform (3000) represents the amplitude change of a speech signal over time, and the vocoder (400) can generate the speech waveform (3000) based on the input acoustic feature data (e.g., Mel spectrogram). In the present invention, for convenience of explanation, the output of the vocoder (400) is referred to as a speech waveform or speech waveform data. The voice waveform data (3000) is a signal representing the amplitude change of a voice signal over time, as described above, and can be stored in formats such as WAV or PCM. The voice waveform data is a sampled digital signal, and each sample can be configured to represent the amplitude of the voice signal at a specific moment in time.
[0039] Meanwhile, the vocoder (400) according to the present invention is a vocoder trained based on deep learning, and can be named a neural vocoder or a universal vocoder.
[0040] The vocoder (400) according to the present invention processes speech signals using deep learning technology and can synthesize and learn features of speech signals using deep learning models such as Convolutional Neural Networks (CNNs) or Recurrent Neural Networks (RNNs). Various models such as UnivNet, WaveNet, WaveGlow, and LPCNet can be applied to the vocoder (400) according to the present invention.
[0041] Meanwhile, a universal vocoder is a vocoder that is not limited to a specific speaker or a specific voice dataset and can handle various speakers and voice data. The vocoder (400) according to the present invention can be configured to learn various voice data so that it can be utilized as a universal vocoder. In the present invention, various universal vocoder models such as WaveNet, WaveGlow, and LPCNet can be applied.
[0042] Meanwhile, the present invention is intended to provide a speech synthesis system capable of flexibly responding to various acoustic features, a control method thereof, and a learning method for the speech synthesis system. More specifically, the present invention is intended to provide a vocoder that can be universally utilized for various acoustic features, and below, the learning method of the vocoder (400) will be examined in more detail.
[0043] As shown in FIGS. 2 and 3, the present invention uses a target speech waveform (or target speech waveform data, 401) and a generated speech waveform (or generated speech waveform data, 404) to train a vocoder (400). In the training stage, the speech synthesis system (1000) may be configured to further include at least one of an acoustic feature extractor (410, or Feature Extraction), a smoothing filter (420, or random smoothing filter (H)), and a learning unit (430). Furthermore, as shown in FIG. 3, in the inference stage, the acoustic feature extractor (410), the smoothing filter (420), and the learning unit (430) may be excluded from the speech synthesis system (1000). In the inference unit, the speech synthesis system (1000) may further include an acoustic model (300) that receives user input (e.g., text) and generates it into acoustic feature data. In the inference unit, the speech synthesis system (1000) can be applied to a service for a user. Below, a method for training a vocoder in the learning unit is described.
[0044] As described, the acoustic feature extractor (410) may be configured to extract acoustic feature data (610) from target voice waveform data (401) corresponding to the correct answer data. For convenience of explanation, the acoustic feature data (610) extracted from the target voice waveform data (401) in the present invention is referred to as the first type of acoustic feature data (610). The first type of acoustic feature data (610) may be stored in the storage unit (200).
[0045] Furthermore, a smoothing filter (or random smoothing filter, 420) can receive a first type of acoustic feature data (610) as input and output a second type of acoustic feature data (620). The smoothing filter (420) can select a filter size for the first type of acoustic feature data (610) smoothing filter and smooth the first type of acoustic feature data (610) with the selected filter size to generate the second type of acoustic feature data (620). The first type of acoustic feature data (610) that has passed through the smoothing filter (420) having the selected filter size can be converted into the second type of acoustic feature data (620). The smoothing filter (420) can be composed of a low-pass filter (LPF). For example, the smoothing filter (420) may be a two-dimensional triangular low-pass filter and may be a filter applied two-dimensionally along the time axis and frequency axis. The triangular filter may have a shape in which the value of the filter decreases linearly as it moves away from the center in each direction. FIG. 5 (a) is a first type of acoustic feature data extracted from a target speech waveform corresponding to the correct data, and FIG. 5 (b), (c), and (d) are examples of a second type of acoustic feature data obtained by smoothing the first type of acoustic feature data.
[0046] According to the city, after passing through a smoothing filter (420) a region (510, 520) of the first type of acoustic feature data shown in FIG. 5 (a), it can be seen that the regions (530, 540, 550, 560) corresponding to the region (510, 520) are smoothed as shown in FIG. 5 (b), (c), and (d).
[0047] As previously described, the first type of acoustic feature data (610) and the second type of acoustic feature data (620) are two-dimensional array data representing a voice signal corresponding to a target voice waveform in terms of the time axis and the frequency axis. The selection of the filter size of the smoothing filter may mean selecting the filter size in the time axis and the filter size in the frequency axis, respectively. The smoothing filter (420) may select the size of the smoothing filter based on the time axis and the frequency axis based on a pre-set filter selection algorithm (or filter application algorithm). In the present invention, the vocoder (400) is trained through a plurality of learning steps, and in the present invention, the smoothing filter may select a different filter size for each learning step where input of the second type of acoustic feature data (620) is required, and process the second type of acoustic feature data (620) with a smoothing filter of a different size applied as input to the vocoder (400).
[0048] In this way, in the present invention, a first type of acoustic feature data (610) is input to a smoothing filter (420), and as the output of the smoothing filter (420), a second type of smoothed acoustic feature data (620) corresponding to the first type of acoustic feature data (610) can be obtained.
[0049] In the learning phase, the vocoder (400) receives acoustic feature data of the first type (610) or acoustic feature data of the second type (620) during multiple learning steps, and can generate a generated voice waveform (or generated voice waveform data, 404) using the received acoustic feature data.
[0050] The learning unit (430) can learn a vocoder using a generated voice waveform (or generated voice waveform data, 404) generated using at least some of the first type of acoustic feature data (610) and the second type of acoustic feature data (620), and a target voice waveform (401).
[0051] When the first type of acoustic feature data (610) is input to the vocoder (400), the vocoder (400) can generate a generated voice waveform (404) using the first type of acoustic feature data (610). Then, the learning unit (430) can train the vocoder (400) by comparing the generated voice waveform (404) generated using the first type of acoustic feature data (610) with the target voice waveform (401) so that the difference between the generated voice waveform (404) and the target voice waveform (401) becomes smaller.
[0052] And, when second type acoustic feature data (620) is input to the vocoder (400), the vocoder (400) can generate a generated voice waveform (404) using the second type acoustic feature data (620). Then, the learning unit (430) can train the vocoder (400) by comparing the generated voice waveform (404) generated using the second type acoustic feature data (620) with the target voice waveform (401) so that the difference between the generated voice waveform (404) and the target voice waveform (401) becomes smaller. In this invention, even when the second type of acoustic feature data (620) is input to the vocoder (400), the learning unit (430) can perform learning between the generated voice waveform (404) and the target voice waveform (401) generated from the vocoder (400) with the second type of acoustic feature data (620) input. In this way, when learning the vocoder (400) in this invention, the correct answer data for the generated voice waveform (404) corresponding to the second type of acoustic feature (420) corresponding to the smoothed data can be set as the target voice waveform (401). Accordingly, the vocoder (400) can generate a high-quality voice waveform even for actual low-quality acoustic feature data.
[0053] Meanwhile, the present invention generates second type acoustic feature data using a smoothing filter (420) and utilizes it for learning the vocoder (400), so it can be expressed as “learning the vocoder (400) using data augmentation.”
[0054] Meanwhile, the learning unit (430) can calculate the loss between the target voice waveform (401) and the generated voice waveform (404) generated by the vocoder (400) using a preset loss function, and perform learning on the vocoder (400) so that the loss is reduced. Learning on the vocoder (400) to reduce the loss can be performed at multiple learning steps in which learning is performed on the vocoder (400). The parameters and weights of the vocoder (400) can be updated at every learning step.
[0055] As such, the learning process for the vocoder (400) according to the present invention may include, as shown in FIG. 4, a process (S100) of extracting a first type of acoustic feature data (610) from a target voice waveform corresponding to the correct answer data, a process (S200) of processing the first type of acoustic feature data (610) obtained by the extraction as an input to a smoothing filter (420), and a process (S300) of obtaining a smoothed second type of acoustic feature data (620) corresponding to the first type of acoustic feature data (610) from the smoothing filter (420). Furthermore, the above learning process may include a process (S400) of learning the vocoder (400) of the speech synthesis system using the generated speech waveform (404) and the target speech waveform (401) generated using at least a portion of the first type of acoustic feature data (610) and the second type of acoustic feature data (620). At this time, the S300 process may be performed optionally when the second type of acoustic feature data (620) needs to be input to the vocoder (400). When the second type of acoustic feature data (620) does not need to be input to the vocoder (400), the first type of acoustic feature data (610) is input to the vocoder (400), and learning between the generated speech waveform (404) and the target speech waveform (401) generated therefrom may be performed.
[0056] Let us examine the learning process for the vocoder (400) in more detail. As illustrated in FIGS. 6 and 7, the vocoder (400) is learned through a plurality of learning steps, and in the present invention, the smoothing filter selects a different filter size for each learning step where input of the second type of acoustic feature data (620) is required, so that the second type of acoustic feature data (620) to which a smoothing filter of a different size is applied can be processed as input to the vocoder (400).
[0057] As shown in the example, in some of the plurality of learning steps, the vocoder (400) is trained using first type acoustic feature data (610), and in the remaining learning steps excluding the aforementioned portion of the plurality of learning steps, the vocoder (400) can be trained using second type acoustic feature data (620). At this time, in at least some of the remaining learning steps, not only the second type acoustic feature data (620) but also the first type acoustic feature data (610) can be input into the vocoder (400) for training.
[0058] Meanwhile, the above-mentioned learning steps may be initial learning steps. In the present invention, the vocoder (400) may be trained on the first type of acoustic feature data (610) during an initial number of learning steps based on the point in time when training for the vocoder (400) begins. In the initial learning steps, training for the vocoder (400) may be performed using a pair of specific first type acoustic feature data (611) and a specific target voice waveform (401a) corresponding to the specific first type acoustic feature data (611) among the first type acoustic feature data and the target voice waveform (401) that are the subjects of training.
[0059] Among multiple learning steps, the number of initial learning steps can be determined by various criteria. For example, the initial learning step may be set to the first 450K steps (450,000 steps). In the initial learning step, training the vocoder (400) with unsmoothed acoustic feature data of the first type is intended to allow the vocoder (400) to learn the basic pattern of the data and generate a stable output. The deep learning model corresponding to the vocoder (400) can learn the basic pattern from the acoustic feature data of the first type corresponding to the target acoustic waveform during the initial learning step. In the case of speech synthesis or synthesis, the model can learn the basic frequency components, rhythm, tone, etc. of the speech at this step.
[0060] And, in the learning step after the initial preset number of times (e.g., remaining learning step), learning for the vocoder (400) can be performed using the second type of acoustic feature data and a pair of specific target voice waveforms (401b) corresponding to the second type of acoustic feature among the target voice waveforms.
[0061] As explained above, in the remaining learning steps, not only the second type of acoustic feature data (620) but also the first type of acoustic feature data (610) can be input into the vocoder (400) and learned. That is, in at least part of the subsequent learning steps (or remaining learning steps), the vocoder (400) can be learned using the first type of acoustic feature data (612), and in the remaining part of the subsequent learning steps (or remaining learning steps), the vocoder (400) can be learned using the second type of acoustic feature data (621, 622).
[0062] Meanwhile, in the present invention, in the remaining learning step among a plurality of learning steps, a specific first type of acoustic feature data to be learned may be selected from among the first type of acoustic feature data. The selected specific first type of acoustic feature data is input to a smoothing filter, and from the smoothing filter, specific second type of acoustic feature data (621, 622) corresponding to the specific first type of acoustic feature data may be obtained. In the remaining learning step, the vocoder (400) may be learned using the specific second type of acoustic feature data and a pair of specific target voice waveforms (401c, 401d) corresponding to the specific first type of acoustic feature data among the target voice waveforms.
[0063] Furthermore, the vocoder (400) can generate a generated voice waveform (404) as illustrated in FIG. 7. The generated voice waveform (404) and the target voice waveform (401) can be input into a learning unit and utilized for learning the vocoder (400). The generated voice waveform data is configured to be generated by a vocoder that has performed learning on at least a portion of the first type of acoustic feature data and the second type of acoustic feature data. In this way, the vocoder can generate generated voice waveforms (404a, 404b, 404c, 404d) corresponding to the input data pairs (first type acoustic data-target voice waveform pairs, or second type acoustic data-target voice waveform pairs) at each learning step. The vocoder (400) can be learned and updated based on the loss between the generated voice waveform and the target voice waveform at each learning step.
[0064] Meanwhile, in the present invention, the learning step may include a process in which a deep learning model corresponding to a vocoder (400) performs waveform generation using a batch of training data, calculates a loss, and then updates weights through the loss. The learning step is performed in batch units. That is, multiple data samples are combined to form a batch, and one step can be performed using this batch.
[0065] Meanwhile, the learning algorithm for training the previously examined vocoder (400) and the filter selection algorithm for selecting the smoothing filter (420) will be examined in detail along with mathematical formulas.
[0066] As an example, the vocoder (G, 400) according to the present invention may be composed of a generator neural network and a discriminator neural network based on a Generative Adversarial Network (GAN). The vocoder (G, 400) may have a structure in which the neural networks learn by opposing each other. The generator may receive data in the form of a Mel-spectogram and generate a speech signal. Furthermore, the generator is trained to receive a sampled noise vector based on actual speech data (correct answer data, target speech waveform (401)) and generate fake speech data (generated speech waveform (404)), and the discriminator may play the role of distinguishing between the fake speech data generated by the generator and the actual speech data. At this time, the discriminator estimates the probability that the given speech is actual speech data, and the generator is trained to generate speech data similar to the actual speech data in order to deceive the discriminator, and the discriminator can distinguish between the actual speech data and the fake speech data more precisely by reducing the loss.
[0067] More specifically, the constructor and discriminator A GAN-based vocoder (G, 400) composed of actual acoustic features Using (Type 1 acoustic feature data, 610), actual speech data You can learn the distribution of (401).
[0068] Specifically, referring to the mathematical formulas illustrated in FIG. 8, as illustrated in the mathematical formulas (a) and (b) of FIG. 8, is a constructor, is a discriminator, is Type 1 acoustic feature data (610, or actual acoustic features), and represents the actual voice data (401, target voice waveform), and z represents the fake voice data generated by the generator (404, generated voice waveform), and and are each constructors and discriminator It can be understood as a loss function. In this case, the generator and discriminator use the cross-entropy loss function, , can calculate.
[0069] As illustrated in the mathematical formulas of Fig. 8 (c) and (d), Is It means distance, is the output of the discriminator, the actual voice waveform It can represent the probability that it is true, also, is data for a fake voice waveform generated by the generator, and ) is fake data generated by the generator by the discriminator can mean the probability that it is true. In this case, to address the situation where learning is not sufficiently achieved using only the loss functions of the generator and discriminator, an auxiliary loss function ( You can use ).
[0070] In this way, a GAN-based vocoder (G, 400) can be trained to generate data similar to actual voice data by adversarially training the generator and the discriminator.
[0071] In the present invention, the vocoder (G, 400) may include a smoothing filter to flexibly respond to various acoustic characteristics and solve the problem of sound quality degradation, thereby enabling the generation of high-quality speech. Here, a random smoothing filter ( 420) is the first type of acoustic feature data ( By augmenting acoustic features in (610), learning on various training data is provided, so that the vocoder (G, 400) can learn various forms of input data and perform high-quality speech synthesis on various inputs during actual use.
[0072] Specifically, the random smoothing filter (420) is convolved with the first type of acoustic feature data (610) to augment the data, and the generated data may be the second type of acoustic feature data (620). Here, the “convolution operation” is an operation used to apply the smoothing filter (420), which is a linear filter, and can generate feature data by combining the input data and the filter.
[0073] For a detailed explanation, referring to mathematical formulas (e), (f), and (g) in FIG. 8, of (a) in FIG. 8 is the first type of acoustic feature data (610) and a smoothing filter ( , 420) may mean acoustic feature data of the second type (620) generated by convolution operation.
[0074] Furthermore, the vocoder (G, 400) of the present invention can be trained by using the first type acoustic feature data (610) and the second type acoustic feature data (620) together.
[0075] Specifically, as illustrated in the mathematical formulas (f) and (g) of FIG. 8, when the vocoder (G, 400) of the present invention is trained with second-type acoustic feature data (620) in the same manner as the mathematical formulas (a) and (b) of FIG. 8, the vocoder (G, 400) can learn various acoustic features, thereby enabling high-quality speech synthesis to be performed even with various inputs during the actual use phase.
[0076] Meanwhile, in the present invention, a two-dimensional triangular low-pass filter may be used as the random smoothing filter (420). A triangular low-pass filter is a filter that passes only components below a specific frequency in the frequency domain and attenuates components above that frequency. The filter can be used in signal processing to limit the frequency band to reduce noise or to pass only signals of a specific band.
[0077] Specifically, referring to FIG. 9, as illustrated in the mathematical formula of FIG. 9 (a), a triangular low-pass filter ( ) is a time frame( ) filter size( and frequency frame ( filter size ( Defined by, where each filter size and can be sampled randomly. More specifically, as illustrated in the mathematical formulas of FIG. 9(b), FIG. 9(c), and FIG. 9(d), and are each and It can be understood as a group of candidates that could become, Except for the case where a uniform distribution can be shown, a uniform distribution can be exhibited. Based on the formulas shown in FIG. 9, the size of the smoothing filter (420) according to the present invention can be selected.
[0078] one side, Is and It can be, and to maintain learning stability, symmetry between the generator and the discriminator can be maintained. and It can be set to an odd number. and The (x,y) point of the input feature can be set to an odd number so that the (x,y) point of the smoothed feature corresponds to the (x,y) point of the smoothed feature. In the convolution operation, an odd-sized filter always has a unique center element in the center, so that symmetry can be maintained for each input position. For example, in a 3x3 filter, (2, 2) can be the center element, and accordingly, when the first type of acoustic feature data (610) is convolutional with the smoothing filter, the size of the second type of acoustic feature data (620), which is the output data, can be maintained uniformly.
[0079] Meanwhile, based on the formula examined above, in the remaining learning step among the plurality of learning steps, the time-axis filter size and the frequency-axis filter size of the smoothing filter (420) may be randomly selected. Furthermore, each filter size of the smoothing filter selected during the remaining learning step may be configured such that at least one of the time-axis filter size and the frequency-axis filter size is different from each other.
[0080] A vocoder trained according to the learning method described above may be configured in the inference stage, excluding a smoothing filter (420) as shown in FIG. 3. In the inference stage, the speech synthesis system and the control method thereof according to the present invention can generate high-quality speech for text input by a user through the following steps, as shown in FIG. 10: receiving (or inputting) text to be converted (or synthesized) from a user terminal (S1000); processing the text as input to an acoustic model (S2000); acquiring acoustic feature data as output to the acoustic model (S3000); processing the acoustic feature data as input to a vocoder trained using the learned acoustic feature data and smoothed feature data obtained by smoothing the learned acoustic feature data (S4000); and acquiring a speech waveform corresponding to the text from the trained vocoder (S5000).
[0081] As seen above, according to the speech synthesis system and the control method thereof and the learning method of the speech synthesis system of the present invention, by providing a neural vocoder that has learned various acoustic features, a single neural vocoder can be utilized for various speech synthesis and conversion.
[0082] Furthermore, according to the speech synthesis system, the control method thereof, and the learning method of the speech synthesis system of the present invention, learning of various learning data can be provided by augmenting various acoustic features using a smoothing filter. Through this, the neural vocoder learns various forms of input data and can maintain high-quality speech synthesis performance even with various inputs during the actual usage phase.
[0083] Furthermore, according to the speech synthesis system, the control method thereof, and the learning method of the speech synthesis system of the present invention, acoustic features of various smoothing levels can be learned to solve the problem of sound quality degradation in actual usage environments. That is, according to the present invention, the ability of a neural vocoder to process smoothed acoustic features can be enhanced, thereby improving the naturalness and quality of the final speech output.
[0084] Meanwhile, computer-readable media include all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SSD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0085] Furthermore, the computer-readable medium may be a server or cloud storage that includes a storage and is accessible to an electronic device via communication. In this case, the computer may download the program according to the present invention from the server or cloud storage via wired or wireless communication.
[0086] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, namely a CPU (Central Processing Unit), and no special limitations are placed on its type.
[0087] Meanwhile, the above detailed description should not be interpreted restrictively in all respects but should be considered exemplary. The scope of the invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the invention are included within the scope of the invention.
Claims
Claim 1 A learning method for a speech synthesis system comprises: a step of extracting first type acoustic feature data from a target speech waveform corresponding to correct answer data; a step of processing the first type acoustic feature data obtained by the extraction as input to a smoothing filter; and a step of obtaining smoothed second type acoustic feature data corresponding to the first type acoustic feature data from the smoothing filter. A method for learning a voice synthesis system, comprising the step of learning a vocoder of the voice synthesis system using a generated voice waveform and a target voice waveform generated using at least a portion of the first type of acoustic feature data and the second type of acoustic feature data, wherein in the learning step, learning of the vocoder is performed through a plurality of learning steps, wherein in some of the plurality of learning steps, the vocoder is learned using the first type of acoustic feature data, and in the remaining learning steps excluding the portion of the learning steps among the plurality of learning steps, the vocoder is learned using the second type of acoustic feature data. Claim 2 A method for learning a speech synthesis system according to claim 1, wherein the learning data for learning the vocoder comprises the first type of acoustic feature data extracted from the target speech waveform and the second type of acoustic feature data obtained from the smoothing filter. Claim 3 delete Claim 4 A method for learning a speech synthesis system according to claim 1, characterized in that, during an initial preset number of learning steps based on the point in time when learning for the vocoder begins, the learning for the vocoder is performed using the first type of acoustic feature data and a specific target speech waveform pair corresponding to the first type of acoustic feature data among the target speech waveforms. Claim 5 A learning method for a speech synthesis system according to claim 4, wherein in the learning step, in the learning step after the initial preset number of times, learning for the vocoder is performed using the second type of acoustic feature data and a specific target speech waveform pair corresponding to the second type of acoustic feature among the target speech waveforms. Claim 6 A method for learning a speech synthesis system according to claim 5, characterized in that, in at least part of the subsequent learning steps, the vocoder is learned using the first type of acoustic feature data, and in the remaining part of the subsequent learning steps, the vocoder is learned using the second type of acoustic feature data. Claim 7 A method for learning a speech synthesis system according to claim 1, wherein in the remaining learning step among the plurality of learning steps, a specific first type of acoustic feature data to be learned is selected among the first type of acoustic feature data, the selected specific first type of acoustic feature data is input to the smoothing filter, and a specific second type of acoustic feature data corresponding to the specific first type of acoustic feature data is obtained from the smoothing filter. Claim 8 A method for learning a speech synthesis system according to claim 7, wherein in the remaining learning step, the vocoder is learned using the specific second type of acoustic feature data and a specific target speech waveform pair corresponding to the specific first type of acoustic feature data among the target speech waveforms. Claim 9 A learning method for a speech synthesis system according to claim 1, wherein the generated speech waveform is generated in the vocoder that has performed learning on at least a portion of the first type of acoustic feature data and the second type of acoustic feature data. Claim 10 A method for learning a speech synthesis system according to claim 9, wherein in the learning step, the loss between the target speech waveform and the generated speech waveform generated by the vocoder is calculated using a preset loss function, and the vocoder is trained to reduce the loss. Claim 11 A learning method for a speech synthesis system according to claim 10, characterized in that the learning of the vocoder to reduce the loss is performed at each of the plurality of learning steps. Claim 12 A learning method for a speech synthesis system according to claim 7, further comprising the step of selecting a filter size of the smoothing filter, and characterized in that acoustic feature data of a specific first type is smoothed using the smoothing filter having the selected filter size. Claim 13 A learning method for a speech synthesis system according to claim 12, wherein the first type of acoustic feature data and the second type of acoustic feature data are two-dimensional array data representing a speech signal corresponding to the target speech waveform in terms of time axis and frequency axis, and in the step of selecting the filter size of the smoothing filter, the time axis filter size and the frequency axis filter size are each selected, and the specific first type of acoustic feature data is smoothed using the smoothing filter having the selected time axis filter size and frequency axis filter size. Claim 14 A learning method for a speech synthesis system according to claim 13, wherein in the remaining learning steps among the plurality of learning steps, the time-axis filter size and the frequency-axis filter size of the smoothing filter are randomly selected. Claim 15 A learning method for a speech synthesis system according to claim 14, wherein each of the filter sizes of the smoothing filter selected during the remaining learning step is characterized in that at least one of the time-axis filter size and the frequency-axis filter size is different from each other. Claim 16 A speech synthesis system comprising a storage unit and a vocoder, wherein the vocoder includes a smoothing filter, and the speech synthesis system extracts first type acoustic feature data from a target speech waveform corresponding to correct answer data and stores it in the storage unit, processes the first type acoustic feature data as an input to the smoothing filter, obtains smoothed second type acoustic feature data corresponding to the first type acoustic feature data from the smoothing filter and stores it in the storage unit, and trains the vocoder of the speech synthesis system using a generated speech waveform created using at least a portion of the first type acoustic feature data and the second type acoustic feature data and the target speech waveform, wherein the training of the vocoder is performed through a plurality of training steps, wherein in some of the plurality of training steps, the vocoder is trained using the first type acoustic feature data, and excluding the some of the plurality of training steps A speech synthesis system characterized in that, in the remaining learning steps, the vocoder is trained using the acoustic feature data of the second type. Claim 17 A program that is executed by one or more processes in an electronic device and stored on a computer-readable medium, wherein the program comprises: a step of extracting first type acoustic feature data from a target voice waveform corresponding to correct answer data; a step of processing the first type acoustic feature data obtained by the extraction as input to a smoothing filter; and a step of obtaining smoothed second type acoustic feature data corresponding to the first type acoustic feature data from the smoothing filter. A program comprising instructions for performing a step of learning a vocoder of a speech synthesis system using a generated speech waveform and a target speech waveform generated using at least a portion of the first type of acoustic feature data and the second type of acoustic feature data, wherein in the learning step, learning of the vocoder is performed through a plurality of learning steps, wherein in some of the plurality of learning steps, the vocoder is learned using the first type of acoustic feature data, and in the remaining learning steps excluding the portion of the plurality of learning steps, the vocoder is learned using the second type of acoustic feature data. Claim 18 A control method for a speech synthesis system comprising: receiving text to be synthesized from a user terminal; processing the text as input to an acoustic model; acquiring acoustic feature data as output to the acoustic model; processing the acoustic feature data as input to a vocoder trained using the trained acoustic feature data and smoothed feature data obtained by smoothing the trained acoustic feature data; and acquiring a speech waveform corresponding to the text from the trained vocoder, wherein the training of the vocoder is performed through a plurality of training steps, wherein in some of the plurality of training steps, the vocoder is trained using the trained acoustic feature data, and in the remaining training steps excluding the portion of the plurality of training steps, the vocoder is trained using the smoothed feature data. Claim 19 A speech synthesis system comprising: an input unit that receives text to be synthesized; an acoustic model that receives said text as input and generates acoustic feature data; and a vocoder trained using said acoustic feature data and smoothed feature data obtained by smoothing said acoustic feature data, wherein said trained vocoder receives said acoustic feature data as input and generates a speech waveform corresponding to said text, and said training of said vocoder is performed through a plurality of training steps, wherein in some of said training steps the vocoder is trained using said acoustic feature data, and in the remaining training steps excluding said some of said training steps the vocoder is trained using said smoothed feature data. Claim 20 A program that is executed by one or more processes in an electronic device and stored on a computer-readable medium, wherein the program comprises instructions for performing the steps of: receiving text to be synthesized from a user terminal; processing the text as input to an acoustic model; acquiring acoustic feature data as output to the acoustic model; processing the acoustic feature data as input to a vocoder trained using the learned acoustic feature data and smoothed feature data obtained by smoothing the learned acoustic feature data; and acquiring a speech waveform corresponding to the text from the trained vocoder, wherein the training of the vocoder is performed through a plurality of learning steps, wherein in some of the plurality of learning steps, the vocoder is trained using the learned acoustic feature data, and in the remaining learning steps excluding the portion of the plurality of learning steps, the vocoder is trained using the smoothed feature data.