Voice waveform generation system, voice waveform generation method and voice waveform generation program

JP2025108262APending Publication Date: 2025-07-23NAT INST OF INFORMATION & COMM TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2024002076
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Existing sequence conversion type end-to-end models for audio waveform generation achieve high performance but fall short in natural speech quality and processing speed, particularly on mobile terminals.

Method used

An audio waveform generation system utilizing a ConvNeXt-based encoder and decoder with depthwise convolutional layers, layer normalization blocks, and Gaussian error linear units, combined with a waveform generation model that includes one-dimensional convolutional layers and transposed convolutional layers, to predict audio waveforms from feature amounts.

Benefits of technology

The system significantly speeds up audio waveform generation while improving quality, achieving faster processing times and enhanced natural speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108262000001_ABST
    Figure 2025108262000001_ABST
Patent Text Reader

Abstract

To provide a series conversion type End-to-end model which can generate a voice waveform faster and improve quality.SOLUTION: A voice waveform generation system includes: an acoustic model for predicting a second feature amount from an inputted first feature amount; and a waveform generation model for predicting a voice waveform from the second feature amount. The acoustic model includes: a decoder for converting the first feature amount into continuous expression; a variance adapter for predicting continuous length of each phoneme from the continuous expression; and an encoder for predicting the second feature amount from output of the variance adapter. Each of the encoder and the decoder includes a depth unit convolution layer, a layer normalization block, a point unit convolution layer and a Gaussian error linear unit.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio waveform generation system, an audio waveform generation method, and an audio waveform generation program.

Background Art

[0002] In recent years, speech synthesis and voice quality conversion technologies using neural networks have advanced significantly. As a result, under certain experimental conditions, high-quality synthesis that is almost indistinguishable from natural speech has become possible.

[0003] Furthermore, a sequence conversion type end-to-end text-to-speech synthesis model that directly generates an audio waveform from text using a single neural network, and a sequence conversion type end-to-end voice quality conversion model that directly generates a converted destination audio waveform from a converted source audio using a single neural network have been proposed.

[0004] More specifically, a Transformer type neural network proposed in a machine translation task is widely used for the encoder and / or decoder of a sequence conversion type text-to-speech synthesis model and / or voice quality conversion model (see Non-Patent Documents 1 and 2, etc.). Since the model using the Transformer type neural network is an autoregressive model, it requires more generation time. However, by learning the input-output alignment with a separate model, a non-autoregressive type high-speed model can be realized (see Non-Patent Documents 3 and 4, etc.). By simultaneously training a non-autoregressive type high-speed model and a neural waveform generation model, a sequence conversion type end-to-end text-to-speech synthesis model and / or a sequence conversion type end-to-end voice quality conversion model using a single neural network can be constructed. Such models have achieved performance superior to conventional models (see Non-Patent Documents 5 and 6, etc.).

Prior Art Documents

Non-Patent Documents

[0005] [Non-Patent Document 1] N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, "Neural speech synthesis with Transformer network," in Proc. AAAI, Jan. 2019, pp. 6706-6713. [Non-Patent Document 2] R. Liu, X. Chen, and X. Wen, "Voice conversion with transformer network," in Proc. ICASSP, May 2020, pp. 7759-7763 [Non-Patent Document 3] Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, "FastSpeech: Fast, robust and controllable text to speech," in Proc. NeurIPS, Dec. 2019, pp. 3165-3174. [Non-Patent Document 4] T. Hayashi, W.-C. Huang, K. Kobayashi, and T. Toda, "Non-autoregressive sequence-to-sequence voice conversion," in Proc. ICASSP, June 2021, pp. 7068-7072. [Non-Patent Document 5] D. Lim, S. Jung, and E. Kim, "JETS: Jointly training Fast-Speech2 and HiFi-GAN for end to end text to speech," in Proc. Interspeech, Sept. 2022, pp. 21-25. [Non-Patent Document 6] T. Okamoto, T. Toda, and H. Kawai, "E2E-S2S-VC: End-to-end sequence-to-sequence voice conversion," in Proc. Interspeech, Aug. 2023, pp. 2043-2047

Summary of the Invention

Problems to be Solved by the Invention

[0006] By using a sequence conversion type end-to-end text-to-speech synthesis model and a sequence conversion type end-to-end voice quality conversion model (hereinafter, also collectively referred to as the "sequence conversion type end-to-end model"), performance exceeding that of conventional models has been achieved, but the quality of natural speech has not been reached yet.

[0007] Also, in the proposed sequence conversion type end-to-end model, even with the processing resources of one core of the CPU, it is possible to generate an audio waveform in real time. However, considering real-time audio waveform generation on a mobile terminal, further speeding up the processing speed is also necessary.

[0008] An object of the present invention is to provide a sequence conversion type end-to-end model that can further speed up audio waveform generation and further improve quality.

Means for Solving the Problems

[0009] An audio waveform generation system according to an embodiment includes an acoustic model that predicts a second feature amount from a first feature amount input thereto, and a waveform generation model that predicts an audio waveform from the second feature amount. The acoustic model includes an encoder for converting the first feature amount into a continuous representation, a variance adapter for predicting the duration of each phoneme from the continuous representation, and a decoder for predicting the second feature amount from the output of the variance adapter. Each of the encoder and the decoder includes a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit.

[0010] The first feature amount may include at least one of a language feature amount generated from text and an acoustic feature amount generated from a mel spectrogram.

[0011] The variance adapter may predict energy and pitch in addition to the duration of each phoneme.

[0012] The waveform generation model may include a one-dimensional convolutional layer, a layer normalization block, a ConvNeXt block, a linearization layer, and a reshaping layer. Each of the ConvNeXt blocks may have the same network structure as the encoder and the decoder.

[0013] The waveform generation model may include a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, and a plurality of transposed convolutional layers arranged between the first one-dimensional convolutional layer and the second one-dimensional convolutional layer.

[0014] The variance adapter and the waveform generation model may be learned together.

[0015] A method for generating an audio waveform according to another embodiment includes predicting a second feature amount from a first input feature amount using an acoustic model, and predicting an audio waveform from the second feature amount using a waveform generation model. The step of predicting the second feature amount includes converting the first feature amount into a continuous representation by an encoder including a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit, predicting the duration of each phoneme from the continuous representation, and predicting the second feature amount from the predicted duration by a decoder including a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit.

[0016] According to yet another embodiment, an audio waveform generation program causes a computer to execute a step of predicting a second feature amount from a first input feature amount using an acoustic model, and a step of predicting an audio waveform from the second feature amount using a waveform generation model. The acoustic model includes an encoder for converting the first feature amount into a continuous representation, a variance adapter for predicting the duration of each phoneme from the continuous representation, and a decoder for predicting the second feature amount from the output of the variance adapter. Each of the encoder and the decoder includes a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit.

Advantages of the Invention

[0017] According to the present invention, a series conversion type End-to-end model is provided that can speed up audio waveform generation and further improve quality.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Best Mode for Carrying Out the Invention

[0019] Embodiments of the present invention will be described in detail with reference to the drawings. For the same or corresponding parts in the drawings, the same reference numerals are given and their descriptions are not repeated.

[0020] A voice waveform generation system including a sequence conversion type End-to-end model (hereinafter also abbreviated as "E2E model") according to the present embodiment is capable of executing a voice synthesis task. The voice synthesis task includes, for example, at least text-to-speech synthesis (TTS) and voice conversion (VC).

[0021] [A. Related Art of Voice Waveform Generation System] First, the related art of the voice waveform generation system according to the present embodiment will be described.

[0022] FIG. 1 is a schematic diagram showing an example of the network configuration of a voice waveform generation system according to the related art. For convenience of explanation, two types of E2E models 3 and 4 are described in parallel in FIG. 1, but it is not necessary to simultaneously implement two types of voice waveform generation systems.

[0023] Referring to FIG. 1, the E2E model 3 corresponds to the JETS model disclosed in Non-Patent Document 5 and predicts a voice waveform from an input of feature quantities (linguistic feature quantities and / or acoustic feature quantities). The E2E model 3 includes an acoustic model 20 (AM: Acoustic model) based on FastSpeech2 and a neural waveform generation model (neural vocoder), HiFi-GAN 30. Both the acoustic model 20 and HiFi-GAN 30 are pre-trained models.

[0024] The E2E model 3 can be used for both text-to-speech (TTS) and voice conversion (VC). When performing the text-to-speech task, it is combined with a language feature generation unit 5 that generates language features from the input text. When performing the voice conversion task, it is combined with an acoustic feature generation unit 6 that generates acoustic features from the mel spectrogram of the converted voice.

[0025] The language feature generation unit 5 includes a text analysis unit 52 and an embedding layer 56. The text analysis unit 52 analyzes the input text 50 and outputs a phoneme sequence 54. The phoneme sequence 54 may include accent information. The embedding layer 56 sequentially arranges the phoneme sequence 54 in a vector space to generate language features.

[0026] The acoustic feature generation unit 6 includes a reshaping layer 62 and a linearization layer 64. The reshaping layer 62 adjusts the waveform indicated by the input mel spectrogram 60. The reshaping layer 62 may compress information with a reduction coefficient r according to the encoder included in the acoustic model 20. e The linearization layer 64 linearizes the waveform (sequence) output from the reshaping layer 62. The linearized sequence becomes the acoustic feature.

[0027] Thus, the features (corresponding to the first features) input to the acoustic model 20 include at least one of the language features generated from the text 50 and the acoustic features generated from the mel spectrogram 60.

[0028] For convenience of explanation, it is depicted that the language features generated by the language feature generation unit 5 and the acoustic features generated by the acoustic feature generation unit 6 are input to the acoustic model 20 through the same path, but individual acoustic models 20 may be prepared according to the language features and acoustic features.

[0029] The acoustic model 20 predicts hidden features (corresponding to the second features) from the input feature quantities (language feature quantities or acoustic feature quantities). The acoustic model 20 includes a Transformer-based encoder 22, a variance adapter 24, and a Transformer-based decoder 26. That is, the encoder 22 and the decoder 26 are FFT-based models with a self-attention mechanism.

[0030] The encoder 22 converts the input feature quantities (discrete representations) into continuous representations. The encoder 22 includes an embedding layer, a self-attention mechanism, and a one-dimensional convolutional layer.

[0031] The variance adapter 24 predicts the duration of each phoneme from the continuous representation. The variance adapter 24 may predict energy and pitch from the continuous representation. The predicted results are added to the output of the encoder 22.

[0032] The decoder 26 predicts hidden features from the output of the variance adapter 24. Similar to the encoder 22, the decoder 26 includes an embedding layer, a self-attention mechanism, and a one-dimensional convolutional layer.

[0033] HiFi-GAN 30 predicts an audio waveform 32 from the hidden features output by the acoustic model 20.

[0034] As shown in FIG. 1, in the E2E model 3, during learning, by adopting an alignment learning framework, during inference, an intermediate mel spectrogram and an external aligner are not required. More specifically, Monotonic Alignment Search (MAS) 90 is used. In the Monotonic Alignment Search 90, the variance adapter 24 is optimized using the output of the encoder 22 and the target mel spectrogram 8. The variance adapter 24 and the voice waveform model (HiFi-GAN 30 or WaveNeXt 40 described later) are learned together. Through learning, an alignment between the target mel spectrogram 8 and the hidden feature (the feature converted from the input text 50 or mel spectrogram 60) is gradually obtained.

[0035] Also, the neural vocoder HiFi-GAN 30 is learned by adversarial training 72 using the HiFi-GAN discriminator 70.

[0036] Next, the E2E model 4 is obtained by changing the neural vocoder from HiFi-GAN 30 to WaveNeXt 40 compared to the E2E model 3. WaveNeXt 40 predicts the voice waveform 32 from the hidden feature output by the acoustic model 20. WaveNeXt 40 is a pre-trained model.

[0037] Here, an example of the network configuration of the neural vocoder will be described.

[0038] FIG. 2 is a schematic diagram showing an example of the network configuration of the neural vocoder. FIG. 2(a) shows an example of the network configuration of HiFi-GAN 30, FIG. 2(b) shows an example of the network configuration of Vocos 40A, and FIG. 2(c) shows an example of the network configuration of WaveNeXt 40. Note that in FIG. 2, T represents the number of frames.

[0039] The HiFi-GAN30 shown in Fig. 2(a) includes a 1D convolutional layer 33 (corresponding to the first 1D convolutional layer), transposed convolutional layers 34, 35, 36, 37, and a 1D convolutional layer 38 (corresponding to the second 1D convolutional layer). Thus, the HiFi-GAN30 includes the 1D convolutional layer 33, the 1D convolutional layer 38, and multiple stages of transposed convolutional layers 34, 35, 36, 37 arranged between the 1D convolutional layer 33 and the 1D convolutional layer 38.

[0040] Each of the transposed convolutional layers 34, 35, 36, 37 has a multiple receptive field (MRF) layer connected to a transposed convolutional block. The waveform is upsampled by the multiple stages of transposed convolutional layers.

[0041] The Vocos40A shown in Fig. 2(b) includes a 1D convolutional layer 43, a layer normalization block 44, a ConvNeXt block 45, a linearization layer 46, and an iSTFT (inverse Short-Time Fourier Transform) layer 47. The linearization layer 46 outputs the amplitude and phase. The iSTFT layer 47 is for upsampling the waveform and generates the waveform based on the amplitude and phase from the linearization layer 46.

[0042] Compared with the Vocos40A shown in Fig. 2(b), the WaveNeXt40 shown in Fig. 2(c) has a learnable linearization layer 48 and a reshape layer 49 arranged instead of the iSTFT layer 47.

[0043] Fig. 3 is a schematic diagram showing an example of the network configuration of the ConvNeXt block 45. Referring to Fig. 3, each of the ConvNeXt blocks 45 includes a depthwise convolutional layer 451, a layer normalization block 452, pointwise convolutional layers 453, 455, a Gaussian error linear unit 454 (GELU), and an addition unit 456.

[0044] The ConvNeXt block 45 shown in FIG. 3 has the same network structure as the ConvNeXt-based encoder 12 and the ConvNeXt-based decoder 16 described later.

[0045] Referring again to FIG. 2, in HiFi-GAN 30, upsampling is realized by a plurality of transposed convolutional layers, while in Vocos 40A shown in FIG. 2(b), a high-resolution STFT spectrum is predicted from the input mel spectrogram by a plurality of ConvNeXt blocks 45. The iSTFT layer 47 directly generates an audio waveform from the high-resolution STFT spectrum.

[0046] By adopting the ConvNeXt block, Vocos 40A can achieve a prediction speed approximately 10 times faster than HiFi-GAN 30, but its quality is inferior compared to HiFi-GAN 30.

[0047] In contrast, WaveNeXt 40 shown in FIG. 2(c) generates an audio waveform from the upsampled STFT spectrum using a learnable linearization layer 48 and a reshaping layer 49. The discriminator used in Vocos 40A is utilized for the learning of the linearization layer 48. That is, the neural vocoder WaveNeXt 40 is trained by adversarial training 82 using the Vocos discriminator 80. By such a learning method, WaveNeXt 40 has both a high processing speed and high-quality audio.

[0048] Vocos 40A and WaveNeXt 40 are jointly trained. In the E2E model 4 shown in FIG. 1, the loss function L of the generator GJETS-WN and the loss function L of the discriminator (Vocos discriminator 80) DJETS-WN can be shown as follows.

[0049] L GJETS-WN =L G +w var l var +w align l align L DJETS-WN =L D Here, L G and L D respectively represent the loss functions in the generator and discriminator of Vocos40A. l var and l align respectively represent the variance loss and alignment loss in the variance adapter 24 of the E2E model 4. w var and w align are the weight coefficients for l var and l align respectively.

[0050] [B. Voice Waveform Generation System According to this Embodiment] FIG. 4 is a schematic diagram showing an example of the network configuration of the voice waveform generation system according to this embodiment. For convenience of explanation, two types of E2E models 1 and 2 are described in parallel in FIG. 4. However, similar to FIG. 1, it is not necessary to simultaneously implement two types of voice waveform generation systems.

[0051] Both the E2E model 1 and the E2E model 2 can be used for both text-to-speech (TTS) and voice conversion (VC). When performing the text-to-speech task, it is combined with a language feature quantity generation unit 5 that generates language feature quantities from the input text. When performing the voice conversion task, it is combined with an acoustic feature quantity generation unit 6 that generates acoustic feature quantities from the mel spectrogram of the source voice.

[0052] The E2E model 1 is obtained by changing the acoustic model 20 to the acoustic model 10 compared to the E2E model 3 shown in FIG. 1. More specifically, the acoustic model 10 includes a ConvNeXt-based encoder 12, a variance adapter 14, and a ConvNeXt-based decoder 16.

[0053] Each of the encoder 12 and the decoder 16 is not based on FFT like the Transformer-based encoder 22 and decoder 26 shown in FIG. 1. Each of the encoder 12 and the decoder 16 includes a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, a Gaussian error linear unit, and an addition unit, similar to the ConvNeXt block 45 shown in FIG. 3. The depthwise convolutional layer corresponds to the sum of loads in the self-attention mechanism of FFT.

[0054] In the E2E model 1, similar to Vocos40A and WaveNeXt40, a ConvNeXt block combined with a 1D convolutional layer is adopted.

[0055] Also, in the E2E model 1, learning of alignment using the discriminator used in HiFi-GAN is adopted. Note that the variance adapter 14 predicts not only the duration of each phoneme but also energy and pitch (fundamental frequency f0).

[0056] By adopting the ConvNeXt-based encoder 12 and decoder 16, the E2E model 1 can achieve faster speech waveform generation and improve the synthesized quality.

[0057] Next, the E2E model 2 is obtained by changing the acoustic model 20 to the acoustic model 10 as compared with the E2E model 4 shown in FIG. 1. More specifically, the acoustic model 10 includes a ConvNeXt-based encoder 12, a variance adapter 14, and a ConvNeXt-based decoder 16.

[0058] Each of the encoder 12 and the decoder 16 is not based on FFT like the Transformer-based encoder 22 and decoder 26 shown in FIG. 1. Similar to the ConvNeXt block 45 shown in FIG. 3, it includes a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, a Gaussian error linear unit, and an addition unit. The depthwise convolutional layer corresponds to the sum of loads in the self-attention mechanism of FFT.

[0059] For the training of WaveNeXt40 of the E2E model 2, the discriminator used in Vocos40A is also utilized. That is, in the E2E model 2 as well, Vocos40A and WaveNeXt40 are jointly trained. A loss function similar to the above-described loss function may be used.

[0060] By adopting the ConvNeXt-based encoder 12 and decoder 16 in the E2E model 2, it is possible to achieve even faster speech waveform generation and improve the synthesized quality.

[0061] [C. Performance Evaluation] Next, an example of the performance evaluation of the speech waveform generation system according to the present embodiment will be described.

[0062] Regarding the E2E model 1 and E2E model 2 shown in FIG. 4, with a sampling frequency of 24 kHz, the performance was evaluated for each condition of text-to-speech (TTS) and voice conversion (VC).

[0063] (1) Dataset As the dataset, data of Japanese speakers (male and female) of the Hi-Fi-CAPTAIN corpus publicly available from the National Institute of Information and Communications Technology (NICT) was used.

[0064] Regarding text-to-speech, a female model was trained using 18,655 parallel sentences and 203 non-parallel sentences. Also, a male model was trained using 18,655 parallel sentences and 203 non-parallel sentences.

[0065] Regarding voice conversion, a conversion model from male to female and a conversion model from female to male were each trained using 18,655 parallel sentence pairs. 100 sentences were used as validation data and test data respectively.

[0066] For the input acoustic feature amounts for the monotonic alignment search and voice quality conversion, a mel spectrogram band-limited to 7,600 Hz was used. The STFT length and the shift length were set to 1,024 samples and 256 samples, respectively.

[0067] (2) Model setting All models were learned and implemented by modifying the JETS-based E2E TTS model implemented in ESPnet2-TTS. The Harvest algorithm was introduced for f0 analysis. For the Japanese E2E TTS model, a G2P function based on pyopenjtalk enhanced with prosody symbols was used. For the model configuration of JETS combined with HiFi-GAN V1, the default settings were adopted except that the sampling frequency was changed from 22.050 kHz to 24 kHz.

[0068] As a condition for E2E-S2S-VC, the reduction coefficient r e was set to 3. For the encoder, decoder, WaveNeXt, and ConvNeXt blocks in E2E model 1 and E2E model 3 according to this embodiment, those implemented in the official implementation of Vocos were adopted. Regarding WaveNeXt, the same ConvNeXt block as that implemented in the official implementation of Vocos was used.

[0069] In E2E model 1 and E2E model 3 according to this embodiment, the input channels, the dimension of the intermediate layer, and the number of ConvNeXt blocks were set to 256 channels, 1,024 channels, and 4, respectively. Also, a probabilistic depth with a weight of 0.2 was introduced only in the ConvNeXt-based encoder and decoder. The weight coefficients w var and w align were both set to 1.0.

[0070] (3) Evaluation criteria As objective evaluations, the mel spectrogram distortion (MCD), the root mean square error (RMSE) of f0, and the character error rate (CER) by a speech recognition model were measured.

[0071] As the subjective evaluation of text-to-speech synthesis, the mean opinion score (MOS) was used. Each subject evaluated on a 5-point scale (14 sentences × 5 conditions × 2 (female + male) × 2 models = 280 samples). As the subjective evaluation of voice conversion, under the same conditions as the subjective evaluation of text-to-speech synthesis, each subject evaluated the speaker similarity of the pair of the target voice sample and the converted voice sample on a 4-point scale (14 sentences × 4 conditions × 2 models = 112 pairs). All the subjects were Japanese speakers, and they listened to the voices through headphones for evaluation.

[0072] (4) Evaluation results The evaluation results of the objective evaluation are shown in the following table.

[0073] [Table 1]

[0074] In the above table, JETS corresponds to the combination of the E2E model 3 and the language feature generation unit 5 in FIG. 1, JETS-WN corresponds to the combination of the E2E model 3 and the language feature generation unit 5 in FIG. 1, CN-JETS corresponds to the combination of the E2E model 1 and the language feature generation unit 5 in FIG. 4, and ConvNeXt-TTS corresponds to the combination of the E2E model 2 and the language feature generation unit 5 in FIG. 4.

[0075] Also, JETS-VC corresponds to the combination of the E2E model 3 and the acoustic feature generation unit 6 in FIG. 1, JETS-WN-VC corresponds to the combination of the E2E model 3 and the acoustic feature generation unit 6 in FIG. 1, CN-JETS-VC corresponds to the combination of the E2E model 1 and the acoustic feature generation unit 6 in FIG. 4, and ConvNeXt-VC corresponds to the combination of the E2E model 2 and the acoustic feature generation unit 6 in FIG. 4.

[0076] Figure 5 is a graph showing an example of the evaluation results of the subjective evaluation of the speech waveform generation system according to the present embodiment. In Figure 5, the number of subjects is 23. Also, the E2E-S2S-TTS condition shows the results in text-to-speech (TTS), and the E2E-S2S-VC condition shows the results in voice conversion (VC). Org indicates the original voice.

[0077] The RTFs of the FastSpeech2-based acoustic model, the ConvNeXt-based acoustic model, the HiFi-GAN-based neural vocoder, and the WaveNeXt-based neural vocoder were 0.03, 0.01, 0.80, and 0.04, respectively.

[0078] It can be seen that the ConvNeXt-based encoders 12 and decoders 16 (RTF) of the E2E model 1 and E2E model 2 according to the present embodiment can achieve about three times faster speed compared to the FFT-based encoder and decoder. In particular, for the E2E model 2, the RTF is 0.05 using a single-core CPU for both text-to-speech (TTS) and voice conversion (VC) (the RTFs of ConvNeXt-TTS and ConvNeXt-VC in the table).

[0079] Also, the E2E model 2 has the lowest character error rate (CER) by the speech recognition model and the lowest mel-spectrogram distortion (MCD). This indicates that the quality and speaker similarity of the synthesized speech are significantly improved compared to the E2E model 4 according to the related art. Furthermore, for the E2E model 1, compared to the E2E model 3 according to the related art, the quality and speaker similarity of the synthesized speech are significantly improved except for the speaker similarity of voice conversion from male to female.

[0080] As described above, it can be seen that the ConvNeXt-based encoder and decoder according to the present embodiment can make the speech waveform generation faster and improve the quality.

[0081] [D. Hardware Configuration Example] Next, a hardware configuration example for realizing the speech waveform generation system according to the present embodiment will be described. The speech waveform generation system according to the present embodiment may be realized using the same computing resources or different computing resources. The computing resources are provided, for example, using a general-purpose computer.

[0082] FIG. 6 is a schematic diagram showing a hardware configuration example for realizing the speech waveform generation system according to the present embodiment.

[0083] Referring to FIG. 6, the information processing apparatus 300 includes, as main hardware components, a CPU (central processing unit) 302, a GPU (graphics processing unit) 304, a main memory 306, an input device 308, a network interface (I / F: interface) 310, a storage 312, an input interface 322, an output interface 324, and an optical drive 326. These components are connected to each other via an internal bus 330.

[0084] The CPU 302 and / or the GPU 304 are processors that execute the processing necessary for realizing the system. A plurality of CPU 302 and GPU 304 may be arranged, or they may have a plurality of cores. Note that since the E2E model according to the present embodiment can perform inference at high speed, the GPU 304 may not be present.

[0085] The main memory 306 is a storage area that temporarily holds (or caches) program codes, work data, etc. when the processor (CPU 302 and / or GPU 304) executes processing, and is composed of, for example, a volatile memory such as DRAM (dynamic random access memory) or SRAM (static random access memory).

[0086] The input device 308 is a device that receives instructions and operations from the user, and is composed of, for example, a keyboard, a mouse, a touch panel, a pen, etc.

[0087] The network interface 310 exchanges data with any information processing device on the Internet or on the intranet. As the network interface 310, for example, any communication method such as Ethernet (registered trademark), wireless LAN (local area network), Bluetooth (registered trademark) can be adopted.

[0088] The input interface 322 receives the audio signal from the microphone 332.

[0089] The output interface 324 outputs the audio signal to the speaker 334.

[0090] The optical drive 326 reads the information stored on the optical disk 328 such as CD-ROM (compact disc read only memory), DVD (digital versatile disc), and outputs it to other components via the internal bus 330. The optical disk 328 is an example of a non-transitory recording medium and is distributed in a state where any program is stored non-volatilely. By the optical drive 326 reading the program from the optical disk 328 and installing it in the storage 312 etc., the computer can function as the information processing device 300.

[0091] The storage 312 stores the programs and data necessary for the realization of the system. The storage 312 is composed of, for example, a non-volatile storage device such as a hard disk, an SSD (solid state drive).

[0092] More specifically, in addition to an OS (operating system) not shown in the figure, the storage 312 stores a feature quantity generation program 314, an acoustic model 316, a neural vocoder 318, a learning program 320, and a corpus 350. Note that the feature quantity generation program 314, the acoustic model 316, and the neural vocoder 318 may be integrated programs. The voice waveform generation program according to the present invention includes at least a part of the feature quantity generation program 314, the acoustic model 316, and the neural vocoder 318.

[0093] The feature quantity generation program 314 includes computer-readable instructions for generating feature quantities (linguistic feature quantities or acoustic feature quantities). The feature quantity generation program 314 may generate feature quantities according to required tasks (text-to-speech synthesis and / or voice quality conversion). The acoustic model 316 includes computer-readable instructions for realizing the acoustic model of the E2E model. The acoustic model 316 may include a definition of a network structure, a set of learned parameters, and hyperparameters. The neural vocoder 318 includes computer-readable instructions for realizing the neural vocoder (HiFi-GAN and / or WaveNeXt) of the E2E model. The neural vocoder 318 may include a definition of a network structure, a set of learned parameters, and hyperparameters. The learning program 320 includes computer-readable instructions for learning the acoustic model and / or the neural vocoder. The learning program 320 may include an algorithm for performing learning and a set of learned parameters of the Vocos discriminator.

[0094] The corpus 350 may include a learning dataset including target mel spectrograms.

[0095] When the processor (CPU 302 and / or GPU 304) executes a program, some of the libraries and functional modules required may be replaced by libraries or functional modules provided as standard by the OS. In this case, the program alone does not include all of the program modules necessary to implement the corresponding functions, but by being installed in the execution environment of the OS, the intended processing can be realized. Furthermore, general-purpose libraries or functional modules that are permitted to be used under a predetermined license may be used. Even a program that does not include such some libraries or functional modules may be included in the technical scope of the present invention.

[0096] Also, these programs may be distributed not only by being stored in and distributed on any of the recording media as described above, but also by being downloaded from a server or the like via the Internet or an intranet.

[0097] FIG. 6 shows a configuration example using a single computer, but is not limited thereto, and a plurality of computers connected via a computer network may explicitly or implicitly cooperate to execute the processing necessary to realize the system.

[0098] All or part of the functions realized by the processor (CPU 302 and / or GPU 304) executing a program may be realized using a hard-wired circuit such as an integrated circuit. For example, it may be realized using an ASIC (application specific integrated circuit) or an FPGA (field-programmable gate array).

[0099] A person skilled in the art will be able to realize the information processing apparatus 300 according to this embodiment by appropriately using the technology corresponding to the era in which the present invention is implemented.

[0100] [E. Processing Procedure] Next, an example of the processing procedure of the system according to the present embodiment will be described.

[0101] (e1: Text-to-Speech Synthesis) FIG. 7 is a flowchart showing an example of the text-to-speech synthesis process by the speech waveform generation system according to the present embodiment. Each step shown in FIG. 7 may be realized by the processor of the information processing apparatus 300 executing the feature amount generation program 314, the acoustic model 316, and the neural vocoder 318.

[0102] Referring to FIG. 7, the information processing apparatus 300 analyzes the input text 50 and outputs a phoneme sequence 54 (text analysis unit 52) (step S100). Subsequently, the information processing apparatus 300 sequentially arranges the phoneme sequence 54 in the vector space to generate a language feature amount (embedding layer 56) (step S102).

[0103] The information processing apparatus 300 uses the ConvNeXt-based acoustic model 10 to predict hidden feature amounts from the language feature amounts. More specifically, the information processing apparatus 300 converts the generated language feature amounts into a continuous representation (encoder 12) (step S104). The information processing apparatus 300 predicts the duration, energy, and pitch (fundamental frequency f0) of each phoneme from the continuous representation (variance adapter 14) (step S106). The decoder 26 predicts hidden feature amounts from the output of the variance adapter 24 (decoder 16) (step S108).

[0104] The information processing apparatus 300 uses a waveform generation model to predict a speech waveform from the hidden feature amounts. More specifically, the information processing apparatus 300 predicts a speech waveform from the hidden feature amounts (HiFi-GAN 30 or WaveNeXt 40) (step S110). Then, the processing below step S100 is repeated.

[0105] (e2: Sound Quality Conversion) FIG. 8 is a flowchart showing an example of a voice quality conversion process by the voice waveform generation system according to the present embodiment. Each step shown in FIG. 7 may be realized by the processor of the information processing apparatus 300 executing the feature quantity generation program 314, the acoustic model 316, and the neural vocoder 318.

[0106] Referring to FIG. 8, the information processing apparatus 300 adjusts the waveform indicated by the input mel spectrogram 60 (reshaping layer 62) (step S200). Subsequently, the information processing apparatus 300 linearizes the adjusted waveform to generate acoustic feature quantities (linearization layer 64) (step S202).

[0107] The information processing apparatus 300 predicts hidden feature quantities from the acoustic feature quantities using the ConvNeXt-based acoustic model 10. More specifically, the information processing apparatus 300 converts the generated acoustic feature quantities into a continuous representation (encoder 12) (step S204). The information processing apparatus 300 predicts the duration, energy, and pitch (fundamental frequency f0) of each phoneme from the continuous representation (variance adapter 14) (step S206). The decoder 26 predicts hidden feature quantities from the output of the variance adapter 24 (decoder 16) (step S208).

[0108] The information processing apparatus 300 predicts a voice waveform from the hidden feature quantities using a waveform generation model. More specifically, the information processing apparatus 300 predicts a voice waveform from the hidden feature quantities (HiFi-GAN 30 or WaveNeXt 40) (step S210). Then, the processes below step S200 are repeated.

[0109] [F. Modification Example] In the above description, a configuration is exemplified in which one of the text 50 and the mel spectrogram 60 is input to the E2E model, but both the text 50 and the mel spectrogram 60 may be input to the E2E model. Also, a plurality of mel spectrograms 60 may be input to the E2E model simultaneously, or in addition to the plurality of mel spectrograms 60, one or more texts 50 may be input to the E2E model simultaneously. Thus, the information input to the E2E model is arbitrarily selected according to the task. As a result, at least one of the language feature amount and the acoustic feature amount is input to the acoustic model.

[0110] [G. Advantages] According to the present embodiment, a cascaded conversion type End-to-end model that can speed up voice waveform generation and further improve quality can be realized.

[0111] The embodiments disclosed this time should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is shown not by the description of the above embodiments but by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims are included.

Explanation of Reference Numerals

[0112] 1, 2, 3, 4 E2E model, 5 language feature generation unit, 6 acoustic feature generation unit, 8 target mel spectrogram, 10, 20, 316 acoustic model, 12, 22 encoder, 14, 24 variance adapter, 16, 26 decoder, 30 HiFi-GAN, 32 audio waveform, 33, 38, 43 1D convolutional layer, 34, 35, 36, 37 transposed convolutional layer, 40 WaveNeXt, 40A Vocos, 44 layer normalization block, 45 ConvNeXt block, 46, 48, 64 linearization layer, 47 iSTFT layer, 49, 62 reshaping layer, 50 text, 52 text analysis unit, 54 phoneme sequence, 56 embedding layer, 60 mel spectrogram, 80 Vocos discriminator, 90 monotonic alignment search, 300 information processing device, 302 CPU, 304 GPU, 306 main memory, 308 input device, 310 network interface, 312 storage, 314 feature generation program, 318 neural vocoder, 320 learning program, 322 input interface, 324 output interface, 326 optical drive, 328 optical disk, 330 internal bus, 332 microphone, 334 speaker, 350 corpus, 451 depthwise unit convolutional layer, 452 layer normalization block, 453, 455 pointwise unit convolutional layer, 454 Gaussian error linear unit, 456 addition unit.

Claims

1. An acoustic model that predicts a second feature quantity from a first input feature quantity, and A waveform generation model that predicts an audio waveform from the second feature quantity, comprising: The acoustic model includes an encoder for converting the first feature quantity into a continuous representation, a variance adapter for predicting the duration of each phoneme from the continuous representation, and a decoder for predicting the second feature quantity from the output of the variance adapter. An audio waveform generation system, wherein each of the encoder and the decoder includes a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit.

2. The audio waveform generation system according to claim 1, wherein the first feature quantity includes at least one of a language feature quantity generated from text and an acoustic feature quantity generated from a mel spectrogram.

3. The waveform generation model includes a one-dimensional convolutional layer, a layer normalization block, a ConvNeXt block, a linearization layer, and a reshaping layer. The audio waveform generation system according to claim 1 or 2, wherein each of the ConvNeXt blocks has the same network structure as the encoder and the decoder.

4. The waveform generation model according to claim 1 or 2, comprising a first one-dimensional convolutional layer, a second one-dimensional convolutional layer, and a plurality of transposed convolutional layers arranged between the first one-dimensional convolutional layer and the second one-dimensional convolutional layer.

5. Predicting a second feature quantity from a first input feature quantity using an acoustic model; and Predicting an audio waveform from the second feature quantity using a waveform generation model, wherein the step of predicting the second feature quantity includes: Converting the first feature quantity into a continuous representation by an encoder including a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit; Predicting the duration of each phoneme from the continuous representation; and Predicting the second feature quantity from the predicted duration by a decoder including a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit. An audio waveform generation method.

6. An audio waveform generation program, which causes a computer to: Predict a second feature quantity from a first input feature quantity using an acoustic model; and Execute a step of predicting an audio waveform from the second feature amount using a waveform generation model. The acoustic model includes an encoder for converting the first feature amount into a continuous representation, a variance adapter for predicting the duration of each phoneme from the continuous representation, and a decoder for predicting the second feature amount from the output of the variance adapter. Each of the encoder and the decoder includes a depthwise convolutional layer, a layer normalization block, a pointwise convolutional layer, and a Gaussian error linear unit, which is an audio waveform generation program.

Citation Information

Cited By

  • Real-time voice changing method, electronic equipment and computer program product

    CN121583269A