Voice waveform generation system, voice waveform generation method, and voice waveform prediction program

The voice waveform generation system addresses the limitations of neural vocoders by using a sound source signal generation and finite impulse response prediction mechanism, achieving comparable processing speed and fundamental frequency controllability to signal processing type vocoders while maintaining neural vocoder quality.

JP2025097122APending Publication Date: 2025-06-30NAT INST OF INFORMATION & COMM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023213233
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-30

AI Technical Summary

Technical Problem

Neural vocoders face challenges in achieving sufficient fundamental frequency controllability and processing speed compared to conventional signal processing type vocoders, while maintaining the quality of neural vocoders.

Method used

The proposed voice waveform generation system includes a sound source signal generation unit that produces components synchronized with the fundamental frequency, a finite impulse response prediction unit that predicts finite impulse responses based on acoustic feature amounts, and a superposition unit that applies these responses to the sound source signal, enhancing controllability and processing speed.

Benefits of technology

This approach allows for a voice waveform generation technique that matches the processing speed of signal processing type vocoders and offers controllability of the fundamental frequency, while maintaining the audio quality of neural vocoders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025097122000001_ABST
    Figure 2025097122000001_ABST
Patent Text Reader

Abstract

To provide a voice waveform generation technique having a processing speed and controllability of a fundamental frequency, which are almost similar to a signal processing vocoder, while maintaining voice quality of a neural vocoder.SOLUTION: A voice waveform generation system 1 includes a voice waveform prediction part 100 for outputting a prediction voice waveform based on an inputted acoustic feature amount. The voice waveform prediction part includes: a sound source signal generation part 102 for generating a sound source signal including a component synchronous with a fundamental frequency; finite impulse response prediction parts 110 and 112 for predicting at least one finite impulse response based on at least a part of the acoustic feature amount; and finite impulse response superimposition parts 114 and 116 for superimposing at least one finite impulse response on the sound source signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio waveform generation system, an audio waveform generation method, and an audio waveform prediction program that output a predicted audio waveform based on input acoustic feature amounts, as well as a learning system for learning an audio waveform prediction model, a learning method for an audio waveform prediction model, and a learning program for learning an audio waveform prediction model.

Background Art

[0002] In recent years, the performance of voice technologies using deep learning has improved dramatically. Along with this, the spread of voice technologies using deep learning has also rapidly expanded. In particular, in text-to-speech synthesis (TTS) technology, by adopting a neural vocoder, the quality of the output voice has reached almost the same level as human speech.

[0003] On the other hand, neural vocoders also have the problem that their processing speed is slower compared to conventional text-to-speech synthesis products (for example, the signal processing type vocoder shown in Non-Patent Document 1). In addition, neural vocoders have limitations in the functions that conventional text-to-speech synthesis products have, particularly in the controllability of the fundamental frequency (F0) representing pitch, which has become a practical problem.

[0004] For example, it has been proposed to control the fundamental frequency by introducing a fundamental frequency-dependent dilated convolutional network (Non-Patent Document 2).

[0005] In addition, it has also been proposed to control the fundamental frequency by a method based on a source-filter model that models a vocal mechanism assuming that a human sound source and vocal tract information are independent (Non-Patent Document 3).

[0006] Furthermore, a method for speeding up processing has been proposed by combining the above two technologies (Non-Patent Document 4).

Prior Art Documents

Non-Patent Documents

[0007] [Non-Patent Document 1] Morise, M., Yokomori, F., and Ozawa, K., "WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications", IEICE Transactions on Information and Systems, vol. 99, no. 7, pp. 1877-1884, 2016. doi:10.1587 / transinf.2015EDP7457. [Non-Patent Document 2] Y.-C. Wu, T. Hayashi, P. L. Tobing, K. Kobayashi and T. Toda, "Quasi-Periodic WaveNet: An Autoregressive Raw Waveform Generative Model With Pitch-Dependent Dilated Convolution Neural Network", in IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1134-1148, 2021, doi: 10.1109 / TASLP.2021.3061245. [Non-Patent Document 3] X. Wang, S. Takaki and J. Yamagishi, "Neural Source-Filter Waveform Models for Statistical Parametric Speech Synthesis", in IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 402-415, 2020, doi: 10.1109 / TASLP.2019.2956145. [Non-Patent Document 4] Yoneyama, Reo et al. "Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder", ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2022): 1-5. [Non-Patent Document 5] S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I.-S. Kweon, and S. Xie, "ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders", Proc. CVPR, pp. 16133-16142, 2023. [Summary of the Invention] [Problems to be Solved by the Invention]

[0008] In neural vocoders, fundamental frequency control is achieved by various devices as described above. However, compared with conventional signal processing type vocoders, sufficient fundamental frequency controllability cannot be obtained. For example, regarding the control range of the fundamental frequency, while there is no limitation in the control range of the signal processing type vocoder, the control range of the neural vocoder is as narrow as about 0.5 to 2 times. Also, regarding the processing speed, even for a high-speed neural vocoder, it is slower than 1 / 5 times that of the signal processing type vocoder. As a reason for not obtaining sufficient processing speed, in existing neural vocoders, it is considered that a deep neural network is directly applied to the source signal that is the basis of the speech waveform. Also, in existing neural vocoders, since the network structure depends on the fundamental frequency, the controllable range also depends on the fundamental frequency.

[0009] An object of the present invention is to provide a voice waveform generation technique that has a processing speed comparable to that of a signal processing type vocoder and controllability of a fundamental frequency while maintaining the voice quality of a neural vocoder.

Means for Solving the Problems

[0010] A voice waveform generation system according to an embodiment includes a voice waveform prediction unit that outputs a predicted voice waveform based on input acoustic feature amounts. The voice waveform prediction unit includes a sound source signal generation unit that generates a sound source signal including components synchronized with the fundamental frequency, a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic feature amounts, and a finite impulse response superposition unit that superimposes at least one finite impulse response on the sound source signal.

[0011] The finite impulse response prediction unit may convolve at least one finite impulse response with the sound source signal.

[0012] The sound source signal may include a noise component independent of the fundamental frequency. The acoustic feature amounts may include an aperiodicity index. The sound source signal may be generated based on the aperiodicity index.

[0013] At least one finite impulse response may include a finite impulse response representing resonance characteristics.

[0014] At least one finite impulse response may further include a finite impulse response representing residual characteristics.

[0015] The acoustic feature amounts may include a mel cepstrum and an aperiodicity index. The finite impulse response representing resonance characteristics may be based on the mel cepstrum. The finite impulse response representing residual characteristics may be based on the mel cepstrum and the aperiodicity index.

[0016] According to another embodiment, a voice waveform generation method is provided that outputs a predicted voice waveform based on input acoustic feature amounts. The voice waveform generation method includes a step of generating a sound source signal including components synchronized with a fundamental frequency, a step of predicting at least one finite impulse response based on at least a part of the acoustic feature amounts, and a step of superimposing at least one finite impulse response on the sound source signal.

[0017] According to yet another embodiment, a voice waveform prediction program is provided that outputs a predicted voice waveform based on input acoustic feature amounts. The voice waveform prediction program causes a computer to perform a step of generating a sound source signal including components synchronized with a fundamental frequency, a step of predicting at least one finite impulse response based on at least a part of the acoustic feature amounts, and a step of superimposing at least one finite impulse response on the sound source signal.

[0018] According to yet another embodiment, a learning system is provided for learning a voice waveform prediction model that outputs a predicted voice waveform based on input acoustic feature amounts. The voice waveform prediction model includes a sound source signal generation unit that generates a sound source signal including components synchronized with a fundamental frequency, a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic feature amounts, and a finite impulse response superimposing unit that superimposes at least one finite impulse response on the sound source signal. The learning system includes a first error calculation unit that calculates an error between a voice waveform and a predicted voice waveform output when an acoustic feature amount extracted from a corpus in which the voice waveform and the acoustic feature amount are associated is input to the voice waveform prediction model, and an update unit that updates first model parameters defining the behavior of the finite impulse response prediction unit and the finite impulse response superimposing unit based on the error calculated by the first error calculation unit.

[0019] The learning system may further include a discriminative feature quantity prediction unit that predicts discriminative feature quantities from each of the predicted speech waveform and the speech waveform, and a second error calculation unit that calculates an error between the discriminative feature quantity predicted from the predicted speech waveform and the discriminative feature quantity predicted from the speech waveform. The update unit may further update a second model parameter that defines the behavior of the discriminative feature quantity prediction unit based on the error calculated by the second error calculation unit.

[0020] The update unit may execute the update of the first model parameter and the update of the second model parameter in a predetermined order.

[0021] According to yet another embodiment, a method for learning an audio waveform prediction model that outputs a predicted audio waveform based on input acoustic feature quantities is provided. The audio waveform prediction model includes a sound source signal generation unit that generates a sound source signal including components synchronized with a fundamental frequency, a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic feature quantities, and a finite impulse response superposition unit that superimposes at least one finite impulse response on the sound source signal. The learning method includes calculating an error between the predicted audio waveform output when the acoustic feature quantities extracted from a corpus in which an audio waveform and acoustic feature quantities are associated are input to the audio waveform prediction model and the audio waveform associated with the extracted acoustic feature quantities, and updating a first model parameter that defines the behavior of the finite impulse response prediction unit and the finite impulse response superposition unit based on the calculated error.

[0022] According to yet another embodiment, a learning program is provided for learning an audio waveform prediction model that outputs a predicted audio waveform based on input acoustic feature quantities. The audio waveform prediction model includes a sound source signal generation unit that generates a sound source signal including components synchronized with a fundamental frequency, a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic feature quantities, and a finite impulse response superposition unit that superimposes at least one finite impulse response on the sound source signal. The learning program causes a computer to calculate an error between a predicted audio waveform output when an acoustic feature quantity extracted from a corpus in which an audio waveform and an acoustic feature quantity are associated is input to the audio waveform prediction model and the audio waveform associated with the extracted acoustic feature quantity, and update first model parameters that define the behavior of the finite impulse response prediction unit and the finite impulse response superposition unit based on the calculated error.

Advantages of the Invention

[0023] According to the present invention, it is possible to provide an audio waveform generation technique having a processing speed and fundamental frequency controllability comparable to those of a signal processing type vocoder while maintaining the audio quality of a neural vocoder.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

[0025] Embodiments of the present invention will be described in detail with reference to the drawings. The same or corresponding parts in the drawings are denoted by the same reference numerals, and their descriptions will not be repeated.

[0026] The audio waveform generation system according to the present embodiment is applicable to an audio synthesis task. The audio synthesis task includes, for example, at least text-to-speech synthesis (TTS) and voice conversion (VC).

[0027] [A. Configuration Example of Audio Waveform Generation System] First, a configuration example of the speech waveform generation system according to the present embodiment will be described.

[0028] (a1: Configuration example) FIG. 1 is a block diagram showing a configuration example of a speech waveform generation system 1 according to the present embodiment. Referring to FIG. 1, the speech waveform generation system 1 includes a speech waveform prediction unit 100 configured as a neural vocoder and an acoustic feature amount determination unit 140 that provides an acoustic feature amount 152.

[0029] The speech waveform prediction unit 100 outputs a predicted speech waveform 160 based on the input acoustic feature amount 152. The speech waveform prediction unit 100 includes a speech waveform prediction model for outputting the predicted speech waveform 160.

[0030] The acoustic feature amount 152 includes a fundamental frequency 1521 (F0) corresponding to the pitch of the sound, an aperiodicity index 1522, and a mel cepstrum 1523. The mel cepstrum 1523 may be the mel cepstrum itself, or may be a mel spectrum, mel frequency cepstrum coefficients (MFCC), or a linear prediction code. The acoustic feature amount 152 may include a voiced / unvoiced flag 1524.

[0031] The acoustic feature amount determination unit 140 outputs the acoustic feature amount 152. For example, the acoustic feature amount determination unit 140 determines the acoustic feature amount 152 corresponding to the speech waveform to be output by referring to the corpus 150 according to the input information 2. In the configuration example shown in FIG. 1, it is applicable to a speech conversion task or the like. In the corpus 150, for example, a speech waveform and an acoustic feature amount indicating the speech waveform are associated with each other.

[0032] The speech waveform prediction unit 100 includes a sound source signal generation unit 102, at least one finite impulse response prediction unit, and a finite impulse response superposition unit.

[0033] At least one finite impulse response prediction unit predicts at least one finite impulse response based on at least a part of the acoustic feature quantity 152. The configuration example shown in FIG. 1 includes two finite impulse response prediction units.

[0034] The finite impulse response superposition unit superimposes at least one finite impulse response on the sound source signal 120. The configuration example shown in FIG. 1 includes two finite impulse response superposition units.

[0035] More specifically, the speech waveform prediction unit 100 includes a first latent feature quantity prediction unit 104, a second latent feature quantity prediction unit 106, a third latent feature quantity prediction unit 108, a first finite impulse response prediction unit 110, a second finite impulse response prediction unit 112, a first finite impulse response superposition unit 114, and a second finite impulse response superposition unit 116.

[0036] The sound source signal generation unit 102 generates a sound source signal 120 including a component synchronized with the fundamental frequency.

[0037] The first latent feature quantity prediction unit 104 is a learned model and predicts the first latent feature quantity 122 from the acoustic feature quantity 152 (aperiodicity index 1522).

[0038] The second latent feature quantity prediction unit 106 is a learned model and predicts the second latent feature quantity 124 from another acoustic feature quantity 152 (Mel cepstrum 1523).

[0039] The third latent feature quantity prediction unit 108 is a learned model and predicts the third latent feature quantity 126 from at least one of the first latent feature quantity 122 and the second latent feature quantity 124.

[0040] The first finite impulse response prediction unit 110 is a learned model that predicts the first finite impulse response 130 from the third latent feature quantity 126. That is, the first finite impulse response prediction unit 110 predicts the first finite impulse response 130 based on at least a part of the acoustic feature quantity 152 (aperiodicity index 1522 and mel cepstrum 1523). The predicted first finite impulse response 130 represents residual characteristics. That is, the first finite impulse response 130 representing residual characteristics is based on the aperiodicity index 1522 and the mel cepstrum 1523.

[0041] The second finite impulse response prediction unit 112 is a learned model that predicts the second finite impulse response 132 from the second latent feature quantity 124. That is, the second finite impulse response prediction unit 112 predicts the second finite impulse response 132 based on at least a part of the acoustic feature quantity 152 (mel cepstrum 1523). The predicted second finite impulse response 132 represents resonance characteristics. That is, the second finite impulse response 132 representing resonance characteristics is based on the mel cepstrum 1523.

[0042] The first finite impulse response superposition unit 114 includes a learned model and superimposes the first finite impulse response 130 on the sound source signal 120. More specifically, the first finite impulse response superposition unit 114 convolves the first finite impulse response 130 with the sound source signal 120 to generate a convolution residual signal.

[0043] The second finite impulse response superposition unit 116 includes a learned model and superimposes the second finite impulse response 132 on the convolution residual signal generated by the first finite impulse response superposition unit 114. More specifically, the second finite impulse response superposition unit 116 convolves the second finite impulse response 132 with the convolution residual signal to generate a predicted speech waveform 160. The generated predicted speech waveform 160 becomes the speech waveform output from the speech waveform prediction unit 100.

[0044] Note that the order in which the finite impulse response is convolved with the sound source signal 120 may be either between the first finite impulse response superposition unit 114 and the second finite impulse response superposition unit 116.

[0045] As shown in FIG. 1, the speech waveform prediction unit 100 generates a sound source signal 120 including components synchronized with the fundamental frequency, and then convolves the generated sound source signal 120 with a first finite impulse response 130 representing residual characteristics and a second finite impulse response 132 representing resonance characteristics, respectively, to generate a speech waveform. In this way, by adopting a method of applying convolution instead of directly applying a deep neural network to the sound source signal 120 that forms the basis of the speech waveform, the generation process of the speech waveform can be accelerated. Also, since the sound source signal 120 and the finite impulse response are clearly separated, the fundamental frequency (F0) can be controlled without limitation. That is, the fundamental frequency 1521 given to the sound source signal generation unit 102 can be freely changed.

[0046] By adopting such a configuration, it is possible to realize a speech waveform generation technique that has a processing speed comparable to that of a signal processing type vocoder and controllability of the fundamental frequency while maintaining the speech quality of the neural vocoder.

[0047] (a2: Sound source signal generation unit) Next, a processing example of the sound source signal generation unit 102 will be described.

[0048] The sound source signal generation unit 102 generates a mixed excitation signal s(t) as the sound source signal 120 by mixing a pulse train according to the fundamental frequency 1521 (F0) and Gaussian noise based on the aperiodicity index 1522. More specifically, the mixed excitation signal s(t) can be expressed as follows.

[0049] [When voiced] s(t)=g p v (k) *p(t)+g n u (k) *n(t) [When unvoiced] s(t)=g nn(t) However, p(t) represents a pulse train following the fundamental frequency 1521, n(t) represents Gaussian noise, and g p represents the gain for p(t), and g n represents the gain for n(t). v (k) and u (k) represent the voiced section and the unvoiced section of the impulse response calculated by applying the inverse Fourier transform to the aperiodicity index 1522 in the k-th frame corresponding to the time t, respectively. Also, the operator "*" represents the convolution operation.

[0050] The gains g p , g n each may be set, for example, in the range of 0.003 to 0.1. In the sound source signal 120, the term including the pulse train p(t) is a component synchronized with the fundamental frequency, and the term including the Gaussian noise n(t) is a noise component independent of the fundamental frequency. Also, the sound source signal 120 is based on the aperiodicity index 1522.

[0051] (a3: Latent Feature Quantity Prediction Unit) Next, an example of the model structure of the latent feature quantity prediction unit (the first latent feature quantity prediction unit 104, the second latent feature quantity prediction unit 106, and the third latent feature quantity prediction unit 108) will be described. The latent feature quantity prediction unit may adopt any trained model (neural network) as long as it can predict the latent feature quantity included in the input information. As the latent feature quantity prediction unit, for example, Causal Convolution can be used. Causal Convolution is an example of a convolutional neural network and is characterized in that only past values are used for prediction.

[0052] FIG. 2 is a block diagram showing an example of the model structure of the latent feature quantity prediction unit shown in FIG. 1. Referring to FIG. 2, each of the latent feature quantity prediction units (the first latent feature quantity prediction unit 104, the second latent feature quantity prediction unit 106, and the third latent feature quantity prediction unit 108) includes one or more Causal ConvNeXt core blocks 1040. The Causal ConvNeXt core block 1040 corresponds to the core module of the ConvNeXt V2 architecture (see Non-Patent Document 5, etc.). FIG. 2 shows a configuration example in which the number of layers of the Causal ConvNeXt core block 1040 is 2, but the number of layers can be arbitrarily set.

[0053] Each of the Causal ConvNeXt core blocks 1040 includes a depthwise convolutional layer 1041 (Causal depthwise 1D-conv) in the input depth direction, a layer normalization 1042, a pointwise convolutional layer 1043 (Pointwise 1D-conv + GELU) with activation of a Gaussian error linear unit (GELU), a global response normalization 1044, and a pointwise convolution 1045 for each output point.

[0054] A causal relationship is applied to the depthwise convolutional layer 1041 in the input depth direction. The kernel size of the depthwise convolutional layer 1041 in the input depth direction may be set to 5, for example.

[0055] For example, the first latent feature quantity prediction unit 104 may predict a 128-channel first latent feature quantity 122 for the input of the aperiodicity index 1522. The second latent feature quantity prediction unit 106 may predict a 256-channel second latent feature quantity 124 for the input of the mel cepstrum 1523. The third latent feature quantity prediction unit 108 may predict a 128-channel third latent feature quantity 126 for the input of the first latent feature quantity 122 and the second latent feature quantity 124.

[0056] (a4: Finite Impulse Response Superposition Section) Next, an example of the model structure of the finite impulse response superposition section (the first finite impulse response superposition section 114 and the second finite impulse response superposition section 116) will be described.

[0057] FIG. 3 is a block diagram showing an example of the model structure of the finite impulse response superposition section shown in FIG. 1. Referring to FIG. 3, each of the finite impulse response superposition sections (the first finite impulse response superposition section 114 and the second finite impulse response superposition section 116) uses a plurality of finite impulse response filters 1142 (FIR (Finite Impulse Response) filters) to convert the input signal x (0) (t) (sound source signal 120 or convolution residual signal) into the output signal x (M) (t).

[0058] Let the impulse response coefficients of each FIR filter be h (1,k) , h (2,k) , … h (M,k) . Then, the intermediate output signals x (1) (t), x (2) (t), …, x (M-1) (t) of each layer can be shown as follows.

[0059] x (m) (t) = h(m, k) * x (m-1) (t) + x (m-1) (t) The impulse response coefficients h (1,k) , h (2,k) , … h (M,k) of the m-th FIR filter are predicted from the latent representation of the auxiliary feature (the first finite impulse response 130 or the second finite impulse response 132) and the immediately preceding impulse response coefficient using an extended causal convolution layer.

[0060] More specifically, each of the finite impulse response superposition sections (the first finite impulse response superposition section 114 and the second finite impulse response superposition section 116) includes a plurality of layers 1140.

[0061] Layer 1140 includes an adder 1141, a finite impulse response filter 1142, a ConvNeXt core block 1143, a linear layer 1144 including a fully-connected (FC) and a Gaussian error linear unit (GELU), a dilated causal convolutional layer 1145 (Dilated Causal Conv1d), a fully-connected layer 1146, and a concatenation layer 1147 (Concatenate).

[0062] For example, the number of channels of the latent representation (the first finite impulse response 130 or the second finite impulse response 132) may be set to 128.

[0063] The kernel size of the dilated causal convolutional layer 1145 may be set to 3, and the kernel sizes of the other layers may be set to 1. Each of the first finite impulse response overlap section 114 and the second finite impulse response overlap section 116 may include 8 layers of finite impulse response filters 1142 with a 256-tap size. The number of dilation sizes of the corresponding dilated causal convolutional layer 1145 for each of the 8 layers may be 1, 2, 4, 8, 1, 2, 4, 8, respectively.

[0064] In the model structure shown in FIG. 3, the impulse response coefficients h (1,k) , h (2,k) , …h (M,k) of the FIR filter are predicted using the aperiodicity index 1522 and the mel cepstrum 1523, but the fundamental frequency 1521 (F0) and / or the sound source signal 120 may be used.

[0065] (a5: First modification example) FIG. 4 is a block diagram showing a configuration example of an audio waveform generation system 1A according to the first modification example of the present embodiment. Referring to FIG. 4, the audio waveform generation system 1A includes an audio waveform prediction unit 100A and an acoustic feature amount determination unit 140.

[0066] The voice waveform prediction unit 100A is different from the voice waveform prediction unit 100 shown in FIG. 1 in that it does not include the first latent feature quantity prediction unit 104. That is, since the first latent feature quantity 122 is not predicted, the third latent feature quantity prediction unit 108 predicts the third latent feature quantity 126 from the second latent feature quantity 124.

[0067] Note that since the sound source signal 120 is generated depending on the aperiodicity index 1522, the voice quality can be maintained even if the first latent feature quantity 122 is not used for the prediction of the third latent feature quantity 126 and the first finite impulse response 130.

[0068] By adopting the configuration shown in FIG. 4, the processing amount of the first latent feature quantity prediction unit 104 is reduced, so that the processing speed can be increased.

[0069] (a6: Second modification example) FIG. 5 is a block diagram showing a configuration example of a voice waveform generation system 1B according to a second modification example of the present embodiment. The voice waveform generation system 1B shown in FIG. 5 can be used, for example, in an end-to-end text-to-speech synthesis system. Referring to FIG. 5, the voice waveform generation system 1B includes an acoustic feature quantity prediction unit 142 and a voice waveform prediction unit 100B.

[0070] The acoustic feature quantity prediction unit 142 predicts the acoustic feature quantity 152 from the input latent feature quantity 4. The latent feature quantity 4 is predicted from the input text or the like using a neural network or the like (not shown).

[0071] The voice waveform prediction unit 100B is different from the voice waveform prediction unit 100A shown in FIG. 4 in that it includes a fourth latent feature quantity prediction unit 128 instead of the third latent feature quantity prediction unit 108. The fourth latent feature quantity prediction unit 128 predicts the third latent feature quantity 126 from the aperiodicity index 1522 and the mel cepstrum 1523.

[0072] That is, in the voice waveform prediction unit 100B, the first finite impulse response 130 and the second finite impulse response 132 are predicted independently of each other.

[0073] By adopting the configuration shown in FIG. 5, a voice waveform can be predicted from the input latent feature amount 4.

[0074] (a7: Third modification example) FIG. 6 is a block diagram showing a configuration example of a voice waveform generation system 1C according to a third modification example of the present embodiment. Referring to FIG. 6, the voice waveform generation system 1C includes a voice waveform prediction unit 100C and an acoustic feature amount determination unit 140.

[0075] The voice waveform prediction unit 100C is different from the voice waveform prediction unit 100B shown in FIG. 5 in that the second latent feature amount prediction unit 106 is not included. Also, the second finite impulse response prediction unit 112 is changed to a second finite impulse response prediction unit 112C.

[0076] The second finite impulse response prediction unit 112C predicts the second finite impulse response 132 from the mel cepstrum 1523. That is, in the voice waveform prediction unit 100C, the second finite impulse response 132 (resonance characteristics) is directly predicted mathematically from the mel cepstrum 1523.

[0077] By adopting the configuration shown in FIG. 6, the processing amount of the second latent feature amount prediction unit 106 is reduced, so the processing speed can be increased.

[0078] (a8: Fourth modification example) FIG. 7 is a block diagram showing a configuration example of a voice waveform generation system 1D according to a fourth modification example of the present embodiment. Referring to FIG. 7, the voice waveform generation system 1D includes a voice waveform prediction unit 100D and an acoustic feature amount determination unit 140.

[0079] The voice waveform prediction unit 100D is different from the voice waveform prediction unit 100B shown in FIG. 5 in that it does not include a configuration (a fourth latent feature quantity prediction unit 128, a first finite impulse response prediction unit 110, and a first finite impulse response superposition unit 114) for predicting the first finite impulse response 130 (residual characteristics).

[0080] By adopting the configuration shown in FIG. 7, the processing amount for predicting the first finite impulse response 130 (residual characteristics) is reduced, so that the processing speed can be increased.

[0081] (a9: The fifth modification example) FIG. 8 is a block diagram showing a configuration example of a voice waveform generation system 1E according to a fifth modification example of the present embodiment. Referring to FIG. 8, the voice waveform generation system 1E includes a voice waveform prediction unit 100E and an acoustic feature quantity determination unit 140.

[0082] The voice waveform prediction unit 100E includes a sound source signal generation unit 102E, a latent feature quantity prediction unit 134, a first finite impulse response prediction unit 110E, a second finite impulse response prediction unit 112E, a first finite impulse response superposition unit 114, a second finite impulse response superposition unit 116, and a signal addition unit 144.

[0083] The latent feature quantity prediction unit 134 is a learned model that predicts a latent feature quantity 136 from an acoustic feature quantity 152 (a mel cepstrum 1523 and an aperiodicity index 1522).

[0084] The first finite impulse response prediction unit 110E is a learned model that predicts a first finite impulse response 130E from the latent feature quantity 136. The first finite impulse response 130E represents the characteristics of an aperiodic signal.

[0085] The second finite impulse response prediction unit 112E is a learned model that predicts a second finite impulse response 132E from the latent feature quantity 136. The second finite impulse response 132E represents the characteristics of a periodic signal.

[0086] The sound source signal generation unit 102E generates an F0 synchronization pulse signal 120P and a noise signal 120N that depend on the fundamental frequency. The excitation signal s P (t) indicating the F0 synchronization pulse signal 120P and the excitation signal s N (t) indicating the noise signal 120N can be expressed as follows, respectively.

[0087] [When there is voice] s P (t)=g p p(t) s N (t)=0 [When there is no voice] s P (t)=0 s N (t)=g n n(t) The first finite impulse response superimposing unit 114 superimposes the first finite impulse response 130E on the F0 synchronization pulse signal 120P. The second finite impulse response superimposing unit 116 superimposes the second finite impulse response 132 on the noise signal 120N.

[0088] The signal addition unit 144 adds the signal generated by the first finite impulse response superimposing unit 114 and the signal generated by the second finite impulse response superimposing unit 116, and outputs the addition result as the predicted speech waveform 160.

[0089] By adopting the configuration shown in FIG. 8, the first finite impulse response 130E showing non-periodic signal characteristics and the second finite impulse response 132E showing periodic signal characteristics can be independently superimposed, so the speech quality can be improved.

[0090] (a10: Other modifications) The present invention is not limited to the configuration example shown in FIG. 1 and the modification examples shown in FIGS. 4 to 8 described above, and any modified configuration may be adopted as long as the finite impulse response is convolved with the sound source signal that is the basis of the speech waveform.

[0091] The acoustic feature quantity 152 provided to the voice waveform prediction unit (neural vocoder) may be determined with reference to the corpus 150 (Figs. 1, 4, 6 to 8), or may be predicted from the latent feature quantity 4 (Fig. 5). That is, in the configuration examples shown in Figs. 1, 4 to 9, the configuration for providing the acoustic feature quantity 152 may be appropriately changed.

[0092] Regarding the sound source signal, a mixed excitation signal (Figs. 1, 4, 6, 7) obtained by mixing a pulse train (F0 synchronous pulse signal) according to the fundamental frequency 1521 (F0) and Gaussian noise (noise signal) based on the aperiodicity index 1522 may be adopted, or a configuration may be adopted such that either the pulse train or the Gaussian noise can be selected, or a configuration (Fig. 8) in which the pulse train and the Gaussian noise are separated may be adopted. That is, the sound source signal that forms the basis of the voice waveform may be any of (1) a mixed excitation signal obtained by mixing an F0 synchronous pulse signal and a noise signal, (2) only the F0 synchronous pulse signal, (3) the noise signal, and (4) the separated F0 synchronous pulse signal and noise signal.

[0093] [B. Learning of Voice Waveform Prediction Unit] Next, the learning of the voice waveform prediction unit included in the voice waveform generation system will be described.

[0094] Any algorithm may be adopted for the learning of the voice waveform prediction unit (neural vocoder). For example, adversarial training using a voice waveform discrimination model can be used. By using adversarial training, a voice waveform prediction unit capable of generating high-quality voice can be generated.

[0095] Fig. 9 is a block diagram showing a configuration example of a learning system 10 for a voice waveform prediction unit according to the present embodiment. The learning system 10 learns (trains) the voice waveform prediction unit 100 (voice waveform prediction model). Referring to Fig. 9, as an example, the learning system 10 adopts a configuration according to a generative adversarial network (GAN).

[0096] More specifically, the learning system 10 includes a voice waveform prediction unit 100 corresponding to a generator, and a first discriminative feature quantity prediction unit 210 and a second discriminative feature quantity prediction unit 220 corresponding to a discriminator.

[0097] With reference to the corpus 150, a voice waveform 156 and an acoustic feature quantity 152 indicating the voice waveform 156 are extracted.

[0098] The voice waveform prediction unit 100 predicts a predicted voice waveform 160 from the acoustic feature quantity 152 from the corpus 150. The behavior of the voice waveform prediction unit 100 (trained model) is defined by the model parameters stored in the voice waveform generation model 170.

[0099] The generation error calculation unit 230 calculates the error between the predicted voice waveform 160 predicted by the voice waveform prediction unit 100 and the voice waveform 156 from the corpus 150. That is, the generation error calculation unit 230 corresponds to a first error calculation unit that calculates the error between the predicted voice waveform 160 output when the acoustic feature quantity 152 extracted from the corpus 150 is input to the voice waveform prediction unit 100 (voice waveform prediction model) and the voice waveform 156 associated with the extracted acoustic feature quantity 152.

[0100] The first discriminative feature quantity prediction unit 210 predicts a first discriminative feature quantity from the predicted voice waveform 160 predicted by the voice waveform prediction unit 100. The second discriminative feature quantity prediction unit 220 predicts a second discriminative feature quantity from the voice waveform 156 from the corpus 150. The behavior of each of the first discriminative feature quantity prediction unit 210 and the second discriminative feature quantity prediction unit 220 is defined by the model parameters stored in the voice waveform discrimination model 270. The first discriminative feature quantity prediction unit 210 and the second discriminative feature quantity prediction unit 220 correspond to discriminative feature quantity prediction units that predict discriminative feature quantities from each of the predicted voice waveform 160 and the voice waveform 156.

[0101] The identification error calculation unit 240 calculates the error between the first identification feature amount predicted by the first identification feature amount prediction unit 210 and the second identification feature amount predicted by the second identification feature amount prediction unit 220. That is, the identification error calculation unit 240 corresponds to a second error calculation unit that calculates the error between the identification feature amount predicted from the predicted speech waveform 160 and the identification feature amount predicted from the speech waveform 156.

[0102] Based on the error calculated by the generation error calculation unit 230 and the error calculated by the identification error calculation unit 240, the model update unit 250 updates (optimizes) the values of the model parameters stored in the speech waveform generation model 170 and the values of the model parameters stored in the speech waveform identification model 270.

[0103] That is, based on the error calculated by the generation error calculation unit 230 (the first error calculation unit), the model update unit 250 updates the first model parameters that define the behavior of the speech waveform prediction unit 100 (finite impulse response prediction unit (the first finite impulse response prediction unit 110 and the second finite impulse response prediction unit 112), finite impulse response superposition unit (the first finite impulse response superposition unit 114 and the second finite impulse response superposition unit 116), and latent feature amount prediction unit (the first latent feature amount prediction unit 104, the second latent feature amount prediction unit 106, and the third latent feature amount prediction unit 108)).

[0104] Also, based on the error calculated by the identification error calculation unit 240 (the second error calculation unit), the model update unit 250 updates the second model parameters that define the behavior of the identification feature amount prediction unit (the first identification feature amount prediction unit 210 and the second identification feature amount prediction unit 220).

[0105] Examples of the loss used by the model update unit 250 to update the model parameters of the speech waveform prediction unit 100 (generator) include, for example, the least squares criterion L adv and the feature matching loss L fm and the mel cepstrum L1 loss L of the speech waveform mel and the source excitation regularization loss L of the residual signal regFour of them may be set as adversarial losses. The objective function L including these four losses G can be shown as follows.

[0106] L G =L adv +λ fm L fm +λ mel L mel +λ reg L reg Here, λ fm , λ mel , λ reg are adjustable hyperparameters. For example, λ fm = 2.0, λ mel = 50.0, λ reg = 20.0 may be used.

[0107] FIG. 10 is a block diagram showing a configuration example of a learning system 10A of an audio waveform prediction unit according to a modification of the present embodiment. Referring to FIG. 10, the learning system 10A of the audio waveform prediction unit includes two audio waveform prediction units 100, a fundamental frequency control unit 280, and an extended generation error calculation unit 260 as compared with the learning system 10 of the audio waveform prediction unit shown in FIG. 9.

[0108] The first audio waveform prediction unit 100-1 is given the acoustic feature amount 152 from the corpus 150.

[0109] The fundamental frequency control unit 280 controls (changes) the fundamental frequency 1521 (F0) included in the acoustic feature amount 152 from the corpus 150. The second audio waveform prediction unit 100-2 is given the acoustic feature amount 152 in which the fundamental frequency 1521 is controlled. That is, the fundamental frequency control unit 280 and the second audio waveform prediction unit 100-2 predict the audio waveform when the fundamental frequency 1521 is controlled.

[0110] The extended generation error calculation unit 260 calculates the error between the predicted speech waveform (with the fundamental frequency controlled) predicted by the second speech waveform prediction unit 100-2 and the speech waveform 156 from the corpus 150. Since it is ideal that only the fundamental frequency is different between the predicted speech waveform predicted by the second speech waveform prediction unit 100-2 and the speech waveform 156 from the corpus 150, feature amounts other than the fundamental frequency may be regarded as losses.

[0111] [C. Performance Evaluation] Next, an example of the performance evaluation of the speech waveform generation system according to the present embodiment will be described. In the following objective evaluation and subjective evaluation, 1000 sentences of male speakers' voices were used.

[0112] (c1: Objective Evaluation) An example of the objective evaluation of the performance of the speech waveform prediction unit 100 (hereinafter referred to as "FIRNet") according to the present embodiment is shown. In the following table, for the fundamental frequency, the results of evaluating the mel cepstrum distortion (MCD [dB]), the root mean square error (RMSE) of the square of log f0, the voiced / unvoiced decision error (VUVE [%]), and the real-time factor (RTF) under seven scaling conditions (1.0, 0.00, 0.25, 0.5, 2.0, 4.00, 8.00) are shown.

[0113] In the following table, "WORLD" indicates the neural vocoder proposed in Non-Patent Document 1, and "SiFi-GAN" indicates the neural vocoder proposed in Non-Patent Document 4.

[0114] [Number]

[0115] Note that RMSE and VUVE could not be calculated for 0.0×f0 and 8.0×f0.

[0116] The speech waveform prediction unit 100 according to this embodiment has a lower MCD at all fundamental frequencies compared to the two prior arts. That is, it can be seen that the speech waveform prediction unit 100 according to this embodiment generates higher-quality speech.

[0117] From the perspective of the processing speed indicated by the RTF, it can be seen that the speech waveform prediction unit 100 according to this embodiment is approximately five times faster than SiFi-GAN. It is considered that the reason is that the speech waveform prediction unit 100 according to this embodiment adopts a classical method using an FIR filter and thus does not require complex processing for time-domain signals. Also, the speech waveform prediction unit 100 according to this embodiment is faster than WORLD at high fundamental frequencies. This is because the generation speed of the speech waveform in WORLD depends on the pitch interval.

[0118] (c2: Subjective evaluation) As a subjective evaluation, a five-level Mean Opinion Score (MOS) test for evaluating the naturalness of the synthesized speech was conducted. Twenty evaluators participated in the test, and the evaluation was performed under seven types of scaling conditions, similar to the above objective evaluation. In addition to the speech waveforms (Original) included in the corpus, the speech waveforms generated by the speech waveform prediction unit 100, WORLD, and SiFi-GAN according to this embodiment were used as evaluation targets. The evaluators evaluated 12 samples for each scaling condition and neural vocoder.

[0119] FIG. 11 is a graph showing an example of the subjective evaluation results of the speech waveform prediction unit 100 according to this embodiment. Referring to FIG. 11, it can be seen that for SiFi-GAN, high speech quality is maintained when the scaling conditions are 1.0, 0.5, and 2.0, but the speech quality significantly deteriorates for other scaling conditions.

[0120] On the other hand, it can be seen that the speech waveform prediction unit 100 according to the present embodiment maintains the same speech quality as WORLD when the scaling conditions are 0.0, 0.2, 0.5, 4.0, and 8.0. Also, when the fundamental frequency is not manipulated (scaling condition is 1.0), the speech waveform prediction unit 100 according to the present embodiment has higher speech quality than WORLD, indicating that it can achieve both excellent robustness against fundamental frequency control and high speech quality.

[0121] (c3: small parenthesis) As shown by the above objective and subjective evaluations, the speech waveform prediction unit 100 according to the present embodiment maintains the quality as a neural vocoder and has fundamental frequency controllability and processing speed equivalent to those of conventional signal processing vocoders.

[0122] [D. Hardware Configuration Example] Next, a hardware configuration example for realizing the speech waveform generation system and learning system according to the present embodiment will be described. The speech waveform generation system and learning system according to the present embodiment may be realized using the same computing resources or different computing resources. The computing resources are provided, for example, using a general-purpose computer.

[0123] FIG. 12 is a schematic diagram showing a hardware configuration example for realizing the system according to the present embodiment.

[0124] Referring to FIG. 12, the information processing apparatus 300 includes, as main hardware components, a CPU (central processing unit) 302, a GPU (graphics processing unit) 304, a main memory 306, an input device 308, a network interface (I / F: interface) 310, a storage 312, an input interface 322, an output interface 324, and an optical drive 326. These components are connected to each other via an internal bus 330.

[0125] The CPU 302 and / or the GPU 304 are processors that execute the processing necessary for the realization of the system. A plurality of CPU 302s and / or GPU 304s may be arranged, or they may have a plurality of cores.

[0126] The main memory 306 is a storage area that temporarily holds (or caches) program codes, work data, etc. when the processor (CPU 302 and / or GPU 304) executes processing, and is composed of, for example, volatile memories such as DRAM (dynamic random access memory) and SRAM (static random access memory).

[0127] The input device 308 is a device that receives instructions, operations, etc. from the user, and is composed of, for example, a keyboard, a mouse, a touch panel, a pen, etc.

[0128] The network interface 310 exchanges data with any information processing device on the Internet or on the intranet. As the network interface 310, for example, any communication method such as Ethernet (registered trademark), wireless LAN (local area network), Bluetooth (registered trademark) can be adopted.

[0129] The input interface 322 receives the audio signal from the microphone 332. The output interface 324 outputs the audio signal to the speaker 334.

[0130] The optical drive 326 reads the information stored on an optical disc 328 such as a CD-ROM (compact disc read only memory) or a DVD (digital versatile disc), and outputs it to other components via the internal bus 330. The optical disc 328 is an example of a non-transitory recording medium and is distributed with any program stored non-volatilely. By having the optical drive 326 read a program from the optical disc 328 and install it in the storage 312 or the like, the computer can function as the information processing apparatus 300.

[0131] The storage 312 stores programs and data necessary for the realization of the system. The storage 312 is composed of, for example, a non-volatile storage device such as a hard disk or an SSD (solid state drive).

[0132] More specifically, the storage 312 stores, in addition to an OS (operating system) not shown, an acoustic feature quantity generation program 314, an audio waveform prediction program 316, a learning program 318, and a corpus 150.

[0133] The acoustic feature quantity generation program 314 includes computer-readable instructions for generating acoustic feature quantities. The acoustic feature quantity generation program 314 realizes the acoustic feature quantity determination unit 140 or the acoustic feature quantity prediction unit 142. The acoustic feature quantity generation program 314 may be a neural network that is a learned model.

[0134] The audio waveform prediction program 316 includes computer-readable instructions for realizing a neural vocoder that is a learned model. The acoustic feature quantity generation program 314 realizes the audio waveform prediction unit 100 (and a modified example of the audio waveform prediction unit). The acoustic feature quantity generation program 314 may include an algorithm, model parameters, and hyperparameters.

[0135] The learning program 318 includes computer-readable instructions for training the acoustic feature quantity generation program 314 (neural vocoder). The learning program 318 realizes a first discriminative feature quantity prediction unit 210, a second discriminative feature quantity prediction unit 220, a generation error calculation unit 230, a discrimination error calculation unit 240, and a model update unit 250. The learning program 318 may be a neural network that is a trained model.

[0136] Note that at least any one of the acoustic feature quantity generation program 314, the speech waveform prediction program 316, and the learning program 318 may include an algorithm, model parameters, and hyperparameters. Further, at least any one of the acoustic feature quantity generation program 314, the speech waveform prediction program 316, and the learning program 318 may be a program in an executable format built from an algorithm, model parameters, and hyperparameters.

[0137] The corpus 150 may include a dataset of speech waveforms and acoustic feature quantities indicating the speech waveforms.

[0138] When the processor (CPU 302 and / or GPU 304) executes a program (translation program), a part of the libraries and functional modules required may be replaced by libraries or functional modules provided as standard by the OS. In this case, the program alone does not include all of the program modules necessary to implement the corresponding functions, but by being installed in the execution environment of the OS, the target processing can be realized. Further, a general-purpose library or functional module that is permitted to be used under a predetermined license may be used. Even a program that does not include such some libraries or functional modules may be included in the technical scope of the present invention.

[0139] In addition, these programs are not only stored and distributed in any of the recording media as described above, but may also be distributed by being downloaded from a server or the like via the Internet or an intranet.

[0140] FIG. 12 shows a configuration example using a single computer, but is not limited thereto, and a plurality of computers connected via a computer network may cooperate explicitly or implicitly to execute processes necessary to realize the system.

[0141] All or part of the functions realized by the processor (CPU 302 and / or GPU 304) executing the program may be realized using a hard-wired circuit such as an integrated circuit. For example, it may be realized using an ASIC (application specific integrated circuit), an FPGA (field-programmable gate array), or the like.

[0142] A person skilled in the art will be able to realize the information processing apparatus 300 according to this embodiment by appropriately using technologies according to the era in which the present invention is implemented.

[0143] [E. Processing Procedure] Next, an example of the processing procedure of the system according to this embodiment will be described.

[0144] (e1: Generation Process of Audio Waveform) FIG. 13 is a flowchart showing an example of the audio waveform generation process by the audio waveform generation system according to this embodiment. The audio waveform generation process includes an audio waveform generation method that outputs a predicted audio waveform 160 based on the input acoustic feature amount 152. Each step shown in FIG. 13 may be realized by the processor of the information processing apparatus 300 executing the acoustic feature amount generation program 314 and the audio waveform prediction program 316.

[0145] Referring to FIG. 13, the information processing apparatus 300 determines an acoustic feature amount 152 according to the input information 2 (step S100).

[0146] The information processing apparatus 300 outputs a predicted speech waveform 160 based on the input acoustic feature amount 152. More specifically, the information processing apparatus 300 generates a sound source signal 120 including a component synchronized with the fundamental frequency 1521 (step S102).

[0147] The information processing apparatus 300 predicts at least one finite impulse response based on at least a part of the acoustic feature amount 152. More specifically, the information processing apparatus 300 predicts a first finite impulse response 130 from the acoustic feature amount 152 (step S104). The information processing apparatus 300 predicts a second finite impulse response 132 from the acoustic feature amount 152 (step S106).

[0148] The information processing apparatus 300 superimposes at least one finite impulse response on the sound source signal 120. More specifically, the information processing apparatus 300 superimposes the first finite impulse response 130 on the sound source signal 120 (step S108). The information processing apparatus 300 further superimposes the second finite impulse response 132 on the sound source signal 120 (step S110). The information processing apparatus 300 outputs the sound source signal 120 after the first finite impulse response 130 and the second finite impulse response 132 are superimposed as a speech waveform (step S112). Then, the processing below step S100 is repeated.

[0149] Note that steps S102 to S106 may be executed in parallel instead of serially. The execution order of steps S102 to S106 may be any. Also, the execution order of step S108 and step S110 may be any.

[0150] (e2: Learning process of the speech waveform prediction unit) FIG. 14 is a flowchart showing an example of a learning method of the speech waveform generation system according to the present embodiment. Each step shown in FIG. 14 may be realized by a processor of the information processing apparatus 300 executing a learning program 318.

[0151] Referring to FIG. 14, the information processing apparatus 300 extracts a pair of a speech waveform 156 and an acoustic feature amount 152 included in the corpus 150 (step S200), and inputs the acoustic feature amount 152 to the speech waveform prediction unit 100 to generate a predicted speech waveform 160 (step S202).

[0152] The information processing apparatus 300 predicts a first discriminant feature amount from the predicted speech waveform 160 (step S204). That is, the information processing apparatus 300 calculates an error between the predicted speech waveform 160 output when the acoustic feature amount 152 extracted from the corpus 150 in which the speech waveform 156 and the acoustic feature amount 152 are associated is input to the speech waveform prediction unit 100 (speech waveform prediction model) and the speech waveform 156 associated with the extracted acoustic feature amount 152.

[0153] Further, the information processing apparatus 300 predicts a second discriminant feature amount from the speech waveform 156 (step S206).

[0154] When the processes of steps S200 to S206 are repeated a predetermined number of times (YES in step S208), the information processing apparatus 300 updates the model parameters of the speech waveform prediction unit 100 based on the error (loss) between the predicted speech waveform 160 and the speech waveform 156 (step S210). That is, the information processing apparatus 300 updates the first model parameters that define the behavior of the speech waveform prediction unit 100 based on the error between the predicted speech waveform 160 and the speech waveform 156.

[0155] The information processing apparatus 300 updates the model parameters of the first identification feature quantity prediction unit 210 and the second identification feature quantity prediction unit 220 based on the error (loss) between the first identification feature quantity and the second identification feature quantity (step S212). That is, the information processing apparatus 300 updates the second model parameters that define the behavior of the identification feature quantity prediction unit (the first identification feature quantity prediction unit 210 and the second identification feature quantity prediction unit 220) based on the error between the first identification feature quantity and the second identification feature quantity.

[0156] Note that the processes of step S210 and step S212 may be alternately executed one of them when the condition of step S208 is satisfied. For example, the information processing apparatus 300 may execute the update of the first model parameters (step S210) and the update of the second model parameters (step S212) in a predetermined order (for example, alternately). Thus, the execution order of step S210 and step S212 may be either.

[0157] The information processing apparatus 300 determines whether or not the learning end condition is satisfied (step S214). The learning end condition may be that the number of updates of the model parameters has reached a predetermined number, or that the error (loss) has decreased to a value equal to or less than a predetermined value.

[0158] If the learning end condition is not satisfied (NO in step S214), the processes below step S200 are repeated.

[0159] If the learning end condition is satisfied (YES in step S214), the learning process ends. The voice waveform prediction unit 100 is configured using the model parameters at this time.

[0160] [F. Advantages] According to this embodiment, for the sound source signal that is the basis of the voice waveform, it is generated in the same manner as a signal processing type vocoder. After predicting the finite impulse response based on the acoustic feature amount, the predicted finite impulse response is superimposed on the sound source signal to generate the voice waveform. By adopting such a mechanism, it is possible to realize a system that has a processing speed comparable to that of a signal processing type vocoder and controllability of the fundamental frequency while maintaining the voice quality of the neural vocoder.

[0161] By using the advantages in this embodiment, it is applicable to all services that require voice synthesis, such as voice translation systems, text-to-speech synthesis, voice quality conversion, and singing voice synthesis.

[0162] The embodiments disclosed this time should be considered as illustrative in all respects and not restrictive. The scope of the present invention is shown not by the description of the above embodiments but by the claims, and it is intended that all changes within the meaning and scope equivalent to the claims are included.

Description of Reference Numerals

[0163] 1, 1A, 1B, 1C, 1D, 1E audio waveform generation system, 2 input information, 4 latent feature quantities, 10, 10A learning system, 12, 120 sound source signal, 100 audio waveform prediction unit, 100, 100A, 100B, 100C, 100D, 100E audio waveform prediction unit, 100-1 first audio waveform prediction unit, 100-2 second audio waveform prediction unit, 102, 102E sound source signal generation unit, 104 first latent feature quantity prediction unit, 106 second latent feature quantity prediction unit, 108 third latent feature quantity prediction unit, 110, 110E first finite impulse response prediction unit, 112, 112C, 112E second finite impulse response prediction unit, 114 first finite impulse response superposition unit, 116 second finite impulse response superposition unit, 120N noise signal, 120P synchronization pulse signal, 122 first latent feature quantity, 124 second latent feature quantity, 126 third latent feature quantity, 128 fourth latent feature quantity prediction unit, 130, 130E first finite impulse response, 132,132E Second finite impulse response, 134 Latent feature predictor, 136 Latent feature, 140 Acoustic feature determination unit, 142 Acoustic feature predictor, 144 Signal adder, 150 Corpus, 152 Acoustic feature, 156 Speech waveform, 160 Predicted speech waveform, 170 Speech waveform generation model, 210 First discriminative feature predictor, 220 Second discriminative feature predictor, 230 Generation error calculator, 240 Discrimination error calculator, 250 Model update unit, 260 Extended generation error calculator, 270 Speech waveform discrimination model, 280 Fundamental frequency control unit, 300 Information processing device, 302 CPU, 304 GPU, 306 Main memory, 308 Input device, 310 Network interface, 312 Storage, 314 Acoustic feature generation program, 316 Speech waveform prediction program, 318 Learning program, 322 Input interface, 324 Output interface, 326 Optical drive, 328 Optical disk, 330 Internal bus, 332 Microphone, 334 Speaker, 1040 Causal ConvNeXt core block, 1041 Convolution layer in the input depth direction, 1042 Layer normalization, 1043 Pointwise convolution layer, 1044 Global response normalization, 1140 Layer, 1141 Adder, 1142 Finite impulse response filter, 1143 ConvNeXt core block, 1144 Linear layer, 1145 Extended causal convolution layer, 1146 Fully connected layer, 1147 Concatenation layer, 1521 Fundamental frequency, 1522 Aperiodicity index, 1523 Mel cepstrum, 1524 Silence flag.,

Claims

1. A voice waveform generation system comprising a voice waveform prediction unit that outputs a predicted voice waveform based on the input acoustic feature amount, wherein the voice waveform prediction unit comprises a sound source signal generation unit that generates a sound source signal including components synchronized with the fundamental frequency, a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic feature amount, and a finite impulse response superposition unit that superimposes the at least one finite impulse response on the sound source signal.

2. The voice waveform generation system according to claim 1, wherein the finite impulse response prediction unit convolves the at least one finite impulse response with the sound source signal.

3. The voice waveform generation system according to claim 1, wherein the sound source signal includes a noise component independent of the fundamental frequency.

4. The voice waveform generation system according to claim 1, wherein the at least one finite impulse response includes a finite impulse response representing resonance characteristics.

5. A voice waveform generation method for outputting a predicted voice waveform based on an input acoustic feature amount, the method comprising: generating a sound source signal including components synchronized with the fundamental frequency; predicting at least one finite impulse response based on at least a part of the acoustic feature amount; and superimposing the at least one finite impulse response on the sound source signal.

6. A voice waveform prediction program for outputting a predicted voice waveform based on an input acoustic feature amount, the program causing a computer to generate a sound source signal including components synchronized with the fundamental frequency; predict at least one finite impulse response based on at least a part of the acoustic feature amount; and superimpose the at least one finite impulse response on the sound source signal.