Speech waveform generation system, speech waveform generation method, and speech waveform prediction program

The system addresses the processing speed and pitch controllability issues of neural vocoders by using an audio waveform prediction unit with sound source signal generation and finite impulse response prediction units, achieving performance comparable to signal processing-based vocoders while maintaining neural vocoder speech quality.

WO2025134624A1PCT designated stage expired Publication Date: 2025-06-26NAT INST OF INFORMATION & COMM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/040604
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-11-15
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Neural vocoders face challenges in processing speed and pitch controllability compared to conventional signal processing-based vocoders, while maintaining the speech quality of neural vocoders.

Method used

The proposed system includes an audio waveform prediction unit that generates a predicted audio waveform based on input acoustic feature quantities, using a sound source signal generation unit, finite impulse response prediction units, and superposition units to achieve faster processing and improved pitch control.

Benefits of technology

This approach enables a speech waveform generation technique with processing speed and pitch controllability comparable to signal processing-based vocoders, while maintaining the speech quality of neural vocoders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024040604_26062025_PF_FP_ABST
    Figure JP2024040604_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This speech waveform generation system includes a speech waveform prediction unit that outputs a predicted speech waveform on the basis of an inputted acoustic feature amount. The speech waveform prediction unit includes: a sound source signal generation unit that generates a sound source signal including a component to be synchronized with a fundamental frequency; a finite impulse response prediction unit that predicts at least one finite impulse response on the basis of at least a portion of the acoustic feature amount; and a finite impulse response superimposition unit that superimposes at least one finite impulse response on the sound source signal.
Need to check novelty before this filing date? Find Prior Art

Description

Voice waveform generation system, voice waveform generation method, and voice waveform prediction program

[0001] The present invention relates to a speech waveform generation system, a speech waveform generation method, and a speech waveform prediction program that output a predicted speech waveform based on input acoustic features, as well as a learning system for training a speech waveform prediction model, a training method for a speech waveform prediction model, and a learning program for training a speech waveform prediction model.

[0002] In recent years, the performance of speech technologies using deep learning has improved dramatically. Accordingly, the use of speech technologies using deep learning has also rapidly spread. In particular, in text-to-speech synthesis (TTS) technology, by adopting neural vocoders, the quality of the output speech has reached a level almost equivalent to that of human speech.

[0003] On the other hand, neural vocoders have the problem of slower processing speed than conventional text-to-speech synthesis products (e.g., the signal processing vocoder described in Non-Patent Document 1). Furthermore, neural vocoders have limitations on the functions of conventional text-to-speech synthesis products, particularly on the controllability of the fundamental frequency (F0) that represents pitch, which poses practical challenges.

[0004] For example, it has been proposed to control the fundamental frequency by introducing a fundamental frequency-dependent dilated convolutional network (Non-Patent Document 2).

[0005] It has also been proposed to control the fundamental frequency using a method based on a source-filter model that models the vocal mechanism, which assumes that the human sound source and vocal tract information are independent (Non-Patent Document 3).

[0006] Furthermore, a method for speeding up processing by combining the above two techniques has been proposed (Non-Patent Document 4).

[0007] Morise, M., Yokomori, F., and Ozawa, K., "WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications", IEICE Transactions on Information and Systems, vol. 99, no. 7, pp. 1877-1884, 2016. doi:10.1587 / transinf.2015EDP7457.Y. -C. Wu, T. Hayashi, P. L. Tobing, K. Kobayashi and T. Toda, "Quasi-Periodic WaveNet: An Autoregressive Raw Waveform Generative Model With Pitch-Dependent Dilated Convolution Neural Network", in IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 1134-1148, 2021, doi: 10.1109 / TASLP.2021.3061245.X. Wang, S. Takaki and J. Yamagishi, "Neural Source-Filter Waveform Models for Statistical Parametric Speech Synthesis", in IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 402-415, 2020, doi: 10.1109 / TASLP.2019.2956145.Yoneyama, Reo et al."Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder", ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2022): 1-5.S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I.-S. Kweon, and S. Xie, "ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders", Proc. CVPR, pp. 16133-16142,2023.

[0008] Neural vocoders achieve fundamental frequency control through various techniques, as described above, but compared to conventional signal-processing vocoders, they do not provide sufficient fundamental frequency control. For example, while signal-processing vocoders have no limit on the fundamental frequency control range, neural vocoders have a narrow control range of approximately 0.5 to 2 times. Furthermore, even high-speed neural vocoders are slow, less than one-fifth the speed of signal-processing vocoders. One possible reason for the lack of sufficient processing speed is that existing neural vocoders directly apply deep neural networks to the sound source signal that forms the basis of the speech waveform. Furthermore, because the network structure of existing neural vocoders depends on the fundamental frequency, the controllable range also depends on the fundamental frequency.

[0009] An object of the present invention is to provide a voice waveform generation technology that maintains the voice quality of a neural vocoder while achieving processing speed and fundamental frequency controllability comparable to that of a signal processing vocoder.

[0010] A speech waveform generation system according to an embodiment includes a speech waveform prediction unit that outputs a predicted speech waveform based on input acoustic features. The speech waveform prediction unit includes a speech source signal generation unit that generates a speech source signal including a component synchronized with a fundamental frequency, a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic features, and a finite impulse response superposition unit that superposes the at least one finite impulse response on the speech source signal.

[0011] The finite impulse response prediction unit may convolve the source signal with at least one finite impulse response.

[0012] The sound source signal may include a noise component independent of the fundamental frequency. The acoustic feature may include an aperiodicity index. The sound source signal may be generated based on the aperiodicity index.

[0013] The at least one finite impulse response may include a finite impulse response representative of a resonance characteristic.

[0014] The at least one finite impulse response may further include a finite impulse response representing a residual characteristic.

[0015] The acoustic features may include a mel-cepstrum and an aperiodicity index. The finite impulse response representing the resonance characteristics may be based on the mel-cepstrum. The finite impulse response representing the residual characteristics may be based on the mel-cepstrum and an aperiodicity index.

[0016] According to another embodiment, there is provided a speech waveform generation method for outputting a predicted speech waveform based on input acoustic features, the speech waveform generation method including the steps of generating a sound source signal including a component synchronized with a fundamental frequency, predicting at least one finite impulse response based at least in part on the acoustic features, and convolving the at least one finite impulse response on the sound source signal.

[0017] According to yet another embodiment, there is provided a speech waveform prediction program that outputs a predicted speech waveform based on input acoustic features. The speech waveform prediction program causes a computer to execute the steps of generating a sound source signal including a component synchronized with a fundamental frequency, predicting at least one finite impulse response based on at least a part of the acoustic features, and superimposing the at least one finite impulse response on the sound source signal.

[0018] According to yet another embodiment, there is provided a training system for training a speech waveform prediction model that outputs a predicted speech waveform based on input acoustic features. The speech waveform prediction model includes: a sound source signal generation unit that generates a sound source signal including a component synchronized with a fundamental frequency; a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a portion of the acoustic features; and a finite impulse response superposition unit that superposes the at least one finite impulse response on the sound source signal. The training system includes: a corpus in which speech waveforms and acoustic features are associated; a first error calculation unit that calculates an error between a predicted speech waveform output when acoustic features extracted from the corpus are input to the speech waveform prediction model and the speech waveform associated with the extracted acoustic features; and an update unit that updates first model parameters that define the behavior of the finite impulse response prediction unit and the finite impulse response superposition unit based on the error calculated by the first error calculation unit.

[0019] The learning system may further include a discriminant feature prediction unit that predicts a discriminant feature from each of the predicted speech waveform and the speech waveform, and a second error calculation unit that calculates an error between the discriminant feature predicted from the predicted speech waveform and the discriminant feature predicted from the speech waveform. The update unit may further update second model parameters that define the behavior of the discriminant feature prediction unit based on the error calculated by the second error calculation unit.

[0020] The update unit may update the first model parameters and the second model parameters in a predetermined order.

[0021] According to yet another embodiment, there is provided a training method for a speech waveform prediction model that outputs a predicted speech waveform based on input acoustic features. The speech waveform prediction model includes: a sound source signal generation unit that generates a sound source signal including a component synchronized with a fundamental frequency; a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic features; and a finite impulse response superposition unit that superposes the at least one finite impulse response on the sound source signal. The training method includes the steps of: calculating an error between a predicted speech waveform output when acoustic features extracted from a corpus in which speech waveforms and acoustic features are associated with each other are input to the speech waveform prediction model; and updating first model parameters that define the behavior of the finite impulse response prediction unit and the finite impulse response superposition unit based on the calculated error.

[0022] According to yet another embodiment, there is provided a training program for training a speech waveform prediction model that outputs a predicted speech waveform based on input acoustic features. The speech waveform prediction model includes a sound source signal generation unit that generates a sound source signal including a component synchronized with a fundamental frequency, a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a part of the acoustic features, and a finite impulse response superposition unit that superposes the at least one finite impulse response on the sound source signal. The training program causes a computer to execute the steps of: calculating an error between a predicted speech waveform output when acoustic features extracted from a corpus in which speech waveforms and acoustic features are associated with each other are input to the speech waveform prediction model; and updating first model parameters that define the behavior of the finite impulse response prediction unit and the finite impulse response superposition unit based on the calculated error.

[0023] According to the present invention, it is possible to provide a voice waveform generation technology that maintains the voice quality of a neural vocoder while achieving processing speed and fundamental frequency controllability comparable to that of a signal processing vocoder.

[0024] FIG. 1 is a block diagram showing an example of the configuration of a speech waveform generation system according to the present embodiment. FIG. 2 is a block diagram showing an example of the model structure of a latent feature prediction unit shown in FIG. 1. FIG. 3 is a block diagram showing an example of the model structure of the finite impulse response superposition unit shown in FIG. 1. FIG. 4 is a block diagram showing an example of the configuration of a speech waveform generation system according to a first modified example of the present embodiment. FIG. 5 is a block diagram showing an example of the configuration of a speech waveform generation system according to a second modified example of the present embodiment. FIG. 6 is a block diagram showing an example of the configuration of a speech waveform generation system according to a third modified example of the present embodiment. FIG. 7 is a block diagram showing an example of the configuration of a speech waveform generation system according to a fourth modified example of the present embodiment. FIG. 8 is a block diagram showing an example of the configuration of a speech waveform generation system according to a fifth modified example of the present embodiment. FIG. 9 is a block diagram showing an example of the configuration of a training system for a speech waveform prediction unit according to the present embodiment. FIG. 10 is a block diagram showing an example of the configuration of a training system for a speech waveform prediction unit according to the modified example of the present embodiment. FIG. 11 is a graph showing an example of subjective evaluation results of a speech waveform prediction unit according to the present embodiment. FIG. 12 is a schematic diagram showing an example of the hardware configuration for realizing a system according to the present embodiment. FIG. 13 is a flowchart showing an example of speech waveform generation processing by the speech waveform generation system according to the present embodiment. FIG. 14 is a flowchart showing an example of a training method for the speech waveform generation system according to the present embodiment.

[0025] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described in detail with reference to the accompanying drawings, in which the same or corresponding parts are designated by the same reference numerals and will not be described repeatedly.

[0026] The speech waveform generation system according to this embodiment is applicable to speech synthesis tasks, which include at least text-to-speech synthesis (TTS) and voice conversion (VC), for example.

[0027] [A. Example of Configuration of Voice Waveform Generation System] First, an example of the configuration of a voice waveform generation system according to the present embodiment will be described.

[0028] (a1: Configuration Example) Fig. 1 is a block diagram showing a configuration example of a speech waveform generation system 1 according to the present embodiment. Referring to Fig. 1, speech waveform generation system 1 includes a speech waveform prediction unit 100 configured as a neural vocoder, and an acoustic feature determination unit 140 that provides acoustic features 152.

[0029] The speech waveform prediction unit 100 outputs a predicted speech waveform 160 based on the input acoustic feature 152. The speech waveform prediction unit 100 includes a speech waveform prediction model for outputting the predicted speech waveform 160.

[0030] The acoustic feature 152 includes a fundamental frequency 1521 (F0) corresponding to the pitch of a sound, an aperiodicity index 1522, and a mel-cepstrum 1523. The mel-cepstrum 1523 may be the mel-cepstrum itself, a mel-spectrum, a mel-frequency cepstrum coefficient (MFCC), or a linear predictive code. The acoustic feature 152 may also include a voiced / unvoiced flag 1524.

[0031] The acoustic feature determination unit 140 outputs acoustic features 152. For example, the acoustic feature determination unit 140 determines acoustic features 152 corresponding to a speech waveform to be output by referring to a corpus 150 in accordance with input information 2. The configuration example shown in Fig. 1 is applicable to a speech conversion task, etc. In the corpus 150, for example, a speech waveform is associated with an acoustic feature indicating the speech waveform.

[0032] The speech waveform prediction unit 100 includes a sound source signal generation unit 102, at least one finite impulse response prediction unit, and a finite impulse response superposition unit.

[0033] At least one finite impulse response predictor predicts at least one finite impulse response based on at least a part of the acoustic feature 152. The exemplary configuration shown in Fig. 1 includes two finite impulse response predictors.

[0034] The finite impulse response superimposing unit superimposes at least one finite impulse response on the sound source signal 120. The configuration example shown in Fig. 1 includes two finite impulse response superimposing units.

[0035] More specifically, the speech waveform prediction unit 100 includes a first latent feature prediction unit 104, a second latent feature prediction unit 106, a third latent feature prediction unit 108, a first finite impulse response prediction unit 110, a second finite impulse response prediction unit 112, a first finite impulse response superimposition unit 114, and a second finite impulse response superimposition unit 116.

[0036] The sound source signal generating unit 102 generates a sound source signal 120 that includes a component synchronized with the fundamental frequency.

[0037] The first latent feature prediction unit 104 is a trained model, and predicts the first latent feature 122 from the acoustic feature 152 (non-periodicity index 1522).

[0038] The second latent feature prediction unit 106 is a trained model, and predicts the second latent feature 124 from another acoustic feature 152 (mel-cepstrum 1523).

[0039] The third latent feature prediction unit 108 is a trained model, and predicts the third latent feature 126 from at least one of the first latent feature 122 and the second latent feature 124 .

[0040] The first finite impulse response prediction unit 110 is a trained model and predicts the first finite impulse response 130 from the third latent feature 126. That is, the first finite impulse response prediction unit 110 predicts the first finite impulse response 130 based on at least a part of the acoustic feature 152 (the aperiodicity index 1522 and the mel-cepstrum 1523). The predicted first finite impulse response 130 represents residual characteristics. That is, the first finite impulse response 130 representing the residual characteristics is based on the aperiodicity index 1522 and the mel-cepstrum 1523.

[0041] The second finite impulse response prediction unit 112 is a trained model, and predicts the second finite impulse response 132 from the second latent feature 124. That is, the second finite impulse response prediction unit 112 predicts the second finite impulse response 132 based on at least a part (mel-cepstrum 1523) of the acoustic feature 152. The predicted second finite impulse response 132 represents a resonance characteristic. That is, the second finite impulse response 132 representing the resonance characteristic is based on the mel-cepstrum 1523.

[0042] The first finite impulse response superimposing unit 114 includes a trained model and superimposes the first finite impulse response 130 on the sound source signal 120. More specifically, the first finite impulse response superimposing unit 114 convolves the first finite impulse response 130 on the sound source signal 120 to generate a convolved residual signal.

[0043] The second finite impulse response superimposing unit 116 includes a trained model and superimposes a second finite impulse response 132 on the convolved residual signal generated by the first finite impulse response superimposing unit 114. More specifically, the second finite impulse response superimposing unit 116 convolves the second finite impulse response 132 on the convolved residual signal to generate a predicted speech waveform 160. The generated predicted speech waveform 160 becomes the speech waveform output from the speech waveform prediction unit 100.

[0044] Note that the order in which the finite impulse responses are convolved with the sound source signal 120 between the first finite impulse response superimposing section 114 and the second finite impulse response superimposing section 116 may be any order.

[0045] As shown in FIG. 1 , the speech waveform prediction unit 100 generates a speech waveform by first generating a sound source signal 120 including a component synchronized with the fundamental frequency, and then convolving the generated sound source signal 120 with a first finite impulse response 130 representing residual characteristics and a second finite impulse response 132 representing resonance characteristics. In this way, by adopting a convolution method rather than a method of directly applying a deep neural network to the sound source signal 120, which is the basis of the speech waveform, the speech waveform generation process can be accelerated. Furthermore, because the sound source signal 120 and the finite impulse responses are clearly separated, the fundamental frequency (F0) can be controlled without any restrictions. In other words, the fundamental frequency 1521 provided to the sound source signal generation unit 102 can be freely changed.

[0046] By adopting such a configuration, it is possible to realize a voice waveform generation technology that maintains the voice quality of a neural vocoder while achieving processing speed and fundamental frequency controllability comparable to that of a signal processing vocoder.

[0047] (a2: Sound Source Signal Generator) Next, an example of processing by the sound source signal generator 102 will be described.

[0048] The excitation signal generation unit 102 generates a mixed excitation signal s(t) as the excitation signal 120 by mixing a pulse train according to a fundamental frequency 1521 (F0) with Gaussian noise based on the aperiodicity index 1522. More specifically, the mixed excitation signal s(t) can be expressed as follows:

[0049] [Voice case] s(t) = g p v (k) *p(t)+g n u (k) *n(t) [in the silent case] s(t) = g n n(t) where p(t) represents a pulse train with a fundamental frequency of 1521, n(t) represents Gaussian noise, and g p denotes the gain for p(t), and g n denotes the gain for n(t). (k) and u (k)indicate the speech interval and silence interval of the impulse response calculated by applying an inverse Fourier transform to the aperiodicity index 1522 in the k-th frame corresponding to time t. The operator "*" indicates a convolution operation.

[0050] Gain g p , g n may be set in the range of 0.003 to 0.1, for example. In the sound source signal 120, the term including the pulse train p(t) is a component synchronized with the fundamental frequency, and the term including the Gaussian noise n(t) is a noise component independent of the fundamental frequency. The sound source signal 120 is also based on the aperiodicity index 1522.

[0051] (a3: Latent Feature Prediction Unit) Next, an example of the model structure of the latent feature prediction unit (first latent feature prediction unit 104, second latent feature prediction unit 106, and third latent feature prediction unit 108) will be described. The latent feature prediction unit may employ any trained model (neural network) as long as it can predict latent features included in input information. As the latent feature prediction unit, for example, causal convolution can be used. Causal convolution is an example of a convolutional neural network, and is characterized in that only past values ​​are used for prediction.

[0052] FIG. 2 is a block diagram showing an example of a model structure of the latent feature prediction unit shown in FIG. 1. Referring to FIG. 2, each of the latent feature prediction units (the first latent feature prediction unit 104, the second latent feature prediction unit 106, and the third latent feature prediction unit 108) includes one or more Causal ConvNeXt core blocks 1040. The Causal ConvNeXt core block 1040 corresponds to a core module of the ConvNeXt V2 architecture (see Non-Patent Document 5, etc.). While FIG. 2 shows an example of a configuration in which the Causal ConvNeXt core block 1040 has two layers, the number of layers can be set arbitrarily.

[0053] Each of the Causal ConvNext core blocks 1040 includes an input depthwise convolution layer 1041 (Causal depthwise 1D-conv), a layer normalization 1042 (Layer normalization), a pointwise convolution layer 1043 (Pointwise 1D-conv + GELU) with Gaussian Error Linear Unit (GELU) activation, a global response normalization 1044 (Global response normalization), and an output pointwise convolution 1045 (Pointwise 1D-conv).

[0054] Causality is applied to the input depthwise convolutional layer 1041. The kernel size of the input depthwise convolutional layer 1041 may be set to, for example, 5.

[0055] For example, the first latent feature prediction unit 104 may predict a 128-channel first latent feature 122 in response to an input of the aperiodicity index 1522. The second latent feature prediction unit 106 may predict a 256-channel second latent feature 124 in response to an input of the mel-cepstrum 1523. The third latent feature prediction unit 108 may predict a 128-channel third latent feature 126 in response to an input of the first latent feature 122 and the second latent feature 124.

[0056] (a4: Finite Impulse Response Superimposing Section) Next, an example of a model structure of the finite impulse response superimposing section (first finite impulse response superimposing section 114 and second finite impulse response superimposing section 116) will be described.

[0057] Fig. 3 is a block diagram showing an example of a model structure of the finite impulse response superposition unit shown in Fig. 1. Referring to Fig. 3, each of the finite impulse response superposition units (first finite impulse response superposition unit 114 and second finite impulse response superposition unit 116) superposes an input signal x using a plurality of finite impulse response filters 1142 (FIR (Finite Impulse Response) filters). (0) (t) (the source signal 120 or the convolution residual signal) into the output signal x (M)Convert to (t).

[0058] The impulse response coefficients of each FIR filter are h (1,k) , h (2,k) , ...h (M,k) Then, the intermediate output signal x (1) (t), x (2) (t), ..., x (M-1) (t) can be shown as follows:

[0059] x (m) (t)=h(m,k)*x (m-1) (t) + x (m-1) (t) Impulse response coefficient h of the mth FIR filter (1,k) , h (2,k) , ...h (M,k) is predicted from the latent representation of the auxiliary feature (first finite impulse response 130 or second finite impulse response 132) and the previous impulse response coefficients using an extended causal convolutional layer.

[0060] More specifically, each of the finite impulse response superposition units (first finite impulse response superposition unit 114 and second finite impulse response superposition unit 116 ) includes a plurality of layers 1140 .

[0061] The layer 1140 includes an adder 1141, a finite impulse response filter 1142, a ConvNext core block 1143, a linear layer 1144 including a fully connected (FC) and a Gaussian error linear unit (GELU), a dilated causal convolution layer 1145 (Dilated Causal Convld), a fully connected layer 1146, and a concatenation layer 1147 (Concatenate).

[0062] For example, the number of channels of the latent representation (first finite impulse response 130 or second finite impulse response 132) may be set to 128.

[0063] The kernel size of the extended causal convolutional layer 1145 may be set to 3, and the kernel sizes of the other layers may be set to 1. Each of the first finite impulse response superposition unit 114 and the second finite impulse response superposition unit 116 may include eight layers of 256-tap finite impulse response filters 1142. The numbers of the extension sizes of the eight corresponding extended causal convolutional layers 1145 may be 1, 2, 4, 8, 1, 2, 4, and 8, respectively.

[0064] In the model structure shown in FIG. 3, the impulse response coefficient h of the FIR filter is calculated using the aperiodicity index 1522 and the mel-cepstrum 1523. (1,k) , h (2,k) , ...h (M,k) is predicted, however, the fundamental frequency 1521 (F0) and / or the source signal 120 may be used.

[0065] (a5: First Modification) Fig. 4 is a block diagram showing an example of the configuration of a speech waveform generation system 1A according to a first modification of the present embodiment. Referring to Fig. 4, speech waveform generation system 1A includes a speech waveform prediction unit 100A and an acoustic feature determination unit 140.

[0066] 1 in that it does not include the first latent feature prediction unit 104. That is, since the first latent feature 122 is not predicted, the third latent feature prediction unit 108 predicts the third latent feature 126 from the second latent feature 124.

[0067] Since the sound source signal 120 is generated depending on the aperiodicity index 1522, the speech quality can be maintained even if the first latent feature 122 is not used to predict the third latent feature 126 and the first finite impulse response 130.

[0068] By employing the configuration shown in FIG. 4, the amount of processing by the first latent feature prediction unit 104 is reduced, and therefore the processing speed can be increased.

[0069] (a6: Second Modification) Fig. 5 is a block diagram showing an example configuration of a speech waveform generation system 1B according to a second modification of the present embodiment. The speech waveform generation system 1B shown in Fig. 5 can be used, for example, in an end-to-end text-to-speech synthesis system. Referring to Fig. 5, the speech waveform generation system 1B includes an acoustic feature prediction unit 142 and a speech waveform prediction unit 100B.

[0070] The acoustic feature prediction unit 142 predicts the acoustic feature 152 from the input latent feature 4. The latent feature 4 is predicted from the input text or the like using a neural network or the like (not shown).

[0071] 4 in that it includes a fourth latent feature prediction unit 128 instead of the third latent feature prediction unit 108. The fourth latent feature prediction unit 128 predicts the third latent feature 126 from the aperiodicity index 1522 and the mel-cepstrum 1523.

[0072] That is, in the speech waveform prediction unit 100B, the first finite impulse response 130 and the second finite impulse response 132 are predicted independently of each other.

[0073] By employing the configuration shown in FIG. 5, it is possible to predict a speech waveform from the input latent feature 4.

[0074] (a7: Third Modification) Fig. 6 is a block diagram showing an example of the configuration of a speech waveform generation system 1C according to a third modification of the present embodiment. Referring to Fig. 6, the speech waveform generation system 1C includes a speech waveform prediction unit 100C and an acoustic feature determination unit 140.

[0075] 5 in that it does not include the second latent feature prediction unit 106. In addition, the second finite impulse response prediction unit 112 has been changed to a second finite impulse response prediction unit 112C.

[0076] The second finite impulse response prediction unit 112C predicts the second finite impulse response 132 from the mel-cepstrum 1523. That is, in the speech waveform prediction unit 100C, the second finite impulse response 132 (resonance characteristics) is mathematically predicted directly from the mel-cepstrum 1523.

[0077] By employing the configuration shown in FIG. 6, the amount of processing by the second latent feature prediction unit 106 is reduced, and therefore the processing speed can be increased.

[0078] (a8: Fourth Modification) Fig. 7 is a block diagram showing an example of the configuration of a speech waveform generation system 1D according to a fourth modification of the present embodiment. Referring to Fig. 7, the speech waveform generation system 1D includes a speech waveform prediction unit 100D and an acoustic feature determination unit 140.

[0079] The speech waveform prediction unit 100D differs from the speech waveform prediction unit 100B shown in FIG. 5 in that it does not include the configuration for predicting a first finite impulse response 130 (residual characteristics) (the fourth latent feature prediction unit 128, the first finite impulse response prediction unit 110, and the first finite impulse response superposition unit 114).

[0080] By employing the configuration shown in FIG. 7, the amount of processing required to predict the first finite impulse response 130 (residual characteristics) is reduced, thereby enabling an increase in processing speed.

[0081] (a9: Fifth Modification) Fig. 8 is a block diagram showing an example of the configuration of a speech waveform generation system 1E according to a fifth modification of the present embodiment. Referring to Fig. 8, the speech waveform generation system 1E includes a speech waveform prediction unit 100E and an acoustic feature determination unit 140.

[0082] The speech waveform prediction unit 100E includes a sound source signal generation unit 102E, a latent feature prediction unit 134, a first finite impulse response prediction unit 110E, a second finite impulse response prediction unit 112E, a first finite impulse response superimposition unit 114, a second finite impulse response superimposition unit 116, and a signal addition unit 144.

[0083] The latent feature prediction unit 134 is a trained model, and predicts the latent feature 136 from the acoustic feature 152 (the mel-cepstrum 1523 and the aperiodicity index 1522).

[0084] The first finite impulse response prediction unit 110E is a trained model, and predicts a first finite impulse response 130E from the latent feature 136. The first finite impulse response 130E represents the characteristics of a non-periodic signal.

[0085] The second finite impulse response prediction unit 112E is a trained model, and predicts a second finite impulse response 132E from the latent feature 136. The second finite impulse response 132E represents the characteristics of a periodic signal.

[0086] The sound source signal generating unit 102E generates an F0 synchronization pulse signal 120P and a noise signal 120N, which depend on the fundamental frequency. P (t), and an excitation signal s representing the noise signal 120N. N (t) can be expressed as follows:

[0087] [Voiced] s P (t) = g p p(t) s N (t) = 0 [silent case] s P (t) = 0 s N (t) = g n The first finite impulse response superimposing section 114 superimposes a first finite impulse response 130E on the F0 synchronization pulse signal 120P. The second finite impulse response superimposing section 116 superimposes a second finite impulse response 132 on the noise signal 120N.

[0088] The signal adder 144 adds the signal generated by the first finite impulse response superimposing section 114 and the signal generated by the second finite impulse response superimposing section 116 together, and outputs the addition result as a predicted speech waveform 160 .

[0089] By adopting the configuration shown in FIG. 8, the first finite impulse response 130E indicating the non-periodic signal characteristics and the second finite impulse response 132E indicating the periodic signal characteristics can be superimposed independently of each other, thereby improving the audio quality.

[0090] (a10: Other Modified Examples) The present invention is not limited to the configuration example shown in FIG. 1 and the modified examples shown in FIGS. 4 to 8, and any modified configuration may be adopted as long as it is a configuration in which a finite impulse response is convolved with a sound source signal that is the basis of an audio waveform.

[0091] The acoustic features 152 provided to the speech waveform prediction unit (neural vocoder) may be determined by referring to the corpus 150 (FIGS. 1, 4, and 6 to 8), or may be predicted from the latent features 4 (FIG. 5). That is, in the configuration examples shown in FIGS. 1 and 4 to 9, the configuration for providing the acoustic features 152 may be changed as appropriate.

[0092] As for the sound source signal, a mixed excitation signal (FIGS. 1, 4, 6, and 7) may be used that mixes a pulse train (F0 synchronization pulse signal) according to the fundamental frequency 1521 (F0) with Gaussian noise (noise signal) based on the aperiodicity index 1522, or it may be possible to select either the pulse train or Gaussian noise, or a configuration (FIG. 8) in which the pulse train and Gaussian noise are separated may be used. In other words, the sound source signal that forms the basis of the speech waveform may be any of (1) a mixed excitation signal that mixes an F0 synchronization pulse signal and a noise signal, (2) only an F0 synchronization pulse signal, (3) a noise signal, or (4) an F0 synchronization pulse signal and a noise signal that are separately output.

[0093] [B. Learning of the Voice Waveform Prediction Unit] Next, learning of the voice waveform prediction unit included in the voice waveform generation system will be described.

[0094] Any algorithm can be used to train the speech waveform prediction unit (neural vocoder), but for example, adversarial training using a speech waveform discrimination model can be used. By using adversarial training, a speech waveform prediction unit that can generate high-quality speech can be generated.

[0095] 9 is a block diagram showing an example of the configuration of a learning system 10 for a speech waveform prediction unit according to the present embodiment. The learning system 10 learns (trains) a speech waveform prediction unit 100 (speech waveform prediction model). Referring to FIG. 9, the learning system 10 employs, as an example, a configuration following a generative adversarial network (GAN).

[0096] More specifically, the learning system 10 includes a speech waveform prediction unit 100 corresponding to a generator, and a first discrimination feature prediction unit 210 and a second discrimination feature prediction unit 220 corresponding to discriminators.

[0097] By referring to the corpus 150, a speech waveform 156 and an acoustic feature 152 representing the speech waveform 156 are extracted.

[0098] The speech waveform prediction unit 100 predicts a predicted speech waveform 160 from acoustic features 152 from a corpus 150. The behavior of the speech waveform prediction unit 100 (trained model) is defined by model parameters stored in a speech waveform generation model 170.

[0099] The generation error calculation unit 230 calculates the error between the predicted speech waveform 160 predicted by the speech waveform prediction unit 100 and the speech waveform 156 from the corpus 150. In other words, the generation error calculation unit 230 corresponds to a first error calculation unit that calculates the error between the predicted speech waveform 160 output when the acoustic feature 152 extracted from the corpus 150 is input to the speech waveform prediction unit 100 (speech waveform prediction model), and the speech waveform 156 associated with the extracted acoustic feature 152.

[0100] The first discrimination feature prediction unit 210 predicts a first discrimination feature from the predicted speech waveform 160 predicted by the speech waveform prediction unit 100. The second discrimination feature prediction unit 220 predicts a second discrimination feature from a speech waveform 156 from the corpus 150. The behavior of each of the first discrimination feature prediction unit 210 and the second discrimination feature prediction unit 220 is defined by model parameters stored in the speech waveform discrimination model 270. The first discrimination feature prediction unit 210 and the second discrimination feature prediction unit 220 correspond to discrimination feature prediction units that predict discrimination features from the predicted speech waveform 160 and the speech waveform 156, respectively.

[0101] The classification error calculation unit 240 calculates the error between the first classification feature predicted by the first classification feature prediction unit 210 and the second classification feature predicted by the second classification feature prediction unit 220. In other words, the classification error calculation unit 240 corresponds to a second error calculation unit that calculates the error between the classification feature predicted from the predicted speech waveform 160 and the classification feature predicted from the speech waveform 156.

[0102] The model update unit 250 updates (optimizes) the values ​​of the model parameters stored in the speech waveform generation model 170 and the values ​​of the model parameters stored in the speech waveform identification model 270 based on the errors calculated by the generation error calculation unit 230 and the errors calculated by the identification error calculation unit 240.

[0103] That is, the model updating unit 250 updates first model parameters that define the behavior of the speech waveform prediction unit 100 (finite impulse response prediction units (first finite impulse response prediction unit 110 and second finite impulse response prediction unit 112), finite impulse response superposition units (first finite impulse response superposition unit 114 and second finite impulse response superposition unit 116), and latent feature prediction units (first latent feature prediction unit 104, second latent feature prediction unit 106, and third latent feature prediction unit 108) based on the error calculated by the generation error calculation unit 230 (first error calculation unit).

[0104] Furthermore, the model update unit 250 updates second model parameters that define the behavior of the discrimination feature prediction units (the first discrimination feature prediction unit 210 and the second discrimination feature prediction unit 220) based on the error calculated by the discrimination error calculation unit 240 (second error calculation unit).

[0105] The model update unit 250 uses the loss to update the model parameters of the speech waveform prediction unit 100 (generator), for example, by using the least squares criterion L adv and the feature matching loss L fm and the Mel-cepstral L1 loss L of the speech waveform mel and the source excitation regularization loss L of the residual signal reg The four losses can be set as the adversarial losses. The objective function L including these four losses is G can be shown as follows:

[0106] L G =L adv +λ fm L fm +λ mel L mel +λ reg L reg where λ fm , λ mel , λ reg is a tunable hyperparameter. For example, λ fm = 2.0, λ mel = 50.0, λ reg = 20.0.

[0107] Fig. 10 is a block diagram showing an example of the configuration of a speech waveform prediction unit training system 10A according to a modified example of the present embodiment. Referring to Fig. 10, compared to the speech waveform prediction unit training system 10 shown in Fig. 9, the speech waveform prediction unit training system 10A includes two speech waveform prediction units 100, a fundamental frequency control unit 280, and an extended generation error calculation unit 260.

[0108] The first speech waveform prediction unit 100-1 is provided with acoustic features 152 from a corpus 150.

[0109] Fundamental frequency control unit 280 controls (changes) fundamental frequency 1521 (F0) included in acoustic feature 152 from corpus 150. Second speech waveform prediction unit 100-2 is provided with acoustic feature 152 with controlled fundamental frequency 1521. That is, fundamental frequency control unit 280 and second speech waveform prediction unit 100-2 predict the speech waveform when fundamental frequency 1521 is controlled.

[0110] The extended generation error calculation unit 260 calculates the error between the predicted speech waveform (with the fundamental frequency controlled) predicted by the second speech waveform prediction unit 100-2 and the speech waveform 156 from the corpus 150. Ideally, only the fundamental frequency differs between the predicted speech waveform predicted by the second speech waveform prediction unit 100-2 and the speech waveform 156 from the corpus 150, so feature quantities other than the fundamental frequency may be treated as loss.

[0111] C. Performance Evaluation Next, an example of performance evaluation of the speech waveform generation system according to the present embodiment will be described. In the following objective and subjective evaluations, 1000 sentences of speech by a male speaker were used.

[0112] (c1: Objective Evaluation) An example of objective evaluation of the performance of speech waveform prediction unit 100 (hereinafter referred to as "FIRNet") according to this embodiment is shown below. The following table shows the results of evaluation of the mel-cepstral distortion (MCD [dB]), the root mean square error (RMSE) of log f0, the voiced / unvoiced sound determination error (VUVE [%]), and the real-time coefficient (RTF) for the fundamental frequency under seven scaling conditions (1.0, 0.00, 0.25, 0.5, 2.0, 4.00, 8.00).

[0113] In the table below, "WORLD" indicates the neural vocoder proposed in Non-Patent Document 1, and "SiFi-GAN" indicates the neural vocoder proposed in Non-Patent Document 4.

[0114]

[0115] RMSE and VUVE are 0.0×f 0 and 8.0 × f 0It was not possible to calculate.

[0116] The speech waveform prediction unit 100 according to this embodiment has a lower MCD at all fundamental frequencies than the two conventional techniques, which means that the speech waveform prediction unit 100 according to this embodiment generates speech of higher quality.

[0117] In terms of the processing speed indicated by RTF, it can be seen that the speech waveform prediction unit 100 according to this embodiment is approximately five times faster than SiFi-GAN. This is thought to be because the speech waveform prediction unit 100 according to this embodiment employs a classical technique using an FIR filter, and therefore does not require complex processing of time-domain signals. Furthermore, the speech waveform prediction unit 100 according to this embodiment is faster than WORLD at high fundamental frequencies. This is because the speech waveform generation speed of WORLD depends on the pitch interval.

[0118] (c2: Subjective Evaluation) For the subjective evaluation, a Mean Opinion Score (MOS) test was conducted on a five-point scale to evaluate the naturalness of the synthesized speech. Twenty evaluators participated in the test, and evaluations were conducted under seven different scaling conditions, similar to the objective evaluation described above. In addition to the speech waveforms (Original) included in the corpus, the speech waveforms generated by the speech waveform prediction unit 100 according to this embodiment, WORLD, and SiFi-GAN were also evaluated. The evaluators evaluated 12 samples for each scaling condition and neural vocoder.

[0119] 11 is a graph showing an example of a subjective evaluation result of the audio waveform prediction unit 100 according to the present embodiment. Referring to FIG. 11, it can be seen that for SiFi-GAN, high audio quality is maintained when the scaling conditions are 1.0, 0.5, and 2.0, but audio quality is significantly degraded under other scaling conditions.

[0120] In contrast, it can be seen that the speech waveform prediction unit 100 according to this embodiment maintains speech quality similar to that of WORLD when the scaling conditions are 0.0, 0.2, 0.5, 4.0, and 8.0. Furthermore, when the fundamental frequency is not manipulated (scaling condition is 1.0), the speech waveform prediction unit 100 according to this embodiment achieves higher speech quality than WORLD, demonstrating that it is possible to achieve both excellent robustness against fundamental frequency control and high speech quality.

[0121] (c3: Summary) As the above objective and subjective evaluations show, the speech waveform prediction unit 100 according to this embodiment has fundamental frequency controllability and processing speed equivalent to conventional signal processing vocoders while maintaining the quality of a neural vocoder.

[0122] [D. Hardware Configuration Example] Next, a hardware configuration example for realizing the voice waveform generation system and learning system according to the present embodiment will be described. The voice waveform generation system and learning system according to the present embodiment may be realized using the same computing resource or different computing resources. The computing resource is provided, for example, using a general-purpose computer.

[0123] FIG. 12 is a schematic diagram showing an example of a hardware configuration for realizing a system according to the present embodiment.

[0124] 12 , the information processing device 300 includes, as main hardware components, a central processing unit (CPU) 302, a graphics processing unit (GPU) 304, a main memory 306, an input device 308, a network interface (I / F) 310, a storage 312, an input interface 322, an output interface 324, and an optical drive 326. These components are connected to one another via an internal bus 330.

[0125] The CPU 302 and / or the GPU 304 are processors that execute processes necessary to implement the system. A plurality of CPUs 302 and GPUs 304 may be provided, and each may have a plurality of cores.

[0126] The main memory 306 is a storage area that temporarily stores (or caches) program code, work data, etc. when the processor (CPU 302 and / or GPU 304) executes processing, and is composed of volatile memory such as DRAM (dynamic random access memory) or SRAM (static random access memory).

[0127] The input device 308 is a device that accepts instructions and operations from the user, and is configured with, for example, a keyboard, a mouse, a touch panel, a pen, and the like.

[0128] The network interface 310 exchanges data with any information processing device on the Internet or an intranet. Any communication method such as Ethernet (registered trademark), a wireless LAN (local area network), or Bluetooth (registered trademark) can be used as the network interface 310.

[0129] The input interface 322 receives an audio signal from a microphone 332. The output interface 324 outputs the audio signal to a speaker 334.

[0130] The optical drive 326 reads information stored on an optical disk 328, such as a CD-ROM (compact disc read only memory) or a DVD (digital versatile disc), and outputs the information to other components via the internal bus 330. The optical disk 328 is an example of a non-transitory recording medium, and is distributed with any program stored therein in a non-volatile manner. The optical drive 326 reads the program from the optical disk 328 and installs it in the storage 312 or the like, causing the computer to function as the information processing device 300.

[0131] The storage 312 stores programs and data necessary for implementing the system. The storage 312 is configured with a non-volatile storage device such as a hard disk or a solid state drive (SSD).

[0132] More specifically, the storage 312 stores an acoustic feature generation program 314, a speech waveform prediction program 316, a learning program 318, and the corpus 150, in addition to an operating system (OS) not shown.

[0133] The acoustic feature generation program 314 includes computer-readable instructions for generating acoustic features. The acoustic feature generation program 314 realizes the acoustic feature determination unit 140 or the acoustic feature prediction unit 142. The acoustic feature generation program 314 may be a neural network that is a trained model.

[0134] The speech waveform prediction program 316 includes computer-readable instructions for implementing a neural vocoder, which is a trained model. The acoustic feature generation program 314 implements the speech waveform prediction unit 100 (and variations of the speech waveform prediction unit). The acoustic feature generation program 314 may include an algorithm, model parameters, and hyperparameters.

[0135] The training program 318 includes computer-readable instructions for training the acoustic feature generation program 314 (neural vocoder). The training program 318 implements the first discriminant feature prediction unit 210, the second discriminant feature prediction unit 220, the generation error calculation unit 230, the discrimination error calculation unit 240, and the model update unit 250. The training program 318 may be a neural network that is a trained model.

[0136] At least one of the acoustic feature generation program 314, the speech waveform prediction program 316, and the learning program 318 may include an algorithm, a model parameter, and a hyperparameter. Also, at least one of the acoustic feature generation program 314, the speech waveform prediction program 316, and the learning program 318 may be an executable program built from an algorithm, a model parameter, and a hyperparameter.

[0137] Corpus 150 may include a data set of speech waveforms and acoustic features that describe the speech waveforms.

[0138] Some of the libraries and functional modules required when the processor (CPU 302 and / or GPU 304) executes a program (interpreter program) may be replaced with libraries or functional modules provided as standard by the OS. In this case, the program itself does not include all of the program modules required to realize the corresponding functions, but the desired processing can be achieved by installing the program in the OS execution environment. Furthermore, general-purpose libraries or functional modules licensed for use under a specific license may be used. Even a program that does not include some of these libraries or functional modules is within the technical scope of the present invention.

[0139] Furthermore, these programs may not only be distributed by being stored on any of the above-mentioned recording media, but may also be distributed by being downloaded from a server or the like via the Internet or an intranet.

[0140] FIG. 12 shows an example of a configuration using a single computer, but this is not limiting; multiple computers connected via a computer network may work together explicitly or implicitly to execute the processing required to realize the system.

[0141] All or part of the functions realized by the processor (CPU 302 and / or GPU 304) executing the programs may be realized using a hard-wired circuit such as an integrated circuit, for example, an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0142] Those skilled in the art will be able to implement the information processing device 300 according to this embodiment by appropriately using technology appropriate for the era in which the present invention is implemented.

[0143] [E. Processing Procedure] Next, an example of a processing procedure of the system according to the present embodiment will be described.

[0144] (e1: Speech waveform generation process) Fig. 13 is a flowchart showing an example of speech waveform generation process by the speech waveform generation system according to the present embodiment. The speech waveform generation process includes a speech waveform generation method that outputs a predicted speech waveform 160 based on input acoustic features 152. Each step shown in Fig. 13 may be realized by the processor of the information processing device 300 executing the acoustic feature generation program 314 and the speech waveform prediction program 316.

[0145] Referring to FIG. 13, information processing device 300 determines acoustic features 152 according to input information 2 (step S100).

[0146] The information processing device 300 outputs a predicted speech waveform 160 based on the input acoustic feature 152. More specifically, the information processing device 300 generates a sound source signal 120 including a component synchronized with a fundamental frequency 1521 (step S102).

[0147] The information processing device 300 predicts at least one finite impulse response based on at least a part of the acoustic features 152. More specifically, the information processing device 300 predicts a first finite impulse response 130 from the acoustic features 152 (step S104). The information processing device 300 predicts a second finite impulse response 132 from the acoustic features 152 (step S106).

[0148] The information processing device 300 convolves at least one finite impulse response on the sound source signal 120. More specifically, the information processing device 300 convolves a first finite impulse response 130 on the sound source signal 120 (step S108). The information processing device 300 further convolves a second finite impulse response 132 on the sound source signal 120 (step S110). The information processing device 300 outputs the sound source signal 120 on which the first finite impulse response 130 and the second finite impulse response 132 have been convolved as a speech waveform (step S112). Then, the processes from step S100 onwards are repeated.

[0149] Note that steps S102 to S106 may be executed in parallel rather than serially. Steps S102 to S106 may be executed in any order. Furthermore, steps S108 and S110 may be executed in any order.

[0150] (e2: Learning process of the speech waveform prediction unit) Fig. 14 is a flowchart showing an example of a learning method for the speech waveform generation system according to the present embodiment. Each step shown in Fig. 14 may be realized by the processor of the information processing device 300 executing the learning program 318.

[0151] Referring to FIG. 14, the information processing device 300 extracts pairs of speech waveforms 156 and acoustic features 152 included in the corpus 150 (step S200), and inputs the acoustic features 152 to the speech waveform prediction unit 100 to generate a predicted speech waveform 160 (step S202).

[0152] The information processing device 300 predicts a first discrimination feature from the predicted speech waveform 160 (step S204). That is, the information processing device 300 calculates the error between the predicted speech waveform 160 output when the acoustic feature 152 extracted from the corpus 150 in which the speech waveform 156 and the acoustic feature 152 are associated is input to the speech waveform prediction unit 100 (speech waveform prediction model), and the speech waveform 156 associated with the extracted acoustic feature 152.

[0153] Furthermore, the information processing device 300 predicts a second discrimination feature from the speech waveform 156 (step S206).

[0154] When the processing of steps S200 to S206 has been repeated a predetermined number of times (YES in step S208), information processing device 300 updates the model parameters of speech waveform prediction unit 100 (step S210) based on the error (loss) between predicted speech waveform 160 and speech waveform 156. That is, information processing device 300 updates the first model parameters that define the behavior of speech waveform prediction unit 100 based on the error between predicted speech waveform 160 and speech waveform 156.

[0155] The information processing device 300 updates the model parameters of the first discrimination feature prediction unit 210 and the second discrimination feature prediction unit 220 based on the error (loss) between the first discrimination feature and the second discrimination feature (step S212). That is, the information processing device 300 updates the second model parameters that define the behavior of the discrimination feature prediction units (the first discrimination feature prediction unit 210 and the second discrimination feature prediction unit 220) based on the error between the first discrimination feature and the second discrimination feature.

[0156] Note that the processes of step S210 and step S212 may be executed alternately when the condition of step S208 is satisfied. For example, the information processing device 300 may execute the update of the first model parameter (step S210) and the update of the second model parameter (step S212) in a predetermined order (e.g., alternately). In this way, the execution order of step S210 and step S212 may be any order.

[0157] The information processing device 300 determines whether a learning termination condition is met (step S214). The learning termination condition may be that the number of updates of the model parameters reaches a predetermined number, or that the error (loss) falls below a predetermined value.

[0158] If the learning termination condition is not met (NO in step S214), the process from step S200 onwards is repeated.

[0159] If the learning end condition is met (YES in step S214), the learning process ends. The speech waveform prediction unit 100 is configured using the model parameters at this point.

[0160] [F. Advantages] According to this embodiment, the sound source signal that forms the basis of the speech waveform is generated using a mechanism similar to that of a signal processing vocoder, and a finite impulse response is predicted based on acoustic features, and the predicted finite impulse response is then superimposed on the sound source signal to generate the speech waveform. By adopting this mechanism, it is possible to realize a system that maintains the speech quality of a neural vocoder while having processing speed and fundamental frequency controllability comparable to that of a signal processing vocoder.

[0161] By utilizing the advantages of this embodiment, it is possible to apply it to all services that require voice synthesis, such as a voice translation system, text-to-speech synthesis, voice quality conversion, and singing voice synthesis.

[0162] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the description of the above embodiments, and is intended to include all modifications within the meaning and scope of the claims.

[0163] 1, 1A, 1B, 1C, 1D, 1E Speech waveform generation system, 2 Input information, 4 Latent feature, 10, 10A Learning system, 12, 120 Sound source signal, 100, 100A, 100B, 100C, 100D, 100E Speech waveform prediction unit, 100-1 First speech waveform prediction unit, 100-2 Second speech waveform prediction unit, 102, 102E Sound source signal generation unit, 104 First latent feature prediction unit, 106 Second latent feature prediction unit, 108 Third latent feature prediction unit, 110, 110E First finite impulse response prediction unit, 112, 112C, 112E Second finite impulse response prediction unit, 114 First finite impulse response superposition unit, 116 Second finite impulse response superposition unit, 120N Noise signal, 120P Synchronization pulse signal, 122, first latent feature, 124, second latent feature, 126, third latent feature, 128, fourth latent feature prediction unit, 130, 130E, first finite impulse response, 132,132E second finite impulse response, 134 latent feature prediction unit, 136 latent feature, 140 acoustic feature determination unit, 142 acoustic feature prediction unit, 144 signal addition unit, 150 corpus, 152 acoustic feature, 156 speech waveform, 160 predicted speech waveform, 170 speech waveform generation model, 210 first discrimination feature prediction unit, 220 second discrimination feature prediction unit, 230 generation error calculation unit, 240 discrimination error calculation unit, 250 model update unit, 260 extended generation error calculation unit, 270 speech waveform discrimination model, 280 fundamental frequency control unit, 300 information processing device, 302 CPU, 304 GPU, 306 main memory, 308 input device, 310 network interface, 312 storage, 314 acoustic feature generation program, 316 speech waveform prediction program, 318 Learning program, 322 input interface, 324 output interface, 326 optical drive, 328 optical disk, 330 internal bus, 332 microphone, 334 speaker, 1040 Causal ConvNeXt core block, 1041 input depthwise convolution layer, 1042 layer normalization, 1043 pointwise convolution layer, 1044 global response normalization, 1140 layer, 1141 adder, 1142 finite impulse response filter, 1143 ConvNeXt core block, 1144 linear layer, 1145 extended causal convolution layer, 1146 fully connected layer, 1147 concatenation layer, 1521 fundamental frequency, 1522 aperiodicity index, 1523 mel-cepstrum, 1524 unvoiced flag.

Claims

1. A speech waveform generation system comprising: a speech waveform prediction unit that outputs a predicted speech waveform based on input acoustic features, the speech waveform prediction unit comprising: a speech source signal generation unit that generates a speech source signal including a component synchronized with a fundamental frequency; a finite impulse response prediction unit that predicts at least one finite impulse response based on at least a portion of the acoustic features; and a finite impulse response superposition unit that superposes the at least one finite impulse response on the speech source signal.

2. The audio waveform generation system according to claim 1, wherein the finite impulse response prediction unit convolves the sound source signal with the at least one finite impulse response.

3. The audio waveform generation system according to claim 1, wherein the sound source signal includes a noise component independent of the fundamental frequency.

4. The audio waveform generation system of claim 1, wherein the at least one finite impulse response includes a finite impulse response representative of a resonant characteristic.

5. A speech waveform generation method for outputting a predicted speech waveform based on input acoustic features, comprising the steps of: generating a sound source signal including a component synchronized with a fundamental frequency; predicting at least one finite impulse response based on at least a portion of the acoustic features; and superimposing the at least one finite impulse response on the sound source signal.

6. A speech waveform prediction program that outputs a predicted speech waveform based on input acoustic features, the speech waveform prediction program causing a computer to execute the steps of: generating a sound source signal including a component synchronized with a fundamental frequency; predicting at least one finite impulse response based on at least a portion of the acoustic features; and superimposing the at least one finite impulse response on the sound source signal.

Citation Information

Patent Citations

  • Speech synthesis system, speech synthesis program and speech synthesis method

    JP2018141915A

  • Speech processing device and speech processing method

    JP2020204651A