Voice synthesis device, voice synthesis method, and program

The voice synthesis device uses a break flag detection system to prevent errors in RNN-based speech synthesis, ensuring natural and efficient speech generation.

JP7709646B2Active Publication Date: 2025-07-17NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023567286
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2025-07-17
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

Autoregressive models like RNNs in speech synthesis face issues with increasing errors in predicted speech waveforms leading to unclear or silent speech, and traditional methods to prevent this impair operation speed.

Method used

A voice synthesis device that combines a recurrent neural network with a break flag detection system, using a break flag conversion unit, break detection unit, and break detection model learning to prevent waveform generation failures while maintaining operation speed.

Benefits of technology

Prevents a decrease in speech naturalness and significant reduction in operation speed by detecting and initializing RNN states to maintain continuous and clear speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709646000045
    Figure 0007709646000045
  • Figure 0007709646000046
    Figure 0007709646000046
  • Figure 0007709646000047
    Figure 0007709646000047
Patent Text Reader

Abstract

The purpose of the present disclosure is to prevent a decrease in the naturalness of speech while preventing the operation speed of waveform generation from being markedly harmed. Thus, the present disclosure is a speech synthesis device that generates a speech waveform in a learning phase, comprising: a waveform generation unit that obtains a speech waveform predicted value for a next time on the basis of a combination of the speech waveform and an acoustic feature quantity, as well as a state of a recurrent neural network; a conversion-to-failure-flag unit that obtains a failure flag on the basis of the speech waveform, the speech waveform predicted value, and a failure flag threshold value; a failure detection unit that obtains a failure flag predicted value on the basis of a sequence of states of the recurrent neural network and a failure detection model; a failure flag discrepancy calculation unit that calculates the discrepancy between the failure flag and the failure flag predicted value; and a failure detection model learning unit that obtains a learned failure detection model on the basis of the discrepancy and the failure detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a speech synthesis device, a speech synthesis method, and a program.

Background Art

[0002] In speech synthesis, a module that converts from acoustic feature quantities such as a spectrum and a pitch representing voice height into a speech waveform is called a vocoder. There are roughly two types of implementation methods for a vocoder. One is a method by signal processing, and techniques such as STRAIGHT (Non-Patent Document 1) and WORLD (Non-Patent Document 2) are well-known. Since these methods express the conversion from acoustic feature quantities to a speech waveform by a mathematical model, learning is not required and the processing speed is high, but the quality is inferior when comparing the analyzed and resynthesized speech with natural speech. The second is a method by a neural network (neural vocoder), and WaveNet is a typical technique thereof (Patent Document 1). While it is possible to synthesize speech of comparable quality even when compared with natural speech, since it is based on a huge convolutional neural network (CNN: Convolutional Neural Network), the amount of calculation is large, the operation is slower than that of the signal processing vocoder, and real-time operation is difficult.

[0003] Therefore, in order to perform real-time operation on a CPU, it is necessary to reduce the amount of calculation. As a main approach, there is WaveRNN that replaces the huge CNN used in WaveNet with a small-scale recurrent neural network (RNN: Recurrent Neural Network) (Patent Document 2). In addition, in LPCNet (Non-Patent Document 3), linear prediction analysis (LPC), which is knowledge of signal processing, is introduced into the process of generating a speech waveform, enabling speech synthesis with a deeper neural network (DNN: Deep Neural Network) that is even smaller than WaveRNN. Thus, in WaveRNN and LPCNet, an RNN is used to realize a small-scale speech synthesis DNN.

Prior Art Documents

Non-Patent Literature

[0004]

Non-Patent Literature 1

Non-Patent Literature 2

Non-Patent Literature 3

Patent Literature

[0005]

Patent Literature 1

Patent Literature 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, in autoregressive models such as RNN, since the predicted speech waveform values are used as the speech waveform values at the next time step, as the series to be predicted becomes longer, the error from the learning phase increases. As a result, waveform generation fails, not only making the speech unclear, but in the worst case, it may become silent. Also, although it is possible to avoid failure by initializing the state variable of the RNN at regular intervals, since the time series information up to that time is initialized, it becomes discontinuous, especially in the case of initialization in the voiced section, which leads to a decrease in the naturalness of the speech.

[0007] Also, for detecting the failure of waveform generation, it is conceivable to execute waveform generation by signal processing with low quality but without failure together with the neural vocoder and compare the results. However, since the vocoder has to be executed by two methods, the operation speed of waveform generation is significantly impaired.

[0008] The present invention has been made in view of the above points, and an object thereof is to prevent a decrease in the naturalness of speech while preventing a significant impairment of the operation speed of waveform generation.

Means for Solving the Problems

[0009] In order to solve the above problems, the invention according to claim 1 is a voice synthesis device that generates a voice waveform in a learning phase, and based on the combined voice waveform and acoustic feature amount, and the state of the recurrent neural network, a waveform generation unit that obtains a predicted value of the voice waveform at the next time; a break flag conversion unit that obtains the break flag based on the voice waveform, the predicted value of the voice waveform, and a threshold value of the break flag indicating whether the voice waveform is broken at each time; a break detection unit that obtains a predicted value of the break flag based on the state sequence of the recurrent neural network and a break detection model; a break flag error calculation unit that calculates an error between the break flag and the predicted value of the break flag; and a break detection model learning unit that obtains a learned break detection model based on the error and the break detection model.

Advantages of the Invention

[0010] As described above, according to the present invention, it is possible to prevent a decrease in the naturalness of voice while preventing a significant reduction in the operation speed of waveform generation.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0013] 〔System Configuration of Embodiment〕 First, the outline of the configuration of the communication system 1 of the present embodiment will be described with reference to FIG. 1. FIG. 1 is a schematic diagram of the communication system according to the present embodiment.

[0014] As shown in FIG. 1, the communication system 1 of the present embodiment is constructed by a voice synthesizer 3 and a communication terminal 5. The communication terminal 5 is managed and used by the user Y.

[0015] Also, the voice synthesizer 3 and the communication terminal 5 can communicate via a communication network 100 such as the Internet. The connection form of the communication network 100 may be either wireless or wired.

[0016] The voice synthesizer 3 is constituted by one or a plurality of computers. When the voice synthesizer 3 is constituted by a plurality of computers, it may be referred to as a "voice synthesizer" or a "voice synthesis system".

[0017] The voice synthesizer 3 is a computer and is a device that generates a voice waveform for voice synthesis using a breakdown detection technique.

[0018] The communication terminal 5 is a computer. In FIG. 1, as an example, a notebook personal computer is shown, but it is not limited to the notebook type and may be a desktop personal computer. Also, the communication terminal may be a smartphone or a tablet terminal. In FIG. 1, the user Y is operating the communication terminal 5.

[0019] 〔Hardware Configuration of Voice Synthesizer and Communication Terminal〕 Next, the hardware configurations of the voice synthesizer 3 and the communication terminal 5 will be described with reference to FIG. 2. FIG. 2 is a hardware configuration diagram of the voice synthesizer and the communication terminal according to the present embodiment. Note that the hardware configurations of the voice synthesizer and the communication terminal are common in the first to fourth embodiments described later.

[0020] As shown in FIG. 2, the voice synthesizer 3 includes a processor 301, a memory 302, an auxiliary storage device 303, a connection device 304, a communication device 305, and a drive device 306. Note that each piece of hardware constituting the voice synthesizer 3 is interconnected via a bus 307.

[0021] The processor 301 serves as a control unit that controls the entire voice synthesizer 3 and has various arithmetic devices such as a CPU (Central Processing Unit). The processor 301 reads various programs onto the memory 302 and executes them. Note that the processor 301 may include a GPGPU (General-purpose computing on graphics processing units).

[0022] The memory 302 has main memory devices such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The processor 301 realizes various functional units described later by executing various programs read onto the memory 302.

[0023] The auxiliary storage device 303 stores various programs and various information (such as the breakdown detection model 30a and the learned breakdown detection model 30b described later) used when the various programs are executed by the processor 301.

[0024] The connection device 304 is a connection device that connects an external device (for example, the display device 310 and the operation device 311) to the voice synthesizer 3.

[0025] The communication device 305 is a communication device for transmitting and receiving various information to and from other devices.

[0026] The drive device 306 is a device for setting the recording medium 330. The recording medium 330 here includes media that optically, electrically, or magnetically record information, such as a CD-ROM (Compact Disc Read-Only Memory), a flexible disk, and a magneto-optical disk. The recording medium 330 may also include semiconductor memories such as a ROM (Read Only Memory) and a flash memory that electrically record information.

[0027] Note that various programs installed in the auxiliary storage device 303 are installed, for example, when the distributed recording medium 330 is set in the drive device 306 and the various programs recorded on the recording medium 330 are read by the drive device 306. Alternatively, various programs installed in the auxiliary storage device 303 may be installed by being downloaded from a network via the communication device 305.

[0028] Also, although the hardware configuration of the communication terminal 5 is shown in FIG. 2, since each configuration is the same except that the reference numerals change from the 300s to the 500s, the descriptions thereof are omitted.

[0029] ●First Embodiment The first embodiment will be described with reference to FIGS. 3 to 6.

[0030] 〔Functional Configuration of Voice Synthesis Device〕 The functional configuration of the voice synthesis device according to the first embodiment will be described with reference to FIGS. 3 and 4.

[0031] <Functional Configuration of Voice Synthesis Device in Learning Phase> FIG. 3 is a functional configuration diagram of the voice synthesis device in the learning phase according to the first embodiment. As shown in FIG. 3, the voice synthesis device 3 includes an input unit 31, a waveform generation unit 32, a conversion unit 33 to a breakdown flag, a breakdown detection unit 34, an error calculation unit 35 of the breakdown flag, and a breakdown detection model learning unit 36.

[0032] Among these, the input unit 31 combines the input voice waveform and acoustic feature amount.

[0033] The waveform generation unit 32 obtains a predicted value of the voice waveform at the next time based on the combined voice waveform and acoustic feature amount, and the state of the recurrent neural network.

[0034] The conversion unit 33 to the breakdown flag obtains a breakdown flag indicating whether the speech waveform at each time is broken based on the speech waveform, the predicted value of the speech waveform, and the threshold value of the breakdown flag.

[0035] The breakdown detection unit 34 obtains a predicted value of the breakdown flag based on the state sequence of the recurrent neural network (RNN: Recurrent Neural Network) and the breakdown detection model 30a.

[0036] The error calculation unit 35 for the breakdown flag calculates the error between the breakdown flag and the predicted value of the breakdown flag.

[0037] The breakdown detection model learning unit 36 obtains a learned breakdown detection model 30b based on the error and the breakdown detection model.

[0038] Note that each of the above functional configurations will be described in detail later.

[0039] <Functional Configuration in the Inference Phase of the Speech Synthesis Device> FIG. 4 is a functional configuration diagram in the inference phase of the speech synthesis device according to the first embodiment. As shown in FIG. 4, the speech synthesis device 3 includes an input unit 31, a waveform generation unit 32, a conversion unit 33 to the breakdown flag, a breakdown detection unit 34, and a state initialization unit 37. Note that the same functional configurations as those in the learning phase are denoted by the same reference numerals and the description thereof is omitted.

[0040] When the breakdown flag (predicted value) indicates that it is predicted to be broken by indicating that it is "broken", the state initialization unit 37 initializes the state of the RNN based on the initial value of the state of the RNN. Note that this functional configuration will be described in detail later.

[0041] [Processing or Operation of the Speech Synthesis Device] Subsequently, with reference to FIGS. 5 and 6, the processing or operation of the speech synthesis device according to the first embodiment will be described.

[0042] <Processing or Operation in the Learning Phase of the Speech Synthesis Device> FIG. 5 is a flowchart showing processing or operations in a learning phase of the speech synthesis device according to the first embodiment.

[0043] First, as shown in FIG. 5, the input unit 31 combines the speech waveform of the learning data at time t

[0044]

Number

[0045] Next, the waveform generation unit 32 generates the speech waveform (predicted value) at the next time based on the combined speech waveform, acoustic feature amount, and the state of the RNN

[0046]

Number

[0047]

Number

[0048] and obtains the speech waveform

[0049]

Number

[0050]

Number

[0051]

Number

[0052] Here, the breakdown flag is a binary flag indicating whether the speech waveform at each time t is broken. In the process (S13), the conversion unit 33 to the breakdown flag compares x with

[0053]

Number

[0054]

Number

[0055]

Number

[0056]

Number

[0057]

Number

[0058]

Number

[0059] Next, the failure flag error calculation unit 35 acquires the failure flag and the failure flag (predicted value), and calculates the error between the failure flag and the failure flag (predicted value) (S15). Since the task of the failure detection model 30a is a classification problem of whether there is a failure, when adopting a DNN as a statistical model, cross-entropy or the like can be used as an error function.

[0060] Next, the failure detection model learning unit 36 obtains a learned failure detection model 30b based on the error calculated by the failure flag error calculation unit 35 and the failure detection model 30a (S16). This process (S16) is achieved by updating the parameters of the failure detection model 30a so as to minimize the error, and error backpropagation is generally used in a DNN. By repeatedly executing the above steps for all of the learning data, the prediction accuracy of the failure detection model 30a is improved. In this way, the inference phase ends.

[0061] <Processing or operation in the inference phase of the speech synthesis device> FIG. 6 is a flowchart showing the processing or operation in the inference phase of the speech synthesis device according to the first embodiment.

[0062] First, as shown in FIG. 6, the input unit 31, in the same manner as the above-described processing (S11), obtains the speech waveform of the learning data at time t

[0063]

Number

[0064] Next, in the same manner as the above-described processing (S12), the waveform generation unit 32 uses the combined speech waveform and acoustic feature amounts, and the state of the RNN

[0065]

Number

[0066]

Number

[0067] Next, the breakdown detection unit 34 obtains the state of the RNN from the waveform generation unit 32

[0068]

Number

[0069]

Number

[0070] Next, when the state initialization unit 37 predicts breakdown by indicating that the breakdown flag (predicted value) indicates "breakdown", based on the initial value of the state of the RNN, the state of the RNN

[0071]

Number

[0072] In the above manner, the inference phase ends.

[0073] <Main effects of the first embodiment> As described above, according to the present embodiment, it is possible to prevent a decrease in the naturalness of speech while preventing a significant reduction in the operation speed of waveform generation. Specifically, it is possible to detect a breakdown in waveform generation only from the information obtained by the operation of the neural vocoder. As a result, by initializing the state variable of the RNN at the timing when the breakdown is detected, it is possible to avoid problems such as the state variable utterance becoming unclear or the voice becoming silent.

[0074] ●Second Embodiment Subsequently, a second embodiment will be described with reference to FIGS. 7 and 8.

[0075] 〔Functional Configuration of Speech Synthesis Device〕 Since the functional configuration of the speech synthesis device 3 according to the present embodiment in the learning phase is the same as that of the speech synthesis device 3 according to the first embodiment in the learning phase, the description thereof will be omitted.

[0076] <Functional Configuration of Speech Synthesis Device in Inference Phase> The speech synthesis device 3 according to the second embodiment further includes a speech waveform buffering unit 41, a speech waveform averaging unit 42, and a speech waveform selection unit 48 with respect to the speech synthesis device 3 according to the first embodiment.

[0077] Among these, the speech waveform buffering unit 41 is constructed in the memory 302 and accumulates speech waveforms (predicted values).

[0078] The speech waveform averaging unit 42 obtains an averaged speech waveform based on the immediately preceding speech waveform (predicted value) accumulated in the buffering unit 41.

[0079] The speech waveform selection unit 48 outputs a speech waveform (predicted value) for predicting the speech waveform at the next time based on the predicted value of the breakdown flag, the speech waveform (predicted value), and the averaged speech waveform.

[0080] Note that each of the above functional configurations will be described in detail hereinafter.

[0081] [Processing or operation of the speech synthesis device] Next, with reference to FIG. 8, the processing or operation of the speech synthesis device according to the second embodiment will be described. The processing or operation of the speech synthesis device 3 according to this embodiment is the same as the processing or operation of the speech synthesis device 3 according to the first embodiment in the learning phase, and only a part of the inference phase is different. Therefore, the description of the processing in the learning phase will be omitted.

[0082] [Processing or operation in the inference phase of the speech synthesis device] FIG. 8 is a flowchart showing the processing or operation in the inference phase of the speech synthesis device according to the second embodiment. Since the processing (S121 to S124) is the same as the processing (S111 to S114) of the first embodiment, the description thereof will be omitted.

[0083] In the process (S122), each time the waveform generation unit 32 obtains the speech waveform (predicted value) at the next time

[0084] [Number] by inputting this speech waveform (predicted value) into the speech waveform buffering unit 41, the speech waveform buffering unit 41 accumulates the speech waveform (predicted value) (S125).

[0085] Next, the speech waveform averaging unit 42 obtains the immediately preceding speech waveform (predicted value) accumulated in the buffering unit 41, and based on this speech waveform (predicted value), an averaged speech waveform

[0086] [Number] is obtained (S126). Here, the speech waveform averaging unit 42 can use a simple moving average that takes the average of the most recent N samples, a weighted moving average or an exponential moving average that emphasizes the speech waveform at the most recent time, and the like.

[0087] Next, the speech waveform selection unit 48 has a breakdown flag (predicted value)

[0088] [Number] Average voice waveform

[0089] [Number] And voice waveform (predicted value)

[0090] [Number] Based on this, it outputs as the voice waveform (predicted value) for predicting the voice waveform at the next time (S127). In this case, at time t + 1, when the breakdown detection unit 34 determines that the breakdown flag

[0091] [Number] is determined to be a breakdown, the voice waveform selection unit 48 does not select the voice waveform (predicted value) obtained by the waveform generation unit 32, but selects and outputs the average voice waveform. As described above, the inference phase ends.

[0092] [Main effects of the second embodiment] As described above, according to the present embodiment, in addition to the effects of the first embodiment, it is possible to eliminate the discontinuity of the voice, which is a side effect when initializing the state h of the RNN. In addition, the information of the voice waveform generated up to that time is accumulated in the state h of the RNN, and the information up to that point is lost due to the initialization of the state h. That is, the continuity of the voice before and after the initialization of the state h is also lost. In the present embodiment, in order to eliminate this discontinuity, the voice waveforms of several samples immediately before initialization are buffered in advance. As the voice waveform at the initialization time, instead of the one predicted by the waveform generation unit 32, the average value of the buffered voice is used. Thereby, while initializing the state h, the continuity with the previous voice can be ensured, and breakdown can be prevented without causing significant quality degradation.

[0093] ● Third Embodiment Next, the third embodiment will be described with reference to FIGS. 9 to 12.

[0094] 〔Functional Configuration of Voice Synthesis Device〕 The voice synthesis device 3 according to this embodiment further includes a statistic calculation unit 49 with respect to the functional configuration in the learning phase and the inference phase of the voice synthesis device 3 according to the first embodiment.

[0095] The statistic calculation unit 49 obtains the statistics of the state of the RNN. This functional configuration will be described in detail later.

[0096] 〔Processing or Operation of Voice Synthesis Device〕 Next, with reference to FIG. 11, the processing or operation of the voice synthesis device according to the third embodiment will be described.

[0097] <Processing or Operation in Learning Phase of Voice Synthesis Device> FIG. 11 is a flowchart showing the processing or operation in the learning phase of the voice synthesis device according to the third embodiment. The processing (S31, S32, S33, S34, S35, S36) corresponds to the processing (S11, S12, S13, S14, S15, S16) in the first embodiment, and since most of them are the same as those in the first embodiment, only the differences will be described.

[0098] In this embodiment, as in the first embodiment, after the processing (S31), the waveform generation unit 32 obtains the voice waveform (predicted value) at the next time (S32). Also, each time this prediction is made, the waveform generation unit 32 obtains the state of the RNN

[0099]

Number

[0100]

Number

[0101] ​In this embodiment, the statistic calculation unit 49 obtains the state of the RNN from the waveform generation unit 32, and calculates the statistic of the state of the RNN

[0102]

Number

[0103]

Number

[0104]

Number

[0105] After that, the statistic calculation unit 49 inputs the statistic of the state of the RNN to the breakdown detection unit 34

[0106]

Number

[0107] <Processing or operation in the inference phase of the speech synthesis device> FIG. 12 is a flowchart showing the processing or operation in the inference phase of the speech synthesis device according to the third embodiment. Note that the processing (S131, S132, S133, S134) corresponds to the processing (S111, S112, S113, S114) in the first embodiment, and since it is mostly the same as the first embodiment, only the differences will be described.

[0108] In this embodiment, similar to the first embodiment, after the processing (S131), the waveform generation unit 32 obtains the speech waveform (predicted value) at the next time

[0109]

Number

[0110]

Number

[0111] Thereafter, the statistic calculation unit 49 inputs the statistic of the state of the RNN to the breakdown detection unit 34

[0112]

Number

[0113] <Main effects of the third embodiment> As described above, according to the present embodiment, in addition to the effects of the first embodiment, the following effects can be achieved. That is, in the breakdown detection process described in the second embodiment, if the state h of the RNN is used as it is, since the dimensionality of the state h is large, the computational complexity is large and the waveform generation operation speed is impaired. On the other hand, according to the present embodiment, the computational complexity required for breakdown detection can be reduced, and the operation speed of waveform generation including breakdown detection can be improved. Further, it can be combined with the second embodiment, and waveform generation that ensures the continuity of the voice while reducing the computational complexity required for breakdown detection is possible.

[0114] ● Fourth Embodiment Next, the fourth embodiment will be described with reference to FIGS. 13 to 16.

[0115] 〔Functional Configuration of Voice Synthesis Device〕 The functional configuration of the voice synthesis device according to the fourth embodiment will be described with reference to FIGS. 13 and 14.

[0116] <Functional Configuration of Voice Synthesis Device in Learning Phase> In the voice synthesis device 3 according to the present embodiment, the conversion unit 33 to the breakdown flag, the breakdown detection unit 34, the error calculation unit 35 of the breakdown flag, and the breakdown detection model learning unit 36 in the voice synthesis device 3 of the first embodiment are respectively replaced by the conversion unit 43 to the index of breakdown detection, the prediction unit 44 of the index of breakdown detection, the error calculation unit 45 of the index of breakdown detection, and the index prediction model learning unit 46 of breakdown detection.

[0117] Among these, the conversion unit 43 to the index of breakdown detection obtains the index of breakdown detection based on the voice waveform and the voice waveform (predicted value).

[0118] The prediction unit 44 of the index of breakdown detection obtains the index of breakdown detection (predicted value) based on the state sequence of the RNN and the index prediction model 40a of breakdown detection.

[0119] The error calculation unit 45 of the index of breakdown detection calculates the error between the index of breakdown detection and the index of breakdown detection (predicted value).

[0120] Based on the error and the index prediction model 40a for flaw detection, the learning unit 46 for the index prediction model of flaw detection obtains the learned index prediction model 40b for flaw detection.

[0121] Each of the above functional configurations will be described in detail later.

[0122] <Functional Configuration in the Inference Phase of the Speech Synthesis Device> In the speech synthesis device 3 according to the present embodiment, the flaw detection unit 34 in the speech synthesis device 3 in the first embodiment is replaced by the prediction unit 44 for the index of flaw detection. Further, instead of inputting the flaw flag (predicted value) to the state initialization unit 37, the index (predicted value) of flaw detection is input, and further, the threshold value f of the flaw flag is input.

[0123] 〔Processing or Operation of Speech Synthesis Device〕 Subsequently, with reference to FIGS. 15 and 16, the processing or operation of the speech synthesis device according to the fourth embodiment will be described.

[0124] <Processing or Operation in the Learning Phase of the Speech Synthesis Device> FIG. 15 is a flowchart showing the processing or operation in the learning phase of the speech synthesis device according to the fourth embodiment. Since the processing (S41, S42) corresponds to the processing (S11, S12) in the first embodiment, only the differences will be described.

[0125] In the present embodiment, the conversion unit 43 to the index of flaw detection executes the above processing (S42) for times t = 1,..., T, thereby obtaining the speech waveform (predicted value)

[0126]

Number

[0127]

Number

[0128]

Number

[0129] Here, the breakdown detection index is the difference between the speech waveform x used to generate the breakdown flag in the first embodiment and its predicted value

[0130]

Number

[0131] Also, when the conversion unit 43 to the breakdown detection index acquires the speech waveform (predicted value)

[0132]

Number

[0133]

Number

[0134]

Number

[0135] Note that the prediction unit 44 of the breakdown detection index may obtain the statistical quantity of the state of the RNN via the statistical quantity calculation unit 49 in the same manner as in the third embodiment

[0136]

Number

[0137] Next, the error calculation unit 45 for the index of breakdown detection acquires the index of breakdown detection and the index of breakdown detection (predicted value), and calculates the error between the index of breakdown detection and the index of breakdown detection (predicted value) (S45). Since the index of breakdown detection and the index of breakdown detection (predicted value) are each continuous values, as a method for calculating the error, the mean squared error or the mean absolute error can be used in the same manner as the process (S14) of the first embodiment.

[0138] Next, the index prediction model learning unit 46 for breakdown detection obtains a learned index prediction model 40b for breakdown detection based on the error calculated by the error calculation unit 45 for the index of breakdown detection and the index prediction model 40a for breakdown detection (S46). This process (S46) is achieved by updating the parameters of the index prediction model 40a for breakdown detection so as to minimize the error, and error backpropagation is generally used in DNN. By repeatedly executing the above procedure for all of the learning data, the prediction accuracy of the index prediction model 40a for breakdown detection is improved. In this way, the inference phase ends.

[0139] <Processing or operation in the inference phase of the speech synthesis device> FIG. is a flowchart showing the processing or operation in the inference phase of the speech synthesis device according to the fourth embodiment. Since the processes (S141, S142) respectively correspond to the processes (S111, S112) in the first embodiment, only the differences will be described.

[0140] In the present embodiment, the conversion unit 43 to the index of breakdown detection obtains the state of the RNN from the waveform generation unit 32, and further obtains the learned index prediction model 40b for breakdown detection, and predicts the index of breakdown detection based on these, thereby obtaining the index of breakdown detection (predicted value)

[0141]

Number

Number

[0142]

Number

[0143] Next, when the index (predicted value) of the breakdown detection is greater than the threshold value f, the state initialization unit 37 determines that the waveform generation "is regarded as broken", and based on the initial value of the state of the RNN, the state of the RNN

[0144] [Number] is initialized.

[0145] In the above manner, the inference phase ends.

[0146] [Main effects of the fourth embodiment] As described above, in the breakdown detection process of the first to third embodiments, when tuning the accuracy of the breakdown detection model learned as the discrimination model, it is necessary to re-learn, such as changing the threshold value f of the breakdown flag in the learning phase or changing the learning conditions including the hyperparameters of the breakdown detection model. Therefore, a great deal of effort is required for tuning to obtain a model that is considered optimal. Further, since the breakdown detection model 30a is learned specifically for the waveform generation unit 32, when the model used in the waveform generation unit 32 changes, the breakdown detection model 30a also has to be retuned accordingly.

[0147] In contrast, in this embodiment, instead of learning a discrimination model that predicts a discrete breakdown flag, it learns as a generation model that predicts an index of continuous breakdown detection. Specifically, the speech synthesis device 3 of the fourth embodiment does not directly predict the breakdown flag from the statistical model, but predicts the value that is the index thereof, and indirectly detects the breakdown based on whether it exceeds the threshold value f. Thereby, it is only necessary to tune the threshold value f without re-learning the model used for breakdown detection, and the problems of the above-described first to third embodiments can be reduced.

[0148] In addition, this embodiment can also be combined with the second or third embodiment, enabling waveform generation that ensures the continuity of voice while reducing the computational complexity required for breakdown detection.

[0149] 〔Supplementary Note〕 The present invention is not limited to the above-described embodiments, and may have configurations or processes (operations) as described below.

[0150] The voice synthesis device 3 can be realized by a computer and a program. However, it is also possible to record this program on a (non-transitory) recording medium or to provide it via the communication network 100.

Explanation of Reference Numerals

[0151] 1 Communication system 3 Voice synthesis device 5 Communication terminal 30a Breakdown detection model 30b Learned breakdown detection model 31 Input unit 32 Waveform generation unit 33 Conversion unit to breakdown flag 34 Breakdown detection unit 35 Error calculation unit for breakdown flag 36 Breakdown detection model learning unit 37 State initialization unit 40a Index prediction model for breakdown detection 40b Learned index prediction model for breakdown detection 41 Buffer unit for voice waveform 42 Average unit for voice waveform 43 Conversion unit to index for breakdown detection 44 Prediction unit for index of breakdown detection 45 Error calculation unit for index of breakdown detection 46 Index prediction model learning unit for breakdown detection

Claims

1. An audio synthesis device that generates an audio waveform in a learning phase, a waveform generation unit that obtains a predicted value of an audio waveform at the next time based on the combined audio waveform and acoustic feature amount, and the state of a recurrent neural network; a conversion unit to the burst flag that obtains the burst flag based on the audio waveform, the predicted value of the audio waveform, and a threshold value of the burst flag indicating whether the audio waveform at each time is burst; a burst detection unit that obtains a predicted value of the burst flag based on the state sequence of the recurrent neural network and a burst detection model; an error calculation unit for the burst flag that calculates an error between the burst flag and the predicted value of the burst flag; a burst detection model learning unit that obtains a learned burst detection model based on the error and the burst detection model; An audio synthesis device having the above.

2. The audio synthesis device according to claim 1, comprising a total amount calculation unit that obtains a statistic of the state of the recurrent neural network based on the state of the recurrent neural network, wherein the burst detection unit obtains a predicted value of the burst flag based on the statistic of the state of the recurrent neural network substituted from the state sequence of the recurrent neural network and the burst detection model.

3. An audio synthesis device that generates an audio waveform in a learning phase, a waveform generation unit that obtains a predicted value of an audio waveform at the next time based on the combined audio waveform and acoustic feature amount, and the state of a recurrent neural network; a conversion unit to an index for burst detection that obtains an index for burst detection based on a difference between the audio waveform and the predicted value of the audio waveform; a prediction unit for an index for burst detection that obtains a predicted value of the index for burst detection based on the state sequence of the recurrent neural network and a prediction model for the index for burst detection; an error calculation unit for the index for burst detection that calculates an error between the index for burst detection and the predicted value of the index for burst detection; a prediction model learning unit for the index for burst detection that obtains a learned burst detection model based on the error and the prediction model for the index for burst detection; An audio synthesis device having the above.

4. An audio synthesis device that generates an audio waveform in an inference phase, a waveform generation unit that obtains a predicted value of an audio waveform at the next time based on the combined audio waveform and acoustic feature amount, and the state of a recurrent neural network; A breakdown detection unit that obtains a predicted value of a breakdown flag indicating whether the speech waveform at each time is broken based on the state of the recurrent neural network and the learned breakdown detection model; When the predicted value of the breakdown flag indicates a breakdown, a state initialization unit that initializes the state of the recurrent neural network based on the initial value of the state of the recurrent neural network; A speech synthesis device having the above.

5. The speech synthesis device according to claim 4, A speech waveform buffering unit that accumulates the predicted value of the speech waveform; A speech waveform averaging unit that obtains an averaged speech waveform based on the predicted value of the immediately preceding speech waveform accumulated in the buffering unit; A speech waveform selection unit that outputs a predicted value of a speech waveform for predicting the speech waveform at the next time based on the predicted value of the breakdown flag, the predicted value of the speech waveform, and the averaged speech waveform; A speech synthesis device having the above.

6. The speech synthesis device according to claim 4, Having a total amount calculation unit that obtains a statistic of the state of the recurrent neural network based on the state of the recurrent neural network, The breakdown detection unit obtains a predicted value of the breakdown flag based on the statistic of the state of the recurrent neural network substituted from the state sequence of the recurrent neural network and the learned breakdown detection model. A speech synthesis device.

7. A speech synthesis device that generates a speech waveform in an inference phase, A waveform generation unit that obtains a predicted value of the speech waveform at the next time based on the combined speech waveform and acoustic feature amount, and the state of the recurrent neural network; A predicted value prediction unit for breakdown detection indicators that obtains a predicted value of a breakdown detection indicator based on the state of the recurrent neural network and a learned breakdown detection indicator prediction model; When the predicted value of the breakdown detection indicator is greater than a threshold value, a state initialization unit that initializes the state of the recurrent neural network based on the initial value of the state of the recurrent neural network; A speech synthesis device having the above.

8. A speech synthesis method executed by a speech synthesis device that generates a speech waveform in a learning phase, The speech synthesis device, A waveform generation process for obtaining a predicted value of a speech waveform at the next time based on the combined speech waveform and acoustic feature amount, and the state of the recurrent neural network; A breakdown flag conversion process for obtaining the breakdown flag based on the voice waveform, the predicted value of the voice waveform, and the threshold of the breakdown flag indicating whether the voice waveform is broken at each time; A breakdown detection process for obtaining a predicted value of the breakdown flag based on the state sequence of the recurrent neural network and the breakdown detection model; An error calculation process for the breakdown flag for calculating the error between the breakdown flag and the predicted value of the breakdown flag; A breakdown detection model learning process for obtaining a learned breakdown detection model based on the error and the breakdown detection model; A voice synthesis method for executing the above.

9. A voice synthesis method executed by a voice synthesis device that generates a voice waveform in an inference phase, wherein the voice synthesis device A waveform generation process for obtaining a predicted value of the voice waveform at the next time based on the combined voice waveform and acoustic feature amount, and the state of the recurrent neural network; A breakdown detection process for obtaining a predicted value of the breakdown flag indicating whether the voice waveform is broken at each time based on the state of the recurrent neural network and the learned breakdown detection model; When the predicted value of the breakdown flag indicates a breakdown, a state initialization process for initializing the state of the recurrent neural network based on the initial value of the state of the recurrent neural network; A voice synthesis method for executing the above.

10. A program for causing a computer to execute the method according to claim 8 or 9.

Citation Information

Patent Citations

  • Voice synthesizer

    JP2021032937A

  • Acoustic feature amount conversion model learning device, method and program, neural vocoder learning device, method and program, and, voice synthesis device, method and program

    JP2021067885A

  • Text-to-speech synthesis method and device using machine learning, and computer-readable storage medium

    JP2021511533A

  • Systems and methods for multi-speaker neural text-to-speech

    US20180336880A1

  • Controlling Expressivity In End-to-End Speech Synthesis Systems

    US20210035551A1