Voice processing device
By performing band segmentation and phase correction on the pulse signal after phase shift, combining the band group delay parameters and band group delay correction parameters, the problems of poor reproducibility and slow generation speed are solved, and high-quality speech synthesis is achieved.
Patent Information
- Application Number
- CN202210403587.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2015-09-16
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2035-09-16
AI Technical Summary
In the prior art, the reproducibility of the speech waveform is poor and it is difficult to generate waveforms at a high speed, especially when waveform generation is generated using group delay feature quantities.
The frequency band segmentation is performed by storing the pulse signal after phase shift, and the delay time calculation unit, the phase calculation unit, the selection unit and the channel filtering unit generate the phase shifted sound source signal, and combine the band group delay parameters and the band group delay correction parameters to improve the reproducibility of the voice waveform.
It realizes high-quality reproducibility and high-speed generation of speech waveforms, and can generate speech waveforms close to the analysis source, improving the effect of speech synthesis.
Smart Images

Figure CN114694632B_ABST
Abstract
Description
[0001] This application is a divisional application of application No. 201580082452.1, filed on September 16, 2015, and entitled “Speech Processing Device, Speech Processing Method, and Speech Processing Program.” Technical Field
[0002] Embodiments of the present invention relate to a speech (audio) processing device. Background Art
[0003] Speech analysis devices that analyze speech waveforms to extract feature parameters, and / or speech synthesis devices that synthesize speech based on the feature parameters obtained by analysis, are widely used in speech processing technologies such as text-to-speech synthesis technology, speech coding technology, and speech recognition technology.
[0004] Prior art literature
[0005] Patent Literature
[0006] Patent Document 1: International Publication No. 2014 / 021318
[0007] Patent Document 2: Japanese Patent Application Laid-Open No. 2013-164572
[0008] Non-patent literature
[0009] Non-patent document 1: Hideki Sakano, "Method of expressing efficiency of short-time phase by time domain smoothing group extension", Journal of the Society of Electronic Information and Communications Technology D-II Vol.J84-D-II, No.4, pp.621-628 Summary of the Invention
[0010] Problems to be solved by the invention
[0011] However, conventional methods have been difficult to use in statistical models, resulting in a mismatch between the reconstructed phase and the phase of the analysis source waveform. Furthermore, using group delay features for waveform generation has been problematic, preventing high-speed waveform generation. The present invention aims to provide a speech processing device, speech processing method, and storage medium that improve the reproducibility of speech waveforms.
[0012] Technical solutions to solve problems
[0013] The speech processing device of the embodiment comprises: a storage unit that stores phase-shifted band pulse signals obtained by band-splitting phase-shifted pulse signals; a delay time calculation unit that calculates the delay time of the phase-shifted band pulse signals based on band group delay parameters in a predetermined frequency range of a group delay spectrum calculated from a phase spectrum of a speech frame at each moment; a phase calculation unit that calculates the phase of a boundary frequency based on the band group delay parameters and a band group delay correction parameter generated from the band group delay parameters to correct phase information; a selection unit that selects the corresponding phase-shifted band pulse signal from the storage unit based on the calculated phase of each frequency band; a superposition unit that generates a phase-shifted sound source signal by delaying and superimposing the selected phase-shifted band pulse signals according to the delay time; and a vocal channel filtering unit that applies a vocal channel filter corresponding to the spectrum parameters calculated for each speech frame of the input speech to output a speech waveform. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a block diagram showing a configuration example of a speech analysis device according to an embodiment.
[0015] Figure 2 The figure illustrates the speech waveform and pitch mark received by the example extraction unit.
[0016] Figure 3 It is a diagram showing an example of processing by the spectrum parameter calculation unit.
[0017] Figure 4 1 and 2 are diagrams showing an example of processing by a phase spectrum calculation unit and a group delay spectrum calculation unit.
[0018] Figure 5 This is a diagram showing an example of creating a frequency scale.
[0019] Figure 6 This is a diagram illustrating the results of analysis based on band group delay parameters.
[0020] Figure 7 This is a diagram illustrating the results of analysis based on the band group delay correction parameters.
[0021] Figure 8 This is a flowchart showing the processing performed by the speech analysis device.
[0022] Figure 9 This is a flowchart showing details of the band group delay parameter calculation procedure.
[0023] Figure 10 This is a flowchart showing details of the band group delay correction parameter calculation procedure.
[0024] Figure 11This is a block diagram showing a first embodiment of a speech synthesis device.
[0025] Figure 12 This is a diagram showing a configuration example of a speech synthesis device that performs inverse Fourier transform and waveform superposition.
[0026] Figure 13 Is to express Figure 2 The figure shows an example of waveform generation corresponding to the interval shown.
[0027] Figure 14 This is a block diagram showing a second embodiment of a speech synthesis device.
[0028] Figure 15 This is a flowchart showing the processing performed by the sound source signal generation unit.
[0029] Figure 16 This is a block diagram showing the configuration of a sound source signal generating unit.
[0030] Figure 17 is a diagram of an example phase-shifted frequency-band pulse signal.
[0031] Figure 18 This is a conceptual diagram showing a selection algorithm performed by a selection unit.
[0032] Figure 19 It is a diagram showing a phase-shifted frequency-band pulse signal.
[0033] Figure 20 : is a diagram showing an example of generating a sound source signal.
[0034] Figure 21 This is a flowchart showing the processing performed by the sound source signal generation unit.
[0035] Figure 22 This figure illustrates an example of a speech waveform generated by including minimum phase correction.
[0036] Figure 23 This is a diagram showing a configuration example of a speech synthesis device using band noise intensity.
[0037] Figure 24 is a graph of noise intensity in example frequency bands.
[0038] Figure 25 This is a diagram showing a configuration example of a speech synthesis device that also uses control based on band noise intensity.
[0039] Figure 26 This is a block diagram showing a third embodiment of a speech synthesis device.
[0040] Figure 27 This is a diagram schematically showing an HMM.
[0041] Figure 28 This is a diagram schematically showing an HMM storage unit.
[0042] Figure 29 This is a diagram schematically showing an HMM learning device.
[0043] Figure 30 This is a diagram showing the processing performed by the analysis unit.
[0044] Figure 31 This is a flowchart showing the processing performed by the HMM learning unit.
[0045] Figure 32 This is a diagram showing an example of constructing an HMM sequence and distribution array. DETAILED DESCRIPTION
[0046] (First Speech Processing Device: Speech Analysis Device)
[0047] Next, a first speech processing device according to an embodiment, that is, a speech analysis device, will be described with reference to the drawings. Figure 1 1 is a block diagram showing a configuration example of the speech analysis device 100 according to the embodiment. Figure 1 As shown, the speech analysis device 100 includes an extraction unit (speech frame extraction unit) 101 , a spectrum parameter calculation unit 102 , a phase spectrum calculation unit 103 , a group delay spectrum calculation unit 104 , a band group delay parameter calculation unit 105 , and a band group delay correction parameter calculation unit 106 .
[0048] The extraction unit 101 receives the input speech and pitch mark, cuts the input speech in frames and outputs it (speech frame extraction). Figure 2 The spectrum parameter calculation unit (first calculation unit) 102 calculates the spectrum parameters based on the speech frame output by the extraction unit 101. The processing example performed by the spectrum parameter calculation unit 102 will be described later. Figure 3 Provide explanation.
[0049] The phase spectrum calculation unit (second calculation unit) 103 calculates the phase spectrum of the speech frame output by the extraction unit 101. The processing example performed by the phase spectrum calculation unit 103 will be described later. Figure 4 The group delay spectrum calculation unit (third calculation unit) 104 calculates the group delay spectrum described later based on the phase spectrum calculated by the phase spectrum calculation unit 103. The processing example performed by the group delay spectrum calculation unit 104 will be described later using Figure 4 (b) is described below.
[0050] The band group delay parameter calculation unit (fourth calculation unit) 105 calculates the band group delay parameter based on the group delay profile calculated by the group delay profile calculation unit 104. An example of the processing performed by the band group delay parameter calculation unit 105 will be described later. Figure 6 The band group delay correction parameter calculation unit (fifth calculation unit) 106 calculates a correction amount (band group delay correction parameter: correction parameter) for correcting the difference between the phase spectrum reconstructed based on the band group delay parameter calculated by the band group delay parameter calculation unit 105 and the phase spectrum calculated by the phase spectrum calculation unit 103. An example of the processing performed by the band group delay correction parameter calculation unit 106 will be described later. Figure 7 Provide explanation.
[0051] Next, the processing performed by the speech analysis device 100 will be described in further detail. Here, the processing performed by the speech analysis device 100 will be described for a case where feature parameter analysis is performed through pitch synchronization analysis.
[0052] The extraction unit 101 receives input speech and pitch label information indicating the center time of each speech frame based on its periodicity. Figure 2 exemplifies the speech waveform and pitch labels received by the extraction unit 101 . Figure 2 The waveform of the speech sound "だ" is shown, and together with the speech waveform, the pitch marker time extracted according to the periodicity of the voiced sound (voiced sound) is shown.
[0053] The following is a sample of a speech frame. Figure 2 The following is an analysis example of the interval shown below (underlined interval). Extraction unit 101 extracts speech frames by multiplying the pitch marker by a window function twice the length of the pitch. The pitch marker is obtained, for example, by extracting the pitch using a pitch extraction device and extracting the peak of the pitch period. Furthermore, silent (unvoiced) intervals without periodicity can be interpolated with pitch markers of a fixed frame rate and / or periodic interval to create a time sequence that serves as the analysis center, thus serving as the pitch marker.
[0054] A Hanning window can be used to extract speech frames. Alternatively, window functions with different characteristics, such as a Hamming window or a Blackman window, may be used. Extraction unit 101 uses the window function to extract the pitch waveform, which forms the unit waveform of a periodic interval, as a speech frame. Furthermore, extraction unit 101 also extracts speech frames in non-periodic intervals, such as silent / unvoiced intervals, by multiplying the window function at times determined by interpolating a fixed frame rate and / or a pitch marker, as described above.
[0055] Furthermore, in this embodiment, the case of using pitch synchronization analysis to extract spectrum parameters, band group delay parameters, and band group delay correction parameters is described as an example, but the present invention is not limited thereto, and parameter extraction may be performed using a fixed frame rate.
[0056] The spectral parameter calculation unit 102 obtains spectral parameters for the speech frames extracted by the extraction unit 101. For example, the spectral parameter calculation unit 102 obtains arbitrary spectral parameters representing the spectral envelope, such as Mel-frequency cepstrum, linear prediction coefficients, Mel-frequency LSP (Line Spectrum Pair), or sine wave models. Furthermore, when performing analysis based on a fixed frame rate rather than pitch-synchronous analysis, parameter extraction can also be performed using these parameters and / or spectral envelope extraction methods implemented by STRAIGHT analysis. Here, spectral parameters based on Mel-frequency LSP are used as an example.
[0057] Figure 3 2 is a diagram showing an example of processing performed by the spectrum parameter calculation unit 102 . Figure 3 (a) represents a speech frame, Figure 3 (b) shows the spectrum obtained by Fourier transform. The spectrum parameter calculation unit 102 applies Mel LSP analysis to the spectrum to obtain Mel LSP coefficients. The 0th order of the Mel LSP coefficient represents the gain term, while the 1st order and above represent the line spectrum frequency on the frequency axis. Grid lines are shown for each LSP frequency. Here, Mel LSP analysis is applied to the 44.1kHz speech. The spectrum envelope obtained in this way becomes the parameter that represents the general shape of the spectrum ( Figure 3 (c)).
[0058] Figure 4 103 and 104. FIG. 104 is a diagram showing an example of processing by the phase spectrum calculation unit 103 and an example of processing by the group delay spectrum calculation unit 104. Figure 4 (a) shows the phase spectrum obtained by the phase spectrum calculation unit 103 through Fourier transform. The phase spectrum is the unwrap spectrum. The phase spectrum calculation unit 103 applies high-pass filtering to both the amplitude and phase to obtain the phase spectrum so that the phase of the DC component is zero.
[0059] The group delay profile calculation unit 104 calculates the delay profile of the Figure 4 The phase spectrum shown in (a) is obtained by the following formula 1 Figure 4 The group delay spectrum shown in (b).
[0060] [Formula 1]
[0061]
[0062] In the above formula 1, τ(ω) represents the group delay spectrum, stands for phase spectrum, with the symbol "'" indicating differentiation. Group delay is the frequency derivative of phase, representing the average time (the center of gravity of the waveform: delay time) of each frequency band in the time domain. The group delay spectrum is equivalent to the differential value of the expanded phase, and therefore ranges from -π to π.
[0063] Here, according to Figure 4 As can be seen from (b), a group delay close to -π is generated at low frequencies. In other words, a difference close to π is generated in the phase spectrum of this frequency. In addition, according to Figure 3 In the amplitude spectrum of (b), a trough can be observed at this frequency position.
[0064] The frequencies where the phase steps occur, representing the boundary frequencies, are due to the opposite signs of the signals at low and high frequencies, which are divided by this frequency. Reproducing discontinuous changes in group delay by including the group delay around π on the frequency axis is crucial for reproducing the source speech waveform and obtaining high-quality analyzed and synthesized speech. Furthermore, the group delay parameter used in speech synthesis is required to be able to reproduce these abrupt changes in group delay.
[0065] The band group delay parameter calculation unit 105 calculates the band group delay parameter based on the group delay parameter calculated by the group delay profile calculation unit 104. The band group delay parameter is a group delay parameter for each predetermined frequency range. This reduces the order of the group delay profile, resulting in a parameter that can be used as a parameter in the statistical model. The band group delay parameter is obtained using the following equation 2.
[0066] [Formula 2]
[0067]
[0068] The band group delay based on the above equation 2 represents the average time in the time domain and represents the offset from the zero-phase waveform. When the average time is calculated from the discrete spectrum, the following equation 3 is used.
[0069] [Formula 3]
[0070]
[0071] Here, the band group delay parameter uses weighting based on the power spectrum, but it is also possible to use only the average of the group delay. In addition, a different calculation method such as weighted averaging based on the amplitude spectrum is also possible, as long as the parameter represents the group delay of each band.
[0072] Thus, the band group delay parameter becomes a parameter representing the group delay in a predetermined frequency range. Thus, as shown in the following equation 4, reconstruction of the group delay based on the band group delay parameter is performed by using the band group delay parameter corresponding to each frequency.
[0073] [Formula 4]
[0074]
[0075] The phase reconstruction based on the generated group delay is obtained by the following equation 5.
[0076] [Formula 5]
[0077]
[0078] The initial value of the phase at ω=0 is 0 due to the high-pass processing mentioned above, but in practice, the phase of the DC component can also be stored and used in advance. b This is the frequency scale used to determine the band group delay. Any frequency scale can be used, but it can be adjusted to suit auditory perception, with finer resolution for low frequencies and coarser resolution for high frequencies.
[0079] Figure 5 This is a diagram showing an example of creating a frequency scale. Figure 5 The frequency scale shown uses a Mel scale with α = 0.35 up to 5kHz, and is expressed as a uniform scale above 5kHz. To improve waveform reproducibility, the group delay parameter is set to represent low frequencies, where power is high, with finer intervals, and high frequencies with coarser intervals. This is because at high frequencies, waveform power decreases, and random phase components due to non-periodic components increase, making it difficult to obtain stable phase parameters. Furthermore, it is known that high-frequency phase has little effect on hearing.
[0080] The control of the random phase component and the component due to pulse excitation is expressed by the intensity of the noise component in each frequency band, which is the strength of the periodic and non-periodic components. When using the output of speech analysis device 100 for speech synthesis, the waveform is generated by including the band noise intensity parameter described later. This allows the phase of high frequencies, where the noise component is strong, to be represented coarsely, reducing the frequency.
[0081] Figure 6 This is an example use Figure 5 The frequency scale shown is a graph of the results obtained by analyzing the group delay parameters based on the frequency band. Figure 6 (a) shows the band group delay parameter obtained by the above formula 3. The band group delay parameter is a weighted average of the group delays of each band. However, it is found that the average group delay cannot reproduce the fluctuations that appear in the group delay spectrum.
[0082] Figure 6 (b) is an example of a phase diagram generated according to the frequency band group delay parameter. Figure 6In the example shown in (b), although the phase inclination can be roughly reproduced, the phase change close to π at low frequency and the step difference of the phase spectrum cannot be captured, and there are parts where the phase spectrum cannot be reproduced.
[0083] An example of waveform generation by inverse Fourier transforming the generated phase and the amplitude spectrum generated by Mel LSP is shown in FIG. Figure 6 (c). The generated waveform becomes: Figure 3 The waveform (a) shows a shape that differs significantly from the analysis source waveform near the center. Thus, when phase is modeled using only the band group delay parameter, the phase steps contained in the speech cannot be captured, resulting in a difference between the regenerated waveform and the analysis source waveform.
[0084] To address this problem, the speech analysis apparatus 100 uses the band group delay parameter and a band group delay correction parameter. The band group delay correction parameter corrects the phase reconstructed based on the band group delay parameter to the phase of the phase spectrum at a predetermined frequency.
[0085] The band group delay correction parameter calculation unit 106 calculates the band group delay correction parameter based on the phase spectrum and the band group delay parameter. The band group delay correction parameter corrects the phase reconstructed using the band group delay parameter to the phase value at the boundary frequency. When the difference (difference) is used as a parameter, it is calculated using the following equation 6.
[0086] [Formula 6]
[0087]
[0088] The first term on the right side of the above formula 6 is the Ω obtained by analyzing the speech. b The second term of the above equation 6 is obtained by using the group delay reconstructed using the band group delay parameter bgrd(b) and the correction parameter bgrdc(b). As shown in the following equation 7, this becomes ω=Ω in the group delay of the above equation 4. b It is expressed by adding the parameters obtained by the correction parameter bgrdc(b) to the boundary.
[0089] [Formula 7]
[0090]
[0091] The phase according to the group delay thus constructed is reconstructed using the above equation 5. In addition, the second term on the right side of the above equation 6 is obtained as follows: After reconstructing the phase to ω=Ω using the above equation 7 and the above equation 5, b After reaching -1, use the bThe phase of the following formula 8 is reconstructed from the frequency band group delay to obtain and used as the Ω b-1 The band group delay parameter and band group delay correction parameter of the frequency band up to Ω b The phase is obtained by reconstructing the frequency band group delay parameter.
[0092] [Formula 8]
[0093]
[0094] In addition, using the above formula 6, the difference between the phase of the second term on the right side and the actual phase is obtained to obtain the band group delay correction parameter. Therefore, at the frequency Ω b Reproduces the actual phase.
[0095] Figure 7 This is a diagram illustrating the results of analysis using the band group delay correction parameters. Figure 7 (a) shows the group delay spectrum obtained by the above formula 7 and reconstructed based on the band group delay parameter and the band group delay correction parameter. Figure 7 (b) shows an example of generating a phase based on the group delay spectrum. Figure 7 As shown in (b), by using the band group delay correction parameter, a phase close to the actual phase can be reconstructed. In particular, in the low-frequency part where the interval of the frequency scale is narrow, Figure 6 The portion where the phase difference is generated in the step-shaped manner in (b) is also included and reproduced.
[0096] Figure 7 (c) shows an example of synthesizing a waveform based on the phase parameters thus reconstructed. Figure 6 The waveform shape of the example shown in (c) is very different from the waveform of the analysis source, but Figure 7 In the example shown in (c), a speech waveform close to the source waveform is generated. While the correction parameter bgrdc in Equation 6 uses phase difference information, it can also be another parameter, such as the phase value at the frequency. For example, any parameter that reproduces the phase at the frequency by combining it with the band group delay parameter will suffice.
[0097] Figure 8This is a flowchart illustrating the processing performed by the speech analysis device 100. The speech analysis device 100 uses a pitch marker loop to calculate parameters corresponding to each pitch marker. First, in the speech frame extraction step, the extraction unit 101 extracts a speech frame (S801). Next, the spectral parameter calculation unit 102 calculates spectral parameters (S802), the phase spectrum calculation unit 103 calculates a phase spectrum (S803), and the group delay spectrum calculation unit 104 calculates a group delay spectrum (S804).
[0098] Next, the band group delay parameter calculation unit 105 calculates the band group delay parameter in a band group delay parameter calculation step ( S805 ). Figure 9 Yes Figure 8 Flowchart showing the details of the band group delay parameter calculation step (S805). Figure 9 As shown, the band group delay parameter calculation unit 105 sets the boundary frequency of the band by looping each band of a predetermined frequency scale (S901), and calculates the band group delay parameter (average group delay) by averaging the group delay using power spectrum weights, etc. as shown in the above formula 3 (S902).
[0099] Next, the band group delay correction parameter calculation unit 106 calculates the band group delay correction parameter ( Figure 8 :S806). Figure 10 Yes Figure 8 Flowchart showing the details of the band group delay correction parameter calculation step (S806). Figure 10 As shown, the band group delay correction parameter calculation unit 106 first sets the band boundary frequency using a loop for each band (S1001). Next, the band group delay correction parameter calculation unit 106 uses the band group delay parameter and the band group delay correction parameters for bands below the current band using Equation 7 and Equation 5 to generate the phase of the boundary frequency (S1002). The band group delay correction parameter calculation unit 106 then calculates the phase spectrum difference parameter using Equation 8 and uses the calculated result as the band group delay correction parameter (S1003).
[0100] In this way, the speech analysis device 100 performs Figure 8 ( Figure 9 、 10 ) is performed to calculate and output the spectrum parameters, band group delay parameters and band group delay correction parameters corresponding to the input speech, thereby improving the reproducibility of the speech waveform when performing speech synthesis.
[0101] (Second Speech Processing Device: Speech Synthesis Device)
[0102] Next, a second speech processing device according to the embodiment, that is, a speech synthesis device, will be described. Figure 11 1 is a block diagram showing a first embodiment of a speech synthesis device (speech synthesis device 1100). Figure 11 As shown, speech synthesis device 1100 includes an amplitude information generation unit 1101, a phase information generation unit 1102, and a speech waveform generation unit 1103. These receive a spectral parameter sequence, a band group delay parameter sequence, a band group delay correction parameter sequence, and timing information of the parameter sequences, and generate a speech waveform (synthesized speech). The parameters input to speech synthesis device 1100 are calculated by speech analysis device 100.
[0103] The amplitude information generation unit 1101 generates amplitude information based on the spectral parameters at each time point. The phase information generation unit 1102 generates phase information based on the band group delay parameter and band group delay correction parameter at each time point. The speech waveform generation unit 1103 generates a speech waveform based on the amplitude information generated by the amplitude information generation unit 1101 and the phase information generated by the phase information generation unit 1102, according to the time information of each parameter.
[0104] Figure 12 This figure shows an example configuration of a speech synthesis device 1200 that performs inverse Fourier transform and waveform superposition. Speech synthesis device 1200 is a specific example configuration of speech synthesis device 1100 and includes an amplitude spectrum calculation unit 1201, a phase spectrum calculation unit 1202, an inverse Fourier transform unit 1203, and a waveform superposition unit 1204. It generates waveforms at each time point through inverse Fourier transform and outputs synthesized speech by superimposing these waveforms.
[0105] More specifically, the amplitude spectrum calculation unit 1201 calculates the amplitude spectrum based on the spectral parameters. For example, when using Mel-LSP as a parameter, the amplitude spectrum calculation unit 1201 verifies the stability of the Mel-LSP, converts it into Mel-LPC coefficients, and calculates the amplitude spectrum based on the Mel-LPC coefficients. The phase spectrum calculation unit 1202 calculates the phase spectrum based on the band group delay parameter and the band group delay correction parameter using Equations 5 and 7 above.
[0106] The inverse Fourier transform unit 1203 performs inverse Fourier transform on the calculated amplitude spectrum and phase spectrum to generate a fundamental waveform. The waveform generated by the inverse Fourier transform unit 1203 is shown in FIG. Figure 7 (c) The waveform superposition unit 1204 performs superposition synthesis on the generated pitch waveforms based on the time information of the parameter sequence to obtain synthesized speech.
[0107] Figure 13 Is to express Figure 2 The figure shows an example of waveform generation corresponding to the interval shown. Figure 13(a) shows Figure 2 The original speech waveform is shown. Figure 13 (b) is the synthesized speech waveform output by the speech synthesis device 1100 (speech synthesis device 1200) based on the frequency band group delay parameter and the frequency band group delay correction parameter. Figure 13 As shown in (a) and (b) of FIG. 1 , the speech synthesis apparatus 1100 can generate a waveform having a shape close to the waveform of the original sound.
[0108] Figure 13 (c) shows the synthesized speech waveform when only the band group delay parameter is used as a comparative example. Figure 13 As shown in (a) and (c) of FIG. 3 , the synthesized speech waveform when only the band group delay parameter is used becomes a waveform having a shape different from that of the original speech.
[0109] In this way, the speech synthesis device 1100 (speech synthesis device 1200) can reproduce the phase characteristics of the original sound by using the frequency band group delay correction parameter in addition to the frequency band group delay parameter, and can make the analysis synthesis waveform close to the shape of the speech waveform of the analysis source, thereby generating a high-quality waveform (improving the reproducibility of the speech waveform).
[0110] Figure 14 This is a block diagram illustrating a second embodiment of a speech synthesis device (speech synthesis device 1400). Speech synthesis device 1400 includes an excitation signal generator 1401 and a vocal tract filter 1402. The excitation signal generator 1401 generates an excitation signal using a band group delay parameter sequence, a band group delay correction parameter sequence, and the timing information of the parameter sequence. The excitation signal is generated using a noise signal for silent intervals and a pulse signal for spoken intervals, without phase control or noise intensity. It has a flat spectrum and is synthesized into a speech waveform by applying the vocal tract filter.
[0111] In the speech synthesis device 1400, the sound source signal generation unit 1401 controls the phase of the pulse component using the band group delay parameter and the band group delay correction parameter. Figure 11 The phase control function of the phase information generator 1102 is implemented by the sound source signal generator 1401. That is, the speech synthesis device 1400 uses the band group delay parameter and the band group delay correction parameter for vocoder-type waveform generation to generate a waveform at high speed.
[0112] One method of controlling the phase of the sound source signal is to use inverse Fourier transform. In this case, the sound source signal generating unit 1401 performs Figure 15That is, at each time point of the characteristic parameters, the sound source signal generation unit 1401 calculates the phase spectrum from the band group delay parameter and the band group delay correction parameter using the above equations 5 and 7 (S1501), performs an inverse Fourier transform with the amplitude set to 1 (S1502), and superimposes the generated waveforms (S1503).
[0113] The vocal tract filter unit 1402 applies a filter determined according to the spectrum parameters to the generated sound source signal to generate a waveform and output a speech waveform (synthesized speech). The vocal tract filter unit 1402 has a function to control the amplitude information. Figure 11 The functions of the amplitude information generating unit 1101 shown in FIG.
[0114] When the phase control is performed as described above, the speech synthesis device 1400 can generate a waveform from the sound source signal. However, since it includes the processing of inverse Fourier transform and filtering operation, it is different from the speech synthesis device 1200 ( Figure 12 ) compared, the processing amount increases and the waveform cannot be generated at a high speed. Therefore, the sound source signal generating unit 1401 is as follows Figure 16 As shown, the structure is such that a sound source signal whose phase is controlled is generated only by processing in the time domain.
[0115] Figure 16 This is a block diagram showing the configuration of the sound source signal generation unit 1401 that generates a sound source signal whose phase is controlled only by processing in the time domain. Figure 16 The sound source signal generation unit 1401 shown prepares phase-shifted frequency-band pulse signals obtained by band-dividing a phase-shifted pulse signal, and generates a sound source waveform by delaying and superimposing the phase-shifted frequency-band pulse signals.
[0116] Specifically, the sound source signal generation unit 1401 first stores signals for each frequency band obtained by phase-shifting and band-dividing the pulse signal in the storage unit 1605. Phase-shifted band pulse signals are signals whose amplitude spectra in the corresponding frequency bands are set to 1 and whose phase spectra are set to constant values. These signals are generated by phase-shifting and band-dividing the pulse signal using the following equation 9.
[0117] [Formula 9]
[0118]
[0119] Here, the frequency band boundary Ω b Determined according to the frequency scale, the phase exist The range is quantized and quantized into P levels. When P is set to 128, 128 band pulse signals of the number of bands are produced according to the step (pitch) of 2π / 128. In this way, the phase-shifted band pulse signal is a signal obtained by dividing the phase-shifted pulse signal into bands, and is selected by the main value of the band and phase during synthesis. When the index of the phase shift of band b is set to ph(b), the phase-shifted band pulse signal produced in this way is expressed as bandpulse b ph(b) (t).
[0120] Figure 17 This is a diagram of an example phase-shifted band pulse signal. The left column shows the pulse signal after phase shifting across the entire band. The upper column shows the case of zero phase, and the lower column shows the phase The second to sixth columns represent the Figure 5 As shown in FIG. 1 , the memory unit 1605 stores the phase-shifted band pulse signals generated by the band division unit 1606 , the phase imparting unit 1607 , and the inverse Fourier transform unit 1608 in advance.
[0121] The delay time calculation unit 1601 calculates the delay time of each frequency band of the phase-shifted frequency band pulse signal based on the frequency band group delay parameter. The frequency band group delay parameter calculated using the above formula 3 represents the average delay time of the frequency band in the time domain, and is integerized using the following formula 10 to obtain the delay time delay(b). The group delay corresponding to the integer delay time is τ int (b) is sought.
[0122] [Equation 10]
[0123]
[0124] The phase calculation unit 1602 calculates the phase at the boundary frequency based on the band group delay parameter and the band group delay correction parameter, which are lower than the calculated frequency band. The phase of the boundary frequency reconstructed based on the parameters is obtained using the above equations 7 and 5. The selection unit 1603 uses the boundary frequency phase and the integer group delay bgrd int (b) The phase of the pulse signal of each frequency band is calculated. And the slope is bgrd int The y-intercept of the straight line (b) is obtained by the following formula 11.
[0125] [Formula 11]
[0126]
[0127] In addition, the selection unit 1603 obtains the phase (hereinafter referred to as "phase(b)") by performing an addition operation or a subtraction operation of 2π so that the main value of the phase obtained by the above formula 11 is in the range of (0≤phase(b)<2π), and obtains the main value of the phase as the number ph(b) of the phase obtained by quantization when producing the phase-shifted frequency band pulse signal (the following formula 12).
[0128] [Equation 12]
[0129]
[0130] Based on this ph(b), a phase-shifted band pulse signal is selected based on the band group delay parameter and the band group delay correction parameter.
[0131] Figure 18 This is a conceptual diagram showing the selection algorithm performed by the selection unit 1603. Here, an example of selecting a phase-shifted frequency-band pulse signal corresponding to a sound source signal in the frequency band of b=1 is shown. The selection unit 1603 generates a frequency band of Ω. b to Ω b+1 The sound source signal is obtained by integerizing the frequency band group delay parameter of the frequency band and obtaining the delay and phase gradient, i.e., the group delay bgrd. int (b) Then, the selection unit 1603 obtains the phase at the boundary frequency generated based on the band group delay parameter and the band group delay correction parameter. And the slope is bgrd int The phase-shifted band pulse signal is selected based on the y-axis intercept phase(b) of the straight line in (b), and ph(b) obtained by quantizing its main value 〈phase(b)〉.
[0132] Figure 19 is a diagram showing a phase-shifted frequency-band pulse signal. Figure 19 As shown in (a), the pulse signal of the entire frequency band based on phase (b) is a signal with a fixed phase (b) and an amplitude of 1. If a time delay is given to it, a fixed group delay corresponding to the delay amount will be generated. Therefore, as shown in Figure 19 As shown in (b), it becomes through phase (b) and the slope is bgrd int (b) is a straight line. Apply a bandpass filter to the linear phase signal of the entire frequency band and cut out Ω b to Ω b+1 The signal obtained in the interval becomes Figure 19 (c) of the amplitude in Ω b to Ω b+1 The boundary Ω is 1 in the interval and 0 in other frequency regions. b The phase is signal.
[0133] Therefore, using Figure 18 The method shown here allows for appropriate selection of phase-shifted pulse signals for each frequency band. The superposition unit 1604 delays the selected phase-shifted frequency band pulse signals by the delay time delay(b) calculated by the delay time calculation unit 1601 and adds them together across the entire frequency band. This generates an excitation signal reflecting the band group delay parameter and the band group delay correction parameter.
[0134] [Formula 13]
[0135]
[0136] Figure 20 : is a diagram showing an example of generating a sound source signal. Figure 20 (a) is the sound source signal of each frequency band, which is a graph showing the waveform obtained by delaying the selected phase-shift pulse signal in five low-frequency bands. The sound source signal generated by adding them up in the entire frequency band is shown in FIG. Figure 20 (b) The phase spectrum of the signal thus generated is shown in Figure 20 (c), the amplitude spectrum is shown in Figure 20 (d).
[0137] Figure 20 The phase spectrum shown in (c) shows the analysis source phase with a thin line and the phase generated using the above equations 5 and 7 with a thick line, which overlaps. In this way, the phase generated by the sound source signal generator 1401 and the phase regenerated based on the parameters essentially overlap, except for a difference due to the different high-frequency expansion, resulting in a phase close to the analysis source phase.
[0138] according to Figure 20 The amplitude spectrum shown in (d) shows that, except for the zero-crossing portion where the phase changes significantly, the spectrum has a nearly flat spectrum with an amplitude of approximately 1.0, indicating that the excitation waveform has been accurately generated. The excitation signal generation unit 1401 synthesizes the excitation signals generated in this manner by superimposing them according to the pitch markers determined by the parameter sequence time information to generate the excitation signal for the entire sentence.
[0139] Figure 21This is a flowchart illustrating the processing performed by the excitation signal generator 1401. The excitation signal generator 1401 loops through the parameter sequence at each time point. In the band pulse delay time calculation step, the delay time is calculated using Equation 10 (S2101). In the boundary frequency phase calculation step, the phase of the boundary frequency is calculated using Equations 5 and 7 (S2102). Next, in the phase-shifted band pulse selection step, the excitation signal generator 1401 selects a phase-shifted band pulse signal stored in the storage unit 1605 using Equations 11 and 12 (S2103). In the delayed phase-shifted band pulse superposition step, the selected phase-shifted band pulse signal is delayed, added, and superimposed to generate an excitation signal (S2104).
[0140] The vocal tract filter unit 1402 applies a vocal tract filter to the sound source signal generated by the sound source signal generator unit 1401 to obtain synthesized speech. In the case of Mel-LSP parameters, the vocal tract filter converts the Mel-LSP parameters into Mel-LPC parameters, performs gain extraction and other processing, and then applies the Mel-LPC filter to generate a waveform.
[0141] Because the influence of the vocal tract filter increases the minimum phase characteristic, the minimum phase correction process can also be applied when calculating the band group delay parameter and the band group delay correction parameter based on the phase of the analysis source. The minimum phase is generated as the imaginary axis: the imaginary axis is obtained by generating an amplitude spectrum using Mel-LSP, performing an inverse Fourier transform on the spectrum based on the logarithmic amplitude spectrum and zero phase, and then Fourier transforming the resulting cepstrum again, doubling the positive components and zeroing the negative components.
[0142] The phase thus obtained is expanded and subtracted from the phase obtained by analyzing the waveform, thereby performing minimum phase correction. Based on the phase spectrum after minimum phase correction, the band group delay parameter and the band group delay correction parameter are obtained. The sound source is generated through the processing of the sound source signal generator 1401 described above, and a filter is applied to obtain synthesized speech that reproduces the phase of the source waveform.
[0143] Figure 22 This figure illustrates an example of a speech waveform generated by including minimum phase correction. Figure 22 (a) is related to Figure 13 (a) Speech waveform of the same analysis source. Figure 22 (b) is an analysis-synthesized waveform based on vocoder-style waveform generation performed by the speech synthesis device 1400. Figure 22 (c) is a vocoder based on a widely used pulse sound source, and in this case, it becomes a minimum phase waveform shape.
[0144] Figure 22The analysis and synthesis waveform obtained by the speech synthesis device 1400 shown in (b) reproduces a waveform close to Figure 22 The waveform of the original sound shown in (a) is also close to Figure 13 The speech waveform of the waveform shown in (b) is shown in FIG. Figure 22 The minimum phase shown in (c) results in a speech waveform in which power is concentrated near the pitch mark, and the shape of the original speech waveform cannot be reproduced.
[0145] In order to compare the processing volume, the processing time for generating a speech waveform of about 30 seconds was measured. Figure 12 The composition is about 9.19 seconds, in the vocoder style Figure 14 In the case of the configuration, the processing time is approximately 0.47 seconds (measured by a computing server with a 2.9GHz CPU). In other words, it was confirmed that the processing time was shortened by approximately 5.1%. In other words, waveform generation using the vocoder method can generate waveforms at high speed.
[0146] This is because it is possible to generate a waveform that reflects the phase characteristics by operating only in the time domain without using the inverse Fourier transform. In the above waveform generation, the sound source is generated, the sound source waveform is superimposed and synthesized, and then the filter is applied, but it is not limited to this. It can also be a different structure such as generating a sound source waveform for each fundamental waveform, applying a filter, generating a fundamental waveform, and superimposing and synthesizing the generated fundamental waveforms. Moreover, as long as the Figure 16 The excitation signal generating unit 1401 based on the phase-shifted band pulse signal may generate the excitation signal based on the band group delay parameter and the band group delay correction parameter.
[0147] Figure 23 It shows the Figure 12 The illustrated diagram shows an example configuration of a speech synthesis device 2300, which incorporates control of separation of noise and periodic components using band noise intensity. Speech synthesis device 2300 is one of the specific components of speech synthesis device 1100. The amplitude spectrum calculation unit 1201 calculates the amplitude spectrum based on the spectral parameter sequence, while the periodic component spectrum calculation unit 2301 and the noise component spectrum calculation unit 2302 separate the periodic component spectrum and the noise component spectrum according to the band noise intensity. Band noise intensity is a parameter that indicates the ratio of noise components in each frequency band of the spectrum. For example, this can be obtained by using a PSHF (Pitch Scaled Harmonic Filter) method to separate speech into periodic and noise components, calculating the noise component ratio for each frequency band, and averaging the ratio for each predetermined frequency band.
[0148] Figure 24 is a graph of noise intensity in example frequency bands. Figure 24 (a) is the signal obtained by separating the speech into periodic components and non-periodic components using PSHF, and obtaining the spectrum of the speech and the spectrum of the non-periodic components of the processing target frame, and obtaining the ratio of the non-periodic components of each frequency, ap(ω). During processing, post-processing such as setting the frequency band with sound to 0 and / or clipping the ratio between 0 and 1 is added to the ratio based on PSHF. The intensity obtained by weighted average of the spectrum scaled according to the noise component ratio thus obtained is Figure 24 The frequency scale is similar to the frequency band group delay, using Figure 5 The scale shown is obtained by the following formula 14.
[0149] [Equation 14]
[0150]
[0151] The noise component spectrum calculation unit 2302 multiplies the noise intensity of each frequency based on the band noise intensity by the spectrum generated based on the spectrum parameters to obtain the noise component spectrum. The periodic component spectrum calculation unit 2301 multiplies by 1.0-bap(b) to obtain the periodic component spectrum from which the noise component spectrum has been removed.
[0152] The noise component waveform generator 2304 generates a noise component waveform by performing an inverse Fourier transform based on the random phase generated from the noise signal and the amplitude spectrum based on the noise component spectrum. The noise component phase can be generated, for example, by generating Gaussian noise with a mean of 0 and a variance of 1, slicing it with a Hanning window of twice the length of the fundamental frequency, and performing a Fourier transform on the slicing windowed Gaussian noise.
[0153] The periodic component waveform generator 2303 generates a periodic component waveform by performing inverse Fourier transform on the phase spectrum calculated by the phase spectrum calculator 1202 based on the band group delay parameter and the band group delay correction parameter and the amplitude spectrum based on the periodic component spectrum.
[0154] The waveform superposition unit 1204 adds the generated noise component waveform and periodic component waveform, and superimposes them according to the timing information of the parameter sequence to obtain synthesized speech.
[0155] In this way, by separating the noise component and the periodic component, the random phase component that is difficult to express as a band group delay parameter can be separated, and the noise component can be generated according to the random phase. In this way, the noise component contained in the silent interval and / or the high-frequency part of the sound friction sound and the sound can be suppressed to become a pulsed sound quality with a screaming feeling. In particular, when the parameters are statistically modeled, if the band group delay and the band group delay correction parameter obtained based on multiple random phase components are averaged, there is a tendency for the average value to be close to 0 and close to the pulsed phase component. By using the band noise intensity together with the band group delay parameter and the band group delay correction parameter, the noise component can be generated according to the random phase, and the periodic component can use the appropriately generated phase, so the sound quality of the synthesized speech is improved.
[0156] Figure 25 This is a diagram showing a configuration example of a vocoder-type speech synthesis device 2500 for realizing high-speed waveform generation, which also uses control based on the intensity of the band noise. The sound source generation of the noise component is performed using a fixed-length band noise signal obtained by pre-band division contained in the band noise signal storage unit 2503. In the speech synthesis device 2500, the band noise signal storage unit 2503 stores the band noise signal, and the noise sound source signal generation unit 2502 controls the amplitude of the band noise signal of each band according to the intensity of the band noise, and adds the amplitude-controlled band noise signals to generate a noise sound source signal. In addition, the speech synthesis device 2500 is Figure 14 A variation of the speech synthesis device 1400 is shown.
[0157] The pulse sound source signal generating unit 2501 generates the phase-shifted frequency band pulse signal stored in the storage unit 1605. Figure 16 The structure shown is a phase-controlled excitation signal. When a delayed phase-shifted band pulse waveform is superimposed, the amplitude of the signal in each band is controlled using the band noise intensity, resulting in an intensity of (1.0 - bap(b)). The speech synthesis device 2500 adds the pulse excitation signal generated in this manner to the noise excitation signal to generate an excitation signal. The vocal tract filter unit 1402 applies a vocal tract filter based on the spectral parameters to obtain synthesized speech.
[0158] Speech synthesis device 2500 and Figure 23Similarly, the illustrated speech synthesis device 2300 generates a noise signal and a periodic signal separately, suppresses the generation of pulse-like noise in the noise component, and generates a sound source by adding the phase-controlled periodic component to the noise component. This allows for speech synthesis with a shape close to that of the analysis source waveform. Furthermore, the speech synthesis device 2500 can generate both the noise sound source and the pulse sound source through processing solely in the time domain, enabling high-speed waveform generation.
[0159] Thus, the first and second embodiments of the speech synthesis device, by using the band group delay parameters and the band group delay correction parameters, can improve the similarity between the reconstructed phase and the phase obtained by analyzing the waveform using the feature parameters with reduced dimensions that can be statistically modeled, and can perform speech synthesis with appropriate phase control based on these parameters. By using the band group delay parameters and the band group delay correction parameters, each speech processing device involved in the embodiment can improve the reproducibility of the waveform and generate the waveform at high speed. Furthermore, in a vocoder-type speech synthesis device, a sound source waveform that has been phase-controlled only through time-domain processing is generated, and waveform generation based on the vocal tract filter can be performed, thereby enabling high-speed generation of a phase-controlled waveform. In addition, by combining the use of the band noise intensity parameters, the speech synthesis device can also improve the reproducibility of the noise component, thereby performing higher-quality speech synthesis.
[0160] Figure 26 This is a block diagram illustrating a third embodiment of a speech synthesis device (speech synthesis device 2600). Speech synthesis device 2600 is obtained by applying the aforementioned band group delay parameters and band group delay correction parameters to a text-to-speech synthesis device. Here, as a text-to-speech synthesis method, in speech synthesis based on an HMM (Hidden Markov Model), a statistical model-based speech synthesis technology, the band group delay parameters and band group delay correction parameters are used as characteristic parameters.
[0161] The speech synthesis device 2600 includes a text analysis unit 2601, an HMM sequence generator 2602, a parameter generator 2603, a waveform generator 2604, and an HMM storage unit 2605. The HMM storage unit (statistical model storage unit) 2605 stores HMMs learned from acoustic feature parameters including a band group delay parameter and a band group delay correction parameter.
[0162] The text parser 2601 parses the input text and obtains information such as pronunciation and stress, thereby generating context information. The HMM sequence generator 2602 generates an HMM sequence corresponding to the input text based on the context information generated from the text and the HMM models stored in the HMM storage 2605. The parameter generator 2603 generates acoustic feature parameters from the HMM sequence. The waveform generator 2604 generates a speech waveform based on the generated feature parameter sequence.
[0163] More specifically, the text analysis unit 2601 generates context information by analyzing the language (speech) of the input text. The text analysis unit 2601 performs morphological analysis on the input text to obtain language (speech) information required for speech synthesis, such as pronunciation information and stress information. The context information is generated based on the obtained pronunciation and language information. Alternatively, the context information can be generated based on separately generated, corrected pronunciation and stress information corresponding to the input text. Context information refers to information used as a unit for classifying speech, such as phonemes, semiphones, or syllable HMMs.
[0164] When phonemes are used as speech units, a sequence of phoneme names can be used as context information, and further, triphones with preceding and succeeding phonemes attached, and / or phoneme information containing two phonemes before and after, phoneme category information representing attributes of phoneme categories based on sound / no sound classification and / or further details, the position of each phoneme within a sentence, within a respiratory group (ventilation unit), within a stressed phrase, the number of moras / stress types in a stressed phrase, the position of moras, the position up to the stress nucleus, information on the presence or absence of a final rising tone, assigned symbol information, and other language (speech) attribute information can be included as context information.
[0165] The HMM sequence generator 2602 generates an HMM sequence corresponding to the input context information based on the HMM information stored in the HMM storage unit 2605. The HMM is a statistical model represented by the state transition probability and the output distribution of each state. When the left-to-right type HMM is used as the HMM, Figure 27 As shown, according to the output distribution N(o|μ i ,Σ i ) and state transition probability a ij (i, j are state indices) modeled as the transition probability to the adjacent state and the value of the own transition probability. ij And using the duration distribution N(d|μ i d ,Σ id ) is called HSMM (Hidden Semi-Markov Model) and is used for long-term modeling.
[0166] The HMM storage unit 2605 stores a model obtained by performing decision tree clustering on the output distribution of each state of the HMM. Figure 28 As shown, the HMM storage unit 2605 stores the decision tree of the model of the characteristic parameters of each state of the HMM and the output distribution of each leaf node of the decision tree, and further stores the decision tree and distribution for continuous long distribution. At each node of the decision tree, there is associated with a question for classifying the distribution, for example, it is classified into questions such as "Is there no sound?", "Is there a sound?", "Is it an accent core?", child nodes when they are consistent with the question, and child nodes when they are not consistent. For the input context information, it is determined whether it is consistent with the question of each node, thereby searching the decision tree and obtaining the leaf node. By using the distribution associated with the obtained leaf node as the output distribution of each state, the HMM corresponding to each speech unit is constructed. Thus, an HMM sequence corresponding to the input context information is produced.
[0167] The HMM stored in the HMM storage unit 2605 is composed of Figure 29 The HMM learning device 2900 shown is used for the above-mentioned operation. The speech corpus storage unit 2901 stores a speech corpus including speech data and context information used for creating an HMM model.
[0168] The analysis unit 2902 analyzes the speech data used for learning and obtains acoustic feature parameters. Here, the speech analysis device 100 described above is used to obtain the band group delay parameters and band group delay correction parameters, which are used together with the spectrum parameters, pitch parameters, and band noise intensity parameters.
[0169] like Figure 30 As shown, the analysis unit 2902 obtains acoustic feature parameters for each speech frame of the speech data. When using pitch-synchronous analysis, the speech frame becomes the parameters at each pitch marker time. When using a fixed frame rate, the feature parameters are extracted by interpolating the acoustic feature parameters of adjacent pitch markers.
[0170] use Figure 1 The speech analysis device 100 shown in FIG. 1 is used to analyze the center time of speech ( Figure 30 The acoustic feature parameters corresponding to the pitch marker position are analyzed to extract the spectrum parameters (Mel LSP), pitch parameters (logarithmic F0), band noise intensity parameters (BAP), band group delay parameters and band group delay correction parameters (BGRD and BGRDC). Furthermore, as the dynamic feature quantities of these parameters, the Δ parameter and Δ 2Parameters are all used as sound feature parameters at each moment.
[0171] The HMM learning unit 2903 learns the HMM based on the characteristic parameters obtained in this way. Figure 31 This is a flowchart showing the processing performed by the HMM learning unit 2903. The HMM learning unit 2903 initializes the phoneme HMM (S3101) and performs maximum likelihood estimation on the phoneme HMM through HSMM learning (S3102), thereby learning the phoneme HMM as the initial model. During maximum likelihood estimation, learning is performed through concatenative learning, which involves probabilistically associating each state with the feature parameters based on the HMM of the entire sentence, which is concatenated to associate the HMM with the sentence, and the acoustic feature parameters corresponding to the sentence.
[0172] Next, the HMM learning unit 2903 initializes the context-dependent HMM using the phoneme HMM (S3103). As described above, the context includes the relevant phonemes, the surrounding phoneme environment, position information within a sentence / stress phrase, the type of stress, the phonological environment such as whether there is a final rising tone, and language information. A model initialized with the relevant phonemes is prepared for the context present in the learning data.
[0173] Then, the HMM learning unit 2903 applies the maximum likelihood estimation based on connection learning to learn the context-dependent HMM (S3104), and applies the state clustering based on the decision tree (S3105). Thus, the HMM learning unit 2903 constructs a decision tree for each state / flow and state continuous long distribution of the HMM. Moreover, the HMM learning unit 2903 learns the rules for classifying the model based on the distribution of each state / flow using the maximum likelihood criterion and / or the MDL (Minimum Description Length) criterion, and constructs Figure 28 Furthermore, during speech synthesis, even when an unknown context not present in the learning data is input, the distribution of each state can be selected along the decision tree to construct the corresponding HMM.
[0174] Finally, the HMM learning unit 2903 performs maximum likelihood estimation on the model after context-dependent clustering to complete model learning (S3106). During clustering, a decision tree is constructed for each stream of each feature value, and together with the spectral parameters (Mel LSP), pitch parameters (logarithmic fundamental frequency), and band noise intensity (BAP), a decision tree for each stream of band group delay and band group delay correction parameters is also constructed. In addition, by constructing a decision tree for the multidimensional distribution of the duration of each state, a duration distribution decision tree based on the HMM is constructed. These obtained HMMs and decision trees are stored in the HMM storage unit 2605.
[0175] HMM sequence generation unit 2602 ( Figure 26 ) An HMM sequence is created based on the input context and the HMM stored in the HMM storage unit 2605. The distribution of each state is repeated according to the number of frames determined by the persistence length distribution, thereby creating a distribution sequence. The generated distribution sequence is a sequence in which the distribution of the number of parameters to be output is arranged.
[0176] The parameter generation unit 2603 generates each parameter using a parameter generation algorithm that takes static and dynamic feature quantities into consideration and is widely used in HMM-based speech synthesis, thereby generating a smooth parameter sequence.
[0177] Figure 32 This is a diagram showing an example of constructing an HMM sequence / distribution column. First, the HMM sequence generator 2602 selects the distribution of each state / stream of the HMM of the input context and the duration distribution to form an HMM sequence. When "preceding phoneme_the phoneme_subsequent phoneme_phoneme position_number of phonemes_moon position_number of moons_stress type" is used as the context to synthesize "aka", since it is a two-moon type 1, the initial phoneme of "a" has a context of "sil_a_k_1_3_1_2_1" because the preceding phoneme is "sil", the phoneme is "a", the subsequent phoneme is "k", the phoneme position is 1, the number of phonemes is 3, the moon position is 1, the number of moons is 2, and the stress type is type 1.
[0178] When following the HMM decision tree, questions such as "Is the phoneme a" or "Is the stress type type 1" are determined at each intermediate node. Leaf node distributions are selected along these questions. The distributions of the Mel-LSP, BAP, BGRD, BGRDC, and LogF0 streams, as well as the distribution of the persistence length, are selected as the HMM states, forming an HMM sequence. This constructs an HMM sequence and distribution sequence for each model unit (e.g., phoneme), and these are arranged across the entire sentence to create a distribution sequence corresponding to the input text.
[0179] The parameter generation unit 2603 generates a parameter sequence based on the generated distribution sequence using a parameter generation algorithm using static / dynamic feature quantities. 2 In the case of dynamic characteristic parameters, the output parameters are obtained by the following method. Characteristic parameter o at time t t Using static characteristic parameter c t and the dynamic characteristic parameter Δc determined based on the characteristic parameters of the previous and next frames t , Δ 2 c t , represented by o t =(c t ′、Δc t ′、Δ2 c t ′). The static feature quantity c that maximizes P(O|J,λ) is t The vector C = (c0′,…,c T-1 ')' is obtained by setting OTM to a T×M-dimensional zero vector and solving the following equation 15.
[0180] [Equation 15]
[0181]
[0182] Where T is the number of frames and J is the state transition sequence. If the relationship between the feature parameter O and the static feature parameter C is related based on the matrix W for calculating the dynamic feature, it can be expressed as O=WC. O becomes a vector of 3TM, C becomes a vector of TM, and W is a matrix of 3TM×TM. Moreover, when μ=(μ s00 ′,…,μ sJ-1Q-1 ′)′、Σ=diag(Σ s00 ′,…,Σ sJ-1Q-1 When ′)′ is set to the mean vector and covariance matrix of the distribution corresponding to the sentence in which the mean vector and diagonal covariance of the output distribution at each moment are arranged, the above formula 15 is solved by solving the equation of the following formula 16 to obtain the optimal feature parameter sequence C.
[0183] [Equation 16]
[0184] W′∑ -1 WC=W′∑ -1 μ …(16)
[0185] This equation is obtained using a method based on Cholesky decomposition. Similarly to the solution used in the time-update algorithm of the RLS filter, a parameter sequence can be generated in chronological order according to the delay time, and generation can also be performed with low latency. Furthermore, the parameter generation process is not limited to the above method; any method for generating characteristic parameters from other distribution sequences, such as interpolation of mean vectors, can also be used.
[0186] The waveform generator 2604 generates a speech waveform based on the parameter sequence thus generated. For example, the waveform generator 2604 synthesizes speech based on the Mel LSP sequence, the logarithmic F0 sequence, the band noise intensity sequence, the band group delay parameter, and the band group delay correction parameter. When these parameters are used, the waveform is generated using the speech synthesis device 1100 or the speech synthesis device 1400 described above. Specifically, Figure 23 The structure based on the inverse Fourier transform shown, or Figure 25 The waveform generation is done by using the high-speed waveform generation of the vocoder shown in the figure. When the band noise intensity is not used, the Figure 12 The speech synthesis device 1200 based on inverse Fourier transform shown, or Figure 14 The speech synthesis device 1400 is shown.
[0187] Through these processes, synthesized speech corresponding to the input context can be obtained, and speech close to the analyzed original speech can be synthesized by using the band group delay parameter and the band group delay correction parameter so that the phase information of the speech waveform is also reflected.
[0188] Furthermore, while the HMM learning unit 2903 described above describes a configuration for performing maximum likelihood estimation of a speaker-dependent model using a corpus of a specific speaker, the present invention is not limited thereto. Alternatively, different configurations may be employed, such as speaker adaptation techniques, model interpolation techniques, or other cluster adaptation learning techniques used to enhance the diversity of HMM speech synthesis. Furthermore, different learning methods, such as distribution parameter estimation using deep neural networks, may also be employed.
[0189] Furthermore, the speech synthesis device 2600 may be configured to further include a feature parameter sequence selection unit between the HMM sequence generation unit 2602 and the parameter generation unit 2603 for selecting a feature parameter sequence. The unit selects feature parameters from the acoustic feature parameters obtained by the analysis unit 2902 using the HMM sequence as a target, and synthesizes a speech waveform based on the selected parameters. In this manner, by selecting acoustic feature parameters, it is possible to suppress sound quality degradation caused by excessive smoothing in HMM speech synthesis, thereby obtaining natural synthesized speech that is closer to actual utterances.
[0190] In this way, by using the band group delay parameter and the band group delay correction parameter as characteristic parameters for speech synthesis, not only can the reproducibility of the waveform be improved, but also the waveform can be generated at high speed.
[0191] Furthermore, the speech synthesis devices described above, such as the speech analysis device 100 and speech synthesis device 1100, can also be implemented using a general-purpose computer as basic hardware. Specifically, the speech analysis device and each speech synthesis device in this embodiment can be implemented by having a processor mounted on a computer execute a program. In this case, implementation can be achieved by pre-installing the program on the computer, storing it on a storage medium such as a CD-ROM, or distributing the program over a network and installing it on the computer as appropriate. Furthermore, implementation can be achieved by utilizing, as appropriate, a memory built into or external to the computer, a hard disk, or a storage medium such as a CD-R, CD-RW, DVD-RAM, or DVD-R. Furthermore, some or all of the speech synthesis devices, such as the speech analysis device 100 and speech synthesis device 1100, can be implemented using either hardware or software.
[0192] In addition, although several embodiments of the present invention have been described in various combinations, these embodiments are provided as examples and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and changes can be made without departing from the scope of the invention. These embodiments and / or their variations are included in the scope and / or spirit of the invention, and are included in the invention described in the claims and their equivalents.
Claims
1. A speech processing device comprising: a storage unit for storing a phase-shifted frequency-band pulse signal obtained by frequency-band-dividing the phase-shifted pulse signal; a delay time calculation unit for calculating a delay time of the phase-shifted band pulse signal based on a band group delay parameter in a predetermined frequency range of a group delay spectrum calculated from a phase spectrum of a speech frame at each time point; a phase calculation unit for calculating a phase of a boundary frequency based on the band group delay parameter and a band group delay correction parameter generated from the band group delay parameter and used to correct phase information; a selection unit that selects a corresponding phase-shifted frequency band pulse signal from the storage unit based on the calculated phase of each frequency band; a superposition unit for delaying the selected phase-shift frequency band pulse signals by the delay time and superimposing the signals to generate a phase-shifted sound source signal; as well as The vocal tract filter unit applies a vocal tract filter corresponding to the spectrum parameters calculated for each speech frame of the input speech, and outputs a speech waveform.
2. The speech processing device according to claim 1, the storage unit, storing a phase-shifted band pulse signal, wherein the phase-shifted band pulse signal is a band pulse signal of each phase obtained by quantizing a main value of the phase into predetermined levels, The selection unit, In each frequency range of the band group delay parameter, the phase of the start frequency of the band is calculated based on the band group delay parameter and the band group delay correction parameter, a delay amount obtained by integerizing the band group delay parameter is calculated, a group delay is calculated based on the delay amount, a phase value at the frequency origin of a straight line passing through the phase of the start frequency is calculated using the group delay calculated based on the delay amount as a slope, and a phase-shifted band pulse signal corresponding to a main value of the calculated phase value is selected. The superposition portion, The phase-shifted frequency band pulse signal delayed according to the delay amount is superimposed.
3. The speech processing device according to claim 1, The device further comprises a band noise signal storage unit for storing the band noise signal obtained by performing the band division. The channel filter unit, A vocal channel filter corresponding to a spectral parameter is applied to a mixed sound source signal, wherein the mixed sound source signal is a signal obtained by mixing the noise signal of each frequency band and the phase-shifted frequency band pulse signal. The noise signal of each frequency band is a signal generated according to the intensity of each frequency band based on a band noise intensity parameter representing the ratio of noise components in a predetermined frequency range, and the noise signal of each frequency band.
4. A speech processing device comprising: a statistical model storage unit storing a statistical model obtained by learning using spectrum parameters calculated for each speech frame of input speech, band group delay parameters in a predetermined frequency range of a group delay spectrum calculated based on a phase spectrum of each speech frame, and band group delay correction parameters for correcting the phase spectrum generated based on the band group delay parameters; a parameter generating unit configured to generate a spectrum parameter, a band group delay parameter, and a band group delay correction parameter corresponding to an arbitrary input text based on context information corresponding to the arbitrary input text and the statistical model stored in the statistical model storing unit; as well as A waveform generating unit generates a waveform based on the spectrum parameter, the band group delay parameter, and the band group delay correction parameter generated by the parameter generating unit.
Citation Information
Patent Citations
Voice feature quantity extraction device, voice feature quantity extraction method, and voice feature quantity extraction program
JP2013164572A
Spectral envelope and group delay inference system and voice signal synthesis system for voice analysis / synthesis
WO2014021318A1
Speech synthesis apparatus, speech synthesis method, speech synthesis program product, and learning apparatus
US20130262087A1
Digital signal sub-band separating / combining apparatus achieving band-separation and band-combining filtering processing with reduced amount of group delay
US6856653B1