Voice processing device, voice processing method and storage medium

By generating amplitude and phase information through a speech analysis device and utilizing band group delay parameters and band group delay correction parameters, the problem of poor speech waveform reproducibility is solved, and high-speed, high-quality speech waveform generation is achieved.

CN114464208BActive Publication Date: 2025-11-14KK TOSHIBA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210141126.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2015-09-16
Publication Date
2025-11-14
Estimated Expiration
2035-09-16

AI Technical Summary

Technical Problem

In existing technologies, speech waveforms have poor reproducibility and slow waveform generation speed, especially when using group delay features, it is difficult to generate high-quality speech waveforms at high speed.

Method used

Amplitude and phase information are generated by a speech analysis device. The phase information is corrected by using the band group delay parameter and the band group delay correction parameter to generate a high-quality speech waveform.

Benefits of technology

It improves the reproducibility of speech waveforms and can generate high-quality speech waveforms that are close to the original sound under high-speed conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114464208B_ABST
    Figure CN114464208B_ABST
Patent Text Reader

Abstract

This not only improves waveform reproducibility but also enables high-speed waveform generation. The speech processing apparatus of this embodiment includes: an amplitude information generation unit that generates amplitude information based on a sequence of spectral parameters calculated for each speech frame of the input speech; a phase information generation unit that generates phase information based on a sequence of band group delay parameters within a predetermined frequency range of the group delay spectrum calculated from the phase spectrum of each speech frame, and a sequence of band group delay correction parameters that corrects the difference between the phase spectrum generated from the band group delay parameter sequence and the phase spectrum of each speech frame; and a speech waveform generation unit that generates a speech waveform at each time determined by time information of a parameter sequence, which serves as time information for each parameter, based on the amplitude information and the phase information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of application No. 201580082452.1, filed on September 16, 2015, entitled "Speech Processing Apparatus, Speech Processing Method and Speech Processing Program". Technical Field

[0002] The embodiments of the present invention relate to a speech (sound) processing apparatus, a speech processing method, and a storage medium. Background Technology

[0003] Speech analysis devices that analyze speech waveforms to extract feature parameters, and / or speech synthesis devices that synthesize speech based on the feature parameters obtained from the analysis, are widely used in speech processing technologies such as text-to-speech synthesis, speech coding, and speech recognition.

[0004] Existing technical documents

[0005] Patent documents

[0006] Patent Document 1: International Publication No. 2014 / 021318

[0007] Patent Document 2: Japanese Patent Application Publication No. 2013-164572

[0008] Non-patent literature

[0009] Non-patent document 1: Hideki Sakano, "Method of expressing efficiency of short-time phase by time domain smoothing group extension", Journal of the Society of Electronic Information and Communications Technology D-II Vol.J84-D-II, No.4, pp.621-628 Summary of the Invention

[0010] The problem that the invention aims to solve

[0011] However, conventional methods have suffered from difficulties in utilizing statistical models and from deviations between the reconstructed phase and the phase of the source waveform. Furthermore, conventional methods have also encountered the problem of not being able to generate waveforms at high speed when using group delay features. The problems to be solved by this invention are to provide a speech processing apparatus, speech processing method, and storage medium that improve the reproducibility of speech waveforms.

[0012] Technical solutions for solving the problem

[0013] The speech processing apparatus of the embodiment includes: an amplitude information generation unit that generates amplitude information based on a sequence of spectral parameters calculated for each speech frame of the input speech; a phase information generation unit that generates phase information based on a sequence of band group delay parameters in a predetermined frequency range of a group delay spectrum calculated from the phase spectrum of each speech frame, and a sequence of band group delay correction parameters that corrects the difference between the phase spectrum generated from the band group delay parameter sequence and the phase spectrum of each speech frame; and a speech waveform generation unit that generates a speech waveform based on the amplitude information and the phase information at each time determined by the parameter sequence time information, which is time information of each parameter. Attached Figure Description

[0014] Figure 1 This is a block diagram illustrating an example configuration of a voice analysis device according to an embodiment.

[0015] Figure 2 This is a diagram of the speech waveform and pitch mark received by the example extraction unit.

[0016] Figure 3 This is a diagram illustrating an example of the processing of the spectral parameter calculation unit.

[0017] Figure 4 This is a diagram showing the processing examples of the phase spectrum calculation unit and the group delay spectrum calculation unit.

[0018] Figure 5 This is a diagram illustrating an example of how a frequency scale is constructed.

[0019] Figure 6 This is a graph showing the results of an example analysis based on the band group delay parameter.

[0020] Figure 7 This is a graph showing the results of an analysis based on the band group delay correction parameters.

[0021] Figure 8 This is a flowchart illustrating the processing performed by the speech analysis device.

[0022] Figure 9 This is a flowchart detailing the steps involved in calculating the band group delay parameter.

[0023] Figure 10 This is a flowchart detailing the steps involved in calculating the band group delay correction parameters.

[0024] Figure 11 This is a block diagram illustrating a first embodiment of a speech synthesis device.

[0025] Figure 12 This is a diagram illustrating an example of the configuration of a speech synthesis device that performs inverse Fourier transform and waveform superposition.

[0026] Figure 13 It means and Figure 2 The diagram shows an example of waveform generation for the interval shown.

[0027] Figure 14 This is a block diagram illustrating a second embodiment of a speech synthesis device.

[0028] Figure 15 This is a flowchart illustrating the processing performed by the sound source signal generation unit.

[0029] Figure 16 This is a block diagram showing the structure of the sound source signal generation unit.

[0030] Figure 17 This is a diagram of an example phase-shifted frequency band pulse signal.

[0031] Figure 18 This is a conceptual diagram representing the selection algorithm performed by the selection department.

[0032] Figure 19 This is a diagram representing a phase-shifted frequency band pulse signal.

[0033] Figure 20 This is a diagram illustrating an example of how a sound source signal is generated.

[0034] Figure 21 This is a flowchart illustrating the processing performed by the sound source signal generation unit.

[0035] Figure 22 This is a diagram of the speech waveform generated by including minimum phase correction as an example.

[0036] Figure 23 This is a diagram illustrating an example of the configuration of a speech synthesis device that utilizes frequency band noise intensity.

[0037] Figure 24 This is a graph showing the noise intensity of an example frequency band.

[0038] Figure 25 This is a diagram illustrating an example of the configuration of a speech synthesis device that also uses control based on the intensity of frequency band noise.

[0039] Figure 26 This is a block diagram illustrating a third embodiment of the speech synthesis device.

[0040] Figure 27 It is a diagram that represents a general outline of an HMM.

[0041] Figure 28 This is a diagram that shows a general outline of the HMM storage section.

[0042] Figure 29 This is a diagram showing a schematic representation of an HMM learning device.

[0043] Figure 30 This is a diagram showing the processing performed by the analysis department.

[0044] Figure 31 This is a flowchart representing the processes performed by the HMM learning department.

[0045] Figure 32 This is a diagram representing an example of the construction of HMM sequences and distributions. Detailed Implementation

[0046] (First speech processing device: speech analysis device)

[0047] Next, with reference to the accompanying drawings, the first speech processing apparatus, namely the speech analysis apparatus, according to the embodiment will be described. Figure 1 This is a block diagram illustrating an example configuration of the voice analysis device 100 according to an embodiment. For example... Figure 1 As shown, the speech analysis device 100 includes an extraction unit (speech frame extraction unit) 101, a spectrum parameter calculation unit 102, a phase spectrum calculation unit 103, a group delay spectrum calculation unit 104, a frequency band group delay parameter calculation unit 105, and a frequency band group delay correction parameter calculation unit 106.

[0048] The extraction unit 101 receives the input speech and pitch markers, and segments and outputs the input speech in frames (speech frame extraction). Examples of the processing performed by the extraction unit 101 will be discussed later. Figure 2 The spectrum parameter calculation unit (first calculation unit) 102 calculates the spectrum parameters based on the speech frames output by the extraction unit 101. An example of the processing performed by the spectrum parameter calculation unit 102 will be used later. Figure 3 Please provide an explanation.

[0049] The phase spectrum calculation unit (second calculation unit) 103 calculates the phase spectrum of the speech frame output by the extraction unit 101. An example of the processing performed by the phase spectrum calculation unit 103 will be used later. Figure 4 (a) will be explained below. The group delay spectrum calculation unit (third calculation unit) 104 calculates the group delay spectrum, which will be described later, based on the phase spectrum calculated by the phase spectrum calculation unit 103. An example of the processing performed by the group delay spectrum calculation unit 104 will be used later. Figure 4 (b) will be explained.

[0050] The band group delay parameter calculation unit (4th calculation unit) 105 calculates the band group delay parameter based on the group delay spectrum calculated by the group delay spectrum calculation unit 104. An example of the processing performed by the band group delay parameter calculation unit 105 will be used later. Figure 6The band group delay correction parameter calculation unit (5th calculation unit) 106 calculates a correction amount (band group delay correction parameter: correction parameter) to correct the difference between the phase spectrum reconstructed from the band group delay parameter calculated by the band group delay parameter calculation unit 105 and the phase spectrum calculated by the phase spectrum calculation unit 103. An example of the processing performed by the band group delay correction parameter calculation unit 106 will be used later. Figure 7 Please provide an explanation.

[0051] Next, the processing performed by the speech analysis device 100 will be described in further detail. Here, regarding the processing performed by the speech analysis device 100, the case of performing feature parameter analysis through pitch synchronization analysis will be explained.

[0052] The extraction unit 101 receives the input speech and the pitch marker information, which represents the center time of each speech frame based on its periodicity. Figure 2 This is a diagram of the speech waveform and pitch marker received by the example extraction unit 101. Figure 2 The waveform of the sound “da” is represented, and together with the sound waveform, the pitch marker time is extracted according to the periodicity of the sound (voiced sound).

[0053] The following are examples of audio frames, illustrating the use of... Figure 2 The analysis example for the interval shown below (the underlined interval) is as follows. The extraction unit 101 extracts speech frames by multiplying the pitch marker by a window function twice the length of the pitch. The pitch marker is obtained, for example, by extracting the pitch using a pitch extraction device and extracting the peak value of the pitch period. In addition, non-periodic silent (unvoiced) intervals can also be processed by interpolating pitch markers with fixed frame rates and / or periodic intervals to create a time series that serves as the analysis center, which is then used as the pitch marker.

[0054] In speech frame extraction, a Hanning window can be used. Alternatively, window functions with different characteristics, such as the Hamming window and the Blackman window, can also be used. The extraction unit 101 uses the window function to slice the pitch waveform of the unit waveform within a periodic interval as the speech frame. Furthermore, in non-periodic intervals such as silent / voiceless intervals, the extraction unit 101 also slices speech frames by multiplying them by the window function according to the time determined by interpolating a fixed frame rate and / or a pitch marker, as described above.

[0055] Furthermore, in this embodiment, the use of pitch synchronization analysis in the extraction of spectrum parameters, band group delay parameters, and band group delay correction parameters is taken as an example, but it is not limited to this, and parameters can also be extracted using a fixed frame rate.

[0056] The spectrum parameter calculation unit 102 calculates the spectrum parameters of the speech frames extracted by the extraction unit 101. For example, the spectrum parameter calculation unit 102 calculates arbitrary spectrum parameters representing the spectral envelope, such as Mel cepstrum, linear prediction coefficients, Mel LSP (Line Spectrum Pair), and sine wave model. Furthermore, in cases where analysis is performed based on a fixed frame rate instead of pitch synchronization analysis, these parameters and / or spectral envelope extraction methods implemented by STRAIGHT analysis can also be used for parameter extraction. Here, as an example, Mel LSP-based spectrum parameters are used.

[0057] Figure 3 This is a diagram illustrating a processing example of the spectrum parameter calculation unit 102. Figure 3 (a) represents the speech frame. Figure 3 (b) shows the spectrum obtained by performing a Fourier transform. The spectrum parameter calculation unit 102 applies Mel-LSP analysis to this spectrum to obtain Mel-LSP coefficients. The 0th order of the Mel-LSP coefficients represents the gain term, while the 1st order and above represent the line spectrum frequencies on the frequency axis. Grid lines are shown for each LSP frequency. Here, Mel-LSP analysis is applied to a 44.1 kHz speech signal. The resulting spectral envelope becomes a parameter representing the approximate shape of the spectrum. Figure 3 (c)).

[0058] Figure 4 This is a diagram showing a processing example of the phase spectrum calculation unit 103 and a processing example of the group delay spectrum calculation unit 104. Figure 4 (a) shows the phase spectrum obtained by the phase spectrum calculation unit 103 through Fourier transform. The phase spectrum is an unwrapped spectrum. The phase spectrum calculation unit 103 obtains the phase spectrum by applying a high-pass filter to both the amplitude and phase, so that the phase of the DC component is 0.

[0059] Group delay spectrum calculation unit 104 according to Figure 4 The phase spectrum shown in (a) is obtained by Equation 1 below. Figure 4 The group delay spectrum shown in (b) is an example.

[0060] Formula 1

[0061]

[0062] In Equation 1 above, τ(ω) represents the group delay spectrum. The symbol "'" represents the phase spectrum, and "'" indicates differentiation. The group delay is the frequency derivative of the phase, and is the average time (centroid of the waveform: delay time) of each frequency band in the time domain. The group delay spectrum is equivalent to the expanded differential of the phase, and therefore ranges from -π to π.

[0063] Here, according to Figure 4As can be seen from (b) of , a group delay close to -π is generated at low frequencies. That is, a difference close to π is generated in the phase spectrum at this frequency. In addition, according to Figure 3 the amplitude spectrum of (b) of , a trough can be observed at this frequency position.

[0064] In the low-frequency and high-frequency bands divided by this frequency, since the signs of the signals are opposite, it becomes such a shape. The frequency at which a step difference (in Japanese: 段差) is generated in the phase represents the boundary frequency. Reproducing the discontinuous change of the group delay including the group delay near π on such a frequency axis is important for reproducing the speech waveform of the analysis source and obtaining high-quality analysis-synthesized speech. In addition, as the group delay parameter used for speech synthesis, a parameter that can reproduce such a sharp change in the group delay is required.

[0065] The band group delay parameter calculation unit 105 calculates the band group delay parameter based on the group delay parameter calculated by the group delay spectrum calculation unit 104. The band group delay parameter is the group delay parameter for each pre-determined frequency range. Thereby, the order of the group delay spectrum is reduced, and it becomes a parameter that can be used as a parameter of a statistical model. The band group delay parameter is obtained by the following formula 2.

[0066]

Formula 2

[0067]

[0068] The band group delay based on the above formula 2 represents the average time in the time domain and represents the offset amount relative to the zero-phase waveform. When obtaining the average time from the discrete spectrum, formula 3 is used.

[0069]

Formula 3

[0070]

[0071] Here, the band group delay parameter uses weighting based on the power spectrum, but it is also possible to use only the average of the group delay. In addition, it can also be a different calculation method such as a weighted average based on the amplitude spectrum, as long as it is a parameter representing the group delay of each frequency band.

[0072] In this way, the band group delay parameter becomes a parameter representing the group delay of a predetermined frequency range. Thereby, as shown in the following formula 4, the reconstruction of the group delay according to the band group delay parameter is performed by using the band group delay parameter corresponding to each frequency.

[0073]

Formula 4

[0074]

[0075] The reconstruction of the phase according to the generated group delay is obtained by the following formula 5.

[0076]

Formula 5

[0077]

[0078] The initial value of the phase at ω = 0 is 0 due to the high-pass processing described above, but in practice, the phase of the DC component can also be saved and used in advance. The Ω they use... b It is the frequency scale used as the boundary of the frequency band when calculating the group delay of the frequency band. The frequency scale can use any scale, but it can be set according to auditory characteristics, with fine intervals for low frequencies and coarse intervals for high frequencies.

[0079] Figure 5 This is a diagram illustrating an example of how a frequency scale is constructed. Figure 5 The frequency scale shown uses a Mel scale of α = 0.35 up to 5 kHz, and an equally spaced scale above 5 kHz. To improve waveform shape reproducibility, the group delay parameter represents the power-enhanced low frequencies with fine intervals and the high frequencies with coarse intervals. This is because at high frequencies, the waveform power decreases, and the random phase components caused by aperiodic components increase, making it impossible to obtain stable phase parameters. Furthermore, it is known that the phase at high frequencies has less impact on hearing.

[0080] The control of the random phase component and the component caused by impulse excitation is represented by the intensity of the noise component in each frequency band, which is a periodic component and an aperiodic component. When performing speech synthesis using the output of the speech analysis device 100, the frequency band noise intensity parameter described later is also included in the generated waveform. As a result, the phase of the high-frequency component with strong noise is coarsely represented, reducing the number of iterations.

[0081] Figure 6 This is an example of usage. Figure 5 The graph shown is a result obtained from the analysis based on the band group delay parameter using the frequency scale. Figure 6 (a) represents the band group delay parameter obtained through Equation 3 above. The band group delay parameter becomes a weighted average of the group delay of each band, but it can be seen that the changes that appear in the group delay spectrum cannot be reproduced in the average group delay.

[0082] Figure 6 (b) is an example of a phase plot generated based on the band group delay parameter. Figure 6 In the example shown in (b), although the phase tilt can be roughly reproduced, the phase change close to π at low frequencies and the step difference of the phase spectrum are not captured, and there are parts of the phase spectrum that cannot be reproduced.

[0083] An example of waveform generation obtained by performing an inverse Fourier transform on the generated phase and the amplitude spectrum generated from the Mel-LSP is shown below. Figure 6(c). The generated waveform becomes: in Figure 3 The waveform in (a) shows a shape near the center that is significantly different from the waveform of the source analysis. Thus, when the phase is modeled using only the band group delay parameter, the regenerated waveform differs from the waveform of the source analysis because the phase step difference contained in the speech cannot be captured.

[0084] To address this issue, the speech analysis device 100 uses a band group delay parameter and a band group delay correction parameter, which corrects the phase reconstructed based on the band group delay parameter at a predetermined frequency to the phase of the phase spectrum at that frequency.

[0085] The band group delay correction parameter calculation unit 106 calculates the band group delay correction parameter based on the phase spectrum and the band group delay parameter. The band group delay correction parameter is a parameter that corrects the phase reconstructed using the band group delay parameter to the phase value at the boundary frequency. When the difference is used as a parameter, it is obtained by the following equation 6.

[0086]

Formula 6

[0087]

[0088] The first term on the right side of Equation 6 above is Ω obtained by analyzing the speech. b The phase. The second term of Equation 6 above is obtained using the group delay reconstructed using the band group delay parameter bgrd(b) and the correction parameter bgrdc(b). As shown in Equation 7 below, this is taken as ω=Ω in the group delay of Equation 4 above. b The boundary is represented by the parameter obtained by adding the correction parameter bgrdc(b).

[0089]

Formula 7

[0090]

[0091] The phase thus constructed based on the group delay is reconstructed using Equation 5 above. Furthermore, the second term on the right-hand side of Equation 6 above is obtained as follows: The phase is reconstructed to ω = Ω using Equations 7 and 5 above. b After -1, use Ω b The phase of Equation 8 is reconstructed from the band group delay and used as the Ω. b-1 Band group delay parameters and band group delay correction parameters, Ω up to the specified frequency band b The phase is obtained by reconstructing the frequency band group delay parameters.

[0092]

Form 8

[0093]

[0094] Furthermore, using Equation 6 above, the difference between the phase of the second term on the right and the actual phase is obtained, thereby determining the band group delay correction parameter. Thus, at frequency Ω... b Reproduce the actual phase.

[0095] Figure 7 This is a graph showing the results of analysis using the band group delay correction parameter. Figure 7 (a) represents the group delay spectrum reconstructed from the band group delay parameter and the band group delay correction parameter obtained from Equation 7 above. Figure 7 (b) shows an example where the phase was generated based on the group's delay spectrum. For example... Figure 7 As shown in (b), a phase close to the actual phase can be reconstructed by using the band group delay correction parameter. This is especially true in the low-frequency range where the frequency scale spacing is narrow. Figure 6 The portion of (b) that produces a ladder-like phase difference is also included in the inland area for reproduction.

[0096] Figure 7 (c) represents an example of a waveform synthesized based on the phase parameters thus reconstructed. Figure 6 In the example shown in (c), the shape of the waveform is very different from that of the waveform of the analysis source, but... Figure 7 In the example shown in (c), a speech waveform close to the source waveform is generated. The correction parameter bgrdc in Equation 6 above uses the phase difference information, but it can also be other parameters such as the phase value at that frequency. For example, any parameter that reproduces the phase at that frequency by combining it with the band group delay parameter will suffice.

[0097] Figure 8 This is a flowchart illustrating the processing performed by the speech analysis device 100. The speech analysis device 100 calculates parameters corresponding to each pitch marker by using a cycle of pitch markers. First, in the speech frame extraction step, the extraction unit 101 extracts a speech frame (S801). Next, the spectrum parameter calculation unit 102 calculates the spectrum parameters in the spectrum parameter calculation step (S802), the phase spectrum calculation unit 103 calculates the phase spectrum in the phase spectrum calculation step (S803), and the group delay spectrum calculation unit 104 calculates the group delay spectrum in the group delay spectrum calculation step (S804).

[0098] Next, the band group delay parameter calculation unit 105 calculates the band group delay parameter in the band group delay parameter calculation step (S805). Figure 9 It means Figure 8 The flowchart shows the details of the step (S805) for calculating the band group delay parameters. Figure 9As shown, the band group delay parameter calculation unit 105 sets the boundary frequency of the band by cycling through each band of the predetermined frequency scale (S901), and calculates the band group delay parameter (average group delay) by averaging the group delay using power spectral weights, etc., as shown in Equation 3 above (S902).

[0099] Next, the band group delay correction parameter calculation unit 106 calculates the band group delay correction parameter in the band group delay correction parameter calculation step. Figure 8 (S806). Figure 10 It means Figure 8 The flowchart shows the details of the step (S806) for calculating the band group delay correction parameters. Figure 10 As shown, the band group delay correction parameter calculation unit 106 first sets the boundary frequency of the band using the cycle of each band (S1001). Next, the band group delay correction parameter calculation unit 106 uses Equations 7 and 5 above, using the band group delay parameter and the band group delay correction parameter of the bands below the current band, to generate the phase of the boundary frequency (S1002). Then, the band group delay correction parameter calculation unit 106 calculates the phase spectrum difference parameter using Equation 8 above, and uses the calculated result as the band group delay correction parameter (S1003).

[0100] Thus, the voice analysis device 100 performs... Figure 8 ( Figure 9 , 10 The processing shown in the figure calculates and outputs the spectral parameters, band group delay parameters, and band group delay correction parameters corresponding to the input speech. Therefore, in the case of speech synthesis, the reproducibility of the speech waveform can be improved.

[0101] (Second speech processing device: speech synthesis device)

[0102] Next, the second speech processing device, namely the speech synthesis device, involved in the embodiment will be described. Figure 11 This is a block diagram illustrating the first embodiment of the speech synthesis device (speech synthesis device 1100). For example... Figure 11 As shown, the speech synthesis device 1100 includes an amplitude information generation unit 1101, a phase information generation unit 1102, and a speech waveform generation unit 1103. It receives a spectral parameter sequence, a band group delay parameter sequence, a band group delay correction parameter sequence, and time information of the parameter sequence to generate a speech waveform (synthesized speech). The parameters input to the speech synthesis device 1100 are calculated by the speech analysis device 100.

[0103] The amplitude information generation unit 1101 generates amplitude information based on the spectral parameters at each time. The phase information generation unit 1102 generates phase information based on the band group delay parameters and band group delay correction parameters at each time. The speech waveform generation unit 1103 generates a speech waveform according to the time information of each parameter, based on the amplitude information generated by the amplitude information generation unit 1101 and the phase information generated by the phase information generation unit 1102.

[0104] Figure 12 This diagram illustrates a configuration example of a speech synthesis apparatus 1200 that performs inverse Fourier transform and waveform superposition. The speech synthesis apparatus 1200 is a specific configuration example of the speech synthesis apparatus 1100, and includes an amplitude spectrum calculation unit 1201, a phase spectrum calculation unit 1202, an inverse Fourier transform unit 1203, and a waveform superposition unit 1204. It generates waveforms at various times through inverse Fourier transform, and outputs synthesized speech by superimposing the generated waveforms.

[0105] More specifically, the amplitude spectrum calculation unit 1201 calculates the amplitude spectrum based on the spectral parameters. For example, when using Mel LSP as a parameter, the amplitude spectrum calculation unit 1201 checks the stability of the Mel LSP, transforms it into Mel LPC coefficients, and calculates the amplitude spectrum based on the Mel LPC coefficients. The phase spectrum calculation unit 1202 calculates the phase spectrum based on the band group delay parameter and the band group delay correction parameter using Equations 5 and 7 above.

[0106] The inverse Fourier transform unit 1203 performs an inverse Fourier transform on the calculated amplitude spectrum and phase spectrum to generate the fundamental waveform. An example of the waveform generated by the inverse Fourier transform unit 1203 is shown below. Figure 7 (c). The waveform superposition unit 1204 superimposes and synthesizes the generated pitch waveform based on the time information of the parameter sequence to obtain synthesized speech.

[0107] Figure 13 It means and Figure 2 The diagram shows an example of waveform generation for the interval shown. Figure 13 (a) shows Figure 2 The original audio waveform is shown. Figure 13 (b) is the synthesized speech waveform output by the speech synthesis device 1100 (speech synthesis device 1200) based on the band group delay parameter and the band group delay correction parameter. For example... Figure 13 As shown in (a) and (b), the speech synthesis device 1100 is able to generate a waveform with a shape close to the original sound waveform.

[0108] Figure 13 (c) is shown as a comparative example, illustrating the synthesized speech waveform using only the band group delay parameter. Figure 13As shown in (a) and (c), the synthesized speech waveform using only the band group delay parameter becomes a waveform with a different shape than the original sound.

[0109] Thus, the speech synthesis device 1100 (speech synthesis device 1200) can reproduce the phase characteristics of the original sound by using a band group delay correction parameter in addition to the band group delay parameter, and can make the analyzed synthesized waveform close to the shape of the speech waveform of the analyzed source, thereby generating a high-quality waveform (improving the reproducibility of the speech waveform).

[0110] Figure 14 This is a block diagram illustrating a second embodiment of the speech synthesis apparatus (speech synthesis apparatus 1400). The speech synthesis apparatus 1400 includes a sound source signal generation unit 1401 and a channel filtering unit 1402. The sound source signal generation unit 1401 generates a sound source signal using a band group delay parameter sequence, a band group delay correction parameter sequence, and timing information of the parameter sequence. The sound source signal is a signal generated by using a noise signal for silent regions and a pulse signal for sound regions without phase control or noise intensity, having a flat spectrum, and synthesized into a speech waveform by applying a channel filter.

[0111] In the speech synthesis apparatus 1400, the sound source signal generation unit 1401 uses a band group delay parameter and a band group delay correction parameter to control the phase of the pulse component. That is, Figure 11 The phase control function of the phase information generation unit 1102 shown is implemented by the sound source signal generation unit 1401. That is, the speech synthesis device 1400 uses the band group delay parameter and the band group delay correction parameter in vocoder-type waveform generation to generate waveforms at high speed.

[0112] One method for phase control of the sound source signal is to use the inverse Fourier transform. In this case, the sound source signal generation unit 1401 performs... Figure 15 The processing is shown. That is, at each moment of the characteristic parameters, the sound source signal generation unit 1401 calculates the phase spectrum based on the frequency band group delay parameter and the frequency band group delay correction parameter using Equation 5 and Equation 7 above (S1501), sets the amplitude to 1 and performs an inverse Fourier transform (S1502), and superimposes the generated waveforms (S1503).

[0113] The channel filtering unit 1402 applies a filter determined based on spectral parameters to the generated sound source signal, generates a waveform, and outputs a speech waveform (synthesized speech). The channel filtering unit 1402 has [specific features] for controlling amplitude information. Figure 11 The amplitude information generation unit 1101 shown has the following functions.

[0114] The speech synthesis device 1400, after phase control as described above, can generate waveforms from sound source signals. However, due to the inclusion of inverse Fourier transform processing and filtering operations, it differs from the speech synthesis device 1200 (…). Figure 12 Compared to [previous version], the increased processing load makes it impossible to generate waveforms at high speed. Therefore, the sound source signal generation unit 1401, as [previous version], [continues to adjust / adjust accordingly]. Figure 16 The configuration shown enables the generation of a sound source signal that has undergone phase control only through time-domain processing.

[0115] Figure 16 This is a block diagram showing the configuration of a sound source signal generation unit 1401 that generates a sound source signal that is phase-controlled only through time-domain processing. Figure 16 The sound source signal generation unit 1401 shown prepares a phase-shifted frequency band pulse signal obtained by frequency band division of the phase-shifted pulse signal, delays the phase-shifted frequency band pulse signal and superimposes it to generate a sound source waveform.

[0116] Specifically, the sound source signal generation unit 1401 first stores in the storage unit 1605 the signals of each frequency band obtained by phase shifting and frequency band division of the pulse signal. The phase-shifted frequency band pulse signal refers to the signal in which the amplitude spectrum of the corresponding frequency band is set to 1 and the phase spectrum is set to a constant value, which becomes the signal of each frequency band obtained by phase shifting and frequency band division of the pulse signal, and is generated using the following formula 9.

[0117]

Form 9

[0118]

[0119] Here, the boundary of the frequency band Ω b Phase is determined based on the frequency scale. exist The range is quantized to the P level. With P = 128, 128 frequency band pulse signals of the number of bands are generated based on a step size of 2π / 128 (pitch). Thus, the phase-shifted frequency band pulse signal is obtained by dividing the phase-shifted pulse signal into frequency bands, and the combination is selected by the principal values ​​of the frequency band and phase. When the exponent of the phase shift of frequency band b is set to ph(b), the phase-shifted frequency band pulse signal generated in this way is represented as bandpulse. b ph(b) (t).

[0120] Figure 17 This is a diagram of an example phase-shifted frequency band pulse signal. The left column shows the phase-shifted pulse signal for the entire frequency band; the upper section represents the 0-phase case, and the lower section represents the phase... The situation is as follows. Columns 2 to 6 represent the situations from... Figure 5The frequency band pulse signal from the low frequency to the 5th frequency band of the scale shown. Thus, the storage unit 1605 pre-stores the phase-shifted frequency band pulse signal generated by the frequency band division unit 1606, the phase assignment unit 1607, and the inverse Fourier transform unit 1608.

[0121] The delay time calculation unit 1601 calculates the delay time of each frequency band of the phase-shifted frequency band pulse signal based on the frequency band group delay parameter. The average delay time of the frequency band is represented in the time domain using the frequency band group delay parameter obtained from Equation 3 above, becoming the delay time delay(b) obtained by integerization using Equation 10 below. The group delay corresponding to the integer delay time is taken as τ. int (b) is requested.

[0122]

Formula 10

[0123]

[0124] The phase calculation unit 1602 calculates the phase at the boundary frequency based on the band group delay parameter (which is lower than the calculated frequency band) and the band group delay correction parameter. The phase at the boundary frequency reconstructed based on the parameters is obtained using Equations 7 and 5 above. The selection unit 1603 uses boundary frequency phase and integer group delay bgrd int (b) Calculate the phase of the pulse signal in each frequency band. This phase is used as the... And the tilt is bgrd int The y-intercept of the line in (b) is obtained by the following formula 11.

[0125]

Formula 11

[0126]

[0127] In addition, the selection unit 1603 performs a 2π addition or subtraction operation to make the principal value of the phase obtained by Equation 11 above fall within the range of (0≤phase(b)<2π) (hereinafter referred to as <phase(b)>), and takes the obtained principal value of the phase as the number ph(b) of the phase obtained by quantization when generating the phase-shifted frequency band pulse signal (Equation 12 below).

[0128]

Formula 12

[0129]

[0130] Based on ph(b), the phase-shifted frequency band pulse signal is selected based on the frequency band group delay parameter and the frequency band group delay correction parameter.

[0131] Figure 18This is a conceptual diagram illustrating the selection algorithm performed by the selection unit 1603. Here, an example of selecting a phase-shifted frequency band pulse signal corresponding to a sound source signal in the frequency band where b=1 is shown. The selection unit 1603 generates a frequency band of Ω... b To Ω b+1 Given a sound source signal, calculate the delay and phase tilt (bgrd) obtained by integerizing the group delay parameter of that frequency band. int (b) Furthermore, the selection unit 1603 obtains the phase at the boundary frequency generated based on the band group delay parameter and the band group delay correction parameter. And the tilt is bgrd int (b) The y-intercept phase(b) of the straight line, and the phase-shifted frequency band pulse signal is selected based on ph(b) obtained by quantizing its principal value <phase(b)>.

[0132] Figure 19 This is a diagram representing a phase-shifted frequency band pulse signal. For example... Figure 19 As shown in (a), the pulse signal across the entire frequency band based on phase (b) is a signal with a fixed phase (b) and an amplitude of 1. If a time delay is applied to it, a fixed group delay corresponding to the amount of delay will occur; therefore, as... Figure 19 As shown in (b), it becomes a phase (b) with an inclination of bgrd. int (b) The straight line. Apply a bandpass filter to the signal with the straight phase across the entire frequency band and cut off Ω. b To Ω b+1 The signal obtained from the interval becomes Figure 19 (c) of, amplitude in Ω b To Ω b+1 The boundary Ω is 1 in the interval and 0 in other frequency regions. b The phase is The signal.

[0133] Therefore, utilizing Figure 18 The method shown can appropriately select phase-shifted pulse signals for each frequency band. The superposition unit 1604 delays the selected phase-shifted frequency band pulse signals according to the delay time delay(b) calculated by the delay time calculation unit 1601, and performs addition operations over the entire frequency band, thereby generating a sound source signal that reflects the frequency band group delay parameter and the frequency band group delay correction parameter.

[0134]

Formula 13

[0135]

[0136] Figure 20 This is a diagram illustrating an example of how a sound source signal is generated. Figure 20(a) is the sound source signal for each frequency band, which is a graph showing waveforms obtained by delaying the selected phase shift pulse signals in 5 frequency bands with low frequencies. The sound source signal obtained by performing addition operations on them across the entire frequency band and generating it is shown in Figure 20 (b). The phase spectrum of the signal thus generated is shown in Figure 20 (c), and the amplitude spectrum is shown in Figure 20 (d).

[0137] Figure 20 The phase spectrum shown in (c) represents the phase of the analysis source with a thin line and overlaps and represents the phase generated using Equation 5 and Equation 7 with a thick line. Thus, the phase generated by the sound source signal generation unit 1401 and the phase regenerated based on the parameters basically overlap, except for the parts where there are differences due to differences in the expansion of high frequencies, and a phase close to the analysis source phase is generated.

[0138] According to Figure 20 the amplitude spectrum shown in (d), it can be seen that except for the parts at the zero-crossings where the phase changes greatly, it becomes a spectrum shape that is close to flat with an amplitude generally of 1.0, and the sound source waveform is correctly generated. The sound source signal generation unit 1401 superimposes and synthesizes the sound source signals thus generated according to the pitch marks determined by the parameter sequence time information to generate the sound source signal for the entire sentence.

[0139] Figure 21 is a flowchart showing the processing performed by the sound source signal generation unit 1401. The sound source signal generation unit 1401 performs a loop for each moment of the parameter sequence, calculates the delay time (S2101) using Equation 10 in the frequency band pulse delay time calculation step, and calculates the phase of the boundary frequency (S2102) using Equation 5 and Equation 7 in the boundary frequency phase calculation step. Then, the sound source signal generation unit 1401 selects the phase shift frequency band pulse signals contained in the storage unit 1605 (S2103) using Equation 11 and Equation 12 in the phase shift frequency band pulse selection step, and delays and performs addition operations and superposition on the selected phase shift frequency band pulse signals to generate the sound source signal (S2104).

[0140] The channel filter unit 1402 applies a channel filter to the sound source signal generated by the sound source signal generation unit 1401 to obtain the synthesized speech. In the case of Mel LSP parameters, the channel filter transforms from Mel LSP parameters to Mel LPC parameters, and after performing processes such as gain separation (括りだし), generates a waveform by applying a Mel LPC filter.

[0141] Because the minimum phase characteristic is increased due to the influence of the channel filter, minimum phase correction can also be applied when determining the band group delay parameter and band group delay correction parameter based on the phase of the analysis source. The minimum phase is generated as the following imaginary axis: the amplitude spectrum is generated based on the Mel LSP, an inverse Fourier transform is performed on the spectrum based on the logarithmic amplitude spectrum and zero phase, and the resulting cepstrum is subjected to another Fourier transform such that the positive components are doubled and the negative components are zero.

[0142] The phase thus obtained is expanded and subtracted from the phase obtained from the waveform analysis, thereby performing minimum phase correction. Based on the phase spectrum after minimum phase correction, the band group delay parameter and the band group delay correction parameter are obtained. A sound source is generated through the processing of the aforementioned sound source signal generation unit 1401, and a filter is applied, thereby obtaining synthesized speech that reproduces the phase of the source waveform.

[0143] Figure 22 This is a diagram of the speech waveform generated by including minimum phase correction as an example. Figure 22 (a) is with Figure 13 (a) Speech waveforms from the same analysis source. Figure 22 (b) is an analytical synthesized waveform based on vocoder-type waveform generation by the speech synthesis device 1400. Figure 22 (c) is a vocoder based on a widely used pulse sound source, which in this case becomes the waveform shape of minimum phase.

[0144] Figure 22 The synthesized waveform obtained by the speech synthesis device 1400, as shown in (b), reproduces a waveform close to... Figure 22 The waveform of the original sound is shown in (a). Additionally, the generated waveform is also close to... Figure 13 The speech waveform shown in (b). In contrast, in Figure 22 The minimum phase shown in (c) becomes the speech waveform with power concentrated near the pitch mark, and fails to reproduce the shape of the original speech waveform.

[0145] In addition, to compare processing volumes, the processing time for generating approximately 30 seconds of speech waveform was measured. Regarding the processing time excluding initial settings such as phase-shifted frequency band pulse generation, the inverse Fourier transform was used... Figure 12 In the case of a vocoder configuration, it takes approximately 9.19 seconds. Figure 14 In the case of this configuration, the processing time is approximately 0.47 seconds (measured using a computing server with a 2.9GHz CPU). This means that the processing time has been reduced by about 5.1%. In other words, waveforms can be generated at high speed using vocoder-based waveform generation.

[0146] This is because waveforms reflecting phase characteristics can be generated without using inverse Fourier transform, relying solely on time-domain operations. The waveform generation described above involves generating a sound source, superimposing and synthesizing the sound source waveforms, and then applying a filter, but is not limited to this. It could also involve generating sound source waveforms for each pitch waveform, applying filters, generating pitch waveforms, and then superimposing and synthesizing the generated pitch waveforms, among other different configurations. Furthermore, as long as it uses… Figure 16 The sound source signal generation unit 1401 shown can generate a sound source signal based on the frequency band group delay parameter and the frequency band group delay correction parameter.

[0147] Figure 23 It shows that it is aimed at Figure 12 The diagram shows a configuration example of a speech synthesis device 2300 obtained by adding control over the separation of noise components and periodic components using frequency band noise intensity to the speech synthesis device 1200. The speech synthesis device 2300 is one specific configuration of the speech synthesis device 1100. The amplitude spectrum calculation unit 1201 calculates the amplitude spectrum based on a spectral parameter sequence, and the periodic component spectrum calculation unit 2301 and the noise component spectrum calculation unit 2302 separate the periodic component spectrum and the noise component spectrum according to the frequency band noise intensity. Frequency band noise intensity is a parameter representing the ratio of noise components in each frequency band of the spectrum. For example, it can be obtained by separating speech into periodic and noise components using a PSHF (Pitch Scaled Harmonic Filter) method, calculating the noise component ratio at each frequency, and averaging it according to a predetermined frequency band.

[0148] Figure 24 This is a graph showing the noise intensity of an example frequency band. Figure 24 (a) is the calculation of the spectrum of the speech and the spectrum of the aperiodic components of the target frame obtained by separating speech into periodic and aperiodic components using PSHF, and the calculation of the ratio of the aperiodic components at each frequency, ap(ω). During processing, post-processing is added to set the frequency bands with sound to 0 and / or clipping the ratio between 0 and 1, based on the PSHF-based ratio. The intensity is then calculated by weighting the frequency-scaled spectrum based on the noise component ratios obtained in this way. Figure 24 The band noise intensity bap(b) is shown in (b). The frequency scale is similar to the band group delay, using... Figure 5 The scale shown is obtained using the following formula 14.

[0149]

Formula 14

[0150]

[0151] The noise component spectrum calculation unit 2302 calculates the noise component spectrum by multiplying the noise intensity of each frequency based on the noise intensity of the frequency band by the spectrum generated according to the spectral parameters. The periodic component spectrum calculation unit 2301 calculates the periodic component spectrum after removing the noise component spectrum by multiplying by 1.0-bap(b).

[0152] The noise component waveform generation unit 2304 generates a noise component waveform by performing an inverse Fourier transform based on a random phase generated from the noise signal and an amplitude spectrum based on the noise component spectrum. The noise component phase can be generated, for example, by generating Gaussian noise with an average of 0 and a variance of 1, cutting it using a Hanning window of twice the length of the pitch, and performing a Fourier transform on the cut-out windowed Gaussian noise.

[0153] The periodic component waveform generation unit 2303 performs an inverse Fourier transform on the phase spectrum calculated by the phase spectrum calculation unit 1202 based on the band group delay parameter and the band group delay correction parameter, and the amplitude spectrum based on the periodic component spectrum, thereby generating the periodic component waveform.

[0154] The waveform superposition unit 1204 adds the generated noise component waveform and periodic component waveform, and superimposes them according to the time information of the parameter sequence to obtain synthesized speech.

[0155] Thus, by separating the noise component and the periodic component, the random phase component, which is difficult to represent as a band group delay parameter, can be separated, and the noise component can be generated based on the random phase. This allows for the suppression of no-voice regions and / or the high-frequency parts of audible fricatives, and the transformation of noise components contained in audible sounds into a pulsed, screaming sound quality. Especially when the parameters are statistically modeled, if the band group delay and band group delay correction parameters calculated from multiple random phase components are averaged, there is a tendency for the average value to approach 0, resulting in a near-pulsating phase component. By using the band noise intensity together with the band group delay parameter and the band group delay correction parameter, noise components can be generated based on random phases, and the periodic component can use an appropriately generated phase, thereby improving the sound quality of the synthesized speech.

[0156] Figure 25This diagram illustrates a configuration example of a vocoder-type speech synthesis device 2500 that also uses frequency band noise intensity control for high-speed waveform generation. Noise source generation utilizes a fixed-length frequency band noise signal, pre-divided by frequency band segmentation, contained in the frequency band noise signal storage unit 2503. In the speech synthesis device 2500, the frequency band noise signal storage unit 2503 stores frequency band noise signals, and the noise source signal generation unit 2502 controls the amplitude of the frequency band noise signals in each frequency band according to the frequency band noise intensity, performing addition operations on the amplitude-controlled frequency band noise signals to generate noise source signals. Furthermore, the speech synthesis device 2500 is... Figure 14 A modified example of the speech synthesis device 1400 shown.

[0157] The pulse sound source signal generation unit 2501 uses the phase-shifted frequency band pulse signal stored in the storage unit 1605 to generate a pulse sound source signal generated by... Figure 16 The illustrated configuration is a phase-controlled sound source signal. Specifically, when superimposing a delayed phase-shifted frequency band pulse waveform, the amplitude of the signal in each frequency band is controlled using the frequency band noise intensity, generating a signal such that the intensity becomes (1.0 - bap(b)). The speech synthesis apparatus 2500 adds the pulse sound source signal thus generated to the noise sound source signal to generate a sound source signal, and applies a channel filter based on spectral parameters in the channel filtering unit 1402 to obtain synthesized speech.

[0158] Speech Synthesis Device 2500 and Figure 23 Similarly, the speech synthesis apparatus 2300 shown generates noise signals and periodic signals respectively, suppresses pulse-like noise generated against the noise component, and adds the phase-controlled periodic component and noise component to generate a sound source. Thus, it can perform speech synthesis with a shape close to that of the analysis source waveform. Furthermore, the speech synthesis apparatus 2500 can calculate both the generation of the noise source and the generation of the pulse source solely through time-domain processing, thereby achieving high-speed waveform generation.

[0159] Thus, the first and second embodiments of the speech synthesis apparatus, by using band group delay parameters and band group delay correction parameters, can statistically model and reduce the dimensionality (dimension) of feature parameters, thereby increasing the similarity between the reconstructed phase and the phase obtained from waveform analysis, and enabling speech synthesis with appropriate phase control based on these parameters. Each speech processing apparatus according to the embodiments, by using band group delay parameters and band group delay correction parameters, can improve waveform reproducibility and generate waveforms at high speed. Furthermore, in vocoder-type speech synthesis apparatuses, generating sound source waveforms with phase control only through time-domain processing enables waveform generation based on channel filters, thereby enabling high-speed generation of phase-controlled waveforms. Additionally, by combining the use of band noise intensity parameters, the speech synthesis apparatus can also improve the reproducibility of noise components, resulting in higher-quality speech synthesis.

[0160] Figure 26 This is a block diagram illustrating a third embodiment of the speech synthesis apparatus (speech synthesis apparatus 2600). The speech synthesis apparatus 2600 is obtained by applying the aforementioned band group delay parameter and band group delay correction parameter to a text-to-speech synthesis apparatus. Here, as a text-to-speech synthesis method, in speech synthesis based on HMM (Hidden Markov Model), a statistical model-based speech synthesis technique, the band group delay parameter and band group delay correction parameter are used as feature parameters.

[0161] The speech synthesis apparatus 2600 includes a text parsing unit 2601, an HMM sequence production unit 2602, a parameter generation unit 2603, a waveform generation unit 2604, and an HMM storage unit 2605. The HMM storage unit (statistical model storage unit) 2605 stores an HMM learned from acoustic feature parameters including band group delay parameters and band group delay correction parameters.

[0162] The text parsing unit 2601 parses the input text and extracts information such as pronunciation and stress, creating contextual information. The HMM sequence creation unit 2602, based on the contextual information created from the text and the HMM model stored in the HMM storage unit 2605, creates an HMM sequence corresponding to the input text. The parameter generation unit 2603 generates acoustic feature parameters based on the HMM sequence. The waveform generation unit 2604 generates a speech waveform based on the generated feature parameter sequence.

[0163] More specifically, the text parsing unit 2601 generates context information through language (speech) parsing of the input text. The text parsing unit 2601 performs lexical analysis on the input text, extracting language (speech) information required for speech synthesis, such as pronunciation information and stress information, and generates context information based on the obtained pronunciation and language information. Alternatively, context information can be generated based on separately generated, corrected pronunciation and stress information corresponding to the input text. Context information refers to information used as units for classifying speech, such as phonemes, semiphonemes, and syllable HMMs.

[0164] When phonemes are used as the units of speech, the sequence of phoneme names can be used as contextual information. This allows for the inclusion of language (speech) attributes such as triphones with antecedent / following phonemes, and / or phoneme information containing two phonemes before and after, phoneme category information representing attributes based on audible / absentness classification and / or further detailed phoneme categories, the position of each phoneme within a sentence, within a breathing group (breathing unit), within a stressed phrase, the number of syllables / stress type of a stressed phrase, the position of the syllables, the position up to the stress nucleus, the presence or absence of a final intonation, and assigned symbolic information, all of which can be used as contextual information.

[0165] The HMM sequence generation unit 2602 generates an HMM sequence corresponding to the input context information based on the HMM information stored in the HMM storage unit 2605. An HMM is a statistical model represented by state transition probabilities and the output distribution of each state. When using a left-to-right type HMM as the HMM, such as... Figure 27 As shown, based on the output distribution N(o|μ) of each state i Σ i ) and state transition probability a ij The model (where i and j are state indices) is modeled as a set of transition probabilities consisting only of the transition probabilities to neighboring states and the state's own transition probability. Here, the state's own transition probability a is replaced by... ij Using a duration distribution N(d|μ) i d Σ i d The model is called HSMM (Hidden Semi-Markov Model) and is used for modeling long durations.

[0166] HMM storage unit 2605 stores the model obtained by decision tree clustering of the output distribution of each state of the HMM. In this case, such as Figure 28As shown, the HMM storage unit 2605 stores the decision tree of the model, which serves as the feature parameters of each state of the HMM, and the output distribution of each leaf node of the decision tree. It also stores the decision tree and distribution used for sustained long distributions. Each node of the decision tree is associated with a question that classifies the distribution, such as questions like "is there no sound?", "is there sound?", or "is there a stress kernel?", and child nodes for cases that match the question and for cases that do not. For the input context information, it is determined whether it matches the question at each node, thereby searching the decision tree to obtain leaf nodes. By using the distribution associated with the obtained leaf nodes as the output distribution of each state, an HMM corresponding to each speech unit is constructed. Thus, an HMM sequence corresponding to the input context information is created.

[0167] The HMM stored in HMM storage unit 2605 is composed of... Figure 29 The HMM learning device 2900 shown is used for this purpose. The speech corpus storage unit 2901 stores a speech corpus containing speech data and contextual information used for creating the HMM model.

[0168] The analysis unit 2902 analyzes the speech data used in the learning process and obtains acoustic feature parameters. Here, the aforementioned speech analysis device 100 is used to obtain the band group delay parameter and the band group delay correction parameter, which are used together with the spectrum parameter, pitch parameter, and band noise intensity parameter.

[0169] like Figure 30 As shown, the analysis unit 2902 obtains acoustic feature parameters in each speech frame of the speech data. When using pitch synchronization analysis, the speech frame becomes the parameter for each pitch marker time. Alternatively, when using a fixed frame rate, feature parameters are extracted using methods such as interpolating the acoustic feature parameters of adjacent pitch markers.

[0170] use Figure 1 The speech analysis device 100 shown corresponds to the center time of speech analysis ( Figure 30 The acoustic feature parameters corresponding to the pitch marker positions (in the middle) are analyzed, and the following parameters are extracted: spectral parameters (Mel LSP), pitch parameters (logarithmic F0), band noise intensity parameters (BAP), band group delay parameters, and band group delay correction parameters (BGRD and BGRDC). Furthermore, as dynamic characteristic quantities of these parameters, the Δ parameter and Δ... 2 All parameters are used as acoustic characteristic parameters at each time point.

[0171] HMM Learning Unit 2903 learns HMM based on the characteristic parameters obtained in this way. Figure 31This is a flowchart illustrating the processing performed by the HMM learning unit 2903. The HMM learning unit 2903 initializes the phoneme HMM (S3101), performs maximum likelihood estimation on the phoneme HMM through HSMM learning (S3102), and learns the phoneme HMM as the initial model. During maximum likelihood estimation, it learns by performing connection learning (connection learning) while probabilistically associating each state with the feature parameters based on the HMM of the entire sentence connected to correspond to the sentence and the acoustic feature parameters corresponding to the sentence.

[0172] Next, the HMM learning unit 2903 initializes the context-dependent HMM using a phoneme HMM (S3103). As the context, as described above, using phonological information such as related phonemes, preceding and following phoneme environments, positional information within sentences / stressed phrases, stress type, and whether there is a final intonation, as well as linguistic information, a model initialized with related phonemes is prepared for the context existing in the learning data.

[0173] Then, the HMM learning unit 2903 applies maximum likelihood estimation based on connection learning to learn the context-dependent HMM (S3104), and applies state grouping based on decision trees (S3105). Thus, the HMM learning unit 2903 constructs decision trees for each state / flow and the state duration distribution of the HMM. Furthermore, based on the distribution of each state / flow, the HMM learning unit 2903 learns rules for classifying the model using the maximum likelihood criterion and / or the MDL (Minimum Description Length) criterion, and constructs... Figure 28 The decision tree is shown. Furthermore, during speech synthesis, even in the presence of unknown contexts not present in the input learning data, the distribution of each state can be selected and the corresponding Hidden Markov Model (HMM) constructed by following the decision tree.

[0174] Finally, the HMM learning unit 2903 performs maximum likelihood estimation on the context-dependent clustered model, completing model learning (S3106). During clustering, a decision tree is constructed for each flow of each feature quantity, along with decision trees for the spectral parameters (Mel LSP), pitch parameters (logarithmic fundamental frequency), and band noise intensity (BAP), as well as for the band group delay and band group delay correction parameters. Furthermore, a decision tree for the sustained-length distribution is constructed for each state in a multidimensional distribution, resulting in a sustained-length distribution decision tree at the HMM level. These derived HMMs and decision trees are stored in the HMM storage unit 2605.

[0175] HMM Sequence Production Department 2602 Figure 26Based on the input context and the HMM stored in the HMM storage unit 2605, an HMM sequence is created. The distribution of each state is repeatedly made according to the number of frames determined by the sustained long distribution, thereby creating a distribution column. The created distribution column is a column that arranges the distribution of the number of parameters to be output.

[0176] The parameter generation unit 2603 uses a parameter generation algorithm that takes into account static / dynamic features and is widely used in HMM-based speech synthesis to generate each parameter, thereby generating a smooth parameter sequence.

[0177] Figure 32 This is a diagram illustrating an example of constructing an HMM sequence / distribution. First, the HMM sequence creation unit 2602 selects the distribution of each state / flow and the duration distribution of the HMM in the input context to construct the HMM sequence. As the context, when "aka" is synthesized using "antecedent phoneme_that phoneme_following phoneme_phoneme position_number of phonemes_short syllable position_number of short syllables_stress type", since it is a two-short syllable type 1, the initial "a" phoneme becomes a context like "sil_a_k_1_3_1_2_1" because the antecedent phoneme is "sil", the phoneme is "a", the following phoneme is "k", the phoneme position is 1, the number of phonemes is 3, the short syllable position is 1, the number of short syllables is 2, and the stress type is type 1.

[0178] As we traverse the decision tree of the Hidden Markov Model (HMM), at each intermediate node, we determine questions such as whether the phoneme is 'a' or whether the stress type is type 1. By selecting the distribution of leaf nodes along these questions, the distributions of each stream (Mel LSP, BAP, BGRD, BGRDC, LogF0) and the distribution of duration are chosen as the states of the HMM, forming the HMM sequence. Thus, the HMM sequence and distribution columns constituting each model unit (e.g., phoneme) are arranged in a sentence to create a distribution column corresponding to the input text.

[0179] The parameter generation unit 2603 generates a parameter sequence based on the constructed distribution series using a parameter generation algorithm that utilizes static / dynamic features. When using Δ and Δ... 2 In the case of dynamic feature parameters, the output parameters are obtained according to the following method. Feature parameter o at time t. t Using static feature parameter c t and the dynamic feature parameters Δc determined based on the feature parameters of the preceding and following frames. t Δ 2 c t , represented as o t =(c t ′、Δc t ′、Δ 2 c t′). The static eigenvalue c that maximizes P(O|J,λ) t The constructed vector C = (c0′, ..., c T-1 The value is obtained by setting OTM as a zero vector of T×M dimensions and solving the equation in Equation 15 below.

[0180]

Formula 15

[0181]

[0182] Where T is the number of frames and J is the state transition sequence. If the relationship between feature parameters O and static feature parameters C is determined by the matrix W used to calculate the dynamic features, it is expressed as O = WC. O becomes a 3TM vector, C becomes a TM vector, and W is a 3TM × TM matrix. Furthermore, when μ = (μ... s00 ′、…、μ sJ-1Q-1 ')'、Σ=diag(Σ s00 ′、…、Σ sJ-1Q-1 Let ′)′ be the average vector of the distribution corresponding to the sentence with all the diagonal covariances arranged at each time point, and the covariance matrix. Equation 15 above is used to solve the equation in Equation 16 below to obtain the optimal feature parameter sequence C.

[0183]

Formula 16

[0184] w′∑ -1 wC=w′∑ -1 μ…(16)

[0185] The equation is obtained using a method based on Cholsky decomposition. Furthermore, similar to the solution used in the time update algorithm for RLS filters, it can generate the parameter sequence in chronological order with the delay time, and can also be generated with low latency. Moreover, the parameter generation process is not limited to the methods described above; it can also use methods such as interpolating the average vector, or any other method to generate feature parameters based on other distributions.

[0186] The waveform generation unit 2604 generates a speech waveform based on the parameter sequence thus generated. For example, the waveform generation unit 2604 synthesizes speech based on a Mel LSP sequence, a logarithmic F0 sequence, a band noise intensity sequence, a band group delay parameter, and a band group delay correction parameter. When using these parameters, the waveform is generated using the aforementioned speech synthesis apparatus 1100 or speech synthesis apparatus 1400. Specifically, using... Figure 23 The structure shown is based on the inverse Fourier transform, or Figure 25 The vocoder-style high-speed waveform generator shown is used for waveform generation. Without using frequency band noise intensity, it will use... Figure 12The speech synthesis device 1200 based on inverse Fourier transform shown, or Figure 14 The speech synthesis device 1400 shown is shown.

[0187] Through these processes, synthesized speech corresponding to the input context can be obtained. By using band group delay parameters and band group delay correction parameters, speech that is close to the analyzed source speech can be synthesized so that the phase information of the speech waveform is also reflected.

[0188] Furthermore, in the HMM learning section 2903 described above, a configuration for performing maximum likelihood estimation of the speaker dependency model using a corpus of a specific speaker was explained, but it is not limited to this. Different configurations can also be used, such as speaker adaptation, model interpolation, and other group adaptation learning, which are techniques used to improve the diversity of HMM speech synthesis. In addition, different learning methods can be used, such as estimating the distribution parameters using deep neural networks.

[0189] Alternatively, the speech synthesis apparatus 2600 can also be configured such that a feature parameter sequence selection unit is provided between the HMM sequence production unit 2602 and the parameter generation unit 2603. This unit selects feature parameters from candidates obtained by the analysis unit 2902 targeting the HMM sequence, and synthesizes a speech waveform based on the selected parameters. In this way, by selecting the acoustic feature parameters, the sound quality degradation caused by excessive smoothing in HMM speech synthesis can be suppressed, resulting in synthesized speech that is closer to natural speech production.

[0190] In this way, by using the band group delay parameter and the band group delay correction parameter as feature parameters for speech synthesis, not only can the reproducibility of the waveform be improved, but the waveform can also be generated at high speed.

[0191] Furthermore, the aforementioned speech analysis device 100 and speech synthesis device 1100, among other speech synthesis devices, can also be implemented using a general-purpose computer device as the basic hardware. That is, the speech analysis device and each speech synthesis device in this embodiment can be implemented by having a program executed by a processor mounted on a computer device. This can be achieved either by pre-installing the program on the computer device, storing it on a storage medium such as a CD-ROM, or distributing the program via a network and appropriately installing it on the computer device. Additionally, it can be implemented using memory, hard disks, or storage media such as CD-R, CD-RW, DVD-RAM, DVD-R, etc., either built into or external to the computer device. Furthermore, some or all of the speech synthesis devices, such as the speech analysis device 100 and speech synthesis device 1100, can be constructed using either hardware or software.

[0192] Furthermore, although several embodiments of the present invention have been described through multiple combinations, these embodiments are presented as examples and are not intended to limit the scope of the invention. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and / or variations thereof are included in the scope and / or spirit of the invention, and are included within the scope of the invention as described in the claims and its equivalents.

Claims

1. A voice processing device, comprising: The amplitude information generation unit generates amplitude information based on the sequence of spectral parameters calculated for each speech frame of the input speech. A phase information generation unit generates phase information based on a frequency band group delay parameter sequence within a predetermined frequency range of the group delay spectrum calculated from the phase spectra of each speech frame, and a frequency band group delay correction parameter sequence that corrects the difference between the phase spectrum generated from the frequency band group delay parameter sequence and the phase spectrum of each speech frame; and The speech waveform generation unit generates a speech waveform based on the amplitude information and the phase information at each time determined by the parameter sequence time information, which serves as time information for each parameter.

2. The voice processing device according to claim 1, The amplitude information generation unit, The amplitude spectrum is calculated based on the spectral parameter sequence at each time point. The phase information generation unit The phase spectrum is calculated based on the frequency band group delay parameter sequence and the frequency band group delay correction parameter sequence. The speech waveform generation unit Based on the amplitude spectrum and the phase spectrum, speech waveforms at each time moment are generated, and the generated speech waveforms at each time moment are superimposed and synthesized to generate a speech waveform.

3. The voice processing device according to claim 2, comprising: The noise component spectrum calculation unit calculates the noise component spectrum based on the amplitude information and the noise intensity of each frequency according to the frequency band noise intensity parameter sequence representing the ratio of noise components in a predetermined frequency range. The periodic component spectrum calculation unit calculates the periodic component spectrum of each frequency based on the amplitude information and the frequency band noise intensity parameter sequence. The periodic component waveform generation unit generates a periodic component waveform based on the periodic component spectrum and the phase spectrum constructed from the frequency band group delay parameter sequence and the frequency band group delay correction parameter sequence. as well as The noise component waveform generation unit generates a noise component waveform based on the noise component spectrum and the phase spectrum corresponding to the noise signal. The speech waveform generation unit Based on the periodic component waveform and the noise component waveform, speech waveforms at each time moment are generated, and the generated speech waveforms at each time moment are superimposed and synthesized to generate a speech waveform.

4. A speech processing method, comprising: The step of generating amplitude information based on the sequence of spectral parameters calculated for each speech frame of the input speech; The step of generating phase information is as follows: a sequence of frequency band group delay parameters in a predetermined frequency range of the group delay spectrum calculated from the phase spectrum of each speech frame, and a sequence of frequency band group delay correction parameters that corrects the difference between the phase spectrum generated from the frequency band group delay parameter sequence and the phase spectrum of each speech frame. as well as The step of generating a speech waveform based on the amplitude information and the phase information at each time point determined by the parameter sequence time information which serves as time information for each parameter.

5. A storage medium storing a speech processing program for causing a computer to perform: The step of generating amplitude information based on the sequence of spectral parameters calculated for each speech frame of the input speech; The steps of generating phase information are as follows: a sequence of band group delay parameters within a predetermined frequency range of the group delay spectrum calculated from the phase spectrum of each speech frame, and a sequence of band group delay correction parameters that corrects the difference between the phase spectrum generated from the band group delay parameter sequence and the phase spectrum of each speech frame; and... The step of generating a speech waveform based on the amplitude information and the phase information at each time point determined by the parameter sequence time information which serves as time information for each parameter.

6. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the speech processing method of claim 4.

Citation Information

Patent Citations

  • Voice feature quantity extraction device, voice feature quantity extraction method, and voice feature quantity extraction program

    JP2013164572A

  • Spectral envelope and group delay inference system and voice signal synthesis system for voice analysis / synthesis

    WO2014021318A1

  • Speech processing device, speech processing method, and speech processing program

    CN107924686A

  • Speech processing device

    CN114694632A