A voice reconstruction method based on vocal tract filtering and glottal excitation
Through the method based on channel filtering and glottal excitation, the audio feature start and end points are marked, the fundamental tone frequency and channel parameters are extracted, and the glottal excitation and channel response are constructed, which solves the complexity and high data requirements of the traditional parameter synthesis method, and realizes efficient speech reconstruction.
Patent Information
- Application Number
- CN202111650490.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-29
AI Technical Summary
Traditional parameter synthesis method has complex algorithms in speech reconstruction and too many parameters are extracted, resulting in large amounts of calculations and high data demand.
By marking the start and end points of the audio voice characteristics, extracting the fundamental tone frequency and channel parameters, constructing glottal excitation and channel responses, using channel filtering and glottal excitation to reconstruct speech, reducing computing volume and data requirements.
It improves the efficiency of speech reconstruction, reduces the computing time, and ensures the accuracy of speech reconstruction.
Smart Images

Figure CN114974271B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a voice reconstruction method based on vocal tract filtering and glottal excitation, and belongs to the technical field of speech synthesis. Background Art
[0002] With the continuous progress of society, speech synthesis technology is widely used in people's daily lives, and its application value has been increasingly emphasized. Whether the synthesized voice can be anthropomorphic and emotional affects the human-computer interaction experience of the device.
[0003] Speech synthesis technology can be mainly divided into three categories: waveform synthesis method, parameter synthesis method, and rule synthesis method. The waveform synthesis method mainly stores the speech waveforms to be synthesized according to different phoneme speech waveforms, and when needed, the required materials are retrieved from the material library, spliced and synthesized, and then output; the parameter synthesis method mainly extracts the parameters of the speech, and synthesizes the required speech signal from the parameter changes; the rule synthesis method stores the acoustic parameters of the speech in the system, forms syllables and words and sentences from phonemes, controls rules such as pitch, rhythm, and stress, and after synthesizing the target text, then converts it into continuous sound waves using the rules.
[0004] With the advantages of small demand for the material speech library, convenient extraction of parameters, and a relatively wide range of prosodic features that the system can adapt to, the parameter synthesis method has developed rapidly in recent years. However, it still has the disadvantages of complex extraction algorithms, excessive extracted parameters, and plain emotions in the synthesized speech.
[0005] The vocalization of humans is achieved by the glottis continuously opening and closing, causing the airflow at the glottis to impact the vocal cords and generate vibrations. These airflows pass through the vocal tract to produce speech. The vocal tract is also constantly changing during speech, so different voices can be heard. The vocalization model mentioned in the present invention is based on the characteristics of the human vocal organs and the principle of speech generation. By extracting the fundamental frequency and vocal tract parameter characteristics of the speaker at different moments in the speech, it simulates the glottal excitation and vocal tract changes during vocalization, and reconstructs the speech signal. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a voice reconstruction method based on vocal tract filtering and glottal excitation to solve the problems of complex extraction algorithms and excessive extracted parameters in the traditional parameter synthesis method.
[0007] The technical solution of the present invention is: a voice reconstruction method based on vocal tract filtering and glottal excitation, characterized in that:
[0008] Step1: Mark the start and end points of the audio voice characteristics, and mark the position information of the speech segments and non-speech segments in the audio according to the flatness of the short-time energy of the detected audio in the frequency domain for use in extracting the fundamental frequency.
[0009] The specific method for marking the start and end points of the marked audio voice features is as follows: After performing frame division and windowing on the audio signal, the positions of the speech segments and non-speech segments in the audio are marked by detecting the flatness of the short-time energy of the audio in the frequency domain, distinguishing the speech segments and non-speech segments in the audio. The extraction result is represented by a two-dimensional array to indicate the endpoint position information of a segment of speech, thereby extracting the start and end points of the audio voice.
[0010] Step2: The fundamental frequency contains the acoustic information of the speaker in the audio. Extract the fundamental frequency of the audio, separate the cepstrum of the acoustic pulse and the cepstrum of the vocal tract response according to the cepstrum method, and extract the fundamental frequency of each frame of the audio.
[0011] Specifically, the quotient of the sampling frequency of the signal and the position of the maximum value in the frequency range after cepstrum is used as a feature. The extraction result is represented by a one-dimensional array to indicate the fundamental frequency of a segment of sample audio.
[0012] Step3: Construct an analog glottal excitation according to the extracted fundamental frequency;
[0013] Specifically, through the sample fundamental frequency extracted in Step2, after interpolation, smoothing, and normalization, the fundamental frequency is input into a voltage-controlled oscillator to output an oscillation signal in the range of 75 - 300 Hz. The oscillation signal is converted into a rectangular wave and delayed by 1 unit for phase subtraction to obtain the required impulse signal. The output signal is represented by a one-dimensional array to indicate the analog glottal excitation of the sample audio.
[0014] Step4: According to the characteristics of the discrete cosine transform, extract the characteristic response of the vocal tract. After performing a fast Fourier transform and taking the logarithmic spectrum on the audio, use the discrete cosine transform to restore the cepstrum data, and finally extract the part with concentrated energy as the analog vocal tract response and use it as the filter data for reconstructing the speech;
[0015] Specifically, perform a fast Fourier transform and a logarithmic operation on the original speech data after frame division, then extract half of the data points (i.e., 128 points) for discrete cosine transform to restore the phase part of the other half of the data, and then perform an inverse fast Fourier transform. Extract 42 points with the most concentrated energy in the oral cavity characteristics as the parameters of the FIR filter, that is, the analog vocal tract response.
[0016] Step5: Reconstruct the speech using the extracted glottal excitation and vocal tract response. Through the glottal excitation response extracted in Step3 and the vocal tract time-domain FIR filter parameters extracted in Step4, make the extracted glottal parameters pass through the FIR filter frame by frame in the form of convolution, and finally stack the data of each frame in a one-dimensional array through inverse frame division, and write the one-dimensional array into an audio file according to the sampling rate of the original speech.
[0017] The beneficial effects of the present invention are as follows: The calculation amount of the extracted vocal tract pulses is small and the operation time is fast. Only half of the data is required to construct the vocal tract parameters, reducing the operation time. The audio endpoint detection reduces the interference of the silent segment speech on the extraction of the reconstruction parameters, improving the operation efficiency. Therefore, aiming at the disadvantages of large calculation amount and high data requirement in the prior art for speech reconstruction, the present invention improves the reconstruction efficiency on the premise of ensuring the accuracy of speech reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the sound production model diagram adopted by the present invention;
[0019] Figure 2 is the overall structural block diagram of the present invention;
[0020] Figure 3 is the cepstrum diagram of a frame of speech signal of the present invention;
[0021] Figure 4 is the fundamental frequency estimation result diagram of the present invention;
[0022] Figure 5 is the waveform diagram of a frame of glottal excitation signal of the present invention;
[0023] Figure 6 is the waveform diagram of a frame of vocal tract parameter of the present invention;
[0024] Figure 7 is the comparison diagram of the spectrogram of the original speech and the reconstructed speech. DETAILED DESCRIPTION OF THE INVENTION
[0025] The present invention will be further described below in conjunction with the drawings and specific embodiments.
[0026] A speech reconstruction method based on vocal tract filtering and glottal excitation, the reconstruction control system diagram of which is as Figure 1 shown. The parameters such as the voiced / silent segment marker, fundamental frequency, glottal parameters, white noise, etc. required for the reconstructed audio are extracted through the parameter extraction module, and the target audio is reconstructed through the FIR filter, so as to solve the problems of complex extraction algorithm and excessive extracted parameters in the traditional parametric synthesis method.
[0027] The present invention is mainly divided into two parts, namely, extracting vocal tract filtering parameters and extracting glottal excitation parameters, and the overall flowchart is as Figure 2 shown.
[0028] The specific technical solution of the present invention is as follows:
[0029] Step1: Mark the start and end points of the audio voice features, and mark the positions of the voiced and unvoiced segments in the audio according to the flatness of the short-time energy of the detected audio in the frequency domain for use in extracting the fundamental frequency;
[0030] Step 2: Extract the fundamental frequency of the audio. Separate the cepstrum of the acoustic pulse and the cepstrum of the vocal tract response according to the cepstrum method, and extract the fundamental frequency of each frame of the audio;
[0031] Step 3: Construct the impulse response simulating the glottal pulse according to the extracted fundamental frequency;
[0032] Step 4: According to the characteristics of the discrete cosine transform, after performing a fast Fourier transform and taking the logarithmic spectrum on the audio, use the discrete cosine transform to restore the cepstrum data, and finally extract the part with concentrated energy as the simulated vocal tract response and use it as the filter data for the reconstructed speech;
[0033] Step 5: Reconstruct the speech using the extracted glottal excitation and vocal tract response.
[0034] The specific content of the said Step 1 is as follows:
[0035] Step 1.1: Perform frame division and windowing on the audio signal in the material library, where the window length is 256 and the frame shift is 128, and perform fast Fourier transform calculation on the short-time speech frame signal after windowing;
[0036] Step 1.2: Calculate the energy of the spectrum of each frame;
[0037] Step 1.3: Calculate the probability density function of each sample point in each frame;
[0038] Step 1.4: Calculate the spectral entropy value of each frame as shown in Equation (1):
[0039]
[0040] In the formula, H(i) is the spectral entropy of the i-th frame, and P(n, i) is the normalized spectral probability density function of the spectral line n in the i-th frame;
[0041] Set the decision threshold. In this embodiment, the threshold is set to 0.12;
[0042] Perform endpoint detection according to the entropy spectrum values of each frame. Values lower than the threshold are represented by 0, considered as silent segments, and values higher than the threshold are represented by 1, considered as voiced segments. The detection result is represented by a one-dimensional array X representing the endpoint detection result, and the length of the array is the number of frames after frame division;
[0043] The specific content of the said Step 2 is as follows:
[0044] Step 2.1: Perform a fast Fourier transform on the framed speech signal x n (m) to obtain the signal X n (k), take the modulus and logarithm of it to obtain the amplitude spectrum as shown in Equation (2):
[0045] E n = 20 log 10 (|X n (k)|) (2)
[0046] Step 2.2: Perform the inverse fast Fourier transform on E n to obtain the cepstrum of the frame signal. As shown in Figure 3 , a peak with an equal interval between harmonics will be shown in the cepstrum. The quotient between the sampling frequency and the peak is the required fundamental frequency. Find the coordinate values S1, S2 between the two peaks and the speech sampling frequency f s , and the fundamental frequency can be obtained according to Equation (3):
[0047]
[0048] In the formula, i is the current i-th frame.
[0049] The estimation of the fundamental frequency of the speech signal is as shown in Figure 4 . The background in the figure is the spectrogram of the speech signal, and the fundamental frequency algorithm of the present invention is extracted more accurately.
[0050] In order to facilitate the estimation of the fundamental frequency value, the present invention performs detection within the range of the cepstrum corresponding to the fundamental frequency of 60 - 500 Hz, that is, searches for the maximum peak coordinate S within the abscissa interval (16, 133) of the cepstrum (i) , and calculates the fundamental frequency according to Equation (4):
[0051]
[0052] Step 2.3: Output the fundamental frequency calculated for each frame to a one-dimensional array, and the length of the array is the total number of frames after the speech signal is framed, as shown in Equation (5), where n is the total number of frames after framing.
[0053] F = [F (1) , F (2) , F (3) , …, F (n) (5)
[0054] The specific content of Step 3 is as follows:
[0055] Step 3.1: Use the cubic spline interpolation method for the fundamental frequency extracted in Step 2.3 to generate a fundamental frequency sequence F that makes the fundamental period transition between frames smoother c , and the length is the product of the frame length and the total number of frames;
[0056] Step 3.2: Feed the interpolated fundamental frequency into the VCO voltage-controlled oscillator, and its expression is as shown in Equation (6):
[0057]
[0058] Step 3.3: Normalize the data output by the VCO as shown in Equation (7);
[0059]
[0060] In the formula, the normalization range is the frequency range from 75 to 300 Hz, and the waveform of the frame glottal excitation signal obtained is as Figure 5 shown.
[0061] Transform z(n) into a rectangular wave r(n), and perform differential decision on the rectangular wave r(n) according to Equation (8) to obtain the glottal pulse excitation;
[0062] R(n) = r(n) - r(n - 1) (8)
[0063] The glottal pulse excitation extracted from a frame of speech signal is as Figure 5 shown, where the abscissa represents the number of sampling points within a frame, the ordinate represents the amplitude value of the glottal pulse excitation, and the number of glottal pulses within a frame is determined by the fundamental period of the current frame.
[0064] The specific content of Step 4 is as follows:
[0065] Step 4.1: Perform FFT transformation on the speech data of each frame, with the number of points being 256, and take the logarithmic spectrum P1 of the first 128 points;
[0066] Step 4.2: Perform discrete cosine transformation on the logarithmic spectrum P1 to obtain P2, and take the data in the range of 1 - 25 in P2 to perform inverse discrete cosine transformation to obtain the matrix P3;
[0067] Step 4.3: Reverse P3 to obtain the matrix P4, and combine P3 and P4 to reconstruct a new matrix P5 = [P3, P4].
[0068] Step 4.4: Perform inverse Fourier transformation on P5 and take the real part to obtain the all - pole filter parameter matrix G of the vocal tract model.
[0069] Step 4.5: Select 42 points with the most concentrated energy in G as the glottal filter parameter matrix G1, and the output result is represented by a two - dimensional array, where the row represents the number of frames and the column represents the vocal tract filter parameters of each frame.
[0070] The vocal tract parameters extracted from a frame of speech signal are as Figure 6 shown. The abscissa represents the number of sampling points, and the ordinate represents the amplitude of the vocal tract parameters.
[0071] The specific content of Step 5 is as follows:
[0072] Step 5.1: In the extracted voiced and unvoiced segment labeling matrix X, when the current frame number is 0, i.e., the unvoiced segment, use random white noise to generate the glottal pulse excitation of the unvoiced segment, extract the vocal tract parameters of the current frame and put them into the FIR filter to reconstruct the speech of the current frame, and update the parameters once for each frame;
[0073] Step 5.2: When the current frame number is 1, i.e., the voiced segment, extract the glottal pulse excitation of the current frame, extract the vocal tract parameters of the current frame and put them into the FIR filter to reconstruct the speech of the current frame of the voiced segment, and update the parameters once per frame;
[0074] Step 5.3: Save the reconstructed speech data of each frame into the matrix W, where each column stores the reconstructed speech signal of each frame, with a total of N rows.
[0075] Step 5.4: Restore the matrix W to the speech signal by inverse frame division. The synthesized speech spectrogram is as follows: Figure 7 As shown. The first sub-graph in the figure is the spectrogram of the original speech, and the second sub-graph is the spectrogram of the reconstructed speech. It can be seen from the figure that the method used in the present invention can well restore the original speech, the relationship between each formant and each harmonic can be well restored in the low frequency, and the various information contained in the speech can be well restored in the high frequency. The reconstructed speech is input into the speech-to-text software, and the text information of the reconstructed speech can be recognized.
[0076] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A voice reconstruction method based on vocal tract filtering and glottal excitation, characterized in that: Step1: Mark the start and end points of the audio voice features. According to the flatness of the short-time energy of the detected audio in the frequency domain, mark the positions of the voiced segments and unvoiced segments in the audio for extracting the fundamental frequency. Step2: Extract the fundamental frequency of the audio. Separate the cepstrum of the acoustic pulse and the cepstrum of the vocal tract response according to the cepstrum method, and extract the fundamental frequency of each frame of the audio. Step3: Construct an analog glottal excitation according to the extracted fundamental frequency. Step4: After performing a fast Fourier transform and taking the logarithmic spectrum on the audio, use the discrete cosine transform to restore the cepstrum data, and finally extract the part with concentrated energy as the analog vocal tract response and use it as the filter data for reconstructing the voice. Step5: Reconstruct the voice using the extracted glottal excitation and vocal tract response. Step3 specifically is: Step 3.1: Generate a pitch frequency sequence F that makes the transition of the pitch period between frames smoother by using cubic spline interpolation for the pitch frequency. c The length is the product of the frame length and the total number of frames. Step3.2: Send the interpolated fundamental frequency into a voltage-controlled oscillator (VCO), and its expression is as shown in Equation (6): Step3.3: Normalize the data output by the VCO, as shown in Equation (7); In the formula, the normalization range is the frequency range from 75 to 300 Hz; Transform z(n) into a rectangular wave r(n), and perform a differential decision on the rectangular wave r(n) according to Equation (8) to obtain the glottal pulse excitation. R(n) = r(n) - r(n - 1) (8) Step4 specifically is: Step4.1: Perform an FFT transform on the voice data of each frame, with 256 points, and take the logarithmic spectrum P1 of the first 128 points. Step4.2: Perform a discrete cosine transform on the logarithmic spectrum P1 to obtain P2, and take the data in the range of 1 - 25 in P2 and perform an inverse discrete cosine transform to obtain the matrix P3. Step4.3: Reverse P3 to obtain the matrix P4, and group P3 and P4 to reconstruct the new matrix P5 = [P3, P4]; Step4.4: Perform an inverse Fourier transform on P5 and take the real part to obtain the vocal tract model all-pole filter parameter matrix G; Step4.5: Take out the 42 points with the most concentrated energy in G as the glottal filtering parameter matrix G1, and the output result is represented by a two-dimensional array, where the rows represent the number of frames and the columns represent the vocal tract filtering parameters of each frame.
2. The voice reconstruction method based on vocal tract filtering and glottal excitation according to claim 1, wherein, In Step1, marking the start and end points of the audio voice features specifically is: After performing frame division and windowing processing on the audio signal, mark the positions of the voiced segments and unvoiced segments in the audio by detecting the flatness of the short-time energy of the audio in the frequency domain, distinguish the voiced segments and unvoiced segments in the audio, and the extraction result is represented by a group of two-dimensional arrays for the endpoint position information of a section of voice, so as to extract the start and end points of the audio voice.
3. The voice reconstruction method based on vocal tract filtering and glottal excitation according to claim 1, characterized in that Step2 specifically is: Use the quotient of the sampling frequency of the signal and the position of the maximum value in its frequency range after cepstrum as a feature, and the extraction result is represented by a group of one-dimensional arrays for the fundamental frequency of a section of sample audio.
Citation Information
Patent Citations
Cepstrum domain pitch period estimation method based on subband SNR weighting
CN109346106A