A formant extraction method for continuous speech based on peak selection
By setting up reference points and resonance peak troughs, combining linear prediction method and full-pole model, using peak selection method to estimate resonance peaks frame by frame, and eliminating false peaks and merged peaks through cyclic processing, the accurate extraction of resonance peak parameters in continuous speech signals is achieved.
Patent Information
- Application Number
- CN202210492452.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-07
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-05-07
AI Technical Summary
The existing technology has problems of false peaks and peak merging in formant estimation, making it difficult to accurately extract the formant parameters of continuous speech.
A peak selection method is adopted to establish a mapping relationship between peaks and reference points by setting up reference points and resonance peak troughs. The resonance peaks are estimated frame by frame by combining the linear prediction method and the full-pole model. The false peaks and merged peaks are eliminated by 100 cyclic averaging processes.
It effectively solves the problems of false peaks and merged peaks in formant estimation, improves the accuracy and robustness of formant parameters, and is suitable for formant extraction of continuous speech signals.
Smart Images

Figure CN115064180B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of speech signal processing and relates to a continuous speech formant extraction method based on peak selection. Background Art
[0002] Speech is an essential means of human communication. Speech excitation and vocal tract transfer functions can be derived from speech signals. The isolated vocal tract transfer functions have a wide range of applications: estimating the pitch period of speech signals using speech excitation, extracting formants using the vocal tract cepstrum, and using vocal tract parameters for parameter encoding and speech recognition. If the vocal tract is considered a resonant cavity, the formant is the cavity's resonant frequency. Formant parameters are crucial for describing speech signal characteristics and are widely used. Accurate and effective formant extraction methods are highly beneficial for speech signal analysis, synthesis, and speech coding.
[0003] Similar to pitch period extraction, accurate formant estimation is very difficult, with the following main problems: 1) False peaks. Under normal circumstances, the maximum value in the spectrum envelope is entirely caused by the formant. However, some peaks that do not meet the formant requirements may be mistakenly evaluated as formants, resulting in false peaks. 2) Formant merging. Adjacent formant frequencies are too close to be distinguished and are mistakenly thought to be a single formant. Existing formant estimation methods mainly include the cepstrum method and the linear prediction method. The cepstrum method extracts formant parameters using the DFT spectrum after performing a discrete Fourier transform on the speech signal. Its disadvantage is that it is easily affected by the harmonics of the fundamental frequency, and the maximum value only appears at the harmonic frequency. The linear prediction method is divided into the root-finding method and the interpolation method. The root-finding method calculates the poles of the full-pole model of the vocal tract response to obtain the conjugate complex roots to obtain the formant parameters. Its disadvantages are that the calculation is large and the root convergence speed is slow. The interpolation method uses parabolic interpolation technology to solve the problem of obtaining formant frequency values when the frequency resolution is low.
[0004] Compared to the cepstrum method, the linear prediction method uses past sample values to predict current or future sample values. By minimizing the prediction error, a unique set of linear prediction coefficients is obtained, thereby establishing a concise speech signal model and enabling relatively accurate formant parameter estimation. Therefore, the linear prediction method is widely used in formant parameter estimation. However, when estimating formants for continuous speech, the linear prediction method is prone to problems such as peak merging and false peaks. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a continuous speech formant extraction method based on peak selection, which can obtain the formant parameters of the continuous speech signal and effectively solve the problems of peak merging and false peaks in the continuous speech formant estimation.
[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions.
[0007] A method for extracting formants from continuous speech based on peak selection comprises the following steps:
[0008] Step 1. Preprocess the input single-frame speech;
[0009] Step 2. Use the linear prediction method to preliminarily estimate the peak value in the spectral envelope of a frame of speech: Use the linear prediction method to obtain the linear prediction coefficient of a frame of speech, and use the all-pole model to establish the transfer function of the vocal tract. At the same time, based on the speech characteristics of most male and female voices, preliminarily estimate the formant frequency and bandwidth in the spectral envelope of a frame of speech.
[0010] Step 3. Set up the reference point and the resonance peak trough, and then use the peak selection method to establish the mapping relationship between the peak and the reference point;
[0011] Step 4. Determine the formant of a frame of speech using the mapping relationship between the peak value and the reference point and the formant trough;
[0012] Step 5: Estimating the formants of the continuous speech.
[0013] The step 1 comprises:
[0014] Using the short-time average amplitude function M n Distinguish the areas of voice and silence in a frame of speech:
[0015]
[0016] Where x n (m) is the nth frame of speech signal obtained after the original speech signal is windowed and framed;
[0017] Then the voiced area is divided into vowel area and non-vowel area according to the total energy of the spectrum and the energy of the 640-2880Hz area.
[0018] Specifically, the step 3 includes:
[0019] Step 3.1. Establish a reference point:
[0020] For the vast majority of male voices, the four reference points are estimated as E1 = 320 Hz, E2 = 1440 Hz, E3 = 2760 Hz, and E4 = 3200 Hz; for the vast majority of female voices, the four reference points are estimated as E1 = 480 Hz, E2 = 1760 Hz, E3 = 3200 Hz, and E4 = 3520 Hz.
[0021] Step 3.2. Set up the resonance peak tank to store the peak value:
[0022] In each frame of speech, three formant slots S1, S2, and S3 are set. An additional formant slot S4 is set. The existence of S4 is only to prevent the possible fourth peak P4 from competing with P3 for the position of S3. In the final formant estimation result, S4 is removed and the peak of S4 is filled.
[0023] Step 3.3. Establish the mapping relationship between reference point and peak value:
[0024] Find all peaks between 120Hz-3600Hz and record the peak frequency and magnitude; the distance from the reference point E i The four nearest peaks that meet the formant conditions are filled into the corresponding four formant slots. If there is only one suitable peak in the entire spectrum, that peak is used to fill all the formant slots. After the formants are initially selected and filled into the corresponding slots, to ensure that every formant slot is filled, the unassigned common peaks and unfilled slots are processed separately.
[0025] Step 3.4. Processing of unassigned peaks:
[0026] Step 3.4.1. If the unallocated peak value is less than 3 / 4 of the peak value allocated to the slot, skip to step 3.4.3.
[0027] Step 3.4.2. If the unassigned peak is greater than 3 / 4 of the peaks assigned to the slots, move the peaks assigned to the slots to the next slot and move the unassigned peak to the current slot.
[0028] Step 3.4.3. If the slot immediately preceding the closest unassigned peak is still unfilled, move the allocated peak in the closest slot to the previous slot to fill it, and then use the unassigned peak to fill the closest slot.
[0029] If there are still unassigned peaks after the above three steps, the unassigned peaks will be discarded;
[0030] Step 3.5. Processing of unfilled slots:
[0031] If slots S1, S2, and S3 are all filled, the formant estimation is performed as usual regardless of whether S4 is filled;
[0032] If the number of peaks is too small or too small, the spectrum is recalculated so that the unit circle contour of the z transform is closer to the two poles, so that the peak value is enhanced before selection; the full-pole channel transfer function is:
[0033]
[0034] Where G is the gain, ak is the linear prediction coefficient;
[0035] Since the result of (2) is z=e j2πn / N (n=0,1,…,N-1) is calculated, then the spectrum recalculation method is:
[0036] In the equation z = e j2πn / N (n=0,1,…,N-1) and then multiply it by a constant coefficient r (r<1) to make the unit circle contour of the z transform closer to the two poles, so that the peak value is enhanced and sharper, and the problem of merging peaks can also be solved.
[0037] Specifically, the step 5 includes:
[0038] Perform pre-emphasis and window pre-processing on the entire continuous speech input and set the length of the time domain and frequency domain; select a random number N from 10 to 1000 i As the number of frames for speech framing, the frame length is calculated from the number of frames; the formant estimation is performed frame by frame using the peak selection algorithm for single-frame speech formant estimation described in steps 1-4; when the number of frames N i After all the corresponding frames are processed, the results are stored and the next random number of frames is selected, and then the formant is estimated frame by frame; the above process is repeated 100 times, and finally the results after 100 experiments are averaged. After the averaged data is low-pass filtered and smoothed, the final result is output, that is, the final continuous speech formant frequency.
[0039] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0040] 1. The present invention establishes a mapping relationship between peak values and frequency reference points, and at the same time performs an average calculation on the formant peaks obtained by applying the peak selection method in the face of continuous speech, thereby eliminating the influence of merged peaks and false peaks.
[0041] 2. This invention combines the linear prediction root-finding method with the peak selection method. Based on the formant continuity criterion, the formant peak of the previous frame is used as the reference point for the next frame, thus solving the problem of slow convergence of the root-finding method. Furthermore, the universality of peak selection overcomes the problem that the linear prediction root-finding method is only suitable when all roots are conjugate complex roots, demonstrating the robustness of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flow chart of a peak selection method for a single-frame speech according to an embodiment of the present invention.
[0043] Figure 2 It is a flow chart of an embodiment of the continuous speech formant extraction method based on peak selection of the present invention. DETAILED DESCRIPTION
[0044] A much less computationally intensive method for determining formant frequency parameters is to select the first three maximum peaks in a frame of speech spectrum as formants. This method works in most cases, but it may fail in some cases. For example: 1. Two or more poles obtained from the denominator of the vocal tract full-pole model are very close in the frequency domain, resulting in only one formant. 2. Frequency shaping may cause one or more peaks to appear in the frequency domain, and these "false peaks" may be identified as formants.
[0045] The method of the present invention utilizes a peak selection method to improve the analysis of continuous speech signals based on the problems of false peaks and merged peaks in formant estimation.
[0046] Formant continuity is one of the criteria for tracking formants. Therefore, it's generally considered optimal if the positions of the three formants P1, P2, and P3 selected in a frame of speech are close to their positions in the previous frame. However, over-reliance on formant continuity is inappropriate for the following reasons: 1. Formant frequencies can change significantly at the interface between nasal and vowel changes. 2. If a frame of speech is incorrect, the formant positions estimated after that frame will also be incorrect. Therefore, this method starts from the closest correct position in the continuous speech and processes and estimates toward positions where accuracy is uncertain. Because formants are sharpest and clearest in vowels, this method considers finding a "reference point" within each vowel. The speech before and after this "reference point" is then processed separately (i.e., forward and backward processing) from a temporal perspective. The formant frequency closest to the "reference point" is used as a reference to determine the formants in the current frame. The starting point is defined as the middle of each vowel.
[0047] The present invention provides a continuous speech formant extraction method based on peak selection. The method finds all peaks of the spectrum envelope in each frame of speech through a linear prediction root-finding method, sets reference points and formant slots, selects peaks and fills the slots according to the mapping relationship between the peaks and the reference points, and performs a new round of selection when processing slots not filled with peaks and unfilled peaks, updates the peaks in the slots, and after the processing is completed, the peaks in the slots are the formants of the frame of speech, and these formants are used as references for the selection of formants in the next frame. At the same time, formant parameters are obtained for the continuous speech, the continuous speech is divided into frames according to different frame numbers, and the above method is used to cycle 100 times to obtain the formant parameters under different frame number tests, and the results after the 100 cycles are averaged and smoothed to obtain the final result.
[0048] The present invention will be further described in detail below with reference to the accompanying drawings.
[0049] The present invention is a continuous speech formant extraction method based on peak selection, such as Figure 1 As shown, the following steps are included:
[0050] Step 1: First, preprocess the input single-frame speech.
[0051] In order to simplify the calculation, the sound and silence areas of a frame of speech are distinguished, and the short-time average amplitude function M is used to calculate the sound and silence areas of a frame of speech. n , which is a function that measures the change in the amplitude of the speech signal. Its advantage is that when calculating small and large sampling values, there will be no large difference due to squaring. Its expression is as follows:
[0052]
[0053] Where: x n (m) is the nth frame speech signal obtained after the original speech signal is windowed and framed.
[0054] At the same time, the voiced area is divided into vowel area and non-vowel area according to the total energy of the spectrum and the energy of the 640-2880Hz area.
[0055] Step 2: Use the linear prediction method to preliminarily estimate the peak value in the spectrum envelope of a frame of speech.
[0056] The linear prediction method is used to obtain the linear prediction coefficients for a frame of speech, and the all-pole model is used to establish the transfer function of the vocal tract. The purpose of obtaining the poles of the transfer function is to provide a preliminary estimate of the formant frequency and bandwidth of the speech signal. Based on the characteristics of the majority of male and female voices, the preliminary estimate of the formant frequency and bandwidth is used to establish reference points at appropriate locations in the spectrum, preparing for the subsequent peak selection method.
[0057] Step 3. Set up the reference point and resonance peak trough, and then use the peak selection method to establish the mapping relationship between the peak and the reference point.
[0058] Step 3.1 Establish reference points:
[0059] For the vast majority of male voices, the four reference points are estimated to be E1 = 320 Hz, E2 = 1440 Hz, E3 = 2760 Hz, and E4 = 3200 Hz. For the vast majority of female voices, the four reference points are estimated to be E1 = 480 Hz, E2 = 1760 Hz, E3 = 3200 Hz, and E4 = 3520 Hz.
[0060] Step 3.2: Set up the resonance peak tank to store the peak value:
[0061] As mentioned above, formants are determined by peak selection. Therefore, in each frame, three formant slots, S1, S2, and S3, are set to fill in the peaks based on the estimated formant frequencies. At the same time, an additional formant slot, S4, is set to track "false peaks" and formants missing due to peak merging. In a frame of speech, S4 does not need to be filled in; it exists only to prevent a possible fourth peak, P4, from competing with P3 for the position of S3. In the final formant estimation result, S4 and the peak that filled S4 are removed.
[0062] Step 3.3: Establish the mapping relationship between the reference point and the peak value:
[0063] First, find all peaks between 120Hz-3600Hz and record the peak frequency and magnitude. i The four nearest peaks that meet the formant criteria are then assigned to the corresponding four formant slots. If the entire spectrum has only one suitable peak, that peak is used to fill all the formant slots. After the initial formant selection and the corresponding slots are filled, unassigned common peaks and unfilled slots are processed separately to ensure that every formant slot is filled.
[0064] Step 3.4: Process unassigned peaks.
[0065] Unassigned peaks are handled as follows:
[0066] Step 3.4.1: If the unallocated peak size is less than 3 / 4 of the peak size allocated to the slots, skip to step 3.4.3.
[0067] Step 3.4.2: If the size of the unassigned peak is larger than 3 / 4 of the size of the peaks assigned to the slots, then the peaks assigned to the slots are moved to the next slot and the unassigned peak is moved to the current slot.
[0068] Step 3.4.3: If the slot immediately preceding the slot closest to the unassigned peak is still unfilled, move the allocated peak in the slot closest to the unassigned peak to the previous slot to fill it, and then use the unassigned peak to fill the nearest slot.
[0069] If there are still unassigned peaks after the above three steps, the unassigned peaks will be discarded.
[0070] Step 3.5: Process unassigned formant slots.
[0071] Unfilled slots are handled as follows:
[0072] If slots S1, S2, and S3 are all filled, the formant estimation is performed as usual regardless of whether S4 is filled.
[0073] If the number of peaks is too small or the peaks are too small, the spectrum is recalculated so that the unit circle contour of the z transform is closer to the two poles, so that the peaks are enhanced before selection.
[0074]
[0075] The above formula is the full-pole channel transfer function. Among them, G is the gain, a k is the linear prediction coefficient. Since the result of (2) is z=e j2πn / N The spectrum is calculated under the value of (n=0,1,…,N-1), so the method of recalculating the spectrum is as follows:
[0076] In the equation z = e j2πn / N (n=0,1,…,N-1) and then multiply it by a constant coefficient r (r<1) to make the unit circle contour of the z transform closer to the two poles, so that the peak value is enhanced and sharper, and the problem of merging peaks can also be solved.
[0077] Step 4. Determine the formant of a frame of speech using the mapping relationship between the peak value and the reference point and the formant trough.
[0078] Finally, determine the formant of a frame of speech: Use the peak selection method to obtain the formant by selecting the peak closest to the reference point. At the same time, establish a mapping relationship between the reference point and the peak value, and process the unused formant slots and peak values according to the mapping relationship to finally obtain the correct formant of a frame of speech. The peak value in the final formant slot is the formant of the frame of speech. At the same time, the formant of this frame is used as a reference for the formant selection of the next frame. If the formant slot of the previous frame is empty, E is still used. i For reference.
[0079] Step 5: Estimating the formants of the continuous speech.
[0080] like Figure 2 As shown: First, the entire continuous speech is pre-processed by pre-emphasis, windowing, etc. and the length of the time domain and frequency domain is set. By selecting a random number N from 10 to 1000 i As the number of frames for speech framing, the frame length is calculated from the number of frames. The peak selection algorithm for single-frame speech formant estimation described in steps 1-4 is used to perform formant estimation frame by frame. i After all frames have been processed, the results are stored and the next step is to select a random number of frames and estimate the formant frame by frame. This process is repeated 100 times. Finally, the results of these 100 experiments are averaged. After low-pass filtering and smoothing the averaged data, the final result, i.e., the final continuous speech formant frequency, is output.
Claims
1. A continuous speech formant extraction method based on peak selection, characterized in that: The following steps are involved: Step 1. Preprocess the input single-frame speech; Step 2. Use the linear prediction method to preliminarily estimate the peak value in the spectrum envelope of a frame of speech; Step 3. Set up the reference point and the resonance peak trough, and then use the peak selection method to establish the mapping relationship between the peak and the reference point; Step 4. Determine the formant of a frame of speech using the mapping relationship between the peak value and the reference point and the formant trough; Step 5. Estimating formants for continuous speech; The step 1 comprises: Using the short-time average amplitude function M n Distinguish the areas of voice and silence in a frame of speech: Where x n (m) is the nth frame of speech signal obtained after the original speech signal is windowed and framed; Then the voiced area is divided into vowel area and non-vowel area according to the total energy of the spectrum and the energy of the 640-2880Hz area.
2. A continuous speech formant extraction method based on peak selection according to claim 1, characterized in that, The linear prediction coefficient of a frame of speech is obtained using the linear prediction method, and the all-pole model is used to establish the transfer function of the vocal tract. At the same time, based on the characteristics of male and female speech, the formant frequency and bandwidth in the spectrum envelope of a frame of speech are preliminarily estimated.
3. A continuous speech formant extraction method based on peak selection according to claim 1, characterized in that: In step 3, the establishment of the reference point and the resonance peak trough, and the establishment of the mapping relationship between the reference point and the peak value include the following steps: Step 3.
1. Establish a reference point: For male speech, the estimates at the four reference points are set to E1 = 320 Hz, E2 = 1440 Hz, E3 = 2760 Hz, and E4 = 3200 Hz; for female speech, the estimates at the four reference points are set to E1 = 480 Hz, E2 = 1760 Hz, E3 = 3200 Hz, and E4 = 3520 Hz. Step 3.
2. Set up the resonance peak tank to store the peak value: In each frame of speech, three formant slots S1, S2, and S3 are set. An additional formant slot S4 is set. The existence of S4 is only to prevent the possible fourth peak P4 from competing with P3 for the position of S3. In the final formant estimation result, S4 is removed and the peak of S4 is filled. Step 3.
3. Establish the mapping relationship between reference point and peak value: Find all peaks between 120Hz-3600Hz and record the peak frequency and magnitude; the distance from the reference point E i The four nearest peaks that meet the formant conditions are filled into the corresponding four formant slots. If there is only one suitable peak in the entire spectrum, that peak is used to fill all the formant slots. After the formants are initially selected and filled into the corresponding slots, to ensure that every formant slot is filled, the unassigned common peaks and unfilled slots are processed separately.
4. A continuous speech formant extraction method based on peak selection according to claim 3, characterized in that: The processing of the unassigned common peak value and the unfilled slot includes: Step 3.
4. Processing of unassigned peaks: Step 3.4.
1. If the unallocated peak value is less than 3 / 4 of the peak value allocated to the slot, skip to step 3.4.
3. Step 3.4.
2. If the unassigned peak is greater than 3 / 4 of the peaks assigned to the slots, move the peaks assigned to the slots to the next slot and move the unassigned peak to the current slot. Step 3.4.
3. If the slot immediately preceding the closest unassigned peak is still unfilled, move the allocated peak in the closest slot to the previous slot to fill it, and then use the unassigned peak to fill the closest slot. If there are still unassigned peaks after steps 3.4.1 to 3.4.3 above, discard the unassigned peaks; Step 3.
5. Processing of unfilled slots: If slots S1, S2, and S3 are all filled, the formant estimation is performed as usual regardless of whether S4 is filled; If the number of peaks is too small or too small, the spectrum is recalculated so that the unit circle contour of the z transform is closer to the two poles, so that the peak value is enhanced before selection; the full-pole channel transfer function is: Where G is the gain, a k is the linear prediction coefficient; Since the result of (2) is z=e j2πn / N The spectrum is calculated under the value of , so the spectrum is recalculated as follows: In the equation z = e j2πn / N Then multiply it by a constant coefficient r, where r<1, so that the unit circle contour of the z transform is closer to the two poles, the peak value is enhanced and sharper, and the problem of merging peaks is also solved; the value of n is 0, 1, 2...N.
5. A continuous speech formant extraction method based on peak selection according to claim 1, characterized in that: In step 5, the formant estimation for the continuous speech includes: Perform pre-emphasis and window pre-processing on the entire continuous speech input and set the length of the time domain and frequency domain; select a random number N from 10 to 1000 i As the number of frames for speech framing, the frame length is calculated from the number of frames; the formant estimation is performed frame by frame using the method described in steps 1-4; when the number of frames N i After all the corresponding frames are processed, the results are stored and the next random number of frames is selected, and then the formant is estimated frame by frame; the above process is repeated 100 times, and finally the results after 100 experiments are averaged. After the averaged data is low-pass filtered and smoothed, the final result, that is, the final continuous speech formant frequency, is output.
Citation Information
Patent Citations
Visualization method of Chinese mandarin complex vowels based on formant frequency
CN101894566A
Methods and apparatus for formant-based voice systems
US20070061145A1
Cited By
Detection system and method for judging singing sound part
CN118078212A