A processing method, experimental system and experimental method for Mandarin nasal-obstructed speech signals

By adjusting the nasal congestion speech signal at low frequency and reconstructing high frequency, the problem of cold nasal congestion speech affecting teaching and work is solved, effective recovery and clarity of speech signals are achieved, and practical cases are provided for teaching.

CN119446122BActive Publication Date: 2025-05-27ZHANGJIAKOU VOCATIONAL & TECH EDUCATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411561295.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-05-27
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

Nasal congestion caused by colds affects teaching and work results. The existing technology can only recognize cold pronunciations but fail to restore normal pronunciation.

Method used

By collecting Mandarin nasal congestion speech signals, detecting and classifying final speech frames, nasal congestion speech recovery processing is performed for nasal congestion speech frames containing n and ng, including adjusting the amplitude spectrum of the low-frequency spectrum, reconstructing the high-frequency amplitude spectrum and combining the phase spectrum to restore the normal speech signal.

Benefits of technology

It realizes effective recovery of nasal congestion pronunciation, improves language clarity in teaching and work, and allows students to understand and master voice signal processing technology through experimental teaching systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446122B_ABST
    Figure CN119446122B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for processing Mandarin nasal voice signals, an experimental system and a method. The processing method includes the following steps: Step S1, collect Mandarin nasal voice signals, detect the vowel voice frames in the nasal voice signals, and perform classification detection on the vowel voice frames to determine whether the vowel voice frame belongs to one of the 16 nasal vowel voice frames containing n or ng, so as to detect the nasal vowel voice frames and determine the category of the nasal vowel voice frame, that is, determine which category of the 16 nasal vowel voice frames the nasal vowel voice frame belongs to; the above 16 nasal vowel voice frames containing n or ng refer to the voice frames of 8 front nasal vowels an, ian, uan, üan, en, uen, in, ün and the voice frames of 8 back nasal vowels ang, iang, uang, eng, ueng, ing, ong, iong; Step S2, perform nasal voice restoration processing on the nasal vowel voice frames, and do not perform processing on other voice frames. The present invention can achieve efficient processing of nasal voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to voice signal processing technology, and particularly to a method for processing nasal voice signals in Mandarin Chinese.

[0002] The present invention also relates to an experimental system and an experimental method, which are used for teaching in courses such as multimedia technology and digital signal processing offered in higher vocational colleges and universities through voice signal processing experiments. Background Art

[0003] Cold, fever, and seasonal allergic rhinitis are common diseases that people often suffer from. In these cases, people often experience symptoms such as nasal congestion and runny nose, which lead to the generation of nasal voice, that is, the voice of speaking is dull and difficult to hear clearly. For teachers, it often affects the listening effect of students and leads to a decline in teaching quality. Similarly, for other people who rely on language for their work, such as announcers and online anchors, it is also an urgent problem to be solved.

[0004] In the prior art, "Research on Feature Extraction and Recognition of the Voices of Cold Patients" mainly extracts a total of 16-dimensional features including formants F1, F2, F3, and 12-dimensional Mel cepstral coefficients, and uses a trained neural network to distinguish between cold and non-cold voices. However, it only involves the recognition of cold voices and does not give a method for restoring normal pronunciation from cold voices. Summary of the Invention

[0005] The object of the present invention is to process nasal voice signals so as to restore nasal voice to normal pronunciation.

[0006] On the other hand, the object of the present invention also relates to applying the processing method to courses such as multimedia technology and digital signal processing in higher vocational colleges and universities, so as to facilitate students' learning, understanding, and mastery of the principles and practices of multimedia technology and voice signal processing technology.

[0007] To this end, the present invention proposes a method for processing Mandarin Chinese nasal voice signals, including the following steps:

[0008] Step S1: Collect Mandarin Chinese nasal voice signals, detect the vowel voice frames in the nasal voice signals, and perform classification detection on the vowel voice frames to determine whether the vowel voice frame belongs to one of the 16 categories of nasal vowel voice frames containing n and ng, so as to detect the nasal vowel voice frames and determine the category of the nasal vowel voice frame, that is, determine which category of the 16 categories of nasal vowel voice frames the nasal vowel voice frame belongs to;

[0009] The above 16 categories of nasal final speech frames containing n and ng refer to the speech frames of 8 front nasal finals an, ian, uan, üan, en, uen, in, ün and the speech frames of 8 back nasal finals ang, iang, uang, eng, ueng, ing, ong, iong;

[0010] Step S2: Perform nasal congestion speech restoration processing on this nasal final speech frame, and do not perform processing on other speech frames;

[0011] Among them, the nasal congestion speech restoration processing in step S2 specifically includes the following steps:

[0012] Step S21: Adjust the low-frequency spectrum amplitude spectrum of this nasal final speech frame to obtain the amplitude spectrum of the nasal final speech signal after low-frequency adjustment;

[0013] Step S22: Reconstruct the high-frequency amplitude spectrum of this nasal final speech frame according to the amplitude spectrum of the nasal final speech signal after low-frequency adjustment obtained in step S21 to obtain the amplitude spectrum after high-low frequency restoration;

[0014] Step S23: Combine the phase spectrum of this nasal final speech frame on the basis of the amplitude spectrum after high-low frequency restoration obtained in step S22 to obtain the restored spectrum, and obtain the restored speech frame time-domain signal according to the restored spectrum.

[0015] In another technical solution, the adjustment of the low-frequency spectrum amplitude spectrum of this nasal final speech frame in step S21 to obtain the amplitude spectrum of the nasal final speech signal after low-frequency adjustment specifically includes:

[0016] Multiply the first low-frequency part of the amplitude spectrum containing n and ng nasal finals in the nasal congestion speech signal by a low-frequency adjustment vector to obtain an adjusted spectrum, and then superimpose the remaining part of the amplitude spectrum of the nasal final speech frame to obtain the amplitude spectrum after low-frequency adjustment.

[0017] In another technical solution, the first low-frequency part in step S21 is specifically that the frequency range of low-frequency adjustment is selected as 0 - 5000Hz;

[0018] The low-frequency adjustment vector in step S21 is output by a trained low-frequency adjustment neural network. Different low-frequency adjustment neural network models are trained for different nasal finals, and 16 low-frequency adjustment neural network models are trained respectively for nasal finals an, ian, uan, üan, en, uen, in, ün, ang, iang, uang, eng, ueng, ing, ong, iong;

[0019] According to the category of the nasalized vowel speech frames determined in step S1, determine the corresponding model from 16 neural network models, input the amplitude spectrum of the nasalized vowel speech frames into the low-frequency adjustment neural network model of the corresponding vowel category, output the low-frequency adjustment vector, multiply the low-frequency adjustment vector by the first low-frequency part of the amplitude spectrum of the nasalized vowel speech frames, and then superimpose the remaining part of the amplitude spectrum of the nasalized vowel speech frames to obtain the amplitude spectrum after low-frequency adjustment.

[0020] In another technical solution, in step 22, obtain the second low-frequency part of the amplitude spectrum of the nasalized vowel speech signal after low-frequency adjustment in step 21, multiply it by the high-frequency reconstruction vector to obtain the reconstructed amplitude spectrum, and then, according to the fundamental frequency, while maintaining the harmonic structure, shift the reconstructed amplitude spectrum to the high-frequency part and superimpose it with the amplitude spectrum of the nasalized vowel speech signal after low-frequency adjustment obtained in step 21, so as to obtain the amplitude spectrum after high-low frequency restoration.

[0021] In another technical solution, the above-mentioned high-frequency reconstruction vector in step S22 is output by a trained high-frequency reconstruction neural network. Different high-frequency reconstruction neural network models are trained for different nasalized vowels. 16 high-frequency reconstruction neural network models are trained respectively for nasalized vowels an, ian, uan, üan, en, uen, in, ün, ang, iang, uang, eng, ueng, ing, ong, iong;

[0022] According to the category of the nasalized vowel speech frames determined in step S1, determine the corresponding model from 16 high-frequency reconstruction neural network models, and input the amplitude spectrum of the nasalized vowel speech signal after low-frequency adjustment of the nasalized vowel speech frames output in step S21 into the high-frequency reconstruction neural network model of the corresponding vowel category to output the high-frequency reconstruction vector.

[0023] In another technical solution, step S1 includes:

[0024] S11, use the VAD endpoint detection algorithm to identify the starting frame of a sentence;

[0025] S12. After identifying the starting frame, determine whether the current frame is a vowel frame or a non-vowel frame.

[0026] S12. Frame by frame, perform classification detection on the vowel speech frames based on the GMM-HMM model to determine whether the vowel speech frame belongs to one of the 16 types of nasalized vowel speech frames containing n or ng:

[0027] Extract 13-dimensional MFCC features for each vowel speech frame as the observation vector of the current frame, obtain the multi-frame MFCC feature vector from the starting frame of this sentence to the end of the current frame as the observation vector sequence, substitute the observation vector sequence into the trained GMM-HMM model, and obtain the result of which nasalized vowel the current frame belongs to;

[0028] S13. When it is detected that the current frame is a frame containing nasal finals n or ng, perform the nasal congestion voice restoration processing described in the foregoing step S2.

[0029] In another technical solution, after detecting the start frame, when it is continuously detected that N frames are all frames containing nasal finals n or ng, perform the nasal congestion voice restoration processing described in step S2 on the current frame and subsequent frames, where N is a positive integer greater than or equal to 3.

[0030] In another technical solution, for the first M nasal final frames after the start frame in a sentence, where M is a positive integer greater than or equal to 5, the low-frequency adjustment vector in step S21 is multiplied by a coefficient to obtain a new low-frequency adjustment vector for nasal congestion voice processing; the high-frequency reconstruction vector in step S22 is multiplied by a coefficient to obtain a new high-frequency reconstruction vector for nasal congestion voice processing; this coefficient is less than 1 and increases frame by frame for the first M frames.

[0031] An experimental teaching system for implementing the foregoing nasal congestion voice signal processing method is also proposed. The system includes: a DSP development board, and an audio acquisition circuit and an audio output circuit.

[0032] The DSP development board includes a processor and an integrated development environment. The DSP development board allows students to download a computer program to the development board for testing and verification. The computer program can implement the foregoing nasal congestion voice signal processing method. The sound collected by the audio acquisition circuit is processed in real time by the DSP development board and directly amplified and output through the audio output circuit, or input into a PC computer for storage.

[0033] An experimental teaching method for the experimental teaching system is also proposed, including:

[0034] The DSP preparation work includes connecting the DSP development board to a computer and configuring the corresponding connection options and parameters, which includes selecting the correct interface connection method, setting serial communication parameters, etc.

[0035] Students open the integrated development environment, create a new project, write signal processing computer program code using the integrated development environment. The computer program code implements the foregoing processing method for Mandarin nasal congestion voice signals. Using the debugging function provided by the development environment, gradually execute the code, view variable and memory status, and perform performance analysis and optimization. After debugging is completed, compile and download the computer program code to the DSP development board for testing and verification. The sound collected by the audio acquisition circuit is processed in real time by the DSP development board and directly amplified and output through the audio output circuit, or input into a PC computer for storage.

[0036] A high-frequency reconstruction vector lookup table and a lookup table for the low-frequency adjustment vector are stored in the DSP development board.

[0037] Technical effect: The contribution of the present invention lies in the use of the recognition of nasal finals containing n and ng in the restoration of cold nasal congestion voice. By combining low-frequency restoration and high-frequency reconstruction, better nasal sound restoration is achieved, and real-time processing is realized through frame-by-frame classification recognition and processing. Description of the drawings

[0038] Figure 1A Is the spectrogram of the normal pronunciation of "si"

[0039] Figure 1B Is the spectrogram of the nasal congestion voice "si"

[0040] Figure 2A Is the spectrogram of the normal pronunciation of "san"

[0041] Figure 2B Is the spectrogram of the nasal congestion voice "san"

[0042] Figure 2C Is the amplitude spectrogram of the normal pronunciation of "san"

[0043] Figure 2D Is the amplitude spectrogram of the nasal congestion voice "san"

[0044] Figure 3 Schematic diagram of the GMM-HMM speech recognition model

[0045] Figure 4 Is the structural diagram of the speech frame in the present invention

[0046] Figure 5 Is the flowchart of the processing method of the first embodiment of the present invention

[0047] Figure 6 Is the flowchart of the processing method of step S2 in the second embodiment of the present invention Detailed implementation manners

[0048]

First embodiment

[0049] A cold and nasal congestion does not affect all types of syllables, but only affects the pronunciation of some types of syllables, that is, those syllables that are pronounced with air passing through the nose.

[0050] Figure 1A Is the spectrogram of the normal pronunciation of "si" in Mandarin, Figure 1B Is the spectrogram of the pronunciation of "si" with a cold and nasal congestion. From the spectrograms of the two, the normal pronunciation of "si" in Mandarin and the pronunciation with a cold and nasal congestion do not differ much, and there is no obvious distinguishable difference in terms of hearing. This is because the pronunciation of "si" in Mandarin includes two parts. The first part is the initial consonant "s (si)", and the second part is the final "i (yi)", where:

[0051] The initial consonant "s (hiss)" belongs to voiceless consonants, which do not involve vocal cord vibration and cavity resonance, and does not require the participation of nasal airflow during the pronunciation process. Nasal congestion has little obvious impact on the pronunciation of the initial consonant "s (hiss)".

[0052] The final vowel "i (yi)" belongs to vowels. A cold with nasal congestion will affect the nasal cavity resonance of the final vowel "i". However, during the pronunciation process of the final vowel "i", most of the airflow flows out through the oral cavity and does not rely too much on the nasal cavity. As a result, nasal congestion also has little obvious impact on the pronunciation of the final vowel "i (yi)". From Figure 1A 、 1B From the spectrogram of the second half of the final vowel "i (yi)", the formant structure of the final vowel "i" only changes slightly.

[0053] According to the record in the paper "Types of Nasal Initial Consonants in Chinese Dialects" (author Li Mingxing) in the Acta Phonologica Sinica, during the pronunciation process of syllables such as m, n, and η in Chinese initial consonants, the frequency of nasal sounds is relatively high. Combining the observation of experimental data, it can be seen that nasal congestion due to a cold has a greater impact on syllables such as m, n, and η, especially the greatest impact on the initial consonant "m". This is because "m" is a closed vowel, and there is no airflow flowing through the oral cavity during pronunciation, and it completely relies on nasal ventilation to achieve pronunciation. When there is nasal congestion, especially when the nasal cavity is completely blocked and one can only breathe through the mouth, since the airflow cannot flow out through the nasal cavity, it will cause the sound of "m" to be extremely difficult to pronounce or even completely inaudible.

[0054] In summary, the degree of influence of nasal congestion due to a cold on different syllables is different. Some syllables are greatly affected, while some are minimally affected or even not affected at all. Based on noticing this phonetic feature, the present invention proposes the following first embodiment.

[0055] Figure 5 It is a flowchart of the processing method of the first embodiment of the present invention. The method includes the following steps:

[0056] Step S1: Collect Mandarin nasal voice signals, detect the vowel voice frames in the nasal voice signals, and perform classification detection on the vowel voice frames to identify nasal vowel voice frames. The nasal vowel voice frames refer to the nasal vowel voice frames containing n and ng, that is, the voice frames of 8 front nasal vowels an, ian, uan, üan, en, uen, in, ün and the voice frames of 8 back nasal vowels ang, iang, uang, eng, ueng, ing, ong, iong. Syllable / phoneme / initial / final classification detection is a mature existing technology in the field of speech recognition. For example, the mixture Gaussian hidden Markov algorithm mentioned in the paper "Design and Implementation of an Embedded Syllable Recognition System" (author: Wang Xiaolong) by Huazhong University of Science and Technology can be used to realize the recognition of vowel syllables. The technology in "Application of Vector Quantization Technology and Hidden Markov Model Method in Vowel Recognition" (author: Wu Jianxiong, Journal of Shanghai Jiao Tong University) can also be used to recognize nasal vowels containing n and ng. The specific method will not be elaborated in this invention.

[0057] Step S2: Perform nasal voice restoration processing on the nasal vowel voice frames, and do not perform processing on other voice frames. In this first embodiment, the present invention realizes that different processing methods should be adopted for different categories of syllables. Syllables with little or no impact are not processed, and only syllables with greater impact (i.e., nasal vowels) are processed. This can not only ensure the restoration effect of nasal voice but also avoid inappropriate processing of syllables with little / no impact, resulting in deteriorated sound quality. In addition, it can also reduce the algorithm complexity, improve the processing efficiency, and the low algorithm complexity also makes it easy for students to understand, thus facilitating teaching.

[0058] The pronunciation of each Chinese character in Mandarin includes an initial part and a final part. From Figure 1A 、 Figure 1B 、 Figure 2A 、 Figure 2B spectrograms, it can be seen that the duration ratio of the initial part is very small, and the duration of the initial part accounts for 1 / 3 or less of the duration of the pronunciation of the entire Chinese character, and its importance is relatively low. Moreover, for voiceless initials such as s, b, d, g, less nasal airflow is used in their pronunciation, so the impact of nasal congestion on these initials is small, and there is no need to process these syllables when performing nasal voice processing.

[0059] The present invention only performs nasal voice restoration processing on the nasal vowel voice frames, that is, only processes the nasal vowel voice frames containing n and ng, that is, the voice frames of 8 front nasal vowels an, ian, uan, üan, en, uen, in, ün and the voice frames of 8 back nasal vowels ang, iang, uang, eng, ueng, ing, ong, iong.

[0060] Although the sound of the initial consonant "m" is most affected by a cold and nasal congestion, since the duration of the initial consonant accounts for a relatively short proportion, its importance is relatively low, and the processing difficulty is relatively high. Therefore, in order to achieve a compromise between the restoration processing effect and the algorithm complexity, the present invention does not process the initial consonant "m".

[0061]

Second Embodiment

[0062] Figure 6 It is a flowchart of the nasal congestion speech restoration processing in step S2, which specifically includes the following steps:

[0063] Step S21: Adjust the low-frequency spectral amplitude spectrum of the nasal vowel speech frame to obtain the amplitude spectrum of the nasal vowel speech signal after low-frequency adjustment;

[0064] Step S22: Reconstruct the high-frequency amplitude spectrum of the nasal vowel speech frame according to the amplitude spectrum of the nasal vowel speech signal after low-frequency adjustment obtained in step S21 to obtain the amplitude spectrum after high-low frequency restoration;

[0065] Step S23: Combine the phase spectrum of the nasal vowel speech frame on the basis of the amplitude spectrum after high-low frequency restoration obtained in step S22 to obtain the restored spectrum, and obtain the restored speech frame time-domain signal according to the restored spectrum;

[0066] The pronunciation of "san" in Mandarin contains the nasal vowel "an". Figure 2A - Figure 2D Shows the spectrogram and amplitude frequency spectrum of the vowel "an" in the speech of "san". Figure 2A Is the spectrogram of the normal pronunciation of "san". Figure 2B Is the spectrogram of the nasal congestion speech of "san". Figure 2C Is the amplitude spectrum diagram of the normal pronunciation of "san". Figure 2D Is the amplitude spectrum diagram of the nasal congestion speech of "san". It can be found by comparison that:

[0067] The normal pronunciation has obvious high-frequency components in the high-frequency band between 6500 - 8000 Hz, and formants are formed in this part (such as Figure 2A 、 Figure 2C ). In the low-frequency band part, the peak position of the spectral amplitude is at the first harmonic position.

[0068] After a cold and nasal congestion, the spectral structure of the vowel an has changed significantly. In the high-frequency part, the high-frequency components between 6500 - 8000 Hz have almost disappeared. The spectral amplitude in the low-frequency part has distorted, and the peak position of the spectral amplitude has changed to the fourth harmonic position.

[0069] It can be seen that compared with normal speech, due to the change of the resonance structure, the amplitude spectrum envelope of the low-frequency component of the nasal congestion speech has distorted, and the high-frequency component section is missing.

[0070] Therefore, in step S2, a method combining the adjustment of the spectral envelope of the low-frequency part and the reconstruction of the high-frequency part is adopted to process the nasal vowels containing n and ng, achieving a good effect of restoring nasal-occluded speech.

[0071] Step S21: Adjust the amplitude spectrum of the low-frequency part of the nasal vowel speech frame to obtain the amplitude spectrum of the nasal vowel speech signal after low-frequency adjustment.

[0072] The low-frequency part is selected as 0 - 5000 Hz, which is called the first low-frequency part. 0 - 5000 Hz contains the main energy of the low-frequency components. In step S21, the first low-frequency part of the amplitude spectrum of the nasal vowel speech frame containing n and ng in the nasal-occluded speech signal is multiplied by the low-frequency adjustment vector Low(1:k) to obtain the adjusted amplitude spectrum, realizing the adjustment of the first low-frequency part.

[0073] The above low-frequency adjustment vector Low(1:k) is output by a trained low-frequency adjustment neural network. Different low-frequency adjustment neural networks are trained for different nasal vowels. For example, the neural network model C_an is trained for the vowel an, the neural network model C_ang is trained for the vowel ang, and so on. 16 low-frequency adjustment neural network models are trained to output low-frequency adjustment vectors, namely C_an, C_ian, C_uan, C_üan, C_en, C_uen, C_in, C_ün, C_ang, C_iang, C_uang, C_eng, C_ueng, C_ing, C_ong, C_iong.

[0074] According to the category of the nasal vowel speech frame (speech frame) determined in step S1 (which nasal vowel belongs to an, ian, uan, üan, en, uen, in, ün, ang, iang, uang, eng, ueng, ing, ong, or iong), the corresponding low-frequency adjustment neural network model (such as C_an) is determined.

[0075] The amplitude spectrum F(1:1024) of the nasal vowel speech frame containing n and ng in the nasal-occluded speech signal is input into the low-frequency adjustment neural network model corresponding to the vowel (such as C_an), and the low-frequency adjustment vector Low(1:k) is output.

[0076] The low-frequency adjustment vector Low(1:k) is multiplied by the first low-frequency part F(1:k) of the amplitude spectrum F(1:1024) of the nasal vowel speech frame, and then combined with the remaining part F((k + 1):1024) of F(1:1024) to obtain the amplitude spectrum F_S21(1:1024) after low-frequency adjustment.

[0077] When training a neural network, a large number of training samples of different people and different nasal finals are collected. The low-frequency amplitude spectrum envelope of the normal nasal final speech frame is divided by the low-frequency amplitude spectrum envelope of the cold and stuffy nose speech frame point by point to obtain the low-frequency adjustment vector samples for training. In one example, the low-frequency amplitude spectrum envelope is the low-frequency amplitude spectrum envelope of the 1-5000 Hz part. The low-frequency adjustment vector is expressed as Low(1:k)=[A1, A2, A3......Ak], where k is the number of FFT frequency points in the first low-frequency part, and A1, A2......Ak are the amplification factors of each frequency point. For example, when the sampling frequency is 16 KHz, the first low-frequency part is selected as 1-5000 Hz, and the number of FFT sampling points is 1024, then k is 320. Taking the above-mentioned low-frequency adjustment vector samples for training as the output of the neural network and the amplitude spectrum envelope of the cold and stuffy nose speech frame as the input of the neural network, 16 low-frequency adjustment neural network models (i.e., C_an, C_ian, C_uan, C_üan, C_en, C_uen, C_in, C_ün, C_ang, C_iang, C_uang, C_eng, C_ueng, C_ing, C_ong, C_iong) are trained to be applicable to cold voices of different nasal finals.

[0078] In this embodiment, different types of nasal sounds output different low-frequency adjustment vectors through different types of neural network models, making the algorithm show good adaptability.

[0079] Step S2 further includes step S22 of reconstructing the high-frequency amplitude spectrum of the nasal final speech frame;

[0080] Since the high-frequency part is almost completely missing and cannot be restored by envelope adjustment. From Figure 2B 、 Figure 2D it can be seen that the high-frequency missing part is mainly above 6 KHz. Since the nasal finals have good harmonic characteristics, the high-frequency part has similar harmonic characteristics to the first low-frequency part. According to this characteristic, the present invention adopts the method of compensating the high-frequency with the low-frequency to achieve high-frequency restoration and reconstruction.

[0081] In step 22, the second low-frequency part of the amplitude spectrum of the nasal final speech signal after low-frequency adjustment obtained in step 21 is multiplied by the high-frequency reconstruction vector to obtain the reconstructed amplitude spectrum. Then, according to the fundamental frequency, while maintaining the harmonic structure, the reconstructed amplitude spectrum is shifted to the high-frequency part and superimposed with the amplitude spectrum F_S21(1:1024) of the nasal final speech signal after low-frequency adjustment obtained in step 21 to obtain the amplitude spectrum after high and low-frequency restoration.

[0082] To achieve the correct reconstruction of the high-frequency harmonic structure without destroying the harmonic structure in the high-frequency part, it is necessary to find the harmonic frequency points in the high-frequency range and then determine the reconstruction frequency band. For example, when the fundamental frequency is 300 Hz, at 7000 Hz, a 1000-Hz range centered approximately on this harmonic frequency point of 7000 Hz is selected as the part of the high-frequency band to be reconstructed, that is, a total of 1000 Hz from 6600 Hz to 7600 Hz is used as the high-frequency part to be reconstructed. The starting frequency of 6600 Hz must be an integer multiple of the fundamental frequency to conform to the harmonic law. Thus, the amplitude spectrum of 6600 Hz - 7600 Hz is restored using the 0 - 1000 Hz part (i.e., the second low-frequency part) of the amplitude spectrum of the nasal vowel speech signal after low-frequency adjustment.

[0083] The high-frequency reconstruction vector High(1:n) is a series of amplification factors High = [A1, A2, A3......An], where the value of n is related to the frequency range, sampling rate, and number of FFT points. For example, when the size of the reconstruction frequency range is selected as 1000 Hz and the sampling rate is 16 KHz, then n = 64 at this time. Correspondingly, the amplitude of each frequency point in the 0 - 1000 Hz frequency range of the amplitude spectrum F_S21(1:1024) of the nasal vowel speech signal after low-frequency adjustment obtained in step 21, that is, the amplitude spectrum F_S21(1:64) of the 1 - 64 frequency points in the amplitude spectrum, is multiplied by the corresponding amplification factor in the high-frequency reconstruction vector High(1:64) respectively, so as to obtain the reconstructed amplitude spectrum F_S21_low(1:n). The reconstructed amplitude spectrum is migrated to the high-frequency reconstruction frequency band of 6400 - 7400 Hz and superimposed with the amplitude spectrum F_S21(1:1024) of the nasal vowel speech signal after low-frequency adjustment obtained in step 21, so as to obtain the amplitude spectrum after high and low frequency restoration.

[0084] The above high-frequency reconstruction vector High[1:n] is output through a trained high-frequency reconstruction neural network. Different high-frequency reconstruction neural networks are trained for different nasal vowels. For example, the neural network model N_an is trained for the vowel an, the neural network model N_ang is trained for the vowel ang, and so on. 16 high-frequency reconstruction neural network models are trained to output high-frequency reconstruction vectors, that is, N_an, N_ian, N_uan, N_üan, N_en, N_uen, N_in, N_ün, N_ang, N_iang, N_uang, N_eng, N_ueng, N_ing, N_ong, N_iong.

[0085] Determine the corresponding high-frequency reconstruction neural network model (such as N_an) according to the category of the nasal vowel speech frames (speech frames) determined in step S1 (which one of an, ian, uan, üan, en, uen, in, ün, ang, iang, uang, eng, ueng, ing, ong, iong the nasal vowel belongs to).

[0086] Input the amplitude spectrum F(1:1024) of the nasal vowel speech frames containing n and ng in the nasal congestion speech signal into the neural network model corresponding to the vowel (such as N_an), and output the high-frequency reconstruction vector High(1:n).

[0087] When training the neural network, collect a large number of training samples of different people and different vowels, and divide the amplitude spectrum envelope within the reconstruction range of the normal nasal vowel speech frames (for example, the frequency band of 6400 - 7400Hz with a bandwidth of 1000Hz near 7000Hz determined according to the fundamental frequency of 300Hz) point by point by the low-frequency amplitude spectrum envelope of 1 - 1000Hz of the cold nasal congestion speech frames, so as to obtain the training high-frequency reconstruction vector samples. Use the above training high-frequency reconstruction vector samples as the output of the neural network, and use the amplitude spectrum envelope of the normal speech frames as the input of the neural network to train 16 high-frequency reconstruction neural network models (that is, N_an, N_ian, N_uan, N_üan, N_en, N_uen, N_in, N_ün, N_ang, N_iang, N_uang, N_eng, N_ueng, N_ing, N_ong, N_iong), so as to be applicable to the cold voice of different nasal vowels.

[0088] In this embodiment, different types of nasal sounds output different high-frequency reconstruction vectors after passing through different types of neural network models, making the algorithm show good adaptability.

[0089] Step S23: Combine the phase spectrum of the nasal vowel speech frames on the basis of the amplitude spectrum after high and low frequency restoration to obtain the restored spectrum, and obtain the restored speech frame time-domain signal according to the restored spectrum.

[0090] The result output by step S22 is the complete amplitude spectrum after high-frequency reconstruction and low-frequency restoration. In order to restore the complete spectrum, the phase spectrum is also needed. The present invention directly combines the phase spectrum of the nasal vowel speech frames with the complete amplitude spectrum, so as to obtain the complete spectrum after high-frequency reconstruction and low-frequency restoration, and then the time-domain signal of the normal speech frame can be restored through IFFT.

[0091] The present invention adopts the high-frequency reconstruction processing of adjusting the low-frequency spectrum amplitude spectrum and compensating the high frequency with the low frequency, and realizes good restoration of the nasal congestion speech.

[0092]

Third Embodiment

[0093] In this embodiment, the consideration of real-time processing is added. In some scenarios of nasal voice restoration, the requirement for real-time processing is relatively high. For example, in the cases of teachers' lectures, radio announcers' live broadcasts, and streamers' live broadcasts, how to restore the normal voice signal in real time is another technical problem to be solved by the present invention.

[0094] The prerequisite for real-time processing is the frame-by-frame real-time detection of nasal compound vowels containing n and ng. In order to complete the frame-by-frame real-time detection of nasal compound vowels, in step S1 of the present invention, a speech recognition technology based on the GMM-HMM model is adopted to complete the classification detection of nasal compound vowels.

[0095] As Figure 3 shown, the speech recognition technology based on the GMM-HMM model uses an Acoustic Model + Lexicon Model + Language Model to achieve speech recognition. Among them, the Acoustic Model can complete the function of judging which nasal compound vowel the current frame belongs to. The Lexicon Model is based on the pronunciation of phonemes combined into words, and the Language Model predicts the corresponding text based on the pronunciation.

[0096] The specific steps to complete the classification detection of nasal compound vowels by using the speech recognition technology based on the GMM-HMM model are as follows:

[0097] S11, use the VAD endpoint detection algorithm to identify the starting frame of a sentence;

[0098] S12. After identifying the starting frame, judge whether the current frame is a compound vowel frame or a non-compound vowel frame.

[0099] Since the present invention only processes nasal compound vowels containing n and ng, such compound vowels (such as Figure 4 the V1, V2, V3 frames shown in...) have higher harmonic characteristics compared to initial consonants (such as Figure 4 the C1 frame to the Cn frames shown in...). Therefore, the present invention first uses the harmonic characteristic to judge whether the current frame is a compound vowel speech frame or a non-compound vowel frame.

[0100] S12. Perform frame-by-frame classification detection of compound vowel speech frames based on the GMM-HMM model to determine whether the compound vowel speech frame belongs to one of the 16 types of nasal compound vowel speech frames containing n and ng, specifically including:

[0101] Extract 13-dimensional MFCC features for each vowel phonetic frame (10-20 ms) as the observation vector of the current frame, and obtain a sequence of multi-frame MFCC feature vectors from the start frame to the end of the current frame of this sentence as the observation vector sequence. Substitute this observation vector sequence into the trained GMM-HMM model to obtain the result of which nasal vowel the current frame belongs to (refer to the thesis of Xidian University, "Research and Implementation of a Speaker-Independent Chinese Continuous Digital Speech Recognition System_Gao Chaohuang.pdf").

[0102] As shown in the figure above, when it is detected that the V1 frame is a vowel, the observation vectors of the C1, C2...Cn, V1 frames of this section of activated speech are collectively substituted into the GMM-HMM model to determine whether the V1 frame is a frame containing the n or ng vowels; when it is detected that the V2 frame is a vowel frame, the observation vectors of the C1, C2...Cn, V1, V2 frames of this section of activated speech are collectively substituted into the GMM-HMM model again to determine whether the V1 frame is a frame containing the n or ng vowels. By this classification and judgment method of substituting frames into the model one by one, the real-time performance of syllable detection is improved.

[0103] S13. When it is detected that the current frame is a frame containing the n or ng nasal vowels, perform the nasal congestion speech restoration process on the nasal vowel phonetic frame in the aforementioned step S2.

[0104]

Fourth Embodiment

[0105] While the above embodiments improve the real-time performance, it may reduce the accuracy of syllable judgment. For the HMM hidden Markov model, the more observation values, the more accurate the classification. Therefore, for the nasal vowels in the first few frames, there may be a problem of misclassification, resulting in the restoration process being performed on frames that are not nasal vowel frames containing n or ng, thus causing an auditory degradation problem.

[0106] In order to reduce the auditory degradation caused by inaccurate syllable detection. The present invention adopts the following two methods.

[0107] First, trailing detection of syllable detection is adopted. Since the duration of the vowel continuously occupies a relatively long time in each word, therefore, only when in a sentence, multiple consecutive frames (for example, N frames, N is preferably a positive integer greater than or equal to 3) containing the n or ng nasal vowels are detected after the start frame. For example, when it is detected that 10 consecutive vowel frames V1, V2, V3.......V10 are all nasal vowel frames, it is considered that the 11th frame is the real nasal vowel phonetic frame, and the restoration process in step S2 is performed on the 11th frame and subsequent nasal vowel phonetic frames, thereby improving the processing accuracy.

[0108] Secondly, a fade-in processing technique is adopted. For the first few nasal vowel frames after the starting frame in a sentence (such as the first M frames, where M is preferably a positive integer greater than or equal to 5), the judgment accuracy of nasal vowels is not high. For these speech frames of nasal vowels, a recovery processing method with a low intervention degree is adopted, that is, the low-frequency adjustment vector and the high-frequency reconstruction vector in step S2 need to be multiplied by a coefficient, and then multiplied by the first and second low-frequency parts of the amplitude spectrum of the nasal vowel speech frame. This coefficient increases frame by frame. For example, for the first nasal vowel frame detected at the starting frame, this coefficient is 0.2, and for the second nasal vowel frame, this coefficient is 0.4, and so on. If it is continuously determined that there are n or ng nasal vowels in 5 consecutive frames, the processing intensity is increased to 1. If a non-nasal vowel frame appears within these 5 frames, step S2 is not executed and the nasal vowel processing is stopped. This way of gradually increasing the processing degree avoids the deterioration of sound quality caused by misprocessing when the judgment in the first few frames is inaccurate.

[0109]

Fifth Embodiment

[0110] Colleges and universities' computer majors often offer courses such as multimedia technology and signal processing. The speech signal processing algorithm can be used as teaching experiment content for students to learn and practice the principles of digital signal processing.

[0111] For this reason, this patent also proposes an experimental teaching system and experimental method for nasal sound processing technology. The system includes a DSP development board, as well as an audio acquisition circuit and an audio output circuit. The preparatory work of the DSP includes connecting the DSP development board to the computer and configuring the corresponding connection options and parameters. This includes selecting the correct interface connection method, setting serial port communication parameters, etc. Among them, the DSP development board includes a processor and an integrated development environment. Students open the integrated development environment, create a new project, use the integrated development environment to write the signal processing algorithm code in the present invention to implement the processing method of the cold and stuffy nose voice in the foregoing solution, use the debugging function provided by the development environment to gradually execute the code, view variable and memory states, and perform performance analysis and optimization. After debugging is completed, compile and download the code to the DSP development board for testing and verification. The sound collected by the audio acquisition circuit is processed in real time through the DSP development board and directly amplified and output through the audio output circuit, or input into a PC computer for storage.

[0112] In order to reduce the programming teaching difficulty of the neural network in the present invention, a look-up table method is adopted to simulate the neural network model. That is, according to the training samples, a look-up table is prefabricated in advance. By using the clustering method, the correspondence between the amplitude spectrum F_Normal(1:1024) of the nasalized vowel speech signal under normal non-stuffy nose conditions and the high-frequency reconstruction vector is prefabricated in the look-up table; then, the correspondence between the amplitude spectrum F_Abnormal(1:1024) of the nasalized vowel speech signal under stuffy nose conditions and the low-frequency adjustment vector is established. For example, the following look-up table is established in advance and stored in the development board:

[0113] Look-up table for high-frequency reconstruction vector of vowel an

[0114] Note: The second low-frequency part is 1000Hz, sampling rate = 16KHz, n = 64)

[0115] 1 F_Normal_1(1:1024) High_AN_1[1:n] 2 F_Normal_2(1:1024) High_AN_2[1:n] 3 F_Normal_3(1:1024) High_AN_3[1:n] 4 F_Normal_4(1:1024) High_AN_4[1:n] ... ...... ......

[0116] Look-up table for low-frequency adjustment vector of vowel an

[0117] Note: The first low-frequency part is 1 - 5000Hz, sampling rate = 16KH, k = 320)

[0118] 1 F_Abnormal_1(1:1024) Low_AN_1[1:k] 2 F_Abnormal_2(1:1024) Low_AN_2[1:k] 3 F_Abnormal_3(1:1024) Low_AN_3[1:k] 4 F_Abnormal_4(1:1024) Low_AN_4[1:k] ... ...... ......

[0119] And so on, look-up tables for the corresponding high-frequency reconstruction vectors and low-frequency adjustment vectors are established for each of the 16 nasal finals respectively. When processing the nasal-obstructed speech signal, when it is determined that the current speech frame is the "an" syllable in Mandarin, look up and compare in the look-up table of the low-frequency adjustment vector for the "an" syllable to find the amplitude spectrum sample with the closest Euclidean distance to the amplitude spectrum of the current nasal-obstructed frame containing the "n" or "ng" nasal finals, such as F_Abnormal_3(1:1024), so as to output the corresponding low-frequency adjusted amplitude spectrum Low_AN_3(1:320) for adjusting the first low-frequency part of the amplitude spectrum. Then, using the amplitude spectrum of the nasal final speech signal F_S21(1:1024) after low-frequency adjustment, look up and compare in the look-up table of the high-frequency reconstruction vector for the "an" syllable to find the amplitude spectrum sample with the closest Euclidean distance to the amplitude spectrum of the current nasal final speech signal after low-frequency adjustment, such as F_Normal_5(1:1024), so as to output the corresponding high-frequency reconstruction vector High_AN_5(1:64). Multiply the second low-frequency part F_S21(1:64) of the amplitude spectrum of the nasal final speech signal F_S21 after low-frequency adjustment by the high-frequency reconstruction vector High_AN_5(1:64) to obtain the reconstructed amplitude spectrum F_S21_low(1:n). When the fundamental frequency is 300 Hz, shift the reconstructed amplitude spectrum to the harmonic frequency spectrum range near 7000 Hz, i.e., the range of 6600 - 7600 Hz, and superimpose it on the amplitude spectrum of the nasal final speech signal F_S21(1:1024) after low-frequency adjustment, so as to obtain the amplitude spectrum after high and low frequency recovery. Then, combine the amplitude spectrum with the phase spectrum of the nasal-obstructed speech frame to restore the original spectrum.

[0120] This embodiment can achieve the purpose of teaching DSP programming development technology and speech signal processing technology, and using the look-up table can also reduce the learning difficulty of students.

Claims

1. A method for processing Mandarin nasal congestion speech signals, characterized in that: The following steps are involved: Step S1, collecting a Mandarin nasal congestion speech signal, detecting a vowel speech frame in the nasal congestion speech signal, and classifying and detecting the vowel speech frame to determine whether the vowel speech frame belongs to one of 16 types of nasal vowel speech frames including n and ng, thereby detecting a nasal vowel speech frame and determining the category of the nasal vowel speech frame, that is, determining which category of the 16 types of nasal vowel speech frames the nasal vowel speech frame belongs to; The above-mentioned 16 types of nasal finals speech frames including n and ng refer to the speech frames of the 8 front nasal finals an, ian, uan, üan, en, uen, in, ün and the speech frames of the 8 back nasal finals ang, iang, uang, eng, ueng, ing, ong, iong; Step S2, performing nasal congestion speech recovery processing on the nasal vowel speech frame, and not performing processing on other speech frames; The nasal congestion speech recovery process in step S2 specifically includes the following steps: Step S21, adjusting the low-frequency spectrum amplitude spectrum of the nasal vowel speech frame to obtain the low-frequency adjusted nasal vowel speech signal amplitude spectrum; Step S22, reconstructing the high-frequency amplitude spectrum of the nasal vowel speech frame according to the amplitude spectrum of the nasal vowel speech signal after the low-frequency adjustment obtained in step S21, and obtaining the amplitude spectrum after high and low frequency restoration; Step S23, combining the phase spectrum of the nasal vowel speech frame based on the high and low frequency restored amplitude spectrum obtained in step S22 to obtain a restored spectrum, and obtaining a restored speech frame time domain signal according to the restored spectrum.

2. The method according to claim 1, characterized in that The step S21 of adjusting the low-frequency spectrum amplitude spectrum of the nasal vowel speech frame to obtain the low-frequency adjusted nasal vowel speech signal amplitude spectrum specifically includes: The first low-frequency part of the amplitude spectrum of the nasal vowels n and ng in the nasal speech signal is multiplied by the low-frequency adjustment vector to obtain the adjusted spectrum, and then the rest of the amplitude spectrum of the nasal vowel speech frame is superimposed to obtain the low-frequency adjusted amplitude spectrum.

3. The method according to claim 2, characterized in that The first low-frequency part in step S21 is specifically to select the frequency range of the low-frequency adjustment to be 0-5000 Hz; The low-frequency adjustment vector in step S21 is outputted through a trained low-frequency adjustment neural network, and different low-frequency adjustment neural network models are trained for different nasal finals, and 16 low-frequency adjustment neural network models are trained for the nasal finals an, ian, uan, üan, en, uen, in, ün, ang, iang, uang, eng, ueng, ing, ong, iong respectively; According to the category of the nasal final speech frame determined in step S1, the corresponding model is determined from 16 neural network models, the amplitude spectrum of the nasal final speech frame is input into the low-frequency adjustment neural network model of the corresponding final category, and a low-frequency adjustment vector is output. The low-frequency adjustment vector is multiplied by the first low-frequency part of the amplitude spectrum of the nasal final speech frame, and then the rest of the amplitude spectrum of the nasal final speech frame is superimposed to obtain the amplitude spectrum after low-frequency adjustment.

4. The method according to claim 1, characterized in that In step S22, the second low-frequency part of the amplitude spectrum of the nasal final speech signal after low-frequency adjustment obtained in step S21 is multiplied by the high-frequency reconstruction vector to obtain the reconstructed amplitude spectrum, and then according to the fundamental frequency, while maintaining the harmonic structure, the reconstructed amplitude spectrum is shifted to the high-frequency part and superimposed with the amplitude spectrum of the nasal final speech signal after low-frequency adjustment obtained in step S21, so as to obtain the amplitude spectrum after high and low frequency restoration.

5. The method according to claim 4, characterized in that The high-frequency reconstruction vector in step S22 is outputted through a trained high-frequency reconstruction neural network, and different high-frequency reconstruction neural network models are trained for different nasal finals, and 16 high-frequency reconstruction neural network models are trained for the nasal finals an, ian, uan, üan, en, uen, in, ün, ang, iang, uang, eng, ueng, ing, ong, iong respectively; According to the category of the nasal final speech frame determined in step S1, the corresponding model is determined from 16 high-frequency reconstruction neural network models, and the low-frequency adjusted nasal final speech signal amplitude spectrum of the nasal final speech frame output in step S21 is input into the high-frequency reconstruction neural network model of the corresponding final category, and a high-frequency reconstruction vector is output.

6. The method according to any one of claims 2 to 5, characterized in that: The step S1 comprises: S11, using VAD endpoint detection algorithm to identify the starting frame of a sentence; S12, after identifying the start frame, determining whether the current frame is a final frame or a non-final frame; S12, performing classification detection based on the GMM-HMM model on the vowel speech frame frame by frame to determine whether the vowel speech frame belongs to one of the 16 types of nasal vowel speech frames including n and ng: Extract 13-dimensional MFCC features for each final speech frame as the observation vector of the current frame, obtain the multi-frame MFCC feature vectors from the start frame of the sentence to the end of the current frame as the observation vector sequence, substitute the observation vector sequence into the trained GMM-HMM model, and obtain the result of which nasal final the current frame belongs to; S13. When it is detected that the current frame is a frame containing nasal finals n or ng, the nasal congestion speech recovery process described in the aforementioned step S2 is performed.

7. The method according to claim 6, characterized in that After the start frame is detected, when N frames are continuously detected to contain nasal finals of n and ng, the nasal speech recovery process described in step S2 is performed on the current frame and subsequent frames, where N is a positive integer greater than or equal to 3.

8. The method according to claim 6, characterized in that For the first M nasal final frames after the starting frame in a sentence, M is a positive integer greater than or equal to 5. The low-frequency adjustment vector in step S21 is multiplied by a coefficient to obtain a new low-frequency adjustment vector for nasal congestion speech processing; the high-frequency reconstruction vector in step S22 is multiplied by a coefficient to obtain a new high-frequency reconstruction vector for nasal congestion speech processing; the coefficient is less than 1 and increases frame by frame for the first M frames.

9. An experimental teaching system for implementing the nasal congestion speech signal processing method according to any one of claims 1 to 8, the system comprising: DSP development board, as well as audio acquisition circuit and audio output circuit; The DSP development board includes a processor and an integrated development environment. The DSP development board allows students to download computer programs to the development board for testing and verification. The computer program can implement the nasal congestion speech signal processing method as described in any one of claims 1 to 8. The sound collected by the audio acquisition circuit is processed in real time by the DSP development board, directly amplified and output through the audio output circuit, or input into a PC for storage.

10. An experimental teaching method for the experimental teaching system according to claim 9, the method comprising: DSP preparation includes connecting the DSP development board to the computer and configuring the corresponding connection options and parameters, including selecting the correct interface connection method, setting serial port communication parameters, etc. The student opens the integrated development environment, creates a new project, and uses the integrated development environment to write a signal processing computer program code, which implements the method for processing Mandarin nasal congestion speech signals described in any one of claims 1 to 8. The student uses the debugging function provided by the development environment to gradually execute the code, view variables and memory status, and perform performance analysis and optimization. After debugging is completed, the student compiles and downloads the computer program code to the DSP development board for testing and verification. The sound collected by the audio acquisition circuit is processed in real time by the DSP development board, directly amplified and output through the audio output circuit, or input into a PC for storage; A high-frequency reconstruction vector lookup table and a low-frequency adjustment vector lookup table are stored in the DSP development board.

Citation Information

Patent Citations

  • Parameter synthesis method and device and perception category measuring method and device of alveolar and velar nasal vowels

    CN105825847A

  • Using method of judging equipment for Chinese nasal final pronunciation barrier patients

    CN107452370A