A language identification method based on language discriminative features of phonemes

By constructing the phoneme distinguishing characteristics and augmenting phoneme set of TIMIT phoneme sets, combining n-gram meta method and Resnet model, the problem of low recognition rate of traditional language recognition methods in real acoustic environments is solved, and high-accurate language recognition is achieved.

CN115019775BActive Publication Date: 2025-06-20KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210096847.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-06-20
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

Traditional language recognition methods are difficult to obtain ideal recognition results in real acoustic environments because the acoustic features include environment and speaker characteristics, which reduces the proportion of language distinctive information.

Method used

By constructing the phoneme distinctive features of the TIMIT phoneme set, a phoneme recognizer is constructed using GMM score determination, a multilingual frame phoneme probability vector is identified, the phoneme set is expanded, and the phoneme vector and probability vector are derived in units of speech segments, the phoneme posterior probability vector of n-gram element method is constructed, and the features are converted into grayscale maps, and language recognition is used using Resnet.

Benefits of technology

The phoneme-based language distinctive characteristics are realized, the accuracy and recognition rate of language recognition are improved, and the number of language expansion is strong.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019775B_ABST
    Figure CN115019775B_ABST
Patent Text Reader

Abstract

The present invention relates to a language recognition method based on phoneme-based language discriminative features, belonging to the technical field of audio signal processing. First, the present invention extracts a phoneme set from the TIMIT dataset, constructs phoneme phonetic discriminative features for the phoneme set, trains and tests a phoneme recognizer using the phoneme phonetic discriminative features, and outputs the frame-level phoneme probability vector of the audio; then, it obtains multi-language corpora from the LibriVox audio database, expands the phoneme set extracted from the TIMIT dataset for the multi-language corpora, and outputs the frame phoneme probability features of the short-time complete semantic speech segments of the languages; finally, it constructs the phoneme probability features of the speech segments according to the frame phoneme probability features of different languages output by the phoneme recognizer, and further constructs the language discriminative features of the speech segments. The present invention can perform language recognition in a classical two-dimensional convolutional neural network and obtain language recognition results with a high recognition rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a language recognition method based on phoneme-based language discriminative features, and belongs to the technical field of audio signal processing. Background Art

[0002] Traditional language recognition methods usually perform speech dimensionality reduction on the frame level of speech, and extract a series of acoustic features of audio, including MFCC features, SDC features, GFCC features, BFCC features, PLP features, LPCC features, i-vector features, etc. The acoustic feature parameters of audio contain rich temporal information of speech and are widely used in most speech and acoustic pattern recognition problems including language recognition.

[0003] As a speech pattern recognition problem, language recognition uses a series of acoustic feature parameters of audio, derivative parameters of acoustic feature parameters of audio, fusion feature parameters of acoustic feature parameters of audio, etc. as the mainstream recognition features. Although the mainstream features have achieved good results in combination with some classification system models under specific corpora, these features are difficult to obtain ideal recognition results in real acoustic environments because these acoustic features contain a lot of environmental features and speaker features, which greatly reduces the proportion of language discriminative information features in the acoustic features.

[0004] Traditional phoneme-based language recognition generally adopts a method divided into three modules: a phoneme recognition module, a phoneme language discriminative feature construction module, and a language information classification module. Among them, the phoneme recognition module directly trains the phoneme set in the way of a neural network, and constructs a phoneme recognizer by using the trained model; this recognition method often adopts the way of inputting speech acoustic features, and the result of phoneme recognition will be affected by factors such as speakers and channels.

[0005] The phoneme language discriminative feature construction module constructs phoneme phonetic features by using phoneme-like elements with coherent acoustic characteristics to replace phonetic phonemes; compared with phonetic phonemes, the speech recognition rate of the phoneme-like elements measured by minimizing the distortion of language segments is greatly reduced.

[0006] The language information classification module, the convolutional neural network based on the two-dimensional speech feature reconstruction of speech features is more superior in classification performance than the Gaussian mixture model GMM, but this two-dimensional speech feature reconstruction is only applied to the speech spectrogram or the two-dimensional spectrogram of speech acoustic features and has not been used in audio speech features. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a language recognition method based on phoneme-based language discriminative features to solve the above problems.

[0008] The technical solution of the present invention is as follows: A language recognition method based on phoneme-based language discriminative features constructs phoneme discriminative features of the TIMIT phoneme set, constructs a phoneme recognizer that outputs a phoneme probability feature vector of the output frame through GMM scoring, further uses the phoneme recognizer to recognize the frame phoneme probability vectors of multiple languages, expands out-of-TIMIT-set phonemes based on the information entropy of the output multi-language frame phoneme probability vectors, and derives the phoneme vector and phoneme probability vector of the speech segment in units of speech segments. The phoneme posterior probability vector combinations of the n-gram method of the speech segment are obtained using the phoneme vector and phoneme probability vector of the speech segment respectively as phoneme discriminative information, constructs multi-language language discriminative features based on the phonetic features of phonemes, and finally converts the constructed phoneme language discriminative information into a grayscale image, and uses the classic residual neural network Resnet for language recognition to obtain a language recognition result with a high recognition rate.

[0009] The specific steps are as follows:

[0010] Stepl: First, obtain LibriVox audio data, and then use short-time spectral entropy, short-time energy, and short-time zero-crossing rate parameters to perform complete semantic short-time speech segment segmentation.

[0011] Step2: Read in the TIMIT dataset and extract the phoneme set according to the manual marking information in the TIMIT dataset.

[0012] Step3: Construct phoneme discriminative features based on the fundamental frequency information and formant frequency information of phonemes in the phoneme set.

[0013] Step4: Use the GMM model to train and test the phoneme discriminative features and construct a frame-level phoneme recognizer.

[0014] Step5: Preprocess and frame the complete semantic short-time speech segment, and then input the frame signal into the phoneme recognizer to output the frame phoneme probability vectors of different languages for the complete semantic short-time speech segment.

[0015] Step6: Based on the TIMIT phoneme set, judge and expand the multi-language phoneme set according to the information entropy of the phoneme probabilities of different language speech frames.

[0016] Step7: First, obtain the phoneme vector and phoneme probability vector of the speech segment according to the frame phoneme probability vector of the speech segment, then obtain the phoneme probability vector of the n-gram method of the speech segment according to the phoneme vector and phoneme probability vector of the speech segment, and finally use the phoneme posterior probability vector combination of the n-gram method of the speech segment as phoneme discriminative information to complete the construction of the phoneme language discriminative features of the speech segment.

[0017] Step 8: First, convert the phoneme language discriminative features of the two-dimensional speech segment into a grayscale image, then use the classic Residual Neural Network (Resnet) for language identification, and finally obtain a language identification result with a high recognition rate.

[0018] The specific content of Step 1 is as follows:

[0019] Step 1.1: Determine an ideal silent segment in the speech segment using the short-time energy threshold, short-time zero-crossing rate threshold, and short-time spectral entropy threshold of microframes with a frame length of 0.025 s and a frame shift of 0.001 s.

[0020] Step 1.2: Determine the syllable boundaries of the speech based on the short-time energy and short-time zero-crossing rate of the found silent segment.

[0021] Step 1.3: Eliminate the silent segment from the audio according to the boundary and perform non-destructive segmentation with a specified duration.

[0022] The specific content of Step 6 is as follows: Input the multi-language speech frame signal set into the GMM phoneme recognizer, calculate the information entropy of the frame phoneme probability vector based on the obtained phoneme probability vector, and determine and expand the multi-language to fit the multi-language phoneme according to the information entropy.

[0023] The specific content of Step 7 is as follows:

[0024] Step 7.1: Calculate the average value of the maximum value pi of the frame phoneme probability vector P(O) of multiple frames bundled by phonemes, and use it as the probability value of the corresponding phoneme in the speech segment phoneme probability vector.

[0025] Step 7.2: Calculate the phoneme probability vector of the n-gram method for the speech segment.

[0026] Step 7.3: Calculate the posterior probability [P l of the l (l = 1, 2, 3) vowel phonemes.

[0027] Step 7.4: Concatenate [P l (l = 1, 2, 3) into a two-dimensional matrix [P] of q×3·q as the phoneme phonetic language discriminative feature of the speech segment.

[0028] The beneficial effects of the present invention are as follows: The present invention constructs a phoneme posterior probability feature with reasonable physical explanations and multi-language discriminability based on phoneme phonetic features. The phoneme expansion based on the TIMIT phoneme set enables the language discriminative feature to have strong scalability in the number of languages. This feature can be used for language identification in a classic two-dimensional convolutional neural network to obtain a language identification result with a high recognition rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is the flowchart of the present invention;

[0030] Figure 2 is the detailed flowchart of the present invention;

[0031] Figure 3 is the flowchart of endpoint detection of the present invention;

[0032] Figure 4 is the speech waveform and spectrogram of short-time segmentation of the complete semantics of audio;

[0033] Figure 5 is the diagram of the language identification classification model of the present invention. Specific embodiments

[0034] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0035] Example 1: As Figure 1-2 shown, a language identification method based on phoneme-based language discriminative features, the specific steps are as follows:

[0036] Step1: Acquisition of multi-language audio data:

[0037] Download multi-language audio files from the LibriVox global free public audiobook dataset. The audio files include English, French, German, Italian, and Spanish, and the duration of each language is at least 20 hours. Use the AudioSegment in the python tool pydub package to uniformly convert the audio sampling frequency to 16000Hz and uniformly transcode the audio file format to mono wav format.

[0038] Step2: Audio preprocessing:

[0039] Including removing the trend term and removing the DC component.

[0040] Step2.1: Removing the trend term:

[0041] The trend term refers to the zero line of the audio signal deviating from the baseline over time. This phenomenon is often caused by unstable factors in the input and output systems of the audio and environmental interference around the microphone. The existence of the trend term causes a linear or slowly varying error in the time series of the audio, deforming the signal autocorrelation function and power spectrum.

[0042] The speech sampling data is {x k}(k = 1, 2,..., n), where n is the number of speech signal sampling data points. The information of removing the trend term is as shown in Equation (1):

[0043]

[0044] In Equation (1), is the speech signal xk The trend term signal of the m-th order least squares fitting. When m equals 0, it represents a DC trend term; when m equals 1, it is a linear trend term; when m is greater than 1, it is a curve trend term.

[0045] Step2.2: Signal DC component elimination and amplitude normalization:

[0046] For the convenience of retrieving reference non-speech segments in subsequent speech processing and setting the thresholds for endpoint detection, all audio needs to be uniformly normalized so that the amplitude of the speech signal is between -1 and 1. The signal DC component elimination is as shown in Equation (2), and the signal amplitude normalization is as shown in Equation (3):

[0047]

[0048]

[0049] Step3: Complete semantic short-time speech segment segmentation:

[0050] Speech segmentation needs to consider the different contexts, grammars, and semantics of different languages, identify the syllable or word boundaries in the speech file, and segment them according to the specified duration.

[0051] Step3.1: Finding the reference non-speech segment:

[0052] Before performing endpoint detection on the speech, a silent segment of the speech (with the non-speech segment length being 0.1 seconds) needs to be found to calculate the short-time energy threshold, short-time zero-crossing rate threshold, and short-time spectral entropy threshold for the double-threshold method of endpoint detection.

[0053] The method used to find the silent segment is as follows: The average frame short-time energy of 0.3 seconds of continuous speech is less than 0.001, and then 0.1 seconds is removed from the beginning and end. When the above method fails to find it, one level lower is searched (the average frame short-time energy is less than 0.002). Since the long-term continuous information of the speech is not considered in the analysis of the speech frames here, the length of the speech analysis frames taken can be as small as possible, and a small frame shift form is used for frame division. Here, the frame length is taken as: 0.025s, and the frame shift is: 0.001s.

[0054] Step3.2: Double-threshold endpoint detection:

[0055] To eliminate the influence of the language environment noise, super short-time energy noise search and double-threshold endpoint detection of short-time zero-crossing rate + short-time energy are used, and the endpoint detection results are used to mark the frame-level speech boundaries in the speech segment.

[0056] Calculate the short-time energy formula of the i-th frame speech signal y i (n) as (L is the frame length, f n is the number of frames):

[0057]

[0058] Define the short-time average zero-crossing rate formula as (where L is the frame length and is the number of frames):

[0059]

[0060] Where:

[0061]

[0062] At a sampling frequency of 16000 Hz, set the frame segmentation parameters: frame length 0.025 s, frame shift 0.01 s; calculate the short-time energy threshold and short-time zero-crossing rate threshold according to the silent segments found above; take the maximum silent length: 15 sampling points, and the minimum speech length 20 sampling points.

[0063] Status judgment:

[0064] Status status, 0 means silent, 1 means may enter the speech segment, 2 means definitely enter the speech segment, 3 means the end of the speech segment, and the process is as Figure 3 shown.

[0065] Step3.3: Complete semantic speech segmentation by frame-by-frame method:

[0066] Extract the speech segments from the frame labels of the detected speech boundaries, and use the frame-by-frame recursive splicing method to judge whether a single speech segment is longer than the speech cutting length. If it is longer, discard it; judge that after the last splicing, a single segment is not too long, and normally splice the next segment; judge that after the last splicing, a single segment is too long, then do not splice this segment and output the speech.

[0067] The waveform and spectrogram of the complete semantic segmentation after removing the silent segment for 3 seconds are as Figure 4 shown. It can be seen from the figure that the speech has been segmented into short-time speech, the speechless segments of the short-time speech have been removed, and the speech has not been damaged. The semantics of the speech are complete and can be used for subsequent short-time processing of speech analysis.

[0068] Step4: Frame the segmented speech segments:

[0069] In order to obtain frame-level audio phonetic features, the speech needs to be framed. It is considered that the framed speech is steady and continuous. For the selection of the frame length and frame shift of the speech frame, generally a frame length of 0.025 seconds and a frame shift of 0.01 seconds are selected to ensure the steady timing information of the speech.

[0070] Step5: Build a phoneme recognizer:

[0071] The most important part of the phoneme recognizer is the phonetic distinctive feature construction module for phonemes. Phonemes are extracted from the TIMIT dataset, and the phonetic features of the phonemes are obtained to train and test the recognition effect of the phoneme recognizer. Finally, the constructed phoneme recognizer is applied to the phoneme recognition of the speech frames in the language recognition dataset to obtain the phoneme probability vector of the speech frames in the language recognition dataset.

[0072] Step5.1: Phoneme set acquisition:

[0073] The TIMIT dataset specifies the specific positions of each phoneme that makes up different speech segments in the speech. All phonemes are extracted according to the audio phoneme labeling document and classified into a phoneme set for the training and testing of the phoneme recognizer.

[0074] The constructed phoneme set is divided into a training set and a test set. The phonetic feature extraction preprocessing, frame segmentation, linear prediction, pitch frequency calculation, and formant calculation are performed on the training set and the test set respectively.

[0075] Step5.2: Phonetic feature extraction preprocessing for phonemes:

[0076] In order to improve the accuracy of phonetic feature extraction for phonemes, more strict endpoint detection needs to be performed on the obtained audio of a single phoneme. The method of energy-entropy ratio is used to remove the audio segments at the head and tail of the speech that do not have vocal cord periodic oscillation and vocal tract resonance.

[0077] Implementation of energy-entropy ratio endpoint detection:

[0078] Calculate the energy of a speech signal with a frame length of N (the energy of the speech signal is greater than the noise energy) as shown in Equation (8):

[0079]

[0080] LE i =log(1 + AMP i / a) Equation (8)

[0081] In Equations (7) and (8), a is a constant. When a takes a relatively large value, when the AMP i amplitude changes drastically, the improved energy LE i changes gently. Appropriately selecting a helps to distinguish between silence and voiceless sounds.

[0082] A speech signal with a frame length of N After FFT transformation, the energy spectrum of the k-th spectral line frequency component f k is Y i (k); the i-th frame signal of the speech The k-th spectral line frequency component f kThe spectral probability density is as shown in Equation (9):

[0083]

[0084] The i-th frame signal of speech The spectral entropy is as shown in Equation (10):

[0085]

[0086] The i-th frame signal of speech The spectral entropy is as shown in Equation (11):

[0087]

[0088] Only one T1 threshold is used for judgment. Judge whether the energy entropy ratio is greater than T1. The part greater than T1 is considered as the candidate value of the effective segment. Then judge whether the length is greater than the minimum length. In the preprocessing of phoneme recognition in the present invention, the minimum length is taken as L min = 10.

[0089] Step5.3: Fundamental frequency calculation:

[0090] To reduce the interference of formants, a preprocessing filter of 60 - 500 Hz is selected. Since the speech signal is not sensitive to phase, an elliptic IIR filter with small computational complexity can be considered.

[0091] For the speech segment, windowing is performed frame by frame (frame length is N): The window function uses the Hamming window, as shown in Equation (12):

[0092]

[0093] The signal x of the i-th frame of speech with frame length N i (n) is windowed as shown in Equation (13):

[0094]

[0095] Calculate the LPC prediction error:

[0096] LPC prediction adopts the method of establishing an all-pole model: The model inputs a periodic pulse or a white noise sequence u(n), and the output is a deterministic signal or a random signal sequence The relationship between them can be expressed by the difference equation (14), where G represents the gain of the p-order all-pole model.

[0097]

[0098] Calculate the predicted value of a frame of speech signal The linear prediction coefficients of the frame of speech signal are obtained by the autocorrelation method as: ​

[0099]

[0100] Find the LPC cepstrum:

[0101] The linear prediction error is as shown in Equation (16):

[0102]

[0103] The LPC cepstrum is as shown in Equation (17). Search in the LPC cepstrum. Find the maximum value in the interval between the maximum and minimum of the pitch period, and that maximum value is the pitch period:

[0104] Pe i = IFFT(2·log 10 (||FFT(e i (n))||)) Equation (17)

[0105] Step5.4: Obtain the formant frequencies:

[0106] To obtain the formants, the same linear prediction method as for obtaining the pitch period is used to obtain the linear prediction coefficients. Different from the preprocessing for obtaining the pitch frequency which uses filtering, the preprocessing for obtaining the formants is pre-emphasis, as shown in Equation (18). The larger the pre-emphasis coefficient a, the more significant the pre-emphasis, and the more significant the reduction in the influence of the glottal pulse; pre-emphasis suppresses the amplitude of the base spectral line, reduces the interference of the fundamental frequency on formant detection, is beneficial to formant detection, leaving only the vocal tract part, which is convenient for analyzing the vocal tract parameters:

[0107]

[0108] The short-time autocorrelation function of the window signal is:

[0109]

[0110] The frequency response of the speech LPC system function model is related to the short-time Fourier transform of the speech through the short-time autocorrelation function. The Fourier transform of the short-time autocorrelation function is equal to the square of the amplitude spectrum of the short-time Fourier transform of the signal.

[0111]

[0112] Comparing the content after the minus sign with the form of the short-time autocorrelation function, it is consistent, reflecting the formant frequency components.

[0113] The energy of the error function is obtained by Parseval's theorem. Combining with the linear gain spectrum, it is known that the square of the linear prediction spectrum implies frequency weighting. The places where the square of the signal spectrum amplitude is large are weighted more for frequency than the places where the square of the signal spectrum amplitude is small.

[0114] Perform weighting on the linear prediction spectrum once again to improve the accuracy of formant extraction:

[0115]

[0116] Based on the frequency response of the speech LPC system function model, the short-time Fourier transform of speech, and the correlation relationship of the short-time autocorrelation function of speech, it can be deduced that at the resonant frequency, it satisfies Increase the discrimination between the energy of formant frequencies and non-formant frequencies.

[0117] Let z -1 =exp(-j2πf / f s ), then the power spectrum P(f) is Equation (22):

[0118]

[0119] The complex roots of the polynomial of the prediction error filter can accurately represent the center frequency and bandwidth of the formants, corresponding to the cascaded steady-state form of the vocal tract transfer function model:

[0120] is an arbitrary complex root, and its conjugate value is also a root. Let the formant frequency corresponding to zi be F i , and the 3dB bandwidth be B i , Equation (24) can be obtained from Equation (23):

[0121]

[0122]

[0123] The p-order linear prediction coefficients correspond to p peaks, and there are p / / 2 peaks that meet the conditions. In order to correctly locate the formants and improve the accuracy of formant recognition, an exhaustive method is proposed to find p / / 2 formant frequencies and formant bandwidths. At this time, the formant frequencies and formant bandwidths cannot be in one-to-one correspondence. Here, the bubble sort indexing method is adopted to label the complex roots with indexes, and finally, the formant frequencies and formant bandwidths with the same indexes corresponding to the complex root labels are obtained.

[0124] Up to this point, the accuracy of formant extraction is relatively sensitive to the linear prediction order, and variable-order LPC is proposed.

[0125] The roots of the linear prediction polynomial are equal to the roots of the non-resonant peaks (root amplitude less than 0.9) plus the roots of the resonant peaks (complex poles closer to the unit circle: root amplitude greater than 0.9), which is equivalent to the roots of the resonant peaks plus the roots corresponding to the radiation model plus the roots corresponding to the glottal pulse shape plus the roots corresponding to other factors of the transmission effect; in order to find out which resonant peak it is, this invention adopts the root-finding method here to set the root amplitude threshold to improve the accuracy of resonant peak identification.

[0126] The first p + 1 values of the autocorrelation function of the speech signal and the autocorrelation function of the impulse response corresponding to the speech production system function are equal; if p is large enough, the frequency response of the all-pole model of the speech system function can approximate the frequency response of the short-time Fourier transform of the signal with an arbitrarily small error (the error is mainly in the region with lower spectral amplitude); p also reflects the smoothness of the spectrum of the prediction filter. The larger p is, the greater the fitting degree is, and the smaller the smoothness of the corresponding prediction filter spectrum is. The LPC order p intuitively reflects the number of linear prediction coefficients of a section of speech (p), and further reflects the number of vocal tract resonance and anti-resonance frequency points (p / 2). The larger p is, the higher the distinguishability of the resonant peaks is. The resonant peak frequencies are screened for the local maxima of the LPC spectral envelope that meet the conditions through the frequency range and bandwidth range (resonant peak frequencies are greater than 150 Hz and less than half of the sampling frequency, and the bandwidth is less than 700 Hz); when the number of roots that meet the conditions is not enough to represent the required number of resonant peaks, the order of the LPC needs to be increased, and in order to reduce the computational amount, variable-order LPC is often used to determine the number of resonances. When p takes a large value, the all-pole model can also be used to fit unvoiced sounds.

[0127] The principle of selecting p: On the premise of retaining the resonant peaks and the basic spectral shape, reduce the value of p to remove most of the spectral features related to the excitation. This invention adopts the method of root amplitude threshold feedback here to realize the root finding of variable-order LPC linear prediction.

[0128] All the roots that meet the conditions are found in this linear prediction. Finally, only the obtained resonant peaks need to be sorted in ascending order, and the resonant peaks that meet the conditions are found.

[0129] The aligned resonant peak frequencies are used as the within-class interval of the resonant peaks. Different resonant peaks are used as different classes. Fisher discrimination is used to smooth the resonant peaks, and finally the bandwidth is used as the discrimination threshold to eliminate the singular point values and output the resonant peaks.

[0130] Step5.5: Construction of phoneme distinguishability features:

[0131] The formant information of phonemes reflects the response of the vocal tract corresponding to the phonemes. Combining it with the pitch can represent the distinctive information of phonemes. Therefore, we can construct phoneme frame-level features, including parameters such as the fundamental period corresponding to the phoneme, the first formant corresponding to the phoneme, the bandwidth of the first formant corresponding to the phoneme, the second formant corresponding to the phoneme, the bandwidth of the second formant corresponding to the phoneme, the third formant corresponding to the phoneme, and the bandwidth of the third formant corresponding to the phoneme.

[0132] Each frame of a phoneme corresponds to a phoneme phonetic feature vector. This vector is a 1-dimensional vector containing 7 elements. This kind of distinctive feature can not only well distinguish phoneme features but also greatly reduce the computational complexity with the advantage of the minimum dimension.

[0133] Step5.6: Training and testing of the phoneme recognizer:

[0134] Input the frame-level phoneme distinctive features into the GMM model for training and testing to generate the scoring situation of each feature frame under the current model. Output the scores of each frame corresponding to the phoneme set in the training model. Through the summary of the scores, output the probability vector of the i-th frame phoneme (the phoneme set is: O = {o1, o2,... o k ) as shown in Equation (25):

[0135] p(phoneme) = [p1, p2,... p k Equation (25)

[0136] After obtaining the frame phoneme probability vector, select the phoneme with the highest score for judging the correct rate of phoneme recognition. Constructing a high-correct-rate phoneme recognizer is the guarantee for language identification applications.

[0137] Step6: Construction of phoneme phonetic language distinctive features for speech segments:

[0138] The basis for constructing phoneme phonetic language distinctive features is that the phoneme distribution in language expression of different languages is unique, which is reflected in the statistical probability of phonemes and also in the posterior probability of phoneme arrangement distribution. Therefore, mainly based on the statistical probability of phonemes and the posterior probability of phoneme arrangement distribution, construct phoneme language distinctive features.

[0139] Step6.1: Expansion of the multi-language phoneme set for the TIMIT phoneme set:

[0140] The TIMIT phoneme set consists of 6,300 sentences from 8 major dialect regions in the United States. This dataset fully considers the diversity of pronunciation and downward compatibility, so a phoneme representation containing 52 phonemes, 6 closures, and 5 identifiers is introduced. Its advantage is that an accurately labeled phoneme set can be extracted from this dataset, and the gender information, speaker, and speaking region information in the phoneme set are relatively sufficient. Its disadvantage is that this dataset is limited to English pronunciation and there will be deviations when fitting other languages. To eliminate the fitting deviation and improve the accuracy of phoneme fitting for different languages, the phoneme set needs to be extended, and the extended phoneme set is q.

[0141] Expansion processing:

[0142] After obtaining the frame phoneme probability vector, calculate the information entropy of the frame phoneme probability vector. For the source space formula (26), the obtained information entropy is formula (27):

[0143]

[0144]

[0145] In the phoneme probability vectors obtained for languages other than English, for each phoneme o i the corresponding probability p i is close to When it indicates that the language phoneme corresponding to this frame cannot be fitted by the unextended phoneme set, so set the threshold to determine H(O)≥α, perform phoneme expansion when. And mark the expanded phonemes to mark the language information of this frame.

[0146] Step6.2: Construction of phoneme posterior probability features for speech segments:

[0147] The phoneme posterior probability features of the speech segment are determined by the frame phoneme probability vector of S6.1. The determination method includes bundling frames with the same phoneme in the speech segment, deriving the phoneme posterior probability of the speech segment, and combining the phoneme posterior probability features of the speech segment.

[0148] Bundling frames with the same phoneme in the speech segment:

[0149] Step6.2.1: Find the maximum value p of the frame phoneme probability vector P(O) i ;

[0150] Step6.2.2: Index o in the phoneme set O according to the i mark in p i ; i ;

[0151] Step6.2.3: Bundle the frames with consecutive identical phonemes o i and consider these frames to correspond to one phoneme phq ;

[0152] Step6.2.4: Obtain the maximum value p of the frame phoneme probability vector P(O) of multiple frames with phoneme bundling i The average value of which is pm q ;

[0153] Step6.2.5: Obtain the phoneme vector ph of the speech segment and the phoneme probability vector pm of the speech segment.

[0154] Step6.2.6: Adopt the n-gram method. According to step5, obtain the single-phoneme phoneme vector ph1 of the speech segment and the single-phoneme probability vector pm1 of the speech segment, the two-phoneme phoneme vector ph2 of the speech segment and the two-phoneme probability vector pm2 of the speech segment, and the single-phoneme phoneme vector ph3 of the speech segment and the single-phoneme probability vector pm3 of the speech segment.

[0155] The posterior probability [P l of the speech segment containing k phonemes and l (l = 1, 2, 3) vowels is derived as shown in Equation (28):

[0156]

[0157] Where: For each [P l holds true.

[0158] Step6.3: Construction of the phonetic language discrimination features of the speech segment:

[0159] The combination [P] of the posterior probabilities of the speech segment phonemes is expressed as Equation (29), which is a two-dimensional matrix of q×3·q:

[0160] [P] = {[P1], [P1], [P3]} Equation (29)

[0161] The constructed [P] is the phonetic language discrimination feature of the speech segment.

[0162] Step7: Language identification model:

[0163] The language identification classification model can obtain a relatively high language identification rate by using a two-dimensional convolutional neural network. This invention adopts the classic residual neural network Resnet. The input feature map is the grayscale map of the combination [P] of the posterior probabilities of the speech segment phonemes, with a size of q*3·q. The identification model is as Figure 5 shown.

[0164] Step8: First, convert the two-dimensional phonetic language discrimination feature of the speech segment into a grayscale map, then adopt the classic residual neural network Resnet for language identification, and finally obtain the language identification result with a relatively high identification rate.

[0165] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A language identification method based on phoneme-based language discriminative features, characterized in that: Step1: First, obtain the LibriVox audio data, and then use short-time spectral entropy, short-time energy, and short-time zero-crossing rate parameters to segment short speech segments with complete semantics; Step2: Read in the TIMIT dataset and extract the phoneme set according to the manual marking information in the TIMIT dataset; Step3: Construct discriminative features of phonemes based on the fundamental frequency information and formant frequency information of phonemes in the phoneme set; Step4: Use the GMM model to train and test the discriminative features of phonemes, and construct a frame-level phoneme recognizer; Step5: Preprocess and frame the short speech segments with complete semantics, and then input the frame signals into the phoneme recognizer to output the frame phoneme probability vectors of short speech segments with complete semantics in different languages; Step6: Based on the TIMIT phoneme set, judge and expand the multi-language phoneme set according to the information entropy of the phoneme probabilities of speech frames in different languages; Step7: First, calculate the phoneme vector and phoneme probability vector of the speech segment according to the frame phoneme probability vector of the speech segment, then calculate the phoneme probability vector of the n-gram method of the speech segment according to the phoneme vector and phoneme probability vector of the speech segment, and finally use the combination of the posterior probability vectors of the n-gram method of the speech segment as the discriminative information of phonemes to complete the construction of the discriminative features of phoneme languages of the speech segment; Step8: First, convert the two-dimensional discriminative features of phoneme languages of the speech segment into a grayscale image, then use the classic residual neural network Resnet for language recognition, and finally obtain a language recognition result with a high recognition rate.

2. The language identification method based on phoneme-based language discriminative features according to claim 1, characterized in that, The specific content of Step1 is as follows: Step1.1: Determine an ideal silent segment in the speech segment by using the short-time energy threshold, short-time zero-crossing rate threshold, and short-time spectral entropy threshold of micro-frames with a frame length of 0.025s and a frame shift of 0.001s; Step1.2: Determine the syllable boundaries of the speech according to the short-time energy and short-time zero-crossing rate of the found silent segment; Step1.3: Eliminate the silent segment from the audio according to the boundary and perform non-destructive segmentation with a specified duration.

3. The language identification method based on phoneme-based language discriminative features according to claim 1, characterized in that, The specific content of Step6 is as follows: Input the multi-language speech frame signal set into the GMM phoneme recognizer, calculate the information entropy of the frame phoneme probability vector according to the obtained phoneme probability vector, and judge and expand the multi-language fitting multi-language phonemes according to the information entropy.

4. The language identification method based on phoneme-based language discriminative features according to claim 1, characterized in that, The specific content of Step7 is as follows: Step 7.1: Find the maximum value p of the frame phoneme probability vectors P(O) of multiple frames with phoneme bundling, and take its average value as the probability value of the corresponding phoneme in the phoneme probability vector of the speech segment; i ​ Step7.2: Calculate the phoneme probability vector of the n-gram method of the speech segment; Step7.3: Obtain the posterior probabilities of l (l = 1, 2, 3) vowel elements [P l ; Step7.4: Concatenate [P l (l = 1, 2, 3) into a two-dimensional matrix [P] of q×3·q as the phonetic language discriminative feature of the speech segment.

Citation Information

Patent Citations

  • System and method for detecting language voice frequency

    CN104681036A

  • Discriminative feature extraction method applied to language identification

    CN106297769A