Cross-language speaker recognition method and system based on voice fusion features

By fusion of pitch period and MFCC features and combining the Teager energy operator, the speech fusion features are extracted, and the problem of degradation of recognition performance among different languages ​​by cross-language speaker recognition system is solved, achieving higher recognition accuracy and robustness.

CN120108401AActive Publication Date: 2025-06-06GLOBAL TONE COMM TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510303490.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-06
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing cross-lingual speaker recognition system has deteriorated recognition performance among different languages, especially in bilingual or multilingual environments, and it is difficult to effectively distinguish the characteristics of different speakers.

Method used

By combining features such as pitch period (Pitch) and Mel frequency cepspectral coefficient (MFCC), combined with the Teager energy operator (TEO), more stable speech fusion features are extracted, thereby improving the accuracy of cross-language speaker recognition.

Benefits of technology

It significantly improves the accuracy of cross-lingual speaker recognition, enhances the robustness and adaptability of the system, maintains high recognition accuracy in multilingual environments, and overcomes the problem of degradation in recognition performance caused by language mismatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108401A_ABST
    Figure CN120108401A_ABST
Patent Text Reader

Abstract

The invention provides a cross-language speaker recognition method and system based on voice fusion features, and relates to the technical field of voice process.The method comprises the steps that band-pass filtering is conducted on original voice signals, and voiced sound segments are extracted and integrated into new voice segments; extracting a pitch period and an MFCC feature in the new speech segment, and connecting in series to form a fusion feature vector; inputting multilingual speech segments of a plurality of testers, extracting fusion features, and training by using K-means clustering and an EM algorithm to obtain a GMM model of each speaker; inputting a test voice segment, extracting fusion features, calculating a likelihood probability score with the GMM model, comparing the score with a preset threshold, and confirming or rejecting the identity of a tester. According to the method, more stable features can be extracted from the voice signal, so that the accuracy of cross-language speaker recognition is improved, the situation of language mismatch of the speaker is effectively handled, the recognition performance is improved, and the method has high adaptability especially in a bilingual or multilingual application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a cross-language speaker recognition method and system based on speech fusion features. Background Art

[0002] Speaker recognition systems have extensive and important applications in defense, finance and other fields. When data is sufficient and high-quality, the system performs well; however, when the registration voice and the test voice do not match, the system performance will be severely degraded. A serious mismatch is language mismatch, that is, the registration voice and the test voice of the same speaker use different languages, which usually occurs in bilingual or multilingual people. Considering that bilingual populations are common not only in my country but also around the world, cross-language speaker verification has a wide range of applications.

[0003] When the registration speech and the test speech are in different languages, the performance of the speaker verification method will be severely degraded. The reason can be largely attributed to the different distribution of acoustic features among different languages. Among them, the difficulty of cross-language speaker recognition lies mainly in the special language factors contained in different languages, such as phonemes, tones, and the tension and relaxation of the vocal organs during pronunciation. To address this problem, it is urgent to propose a method based on the fusion of MFCC parameters based on human auditory features and fundamental pitch parameters related to human vocal tract information. The fundamental pitch period is related to the speaker's vocal organs, pronunciation habits, and text, but in a statistical sense, the distribution of the fundamental pitch period of different speakers is different, which can be used to distinguish different speakers. Summary of the invention

[0004] In view of this, the purpose of the present invention is to propose a cross-language speaker recognition method and system based on speech fusion features, aiming to extract more stable features from speech signals by fusing features such as pitch and Mel-Frequency Cepstral Coefficients (MFCC) and combining Teager Energy Operator (TEO), thereby improving the accuracy of cross-language speaker recognition. This method can effectively deal with the situation of speaker language mismatch and improve recognition performance, especially in bilingual or multilingual application scenarios, and has strong adaptability.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] Based on the above objectives, in a first aspect, the present invention provides a cross-language speaker recognition method based on speech fusion features, comprising the following steps:

[0007] Perform bandpass filtering on the original speech signal, extract the voiced segments and integrate them into new speech segments;

[0008] Extract the pitch period and MFCC features in the new speech segment and concatenate them to form a fusion feature vector;

[0009] Input multilingual speech segments of multiple testers, extract fusion features, and use K-means clustering and EM algorithm to train the Gaussian Mixture Model (GMM) of each speaker;

[0010] Input the test speech segment, extract the fusion features, calculate the likelihood probability score with the GMM model, compare it with the preset threshold, and confirm or reject the identity of the tester.

[0011] As a further solution of the present invention, bandpass filtering is performed on the original speech signal to extract the voiced segments and integrate them into new speech segments, including the following steps:

[0012] The original speech signal is filtered to remove 50Hz DC noise and high-frequency noise greater than 4kHz, and a bandpass filter with a frequency band of 60Hz to 4kHz is used for bandpass filtering;

[0013] Perform Fast Fourier Transform (FFT) on the filtered signal and calculate the Teager Energy Operator (TEO);

[0014] The voiced segments are extracted according to the preset threshold, the speech frame numbers corresponding to the voiced segments are obtained, and the voiced segments are connected in series to form new speech segments as the original speech for extracting fusion features.

[0015] As a further solution of the present invention, the pitch period and MFCC features in the new speech segment are extracted and connected in series to form a fusion feature vector, including the following steps:

[0016] Extract and integrate voiced segments by calculating the TEO value of speech signals;

[0017] Extract the pitch period and MFCC features in the new speech segment;

[0018] The obtained 1-dimensional pitch period and MFCC are connected in series to form a new fusion feature vector.

[0019] As a further solution of the present invention, training a GMM model for each speaker includes the following steps:

[0020] Input multilingual speech segments of multiple testers and extract fusion features from the speech segments of each speaker;

[0021] Use K-means clustering method to cluster multilingual speech data and obtain the initial GMM parameters for each speaker;

[0022] The GMM model of each speaker's speech segment is obtained by training the Expectation-Maximization Algorithm (EM) to fit the speech characteristics of each speaker.

[0023] As a further solution of the present invention, confirming or rejecting the identity of the tester includes the following steps:

[0024] In the testing phase, the tester's speech segment is input and the fusion features of the tester are extracted;

[0025] Calculate the likelihood probability score between the test subject's speech features and the GMM model of the known speaker model;

[0026] The calculated likelihood probability score is compared with the set threshold. When the score is greater than the threshold, the test subject is considered a known speaker; when the score is less than the threshold, the test subject is rejected as a known speaker and the recognition result is output.

[0027] In a second aspect, the present invention further provides a cross-language speaker recognition system based on speech fusion features, comprising the following components:

[0028] Voiced segment extraction module: used to extract and integrate voiced segments;

[0029] Fusion feature extraction module: used to extract pitch period and MFCC features and form a fusion feature vector;

[0030] GMM training module: used to train Gaussian mixture model (GMM);

[0031] Speaker verification module: used to calculate the likelihood probability score between the test speech and the GMM model, and compare it with the preset threshold to confirm or reject the identity of the tester.

[0032] As a further solution of the present invention, the voiced segment extraction module includes:

[0033] Bandpass filter: used to perform 60Hz~4kHz bandpass filtering on the original speech signal;

[0034] FFT calculation unit: used to perform fast Fourier transform (FFT) on the filtered signal;

[0035] TEO calculation unit: used to calculate the Teager Energy Operator (TEO);

[0036] The voiced segment integration unit is used to extract the voiced segments according to the preset threshold value, and integrate the voiced segments into a new speech segment in series.

[0037] As a further solution of the present invention, the fusion feature extraction module includes:

[0038] Pitch period extraction unit: used to extract the pitch period from the new speech segment;

[0039] MFCC extraction unit: used to extract MFCC features from new speech segments;

[0040] Feature fusion unit: used to concatenate the pitch period and MFCC features to form a fused feature vector.

[0041] As a further solution of the present invention, the GMM training module includes:

[0042] Fusion feature extraction unit: used to input multilingual speech segments of multiple testers and extract the fusion features of each speaker;

[0043] K-means clustering unit: used to initialize GMM parameters;

[0044] EM algorithm unit: used to train the GMM model for each speaker.

[0045] As a further solution of the present invention, the speaker confirmation module includes:

[0046] Fusion feature extraction unit: used to input the test speech segment and extract the fusion feature;

[0047] Likelihood probability calculation unit: used to calculate the likelihood probability score of the test speech and the GMM model;

[0048] Threshold comparison unit: used to compare with the preset threshold to confirm or reject the identity of the tester.

[0049] Compared with the prior art, the cross-language speaker recognition method and system based on speech fusion features proposed in the present invention has the following beneficial effects:

[0050] 1. The present invention can effectively reduce the feature differences between different languages ​​and significantly improve the accuracy of cross-language speaker recognition by integrating multiple features including pitch period and Mel frequency cepstral coefficient (MFCC). In a bilingual or multilingual environment, the system can better adapt to the pronunciation habits and acoustic characteristics of different languages, overcoming the problem of reduced recognition performance caused by language mismatch.

[0051] 2. When faced with factors such as pronunciation differences, phoneme changes, and tone differences in different languages, the system of the present invention integrates the pitch period and MFCC features to make the recognition system more robust to changes in speech signals. In particular, when there are large differences in pronunciation habits across languages, it can still maintain a high recognition accuracy, significantly improving the stability and reliability of the system.

[0052] 3. This method uses an efficient Teager energy operator and MFCC calculation method in the feature extraction process, which reduces the amount of calculation and improves the real-time performance of the entire process. This feature enables this method to play a role in application scenarios that require rapid response, and is suitable for various fields with high real-time requirements, such as security monitoring, financial transactions, and intelligent security. By combining the pitch period with the MFCC features and using the Gaussian mixture model (GMM) for modeling, the present invention can more accurately capture the personalized characteristics of the speaker, especially in a multilingual environment, and improve the accuracy of speaker recognition. This method can achieve a balance between different languages, avoiding the problem of a sudden drop in recognition performance when switching between languages ​​in traditional methods.

[0053] 4. The present invention optimizes and adjusts the parameter settings in the feature extraction process through experiments, so that the system can adaptively adjust the feature extraction method according to different voice data, further improving the adaptability and recognition effect in a cross-language environment. In addition, the dynamic adjustment of the threshold enables the system to optimize recognition decisions based on real-time data, improving the flexibility and accuracy of recognition. The present invention is suitable for identity authentication and security monitoring in a multilingual environment, and is particularly suitable for applications in bilingual or multilingual areas. At the same time, the present invention has a wide range of application fields, especially in fields with high security requirements such as national defense, finance, and public security, and can provide users with efficient and accurate speaker recognition technical support.

[0054] In summary, the cross-language speaker recognition method and system based on speech fusion features provided by the present invention have excellent cross-language adaptability, robustness, real-time and high-precision recognition capabilities, can effectively improve the speaker recognition performance in a multilingual environment, adapt to a wider range of application scenarios, and have strong market competitiveness and application value.

[0055] These and other aspects of the present application will be more concise and understandable in the following description of the embodiments. It should be understood that the above general description and the following detailed description are only exemplary and explanatory and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following briefly introduces the drawings required for use in the exemplary embodiments or related technical descriptions. The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0057] Figure 1 The present invention is a flowchart of a method for cross-language speaker recognition based on speech fusion features according to an embodiment of the present invention.

[0058] Figure 2 The present invention is a fusion feature flow chart of a cross-language speaker recognition method based on speech fusion features in an embodiment of the present invention.

[0059] Figure 3 The present invention is a flowchart of the testing phase in a cross-language speaker recognition method based on speech fusion features according to an embodiment of the present invention. DETAILED DESCRIPTION

[0060] Below, the present application is further described in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.

[0061] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in combination with specific embodiments and with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0062] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are intended to distinguish two non-identical entities or non-identical parameters with the same name. It can be seen that "first" and "second" are only for the convenience of expression and should not be understood as limitations on the embodiments of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, other steps or units inherent to a process, method, system, product or device that includes a series of steps or units.

[0063] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0064] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0065] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0066] The difficulty of cross-language speaker recognition lies mainly in the special language factors contained in different languages, such as phonemes, tones, and the tension and relaxation of the vocal organs during pronunciation. To address this problem, the present invention proposes a cross-language speaker recognition method based on speech fusion features. The pitch period is related to the speaker's vocal organs, pronunciation habits, and text, but in a statistical sense, the pitch period distribution of different speakers is different, which can be used to distinguish different speakers. The Teager energy operator has excellent time resolution, is simple and fast, and can track the time-varying part of the speech signal in real time. Through a more accurate instruction table, a more accurate threshold h can be obtained.

[0067] See also Figure 1 As shown, an embodiment of the present invention provides a cross-language speaker recognition method based on speech fusion features, the method comprising the following steps:

[0068] Step 1: Bandpass filter the original speech signal, extract the voiced segments and integrate them into new speech segments;

[0069] Step 2: extract the pitch period and MFCC features in the new speech segment and connect them in series to form a fusion feature vector;

[0070] Step 3: Input multilingual speech segments of multiple testers, extract fusion features, and use K-means clustering and EM algorithm to train the Gaussian Mixture Model (GMM) of each speaker;

[0071] Step 4: Input the test speech segment, extract the fusion features, calculate the likelihood probability score with the GMM model, compare it with the preset threshold, and confirm or reject the identity of the tester.

[0072] In the present invention, voiced segments are extracted from the original speech, the voiced segments are connected in series to form a new speech segment, the pitch period and MFCC in the new speech segment are extracted, and the obtained pitch period and MFCC are used to form a fusion feature F t ; Calculate F t and the GMM model parameters λ of the assumed speaker i Likelihood Probability Score And compare it with the threshold. If it is greater than the threshold, the tester is accepted as the speaker, otherwise it is rejected and the confirmation result is output. Among them, the GMM model parameter λ i The way to obtain is:

[0073] Input N testers' multilingual voices;

[0074] Calculate and extract fusion features of each speaker's multilingual speech;

[0075] The initialization parameter λ of the speaker model is obtained according to the K-means clustering method0

[0076] According to the initialization parameter λ 0 , the GMM model with the mixing degree M = 128 for each speaker is calculated using the EM algorithm, G L ={λ 1 , …, λ i , …, λ N},λ i is the GMM model parameter of the i-th tester, output G L ; G L is a set of GMM model parameters for all testers.

[0077] In an embodiment of the present invention, in step 1, extracting and integrating voiced segments, bandpass filtering is performed on the original speech signal s(n), and the voiced segments are extracted and integrated into new speech segments, including the following steps:

[0078] Step 1: Filter out 50Hz DC noise and high-frequency noise greater than 4kHz from the original speech signal, and perform bandpass filtering with a bandpass filter H(n) in the 60Hz to 4kHz band; wherein:

[0079] x(n)=s(n)×H(n), n=1,2,…,N;

[0080] Where x(n) is a signal with a limited frequency band; s(n) is the original speech signal; H(n) is a bandpass filter with a frequency band of 60Hz-4kHz, which is used to perform bandpass filtering on the original speech signal s(n) to filter out the noise outside the specified frequency range.

[0081] Step 2: Perform Fast Fourier Transform (FFT) on the filtered signal and calculate the Teager Energy Operator (TEO); wherein, perform Fast Fourier Transform on the band-limited signal X(n) to obtain:

[0082]

[0083] Where X(k) is the complex value corresponding to the kth frequency point in the frequency domain after the fast Fourier transform (FFT) of the limited frequency band signal X(n), which represents the amplitude and phase information of the frequency point; N is the total number of sampling points of the discrete signal X(n); j is the imaginary unit, which is used to represent complex number operations in Fourier transform; k is the frequency point number in the frequency domain, ranging from 1 to N, corresponding to different frequency components;

[0084] When calculating the Teager energy operator, the Teager energy operator can effectively track the instantaneous energy of the signal and is used for time-varying signal analysis of speech. For a finite-band signal x(n), the operator can be approximately expressed as:

[0085] φ[x(n)]=[x(n)] 2 -x(n+1)x(n-1)

[0086] Where, φ[x(n)] is the calculation result of the Teager energy operator of the finite-band signal X(n), which is used to track the instantaneous energy of the signal; x(n+1) is the sampled value of the signal X(n) at time n+1; x(n-1) is the sampled value of the signal X(n) at time n-1;

[0087] TEO extracts the envelope by calculating the three adjacent sampling points of the detected waveform. It has excellent time resolution and is simple and fast. It can track the time-varying part of the speech signal in real time. The calculation formula of this step is:

[0088] t(k)=φ[X(k)]

[0089] Where t(k) is the result obtained by calculating the Teager energy operator on the frequency domain signal X(k) after fast Fourier transformation, which is used for subsequent operations such as extracting voiced segments.

[0090] Step 3: extract the voiced segments according to the preset threshold, obtain the speech frame numbers corresponding to the voiced segments, and connect the voiced segments in series to form new speech segments as the original speech for extracting fusion features.

[0091] Among them, the voiced segment v(k) is selected. Since the voiced sound t(k) is greater than the unvoiced sound and silence, a threshold h can be used to extract the voiced segment:

[0092]

[0093] The voiced speech threshold is h=0.02, and the speech frame number corresponding to the voiced speech segment is taken. The voiced speech segments are connected in series and integrated into a new speech segment as the original speech for extracting fusion features.

[0094] In the embodiment, in step 2, see Figure 2 As shown, the fundamental pitch period and MFCC features in the new speech segment are extracted and connected in series to form a fusion feature vector, including the following steps:

[0095] Extract and integrate voiced segments by calculating the TEO value of speech signals;

[0096] Extract the pitch period and MFCC features in the new speech segment;

[0097] The obtained 1-dimensional pitch period and MFCC are connected in series to form a new fusion feature vector.

[0098] In step 3, the GMM model of each speaker is trained, including the following steps:

[0099] Input multilingual speech segments of multiple testers and extract fusion features from the speech segments of each speaker;

[0100] Use K-means clustering method to cluster multilingual speech data and obtain the initial GMM parameters for each speaker;

[0101] The GMM model of each speaker's speech segment is obtained by training the Expectation-Maximization Algorithm (EM) to fit the speech characteristics of each speaker.

[0102] In this embodiment, after the fusion features are obtained, they are applied to the training of the model, wherein N tester's English / Korean / Japanese etc. speech segments S are input:

[0103] S={s i}, i = 1, 2, ..., N;

[0104] l∈{English,Korean,Japenese,Mongolian}

[0105] In the formula, s i is the set of all speech segments of the i-th tester, i ranges from 1 to N, and N is the number of testers; is the speech segment of the lth language of the ith tester, l represents the language type, j represents the number of frames in the lth language speech segment of the tester; l represents the language type.

[0106] 1. Calculate and extract the fusion feature parameters of each speaker's multilingual speech,

[0107] F={F 1 , F 2 , …, F i , …, F n}; F i ={f 1 , …, f t}, where T is the number of frames; F is the set of multilingual speech fusion features of all testers; F n is the fusion feature set of the multilingual speech of the nth tester; F i is the fusion feature set of the multilingual speech of the i-th tester; f t is the t-th frame feature in the fusion feature set of the i-th tester, t represents the frame number, ranging from 1 to T, and T is the frame number;

[0108] 2. Get the initialization parameter λ of the speaker model based on the K-means clustering method 0 ;

[0109] 3. According to the initialization parameter λ 0 , the GMM model with the mixing degree M = 128 for each speaker is calculated using the EM algorithm, G L ={λ 1 , …, λ i , …, λ N},λ i is the GMM model parameter of the i-th tester, output G L .

[0110] The model confirmation method is as follows:

[0111] Input a Chinese speech segment of a tester t , GMM model G of N speaker models L , threshold Threshold (Threshold = 20);

[0112] 1. Calculate and extract the fusion feature parameter F of the tester's Chinese speech t ;

[0113] 2. Calculate F t and the GMM model parameters λ of the assumed speaker i Likelihood Probability Score And compare it with the threshold. If it is greater than the threshold, the test person is accepted as the speaker, otherwise it is rejected and the confirmation result is output.

[0114] See also Figure 3 As shown, in this embodiment, Gaussian mixture model (GMM) is used as the main feature model for cross-language speaker recognition, and the GMM model of the speaker is obtained by training using the K-means clustering method and the EM algorithm. In the test phase of the present invention, the maximum likelihood function method is used to calculate the test speech and the GMM model score, and then the judgment result is obtained by comparing with the threshold.

[0115] In this embodiment, in step 4, confirming or rejecting the identity of the tester includes the following steps:

[0116] In the testing phase, the tester's speech segment is input and the fusion features of the tester are extracted;

[0117] Calculate the likelihood probability score between the tester's speech features and the GMM model of the known speaker model;

[0118] The calculated likelihood probability score is compared with the set threshold. When the score is greater than the threshold, the test subject is considered a known speaker; when the score is less than the threshold, the test subject is rejected as a known speaker and the recognition result is output.

[0119] The present invention can effectively reduce the feature differences between different languages ​​and significantly improve the accuracy of cross-language speaker recognition by integrating multiple features including pitch and Mel frequency cepstral coefficients (MFCC). In a bilingual or multilingual environment, it can better adapt to the pronunciation habits and acoustic characteristics of different languages, overcoming the problem of reduced recognition performance caused by language mismatch.

[0120] The method of the present invention adopts an efficient Teager energy operator and MFCC calculation method in the feature extraction process, so that the amount of calculation is small, thereby improving the real-time performance of the entire process. This feature enables the method to play a role in application scenarios that require rapid response, and is suitable for various fields with high real-time requirements, such as security monitoring, financial transactions, and intelligent security. By combining the pitch period with the MFCC features and using the Gaussian mixture model (GMM) for modeling, the present invention can more accurately capture the personalized characteristics of the speaker, especially in a multilingual environment, and improve the accuracy of speaker recognition. The method can achieve a balance between different languages, avoiding the problem of a sudden drop in recognition performance of traditional methods when switching between languages.

[0121] It should be noted that the above figures are only schematic illustrations of the processes included in the method according to an exemplary embodiment of the present invention, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0122] It should be understood that, although described in a certain order, these steps are not necessarily performed in sequence in the above order. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, a part of the steps of the present embodiment may include a plurality of steps or a plurality of stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a part of the steps or stages in other steps or other steps.

[0123] The embodiment of the present invention further provides a cross-language speaker recognition system based on speech fusion features, which is used to execute the above-mentioned cross-language speaker recognition method. The cross-language speaker recognition system includes:

[0124] Voiced segment extraction module: used to extract and integrate voiced segments;

[0125] Fusion feature extraction module: used to extract pitch period and MFCC features and form a fusion feature vector;

[0126] GMM training module: used to train Gaussian mixture model (GMM);

[0127] Speaker verification module: used to calculate the likelihood probability score between the test speech and the GMM model, and compare it with the preset threshold to confirm or reject the identity of the tester.

[0128] Wherein, the voiced segment extraction module includes:

[0129] Bandpass filter: used to perform 60Hz~4kHz bandpass filtering on the original speech signal;

[0130] FFT calculation unit: used to perform fast Fourier transform (FFT) on the filtered signal;

[0131] TEO calculation unit: used to calculate the Teager Energy Operator (TEO);

[0132] The voiced segment integration unit is used to extract the voiced segments according to the preset threshold value, and integrate the voiced segments into a new speech segment in series.

[0133] Wherein, the fusion feature extraction module includes:

[0134] Pitch period extraction unit: used to extract the pitch period from the new speech segment;

[0135] MFCC extraction unit: used to extract MFCC features from new speech segments;

[0136] Feature fusion unit: used to concatenate the pitch period and MFCC features to form a fused feature vector.

[0137] Wherein, the GMM training module includes:

[0138] Fusion feature extraction unit: used to input multilingual speech segments of multiple testers and extract the fusion features of each speaker;

[0139] K-means clustering unit: used to initialize GMM parameters;

[0140] EM algorithm unit: used to train the GMM model for each speaker.

[0141] Wherein, the speaker confirmation module includes:

[0142] Fusion feature extraction unit: used to input the test speech segment and extract the fusion feature;

[0143] Likelihood probability calculation unit: used to calculate the likelihood probability score of the test speech and the GMM model;

[0144] Threshold comparison unit: used to compare with the preset threshold to confirm or reject the identity of the tester.

[0145] The present invention optimizes and adjusts the parameter settings in the feature extraction process through experiments, so that the system can adaptively adjust the feature extraction method according to different voice data, further improving the adaptability and recognition effect in a cross-language environment. In addition, the dynamic adjustment of the threshold enables the system to optimize the recognition decision according to real-time data, improving the flexibility and accuracy of recognition. The present invention is suitable for identity authentication and security monitoring in a multilingual environment, and is particularly suitable for applications in bilingual or multilingual areas. At the same time, the present invention has a wide range of application fields, especially in fields with high security requirements such as national defense, finance, and public security, and can provide users with efficient and accurate speaker recognition technical support.

[0146] In the face of factors such as pronunciation differences, phoneme changes, and tone differences in different languages, the system of the present invention integrates the pitch period and MFCC features to make the recognition system more robust to changes in speech signals. In particular, when there are large differences in pronunciation habits across languages, it can still maintain a high recognition accuracy, significantly improving the stability and reliability of the system.

[0147] Through the above detailed steps, the cross-language speaker recognition system based on speech fusion features of the present invention is used to execute the steps of the cross-language speaker recognition method based on speech fusion features in the above embodiment, which will not be repeated here. The cross-language speaker recognition method and system based on speech fusion features provided by the present invention have excellent cross-language adaptability, robustness, real-time and high-precision recognition capabilities, can effectively improve the speaker recognition performance in a multilingual environment, adapt to a wider range of application scenarios, and have strong market competitiveness and application value.

[0148] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications may be made without departing from the scope disclosed in the embodiments of the present invention as defined in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present invention may be described or required in individual form, they may also be understood as multiple unless explicitly limited to the singular.

[0149] It should be understood that, as used herein, the singular form "a" or "an" is intended to include the plural form as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" refers to any and all possible combinations of one or more of the items listed in association. The serial numbers of the embodiments disclosed in the above embodiments of the present invention are for description only and do not represent the advantages and disadvantages of the embodiments.

[0150] A person skilled in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the disclosure of the embodiments of the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and there are many other changes in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of simplicity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the protection scope of the embodiments of the present invention.

Claims

1. A cross-language speaker recognition method based on speech fusion features, characterized in that: The method comprises the following steps: Perform bandpass filtering on the original speech signal, extract the voiced segments and integrate them into new speech segments; Extract the pitch period and MFCC features in the new speech segment and concatenate them to form a fusion feature vector; Input multilingual speech segments of multiple testers, extract fusion features, and use K-means clustering and EM algorithm to train the GMM model of each speaker; Input the test speech segment, extract the fusion features, calculate the likelihood probability score with the GMM model, compare it with the preset threshold, and confirm or reject the identity of the tester.

2. The cross-language speaker recognition method based on speech fusion features as claimed in claim 1, characterized in that: The original speech signal is bandpass filtered to extract the voiced segments and integrate them into new speech segments, including the following steps: The original speech signal is filtered to remove 50Hz DC noise and high-frequency noise greater than 4kHz, and a bandpass filter with a frequency band of 60Hz to 4kHz is used for bandpass filtering; Perform fast Fourier transform on the filtered signal and calculate the Teager energy operator; The voiced segments are extracted according to the preset threshold, the speech frame numbers corresponding to the voiced segments are obtained, and the voiced segments are connected in series to form new speech segments as the original speech for extracting fusion features.

3. The cross-language speaker recognition method based on speech fusion features as claimed in claim 2, characterized in that: Extract the pitch period and MFCC features in the new speech segment and concatenate them to form a fusion feature vector, including the following steps: Extract and integrate voiced segments by calculating the TEO value of speech signals; Extract the pitch period and MFCC features in the new speech segment; The obtained 1-dimensional pitch period and MFCC are connected in series to form a new fusion feature vector.

4. The cross-language speaker recognition method based on speech fusion features as claimed in claim 3, characterized in that: Training the GMM model for each speaker includes the following steps: Input multilingual speech segments of multiple testers and extract fusion features from the speech segments of each speaker; Use K-means clustering method to cluster multilingual speech data and obtain the initial GMM parameters for each speaker; The GMM model of each speaker's speech segment is obtained through EM algorithm training to fit the speech characteristics of each speaker.

5. The cross-language speaker recognition method based on speech fusion features as claimed in claim 4, characterized in that: Confirming or denying the tester's identity includes the following steps: In the testing phase, the tester's speech segment is input and the fusion features of the tester are extracted; Calculate the likelihood probability score between the test subject's speech features and the GMM model of the known speaker model; The calculated likelihood probability score is compared with the set threshold. When the score is greater than the threshold, the test subject is considered a known speaker; when the score is less than the threshold, the test subject is rejected as a known speaker and the recognition result is output.

6. A cross-language speaker recognition system based on speech fusion features, characterized in that: Used to execute the cross-language speaker recognition method based on speech fusion features as described in any one of claims 1 to 5, the cross-language speaker recognition system comprises: Voiced segment extraction module: used to extract and integrate voiced segments; Fusion feature extraction module: used to extract pitch period and MFCC features and form a fusion feature vector; GMM training module: used to train Gaussian mixture models; Speaker verification module: used to calculate the likelihood probability score between the test speech and the GMM model, and compare it with the preset threshold to confirm or reject the identity of the tester.

7. The cross-language speaker recognition system based on speech fusion features as claimed in claim 6, characterized in that: The voiced segment extraction module comprises: Bandpass filter: used to perform 60Hz~4kHz bandpass filtering on the original speech signal; FFT calculation unit: used to perform fast Fourier transform on the filtered signal; TEO calculation unit: used to calculate the Teager energy operator; The voiced segment integration unit is used to extract the voiced segments according to the preset threshold value, and integrate the voiced segments into a new speech segment in series.

8. The cross-language speaker recognition system based on speech fusion features as claimed in claim 6, characterized in that: The fusion feature extraction module comprises: Pitch period extraction unit: used to extract the pitch period from the new speech segment; MFCC extraction unit: used to extract MFCC features from new speech segments; Feature fusion unit: used to concatenate the pitch period and MFCC features to form a fused feature vector.

9. The cross-language speaker recognition system based on speech fusion features as claimed in claim 6, characterized in that: The GMM training module includes: Fusion feature extraction unit: used to input multilingual speech segments of multiple testers and extract the fusion features of each speaker; K-means clustering unit: used to initialize GMM parameters; EM algorithm unit: used to train the GMM model for each speaker.

10. The cross-language speaker recognition system based on speech fusion features as claimed in claim 6, characterized in that: The speaker confirmation module comprises: Fusion feature extraction unit: used to input the test speech segment and extract the fusion feature; Likelihood probability calculation unit: used to calculate the likelihood probability score of the test speech and the GMM model; Threshold comparison unit: used to compare with the preset threshold to confirm or reject the identity of the tester.

Citation Information

Patent Citations

  • A speaker identification system for remote Chinese teaching

    CN101241699A

  • Children speech emotion recognition method

    CN101685634A

  • Voiceprint identification method based on Gauss mixing model and system thereof

    CN102324232A

  • Method and device for identifying ages

    CN104700843A

  • Vocal cord anomaly detection method based on acoustic phonetic features

    CN106941005A