Voiceprint recognition method and system based on multi-expert model

CN120748411APending Publication Date: 2025-10-03NANJING LONGYUAN INFORMATION TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510906553.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-03

Smart Images

  • Figure CN120748411A_ABST
    Figure CN120748411A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voiceprint recognition, in particular to a voiceprint recognition method and system based on a multi-expert model. The method comprises the following steps: acquiring a voiceprint signal of a user; carrying out noise reduction and normalization operation on the collected voiceprint signals to prepare for subsequent feature extraction; extracting features of multiple dimensions from the pre-processed voiceprint signals, and inputting the features into a multi-expert model; analyzing the plurality of voiceprint features by using a multi-expert model, and outputting an identification result; dynamically adjusting the weight of each model according to the output confidence of each expert model and the current environment condition, and generating a final voiceprint recognition result; a final voiceprint recognition result is fed back to a user and an application system, and multiple application scenes are supported; by adopting the above mode, comprehensive evaluation can be performed on the voiceprint signal from multiple angles, and the accuracy and robustness of voiceprint recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voiceprint recognition, and in particular to a voiceprint recognition method and system based on a multi-expert model. Background Art

[0002] Speaker recognition, also known as speaker identification or voice recognition, is a biometric recognition technology that identifies a speaker by analyzing unique features in voice signals. Voiceprint recognition can be divided into two categories: speaker verification and speaker identification. The former is used to verify that a person is who they claim to be, while the latter is used to determine the identity of an unknown speaker from a group of known speakers. The core of voiceprint recognition lies in extracting and utilizing unique features in voice signals.

[0003] Traditional voiceprint recognition methods mainly rely on a single model for feature extraction and classification. Although they can achieve a high accuracy rate under certain specific conditions, it is difficult to take all aspects into account when faced with multi-dimensional features. Voice color, gender, age and frequency features may show different importance in different environments. A single model is difficult to flexibly adjust the weight, which affects the accuracy of recognition. Summary of the Invention

[0004] The purpose of the present invention is to provide a voiceprint recognition method and system based on a multi-expert model, aiming to solve the technical problem in the prior art that voiceprint recognition is difficult to take into account all aspects when facing multi-dimensional features, and timbre, gender, age and frequency features may show different importance in different environments. A single model is difficult to flexibly adjust the weights, thereby affecting the accuracy of recognition.

[0005] To achieve the above objectives, the present invention adopts a voiceprint recognition method based on a multi-expert model, comprising the following steps:

[0006] Collect the user's voiceprint signal;

[0007] Perform noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction;

[0008] Extract features of multiple dimensions from the preprocessed voiceprint signal and input the features into the multi-expert model;

[0009] Use multiple expert models to analyze multiple voiceprint features and output recognition results;

[0010] Dynamically adjust the weight of each model based on the output confidence of each expert model and the current environmental conditions, and generate the final voiceprint recognition result;

[0011] The final voiceprint recognition results are fed back to users and application systems, and support multiple application scenarios.

[0012] Among them, in the step of performing noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction:

[0013] The noise reduction methods include spectral subtraction and Wiener filtering, and the normalization processing adopts mean variance normalization and maximum and minimum normalization.

[0014] Among them, in the step of extracting features of multiple dimensions from the preprocessed voiceprint signal and inputting the features into the multi-expert model:

[0015] The features include timbre, gender, age and frequency. Short-time Fourier transform (STFT) is used to extract frequency features, Mel-frequency cepstral coefficients (MFCC) are used to extract timbre features, and Gaussian mixture model (GMM) is used to estimate gender and age.

[0016] Among them, in the step of using multiple expert models to analyze multiple voiceprint features and output recognition results:

[0017] The multi-expert model includes timbre expert model, gender expert model, age expert model and frequency expert model.

[0018] Among them, before using the multi-expert model to analyze multiple voiceprint features and output the recognition results:

[0019] Train multiple expert models so that each expert model undergoes specialized training to ensure that each achieves high-precision feature extraction and classification in the responsible dimension. The training process includes data preparation, model selection, model training, model verification and model optimization.

[0020] Among them, in the step of feeding back the final voiceprint recognition results to users and application systems and supporting multiple application scenarios:

[0021] Application scenarios include authentication, voice assistant, healthcare and smart home.

[0022] The present invention also provides a voiceprint recognition system based on a multi-expert model, comprising a voiceprint acquisition module, a preprocessing module, a feature extraction module, a multi-expert model module, an adaptive decision module and an output module; wherein:

[0023] The voiceprint collection module is used to collect the user's voiceprint signal;

[0024] The pre-processing module is used to perform noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction;

[0025] The feature extraction module is used to extract features of multiple dimensions from the preprocessed voiceprint signal and input the features into the multi-expert model;

[0026] The multi-expert model module is used to analyze multiple voiceprint features using multiple expert models and output recognition results;

[0027] The adaptive decision module is used to dynamically adjust the weight of each model according to the output confidence of each expert model and the current environmental conditions, and generate the final voiceprint recognition result;

[0028] The output module is used to feed back the final voiceprint recognition results to the user and the application system, and supports multiple application scenarios.

[0029] The multi-expert model module includes a timbre expert model, a gender expert model, an age expert model and a frequency expert model; wherein:

[0030] The timbre expert model is responsible for analyzing the timbre characteristics of the voiceprint;

[0031] The gender expert model is used to determine the gender characteristics of the voiceprint and distinguish between male and female voices;

[0032] The age expert model is used to estimate the age characteristics of the voiceprint and identify voices of different age groups;

[0033] The frequency expert model is responsible for analyzing the frequency characteristics of the voiceprint.

[0034] The present invention discloses a voiceprint recognition method and system based on a multi-expert model. The method first collects the user's voiceprint signal, then performs noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction, extracts features of multiple dimensions from the preprocessed voiceprint signal, and inputs the features into the multi-expert model. The multi-expert model is then used to analyze multiple voiceprint features and output recognition results. Subsequently, the weight of each model is dynamically adjusted according to the output confidence of each expert model and the current environmental conditions, and the final voiceprint recognition result is generated. Finally, the final voiceprint recognition result is fed back to the user and the application system, and supports multiple application scenarios. By adopting the above method and constructing a multi-expert model, it is possible to comprehensively evaluate the voiceprint signal from multiple angles, thereby improving the accuracy and robustness of voiceprint recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 It is a flowchart of the steps of the voiceprint recognition method based on the multi-expert model of the present invention.

[0037] Figure 2 It is a flow chart of the steps in the feature extraction process of the present invention.

[0038] Figure 3 It is a step flow chart of S500 of the present invention.

[0039] Figure 4 It is a schematic diagram of the principle of the voiceprint recognition system based on multiple expert models of the present invention.

[0040] 701-Voiceprint acquisition module, 702-Preprocessing module, 703-Feature extraction module, 704-Multi-expert model module, 705-Adaptive decision module, 706-Output module, 7041-Timbre expert model, 7042-Gender expert model, 7043-Age expert model, 7044-Frequency expert model. DETAILED DESCRIPTION

[0041] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.

[0042] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0043] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0044] See also Figures 1 to 3 The present invention provides a voiceprint recognition method based on a multi-expert model, comprising the following steps:

[0045] S100: collecting the user's voiceprint signal;

[0046] S200: Perform noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction;

[0047] S300: Extracting features of multiple dimensions from the preprocessed voiceprint signal and inputting the features into the multi-expert model;

[0048] S400: Analyze multiple voiceprint features using multiple expert models and output recognition results;

[0049] S500: Dynamically adjust the weight of each model based on the output confidence of each expert model and the current environmental conditions, and generate the final voiceprint recognition result;

[0050] S600: Feeds back the final voiceprint recognition results to users and application systems, and supports multiple application scenarios.

[0051] In this implementation, the user's voiceprint signal is first collected. Noise reduction and normalization are then performed on the collected voiceprint signal to prepare for subsequent feature extraction. Multi-dimensional features are extracted from the preprocessed voiceprint signal and input into a multi-expert model. The multi-expert model is then used to analyze the multiple voiceprint features and output recognition results. The weights of each model are then dynamically adjusted based on the output confidence of each expert model and the current environmental conditions, generating a final voiceprint recognition result. The final voiceprint recognition result is then fed back to the user and the application system, supporting a variety of application scenarios. The output probability of each expert model is used as its confidence. A higher confidence indicates a better recognition performance under the current environment. The weight of each model is dynamically adjusted based on its confidence. Specifically, the system uses a softmax function to normalize the confidence of each model and generate its weight. Assuming there are \(N\) expert models and the confidence of the \(i\)th model is \(p_i\), its weight \(wi\) can be expressed as:

[0052] \[w_i=\frac{e^{p_i}}{\sum_{j=1}^{N}e^{p_j}}\]

[0053] Then, according to the weight of each model, their outputs are weighted averaged to generate the final voiceprint recognition result. Assuming that the output of the \(i\)th model is \(y_i\), the final recognition result \(Y\) can be expressed as:

[0054] \[Y=\sum_{i=1}^{N}w_i y_i\];

[0055] By adopting the above method and constructing a multi-expert model, it is possible to comprehensively evaluate voiceprint signals from multiple angles, thereby improving the accuracy and robustness of voiceprint recognition.

[0056] Furthermore, in the step of performing noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction:

[0057] The noise reduction methods include spectral subtraction and Wiener filtering, and the normalization processing adopts mean variance normalization and maximum and minimum normalization.

[0058] In this embodiment, in the preprocessing stage, the system first performs noise reduction processing on the voiceprint signal to remove environmental noise and interference signals, wherein the noise reduction method includes spectral subtraction and Wiener filtering. Then, the system normalizes the signal to ensure that different voiceprint signals are compared at the same scale, wherein the normalization processing can adopt the mean variance normalization and maximum and minimum normalization methods.

[0059] Furthermore, in the step of extracting features of multiple dimensions from the preprocessed voiceprint signal and inputting the features into the multi-expert model:

[0060] The features include timbre, gender, age and frequency. Short-time Fourier transform (STFT) is used to extract frequency features, Mel-frequency cepstral coefficients (MFCC) are used to extract timbre features, and Gaussian mixture model (GMM) is used to estimate gender and age.

[0061] In this embodiment, STFT is used to convert the voiceprint signal into the frequency domain to extract frequency features such as fundamental frequency and harmonics. STFT can provide a spectrum with high time-frequency resolution, which is convenient for subsequent analysis. MFCC is used to extract the timbre characteristics of the voiceprint. MFCC is a commonly used audio feature extraction method that can effectively characterize the pitch and sound quality of the voiceprint. GMM is used to estimate the gender and age characteristics of the voiceprint. GMM is a probabilistic model that can estimate the gender and age distribution of the voiceprint by learning a large amount of sample data.

[0062] Furthermore, in the step of using multiple expert models to analyze multiple voiceprint features and output recognition results:

[0063] The multi-expert model includes timbre expert model, gender expert model, age expert model and frequency expert model.

[0064] In this embodiment, the timbre expert model is responsible for analyzing the timbre characteristics of the voiceprint, the gender expert model is responsible for determining the gender characteristics of the voiceprint and distinguishing male and female voices, the age expert model is responsible for estimating the age characteristics of the voiceprint and identifying voices of different age groups, and the frequency expert model is responsible for analyzing the frequency characteristics of the voiceprint; wherein:

[0065] Tone Expert Model:

[0066] Feature extraction: Mel-frequency cepstral coefficients (MFCC) are used as timbre features, and 13-dimensional MFCC features and their first-order and second-order differences are extracted.

[0067] import librosadefextract_mfcc(audio_signal,sr=16000):

[0068] mfcc=librosa.feature.mfcc(y=audio_signal,sr=sr,n_mfcc=13)

[0069] delta_mfcc=librosa.feature.delta(mfcc)

[0070] delta2_mfcc=librosa.feature.delta(mfcc,order=2)

[0071] returnnp.concatenate((mfcc,delta_mfcc,delta2_mfcc),axis=0)

[0072] Model training: Use Gaussian mixture model (GMM) or deep neural network (DNN) to model timbre features and train timbre classifiers.

[0073] from sklearn.mixture import GaussianMixture

[0074] deftrain_gmm(features,n_components=16):

[0075] gmm=GaussianMixture(n_components=n_components)

[0076] gmm.fit(features)

[0077] return gmm

[0078] Gender Expert Model:

[0079] Feature extraction: Gender classification is performed based on features such as fundamental frequency (F0) and spectral centroid.

[0080] defextract_f0_and_spectral_centroid(audio_signal,sr=16000):

[0081] f0,voiced_flag,voiced_probs=librosa.pyin(audio_signal,fmin=librosa.note_to_hz('C2'),fmax=librosa.note_to_hz('C7'))

[0082] spectral_centroid=librosa.feature.spectral_centroid(y=audio_signal,sr=sr)

[0083] return f0,spectral_centroid

[0084] Model training: Use support vector machine (SVM) or convolutional neural network (CNN) to classify gender features.

[0085] from sklearn.svm import SVC

[0086] deftrain_svm(features,labels):

[0087] svm=SVC(kernel='linear')

[0088] svm.fit(features,labels)

[0089] return svm

[0090] Age Expert Model:

[0091] Feature extraction: Age estimation is performed using features such as the harmonic-to-noise ratio (HNR) and spectral tilt of speech signals.

[0092] defextract_hnr_and_spectral_tilt(audio_signal,sr=16000):

[0093] hnr=librosa.effects.harmonic(audio_signal) /

[0094] librosa.effects.percussive(audio_signal)

[0095] spectral_tilt=np.mean(librosa.feature.spectral_slope(y=audio_signal,sr=sr))

[0096] returnhnr,spectral_tilt

[0097] Model training: Use a regression model (such as linear regression or random forest) to predict age features.

[0098] from sklearn.ensemble import RandomForestRegressor

[0099] deftrain_random_forest(features,labels):

[0100] rf=RandomForestRegressor(n_estimators=100)

[0101] rf.fit(features,labels)

[0102] return rf

[0103] Frequency Expert Model:

[0104] Feature extraction: Extract spectral features through short-time Fourier transform (STFT) and analyze the frequency distribution of speech signals.

[0105] defextract_stft_features(audio_signal,sr=16000):

[0106] stft=librosa.stft(audio_signal)

[0107] magnitude,phase=np.abs(stft),np.angle(stft)

[0108] return magnitude,phase

[0109] Model training: Use deep neural networks (DNNs) to model frequency features and identify specific frequency patterns.

[0110] from tensorflow.keras.models import Sequential

[0111] from tensorflow.keras.layers import Dense,Dropout

[0112] defbuild_dnn(input_shape,output_units):

[0113] model=Sequential()

[0114] model.add(Dense(256,activation='relu',input_shape=input_shape))

[0115] model.add(Dropout(0.5))

[0116] model.add(Dense(128,activation='relu'))

[0117] model.add(Dropout(0.5))

[0118] model.add(Dense(output_units,activation='softmax'))

[0119] model.compile(optimizer='adam',loss='categorical_crossentropy',metrics=['accuracy'])

[0120] return model

[0121] Furthermore, before using the multi-expert model to analyze multiple voiceprint features and output recognition results:

[0122] Train multiple expert models so that each expert model undergoes specialized training to ensure that each achieves high-precision feature extraction and classification in the responsible dimension. The training process includes data preparation, model selection, model training, model verification and model optimization.

[0123] In this embodiment, a large amount of voiceprint data is first collected and labeled, including information on dimensions such as timbre, gender, age, and frequency. Then, a model structure suitable for each dimension is selected, such as deep neural network (DNN) and convolutional neural network (CNN). The labeled data sets are used to train each expert model, and the model parameters are optimized through the back propagation algorithm. The trained model is verified using the cross-validation method to evaluate its performance on different data sets. Finally, the model parameters are adjusted according to the verification results to optimize the model performance.

[0124] Furthermore, in the step of feeding back the final voiceprint recognition results to the user and the application system and supporting multiple application scenarios:

[0125] Application scenarios include authentication, voice assistant, healthcare and smart home.

[0126] In this embodiment, in the fields of finance, security, etc., voiceprint recognition can be used for user identity authentication to improve the security of the system; in scenarios such as smart homes and vehicle systems, voiceprint recognition can be used for personalized services to enhance user experience; in smart home systems, voiceprint recognition can be used to control home appliances and realize voice interaction; in the field of medical health, voiceprint recognition can be used for auxiliary diagnosis, such as judging the health status of patients through voiceprint characteristics.

[0127] See also Figure 4The present invention also provides a voiceprint recognition system based on a multi-expert model, comprising a voiceprint acquisition module 701, a preprocessing module 702, a feature extraction module 703, a multi-expert model module 704, an adaptive decision module 705, and an output module 706; wherein:

[0128] The voiceprint collection module 701 is used to collect the user's voiceprint signal;

[0129] The pre-processing module 702 is used to perform noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction;

[0130] The feature extraction module 703 is used to extract features of multiple dimensions from the pre-processed voiceprint signal and input the features into the multi-expert model;

[0131] The multi-expert model module 704 is responsible for analyzing multiple voiceprint features and outputting recognition results;

[0132] The adaptive decision module 705 is used to dynamically adjust the weight of each model according to the output confidence of each expert model and the current environmental conditions, and generate the final voiceprint recognition result;

[0133] The output module 706 is used to feed back the final voiceprint recognition result to the user and the application system, and supports multiple application scenarios.

[0134] In this embodiment, the user's voiceprint signal is first collected by the voiceprint collection module 701, and the preprocessing module 702 performs noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction. Then, the feature extraction module 703 extracts features of multiple dimensions from the preprocessed voiceprint signal and inputs the features into the multi-expert model. The multi-expert model module 704 is responsible for analyzing multiple voiceprint features and outputting recognition results. The adaptive decision module 705 is used to dynamically adjust the weight of each model according to the output confidence of each expert model and the current environmental conditions, and generate the final voiceprint recognition result. Finally, the output module 706 is used to feed back the final voiceprint recognition result to the user and the application system, and supports multiple application scenarios.

[0135] Furthermore, the multi-expert model module 704 includes a timbre expert model 7041, a gender expert model 7042, an age expert model 7043, and a frequency expert model 7044; wherein:

[0136] The timbre expert model 7041 is responsible for analyzing the timbre characteristics of the voiceprint;

[0137] The gender expert model 7042 is responsible for determining the gender characteristics of the voiceprint and distinguishing between male and female voices;

[0138] The age expert model 7043 is responsible for estimating the age characteristics of the voiceprint and identifying voices of different age groups;

[0139] The frequency expert model 7044 is responsible for analyzing the frequency characteristics of the voiceprint.

[0140] In this embodiment, the timbre expert model 7041 is responsible for analyzing the timbre characteristics of the voiceprint, the gender expert model 7042 is responsible for determining the gender characteristics of the voiceprint and distinguishing between male and female voices, the age expert model 7043 is responsible for estimating the age characteristics of the voiceprint and identifying voices of different age groups, and the frequency expert model 7044 is responsible for analyzing the frequency characteristics of the voiceprint.

[0141] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed in this application.

[0142] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. A voiceprint recognition method based on a multi-expert model, characterized in that: The steps include: Collect the user's voiceprint signal; Perform noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction; Extract features of multiple dimensions from the preprocessed voiceprint signal and input the features into the multi-expert model; Use multiple expert models to analyze multiple voiceprint features and output recognition results; Dynamically adjust the weight of each model based on the output confidence of each expert model and the current environmental conditions, and generate the final voiceprint recognition result; The final voiceprint recognition results are fed back to users and application systems, and support multiple application scenarios.

2. The voiceprint recognition method based on a multi-expert model according to claim 1, characterized in that: In the steps of performing noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction: The noise reduction methods include spectral subtraction and Wiener filtering, and the normalization processing adopts mean variance normalization and maximum and minimum normalization.

3. The voiceprint recognition method based on a multi-expert model according to claim 1, characterized in that: In the step of extracting features of multiple dimensions from the preprocessed voiceprint signal and inputting the features into the multi-expert model: The features include timbre, gender, age and frequency. Short-time Fourier transform is used to extract frequency features, Mel-frequency cepstral coefficients are used to extract timbre features, and Gaussian mixture model is used to estimate gender and age.

4. The voiceprint recognition method based on a multi-expert model according to claim 1, characterized in that: In the step of using multiple expert models to analyze multiple voiceprint features and output recognition results: The multi-expert model includes timbre expert model, gender expert model, age expert model and frequency expert model.

5. The voiceprint recognition method based on a multi-expert model according to claim 4, characterized in that: Before using the multi-expert model to analyze multiple voiceprint features and output recognition results: Train multiple expert models so that each expert model undergoes specialized training to ensure that each achieves high-precision feature extraction and classification in the responsible dimension. The training process includes data preparation, model selection, model training, model verification and model optimization.

6. The voiceprint recognition method based on a multi-expert model according to claim 1, characterized in that: In the step of feeding back the final voiceprint recognition results to users and application systems and supporting multiple application scenarios: Application scenarios include authentication, voice assistant, healthcare and smart home.

7. A voiceprint recognition system based on a multi-expert model, applied to the voiceprint recognition method based on a multi-expert model as claimed in claim 1, characterized in that: It includes voiceprint collection module, preprocessing module, feature extraction module, multi-expert model module, adaptive decision module and output module; among which: The voiceprint collection module is used to collect the user's voiceprint signal; The pre-processing module is used to perform noise reduction and normalization operations on the collected voiceprint signal to prepare for subsequent feature extraction; The feature extraction module is used to extract features of multiple dimensions from the preprocessed voiceprint signal and input the features into the multi-expert model; The multi-expert model module is used to analyze multiple voiceprint features using multiple expert models and output recognition results; The adaptive decision module is used to dynamically adjust the weight of each model according to the output confidence of each expert model and the current environmental conditions, and generate the final voiceprint recognition result; The output module is used to feed back the final voiceprint recognition results to the user and the application system, and supports multiple application scenarios.

8. The voiceprint recognition system based on a multi-expert model according to claim 7, characterized in that: The multi-expert model module includes a timbre expert model, a gender expert model, an age expert model and a frequency expert model; wherein: The timbre expert model is responsible for analyzing the timbre characteristics of the voiceprint; The gender expert model is used to determine the gender characteristics of the voiceprint and distinguish between male and female voices; The age expert model is used to estimate the age characteristics of the voiceprint and identify voices of different age groups; The frequency expert model is responsible for analyzing the frequency characteristics of the voiceprint.

Citation Information

Patent Citations

  • Multi-dimensional identity information identification method and device, computer equipment and storage medium

    CN110443137A

  • Household voiceprint recognition method based on multi-feature parameter fusion

    CN111785285A

  • Voiceprint recognition method and device, equipment and storage medium

    CN116246635A

  • Information identification method, device and equipment and storage medium thereof

    CN119207426A

  • Dynamic equipment quality evaluation method based on multi-agent collaborative reinforcement learning

    CN119671383A