Gender, age and accent joint detection method and system based on speaker classification

By constructing a speaker database and using deep neural network model for training, combined with the additional angle margin loss function, joint optimization of gender, age, and accent in speech data is achieved, solving the problems of high data acquisition and deployment costs and insufficient model correlation in traditional methods, and improving classification accuracy.

CN118522291BActive Publication Date: 2025-06-06PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410591017.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-06-06
Estimated Expiration
2044-05-13

AI Technical Summary

Technical Problem

Traditional methods of predicting the age, gender, and accent of the speaker in voice data require the construction of databases and training models separately, resulting in high data acquisition and deployment costs, and lack of correlation between models, affecting classification accuracy.

Method used

A joint detection method for gender, age and accent based on speaker classification is proposed. By constructing a speaker database and using a deep neural network model for training, an additional angle margin loss function is used to improve the distinguishing ability of the model, and joint optimization of gender, age and accent is achieved.

Benefits of technology

It reduces the cost of data collection and deployment, uses a single model to predict age, gender, and accent simultaneously, improves the accuracy of model classification, and realizes joint optimization of gender, age, and accent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118522291B_ABST
    Figure CN118522291B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for joint detection of gender, age and accent based on speaker classification, the method comprising: constructing a speaker database; constructing a deep neural network speaker classification model based on the speaker database, and using an additional angular margin loss function to train the deep neural network speaker classification model; performing a forward neural network calculation on the input voice data to obtain the posterior probability of the speaker label; obtaining the speaker label corresponding to the input voice according to the posterior probability, and outputting the gender, age and accent information corresponding to the speaker label. The present invention reduces the cost of data collection by constructing a speaker database, uses one model to simultaneously predict age, gender and accent, and only deploys one model when applied, saving computing resources; can simultaneously predict the corresponding gender, age, accent and other information based on voice data, realizes the joint optimization of age, gender and accent, promotes each other, and effectively improves the accuracy of model classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular, to a user portrait prediction technology in the field of speech big data analysis; specifically, to a gender, age, and accent joint detection method and system based on speaker classification. Background Art

[0002] With the popularization of emerging technologies such as AI and 5G, the market demand for intelligent customer service has increased. When manual customer service reception and office space are suspended, online consultations have gradually increased, and a large amount of voice data is generated every day. A lot of valuable information can be mined from these voice data to help companies develop better. At present, the customer service center is transforming from a cost center to a value center.

[0003] However, since voice data is unstructured and difficult to analyze, automatic speech recognition (ASR) is usually used to recognize voice data as text data, and then natural language processing (NLP) is used to mine knowledge. In addition, the user's emotions, age, gender, accent and other information can be judged from voice data to portray user portraits, provide personalized services, and improve customer satisfaction with the company.

[0004] Predicting the speaker's age, gender, accent and other information based on voice data can be regarded as a classification problem and implemented using statistical learning-based methods, such as mixed Gaussian models, deep neural network models, etc. It is generally divided into three steps: first, building a corresponding database, such as an age database, gender database, accent database, etc.; second, training a classification model based on the built database; third, predicting the age, gender, accent, etc. of the input voice based on the trained classification model.

[0005] Traditional prediction methods generally use different models to predict age, gender, and accent respectively. This prediction method has the following three shortcomings:

[0006] First, it is necessary to build age database, gender database and accent database separately, and the cost of data collection is high;

[0007] Second, three different models need to be trained and deployed, which results in high deployment costs.

[0008] 3. The models for predicting age, gender, and accent are independent of each other, and the correlation among the three models is not considered. The performance of the models, such as their correlation, needs to be further improved. Summary of the invention

[0009] In view of this, the purpose of the present invention is to propose a method and system for joint detection of gender, age and accent based on speaker classification, which reduces the cost of data collection by building a speaker database, uses one model to simultaneously predict age, gender and accent, and saves computing resources; based on speech data, the corresponding gender, age, accent and other information are simultaneously predicted to achieve joint optimization of age, gender and accent, promote each other, and improve the accuracy of model classification.

[0010] The present invention provides a method for joint detection of gender, age and accent based on speaker classification, comprising the following steps:

[0011] S1. Collecting speech data of a set number of speakers, and recording the gender, age, and accent-related information of the speakers, to construct a speaker database, wherein the speaker database includes: the speech data of the speakers and the corresponding speaker label information; the speaker label information includes: the gender, age, and accent information of the speakers;

[0012] S2. Using the collected speech data and the corresponding speaker label information, a deep neural network speaker classification model based on the speaker database is constructed, and an additive angular margin loss function (AAM) is used to train the deep neural network speaker classification model; the expression of the additive angular margin loss function is:

[0013]

[0014] In formula (1), θ is the cosine similarity angle between the embedding vector output by the model and the embedding vector corresponding to the real speaker category, m a is the angular margin parameter, y is the index of the real speaker class, θ i is the cosine similarity angle between the model output and each category embedding vector;

[0015] The present invention optimizes model parameters by minimizing the loss function during the training of the deep neural network speaker classification model using training set data. The loss function includes not only the traditional cross entropy loss but also the AAM loss to ensure that the model can accurately distinguish different speaker categories.

[0016] The angular margin parameter m of the additional angular margin loss function (AAM) a Used to increase the angular distance between a sample and its nearest neighbor during model training, making the model more discriminative for different speakers.

[0017] In order to improve the model's discriminative ability, especially in tasks such as speaker classification that have obvious class distinction, the model is trained using the additive angle margin loss function AAM. The AAM loss function aims to increase the distinction between the model output embedding vector and its corresponding class embedding vector.

[0018] The AAM loss function maximizes the cosine similarity between the sample and its corresponding category while minimizing the cosine similarity with other categories, which can encourage the model to generate more dispersed speaker embeddings, thereby improving the accuracy of speaker classification.

[0019] Preferably, after the model training is completed, the model is evaluated using the validation set and the test set to ensure that the model has good generalization ability, and the model is adjusted and optimized as necessary.

[0020] S3. According to the trained deep neural network speaker classification model, the input speech data is subjected to a forward neural network calculation to obtain the posterior probability of the speaker label;

[0021] S4. Obtain a speaker label corresponding to the input speech according to the posterior probability, and output the gender, age, and accent information corresponding to the speaker label.

[0022] Furthermore, the S3 step includes the following steps:

[0023] S31, inputting the speech data into the trained deep neural network speaker classification model;

[0024] S32, the deep neural network speaker classification model processes the input speech data through a forward propagation algorithm, and extracts features and performs classification using parameters learned by the model in the training phase;

[0025] S33. In the output layer of the deep neural network speaker classification model, a softmax activation function is used to convert the original output logits into a probability distribution; the expression of the softmax activation function is:

[0026]

[0027] In formula (2), P(C = k|x) is the posterior probability that the speaker label is category k given the input x;

[0028] f k (x) is the original prediction value of the model for sample x in category k, that is, logits;

[0029] Denominator It is the sum of the index values ​​of all categories, ensuring that the sum of the probabilities of all categories is 1;

[0030] S34, through steps S31-S33, the deep neural network speaker classification model generates a probability distribution for each input sample, and the probability distribution covers all possible speaker labels;

[0031] S35. According to the probability distribution, select the label with the highest probability as the speaker label predicted by the model;

[0032] S36. The deep neural network speaker classification model outputs the speaker label and its posterior probability corresponding to each input speech data.

[0033] Furthermore, the method of outputting the gender, age, and accent information corresponding to the speaker tag in step S4 includes:

[0034] Based on the mapping relationship between labels and features established during the model training and database construction phase, the speaker labels are converted into specific gender, age, and accent information, and these gender, age, and accent information are output. Since the model directly outputs speaker labels, a mapping process is required.

[0035] Furthermore, after the step S4, the following steps are further included:

[0036] Deploy the trained deep neural network speaker classification model to the actual application environment to classify the speakers of new speech data and predict the corresponding gender, age, and accent information based on the speaker classification.

[0037] Furthermore, the deep neural network speaker classification model used in step S2 includes: any one of an ECAPA-TDNN model and a UIS-RNN model.

[0038] The ECAPA-TDNN model enhances the time-delay neural network TDNN framework by introducing Res2Block, squeeze-and-excitation (SE) modules and Adaptive Statistics Pooling (ASP) pooling layers to improve the ability to capture and distinguish the characteristics of different speakers.

[0039] The UIS-RNN model establishes a recurrent neural network (RNN) for each speaker, which is trained through supervised learning and can be continuously updated and adapted to new speakers.

[0040] The present invention also provides a gender, age, and accent joint detection system based on speaker classification, which implements the gender, age, and accent joint detection method based on speaker classification as described above, including:

[0041] The speaker database building module is used to collect the voice data of a set number of speakers, and record the gender, age, and accent-related information of the speakers to build a speaker database, wherein the speaker database includes: the speaker's voice data and the corresponding speaker tag information; the speaker tag information includes: the speaker's gender, age, and accent information;

[0042] Constructing a deep neural network speaker classification model module: for constructing a deep neural network speaker classification model based on the speaker database using the collected speech data and the corresponding speaker label information, and using the additional angle margin loss function AAM to train the deep neural network speaker classification model;

[0043] Calculate posterior probability module: used to perform forward neural network calculation on the input speech data according to the trained deep neural network speaker classification model to obtain the posterior probability of the speaker label;

[0044] Output speaker label corresponding information module: used to obtain the speaker label corresponding to the input speech according to the posterior probability, and output the gender, age, and accent information corresponding to the speaker label.

[0045] Furthermore, the module for calculating the posterior probability includes:

[0046] Model input unit: used to input speech data into the trained deep neural network speaker classification model;

[0047] Forward propagation unit: used by the deep neural network speaker classification model to process the input speech data through the forward propagation algorithm, and extract features and perform classification using the parameters learned by the model during the training phase;

[0048] Output layer activation unit: used in the output layer of the deep neural network speaker classification model, using the softmax activation function to convert the original output logits into a probability distribution;

[0049] Obtaining posterior probability unit: used to generate a probability distribution for each input sample by the deep neural network speaker classification model through steps S31-S33, and the probability distribution covers all possible speaker labels;

[0050] A label selection unit: used for selecting the label with the highest probability as the speaker label predicted by the model according to the probability distribution;

[0051] Output label unit: used by the deep neural network speaker classification model to output the speaker label and its posterior probability corresponding to each input speech data.

[0052] Furthermore, the module for outputting speaker label corresponding information includes:

[0053] Gender, age, and accent information mapping unit: used to convert the speaker label into specific gender, age, and accent information based on the mapping relationship between labels and features established in the model training and database construction stages, and output the gender, age, and accent information.

[0054] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method for joint detection of gender, age and accent based on speaker classification as described above are implemented.

[0055] The present invention also provides a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method for joint detection of gender, age, and accent based on speaker classification as described above are implemented.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] The gender, age, and accent joint detection method and system based on speaker classification provided by the present invention reduces the cost of data collection by constructing a speaker database, and uses one model to simultaneously predict age, gender, and accent. Only one model is deployed during application, saving computing resources. The corresponding gender, age, accent and other information can be simultaneously predicted based on voice data, thereby achieving joint optimization of age, gender, and accent, promoting each other, and effectively improving the accuracy of model classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the following detailed description of the preferred embodiment.The drawings are only for the purpose of illustrating the preferred embodiments and are not to be construed as limiting the invention.

[0059] In the attached picture:

[0060] Figure 1 A schematic diagram of the basic flow of a method for joint detection of gender, age and accent based on speaker classification according to an embodiment of the present invention;

[0061] Figure 2 Schematic diagram of the processing flow of the deep neural network speaker classification model in an embodiment of the present invention;

[0062] Figure 3 This is a flow chart of the method for joint detection of gender, age and accent based on speaker classification of the present invention;

[0063] Figure 4 is a flow chart of step S3 of the present invention;

[0064] Figure 5 The figure is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0065] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and products consistent with some aspects of the present disclosure as detailed in the appended claims.

[0066] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms of "a", "said" and "the" used in this disclosure and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0067] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0068] The embodiments of the present invention are described in further detail below.

[0069] The embodiment of the present invention provides a method for joint detection of gender, age and accent based on speaker classification, see Figure 3 As shown, the following steps are included:

[0070] S1. Collecting speech data of a set number of speakers, and recording the gender, age, and accent-related information of the speakers, to construct a speaker database, wherein the speaker database includes: the speech data of the speakers and the corresponding speaker label information; the speaker label information includes: the gender, age, and accent information of the speakers;

[0071] S2. Using the collected speech data and the corresponding speaker label information, a deep neural network speaker classification model based on the speaker database is constructed (e.g. Figure 2As shown), the deep neural network speaker classification model is trained using an additional angular margin loss function (Additive Angular Margin, AAM); the expression of the additional angular margin loss function is:

[0072]

[0073] In formula (1), θ is the cosine similarity angle between the embedding vector output by the model and the embedding vector corresponding to the real speaker category, m a is the angular margin parameter, y is the index of the real speaker class, θ i is the cosine similarity angle between the model output and each category embedding vector;

[0074] The deep neural network speaker classification model used in this embodiment is the ECAPA-TDNN model.

[0075] The ECAPA-TDNN model enhances the time-delay neural network TDNN framework by introducing Res2Block, squeeze-and-excitation (SE) modules and Adaptive Statistics Pooling (ASP) pooling layers, thereby improving the ability to capture and distinguish the characteristics of different speakers.

[0076] In the process of training a deep neural network speaker classification model using training set data, the embodiment of the present invention optimizes model parameters by minimizing the loss function. The loss function includes not only the traditional cross entropy loss but also the AAM loss to ensure that the model can accurately distinguish different speaker categories.

[0077] The angular margin parameter m of the additional angular margin loss function (AAM) a Used to increase the angular distance between a sample and its nearest neighbor during model training, making the model more discriminative for different speakers.

[0078] In order to improve the model's discriminative ability, especially in tasks such as speaker classification that have obvious class distinction, the model is trained using the additive angle margin loss function AAM. The AAM loss function aims to increase the distinction between the model output embedding vector and its corresponding class embedding vector.

[0079] The AAM loss function maximizes the cosine similarity between the sample and its corresponding category while minimizing the cosine similarity with other categories, which can encourage the model to generate more dispersed speaker embeddings, thereby improving the accuracy of speaker classification.

[0080] Preferably, after the model training is completed, the model is evaluated using the validation set and the test set to ensure that the model has good generalization ability, and the model is adjusted and optimized as necessary.

[0081] S3. According to the trained deep neural network speaker classification model, the input speech data is subjected to a forward neural network calculation to obtain the posterior probability of the speaker label;

[0082] The S3 step includes the following steps (see Figure 4 shown):

[0083] S31, inputting the speech data into the trained deep neural network speaker classification model;

[0084] S32, the deep neural network speaker classification model processes the input speech data through a forward propagation algorithm, and extracts features and performs classification using parameters learned by the model in the training phase;

[0085] S33. In the output layer of the deep neural network speaker classification model, a softmax activation function is used to convert the original output logits into a probability distribution; the expression of the softmax activation function is:

[0086]

[0087] In formula (2), P(C = k|x) is the posterior probability that the speaker label is category k given the input x;

[0088] f k (x) is the original prediction value of the model for sample x in category k, that is, logits;

[0089] Denominator It is the sum of the index values ​​of all categories, ensuring that the sum of the probabilities of all categories is 1;

[0090] S34, through steps S31-S33, the deep neural network speaker classification model generates a probability distribution for each input sample, and the probability distribution covers all possible speaker labels;

[0091] S35. According to the probability distribution, select the label with the highest probability as the speaker label predicted by the model;

[0092] S36. The deep neural network speaker classification model outputs the speaker label and its posterior probability corresponding to each input speech data.

[0093] S4. Obtain a speaker label corresponding to the input speech according to the posterior probability, and output the gender, age, and accent information corresponding to the speaker label.

[0094] The method of outputting the gender, age, and accent information corresponding to the speaker label includes:

[0095] Based on the mapping relationship between labels and features established in the model training and database construction stages, the speaker labels are converted into specific gender, age, and accent information, and the gender, age, and accent information are output.

[0096] Preferably, the trained deep neural network speaker classification model is deployed in an actual application environment to perform speaker classification on new speech data, and predict the corresponding gender, age, and accent information based on the speaker classification.

[0097] Figure 1 The basic process of the gender, age and accent joint detection method based on speaker classification in this embodiment is shown.

[0098] The present invention also provides a gender, age, and accent joint detection system based on speaker classification, which implements the gender, age, and accent joint detection method based on speaker classification as described above, including:

[0099] The speaker database building module is used to collect the voice data of a set number of speakers, and record the gender, age, and accent-related information of the speakers to build a speaker database, wherein the speaker database includes: the speaker's voice data and the corresponding speaker tag information; the speaker tag information includes: the speaker's gender, age, and accent information;

[0100] Constructing a deep neural network speaker classification model module: for constructing a deep neural network speaker classification model based on the speaker database using the collected speech data and the corresponding speaker label information, and using the additional angle margin loss function AAM to train the deep neural network speaker classification model;

[0101] Calculate posterior probability module: used to perform forward neural network calculation on the input speech data according to the trained deep neural network speaker classification model to obtain the posterior probability of the speaker label;

[0102] Output speaker label corresponding information module: used to obtain the speaker label corresponding to the input speech according to the posterior probability, and output the gender, age, and accent information corresponding to the speaker label.

[0103] Furthermore, the module for calculating the posterior probability includes:

[0104] Model input unit: used to input speech data into the trained deep neural network speaker classification model;

[0105] Forward propagation unit: used by the deep neural network speaker classification model to process the input speech data through the forward propagation algorithm, and extract features and perform classification using the parameters learned by the model during the training phase;

[0106] Output layer activation unit: used in the output layer of the deep neural network speaker classification model, using the softmax activation function to convert the original output logits into a probability distribution;

[0107] Obtaining posterior probability unit: used to generate a probability distribution for each input sample by the deep neural network speaker classification model through steps S31-S33, and the probability distribution covers all possible speaker labels;

[0108] A label selection unit: used for selecting the label with the highest probability as the speaker label predicted by the model according to the probability distribution;

[0109] Output label unit: used by the deep neural network speaker classification model to output the speaker label and its posterior probability corresponding to each input speech data.

[0110] The module for outputting speaker label corresponding information includes:

[0111] Gender, age, and accent information mapping unit: used to convert the speaker label into specific gender, age, and accent information based on the mapping relationship between labels and features established in the model training and database construction stages, and output the gender, age, and accent information.

[0112] Application Examples

[0113] Assume that the speaker labels output by the model correspond to the following mapping:

[0114] Speaker label 1: male, 20-30 years old, Mandarin accent;

[0115] Speaker label 2: female, 30-40 years old, Cantonese accent;

[0116] Speaker label 3: male, over 40 years old, Jiangxi accent;

[0117] Now there is a speech sample. After being processed by the model, the following probability distribution of speaker labels is obtained:

[0118] Speaker label 1: probability 0.8;

[0119] Speaker label 2: probability 0.15;

[0120] Speaker label 3: probability 0.05;

[0121] Based on these probabilities, the following steps are performed:

[0122] 1. Select the most likely speaker label: Since speaker label 1 has the highest probability (0.8), speaker label 1 is selected as the predicted speaker label.

[0123] 2. Mapping to specific speaker information: According to the mapping relationship, speaker label 1 corresponds to "male, 20-30 years old, Mandarin accent".

[0124] 3. Output: The system finally outputs the following information: "The speaker of this voice sample is a male, aged between 20 and 30, with a Mandarin accent."

[0125] 4. Application results: The output speaker information can be used in intelligent customer service systems to provide more personalized services. For example, the system selects a young male customer service representative to communicate with customers in Mandarin.

[0126] An embodiment of the present invention further provides a computer device, Figure 5 is a schematic diagram of the structure of a computer device provided by an embodiment of the present invention; see the accompanying drawings Figure 5 As shown, the computer device includes: an input device 23, an output device 24, a memory 22 and a processor 21; the memory 22 is used to store one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the gender, age, and accent joint detection method based on speaker classification as provided in the above embodiment; wherein the input device 23, the output device 24, the memory 22 and the processor 21 can be connected by a bus or other means, Figure 5 The example of connecting through bus is taken in the following.

[0127] The memory 22 is a readable and writable storage medium of a computing device, which can be used to store software programs and computer executable programs, such as program instructions corresponding to the gender, age, and accent joint detection method based on speaker classification described in the embodiment of the present invention; the memory 22 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required for a function; the data storage area can store data created according to the use of the device, etc.; in addition, the memory 22 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device; in some instances, the memory 22 can further include a memory remotely arranged relative to the processor 21, and these remote memories can be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0128] The input device 23 may be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the device; the output device 24 may include a display device such as a display screen.

[0129] The processor 21 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 22, that is, realizes the above-mentioned gender, age and accent joint detection method based on speaker classification.

[0130] The computer device provided above can be used to execute the gender, age, and accent joint detection method based on speaker classification provided in the above embodiment, and has corresponding functions and beneficial effects.

[0131] The embodiment of the present invention also provides a storage medium containing computer executable instructions, which are used to perform the gender, age, and accent joint detection method based on speaker classification as provided in the above embodiment when executed by a computer processor. The storage medium is any of various types of memory devices or storage devices, and the storage medium includes: installation media, such as CD-ROM, floppy disk or tape device; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disk or optical storage); registers or other similar types of memory elements, etc.; the storage medium may also include other types of memory or combinations thereof; in addition, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system, which is connected to the first computer system via a network (such as the Internet); the second computer system may provide program instructions to the first computer for execution. The storage medium includes two or more storage media that can reside in different locations (for example, in different computer systems connected via a network). The storage medium can store program instructions (for example, specifically implemented as a computer program) that can be executed by one or more processors.

[0132] Of course, the storage medium containing computer executable instructions provided in an embodiment of the present invention is not limited to the gender, age, and accent joint detection method based on speaker classification as described in the above embodiment, and can also execute related operations in the gender, age, and accent joint detection method based on speaker classification provided in any embodiment of the present invention.

[0133] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

[0134] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for joint detection of gender, age and accent based on speaker classification, characterized in that: The following steps are involved: S1. Collecting speech data of a set number of speakers, and recording the gender, age, and accent-related information of the speakers, to construct a speaker database, wherein the speaker database includes: the speech data of the speakers and the corresponding speaker label information; the speaker label information includes: the gender, age, and accent information of the speakers; S2. Using the collected speech data and the corresponding speaker label information, a deep neural network speaker classification model based on the speaker database is constructed, and the additional angular margin loss function AAM is used to train the deep neural network speaker classification model; the additional angular margin loss function is expressed as: (1) In formula (1), is the cosine similarity angle between the embedding vector output by the model and the embedding vector corresponding to the true speaker category, is the angular margin parameter, y It is an index of real speaking human categories. is the cosine similarity angle between the model output and each category embedding vector; S3. According to the trained deep neural network speaker classification model, the input speech data is subjected to a forward neural network calculation to obtain the posterior probability of the speaker label; S4, obtaining a speaker label corresponding to the input speech according to the posterior probability, and outputting the gender, age, and accent information corresponding to the speaker label; The S3 step includes the following steps: S31, inputting the speech data into the trained deep neural network speaker classification model; S32, the deep neural network speaker classification model processes the input speech data through a forward propagation algorithm, and extracts features and performs classification using parameters learned by the model in the training phase; S33. In the output layer of the deep neural network speaker classification model, a softmax activation function is used to convert the original output logits into a probability distribution; the expression of the softmax activation function is: (2) In formula (2), is the posterior probability that the speaker label is category k given the input x; It is the original prediction value of the model for sample x in category k, that is, logits; Denominator It is the sum of the index values ​​of all categories, ensuring that the sum of the probabilities of all categories is 1; S34, through steps S31-S33, the deep neural network speaker classification model generates a probability distribution for each input sample, and the probability distribution covers all possible speaker labels; S35. According to the probability distribution, select the label with the highest probability as the speaker label predicted by the model; S36, the deep neural network speaker classification model outputs the speaker label and its posterior probability corresponding to each input voice data; The method of outputting the gender, age, and accent information corresponding to the speaker tag in step S4 includes: Based on the mapping relationship between labels and features established in the model training and database construction phases, the speaker labels are converted into specific gender, age, and accent information, and the gender, age, and accent information are output; The S4 step further includes: Deploy the trained deep neural network speaker classification model to the actual application environment to classify the speakers of new speech data and predict the corresponding gender, age, and accent information based on the speaker classification; The deep neural network speaker classification model used in step S2 includes: any one of the ECAPA-TDNN model and the UIS-RNN model; The UIS-RNN model establishes a recurrent neural network (RNN) for each speaker, which is trained through supervised learning and can be continuously updated and adapted to new speakers.

2. A system for joint detection of gender, age and accent based on speaker classification, which implements the method for joint detection of gender, age and accent based on speaker classification as claimed in claim 1, characterized in that: include: The speaker database building module is used to collect the voice data of a set number of speakers, and record the gender, age, and accent-related information of the speakers to build a speaker database, wherein the speaker database includes: the speaker's voice data and the corresponding speaker tag information; the speaker tag information includes: the speaker's gender, age, and accent information; Constructing a deep neural network speaker classification model module: for constructing a deep neural network speaker classification model based on the speaker database using the collected speech data and the corresponding speaker label information, and using the additional angle margin loss function AAM to train the deep neural network speaker classification model; Calculate posterior probability module: used to perform forward neural network calculation on the input speech data according to the trained deep neural network speaker classification model to obtain the posterior probability of the speaker label; Output speaker label corresponding information module: used to obtain the speaker label corresponding to the input speech according to the posterior probability, and output the gender, age, and accent information corresponding to the speaker label.

3. The gender, age and accent joint detection system based on speaker classification according to claim 2 is characterized in that: The module for calculating posterior probability comprises: Model input unit: used to input speech data into the trained deep neural network speaker classification model; Forward propagation unit: used by the deep neural network speaker classification model to process the input speech data through the forward propagation algorithm, and extract features and perform classification using the parameters learned by the model during the training phase; Output layer activation unit: used in the output layer of the deep neural network speaker classification model, using the softmax activation function to convert the original output logits into a probability distribution; Obtaining posterior probability unit: used to generate a probability distribution for each input sample by the deep neural network speaker classification model through steps S31-S33, and the probability distribution covers all possible speaker labels; A label selection unit: used for selecting the label with the highest probability as the speaker label predicted by the model according to the probability distribution; Output label unit: used by the deep neural network speaker classification model to output the speaker label and its posterior probability corresponding to each input speech data.

4. The gender, age and accent joint detection system based on speaker classification according to claim 2 is characterized in that: The module for outputting speaker label corresponding information includes: Gender, age, and accent information mapping unit: used to convert the speaker label into specific gender, age, and accent information based on the mapping relationship between labels and features established in the model training and database construction stages, and output the gender, age, and accent information.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for joint detection of gender, age and accent based on speaker classification described in claim 1 are implemented.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for joint detection of gender, age, and accent based on speaker classification as described in claim 1 are implemented.

Citation Information

Patent Citations

  • Accent classification method based on deep neural network and model thereof

    CN112992119A

  • Short-time voice speaker recognition system and method based on deep learning

    CN114822559A