Preprocessing depolarization method for speech recognition system
By using undersampling and SMOTE technology to process the data in the speech recognition system, and combining the bias evaluation method, the bias problem caused by the unbalanced data set in the speech recognition system is solved, achieving higher recognition accuracy and fairness.
Patent Information
- Application Number
- CN202510075782.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively eliminate the bias problem caused by unbalanced data sets in speech recognition systems, resulting in poor performance in the recognition accuracy of a few types of samples, and the debiasing method is difficult to ensure the recognition accuracy.
A preprocessing and debiasing method for speech recognition systems is proposed. Multi-class and minor-class data are sampled through undersampling and SMOTE technology to form a balanced data set, and the discriminator and ASR are trained, combining bias evaluation methods and evaluation indicators to achieve comprehensive bias evaluation.
Effectively reduce or eliminate biases in the data, improve the model's recognition accuracy of different attribute groups, and significantly improve the fairness, accuracy and credibility of the speech recognition system.
Smart Images

Figure CN119964557A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and aims to provide a preprocessing debiasing method for a speech recognition system. Background Art
[0002] Artificial intelligence is developing rapidly and gradually penetrating into all aspects of life. It plays an increasingly important role in influencing personal decision-making applications, including image classification, job recruitment, loan applications, etc., bringing extensive benefits to human society. As a typical application of artificial intelligence, the widespread application of speech recognition is revolutionizing the way people of all ages and genders live and work:
[0003] (1) In the healthcare sector, doctors and nurses can use speech recognition to record medical records, develop treatment plans, and even perform surgical operations;
[0004] (2) In the automotive industry, drivers can use voice recognition to control vehicles, navigate, make calls, and obtain real-time traffic information;
[0005] (3) In the education industry, speech recognition can help students practice pronunciation, improve language skills, etc.
[0006] However, while AI brings benefits, it also faces constant moral, ethical and fairness issues. Because of the use of unequal data or improper algorithm design, AI may be affected by certain socio-demographic factors (such as race, gender and age, etc.), leading to discrimination and prejudice against minority groups. Bias in speech recognition systems can also have a profound impact on society. As the use of speech recognition systems in various fields continues to increase, when speech recognition systems are biased towards a particular group, people who do not belong to that group may suffer unfair treatment during use, resulting in problems such as recognition errors, poor communication and information distortion. This unfairness may undermine social and political stability and trigger a crisis of confidence in AI.
[0007] Unequal training data sets are an important source of bias in speech recognition systems. When multi-class data dominates, the model tends to overfit the speech features of the multi-class group, thereby ignoring the minority group. This bias will significantly reduce the recognition accuracy of minority data, further exacerbating the unfairness. In addition, data imbalance will also limit the generalization ability of the model. Especially in practical applications, the model may not be able to effectively recognize the speech of the minority group, or even mistake it for noise. Overall, unbalanced training data sets not only weaken the fairness of the speech recognition system, but may also deepen the technological gap between different groups and bring unfair results.
[0008] Existing technologies are still insufficient in the research of debiasing methods for speech recognition systems. There is a lack of systematic and targeted solutions, and it is difficult to effectively eliminate bias problems caused by sensitive attributes such as gender and race. At the same time, existing methods often cannot guarantee recognition accuracy during the debiasing process, which can easily lead to a decline in model performance, especially in the recognition accuracy of minority samples. How to eliminate the bias of speech recognition systems caused by unbalanced data sets is an urgent problem to be solved. Summary of the invention
[0009] In order to solve the above problems in the prior art, the present invention proposes a preprocessing debiasing method for a speech recognition system, comprising:
[0010] S1, data sampling;
[0011] S2, model training;
[0012] S3, classification and identification;
[0013] S4, bias assessment;
[0014] S5. Comparative analysis.
[0015] Preferably, the process of sampling unbalanced data includes:
[0016] Step 1: Calculate the ratio C between multi-class data A1 and minority data B1, so as to set the ratio of undersampling and SMOTE so that the amount of data after sampling is balanced;
[0017] Step 2: Undersample the multi-class data A1 to obtain multi-class data A2;
[0018] Step 3: Perform SMOTE on the minority class data B1 to obtain the minority class data B2.
[0019] Furthermore, preprocessing debiasing is a technique used in machine learning and artificial intelligence systems. It aims to reduce or eliminate bias in the data by debiasing the data before it is input into the model, so as to prevent the bias from being further amplified in subsequent model training. Multi-class and few-class refer to samples of different categories under the same sensitive attribute, in which one sample has more data and the other has less data. Sensitive attributes include age, race, gender, disease, and religious beliefs.
[0020] Furthermore, undersampling refers to randomly selecting and deleting a set proportion of samples. The SMOTE process includes:
[0021] Step 1: Initialize the new dataset to be empty to store the generated synthetic minority class samples;
[0022] Step 2: Select the number of synthetic samples to be generated, that is, according to the imbalance of the data set, determine how many new minority class samples need to be generated to achieve the expected balance ratio;
[0023] Step 3: Select the nearest neighbor for each minority class sample. For each minority class sample, calculate the Euclidean distance between it and other minority class samples, and then select the k nearest neighbor minority class samples of each sample according to the set number of neighbors (usually K = 5);
[0024] Step 4: Randomly select a neighbor sample from the k nearest neighbors selected in step 3 to generate a new sample. The formula is:
[0025] New sample = original sample + λ × (neighborhood sample - original sample)
[0026] Among them, λ is a random number between 0 and 1, which is used to control the position of the new sample to ensure that it is located at a random point between the original sample and the neighbor sample;
[0027] Step 5: Repeat steps 3 and 4 according to the number of synthetic samples that need to be generated in step 2 until a sufficient number of synthetic minority class samples are generated;
[0028] Step 6: Merge the generated synthetic minority class samples with the samples in the original dataset to form a new, more balanced dataset.
[0029] Preferably, the models to be trained include the discriminator and ASR, and the training process includes:
[0030] Step 1: Train the baseline ASR using the initial multi-class data A1 and the few-class data B1;
[0031] Step 2: Use the under-sampled multi-class data A2 to train the A-class ASR;
[0032] Step 3: Use the minority data B2 after SMOTE to train the B-class ASR;
[0033] Step 4: Train the discriminator using the multi-class data A2 and the few-class data B2.
[0034] Furthermore, the discriminator is trained using CNN to identify specific attributes of the input speech data, such as gender, age, etc. Taking gender attributes as an example, the main process is: first, extract features such as MFCC from the speech to capture the spectral information of the speech; then, use the CNN model to automatically extract local features through the convolution layer, and classify through the fully connected layer, and finally output whether the speech belongs to male or female.
[0035] Furthermore, ASR refers to the technology of converting speech signals into text, and its main process is: first, the system accepts speech signals as input; then, the speech signals are preprocessed and noise is removed; then, the speech signals are converted into feature vectors using feature extraction methods; finally, the features are decoded based on machine learning algorithms to identify the corresponding text.
[0036] Preferably, the classification and recognition process is: first, using a discriminator to discriminate the input voice data and output its corresponding category, such as male or female; then, selecting the corresponding ASR for recognition.
[0037] Preferably, bias assessment includes construction of an assessment dataset, setting of an assessment method and an assessment indicator.
[0038] Furthermore, the construction process of the evaluation data set (taking gender as an example) is as follows:
[0039] Step 1: Obtain a mainstream speech dataset, such as TIMIT;
[0040] Step 2: Data standardization and cleaning;
[0041] Step 3: Extract data from different datasets in proportion;
[0042] Step 4: Define matching rules to ensure that the extracted male and female voice data match on other attributes (such as age, race, etc.);
[0043] Step 5: Based on step 4, construct an evaluation dataset for the sensitive attribute of gender.
[0044] The matching rule uses Euclidean distance to measure the similarity of attributes other than gender, and uses a greedy algorithm to match samples. The process of the greedy algorithm is: for each male sample, select the closest female sample to ensure that the difference in attributes other than gender is minimal; if the match is successful, match the male sample with the female sample; after each pairing, mark the matched samples to avoid repeated matching. The Euclidean distance formula is
[0045]
[0046] Among them, R1 and R2 are other attributes 1 of sample 1 and sample 2, and Y1 and Y2 are other attributes 2 of sample 1 and sample 2.
[0047] Furthermore, the evaluation method process is as follows: First, use ASR to recognize the speech data, obtain the recognized text, and based on the original question and the recognized question of the speech data, obtain the WER of each piece of data, count the WER of all data, and calculate the UWER (Universal Word Error Rate); then, make a judgment according to the recognized text and WER of each piece of speech. Having text and WER < UWER is TP (True Positive), not getting text or getting text but WER > UWER is FN (False Negative); finally, calculate the evaluation metrics and conduct bias evaluation. The calculation formula of WER is:
[0048]
[0049] Among them, S represents the number of substitution operations, D represents the number of deletion operations, I represents the number of insertion operations, N represents the total number of words in the reference text, and a low value of WER indicates more accurate recognition by the speech recognition system. The calculation formula of UWER is:
[0050]
[0051] Among them, n represents the number of speech data in the dataset, WER represents the word error rate corresponding to the speech data, and i represents the i-th piece of data in the dataset. Through the universal word error rate, it can be determined whether each piece of speech data is successfully recognized and predicted; when the word error rate of a piece of speech data is lower than UWER, it indicates that the speech recognition prediction is successful; when the word error rate of a piece of speech data is higher than UWER, it indicates that the speech recognition prediction fails.
[0052] Furthermore, the evaluation metrics are divided into sensitivity analysis, statistical parity, equality of opportunity, and positive predictive value.
[0053] Sensitivity analysis refers to observing whether there are obvious unfair phenomena by comparing the performance of the algorithm on different groups (such as gender, race, etc.). For the speech recognition system, the fairness of the system can be evaluated by calculating the average word error rate of each category. By comparing the average WER values of different categories, we can quantify the performance differences of different groups in the speech recognition process, and then identify whether there is bias or unfairness in ASR.
[0054] Statistical parity, also known as demographic parity, is one of the important metrics for evaluating fairness in the fields of machine learning and artificial intelligence. It means that in statistical data analysis, it is ensured that there are no obvious inequalities or discriminatory differences in the data distribution between different groups. The statistical parity formula is:
[0055]
[0056] P represents probability, represents the prediction variable (whether the result of the speech recognition system is predicted successfully), with a value of 1 indicating a successful prediction and a value of 0 indicating a failed prediction. A represents a binary sensitive attribute, including gender, age, and race, with a value of 1 indicating one category under the sensitive attribute and a value of 0 indicating another category under the sensitive attribute. Equal opportunity is a specific case of equal opportunity, and is a fairness measure designed to ensure that the speech recognition system has the same recognition accuracy for different groups when they are actually positive examples. The formula for equal opportunity is:
[0057]
[0058] Y represents the actual variable, with a value of 1 for human voice and a value of 0 for non-human voice.
[0059] Positive predictive value is a key indicator for evaluating model accuracy and effectiveness. When the performance of the model is inconsistent in different subgroups, these indicators can help identify the bias of the model. The bias evaluation of the speech recognition system focuses more on the actual positive examples, so the TPR and FNR in the positive predictive value are used. TPR is also called recall rate, which refers to the proportion of all positive samples that the model correctly identifies as positive. It is used to measure the model's ability to recognize positive samples. The formula is:
[0060]
[0061] FNR refers to the proportion of all positive classes that the model mistakenly identifies as negative classes, which measures the false negative rate of the model. The formula is:
[0062]
[0063] Preferably, the bias assessment results of the baseline ASR and the debiased model are compared and analyzed.
[0064] Beneficial effects of the present invention:
[0065] The present invention proposes a preprocessing debiasing method for speech recognition system. First, multi-class data and minority data are sampled and processed respectively by undersampling and SMOTE, and then the original data and the sampled data are used to train the discriminator and ASR respectively, and then the trained discriminator and ASR are used for subsequent classification and recognition. Then, by constructing a balanced and fair speech evaluation data set, a bias evaluation method and evaluation index for speech recognition system are proposed, so as to achieve a comprehensive bias evaluation. Finally, the effectiveness of the debiasing method is evaluated by comparing and analyzing the performance of the benchmark ASR model trained with unbalanced data and the model after debiasing in the bias evaluation. Compared with the existing debiasing method, this method is designed according to the characteristics of the speech recognition system, and provides a more complete and efficient solution in data processing and evaluation methods. It can not only accurately handle the bias problem in speech data, but also effectively improve the recognition accuracy of the model for different attribute groups, thereby significantly improving the fairness, accuracy and credibility of the speech recognition system in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a framework structure diagram proposed by the present invention;
[0067] Figure 2 is a flow chart of bias assessment of a speech recognition system proposed by the present invention;
[0068] Figure 3 It is the overall flow chart proposed by the present invention. DETAILED DESCRIPTION
[0069] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] The present invention provides a preprocessing debiasing method for a speech recognition system, the framework is as follows Figure 1 As shown, the method includes five stages: data sampling, model training, classification identification, bias assessment and comparative analysis.
[0071] In this embodiment, a preprocessing debiasing method for a speech recognition system is provided. The overall process is as follows: Figure 3 As shown, the method includes:
[0072] S1, data sampling;
[0073] S2, model training;
[0074] S3, classification and identification;
[0075] S4, bias assessment;
[0076] S5. Comparative analysis.
[0077] The process of obtaining imbalanced data includes: collecting data sets from official and authoritative platforms; screening according to the sensitive attribute annotations in the data sets, and selecting data with the same sensitive attributes; under the same sensitive attribute, giving priority to obtaining data of a certain category to form multi-category data; at the same time, obtaining a smaller amount of data of another category to form a small number of categories of data. The above-mentioned sensitive attributes refer to attributes that can reveal individual or group characteristics in data analysis, machine learning or social science research. These characteristics usually have certain social, moral or legal significance. Common sensitive attributes include gender, race, age and religion, etc.
[0078] Classifying imbalanced data includes: calculating the ratio C between multi-class data A1 and minority data B1; based on the ratio, undersampling multi-class data A1 to obtain multi-class data A2; based on the ratio, performing SMOTE on minority data B1 to obtain minority data B2. The above undersampling refers to randomly deleting samples of a set ratio. The above SMOTE process includes:
[0079] Step 1: Initialize the new dataset to be empty to store the generated synthetic minority class samples;
[0080] Step 2: Select the number of synthetic samples to be generated, that is, according to the imbalance of the data set, determine how many new minority class samples need to be generated to achieve the expected balance ratio;
[0081] Step 3: Select the nearest neighbor for each minority class sample. For each minority class sample, calculate the Euclidean distance between it and other minority class samples, and then select the k nearest neighbor minority class samples of each sample according to the set number of neighbors (usually K = 5);
[0082] Step 4: Randomly select a neighbor sample from the k nearest neighbors selected in step 3, using the formula new sample = original sample + λ × (neighbor sample - original sample)
[0083] Generate new samples, where λ is a random number between 0 and 1, which is used to control the position of the new sample and ensure that it is located at a random point between the original sample and the neighboring sample;
[0084] Step 5: Repeat steps 3 and 4 according to the number of synthetic samples to be generated determined in step 2 until a sufficient number of synthetic minority class samples are generated;
[0085] Step 6: Merge the generated synthetic minority class samples with the samples in the original dataset to form a new, more balanced dataset.
[0086] In this embodiment, the discriminator and ASR are trained using the sampled data. The discriminator is trained using CNN to identify specific attributes of the input voice data, such as gender, age, etc. Taking gender attributes as an example, the main process is: first, extract features such as MFCC from the voice to capture the spectral information of the voice; then, use the CNN model to automatically extract local features through the convolution layer, and classify through the fully connected layer, and finally output whether the voice belongs to male or female. ASR refers to the technology of converting voice signals into text, and its main process is: first, the system accepts voice signals as input; then, preprocesses and removes noise from the voice signal; then, uses feature extraction methods to convert the voice signal into feature vectors; finally, decodes the features based on the machine learning algorithm to identify the corresponding text. The specific process of training is:
[0087] Step 1: Train the baseline ASR using the initial multi-class data A1 and the few-class data B1;
[0088] Step 2: Use the under-sampled multi-class data A2 to train the A-class ASR;
[0089] Step 3: Use the minority data B2 after SMOTE to train the B-class ASR;
[0090] Step 4: Train the discriminator using the multi-class data A2 and the few-class data B2.
[0091] In this embodiment, classification and recognition are performed based on the trained discriminator and ASR, and the process is as follows: first, the discriminator is used to discriminate the input voice data and output its corresponding category, such as male or female; then the corresponding ASR is selected for recognition.
[0092] In this embodiment, bias evaluation is performed on the baseline ASR and the debiased model, and the bias evaluation includes the construction of an evaluation dataset, an evaluation method, and the setting of evaluation indicators.
[0093] Preferably, the construction process of the evaluation data set (taking gender as an example) is:
[0094] Step 1: Obtain a mainstream speech dataset, such as TIMIT;
[0095] Step 2: Data standardization and cleaning;
[0096] Step 3: Extract data from different datasets in proportion;
[0097] Step 4: Define matching rules to ensure the matching of the extracted male and female voice data in other attributes (such as age, race, etc.);
[0098] Step 5: According to Step 4, construct an evaluation data set for the sensitive attribute of gender.
[0099] The matching rules use the Euclidean distance to measure the similarity of attributes other than gender and use the greedy algorithm to perform sample matching. The process of the greedy algorithm is as follows: for each male sample, select the female sample with the closest distance to ensure the smallest difference in attributes other than gender; if the matching is successful, match this male sample and the female sample; after each pairing, mark the matched samples to avoid repeated matching. The Euclidean distance formula is
[0100]
[0101] where R1 and R2 are other attribute 1 of sample 1 and sample 2, and Y1 and Y2 are other attribute 2 of sample 1 and sample 2.
[0102] Preferably, the evaluation method is as Figure 2 shown. The process is as follows: First, use ASR to identify the voice data to obtain the recognized text. Based on the original question and the recognized question of the voice data, obtain the WER of each piece of data, count the WER of all data, and calculate the UWER (Universal Word Error Rate); then, make a judgment according to the recognized text and WER of each voice. Having text and WER < UWER is TP (True Positive), not getting text or getting text but WER > UWER is FN (False Negative); finally, calculate the evaluation index and conduct bias evaluation. The calculation formula of WER is:
[0103]
[0104] where S represents the number of substitution operations, D represents the number of deletion operations, I represents the number of insertion operations, N represents the total number of words in the reference text, and a low value of WER indicates more accurate recognition by the speech recognition system. The calculation formula of UWER is:
[0105]
[0106] Among them, n represents the number of voice data in the data set, WER represents the word error rate corresponding to the voice data, and i represents the number of data in the data set. The global word error rate can be used to determine whether each voice data is successfully recognized and predicted; when the word error rate of a voice data is lower than UWER, it indicates that the voice recognition prediction is successful; when the word error rate of a voice data is higher than UWER, it indicates that the voice recognition prediction fails.
[0107] Preferably, the evaluation indicators are divided into sensitivity analysis, statistical equality, equality of opportunity and positive predictive value.
[0108] Sensitivity analysis refers to comparing the performance of algorithms on groups of different categories (such as gender, race, etc.) to observe whether there is obvious unfairness. For speech recognition systems, the fairness of the system can be evaluated by calculating the average word error rate for each category. By comparing the average WER values of different categories, we can quantify the performance differences between different groups in the speech recognition process and then identify whether there is bias or unfairness in ASR.
[0109] Statistical equality, also known as demographic equality, is one of the important indicators used to evaluate fairness in the field of machine learning and artificial intelligence. It refers to ensuring that there is no obvious inequality or discriminatory difference in the distribution of data between different groups in statistical data analysis. The formula for statistical equality is:
[0110]
[0111] P represents probability, represents the prediction variable (whether the result of the speech recognition system is predicted successfully), with a value of 1 indicating a successful prediction and a value of 0 indicating a failed prediction. A represents a binary sensitive attribute, including gender, age, and race, with a value of 1 indicating one category under the sensitive attribute and a value of 0 indicating another category under the sensitive attribute. Equal opportunity is a specific case of equal opportunity, and is a fairness measure designed to ensure that the speech recognition system has the same recognition accuracy for different groups when they are actually positive examples. The formula for equal opportunity is:
[0112]
[0113] Y represents the actual variable, with a value of 1 for human voice and a value of 0 for non-human voice.
[0114] Positive predictive value is a key indicator for evaluating model accuracy and effectiveness. When the performance of the model is inconsistent in different subgroups, these indicators can help identify model bias issues. The bias assessment of the speech recognition system focuses more on the actual positive examples, so the TPR (True Positive Rate) and FNR (False Negative Rate) in the positive predictive value are used. TPR is also called recall rate, which refers to the proportion of all positive samples that the model correctly identifies as positive. It is used to measure the model's ability to recognize positive samples. The formula is:
[0115]
[0116] FNR refers to the proportion of all positive classes that the model mistakenly identifies as negative classes, which measures the false negative rate of the model. The formula is:
[0117]
[0118] In this embodiment, the bias assessment results of the baseline ASR and the debiased model are compared and analyzed.
[0119] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A preprocessing debiasing method for a speech recognition system, characterized in that: include: S1, data sampling; S2, model training; S3, classification and identification; S4, bias assessment; S5. Comparative analysis.
2. A preprocessing debiasing method for a speech recognition system according to claim 1, characterized in that: Preprocessing debiasing is a technique used in machine learning and artificial intelligence systems. It aims to reduce or eliminate bias in the data by debiasing the data before the training data is input into the model, so as to prevent the bias from being further amplified in subsequent model training. Multi-class and minority-class refer to samples of different categories under the same sensitive attribute, in which one sample has more data and the other has less data. Sensitive attributes include age, race, gender, disease, and religious beliefs. Sampling methods include undersampling and synthetic minority oversampling technique (SMOTE). The process of sampling data includes: Step 1: Calculate the ratio C between multi-class data A1 and minority data B1, so as to set the ratio of undersampling and SMOTE so that the amount of data after sampling is balanced; Step 2: Undersample the multi-class data A1 to obtain multi-class data A2; Step 3: Perform SMOTE on the minority class data B1 to obtain the minority class data B2.
3. A preprocessing debiasing method for a speech recognition system according to claim 2, characterized in that: Undersampling refers to randomly selecting and deleting samples of a set proportion. The SMOTE process includes: Step 1: Initialize the new dataset to be empty to store the generated synthetic minority class samples; Step 2: Select the number of synthetic samples to be generated, that is, according to the imbalance of the data set, determine how many new minority class samples need to be generated to achieve the expected balance ratio; Step 3: Select the nearest neighbor for each minority class sample. For each minority class sample, calculate the Euclidean distance between it and other minority class samples, and then select the k nearest neighbor minority class samples of each sample according to the set number of neighbors (usually K = 5); Step 4: Randomly select a neighbor sample from the k nearest neighbors selected in step 3 to generate a new sample. The formula is: New sample = original sample + λ × (neighborhood sample - original sample) Among them, λ is a random number between 0 and 1, which is used to control the position of the new sample to ensure that it is located at a random point between the original sample and the neighbor sample; Step 5: According to the number of synthetic samples required to be generated in step 2, repeat steps 3 and 4 until a sufficient number of synthetic minority class samples are generated; Step 6: Merge the generated synthetic minority class samples with the samples in the original dataset to form a new, more balanced dataset.
4. The preprocessing debiasing method for a speech recognition system according to claim 1, characterized in that: The models that need to be trained include the discriminator and automatic speech recognition (ASR). The training process includes: Step 1: Train the baseline ASR using the initial multi-class data A1 and the few-class data B1; Step 2: Use the under-sampled multi-class data A2 to train the A-class ASR; Step 3: Use the minority data B2 obtained by SMOTE to train the B-class ASR; Step 4: Train the discriminator using the multi-class data A2 and the few-class data B2.
5. A preprocessing debiasing method for a speech recognition system according to claim 4, characterized in that: The discriminator is trained using Convolutional Neural Networks (CNN) to identify specific attributes of input speech data, such as gender and age. Taking gender attributes as an example, its main process is: first, extract features such as Mel Frequency Cepstrum Coefficient (MFCC) from the speech to capture the spectral information of the speech; then, use the CNN model to automatically extract local features through the convolution layer, and classify through the fully connected layer, and finally output whether the speech belongs to male or female. ASR refers to the technology of converting speech signals into text. Its main process is: first, the system accepts speech signals as input; then, the speech signals are preprocessed and noise removed; then, the speech signals are converted into feature vectors using feature extraction methods; finally, the features are decoded based on machine learning algorithms to identify the corresponding text.
6. A preprocessing debiasing method for a speech recognition system according to claim 1, characterized in that: The discrimination process is as follows: first, use the discriminator to discriminate the input voice data and output its corresponding category, such as male or female; then, select the corresponding ASR for recognition.
7. A preprocessing debiasing method for a speech recognition system according to claim 1, characterized in that: Bias assessment includes the construction of evaluation datasets, evaluation methods, and the setting of evaluation metrics.
8. A preprocessing debiasing method for a speech recognition system according to claim 7, characterized in that: Taking gender as an example, the construction process of the evaluation dataset is as follows: Step 1: Obtain a mainstream speech dataset, such as TIMIT; Step 2: Data standardization and cleaning; Step 3: Extract data from different datasets in proportion; Step 4: Define matching rules to ensure that the extracted male and female voice data match on other attributes (such as age, race, etc.); Step 5: Repeat step 4 to construct an evaluation dataset for the sensitive attribute of gender.
9. A preprocessing debiasing method for a speech recognition system according to claim 8, characterized in that: The matching rule uses Euclidean distance to measure the similarity of attributes other than gender, and uses a greedy algorithm to match samples. The process of the greedy algorithm is: for each male sample, select the closest female sample to ensure that the difference in attributes other than gender is minimal; if the match is successful, match the male sample with the female sample; after each pairing, mark the matched samples to avoid repeated matching. The Euclidean distance formula is: Among them, R1 and R2 are other attributes 1 of sample 1 and sample 2, and Y1 and Y2 are other attributes 2 of sample 1 and sample 2.
10. A preprocessing debiasing method for a speech recognition system according to claim 7, characterized in that: Traditional evaluation methods mainly focus on binary classification models, such as recommendation systems, and are insufficient in covering non-binary classification models. The speech recognition system is a typical non-binary classification model, and traditional evaluation methods are not well-suited. By introducing the concept of Word Error Rate (WER), redefining True Positive (TP) and False Negative (FN), a bias evaluation method more suitable for speech recognition systems is proposed. The method process is as follows: First, use ASR to recognize speech data to obtain the recognized text. Based on the original text and the recognized text of the speech data, obtain the WER of each piece of data, count the WER of all data, and calculate the Universal Word Error Rate (UWER); then, make a judgment according to the recognized text and WER of each piece of speech. Having text and WER < UWER is TP, not getting text or getting text but WER > UWER is FN; finally, calculate the evaluation metrics for bias evaluation. The calculation formula of WER is: Among them, S represents the number of substitution operations, D represents the number of deletion operations, I represents the number of insertion operations, N represents the total number of words in the reference text, and a low value of WER indicates more accurate recognition by the speech recognition system. The calculation formula of UWER is: Among them, n represents the number of speech data in the dataset, WER represents the word error rate corresponding to the speech data, and i represents the i-th piece of data in the dataset. Through the Universal Word Error Rate, it can be determined whether each piece of speech data is successfully recognized and predicted; when the word error rate of a piece of speech data is lower than UWER, it indicates that the speech recognition prediction is successful; when the word error rate of a piece of speech data is higher than UWER, it indicates that the speech recognition prediction fails.
11. A preprocessing debiasing method for a speech recognition system according to claim 7, characterized in that: The evaluation metrics are divided into sensitivity analysis, statistical parity, equality of opportunity, and positive predictive value.
12. A preprocessing debiasing method for a speech recognition system according to claim 11, characterized in that: Sensitivity analysis refers to observing whether there are obvious unfair phenomena by comparing the performance of the algorithm on different groups (such as gender, race, etc.). For the speech recognition system, the fairness of the system can be evaluated by calculating the average word error rate of each category. By comparing the average WER values of different categories, we can quantify the performance differences of different groups in the speech recognition process, and then identify whether there is bias or unfairness in ASR.
13. A preprocessing debiasing method for a speech recognition system according to claim 11, characterized in that: Statistical parity, also known as demographic parity, is one of the important metrics for evaluating fairness in the fields of machine learning and artificial intelligence. It means that in statistical data analysis, it is ensured that there are no obvious unequal or discriminatory differences in the data distribution among different groups. The formula for statistical parity is: Where P represents probability, Represents the prediction variable (whether the result of the speech recognition system is predicted successfully). A value of 1 indicates a successful prediction, and a value of 0 indicates a failed prediction. A Represents binary sensitive attributes, including gender, age, and race. Its value is 1 for one category under the sensitive attribute, and 0 for another category under the sensitive attribute. Equal opportunity is a fairness measure that aims to ensure that the speech recognition system has the same recognition accuracy for different groups when they are actually positive examples. The formula for equal opportunity is: in, Y Represents the actual variable, with a value of 1 for human voice and a value of 0 for non-human voice.
14. A preprocessing debiasing method for a speech recognition system according to claim 11, characterized in that: Positive predictive value is a key indicator for evaluating model accuracy and effectiveness. When the performance of the model is inconsistent in different subgroups, these indicators can help identify the bias of the model. The bias evaluation of the speech recognition system focuses more on the actual positive examples, so the true positive rate (True Positive Rate, TPR) and false negative rate (False Negative Rate, FNR) in the positive predictive value are used. TPR is also called recall rate, which refers to the proportion of all positive samples that the model correctly identifies as positive. It is used to measure the model's ability to recognize positive samples. The formula is: FNR refers to the proportion of all positive classes that the model mistakenly identifies as negative classes, which measures the false negative rate of the model. The formula is:
15. The preprocessing debiasing method for a speech recognition system according to claim 1, characterized in that: Comparative analysis of bias evaluation results of baseline ASR and debiased ASR.