Speech recognition model acquisition method and device, computer equipment, readable storage medium and program product

By cleaning and pseudo-label prediction of unlabeled speech training data, combining text error correction and speech correction, the pre-trained speech recognition model is adjusted, and the problem of dependence on manual labeled data in the existing technology is solved, and the training accuracy and generalization ability of the speech recognition model are improved.

CN120148488APending Publication Date: 2025-06-13GUANGZHOU QUYAN NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510270528.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing speech recognition models rely on a large amount of manual annotation data during training, which makes data acquisition time-consuming and laborious and expensive. Especially in specific fields or low-resource languages, there is a lack of annotation data, resulting in low generalization capabilities of the model.

Method used

By obtaining unlabeled speech training data, cleaning the data, using the pre-trained speech recognition model for pseudo-label prediction, then text correction and speech correction of the pseudo-label, and finally adjusting the pre-trained model based on the modified data to obtain the target speech recognition model.

Benefits of technology

This method can effectively reduce the dependence on manual labeled data, improve the training accuracy and generalization ability of speech recognition models, especially in the scenario of unlabeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148488A_ABST
    Figure CN120148488A_ABST
Patent Text Reader

Abstract

The invention relates to a voice recognition model acquisition method and device, computer equipment, a computer readable storage medium and a computer program product, relates to the technical field of voice recognition, and can improve the training precision of a voice recognition model and the generalization ability of the model in an unlabeled application scene. The method comprises the following steps: acquiring unmarked voice training data; performing data cleaning on the voice training data to obtain cleaned voice training data; obtaining a pre-training voice recognition model, and performing pseudo-tag prediction on the cleaned voice training data through the model to obtain first voice training data with a tag; performing text error correction on a pseudo tag in the first voice training data with the tag, and performing voice correction on voice training data in the first voice training data with the tag to obtain second voice training data with the tag; and adjusting the pre-trained speech recognition model according to the second speech training data with the label to obtain a target speech recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and particularly to a method, device, computer device, computer-readable storage medium, and computer program product for obtaining a speech recognition model. Background Art

[0002] With the continuous development of deep learning technology, the performance of speech recognition models has been significantly improved. Especially the application of large-scale pre-trained models has enabled speech recognition to show excellent effects in multiple scenarios.

[0003] However, the training and optimization of these models usually rely on a large amount of manually labeled data. The acquisition of labeled data is not only time-consuming and laborious but also costly. In speech recognition tasks for specific fields or low-resource languages, the problem of lack of labeled data is particularly prominent. As a result, when adjusting the pre-trained speech recognition model, the model makes initial prediction errors and has low confidence during the iterative training process, ultimately leading to poor generalization ability of the obtained speech recognition model. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for obtaining a speech recognition model.

[0005] In a first aspect, this application provides a method for obtaining a speech recognition model, including:

[0006] Obtain unlabeled speech training data;

[0007] Clean the speech training data to obtain the speech training data after data cleaning;

[0008] Obtain a pre-trained speech recognition model, and use the pre-trained speech recognition model to perform pseudo-label prediction on the speech training data after data cleaning to obtain the first speech training data with labels;

[0009] Correct the text of the pseudo-labels in the first speech training data with labels, and correct the speech of the speech training data in the first speech training data with labels to obtain the second speech training data with labels;

[0010] Adjust the pre-trained speech recognition model according to the second speech training data with labels to obtain the target speech recognition model.

[0011] In one of the embodiments, the adjusting the pre-trained speech recognition model according to the second speech training data with labels to obtain the target speech recognition model includes:

[0012] Adjust the pre-trained speech recognition model according to the labeled second speech training data to obtain an adjusted speech recognition model;

[0013] Perform pseudo-label prediction on the second speech training data through the adjusted speech recognition model to obtain updated pseudo-labels, and determine the difference value between the updated pseudo-labels and the pseudo-labels before updating;

[0014] In the case where the difference value is greater than a preset difference value threshold, return to execute the step of performing text correction and speech correction on the labeled first speech training data to continue adjusting the pre-trained speech recognition model until the difference value is less than the preset difference value threshold to obtain a target speech recognition model.

[0015] In one embodiment, the speech training data includes a speech part and a non-speech part;

[0016] The data cleaning of the speech training data to obtain data-cleaned speech training data includes:

[0017] Perform noise reduction processing on the speech training data to obtain noise-reduced speech training data;

[0018] Extract the data segments of the speech part in the noise-reduced speech training data, and obtain data-cleaned speech training data according to the extracted data segments of the speech part.

[0019] In one embodiment, the obtaining data-cleaned speech training data according to the extracted data segments of the speech part includes:

[0020] Obtain a speech quality evaluation model, and determine the speech quality score of the data segments of the speech part through the speech quality evaluation model;

[0021] Remove the data segments of the speech part with a speech quality score less than a preset speech quality score threshold to obtain data-cleaned speech training data.

[0022] In one embodiment, the text correction of the pseudo-labels in the labeled first speech training data includes:

[0023] Obtain a language model, and in the labeled first speech training data, determine the pseudo-labels to be corrected with text expressions that do not conform to the preset expression through the language model;

[0024] Perform correction on the text corresponding to the pseudo-labels to be corrected.

[0025] In one embodiment, the speech training data in the labeled first speech training data is corrected to obtain labeled second speech training data, including:

[0026] Calculating the perplexity of the pseudo-labels in the labeled first speech training data to obtain a perplexity calculation result;

[0027] Determining the labeled first speech training data corresponding to the pseudo-labels with the perplexity calculation result greater than a preset perplexity threshold as the labeled speech training data to be corrected;

[0028] Determining the text content of the pseudo-labels in the speech training data to be corrected, and generating corresponding target speech training data using the text content;

[0029] Replacing the speech training data of the labeled speech training data to be corrected with the target speech training data to obtain the labeled second speech training data.

[0030] In a second aspect, the present application further provides a speech recognition model acquisition device, including:

[0031] A speech training data acquisition module, configured to acquire unlabeled speech training data;

[0032] A data cleaning module, configured to clean the speech training data to obtain the speech training data after data cleaning;

[0033] A pseudo-label prediction module, configured to obtain a pre-trained speech recognition model, and perform pseudo-label prediction on the speech training data after data cleaning through the pre-trained speech recognition model to obtain labeled first speech training data;

[0034] A text error correction and speech correction module, configured to perform text error correction on the pseudo-labels in the labeled first speech training data, and correct the speech training data in the labeled first speech training data to obtain labeled second speech training data;

[0035] A target speech recognition model determination module, configured to adjust the pre-trained speech recognition model according to the labeled second speech training data to obtain a target speech recognition model.

[0036] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0037] Acquiring unlabeled speech training data;

[0038] Clean the speech training data to obtain the speech training data after data cleaning;

[0039] Obtain a pre-trained speech recognition model, and use the pre-trained speech recognition model to perform pseudo-label prediction on the speech training data after data cleaning to obtain the first speech training data with labels;

[0040] Perform text error correction on the pseudo-labels in the first speech training data with labels, and perform speech correction on the speech training data in the first speech training data with labels to obtain the second speech training data with labels;

[0041] Adjust the pre-trained speech recognition model according to the second speech training data with labels to obtain a target speech recognition model.

[0042] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0043] Obtain unlabeled speech training data;

[0044] Clean the speech training data to obtain the speech training data after data cleaning;

[0045] Obtain a pre-trained speech recognition model, and use the pre-trained speech recognition model to perform pseudo-label prediction on the speech training data after data cleaning to obtain the first speech training data with labels;

[0046] Perform text error correction on the pseudo-labels in the first speech training data with labels, and perform speech correction on the speech training data in the first speech training data with labels to obtain the second speech training data with labels;

[0047] Adjust the pre-trained speech recognition model according to the second speech training data with labels to obtain a target speech recognition model.

[0048] In a fifth aspect, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0049] Obtain unlabeled speech training data;

[0050] Clean the speech training data to obtain the speech training data after data cleaning;

[0051] Obtain a pre-trained speech recognition model, and use the pre-trained speech recognition model to perform pseudo-label prediction on the speech training data after data cleaning to obtain the first speech training data with labels;

[0052] Perform text error correction on the pseudo-labels in the labeled first voice training data, and perform voice correction on the voice training data in the labeled first voice training data to obtain the labeled second voice training data;

[0053] Adjust the pre-trained voice recognition model according to the labeled second voice training data to obtain the target voice recognition model.

[0054] The above voice recognition model acquisition method, device, computer device, computer-readable storage medium and computer program product obtain unlabeled voice training data; perform data cleaning on the voice training data to obtain the voice training data after data cleaning; obtain a pre-trained voice recognition model, and use the pre-trained voice recognition model to perform pseudo-label prediction on the voice training data after data cleaning to obtain the labeled first voice training data; perform text error correction on the pseudo-labels in the labeled first voice training data, and perform voice correction on the voice training data in the labeled first voice training data to obtain the labeled second voice training data; adjust the pre-trained voice recognition model according to the labeled second voice training data to obtain the target voice recognition model. In this application, by performing data cleaning on the unlabeled voice training data to remove invalid data, the unlabeled voice training data becomes purer. Further, by performing pseudo-label prediction on the cleaned unlabeled data, high-quality pseudo-label data is generated, avoiding the need for manual labeling, and performing text error correction and voice correction on the labeled voice training data, greatly retaining high-quality voice training data. By using this high-quality voice training data to adjust the pre-trained voice recognition model to obtain the target voice recognition model, the training accuracy of the voice recognition model in the unlabeled application scenario is improved, and the generalization ability of the voice recognition model is enhanced. Description of the Drawings

[0055] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0056] Figure 1 It is a schematic flowchart of a method for obtaining a voice recognition model in one embodiment;

[0057] Figure 2 It is a schematic flowchart of a method for obtaining a voice recognition model in another embodiment;

[0058] Figure 3 It is a structural block diagram of a voice recognition model acquisition device in an embodiment;

[0059] Figure 4 It is an internal structure diagram of a computer device in an embodiment. Specific implementation manners

[0060] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0061] In one embodiment, as Figure 1 shown, a method for acquiring a voice recognition model is provided. In this embodiment, it is exemplified that the method is applied to a server. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0062] Step S102, acquire unlabeled voice training data.

[0063] Among them, the unlabeled voice training data is the original voice data that has not been manually labeled. The data comes from a variety of scenarios and environments, including but not limited to telephone recordings, conference voices, broadcast content, etc., and is not directly associated with the corresponding text labels.

[0064] Exemplarily, during the process of the server acquiring the unlabeled voice training data, it will cover different voice scenarios, accents of different voice speakers, speech rates, etc., and collect voice data with different background noises to ensure the representativeness and diversity of the voice training data.

[0065] Step S104, perform data cleaning on the voice training data to obtain the voice training data after data cleaning.

[0066] Exemplarily, after the server receives the original voice training data, it first performs data cleaning. The main purpose of the cleaning is to remove invalid parts, such as non-voice parts like noise and silence, so as to improve the data quality. The data cleaning operations include but are not limited to denoising, voice part extraction, and silence elimination, etc., to ensure that the data used subsequently has high effectiveness and accuracy.

[0067] Step S106, acquire a pre-trained voice recognition model, and through the pre-trained voice recognition model, perform pseudo-label prediction on the voice training data after data cleaning to obtain the first voice training data with labels.

[0068] Among them, the pre-trained speech recognition model is a speech recognition model trained through a large-scale speech dataset, which can recognize the corresponding text information from the speech signal.

[0069] Pseudo-label prediction is a process of automatically labeling new unlabeled speech training data using an existing pre-trained speech recognition model without manually labeled data. Through this process, the model generates prediction labels for the input speech data, and these labels can be used as pseudo-labels for training to fine-tune the speech recognition model.

[0070] The labeled first speech training data is the speech data with attached prediction labels obtained through the pseudo-label prediction step.

[0071] Exemplarily, the server first obtains the pre-trained speech recognition model that has been trained through the network or the local database. The server loads and applies this pre-trained speech recognition model to the speech training data after data cleaning. Subsequently, the server performs pseudo-label prediction on these cleaned data through the pre-trained model. The model predicts the corresponding text labels as pseudo-labels based on the input speech signal, and matches these pseudo-labels with the speech data to form the labeled first speech training data.

[0072] Step S108: Correct the text of the pseudo-labels in the labeled first speech training data and correct the speech of the speech training data in the labeled first speech training data to obtain the labeled second speech training data.

[0073] Among them, text correction refers to correcting the text content in the pseudo-labels to make it conform to the grammar, spelling, and semantic norms of the language. Speech correction refers to adjusting or supplementing the speech training data itself to improve the speech quality or correct possible errors in the model recognition process, which can be carried out through means such as speech synthesis and audio enhancement. The labeled second speech training data refers to the new round of labeled speech training data obtained after text correction and speech correction processing.

[0074] Exemplarily, the server first receives the labeled first speech training data from the previous step. This first speech training data has undergone pseudo-label prediction through the pre-trained model. However, due to the error of the pre-trained speech recognition model, the pseudo-labels contain spelling, grammar, or semantic errors. The server performs text correction on the error-prone pseudo-labels through a preset algorithm to ensure the accuracy of the labels and the normativity of the language, thereby improving the quality of the training data.

[0075] Meanwhile, due to the influence of noise, uneven volume, etc. during the collection of speech training data, the server uses speech processing technologies (such as audio enhancement or speech synthesis, etc.) to correct the speech data to ensure its clarity and easy recognition. The corrected data then forms the labeled second speech training data for further training.

[0076] Step S110: Adjust the pre-trained speech recognition model according to the labeled second speech training data to obtain the target speech recognition model.

[0077] Exemplarily, the server obtains the labeled second speech training data, and the server uses the labeled second speech training data to adjust the parameters of the pre-trained speech recognition model so that the model can better process the specific speech features in the labeled data. On the basis of retaining the pre-trained speech recognition model, combined with the new second speech training data, the model can adapt to specific application scenarios and improve the recognition accuracy and robustness.

[0078] In this embodiment, by performing data cleaning on the unlabeled speech training data to remove invalid data, the unlabeled speech training data becomes purer. Further, by performing pseudo-label prediction on the cleaned unlabeled data, high-quality pseudo-label data is generated, avoiding the need for manual annotation, and performing text error correction and speech correction on the labeled speech training data, greatly retaining high-quality speech training data. By using the high-quality speech training data to adjust the pre-trained speech recognition model, the target speech recognition model is obtained, thereby improving the training accuracy of the speech recognition model in unlabeled application scenarios and enhancing the generalization ability of the speech recognition model.

[0079] In an exemplary embodiment, step S110 adjusts the pre-trained speech recognition model according to the labeled second speech training data to obtain the target speech recognition model, including: adjusting the pre-trained speech recognition model according to the labeled second speech training data to obtain the adjusted speech recognition model; performing pseudo-label prediction on the second speech training data through the adjusted speech recognition model to obtain the updated pseudo-labels, and determining the difference value between the updated pseudo-labels and the pre-updated pseudo-labels; in the case where the difference value is greater than the preset difference value threshold, return to execute the step of performing text error correction and speech correction on the first pseudo-labeled data to continue adjusting the pre-trained speech recognition model until the difference value is less than the preset difference value threshold to obtain the target speech recognition model.

[0080] Specifically, the server inputs the labeled second voice training data into the pre-trained voice recognition model for fine-tuning training. During the fine-tuning process, the server continuously adjusts the model's parameters by extracting data features and optimizing the loss function, making its performance on this specific dataset more excellent. The server uses the adjusted voice recognition model. The server inputs each audio sample of the second voice training data into the adjusted model, and the model generates new pseudo-labels according to the input voice data. The server compares the generated updated pseudo-labels with the previous pseudo-labels and calculates the difference value between the two. Optionally, the difference value can be obtained by calculating the edit distance of the pseudo-label text or other similarity measurement methods. If the difference value is greater than the preset threshold, the server will backtrack to the labeled first voice training data, use the language model for text correction, and use the voice correction algorithm to adjust the audio data to ensure the data quality. The corrected data will be input into the model for retraining. This process will continue until the difference value of the pseudo-labels is less than the preset threshold, indicating that the model has reached a stable state, and finally, the target voice recognition model is generated.

[0081] In this embodiment, by calculating the difference value of the pseudo-labels before and after the update and automatically adjusting and iteratively training the pre-trained recognition model according to the difference value, the error and manual intervention in the model training process are reduced, and the training efficiency is improved.

[0082] In an exemplary embodiment, the voice training data includes a voice part and a non-voice part; step S104 performs data cleaning on the voice training data to obtain the voice training data after data cleaning, including: performing noise reduction processing on the voice training data to obtain the voice training data after noise reduction; extracting the data segments of the voice part in the voice training data after noise reduction, and obtaining the voice training data after data cleaning according to the extracted data segments of the voice part.

[0083] In one embodiment, the unlabeled voice training data obtained by the server contains multiple audio segments, and these multiple audio segments contain a voice part and a non-voice part. The non-voice part refers to the part without sound or with relatively high background noise, while the voice part is the part containing actual voice information. The server identifies and separates these two parts for subsequent processing.

[0084] Specifically, the server processes the noise in the speech training data. The server can adopt noise reduction algorithms such as spectral subtraction and filter filtering. By analyzing the spectral characteristics of the audio signal, the components of the non-speech signal are removed. After the noise reduction process, the audio data obtained by the server will be cleaner, the noise components are significantly reduced, and the speech information is more prominent. The server further processes the noise-reduced speech training data to extract the effective speech part. Optionally, the server uses the Voice Activity Detection (VAD) algorithm to analyze the energy change of the audio signal, identify and remove the non-speech part in the speech data, and retain the effective speech part. The server extracts the data segment of the speech part to further obtain the cleaned speech training data set based on the data segment of the speech part.

[0085] In this embodiment, through the processing of noise reduction and extraction of the data segment of the speech part, the non-speech part is removed, ensuring that the training data set only contains effective speech signals, effectively reducing the errors caused by noise and silence during the training process, helping the model to better learn the speech features, thereby improving the accuracy of the speech recognition model and enhancing the adaptability to the actual application scenario.

[0086] In an exemplary embodiment, based on the extracted data segment of the speech part, the cleaned speech training data is obtained, including: obtaining a speech quality evaluation model, and determining the speech quality score of the data segment of the speech part through the speech quality evaluation model; removing the data in the data segment of the speech part whose speech quality score is less than the preset speech quality score threshold to obtain the cleaned speech training data.

[0087] The speech quality evaluation model refers to a model that evaluates the quality of speech training data through a preset algorithm. It can score the speech data according to factors such as audio clarity, noise level, and speech signal strength, and provide a decision basis for further processing. The preset speech quality score threshold is a scoring threshold value set by the user or the system according to the actual application requirements. In some embodiments, the speech training data with a speech quality score lower than the preset speech quality score threshold will be considered to have unqualified speech quality, and the server will remove it.

[0088] Specifically, the server analyzes various features (such as signal-to-noise ratio, volume, speech clarity, etc.) of the data segment of the speech part through the speech quality evaluation model, and evaluates the quality of the audio according to each feature to obtain the speech quality score corresponding to the data segment of the speech part. Optionally, the speech quality evaluation model can be pre-trained and can be customized and adjusted in the application scenario.

[0089] The server will screen the scoring results according to the preset voice quality scoring threshold. If the score of a data segment of a certain voice part is lower than the preset threshold, it indicates that there is strong noise or other interference in this segment. The server will eliminate the data segments of this low-quality voice part and retain those data segments of the voice parts with scores higher than the threshold.

[0090] In this embodiment, through the voice quality evaluation model, low-quality voice segments can be automatically screened out and eliminated. While avoiding the cumbersome work of manually checking and screening data one by one, it also avoids the negative impact of low-quality data on the model training process, ensuring that the trained model can achieve a high recognition accuracy in actual applications.

[0091] In an exemplary embodiment, correcting the text of the pseudo-labels in the labeled first voice training data in step S108 includes: obtaining a language model, and in the labeled first voice training data, determining the pseudo-labels to be corrected whose text expressions do not conform to the preset expression mode through the language model; correcting the text corresponding to the pseudo-labels to be corrected.

[0092] Among them, the language model refers to a model trained through a large amount of language data, which can predict and generate natural language text according to the context, can be used to identify the relationships between words, judge the reasonableness of words in a sentence, and predict appropriate words or phrases to correct incorrect text expressions.

[0093] Specifically, the server obtains the trained language model and applies the language model to the labeled first voice training data. The server uses the language model to check whether the text expression of each pseudo-label conforms to the preset expression mode. If the text expression of a certain pseudo-label does not conform to the preset expression mode (for example, there are grammar errors, spelling errors or context mismatches), the language model will mark these pseudo-labels as objects to be corrected, that is, the pseudo-labels to be corrected. The language model will analyze the text of the pseudo-labels to be corrected and modify it according to the rules, vocabulary and grammar structure in the preset expression mode, so that the pseudo-labels after error correction will become more standardized and conform to the predetermined expression standard, thereby improving the quality and accuracy of the data.

[0094] Optionally, before the server corrects the errors through the language model, it screens the labeled first voice training data and corrects the screened first voice training data. Specifically, by imposing conditional restrictions on the pronunciation and amplitude of the labeled first voice data, the voice data that does not meet the preset pronunciation conditions and amplitude conditions in the labeled first voice data is screened out, and then the screened voice training data is corrected through the language model. Only the data that does not meet the preset conditions is corrected, reducing the training time while ensuring the high quality of the data.

[0095] In this embodiment, by using a language model to correct the text of the pseudo-labels in the labeled first speech training data, the quality and accuracy of the speech training data are ensured, the training process of the speech recognition model is optimized, and the speech recognition efficiency and accuracy are improved.

[0096] In an exemplary embodiment, in step S108, the speech training data in the labeled first speech training data is corrected to obtain labeled second speech training data, including: calculating the perplexity of the pseudo-labels in the labeled first speech training data to obtain a perplexity calculation result; determining the first speech training data corresponding to the pseudo-labels with the perplexity calculation result greater than a preset perplexity threshold as the labeled speech training data to be corrected; determining the text content of the pseudo-labels in the labeled speech training data to be corrected, and generating corresponding target speech training data using the text content; and replacing the speech training data in the labeled speech training data to be corrected with the target speech training data to obtain the labeled second speech training data.

[0097] Among them, perplexity is used to represent the difficulty of predicting a text sequence.

[0098] Specifically, the server processes each pseudo-label text in the first speech training data through a language model and calculates the perplexity calculation result of each pseudo-label. The perplexity calculation result represents the prediction accuracy of the pseudo-label in the language model. The higher the value, the greater the uncertainty of the speech data corresponding to the label. The server filters out the pseudo-labels with the perplexity calculation result greater than the threshold according to the preset perplexity threshold. At the same time, the speech training data corresponding to the pseudo-label will be recognized as speech training data with a high error rate, and the speech training data will be marked as speech training data to be corrected.

[0099] The server further analyzes the pseudo-labels in the speech training data to be corrected and extracts their text content. The server uses this text content to generate corresponding target speech training data through a speech synthesis model. In one embodiment, the server optimizes the grammar of the text content through a speech synthesis model and sets a higher speech quality standard for the generated speech to ensure the accuracy and clarity of the generated target speech training data. The server replaces the original speech training data to be corrected with the target speech training data generated through the text content to form new labeled second speech training data.

[0100] In this embodiment, by correcting the speech training data with a high perplexity of pseudo-labels, target speech training data with high accuracy and high clarity is generated, enabling the pre-trained speech recognition model to be further trained with high-quality data, improving the training effect of the model, and thus enhancing the generalization ability of the speech recognition model.

[0101] To enable those skilled in the art to better understand the above steps, the embodiments of the present application are exemplarily described below by way of an example. Before the exemplary description of the embodiments of the present application, the related technologies of the present application are introduced first. However, it should be understood that the embodiments of the present application are not limited thereto.

[0102] In an embodiment of the related technology, such as in a method for training a speech recognition model based on a small amount of labeled data, an initial model is trained by using a small amount of manually labeled data, and the initial model can provide a basic recognition effect. Then, the initial model is used to predict unlabeled data to generate pseudo-labels, and the labeled data is obtained. The pseudo-labels are the texts predicted by the model. The obtained labeled data is merged with the small amount of manually labeled data to form an extended data set, and a new model is retrained based on the extended data set. The new model will be used as the initial model for the next round of training to continue predicting labels for unlabeled data. This process is repeated until the difference between the currently generated pseudo-labels and the pseudo-labels generated in the previous round is less than a preset threshold.

[0103] In the process of generating pseudo-labels, the pseudo-labels can be further screened. Specifically, based on the confidence score of the model, the data with lower scores is screened out and excluded from participating in subsequent training. In addition, a confidence prediction network can be used to evaluate the confidence of each data point, and only the data with higher confidence is selected for subsequent training, so as to improve the data quality and optimize the model training process.

[0104] In an exemplary embodiment of the present application, as Figure 2 shown, in step S201, the server obtains unlabeled speech data. The unlabeled speech training data can come from real speech interactions, such as customer service call recordings, conference recordings, or other speech application scenarios.

[0105] The server performs data cleaning on the obtained unlabeled speech data. The data cleaning includes step S202 of performing noise reduction and extracting data segments of the speech part from the obtained unlabeled speech data, and step S203 of performing speech quality evaluation and screening on the extracted data segments of the speech part. Specifically, the server performs noise reduction and extracts speech part data from the collected unlabeled speech data. At this stage, the server preprocesses the original speech data through a noise reduction algorithm to remove the background noise therein and improve the quality of the speech signal. In the data after the noise reduction process, the server segments the audio data through a VAD model. The VAD model analyzes the audio signal, identifies and separates the parts containing speech, and removes the non-speech parts. Subsequently, the server performs speech quality evaluation and screening on the extracted data segments of the speech part. Next, the server uses a speech quality evaluation model to score the audio segments after VAD processing. This model comprehensively considers factors such as the clarity of the speech signal, the noise level, and the volume size, and scores each data segment of the speech part. If the quality evaluation result of the speech training data is lower than the preset quality score threshold, this data will be excluded to avoid affecting the subsequent model training process.

[0106] In step S204, the server performs pseudo-label prediction on the data after data cleaning to obtain the first labeled speech training data. The server performs speech recognition on the speech data after data cleaning by obtaining and using a pre-trained speech recognition model, generates pseudo-labels, and labels them on the speech training data to obtain the first labeled speech training data, providing preliminary labeling information for subsequent text correction and speech correction.

[0107] The server performs text correction and speech correction on the first labeled speech training data to obtain the second labeled speech training data. Specifically, in step S205, the server corrects the pseudo-labels in the first labeled speech training data through a language model. The language model identifies the parts of the pseudo-label text that do not conform to the preset standard, such as grammar errors, spelling mistakes, or improper use of vocabulary, etc., and automatically performs text correction to correct the parts that do not conform to the preset expression. Specifically, in step S206, the server calculates the perplexity of the pseudo-labels for the first labeled speech data through a language model and performs speech correction on the first speech training data according to the perplexity calculation result and a speech generation model. After correcting the pseudo-label text, the server continues to use the language model. For the speech training data, the server calculates the perplexity of the pseudo-labels to determine whether there are deviations in the speech content of the pseudo-label text. For the pseudo-labels with a perplexity higher than the preset perplexity threshold, the server further corrects the speech training data corresponding to the pseudo-labels. The target speech training data corresponding to the text can be generated through a speech generation model and replace the original speech training data to obtain the second labeled speech training data.

[0108] Step S207: The server adjusts the pre-trained speech recognition model according to the tagged second speech training data. After performing text correction and speech modification, the server optimizes the original model parameters by fine-tuning training, using the second speech training data as new training data to make it more adaptable to the speech features of a specific scenario.

[0109] Step S208: The server uses the adjusted pre-trained speech recognition model to perform pseudo-label prediction on the second speech training data and evaluate the difference in pseudo-labels before and after the update. The server uses the adjusted cloud training speech recognition model to perform pseudo-label prediction on the second speech training data to generate updated pseudo-labels. At this time, the server determines the effect of model optimization by calculating the difference in pseudo-labels before and after the update. The difference value can be calculated through text similarity or other metrics. If the difference value after pseudo-label update is still greater than the preset difference value threshold, it indicates that the model has not reached the expected accuracy and still needs to be further optimized. Return to step S205.

[0110] Step S209: Determine the target speech recognition model. Specifically, the pre-trained speech recognition model corresponding to the situation where the difference value is less than the preset difference value threshold in step S208 is determined as the target speech recognition model.

[0111] In this embodiment, through the pre-trained speech recognition model, speech noise reduction algorithm, VAD model, speech quality evaluation model, language model, and speech synthesis model, unlabeled data in the application scenario is segmented, corrected, error-corrected, predicted, etc., to obtain high-quality pseudo-labeled data. Then, the model is iteratively trained using this pseudo-labeled data, thus achieving good recognition effects in multiple application scenarios. Compared with the dependence on manual annotation in the related art, this application uses unlabeled data and performs a series of processes on the unlabeled data, reducing the dependence on manual annotation while achieving the adjustment accuracy of the pre-trained speech recognition model, and further improving the generalization ability of the speech recognition model.

[0112] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps in other steps.

[0113] Based on the same inventive concept, an embodiment of the present application further provides a voice recognition acquisition device for implementing the voice recognition acquisition method involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the voice recognition acquisition device provided below can refer to the limitations on the voice recognition acquisition method in the above text, and will not be repeated here.

[0114] In an exemplary embodiment, as Figure 3 shown, a voice recognition acquisition device is provided, including: a voice training data acquisition module 310, a data cleaning module 320, a pseudo-label prediction module 330, a text error correction and voice correction module 340, and a target voice recognition model determination module 350, where:

[0115] The voice training data acquisition module 310 is configured to acquire unlabeled voice training data;

[0116] The data cleaning module 320 is configured to perform data cleaning on the voice training data to obtain the voice training data after data cleaning;

[0117] The pseudo-label prediction module 330 is configured to acquire a pre-trained voice recognition model, and perform pseudo-label prediction on the voice training data after data cleaning through the pre-trained voice recognition model to obtain the first voice training data with labels;

[0118] The text error correction and voice correction module 340 is configured to perform text error correction on the pseudo-labels in the first voice training data with labels, and perform voice correction on the voice training data in the first voice training data with labels to obtain the second voice training data with labels;

[0119] The target voice recognition model determination module 350 is configured to adjust the pre-trained voice recognition model according to the second voice training data with labels to obtain a target voice recognition model.

[0120] In one embodiment, the target speech recognition model determination module 350 is also used to adjust the pre-trained speech recognition model according to the labeled second speech training data to obtain an adjusted speech recognition model; perform pseudo-label prediction on the second speech training data through the adjusted speech recognition model to obtain an updated pseudo-label, and determine the difference between the updated pseudo-label and the pseudo-label before the update; when the difference value is greater than a preset difference value threshold, return to execute the step of performing text correction and speech correction on the labeled first speech training data to continue adjusting the pre-trained speech recognition model until the difference value is less than the preset difference value threshold to obtain the target speech recognition model.

[0121] In one embodiment, the speech training data includes a speech part and a non-speech part, and the data cleaning module 320 is further used to perform noise reduction processing on the speech training data to obtain noise-reduced speech training data; extract data segments of the speech part in the noise-reduced speech training data, and obtain the speech training data after data cleaning based on the extracted data segments of the speech part.

[0122] In one embodiment, the data cleaning module 320 is also used to obtain a speech quality assessment model, and determine the speech quality score of the data segment of the speech part through the speech quality assessment model; remove the data in the data segment of the speech part whose speech quality score is less than a preset speech quality score threshold, and obtain speech training data after data cleaning.

[0123] In one embodiment, the text correction and speech modification module 340 is also used to obtain a language model, and determine, in the labeled first speech training data, a pseudo-label to be corrected whose text expression does not conform to a preset expression through the language model; and correct the text corresponding to the pseudo-label to be corrected.

[0124] In one embodiment, the text error correction and speech correction module 340 is further used to perform perplexity calculation on the pseudo-labels in the labeled first speech training data to obtain perplexity calculation results; determine the first speech training data corresponding to the pseudo-labels whose perplexity calculation results are greater than a preset perplexity threshold as labeled speech training data to be corrected; determine the text content of the pseudo-labels in the labeled speech training data to be corrected, and generate corresponding target speech training data using the text content; replace the speech training data in the labeled speech training data to be corrected with the target speech training data to obtain the labeled second speech training data. .

[0125] Each module in the above voice recognition acquisition device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0126] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store unlabeled voice training data, labeled first voice training data, and labeled second voice training data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for obtaining a voice recognition model.

[0127] Those skilled in the art can understand that Figure 4 the structure shown in

[0128] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0129] In an embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0130] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0132] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0133] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0134] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A method for acquiring a speech recognition model, characterized in that: The method comprises: Obtain unlabeled speech training data; Performing data cleaning on the speech training data to obtain speech training data after data cleaning; Obtain a pre-trained speech recognition model, and use the pre-trained speech recognition model to perform pseudo-label prediction on the speech training data after the data cleaning to obtain first speech training data with labels; Performing text error correction on the pseudo-labels in the first labeled speech training data, and performing speech correction on the speech training data in the first labeled speech training data, to obtain second labeled speech training data; The pre-trained speech recognition model is adjusted according to the labeled second speech training data to obtain a target speech recognition model.

2. The method according to claim 1, characterized in that: The step of adjusting the pre-trained speech recognition model according to the labeled second speech training data to obtain a target speech recognition model includes: According to the labeled second speech training data, the pre-trained speech recognition model is adjusted to obtain an adjusted speech recognition model; Performing pseudo-label prediction on the second speech training data by using the adjusted speech recognition model to obtain updated pseudo-labels, and determining a difference between the updated pseudo-labels and the pseudo-labels before the update; When the difference value is greater than the preset difference value threshold, return to the step of performing text error correction and speech correction on the labeled first speech training data to continue adjusting the pre-trained speech recognition model until the difference value is less than the preset difference value threshold, thereby obtaining a target speech recognition model.

3. The method according to claim 1, characterized in that The speech training data includes a speech part and a non-speech part; The step of performing data cleaning on the speech training data to obtain the cleaned speech training data includes: Performing noise reduction processing on the speech training data to obtain noise-reduced speech training data; Extracting data segments of the speech part from the noise-reduced speech training data, and obtaining the speech training data after data cleaning according to the extracted data segments of the speech part.

4. The method according to claim 3, characterized in that The method of obtaining speech training data after data cleaning based on the extracted speech data segments comprises: Acquire a speech quality assessment model, and determine a speech quality score of the data segment of the speech portion by using the speech quality assessment model; The data whose speech quality score is less than a preset speech quality score threshold in the data segment of the speech part is eliminated to obtain speech training data after data cleaning.

5. The method according to claim 1, characterized in that The performing text error correction on the pseudo-labels in the labeled first speech training data comprises: Acquire a language model, and determine, in the labeled first speech training data, pseudo labels to be corrected whose text expressions do not conform to a preset expression through the language model; Correct the text corresponding to the pseudo-label to be corrected.

6. The method according to claim 1, characterized in that The performing speech correction on the speech training data in the first speech training data with labels to obtain second speech training data with labels includes: Performing perplexity calculation on the pseudo-labels in the labeled first speech training data to obtain a perplexity calculation result; Determine the first speech training data corresponding to the pseudo-label whose perplexity calculation result is greater than a preset perplexity threshold as the labeled speech training data to be corrected; Determining the text content of the pseudo-label in the labeled speech training data to be corrected, and generating corresponding target speech training data using the text content; The target speech training data replaces the speech training data in the labeled speech training data to be corrected to obtain the labeled second speech training data.

7. A speech recognition model acquisition device, characterized in that: The device comprises: A speech training data acquisition module is used to acquire unlabeled speech training data; A data cleaning module, used for cleaning the speech training data to obtain the cleaned speech training data; A pseudo-label prediction module is used to obtain a pre-trained speech recognition model, and perform pseudo-label prediction on the speech training data after the data cleaning through the pre-trained speech recognition model to obtain first speech training data with labels; A text error correction and speech correction module, used for performing text error correction on the pseudo-labels in the first labeled speech training data, and performing speech correction on the speech training data in the first labeled speech training data, to obtain second labeled speech training data; The target speech recognition model determination module is used to adjust the pre-trained speech recognition model according to the labeled second speech training data to obtain a target speech recognition model.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Self-evolution method and system of speech recognition model

    CN120808760A