Speech corpus generation method, device, computer device and storage medium

By obtaining unlabeled audio collections and using multiple speech recognition models for transcription and screening, the problem of low quality of speech corpus is solved, and efficient and accurate speech corpus generation and model optimization are achieved.

CN114141235BActive Publication Date: 2025-07-08ZHAOLIAN CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111249234.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-07-08
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

Existing speech recognition technology requires a large number of manually labeled speech corpus when training neural network models, resulting in lower quality and time-consuming and error-consuming in labeling.

Method used

By obtaining unlabeled audio sets, transcribe using the speech recognition model to be optimized, optimal and referenced speech recognition models, calculate the difference information and loss values, filter out high-quality audio and store it in the speech corpus, and optimize the speech recognition model to improve the labeling accuracy.

Benefits of technology

Improve the generation efficiency and quality of the voice corpus, ensure the accuracy of the label text, and reduce the need for manual labeling through automatic labeling and model optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114141235B_ABST
    Figure CN114141235B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, and storage medium for generating a speech corpus. The method includes: for each audio in the audio set, performing speech transcription on the audio through an speech recognition model to be optimized, an optimal speech recognition model, and a reference speech recognition model respectively, to obtain corresponding first transcription text, optimal transcription text, and second transcription text; the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized; respectively determining the difference information between the first transcription text and the second transcription text compared with the optimal transcription text; inputting the audio and the optimal transcription text into the speech recognition model to be optimized to obtain a loss value aligned at the phoneme level; screening out the audios that meet the high-quality conditions from the audio set; storing the screened audios and the corresponding optimal transcription texts into the speech corpus; and optimizing the speech recognition model to be optimized based on the data in the speech corpus. Therefore, the efficiency and quality of generating the speech corpus are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, apparatus, computer device, and storage medium for generating a speech corpus. Background Art

[0002] With the development of artificial intelligence, speech recognition technology has emerged, which is used to recognize user speech and obtain the corresponding text content of the speech, so as to further interact with the user. At present, a large number of labeled speech corpora are required during the development of speech recognition technology to train neural network models. However, preparing speech and manually annotating the speech takes a lot of time, and the annotation is prone to errors. Therefore, the quality of the speech corpus is relatively low. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium for generating a speech corpus that can improve the quality of the speech corpus.

[0004] A method for generating a speech corpus, the method comprising:

[0005] Obtain an audio set; the audio set includes audio without corresponding transcribed text being labeled;

[0006] For each audio in the audio set, perform speech transcription on the audio through an speech recognition model to be optimized, an optimal speech recognition model, and a reference speech recognition model respectively, to obtain corresponding first transcribed text, optimal transcribed text, and second transcribed text; the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized;

[0007] Determine the difference information between the first transcribed text and the second transcribed text compared to the optimal transcribed text respectively;

[0008] Input the audio and the optimal transcribed text into the speech recognition model to be optimized to obtain a loss value at the phoneme level alignment;

[0009] Based on the difference information and the loss value corresponding to each audio, screen out the audio that meets the high-quality conditions from the audio set;

[0010] Store the screened audio and the corresponding optimal transcribed text as a set of data into the speech corpus;

[0011] Optimize the speech recognition model to be optimized based on each set of data in the speech corpus.

[0012] In one embodiment, the obtaining the audio set includes: obtaining an initial audio set;

[0013] Input the audios in the initial audio set into the optimal speech recognition model respectively to obtain the optimal transcription text corresponding to the audio and the relevant information of the audio;

[0014] Based on the optimal transcription text and the audio relevant information, perform preliminary screening processing on the audios in the initial audio set to obtain an audio set.

[0015] In one embodiment, the method further includes:

[0016] Input the first transcription text into a statistical language model to obtain a perplexity metric;

[0017] The screening out of the audios that meet the high-quality conditions from the audio set based on the difference information and the loss value includes:

[0018] Based on the difference information, the loss value, and the perplexity metric, screen out the audios that meet the high-quality conditions from the audio set.

[0019] In one embodiment, the difference information includes the word error rate of the first transcription text and the second transcription text compared with the optimal transcription text; the respectively determining the difference information of the first transcription text and the second transcription text compared with the optimal transcription text includes:

[0020] For any one of the first transcription text and the second transcription text, taking the optimal transcription text as a standard, respectively determine the number of words added, reduced, and replaced in the any one of the transcription texts compared with the optimal transcription text;

[0021] Based on the number of words, obtain the word error rate of the any one of the transcription texts compared with the optimal transcription text.

[0022] In one embodiment, the screening out of the audios that meet the high-quality conditions from the audio set based on the difference information and the loss value corresponding to each audio includes:

[0023] For each audio, determine the degree of proximity between the word error rates of the corresponding first transcription text and the second transcription text compared with the optimal transcription text;

[0024] If the degree of proximity meets a preset proximity condition and the loss value is less than a preset loss value threshold, then determine that the audio meets the high-quality conditions.

[0025] In one embodiment, the method further includes:

[0026] For the unfiltered audio, obtain at least two manually annotated texts annotated for the audio, where the manually annotated texts are obtained by different annotators annotating the audio separately;

[0027] Detect whether the manually annotated texts are consistent;

[0028] If they are inconsistent, trigger a recheck annotation for the audio to obtain a recheck annotation text;

[0029] Store the audio and the recheck annotation text in the speech corpus.

[0030] In one embodiment, the difference information and the loss value are included in the screening information set corresponding to the audio;

[0031] The method further includes:

[0032] For the unfiltered audio, obtain audio-related information of the audio;

[0033] Input the audio-related information and the screening information set into the trained audio quality model to obtain the annotation difficulty corresponding to the audio;

[0034] Based on the annotation difficulty, group the unfiltered audio so that the corresponding annotators can annotate each group of audio to obtain manually annotated texts.

[0035] A speech corpus generation device, the device includes:

[0036] An acquisition module, configured to acquire audio; acquire at least two transcription texts obtained for the audio respectively based on at least two speech recognition models, where the speech recognition models include an optimal speech recognition model and a speech recognition model to be optimized; the transcription texts include the optimal transcription text output by the optimal speech recognition model;

[0037] A screening information calculation module, configured to obtain a word error rate set based on the transcription texts; input the audio and the optimal transcription text into the speech recognition model to be optimized to obtain a loss value index;

[0038] A screening module, configured to perform screening processing on the audio based on a screening information set; if meeting a preset requirement, put the audio and the optimal transcription text into the speech corpus; the screening information set includes the word error rate set and the loss value index;

[0039] A continuous optimization module, configured to continuously optimize the speech recognition model to be optimized based on the speech corpus.

[0040] A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the steps of the above-mentioned speech corpus generation method.

[0041] A computer-readable storage medium stores a computer program thereon, and the computer program is executed by a processor to perform the steps of the above-mentioned speech corpus generation method.

[0042] In the above-mentioned speech corpus generation method, device, computer device and storage medium, by obtaining an audio set including audio without corresponding transcribed text. For each audio in the audio set, perform speech transcription on the audio through an speech recognition model to be optimized, an optimal speech recognition model and a reference speech recognition model respectively, to obtain corresponding first transcribed text, optimal transcribed text and second transcribed text, wherein the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized. Determine the difference information between the first transcribed text and the second transcribed text compared with the optimal transcribed text respectively; then input the audio and the optimal transcribed text into the speech recognition model to be optimized to obtain a loss value aligned at the phoneme level; based on the difference information and loss value corresponding to each audio, screen out the audio that meets the high-quality conditions from the audio set. Store the screened audio and the corresponding optimal transcribed text as a set of data in the speech corpus, thus ensuring the accuracy of the optimal transcribed text, that is, the annotated text. And optimize the speech recognition model to be optimized based on each set of data in the speech corpus to continuously optimize the speech recognition model to be optimized, so as to increase the accuracy of the screened corpus data. Therefore, through the automatic annotation of speech, the screening of annotated text, and the continuous optimization of the speech recognition model, the efficiency and quality of speech corpus generation are improved. Description of the Drawings

[0043] Figure 1 It is an application environment diagram of the speech corpus generation method in an embodiment;

[0044] Figure 2 It is a schematic flowchart of the speech corpus generation method in an embodiment;

[0045] Figure 3 It is an overall framework diagram of the speech corpus generation method in an embodiment;

[0046] Figure 4 It is a structural block diagram of the speech corpus generation device in an embodiment;

[0047] Figure 5 It is a structural block diagram of the audio grouping module in an embodiment;

[0048] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0049] In order to make the objectives, technical solutions, and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] The method for generating a speech corpus provided by the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 110 communicates with the server 120 through a network. Among them, the terminal 110 can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices, and the server 120 can be implemented by an independent server or a server cluster composed of multiple servers.

[0051] The terminal 110 can collect audio and initially screen the audio to obtain an audio set including audio without corresponding transcribed text. The terminal 110 sends the audio set to the server 120. The server 120 obtains the audio set; for each audio in the audio set, speech transcription is performed on the audio through a speech recognition model to be optimized, an optimal speech recognition model, and a reference speech recognition model respectively to obtain corresponding first transcribed text, optimal transcribed text, and second transcribed text; the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized. The server 120 respectively determines the difference information between the first transcribed text and the second transcribed text compared with the optimal transcribed text; inputs the audio and the optimal transcribed text into the speech recognition model to be optimized to obtain a loss value aligned at the phoneme level. The server 120 screens out the audio that meets the high-quality conditions from the audio set based on the difference information and the loss value corresponding to each audio. The server 120 stores the screened audio and the corresponding optimal transcribed text as a set of data in the speech corpus; optimizes the speech recognition model to be optimized based on the data in each group in the speech corpus.

[0052] In one embodiment, the terminal 110 can also be replaced by a server, and there is no limitation in this regard.

[0053] In one embodiment, as Figure 2 shown, a method for generating a speech corpus is provided. In this embodiment, an example is given where this method is applied to a server. It can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0054] S202, obtain an audio set; the audio set includes audio without corresponding transcribed texts; for each audio in the audio set, perform speech transcription on the audio through an speech recognition model to be optimized, an optimal speech recognition model, and a reference speech recognition model respectively, to obtain corresponding first transcribed texts, optimal transcribed texts, and second transcribed texts; the recognition accuracy rate of the reference speech recognition model is lower than that of the speech recognition model to be optimized.

[0055] Among them, the speech recognition model to be optimized is the speech recognition model that needs to be optimized in this method. The optimal speech recognition model is the speech recognition model with the highest speech recognition accuracy rate among the three speech recognition models used in this method (that is, the speech recognition model to be optimized, the optimal speech recognition model, and the reference speech recognition model). The recognition accuracy rate of the reference speech recognition model is lower than that of the speech recognition model to be optimized, that is, the recognition accuracy rate of the reference speech recognition model is the lowest among the three speech recognition models.

[0056] Among them, the transcribed text refers to the text generated after performing speech recognition on the audio, that is, the text obtained by transcribing the words expressed in the audio into text format. The transcribed text can also be called the annotated text of the audio.

[0057] Specifically, the server obtains an audio set; the audio set includes multiple audio without corresponding transcribed texts. For each audio in the audio set, the server inputs the audio into the speech recognition model to be optimized, the optimal speech recognition model, and the reference speech recognition model respectively for speech transcription, to obtain corresponding first transcribed texts, optimal transcribed texts, and second transcribed texts. It can be understood that the first transcribed text is generated by the speech recognition model to be optimized for speech transcription of the audio, the optimal transcribed text is generated by the optimal speech recognition model for speech transcription of the audio, and the second transcribed text is generated by the reference speech recognition model for speech transcription of the audio. The recognition accuracy rate of the optimal transcribed text is the highest among the three transcribed texts, and the recognition accuracy rate of the second transcribed text is the lowest among the three transcribed texts.

[0058] In one embodiment, the audio set can be initially screened based on the optimal speech recognition model, and the optimal transcribed text is obtained during the initial screening process. In other embodiments, the audio set can also be an original audio set without initial screening, and this is not limited.

[0059] In one embodiment, when the server preliminarily screens the audio set based on the optimal speech recognition model, it involves generating the optimal transcription text by transcribing each audio in the audio set. Therefore, after screening out the audio set, the server can directly obtain the optimal transcription text generated for each audio during the screening process, without using the optimal speech recognition model for secondary transcription. Instead, it only needs to transcribe the audio in the screened audio set through the speech recognition model to be optimized and the reference speech recognition model respectively to obtain the corresponding first transcription text and second transcription text, thereby saving computer processing resources.

[0060] S204. Respectively determine the difference information between the first transcription text and the second transcription text compared with the optimal transcription text; input the audio and the optimal transcription text into the speech recognition model to be optimized to obtain the loss value of phoneme-level alignment.

[0061] Among them, the difference information refers to the difference information related to the text content between multiple transcription texts.

[0062] In one embodiment, the difference information can be calculated based on the difference of characters between texts.

[0063] In another embodiment, the difference information can be calculated respectively based on the differences of characters and punctuation marks between texts.

[0064] In another embodiment, the difference information can be calculated respectively based on the differences of characters, the corresponding pinyin of characters, and punctuation marks between texts.

[0065] Among them, a phoneme refers to the smallest speech unit divided according to the natural attributes of speech. The loss value of phoneme-level alignment refers to the loss value between the audio and the transcription text at the phoneme level.

[0066] Specifically, the server calculates the difference information between the first transcription text and the optimal transcription text, and calculates the difference information between the second transcription text and the optimal transcription text. The server inputs the audio and the optimal transcription text into the speech recognition model to be optimized to obtain the loss value between the audio and the transcription text at the phoneme level. It can be understood that the difference information and loss value obtained by the server are used to screen the audio set.

[0067] S206. Based on the difference information and loss value corresponding to each audio, screen out the audio that meets the high-quality conditions from the audio set; store the screened audio and the corresponding optimal transcription text as a group of data in the speech corpus.

[0068] Among them, the speech corpus is used to store speech corpora, that is, to store a set of multiple audios and corresponding transcription texts.

[0069] Specifically, for each audio in the audio set, the server performs a screening process on the audio based on the difference information and loss value corresponding to each audio to determine whether the audio meets the preset high-quality condition; if so, the audio and the corresponding optimal transcription text are stored in the speech corpus as a set of data. It can be understood that the audio that meets the high-quality condition is selected based on the difference information and loss value to ensure the accuracy of the transcription text (i.e., the labeled text) in the speech corpus.

[0070] In one embodiment, for at least some of the audios that do not meet the high-quality condition and are not screened out, after manual annotation processing, the manually annotated text and the audio are put into the speech corpus as a set of data.

[0071] In one embodiment, for the audios that do not meet the high-quality condition and are not screened out, the audios that meet the medium-quality condition can be selected from them based on the difference information and loss value, and after manual annotation processing, the manually annotated text and the audio are put into the speech corpus as a set of data.

[0072] In one implementation, the server can discard the audios that do not meet the high-quality condition and do not meet the medium-quality condition.

[0073] S208, optimize the speech recognition model to be optimized based on each set of data in the speech corpus.

[0074] Specifically, multiple audios and corresponding transcription texts are stored in the speech corpus, and a set of data includes an audio and the corresponding transcription text. The server obtains each set of data in the speech corpus and trains the speech recognition model to be optimized based on each set of data for further optimizing the speech recognition model.

[0075] It can be understood that after optimizing the speech recognition model to be optimized, when using the optimized speech recognition model to perform automatic transcription and annotation processing on the next batch of audio sets, it can be more accurate, that is, it can make the first transcription text obtained in step S202 have a higher recognition accuracy, and the loss value at the phoneme level obtained in step 204 has a better result, thereby improving the accuracy of the entire text automatic annotation processing (i.e., the process of automatically adding transcription text to the audio), and further reducing the manual annotation processing and improving the annotation efficiency.

[0076] The above-mentioned method, apparatus, computer device, and storage medium for generating a speech corpus obtain an audio set including audio without corresponding transcribed text. For each audio in the audio set, speech transcription is performed on the audio through an speech recognition model to be optimized, an optimal speech recognition model, and a reference speech recognition model, respectively, to obtain corresponding first transcribed text, optimal transcribed text, and second transcribed text, where the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized. The difference information between the first transcribed text and the second transcribed text compared with the optimal transcribed text is determined respectively; then the audio and the optimal transcribed text are input into the speech recognition model to be optimized to obtain a loss value aligned at the phoneme level; based on the difference information and the loss value corresponding to each audio, the audio that meets the high-quality condition is screened out from the audio set. The screened audio and the corresponding optimal transcribed text are stored in the speech corpus as a set of data, thus ensuring the accuracy of the optimal transcribed text, that is, the annotated text. And the speech recognition model to be optimized is optimized based on each set of data in the speech corpus to continuously optimize the speech recognition model to be optimized, so as to increase the accuracy of the screened corpus data. Therefore, by automatically annotating speech, screening annotated text, and continuously optimizing the speech recognition model, the generation efficiency and quality of the speech corpus are improved.

[0077] In one embodiment, obtaining the audio set includes: obtaining an initial audio set; inputting the audios in the initial audio set into the optimal speech recognition model respectively to obtain the optimal transcribed text corresponding to the audio and the relevant information of the audio; and performing preliminary screening processing on the audios in the initial audio set based on the optimal transcribed text and the audio relevant information to obtain the audio set.

[0078] In one embodiment, before obtaining the initial audio set, the method further includes: obtaining a long audio containing voice tracks of two dialogue roles, performing voice activity detection (VAD) processing on the audio to obtain a plurality of short audios each containing a voice track of one dialogue role, and putting the short audios into the initial audio set.

[0079] In one embodiment, the two dialogue roles corresponding to the long audio include a customer and an agent, and the long audio is obtained by recording the dialogue between the customer and the agent.

[0080] Wherein, the relevant information of the audio includes the acoustic feature information of the audio.

[0081] In one embodiment, the acoustic feature information may include at least one of duration, speech rate, pitch, etc.

[0082] Specifically, the server obtains an initial audio set containing multiple audios, inputs the audios in the initial audio set into the optimal speech recognition model respectively, and obtains the optimal transcription text corresponding to the audio and the relevant information of the audio, including the duration, speech rate, etc. The server performs preliminary screening on the audios in the initial audio set based on the optimal transcription text and the audio relevant information, discards the audios that do not meet the requirements, and obtains an audio set.

[0083] In one embodiment, the server can perform preliminary screening on the audios in the initial audio set based on the optimal transcription text, the speaker information of the audio, and the audio relevant information, and obtain an audio set. Among them, the speaker information includes at least one of gender information, dialogue role information, etc.

[0084] In one embodiment, the preliminary screening process may include at least one of screening in terms of phoneme distribution diversity dimension, speaker diversity dimension, acoustic feature dimension, etc.

[0085] In one embodiment, the screening in terms of phoneme distribution diversity dimension includes: obtaining the total phoneme distribution information of the speech corpus; obtaining the target phoneme distribution information of the optimal transcription text, calculating the contribution value of the audio in maintaining the phoneme diversity of the speech corpus according to the target phoneme distribution information and the total phoneme distribution information, and if the contribution value is less than the preset threshold, discarding the audio. It can be understood that screening the audio in terms of phoneme distribution diversity dimension enables the speech corpus to have a diverse phoneme distribution.

[0086] In one embodiment, the screening in terms of acoustic feature dimension includes: obtaining the duration and speech rate values in the relevant information of the audio; judging whether the duration of the audio is greater than the preset duration threshold, if so, discarding the audio, if not, then judging whether the speech rate of the audio is greater than the preset speech rate threshold, if so, discarding the audio. For example, due to the limitation of the speech recognition model for the input of speech duration, the server discards the audios with a duration exceeding 10S; the server also judges the speech rate of the audio and discards the audios with an average speech rate greater than 12 words per second. It can be understood that screening the audio in terms of acoustic feature dimension ensures the speech quality and usability of the speech corpus.

[0087] In one embodiment, the screening in terms of speaker diversity dimension includes: if the dialogue role corresponding to the audio is an agent, obtaining the agent identifier included in the speaker information, obtaining the occupancy ratio of the audios with the same agent identifier in the speech corpus, and if the occupancy ratio is greater than the preset ratio threshold, discarding the audio. It can be understood that screening the audio in terms of speaker diversity dimension ensures that the speech phenomena in the speech corpus are rich.

[0088] In this embodiment, the server performs preliminary screening on the audio in the initial audio set based on the optimal transcription text and audio-related information to obtain an audio set, so as to ensure the speech quality, diverse phoneme distribution, and rich enough speech phenomena of the audio in the speech corpus.

[0089] In one embodiment, the method further includes: inputting the first transcription text into a statistical language model to obtain a perplexity metric; screening out the audio that meets the high-quality conditions from the audio set based on the difference information and the loss value includes: screening out the audio that meets the high-quality conditions from the audio set based on the difference information, the loss value, and the perplexity metric.

[0090] Among them, the statistical language model is a basic model of Natural Language Processing (NLP), which is used to obtain the PPL (perplexity) value of the audio. The perplexity metric, that is, the perplexity value, is used to represent the magnitude of the perplexity of the audio.

[0091] In one embodiment, the statistical language model used is an N-gram model (N-gram model) or a neural language model (NLM, Neural Language Model, a type of language model used to overcome the curse of dimensionality, which models natural language sequences using the distributed representation of words). In one embodiment, the N-gram model used is a 5-gram model.

[0092] Specifically, the server inputs the first transcription text into a statistical language model to obtain a perplexity metric. For each audio in the audio set, the server screens out the audio that meets the high-quality conditions from the audio set based on the corresponding difference information, loss value, and perplexity metric, and puts the audio and the corresponding transcription text into the speech corpus. After performing manual annotation processing on the audio that meets the medium-quality conditions, it is put into the speech corpus; and the low-quality audio is discarded. It can be understood that by screening the audio set, the audio quality of the speech corpus and the accuracy of the transcription text, that is, the annotation text, are ensured.

[0093] In one embodiment, when the difference information and the loss value corresponding to the audio meet the preset conditions, the perplexity metric is further compared with a preset perplexity threshold. If it is less than the preset perplexity threshold, the audio meets the high-quality conditions and is directly put into the speech corpus.

[0094] In this embodiment, based on the difference information, the loss value, and the perplexity metric, the audio that meets the high-quality conditions is screened out from the audio set to ensure the audio quality of the speech corpus and the accuracy of the transcription text, that is, the annotation text.

[0095] In one embodiment, the difference information includes the word error rates of the first transcribed text and the second transcribed text compared to the optimal transcribed text. Determining the difference information of the first transcribed text and the second transcribed text compared to the optimal transcribed text respectively includes: for any one of the first transcribed text and the second transcribed text, taking the optimal transcribed text as the standard, respectively determining the number of words added, reduced, and replaced in any one of the transcribed texts compared to the optimal transcribed text; and obtaining the word error rate of any one of the transcribed texts compared to the optimal transcribed text based on the number of words.

[0096] Among them, the word error rate refers to the ratio value obtained by comparing the recognized words with the words in the standard sentence and calculating the ratio between the number of different words and the number of words in the standard sentence.

[0097] Specifically, the server takes the optimal transcribed text as the standard for the first transcribed text, determines the number of words added, reduced, and replaced in the first transcribed text compared to the optimal transcribed text, and then obtains the word error rate of the first transcribed text by taking the ratio of the number of words to the number of words in the optimal transcribed text. Similarly, the server takes the optimal transcribed text as the standard for the second transcribed text, determines the number of words added, reduced, and replaced in the second transcribed text compared to the optimal transcribed text, and then obtains the word error rate of the second transcribed text by taking the ratio of the number of words to the number of words in the optimal transcribed text. It can be understood that the word error rate can represent the difference value of the transcribed text relative to the optimal transcribed text.

[0098] In one embodiment, the statistical formula for the word error rate is

[0099] Word error rate = ((number of inserted words + number of replaced words + number of reduced words) / total number of words in the standard sentence) * 100%

[0100] For example, if the content of the optimal transcribed text is "Have you eaten?", and the content of the first transcribed text is "Have you eaten eaten?", then the number of added words is 1, and the obtained word error rate is 33%. If the content of the first transcribed text is "Have eaten", then the number of reduced words is 1, and the obtained word error rate is 33%. If the content of the first transcribed text is "Have you eaten?", then the number of replaced words is 1, and the obtained word error rate is 33%.

[0101] For example, if the content of the optimal transcribed text is "Did you have breakfast this morning?", and the content of the first transcribed text is "Did you have lunch", then the number of reduced words is 2, and the number of replaced words is 1, and the obtained word error rate is (1 + 2) / 6 * 100% = 50%.

[0102] In this embodiment, for any one of the first transcription text and the second transcription text, taking the optimal transcription text as the standard, the character error rate of any one of the transcription texts compared with the optimal transcription text is determined respectively, so as to obtain the differences between the three transcription texts output by the speech recognition model to be optimized, the optimal recognition model and the reference recognition model, so as to accurately measure whether the transcription text of the audio is accurate.

[0103] In one embodiment, based on the difference information and loss value corresponding to each audio, screening out the audio that meets the high-quality conditions from the audio set includes: for each audio, determining the closeness between the character error rates of the corresponding first transcription text and the second transcription text compared with the optimal transcription text; if the closeness meets the preset closeness condition and the loss value is less than the preset loss value threshold, it is determined that the audio meets the high-quality conditions.

[0104] Specifically, the server calculates the character error rate of the first transcription text compared with the optimal transcription text and the character error rate of the second transcription text compared with the optimal transcription text for each audio in the audio set, and calculates the absolute value of the difference between the two character error rates, that is, the closeness. If the absolute value meets the preset absolute value threshold and the loss value is less than the preset loss value threshold, it is determined that the audio meets the high-quality conditions.

[0105] In another embodiment, the server determines the closeness between the character error rates of the corresponding first transcription text and the second transcription text compared with the optimal transcription text and the maximum value between the two character error rates for each audio; if the closeness meets the preset closeness condition, the maximum value of the character error rate is less than the preset character error rate threshold, and the loss value is less than the preset loss value threshold, it is determined that the audio meets the high-quality conditions.

[0106] In another embodiment, the server determines the closeness between the character error rates of the corresponding first transcription text and the second transcription text compared with the optimal transcription text and the maximum value between the two character error rates for each audio; if the closeness meets the preset closeness condition, the maximum value of the character error rate is less than the preset threshold, the loss value is less than the preset loss value threshold, and the perplexity index is less than the preset perplexity threshold, it is determined that the audio meets the high-quality conditions.

[0107] In this embodiment, the server determines the closeness between the character error rates of the corresponding first transcription text and the second transcription text compared with the optimal transcription text for each audio; if the closeness meets the preset closeness condition and the loss value is less than the preset loss value threshold, it is determined that the audio meets the high-quality conditions to ensure the accuracy of the transcription text of the selected audio.

[0108] In one embodiment, the method further includes: for the unselected audio, obtaining at least two pieces of manually annotated text annotated for the audio, where the manually annotated text is obtained by different annotators annotating the audio respectively; detecting whether the manually annotated texts are consistent; if they are not consistent, triggering a recheck annotation for the audio to obtain a recheck annotation text; and storing the audio and the recheck annotation text in a speech corpus.

[0109] Among them, the recheck annotation refers to the recheck annotator annotating the audio again. The unselected audio refers to the audio that is selected and does not meet the high-quality conditions and is not discarded during the process of step S206.

[0110] Specifically, for each unselected audio, the server assigns the audio to at least two annotators for annotation. After the annotators complete the annotation, at least two pieces of manually annotated text are generated. The server obtains at least two pieces of manually annotated text for the audio and detects whether the texts in the manually annotated texts are consistent. If they are consistent, the audio and the manually annotated text are put into the speech corpus as a group of data; if they are not consistent, the audio is assigned to another recheck annotator for recheck annotation to obtain a recheck annotation text, and the audio and the recheck annotation text are stored in the speech corpus as a group of data.

[0111] In one embodiment, the server can also evaluate the work efficiency of the annotators based on the step of whether the audio needs recheck annotation.

[0112] In one embodiment, the annotation of the audio is managed based on an annotation server. Specifically, the annotation server obtains the audio that needs to be manually annotated, and the annotation administrator uses the annotation server and assigns the audio to different annotators based on the annotation difficulty of the audio. The annotation server stores the grouping information of the audio. After different annotators complete the annotation of the audio on the workbench, the annotation server reads at least two pieces of annotation text for the same audio. If they are not consistent, it prompts the annotation administrator with information indicating that the audio annotation is incorrect, so that the annotation administrator can perform a recheck annotation process on the audio.

[0113] In this embodiment, the server obtains at least two pieces of manually annotated text for the audio, detects whether the manually annotated texts are consistent; if they are not consistent, it triggers a recheck annotation for the audio to obtain a recheck annotation text; and stores the audio and the recheck annotation text in the speech corpus to ensure the accuracy of the transcribed text in the speech corpus.

[0114] In one embodiment, the difference information and the loss value are included in the screening information set corresponding to the audio; the method further includes: for the un-screened audio, obtaining the audio-related information of the audio; inputting the audio-related information and the screening information set into the trained audio quality model to obtain the annotation difficulty corresponding to the audio; and grouping the un-screened audio based on the annotation difficulty, so that the corresponding annotators can annotate each group of audio to obtain the manually annotated text.

[0115] Among them, the annotation difficulty is used to represent the difficulty level of manually annotating the audio. The screening set is a set including input information for screening such as difference information and loss value.

[0116] In one embodiment, the screening information set includes difference information, loss value, and perplexity metric.

[0117] Specifically, different annotators have different annotation efficiencies and annotation capabilities. The server obtains the audio-related information of the audio for the un-screened audio. Inputting the audio-related information and the screening information set into the trained audio quality model to obtain the annotation difficulty corresponding to the audio; based on the annotation difficulty and the annotation efficiency and annotation capabilities of the annotators, copying and grouping the un-screened audio, and different groups belong to different annotators and different groups may include the same audio, so that the standard personnel can more accurately manually annotate the audio.

[0118] In another embodiment, the server obtains the audio-related information and the speaker information of the audio for the un-screened audio; inputs the audio-related information, the speaker information, and the screening information set into the trained audio quality model to obtain the annotation difficulty.

[0119] In one embodiment, before the training of the audio quality model is completed, the server obtains the standard difficulty evaluated manually corresponding to the audio, and uses the standard difficulty, the audio-related information, and the screening information set to train the audio quality model to obtain the trained audio quality model.

[0120] In one embodiment, the server can evaluate the annotators based on the annotation difficulty of the audio and whether the annotation results of the annotators are accurate.

[0121] In this embodiment, for the un-screened audio, based on the audio quality model, the annotation difficulty corresponding to the audio is obtained; based on the annotation difficulty, the un-screened audio is grouped, so that the corresponding annotators can annotate each group of audio to obtain the manually annotated text, thus ensuring the accuracy of the manually annotated text.

[0122] In one embodiment, such as Figure 3, a general framework diagram of the speech corpus generation method is described. Among them, the framework includes a dataset part, processing steps, and a model. The dataset part is divided into a long audio set used in the preliminary screening step, an audio set obtained after performing the preliminary screening, and a speech corpus. During the preliminary screening process, the server obtains long audio, performs VAD processing to obtain multiple short audio, and based on at least one of phoneme distribution, speaker information, and acoustic features, etc., after preliminary screening of the short audio, an audio set for sample grading is obtained. In the sample grading step, the server obtains the audio in the audio set, inputs it into multiple speech recognition models to obtain multiple transcription texts, and based on the transcription texts, obtains difference information. The audio and the optimal transcription text are input into the speech recognition model to be optimized to obtain a loss value. The transcription text output by the speech recognition model to be optimized is input into a statistical language model to obtain a perplexity metric. Finally, the server screens the audio based on the loss value, difference information, etc., puts the audio that meets high quality into the speech corpus, and performs manual annotation processing on the audio that meets medium quality conditions. During the manual annotation processing, the server first obtains the audio annotation difficulty based on the audio quality model, then groups the audio based on the audio annotation difficulty, and different annotators perform the annotation. If multiple annotation texts are inconsistent, after review processing, the audio and the final annotation text are put into the speech corpus. The server also continuously optimizes the speech recognition model to be optimized based on the speech corpus, and trains the audio quality model based on the difference information, loss value, etc. obtained in the sample grading.

[0123] It should be understood that although the steps in the flowcharts in some embodiments of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0124] In one embodiment, as Figure 4 shown, a speech corpus generation device 400 is provided, including: an acquisition module 402, an annotation module 404, a feature extraction module 406, and a screening module 408, where:

[0125] The acquisition module 402 is used to acquire an audio set; the audio set includes audio without corresponding transcription texts annotated.

[0126] The annotation module 404 is configured to perform speech transcription on each audio in the audio set through the speech recognition model to be optimized, the optimal speech recognition model, and the reference speech recognition model respectively, so as to obtain corresponding first transcription texts, optimal transcription texts, and second transcription texts; the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized.

[0127] The feature extraction module 406 is configured to respectively determine the difference information between the first transcription text and the second transcription text compared with the optimal transcription text; input the audio and the optimal transcription text into the speech recognition model to be optimized, and obtain the loss value of phoneme-level alignment.

[0128] The screening module 408 is configured to screen out the audios that meet the high-quality conditions from the audio set based on the difference information and loss value corresponding to each audio; store the screened audios and the corresponding optimal transcription texts as a set of data in the speech corpus.

[0129] The continuous optimization module 410 is configured to optimize the speech recognition model to be optimized based on each set of data in the speech corpus.

[0130] In one embodiment, the acquisition module 402 is further configured to: acquire an initial audio set; input the audios in the initial audio set into the optimal speech recognition model respectively, and obtain the optimal transcription texts corresponding to the audios and the relevant information of the audios; perform preliminary screening processing on the audios in the initial audio set based on the optimal transcription texts and audio relevant information, so as to obtain an audio set.

[0131] In one embodiment, the feature extraction module 404 is further configured to: input the first transcription text into a statistical language model to obtain a perplexity metric; screening out the audios that meet the high-quality conditions from the audio set based on the difference information and loss value includes: screening out the audios that meet the high-quality conditions from the audio set based on the difference information, loss value, and perplexity metric.

[0132] In one embodiment, the difference information includes the word error rate of the first transcription text and the second transcription text compared with the optimal transcription text; the feature extraction module 404 is further configured to: for any one of the first transcription text and the second transcription text, taking the optimal transcription text as a standard, respectively determine the number of words added, reduced, and replaced in any one of the transcription texts compared with the optimal transcription text; obtain the word error rate of any one of the transcription texts compared with the optimal transcription text based on the number of words.

[0133] In one embodiment, the screening module 408 is further configured to: for each audio, determine the degree of closeness between the word error rates of the corresponding first transcription text and the second transcription text compared with the optimal transcription text; if the degree of closeness meets the preset closeness condition and the loss value is less than the preset loss value threshold, it is determined that the audio meets the high-quality conditions.

[0134] In one embodiment, the screening module 408 is further configured to: for the un-screened audio, obtain at least two pieces of manually annotated text annotated for the audio, where the manually annotated text is obtained by different annotators annotating the audio respectively; detect whether the manually annotated texts are consistent; if not, trigger a review annotation for the audio to obtain a review annotated text; and store the audio and the review annotated text in the speech corpus.

[0135] In one embodiment, the difference information and the loss value are included in the screening information set corresponding to the audio; the speech corpus device 400 includes an audio grouping module 500, and the audio grouping module 500 includes an annotation difficulty acquisition module 502 and a grouping module 504, where:

[0136] The annotation difficulty acquisition module 502 is configured to, for the un-screened audio, obtain audio-related information of the audio; input the audio-related information and the screening information set into the trained audio quality model to obtain the annotation difficulty corresponding to the audio.

[0137] The grouping module 504 is configured to group the un-screened audio based on the annotation difficulty, so that the corresponding annotator annotates each group of audio to obtain the manually annotated text.

[0138] The above speech corpus generation device obtains an audio set including audio without corresponding transcribed text. For each audio in the audio set, the audio is respectively subjected to speech transcription by the speech recognition model to be optimized, the optimal speech recognition model, and the reference speech recognition model to obtain the corresponding first transcribed text, optimal transcribed text, and second transcribed text, where the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized. The difference information between the first transcribed text and the second transcribed text compared with the optimal transcribed text is respectively determined; then the audio and the optimal transcribed text are input into the speech recognition model to be optimized to obtain the loss value at the phoneme level alignment; based on the difference information and the loss value corresponding to each audio, the audio that meets the high-quality condition is screened out from the audio set. The screened audio and the corresponding optimal transcribed text are stored in the speech corpus as a set of data, thus ensuring the accuracy of the optimal transcribed text, that is, the annotated text. And the speech recognition model to be optimized is optimized based on each set of data in the speech corpus to realize the continuous optimization of the speech recognition model to be optimized, so as to increase the accuracy of the screened corpus data. Therefore, through the automatic annotation of speech, the screening of the annotated text, and the continuous optimization of the speech recognition model, the efficiency and quality of the speech corpus generation are improved.

[0139] For the specific limitations of the above voice corpus generation device, reference can be made to the limitations of the above voice corpus generation method in the foregoing text, which will not be elaborated herein. Each module in the above voice corpus generation device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0140] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, and a network interface connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store voice corpus data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a voice corpus generation method.

[0141] Those skilled in the art can understand that Figure 6 the structure shown in

[0142] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0143] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0144] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0145] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0146] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for generating a speech corpus, characterized in that The method includes: Obtaining an audio set; the audio set includes audio for which corresponding transcription texts are not labeled. For each audio in the audio set, performing speech transcription on the audio respectively through an speech recognition model to be optimized, an optimal speech recognition model, and a reference speech recognition model, to obtain corresponding first transcription texts, optimal transcription texts, and second transcription texts; the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized. Respectively determining the difference information between the first transcription text and the second transcription text compared to the optimal transcription text. Inputting the audio and the optimal transcription text into the speech recognition model to be optimized to obtain a loss value at the phoneme level alignment. Based on the difference information and the loss value corresponding to each audio, screening out the audio that meets the high-quality conditions from the audio set. Storing the screened audio and the corresponding optimal transcription text as a set of data in a speech corpus. Optimizing the speech recognition model to be optimized based on each set of data in the speech corpus.

2. The method according to claim 1, characterized in that The obtaining the audio set includes: Obtaining an initial audio set. Inputting the audio in the initial audio set into the optimal speech recognition model respectively to obtain the optimal transcription text corresponding to the audio and the relevant information of the audio. Performing a preliminary screening process on the audio in the initial audio set based on the optimal transcription text and the audio relevant information to obtain an audio set.

3. The method according to claim 1, wherein The method further includes: Inputting the first transcription text into a statistical language model to obtain a perplexity metric. The screening out the audio that meets the high-quality conditions from the audio set based on the difference information and the loss value includes: Screening out the audio that meets the high-quality conditions from the audio set based on the difference information, the loss value, and the perplexity metric.

4. The method according to claim 1, wherein The difference information includes the word error rate of the first transcription text and the second transcription text compared to the optimal transcription text; the respectively determining the difference information between the first transcription text and the second transcription text compared to the optimal transcription text includes: For any one of the first transcription text and the second transcription text, taking the optimal transcription text as a standard, respectively determining the number of words added, reduced, and replaced in the any one of the transcription texts compared to the optimal transcription text. Obtaining the word error rate of the any one of the transcription texts compared to the optimal transcription text based on the number of words.

5. The method according to claim 4, characterized in that The screening out the audio that meets the high-quality conditions from the audio set based on the difference information and the loss value corresponding to each audio includes: For each audio, determining the degree of proximity between the word error rates of the corresponding first transcription text and the second transcription text compared to the optimal transcription text. If the degree of proximity meets a preset proximity condition and the loss value is less than a preset loss value threshold, it is determined that the audio meets the high-quality conditions.

6. The method according to claim 1, characterized in that The method further includes: For the audio not screened out, obtaining at least two manually annotated texts annotated for the audio, and the manually annotated texts are obtained by different annotators respectively annotating the audio. Detecting whether the manually annotated texts are consistent. If they are inconsistent, trigger a review and annotation of the audio to obtain a review annotation text; Store the audio and the review annotation text in the speech corpus.

7. The method according to any one of claims 1 to 6, characterized in that, The difference information and the loss value are included in the screening information set corresponding to the audio; The method further includes: For the audio that has not been screened out, obtain the audio-related information of the audio; Input the audio-related information and the screening information set into the trained audio quality model to obtain the annotation difficulty corresponding to the audio; Based on the annotation difficulty, group the audio that has not been screened out so that the corresponding annotators can annotate each group of audio to obtain an artificial annotation text.

8. A voice corpus generation device, characterized in that, The device includes: An acquisition module, configured to acquire an audio set; the audio set includes audio for which no corresponding transcription text has been annotated; An annotation module, configured to, for each audio in the audio set, perform speech transcription on the audio through a speech recognition model to be optimized, an optimal speech recognition model, and a reference speech recognition model respectively, to obtain corresponding first transcription texts, optimal transcription texts, and second transcription texts; the recognition accuracy of the reference speech recognition model is lower than that of the speech recognition model to be optimized; A feature extraction module, configured to respectively determine the difference information between the first transcription text and the second transcription text compared with the optimal transcription text; input the audio and the optimal transcription text into the speech recognition model to be optimized to obtain a loss value of phoneme-level alignment; A screening module, configured to screen out the audio that meets the high-quality conditions from the audio set based on the difference information and the loss value corresponding to each audio; store the screened audio and the corresponding optimal transcription text as a set of data in the speech corpus; A continuous optimization module, configured to optimize the speech recognition model to be optimized based on each set of data in the speech corpus.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Audio corpus screening method and device for speech recognition and computer equipment

    CN110263322A

  • Voice recognition and model training method and device, equipment and storage medium

    CN111243576A