A method, apparatus, device, and medium for obtaining speech training data

By splitting and denoising the multi-channel audio, splitting it into a single speaker segment, voice recognition and sound quality enhancement, the problem of poor training data quality in speech synthesis is solved, and high-quality voice training data acquisition is achieved.

CN119993196BActive Publication Date: 2025-07-04BEIJING YUNSHANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510147634.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-04
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

The existing speech synthesis technology has high demand on training data scale, and there are background sound, noise, speakers and sound quality problems, resulting in unnatural speech generation effects and poor training data quality.

Method used

Through multi-channel audio splitting into a single channel, background music and noise are removed, multi-person conversations are split into a single speaker segment using voice activity detection technology, language recognition and speech recognition are performed, punctuation is added, and synthetic speech automatic evaluation quality evaluation model is used to filter and sound quality enhancement processing to obtain high-quality speech training data.

Benefits of technology

The quality of speech training data is improved, the naturalness and consistency of speech generation is ensured, the damage to human voices is reduced by noise, the sound quality is improved, and high-quality speech training data is obtained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993196B_ABST
    Figure CN119993196B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and medium for obtaining speech training data, which relates to the field of intelligent speech technology. The method includes: splitting multi-channel audio into single channels; removing background music and background noise; splitting multi-person conversation audio into single speaker segments; adding punctuation; and enhancing the audio quality of audio with poor quality scores, so as to obtain speech training data with high corpus quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent voice technology, and particularly to a method, device, equipment and medium for obtaining voice training data. Background Art

[0002] In human-computer interaction, voice interaction is one of the most natural ways, which is very convenient and easy to understand. In recent years, the development of artificial intelligence technology has advanced by leaps and bounds. At present, machines can already recognize and understand the internal meaning of voices and make corresponding responses in a timely manner, such as playing voices, calling relevant skills, etc. In this interaction process, the effects of voice synthesis and voice cloning greatly affect the user experience. It is necessary to provide multi-language and mixed-language support, and control the speech rate, emotion, timbre, and sound quality of the generated voice.

[0003] The current voice synthesis is based on the VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) model framework, such as Bert-VITS2. However, it has problems such as mechanical and unnatural pronunciation and inability to clone timbre without samples. In recent years, the development of large language models has been in full swing. Directions such as images and voices are constantly leveraging the framework of large language models to improve the algorithm performance in their respective directions. This includes voice synthesis and voice cloning. The latest voice generation method is a large voice generation model based on voice quantization coding, which is a new technology that deeply integrates text understanding and voice generation. It discretizes the encoding of voices and generates natural and fluent voices with the help of large model technology. Compared with traditional methods, its naturalness is like that of a real person. However, in terms of the scale of training data, it requires a large amount of data, usually tens of thousands of hours of training data. The magnitude of its training data is already comparable to that of speech recognition. Speech recognition requires generalization ability, so it is not sensitive to problems such as background noise, noise, multiple speakers, and sound quality in the data. However, voice generation requires control of speech rate, emotion, timbre, and sound quality, so its requirements for training corpora are relatively high.

[0004] Currently, corpora are mostly obtained from audio platforms, video platforms, and open-source data sets, etc., which have the defect of poor corpus quality. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, equipment and medium for obtaining voice training data, which can obtain voice training data with high corpus quality.

[0006] To achieve the above purpose, this application provides the following solutions:

[0007] In the first aspect, this application provides a method for obtaining voice training data, including:

[0008] Obtain a large amount of audio data;

[0009] Split the multi-channel audio data in the audio data into single-channel audio data;

[0010] Remove the background music and background noise in all the single-channel audio data to obtain denoised single-channel audio data;

[0011] According to the speaker log, use the voice activity detection technology to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments, where the speaker log includes the speaking time and speaking order of each speaker;

[0012] Screen out the single-speaker audio segments that meet the preset duration;

[0013] Perform language recognition and speech recognition on the single-speaker audio segments that meet the preset duration in sequence to obtain the text of the recognized audio segments;

[0014] Add punctuation to the text of the recognized audio segments to obtain the text with punctuation added;

[0015] Use the synthetic speech automatic evaluation quality evaluation model to evaluate the quality of the audio segments corresponding to the text with punctuation added;

[0016] Perform sound quality enhancement processing on the audio segments with a quality evaluation less than the quality threshold to obtain enhanced audio segments, where the sound quality enhancement processing includes replacing the timbre, restoring audio distortion, and expanding the audio bandwidth;

[0017] Use the enhanced audio segments as speech training data.

[0018] In a second aspect, the present application provides a device for obtaining speech training data, including:

[0019] A large amount of audio data acquisition module, configured to obtain a large amount of audio data;

[0020] A first splitting module, configured to split the multi-channel audio data in the audio data into single-channel audio data;

[0021] A noise removal module, configured to remove the background music and background noise in all the single-channel audio data to obtain denoised single-channel audio data;

[0022] A second splitting module, configured to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments according to the speaker log by using the voice activity detection technology, where the speaker log includes the speaking time and speaking order of each speaker;

[0023] A screening module for screening out the single-speaker audio segments that meet the preset duration;

[0024] An identification module for sequentially performing language identification and speech recognition on the single-speaker audio segments that meet the preset duration to obtain the text of the identified audio segments;

[0025] A punctuation adding module for adding punctuation to the text of the identified audio segments to obtain the text with added punctuation;

[0026] A quality assessment module for automatically evaluating the quality of the audio segments corresponding to the text with added punctuation by using a synthesized speech automatic quality assessment model;

[0027] A sound quality enhancement module for performing sound quality enhancement processing on the audio segments with a quality assessment less than the quality threshold to obtain the enhanced audio segments, wherein the sound quality enhancement processing includes replacing the timbre, restoring audio distortion, and expanding the audio bandwidth;

[0028] A speech training data acquisition module for using the enhanced audio segments as speech training data.

[0029] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method for acquiring speech training data described in the first aspect above.

[0030] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for acquiring speech training data described in the first aspect above is implemented.

[0031] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0032] The present application provides a method, device, equipment, and medium for acquiring speech training data. The method includes: splitting multi-channel audio into single channels; removing background music and background noise; splitting multi-person dialogue audio into single-speaker segments; adding punctuation; and performing sound quality enhancement on audio with a poor quality score, so as to obtain speech training data with high corpus quality. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 Schematic flowchart of a method for obtaining speech training data provided in Embodiment 1 of this application;

[0035] Figure 2 Schematic diagram of the structure of a computer device provided in Embodiment 3 of this application. Detailed implementation manners

[0036] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0037] To make the above objects, features, and advantages of this application more obvious and understandable, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific implementation manners.

[0038] Embodiment 1

[0039] Since the massive audio data obtained from audio platforms, video platforms, open-source data sets, etc. has complex scenarios and generally includes background sounds, noises, multi-person conversations, and uneven audio quality.

[0040] In response to this, this embodiment provides a method for obtaining speech training data, including:

[0041] S1: Obtain massive audio data.

[0042] S2: Split the multi-channel audio data in the audio data into single-channel audio data.

[0043] S3: Remove the background music and background noise in all the single-channel audio data to obtain denoised single-channel audio data.

[0044] S31: Use a noise evaluation model to evaluate whether the single-channel audio data needs noise reduction to obtain an evaluation result.

[0045] S32: When the evaluation result is yes, use a speech noise reduction model to remove the background noise in the single-channel audio data.

[0046] S4: According to the speaker log, use speech activity detection technology to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments, where the speaker log includes the speaking time and speaking order of each speaker.

[0047] S5: Screen out the single-speaker audio segments that meet the preset duration.

[0048] S6: Perform language identification and speech recognition on the single-speaker audio segments that meet the preset duration in sequence to obtain the text of the recognized audio segments.

[0049] S7: Add punctuation to the text of the recognized audio segments to obtain the text with punctuation added.

[0050] S8: Use a synthetic speech automatic evaluation quality assessment model to evaluate the quality of the audio segments corresponding to the text with punctuation added.

[0051] S9: Perform sound quality enhancement processing on the audio segments with a quality assessment less than the quality threshold to obtain the enhanced audio segments. Among them, the sound quality enhancement processing includes replacing the timbre, restoring audio distortion, and expanding the audio bandwidth.

[0052] S10: Use the enhanced audio segments as speech training data.

[0053] The following Figure 1 specifically explains the conceptual process of the method for obtaining speech training data provided in this embodiment.

[0054] The processing flow of the method for obtaining speech training data is as follows:

[0055] 1) Convert multi-channel audio to single-channel.

[0056] 2) Remove background music.

[0057] 3) Remove background noise. In theory, an ideal speech noise reduction model should be able to identify and remove background noise while retaining the human voice. However, in practical applications, it is inevitable to accidentally damage the human voice, resulting in the inability to obtain high-quality human voices. By first using a noise assessment model to evaluate whether the audio needs noise reduction, and then using a speech noise reduction model to identify and remove background noise, the proportion of high-quality human voices in the corpus can be improved.

[0058] 4) According to the speaker log, use VAD (Voice Activity Detection technology) to split the multi-person conversation audio into single-speaker segments.

[0059] 5) Perform language identification and speech recognition on the audio segments with qualified durations, and obtain the timestamp of each word or syllable in the ASR (Automatic Speech Recognition) recognition result. Screen the audio segments through the ASR recognition result.

[0060] However, there are errors in ASR recognition itself. How to improve the accuracy of the transcript (i.e., the transcribed text) is very important.

[0061] a) Recognition quality assessment of audio segments with existing transcriptions: For the audio segments of the existing transcript, text regularization is performed, that is, convert 2019 to two thousand and nineteen. In this embodiment, methods such as WFST (Weighted Finite-State Transducers) based on grammar rules or neural network-based models can be used for text regularization; calculate the word error rate or character error rate by computing the edit distance. Since the current mainstream Chinese recognition models directly use Chinese characters for modeling, there are cases of recognizing other homophonic characters. Therefore, in this embodiment, the edit distance of pinyin is used to calculate the word error rate and character error rate.

[0062] b) Recognition quality assessment of audio segments without transcriptions: For the audio segments without transcript, in this embodiment, the reliability of the recognition result is judged through ASR confidence estimation. Among them, confidence estimation can adopt acoustic information, such as the character-level ASR confidence estimation method based on entropy, or a language model can be used to score the ASR recognition result, or both methods can be used jointly.

[0063] c) For the audio segments with recognition quality greater than the quality threshold, calculate the average pronunciation duration of each character in the audio segment, judge whether it is reasonable, and filter out unreasonable audio segments.

[0064] 6) Add punctuation to the ASR recognition result. Usually, it is expected to control pauses through punctuation during speech generation, but the model for adding punctuation to the text (i.e., the text punctuation model) usually does not refer to the acoustic results of ASR, so the punctuation and pauses cannot be consistent. In this embodiment, the ASR acoustic information, that is, the timestamp interval of each character, can be used to judge whether to add or delete punctuation.

[0065] 7) Screen through MOS (Mean Opinion Score). Usually, the crowdsourcing form is adopted, which is costly for a large-scale training set. In this embodiment, the MOS value is estimated through a synthetic speech automatic quality evaluation model to further ensure the audio quality.

[0066] 8) In view of the phenomenon that a large number of audio with unqualified sound quality are discarded, this embodiment enhances the sound quality of the audio with poor quality scores. It includes replacing with high-quality timbres, restoring audio distortion, and expanding the audio bandwidth to improve the perceived audio quality.

[0067] Generally speaking, the method for obtaining the above-mentioned speech training data in this embodiment is to denoise and separate the massive speech data obtained into high-quality corpora of a single speaker. The method for obtaining the above-mentioned speech training data provided in this embodiment can achieve: 1. reducing the damage of noise to human voices; 2. judging whether the ASR results are accurately recognized; 3. making the punctuation consistent with the pauses in the audio; 4. obtaining the MOS score using a neural network; 5. improving the sound quality of audio with poor sound quality, so as to obtain speech training data with high corpus quality.

[0068] Embodiment 2

[0069] This embodiment provides an apparatus for obtaining speech training data, including:

[0070] A massive audio data acquisition module, configured to acquire massive audio data;

[0071] A first splitting module, configured to split the multi-channel audio data in the audio data into single-channel audio data;

[0072] A noise removal module, configured to remove background music and background noise in all the single-channel audio data to obtain denoised single-channel audio data;

[0073] A second splitting module, configured to split the multi-person dialogue audio in the denoised single-channel audio data into single-speaker audio segments according to the speaker log by using voice activity detection technology, where the speaker log includes the speaking time and speaking order of each speaker;

[0074] A screening module, configured to screen out the single-speaker audio segments that meet a preset duration;

[0075] An identification module, configured to sequentially perform language identification and speech recognition on the single-speaker audio segments that meet the preset duration to obtain the text of the identified audio segments;

[0076] A punctuation adding module, configured to add punctuation to the text of the identified audio segments to obtain text with added punctuation;

[0077] A quality evaluation module, configured to evaluate the quality of the audio segments corresponding to the text with added punctuation by using a synthetic speech automatic evaluation quality measurement model;

[0078] A sound quality enhancement module, configured to perform sound quality enhancement processing on the audio segments with a quality evaluation less than a quality threshold to obtain enhanced audio segments, where the sound quality enhancement processing includes replacing the timbre, restoring audio distortion, and expanding the audio bandwidth;

[0079] A speech training data acquisition module, configured to use the enhanced audio segments as speech training data.

[0080] Example 3

[0081] This embodiment provides a computer device, which can be a server or a terminal, and its internal structure diagram can be as shown in Figure 2 . The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store voice training data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the method for obtaining voice training data provided in Embodiment 1.

[0082] Those skilled in the art can understand that Figure 2 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor, and a computer program is stored in the memory. When the processor executes the computer program, it implements the steps in the above method embodiments.

[0083] Example 4

[0084] This embodiment provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method for obtaining voice training data provided in Embodiment 1 above.

[0085] Example 5

[0086] This embodiment provides a computer program product including a computer program, which when executed by a processor implements the method for obtaining voice training data provided in Embodiment 1 above.

[0087] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0088] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0089] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., and are not limited thereto.

[0090] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0091] In this text, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for obtaining speech training data, characterized in that, The method for obtaining the speech training data includes: Obtaining a large amount of audio data; Splitting the multi-channel audio data in the audio data into single-channel audio data; Removing the background music and background noise in all the single-channel audio data to obtain denoised single-channel audio data; According to the speaker log, using the voice activity detection technology to split the multi-person dialogue audio in the denoised single-channel audio data into single-speaker audio segments, where the speaker log includes the speaking time and speaking order of each speaker; Screening out the single-speaker audio segments that meet the preset duration; Successively performing language recognition and speech recognition on the single-speaker audio segments that meet the preset duration to obtain the text of the recognized audio segments; Adding punctuation to the text of the recognized audio segments to obtain the text with punctuation added; Using a synthetic speech automatic evaluation quality evaluation model to evaluate the quality of the audio segments corresponding to the text with punctuation added; Performing sound quality enhancement processing on the audio segments with a quality evaluation less than the quality threshold to obtain enhanced audio segments, where the sound quality enhancement processing includes replacing the timbre, restoring audio distortion, and expanding the audio bandwidth; Using the enhanced audio segments as speech training data; After performing the step of "successively performing language recognition and speech recognition on the single-speaker audio segments that meet the preset duration to obtain the text of the recognized audio segments", the method for obtaining the speech training data further includes: Performing text regularization processing on the audio segments with existing transcribed text, where the audio segments with existing transcribed text are the audio segments in the single-speaker audio segments that meet the preset duration, to obtain the regularized transcribed text; Calculating the edit distance between the pinyin of the regularized transcribed text and the pinyin of the corresponding text of the recognized audio segments; Evaluating the first recognition quality according to the edit distance; Filtering out the text of the recognized audio segments with the first recognition quality less than the first quality threshold to obtain the text of the first recognized audio segments; After performing the step of "successively performing language recognition and speech recognition on the single-speaker audio segments that meet the preset duration to obtain the text of the recognized audio segments", the method for obtaining the speech training data further includes: Calculating the ASR confidence between the audio segments without transcribed text and the corresponding text of the recognized audio segments, where the audio segments without transcribed text are the audio segments in the single-speaker audio segments that meet the preset duration; Evaluating the second recognition quality according to the ASR confidence; Filtering out the text of the recognized audio segments with the second recognition quality less than the second quality threshold to obtain the text of the second recognized audio segments.

2. The method for obtaining speech training data according to claim 1, wherein The process of removing the background noise in all the single-channel audio data specifically includes: Using a noise evaluation model to evaluate whether the single-channel audio data needs noise reduction to obtain an evaluation result; When the evaluation result is yes, using a speech noise reduction model to remove the background noise in the single-channel audio data.

3. The method for obtaining speech training data according to claim 1, wherein After performing the steps of "filtering the text of the recognized post-audio segment with the first recognition quality less than the first quality threshold to obtain the first recognized post-audio segment text" and "filtering the text of the recognized post-audio segment with the second recognition quality less than the second quality threshold to obtain the second recognized post-audio segment text", the method for obtaining the speech training data further includes: Calculating the average pronunciation duration of each word in the target audio segment, where the target audio segment is the text of the first recognized post-audio segment and the text of the second recognized post-audio segment; Filtering the target audio segment according to the average pronunciation duration of each word and the duration threshold.

4. The method for obtaining speech training data according to claim 1, wherein Adding punctuation to the text of the recognized post-audio segment to obtain the text with punctuation added, specifically including: Obtaining the ASR acoustic information of the audio segment corresponding to the text of the recognized post-audio segment, where the ASR acoustic information is the timestamp interval between adjacent words in the corresponding audio segment; Based on the ASR acoustic information, adjusting the punctuation position through a text punctuation model, where the punctuation position includes the insertion position and the deletion position.

5. The method for obtaining speech training data according to claim 1, wherein Evaluating the quality of the audio segment corresponding to the text with punctuation added, specifically including: Evaluating the quality of the audio segment corresponding to the text with punctuation added by using a synthetic speech automatic evaluation quality measurement model.

6. An apparatus for obtaining speech training data, characterized in that, The apparatus for obtaining the speech training data includes: A massive audio data acquisition module for acquiring massive audio data; A first splitting module for splitting the multi-channel audio data in the audio data into single-channel audio data; A noise removal module for removing background music and background noise in all the single-channel audio data to obtain denoised single-channel audio data; A second splitting module for splitting the multi-person dialogue audio in the denoised single-channel audio data into single-speaker audio segments according to the speaker log by using a voice activity detection technique, where the speaker log includes the speaking time and speaking order of each speaker; A screening module for screening out the single-speaker audio segments that meet the preset duration; An identification module for sequentially performing language identification and speech recognition on the single-speaker audio segments that meet the preset duration to obtain the text of the recognized post-audio segment; A punctuation adding module for adding punctuation to the text of the recognized post-audio segment to obtain the text with punctuation added; A quality evaluation module for evaluating the quality of the audio segment corresponding to the text with punctuation added by using a synthetic speech automatic evaluation quality measurement model; A sound quality enhancement module for performing sound quality enhancement processing on the audio segments with quality evaluation less than the quality threshold to obtain enhanced audio segments, where the sound quality enhancement processing includes replacing the timbre, restoring audio distortion, and expanding the audio bandwidth; A speech training data acquisition module for using the enhanced audio segments as speech training data; The first audio segment text processing module is used to perform text regularization processing on the audio segments of the existing transcribed text to obtain the regularized transcribed text. Among them, the audio segments of the existing transcribed text are the audio segments in the single speaker audio segments that meet the preset duration; calculate the edit distance between the pinyin of the regularized transcribed text and the pinyin of the corresponding recognized audio segment text; evaluate the first recognition quality according to the edit distance; filter the recognized audio segment text with the first recognition quality less than the first quality threshold to obtain the first recognized audio segment text. The second audio segment text processing module is used to calculate the ASR confidence between the audio segments without transcribed text and the corresponding recognized audio segment text. Among them, the audio segments without transcribed text are the audio segments in the single speaker audio segments that meet the preset duration; evaluate the second recognition quality according to the ASR confidence; filter the recognized audio segment text with the second recognition quality less than the second quality threshold to obtain the second recognized audio segment text.

7. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method for obtaining speech training data according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for obtaining speech training data according to any one of claims 1-5.

Citation Information

Patent Citations

  • Audio quality evaluation method and device, electronic equipment and storage medium

    CN113782036A

  • Voice library training data analysis method and device

    CN113889096A