Voice training data acquisition method and device, equipment and medium
By performing multi-step processing of massive audio data, including denoising, multi-person conversation splitting, speech recognition and sound quality enhancement, the problems of unnatural pronunciation and difficulty in cloning tone in existing speech synthesis technology are solved, and high-quality speech training data are obtained.
Patent Information
- Application Number
- CN202510147634.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing speech synthesis technology has problems such as unnatural pronunciation and inability to clone timbre in zero samples, and the obtained speech training data is of poor quality.
By obtaining massive audio data, splitting multi-channel audio into a single channel, removing background music and noise, splitting multi-person conversation into a single speaker audio clip, performing language recognition and speech recognition, adding punctuation, and performing quality evaluation and sound quality enhancement through synthesizing speech automatic evaluation quality evaluation model.
Speech training data with high corpus quality was obtained, which solved the naturalness and tone cloning problems in speech synthesis, and improved the quality of the training data.
Smart Images

Figure CN119993196A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent voice technology, and in particular to a method, device, equipment and medium for acquiring voice training data. Background Art
[0002] In human-computer interaction, voice interaction is the most natural way, which is very convenient and easy to understand. In recent years, artificial intelligence technology has developed by leaps and bounds. At present, machines can recognize and understand the inner meaning of voice and make corresponding responses in time, such as playing voice and calling related skills. In this interactive process, the effects of speech synthesis and sound cloning have a great impact on user experience. It is necessary to provide support for multiple languages and mixed languages, and control the speed, emotion, timbre and sound quality of the generated voice.
[0003] The current speech synthesis is based on the VITS (Variational Inference with adversariallearning for end-to-end Text-to-Speech) model framework. For example, Bert-VITS2. However, it has problems such as mechanical and unnatural pronunciation and inability to clone timbre with zero samples. In recent years, the development of language big models has been in full swing. The image, speech and other directions are constantly using the framework of language big models to improve the algorithm performance in their directions, including speech synthesis and sound cloning. The latest speech generation method is a speech generation big model based on speech quantization coding. It is a new technology that deeply integrates text understanding and speech generation. It discretizes speech and generates natural and fluent speech with the help of big model technology. Compared with traditional methods, its naturalness is like that of a real person. However, in terms of the scale of training data, it requires massive data, usually tens of thousands of hours of training data. The scale of its training data is comparable to that of speech recognition. Speech recognition requires generalization ability, so it is not sensitive to background sound, noise, multiple speakers, sound quality and other issues in the data. However, speech generation requires control of speech speed, emotion, timbre and sound quality, so its training corpus requirements are relatively high.
[0004] Currently, most corpora are obtained from audio platforms, video platforms, and open source data sets, which results in poor corpus quality. Summary of the invention
[0005] The purpose of this application is to provide a method, device, equipment and medium for obtaining speech training data, which can obtain speech training data of corpus quality.
[0006] To achieve the above objectives, this application provides the following solutions:
[0007] In a first aspect, the present application provides a method for acquiring speech training data, comprising:
[0008] Get massive audio data;
[0009] Splitting the multi-channel audio data in the audio data into single-channel audio data;
[0010] Removing background music and background noise from all the single-channel audio data to obtain denoised single-channel audio data;
[0011] According to the speaker log, voice activity detection technology is used to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments, wherein the speaker log includes the speaking time and speaking order of each speaker;
[0012] Filter out the single speaker audio segment that meets the preset duration;
[0013] Performing language recognition and speech recognition on the single speaker audio segment that meets the preset time length in sequence to obtain a recognized audio segment text;
[0014] Adding punctuation to the recognized audio segment text to obtain punctuation-added text;
[0015] Using a synthetic speech automatic assessment quality evaluation model to perform quality assessment on the audio segment corresponding to the punctuation-added text;
[0016] Performing sound quality enhancement processing on the audio segment whose quality evaluation is less than the quality threshold to obtain an enhanced audio segment, wherein the sound quality enhancement processing includes replacing timbre, restoring audio distortion, and expanding audio bandwidth;
[0017] The enhanced audio segment is used as speech training data.
[0018] In a second aspect, the present application provides a device for acquiring speech training data, comprising:
[0019] A massive audio data acquisition module, used to acquire massive audio data;
[0020] A first splitting module, used for splitting the multi-channel audio data in the audio data into single-channel audio data;
[0021] A denoising module, used to remove background music and background noise from all the single-channel audio data to obtain denoised single-channel audio data;
[0022] A second splitting module is used to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments according to the speaker log and using voice activity detection technology, wherein the speaker log includes the speaking time and speaking order of each speaker;
[0023] A screening module, used to screen out the audio segment of the single speaker that meets the preset time length;
[0024] A recognition module, used to perform language recognition and voice recognition on the single speaker audio segment that meets the preset time length in sequence to obtain a recognized audio segment text;
[0025] A punctuation adding module, used for adding punctuation to the recognized audio segment text to obtain a punctuation added text;
[0026] A quality assessment module, used to perform quality assessment on the audio segment corresponding to the punctuation-added text using a synthetic speech automatic assessment quality evaluation model;
[0027] A sound quality enhancement module, configured to perform sound quality enhancement processing on the audio segment whose quality evaluation is less than the quality threshold to obtain an enhanced audio segment, wherein the sound quality enhancement processing includes replacing timbre, restoring audio distortion and expanding audio bandwidth;
[0028] The speech training data acquisition module is used to use the enhanced audio segment as speech training data.
[0029] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for acquiring speech training data described in the first aspect above.
[0030] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for acquiring speech training data described in the first aspect above.
[0031] According to the specific embodiments provided in this application, this application has the following technical effects:
[0032] The present application provides a method, apparatus, device and medium for acquiring speech training data, the method comprising: splitting multi-channel audio into single channels; removing background music and background noise; splitting multi-person conversation audio into single speaker segments; adding punctuation; and enhancing the sound quality of audio with poor quality scores, so as to obtain speech training data of corpus quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0034] Figure 1 A flowchart of a method for acquiring speech training data provided in Example 1 of the present application;
[0035] Figure 2 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0037] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0038] Example 1
[0039] Due to the massive amount of audio data obtained from audio platforms, video platforms, open source data sets, etc., the scenes are complex and generally include background sounds, noise, multi-person conversations, and uneven sound quality.
[0040] In this regard, this embodiment provides a method for acquiring speech training data, including:
[0041] S1: Obtain massive audio data.
[0042] S2: Split the multi-channel audio data in the audio data into single-channel audio data.
[0043] S3: removing background music and background noise from all the single-channel audio data to obtain denoised single-channel audio data.
[0044] S31: Using a noise evaluation model to evaluate whether the single-channel audio data needs noise reduction, and obtaining an evaluation result.
[0045] S32: When the evaluation result is yes, the background noise in the single-channel audio data is removed using a speech noise reduction model.
[0046] S4: According to the speaker log, voice activity detection technology is used to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments, wherein the speaker log includes the speaking time and speaking order of each speaker.
[0047] S5: Filter out the audio segment of the single speaker that meets the preset time length.
[0048] S6: Performing language recognition and voice recognition on the single speaker audio segment that meets the preset time length in sequence to obtain a recognized audio segment text.
[0049] S7: adding punctuation marks to the recognized audio segment text to obtain punctuation-added text.
[0050] S8: Using a synthetic speech automatic assessment quality evaluation model to perform quality assessment on the audio segment corresponding to the punctuation-added text.
[0051] S9: Performing sound quality enhancement processing on the audio segment whose quality evaluation is less than the quality threshold to obtain an enhanced audio segment, wherein the sound quality enhancement processing includes replacing timbre, restoring audio distortion, and expanding audio bandwidth.
[0052] S10: Using the enhanced audio segment as speech training data.
[0053] The following combination Figure 1 The conceptual process of the method for acquiring speech training data provided in this embodiment is specifically explained.
[0054] The processing flow of the method for obtaining speech training data is as follows:
[0055] 1) Multi-channel audio is converted to single channel.
[0056] 2) Remove background music.
[0057] 3) Remove background noise. In theory, an ideal speech noise reduction model should be able to identify and remove background noise while retaining human voices. However, in practical applications, human voices are inevitably accidentally damaged, making it impossible to obtain high-quality human voices. By first using a noise assessment model to assess whether the audio needs noise reduction, and then using a speech noise reduction model to identify and remove background noise, the proportion of high-quality human voices in the corpus can be increased.
[0058] 4) Based on the speaker log, VAD (Voice Activity Detection) is used to split the multi-person conversation audio into single speaker segments.
[0059] 5) Perform language recognition and speech recognition on the audio clips with qualified duration, obtain the timestamp of each word or syllable in the ASR (Automatic Speech Recognition) recognition result, and filter the audio clips according to the ASR recognition result.
[0060] However, ASR recognition itself has errors, so it is very important to improve the accuracy of the transcript.
[0061] a) Recognition quality assessment of audio clips with existing transcribed texts: For audio clips with existing transcripts, text regularization is performed, that is, 2019 is converted to 2019. This embodiment can use methods such as WFST (Weighted Finite-State Transducers) based on grammatical rules or neural network-based models to perform text regularization; the word error rate or character error rate is calculated by calculating the edit distance. Since the current mainstream Chinese recognition model directly uses Chinese characters for modeling, there are situations where other homophones are recognized. For this, this embodiment uses the edit distance of pinyin to calculate the word error rate and character error rate.
[0062] b) Recognition quality assessment of audio clips without transcripts: For audio clips without transcripts, this embodiment determines the reliability of the recognition results through ASR confidence estimation. The confidence estimation can use acoustic information, such as the word-level ASR confidence estimation method based on entropy, or use a language model to score the ASR recognition results, or use the two methods together.
[0063] c) For audio segments whose recognition quality is greater than the quality threshold, the average pronunciation duration of each word in the audio segment is calculated to determine whether it is reasonable, and unreasonable audio segments are filtered out.
[0064] 6) Add punctuation to the ASR recognition results. Usually, when generating speech, it is expected that pauses are controlled by punctuation. However, the model for adding punctuation to text (i.e., the text punctuation model) usually does not refer to the acoustic results of ASR, so the punctuation and pauses cannot be consistent. This embodiment can determine whether to add or delete punctuation through ASR acoustic information, i.e., the timestamp interval of each word.
[0065] 7) Screening by MOS (Mean Opinion Score). Crowdsourcing is usually used, which is costly for massive training sets. This embodiment estimates the MOS value through a synthetic speech automatic quality evaluation model to further ensure audio quality.
[0066] 8) In order to address the phenomenon that a large amount of audio with substandard sound quality is discarded, this embodiment enhances the sound quality of audio with poor quality scores, including replacing high-quality timbre, restoring audio distortion, and expanding audio bandwidth to improve perceived audio quality.
[0067] The method for acquiring the speech training data in this embodiment is generally to denoise and separate the acquired massive speech data into high-quality corpus of a single speaker. The method for acquiring the speech training data provided in this embodiment can achieve: 1. reduce the damage of noise to human voice; 2. determine whether the ASR result is accurate; 3. make punctuation and pauses of audio consistent; 4. use neural network to obtain MOS score; 5. improve the sound quality of audio with poor sound quality, so as to obtain speech training data of corpus quality.
[0068] Example 2
[0069] This embodiment provides a device for acquiring speech training data, including:
[0070] A massive audio data acquisition module, used to acquire massive audio data;
[0071] A first splitting module, used for splitting the multi-channel audio data in the audio data into single-channel audio data;
[0072] A denoising module, used to remove background music and background noise from all the single-channel audio data to obtain denoised single-channel audio data;
[0073] A second splitting module is used to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments according to the speaker log and using voice activity detection technology, wherein the speaker log includes the speaking time and speaking order of each speaker;
[0074] A screening module, used to screen out the audio segment of the single speaker that meets the preset time length;
[0075] A recognition module, used to perform language recognition and voice recognition on the single speaker audio segment that meets the preset time length in sequence to obtain a recognized audio segment text;
[0076] A punctuation adding module, used for adding punctuation to the recognized audio segment text to obtain a punctuation added text;
[0077] A quality assessment module, used to perform quality assessment on the audio segment corresponding to the punctuation-added text using a synthetic speech automatic assessment quality evaluation model;
[0078] A sound quality enhancement module, configured to perform sound quality enhancement processing on the audio segment whose quality evaluation is less than the quality threshold to obtain an enhanced audio segment, wherein the sound quality enhancement processing includes replacing timbre, restoring audio distortion and expanding audio bandwidth;
[0079] The speech training data acquisition module is used to use the enhanced audio segment as speech training data.
[0080] Example 3
[0081] This embodiment provides a computer device, which may be a server or a terminal. Its internal structure diagram may be as follows: Figure 2 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store voice training data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the method for obtaining voice training data provided in Example 1 is implemented.
[0082] Those skilled in the art will understand that Figure 2 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0083] Example 4
[0084] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method for acquiring speech training data provided in the above embodiment 1 is implemented.
[0085] Example 5
[0086] This embodiment provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for acquiring speech training data provided in the above embodiment 1 is implemented.
[0087] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0088] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0089] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0090] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0091] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for acquiring speech training data, characterized in that: The method for acquiring the speech training data comprises: Get massive audio data; Splitting the multi-channel audio data in the audio data into single-channel audio data; Removing background music and background noise from all the single-channel audio data to obtain denoised single-channel audio data; According to the speaker log, voice activity detection technology is used to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments, wherein the speaker log includes the speaking time and speaking order of each speaker; Filter out the single speaker audio segment that meets the preset duration; Performing language recognition and speech recognition on the single speaker audio segment that meets the preset time length in sequence to obtain a recognized audio segment text; Adding punctuation to the recognized audio segment text to obtain punctuation-added text; Using a synthetic speech automatic assessment quality evaluation model to perform quality assessment on the audio segment corresponding to the punctuation-added text; Performing sound quality enhancement processing on the audio segment whose quality evaluation is less than the quality threshold to obtain an enhanced audio segment, wherein the sound quality enhancement processing includes replacing timbre, restoring audio distortion, and expanding audio bandwidth; The enhanced audio segment is used as speech training data.
2. The method for acquiring speech training data according to claim 1, characterized in that: The process of removing background noise from all the single-channel audio data specifically includes: Using a noise evaluation model to evaluate whether the single-channel audio data needs noise reduction, and obtaining an evaluation result; When the evaluation result is yes, the background noise in the single-channel audio data is removed using a speech noise reduction model.
3. The method for acquiring speech training data according to claim 1, characterized in that: After executing the step of "performing language recognition and speech recognition on the single speaker audio segment that meets the preset time length in sequence to obtain the recognized audio segment text", the method for obtaining speech training data further includes: Performing text regularization processing on the audio segment of the existing transcribed text to obtain the regularized transcribed text, wherein the audio segment of the existing transcribed text is an audio segment of the single speaker audio segment that meets the preset time length; Calculating the edit distance between the pinyin of the regularized transcribed text and the pinyin of the corresponding recognized audio segment text; evaluating a first recognition quality according to the edit distance; The recognized audio segment texts whose first evaluation quality is less than a first quality threshold are filtered to obtain a first recognized audio segment text.
4. The method for acquiring speech training data according to claim 3, characterized in that: After executing the step of "performing language recognition and speech recognition on the single speaker audio segment that meets the preset time length in sequence to obtain the recognized audio segment text", the method for obtaining speech training data further includes: Calculating an ASR confidence between an audio segment without a transcribed text and a corresponding recognized audio segment text, wherein the audio segment without a transcribed text is an audio segment in the single speaker audio segment that meets a preset time length; evaluating a second recognition quality according to the ASR confidence; The recognized audio segment texts whose second evaluation quality is less than a second quality threshold are filtered to obtain a second recognized audio segment text.
5. The method for acquiring speech training data according to claim 4, characterized in that: After performing the step of "filtering the recognized audio segment texts whose first evaluation quality is less than the first quality threshold to obtain the first recognized audio segment text" and the step of "filtering the recognized audio segment texts whose second evaluation quality is less than the second quality threshold to obtain the second recognized audio segment text", the method for acquiring speech training data further includes: Calculating an average pronunciation duration of each word in a target audio segment, wherein the target audio segment is the first recognized audio segment text and the second recognized audio segment text; The target audio segment is filtered according to the average pronunciation duration of each word and the duration threshold.
6. The method for acquiring speech training data according to claim 1, characterized in that: Adding punctuation to the recognized audio segment text to obtain punctuation-added text specifically includes: Acquiring ASR acoustic information of an audio segment corresponding to the recognized audio segment text, wherein the ASR acoustic information is a timestamp interval between adjacent words in the corresponding audio segment; Based on the ASR acoustic information, the punctuation position is adjusted through a text punctuation model, wherein the punctuation position includes an insertion position and a deletion position.
7. The method for acquiring speech training data according to claim 1, characterized in that: The quality of the audio segment corresponding to the punctuation-added text is evaluated, specifically including: A synthetic speech automatic assessment quality evaluation model is used to perform quality assessment on the audio segment corresponding to the punctuation-added text.
8. A device for acquiring speech training data, characterized in that: The device for acquiring speech training data comprises: A massive audio data acquisition module, used to acquire massive audio data; A first splitting module, used for splitting the multi-channel audio data in the audio data into single-channel audio data; A denoising module, used to remove background music and background noise from all the single-channel audio data to obtain denoised single-channel audio data; A second splitting module is used to split the multi-person conversation audio in the denoised single-channel audio data into single-speaker audio segments according to the speaker log and using voice activity detection technology, wherein the speaker log includes the speaking time and speaking order of each speaker; A screening module, used to screen out the audio segment of the single speaker that meets a preset time length; A recognition module, used to perform language recognition and voice recognition on the single speaker audio segment that meets the preset time length in sequence to obtain a recognized audio segment text; A punctuation adding module, used for adding punctuation to the recognized audio segment text to obtain a punctuation added text; A quality assessment module, used to perform quality assessment on the audio segment corresponding to the punctuation-added text using a synthetic speech automatic assessment quality evaluation model; A sound quality enhancement module, configured to perform sound quality enhancement processing on the audio segment whose quality evaluation is less than the quality threshold to obtain an enhanced audio segment, wherein the sound quality enhancement processing includes replacing timbre, restoring audio distortion and expanding audio bandwidth; The speech training data acquisition module is used to use the enhanced audio segment as speech training data.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for acquiring speech training data according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for acquiring speech training data according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Audio quality evaluation method and device, electronic equipment and storage medium
CN113782036A
Voice library training data analysis method and device
CN113889096A
Training data screening method, system and device, and medium
CN113901992A
Sound synthesis training data acquisition method and device, server and storage medium
CN116469367A
Application programming interface for indicating number of wireless cells
CN116709372A
Cited By
Prompt tone acquisition method for voice cloning
CN121260177A