Audio data processing method, model training method, device, equipment and product
By performing multiple rounds of processing on audio data and progressively improving the screening criteria, the problem of efficiently extracting high-quality text-audio pairs from massive audio and video resources was solved, thereby improving data processing efficiency and the training effect of speech synthesis models.
Patent Information
- Application Number
- CN202511051818.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies for extracting high-quality text-audio pairs from massive audio and video resources suffer from problems such as low data processing efficiency, insufficient accuracy of automatic transcription in multiple languages, limited long audio segmentation strategies, and resource waste due to lag in the quality screening process.
By processing audio data in multiple rounds, each round including parsing and filtering, and gradually increasing the filtering criteria, low-quality or abnormal data is gradually eliminated, thus achieving a data processing method of "processing and filtering simultaneously".
It improves the efficiency of audio data processing, reduces the data load in subsequent processing stages, ensures the quality and consistency of training data, and enhances the training effect of speech synthesis models.
Smart Images

Figure CN120808789A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, and in particular to an audio data processing method and device, a model training method and device, and a product. BACKGROUND
[0002] Training of a speech synthesis model requires high-quality text-audio pair data. In order to extract high-quality text-audio pairs from a large amount of audio-video resources, a series of data processing procedures such as audio purification and speech recognition to convert the original audio-video data into text are required. Subsequently, the preliminary processed data need to be further screened to ensure the quality of the text-audio pairs.
[0003] However, this screening method often requires a large amount of time to process data that is ultimately proven to be ineffective, thus reducing the overall data processing efficiency. SUMMARY
[0004] Based on the above technical status, the present application provides an audio data processing method, a model training method, a device, an apparatus and a product, which can improve the audio data processing efficiency.
[0005] To achieve the above technical purpose, the present application specifically proposes the following technical solutions:
[0006] According to a first aspect of an embodiment of the present application, an audio data processing method is provided, comprising: obtaining audio data to be processed; sequentially performing N rounds of audio data processing on the audio data to be processed to obtain target audio data; the target audio data is used to train a speech synthesis model, and N is a positive integer; wherein each round of audio data processing in the N rounds of audio data processing includes analyzing the audio data and screening the analyzed audio data; and the screening standards of the screening mechanisms corresponding to the N rounds of audio data processing gradually increase.
[0007] In some implementations, the audio data to be processed includes audio segments; and the screening mechanism corresponding to the N rounds of audio data processing includes at least one of a screening mechanism based on audio duration, a screening mechanism based on audio quality, a screening mechanism based on the number of speaker transitions, a screening mechanism based on comparison between multiple text transcription results, and a screening mechanism based on the audio duration of the word segmentation in each audio segment.
[0008] In some implementations, the N rounds of audio data processing are sequentially performed on the audio data to be processed to obtain target audio data, including: in a first round of audio data processing, the audio data to be processed is segmented into multiple audio segments using different segmentation granularities to obtain an audio set after the first round of audio data processing; and audio segments with a duration less than a first preset duration and / or audio segments with a duration greater than a second preset duration are deleted from the audio set after the first round of audio data processing.
[0009] In some implementations, the audio data to be processed is segmented into multiple audio segments using different segmentation granularities to obtain an audio set after the first round of audio data processing, including: the audio data to be processed is segmented into audio segments using a first segmentation granularity to obtain an initial audio set, the initial audio set including each audio segment; if a proportion of target audio segments in the initial audio set is greater than or equal to a preset proportion, the target audio segments are segmented using a second segmentation granularity, and the initial audio set is updated until the proportion of the target audio segments in the updated initial audio set is less than the preset proportion, to obtain an audio set after the first round of audio data processing, the first segmentation granularity being greater than the second segmentation granularity; and the target audio segments are audio segments with a duration greater than or equal to the second preset duration.
[0010] In some implementations, the N rounds of audio data processing are sequentially performed on the audio data to be processed to obtain target audio data, further including: in a second round of audio data processing, a quality index value corresponding to each audio segment in the audio set after the first round of audio data processing is determined to obtain an audio set after the second round of audio data processing; and audio segments with a quality index value that does not meet a preset quality requirement are deleted from the audio set after the second round of audio data processing.
[0011] In some implementations, the N rounds of audio data processing are sequentially performed on the audio data to be processed to obtain target audio data, further including: in a third round of audio data processing, speaker detection is performed on each audio segment in the audio set after the second round of audio data processing to obtain an audio set after the third round of audio data processing, the audio set after the second round of audio data processing including each audio segment and a corresponding number of speaker changes and a corresponding probability of speaker change; and audio segments with a number of speaker changes that does not meet a preset number requirement, or with a number of speaker changes that does not meet the preset number requirement and a corresponding probability of speaker change that does not meet a preset probability threshold, are deleted from the audio set after the third round of audio data processing.
[0012] In some implementations, the N rounds of audio data processing are sequentially performed on the audio data to be processed to obtain target audio data, and further comprising: in the fourth round of audio data processing, using at least two speech recognition models to perform speech recognition on the audio set after the third round of audio data processing respectively to obtain at least two sets of transcribed texts; based on comparison between the at least two sets of transcribed texts, obtaining the audio set after the fourth round of audio data processing and the corresponding set of transcribed texts.
[0013] In some implementations, the using at least two speech recognition models to perform speech recognition on the audio set after the third round of audio data processing respectively to obtain at least two sets of transcribed texts comprises: using a first speech recognition model and a second speech recognition model to perform speech recognition on the audio set after the third round of audio data processing respectively to obtain a first set of transcribed texts and a second set of transcribed texts; and based on comparison between the at least two sets of transcribed texts, obtaining the audio set after the fourth round of audio data processing and the corresponding set of transcribed texts comprises: performing a first preprocessing operation on the first set of transcribed texts and the second set of transcribed texts respectively to obtain a preprocessed first set of transcribed texts and a preprocessed second set of transcribed texts; and in a case where a comparison result between the preprocessed first set of transcribed texts and the preprocessed second set of transcribed texts does not meet a preset requirement, performing a second preprocessing operation on the preprocessed first set of transcribed texts and the preprocessed second set of transcribed texts respectively until a comparison result between a new preprocessed first set of transcribed texts and a new preprocessed second set of transcribed texts meets the preset requirement, to obtain the audio set after the fourth round of audio data processing and the corresponding set of transcribed texts.
[0014] In some implementations, the N rounds of audio data processing are sequentially performed on the audio data to be processed to obtain target audio data, and further comprising: in the fifth round of audio data processing, performing alignment processing on the audio set after the fourth round of audio data processing and the corresponding set of transcribed texts to obtain an alignment result between the audio set after the fourth round of audio data processing and the corresponding set of transcribed texts, the alignment result comprising each word piece in each audio segment in the audio set after the fourth round of audio data processing and a corresponding audio duration; and for each audio segment in the audio set after the fourth round of audio data processing, if a number of target word pieces in each word piece in the audio segment is greater than a preset number of word pieces, or an average value of audio durations of each word piece in the audio segment is greater than a first preset average value or less than a second preset average value, the audio segment is deleted; and the target word piece is a word piece with an audio duration greater than a preset duration.
[0015] According to a second aspect of the embodiments of the present application, a method for training a speech synthesis model is provided, including obtaining audio sample data, the audio sample data being obtained by using any one of the methods of the first aspect; and training a speech synthesis model based on the audio sample data.
[0016] According to a third aspect of the embodiments of the present application, an audio data processing apparatus is provided, including: an obtaining unit configured to obtain audio data to be processed; and a processing unit configured to sequentially perform N rounds of audio data processing on the audio data to be processed to obtain target audio data, the target audio data being used to train a speech synthesis model, N being a positive integer; wherein each round of audio data processing in the N rounds of audio data processing includes audio data analysis and filtering of the analyzed audio data; and the filtering criteria of the filtering mechanisms corresponding to the N rounds of audio data processing gradually increase.
[0017] According to a fourth aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor; the memory is connected to the processor and is configured to store programs; and the processor is configured to implement the audio data processing method of the first aspect or the training method of the speech synthesis model of the second aspect by running the programs in the memory.
[0018] According to a fifth aspect of the embodiments of the present application, a computer program product is provided, including computer program instructions, which, when executed by a processor, cause the processor to perform the audio data processing method of the first aspect or the training method of the speech synthesis model of the second aspect.
[0019] The audio data processing method, model training method, apparatus, device and product provided by the embodiments of the present application gradually filter out target audio data used to train a speech synthesis model by sequentially performing N rounds of audio data processing on the obtained audio data to be processed, wherein each round of audio data processing includes audio data analysis, and after each round of analysis, the analyzed audio data is filtered, and the filtering criteria of the filtering mechanisms corresponding to the N rounds of audio data processing gradually increase. By introducing a data filtering mechanism after each round of analysis in the N rounds of audio data processing, a data processing mode of "processing and filtering simultaneously" can be realized, low-quality or abnormal data can be effectively eliminated in the early stage of data processing, the data load of subsequent processing links is reduced, and the overall data processing efficiency is gradually improved. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the accompanying drawings in the following description only only the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.
[0021] Figure 1 A flow chart of an audio data processing method provided by an embodiment of the present application is shown in FIG. 4.
[0022] Figure 2 A flow chart of first round audio data processing in N rounds of audio data processing provided by an embodiment of the present application is shown in FIG. 5.
[0023] Figure 3 A flow chart of second round audio data processing in N rounds of audio data processing provided by an embodiment of the present application is shown in FIG. 6.
[0024] Figure 4 A flow chart of third round audio data processing in N rounds of audio data processing provided by an embodiment of the present application is shown in FIG. 7.
[0025] Figure 5 A flow chart of fourth round audio data processing in N rounds of audio data processing provided by an embodiment of the present application is shown in FIG. 8.
[0026] Figure 6 A flow chart of fifth round audio data processing in N rounds of audio data processing provided by an embodiment of the present application is shown in FIG. 9.
[0027] Figure 7 A structural schematic diagram of an audio data processing device provided by an embodiment of the present application is shown in FIG. 10.
[0028] Figure 8 A structural schematic diagram of a speech synthesis model training device provided by an embodiment of the present application is shown in FIG. 11.
[0029] Figure 9 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 12. DETAILED DESCRIPTION
[0030] The technical solutions provided by the embodiments of the present application are suitable for cross-language real-time communication systems, bilingual subtitle automatic generation of multimedia content, localization processing of commodity information on cross-border e-commerce platforms, and large-scale online translation services and other scenarios.
[0031] The technical solutions provided by the embodiments of the present application can be exemplarily applied to hardware devices such as processors, electronic devices, servers (including cloud servers), or packaged into software programs to be run. When the hardware devices execute the processing process of the technical solutions of the embodiments of the present application, or the above software programs are run, the automatic splitting of the target task and the automatic calling of the application program interface required by the task can be realized, and the purpose of completing the target task can be achieved. The embodiments of the present application only exemplarily introduce the specific processing process of the technical solutions of the present application, and do not limit the specific implementation form of the technical solutions of the present application. Any technical implementation form that can execute the processing process of the technical solutions of the present application can be adopted by the embodiments of the present application.
[0032] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0033] Before introducing the solutions of the present application, the related art is first introduced:
[0034] The traditional multi-lingual speech synthesis system has a small demand for the number of training data, but has a very high requirement for the quality. In order to obtain high-quality text-audio pairs that meet the requirements, a professional speaker of the target language usually needs to record the audio piece by piece in a professional recording environment. This way not only has high cost and low efficiency, but also the final obtained data set is limited in size, usually only tens to hundreds of hours. In addition, each piece of audio needs to be carefully annotated by a professional person with corresponding language background to ensure the high consistency of the text and the audio. The overall process needs to rely heavily on manual intervention, which is time-consuming and laborious, especially for some small languages with scarce resources, the lack of professional speakers and language experts further increases the difficulty of data acquisition.
[0035] With the development of large language models (LLM), the speech synthesis task has gradually evolved from the traditional speech synthesis system to the new generation system based on large model architecture. Under this background, the demand for training data of multi-lingual speech synthesis system has changed, i.e. from relying on a small amount of high-quality data to requiring a large amount of data (from tens of hours to tens of thousands or even hundreds of thousands of hours), although the quality requirement of a single sample is relaxed, but the overall data quality still needs to meet the basic standard of model training.
[0036] Under this new demand, the traditional piece-by-piece recording and manual annotation method has been difficult to support large-scale data collection. Although the audio and video data obtained from public channels is large in quantity, its quality is uneven, and often lacks corresponding text content, which is a typical unsupervised audio data.
[0037] Among the currently disclosed technical solutions, there are two types of systems that can be used to process massive unsupervised audio-video data:
[0038] The first type is mainly aimed at the speech recognition field and is used to train a large speech recognition model. Since the speech recognition model has relatively low requirements for data quality, the processing flow of this type of system is relatively simple. Representative technologies such as Gigaspeech1&2 have a typical flow that includes: raw data collection → preprocessing (unified to 16kHz single-channel WAV format) → automatic transcription (using Whisper large-v3) → forced alignment (obtaining word-level timestamps through an alignment model) → text regularization → quality screening (character set, language confidence, audio duration, deduplication, etc.) → output of available text-audio pairs.
[0039] The second type is more focused on the speech synthesis scenario, and the processed data is used for training a large speech synthesis model. This type of system is more complex than the first type, and one of the representative technologies is Emilia, which has a flow that includes: raw data collection → format and audio parameter standardization (sampling rate, bit depth, volume, waveform energy normalization) → source separation (using UVR tools to remove background noise) → speaker segmentation (detecting speaker switching points and clustering) → VAD-based audio segmentation (retaining sub-audios of 3-30 seconds in length) → automatic transcription (using WhisperX) → quality screening (DNSMOS P.835 OVRL score and audio duration filtering) → output of available text-audio pairs.
[0040] According to the processing flow of the above two types of systems, the following defects still exist:
[0041] 1. Inadequate accuracy of multi-language automatic transcription
[0042] Currently, the text transcription process generally relies on a single speech recognition model, which has obvious limitations in multi-language, especially in some small language scenarios. Due to the insufficient coverage and recognition accuracy of the model for related languages, the accuracy of the transcription result is often low, making it difficult to ensure the consistency of the text and audio. If such low-quality data is used for training a speech synthesis model, the model may produce inaccurate pronunciation, word confusion, and other problems in subsequent applications, thereby affecting the overall quality and naturalness of the synthesized speech.
[0043] 2. Single long audio segmentation strategy, affecting data quality and utilization
[0044] Generally, the method of relying only on a single voice activity detection (VAD) model to segment long audio tends to be too mechanical. This processing method can lead to two extreme cases: on the one hand, the audio may be segmented too fragmented, thereby destroying the original semantic integrity; on the other hand, if the granularity of segmentation is too rough, it may lead to an average sentence length that is too long, which is also not conducive to the training of a speech synthesis model. Even if a strict time limit (such as the 3 to 30 second audio retention standard in the Emilia system) is set to filter the data, it may result in a large amount of actually eligible data being unnecessarily discarded due to overly stringent screening conditions.
[0045] 3. The quality screening link lags, resulting in waste of resources in invalid data processing
[0046] The quality screening module is usually located at the end of the entire processing flow, which means that all raw audio data needs to go through complete processing of audio processing, text transcription, speaker analysis, VAD segmentation, and other links before entering the final quality screening stage. However, in practice, a considerable portion of audio contains invalid content, such as silent passages, non-human voice interference, or extremely poor quality segments. Using the "process first, then screen" process not only results in a large amount of invalid data being repeatedly loaded and calculated in the entire processing chain, but also significantly increases the consumption of computing resources and processing time, reducing the overall processing efficiency.
[0047] Therefore, the embodiments of the present application are dedicated to providing an audio data processing method, a model training method, an apparatus, a device and a product. By sequentially performing N rounds of processing operations on the audio data to be processed, target audio data that can be used for training of a speech synthesis model is generated. In the process of processing the audio data, an analysis operation is performed on the audio data, and after each round of analysis operation, a screening mechanism for the analyzed audio data is introduced to improve the overall processing efficiency of the audio data. In the following embodiments, each is described in detail.
[0048] Exemplary method
[0049] Figure 1 A flowchart of an audio data processing method provided by an embodiment of the present application. As shown in Figure 1 the audio data processing method provided by the embodiment of the present application includes steps S101-S102:
[0050] S101, obtaining audio data to be processed.
[0051] The audio data to be processed can be obtained from original audio resources or original audio and video resources containing speech in a public channel, and a series of preprocessing operations are performed thereon. The preprocessing operations include format conversion, uniform naming, uniform storage method, and voice separation operation, etc.
[0052] Specifically, the original audio resources or original audio and video resources can be uniformly formatted by using an open source format conversion tool ffmpeg: converted into single-channel, 16khz sampling rate, 16bit bit depth, WAV format standard audio data. At the same time, all audio data corresponding to the files need to be named according to uniform rules, the format is: source (English) - batch (English) - file name (number).wav, to ensure that the file name only contains English letters and numbers, to improve the compatibility and stability of subsequent module processing. In addition, the duration of the audio file also needs to be counted and stored in batches in units of 100h to facilitate the parallel processing and management of subsequent large-scale tasks. Finally, the open source voice separation algorithm, such as demus, is used to separate the human voice and background sound in the audio and retain the human voice part to improve the audio quality.
[0053] S102, sequentially performing N rounds of audio data processing on the audio data to be processed to obtain target audio data.
[0054] The target audio data in this embodiment is used to train a speech synthesis model, which can be a multi-lingual speech synthesis model. In the case of a multi-lingual speech synthesis model, the audio data to be processed includes audio data corresponding to each language, and the audio data to be processed is also processed when the audio data to be processed is processed.
[0055] In some embodiments, the audio data processing includes parsing the audio data, and at least one of the first N-1 rounds of audio data processing includes filtering the parsed audio data.
[0056] In this embodiment, N represents the number of multiple rounds of screening set in the entire processing flow, and N is a positive integer.
[0057] For each round of audio data processing in the N rounds of audio data processing, the audio data to be processed first needs to be parsed, and the parsing operation includes at least two or more combinations: audio segment cutting, audio quality evaluation, speaker detection, audio transcription text, and audio and text alignment processing. Through these parsing steps, the basic features and semantic information of the audio can be extracted to provide a basis for subsequent data filtering.
[0058] In the first N-1 rounds of audio data processing, at least one round of data screening operation based on the analysis result is included. Specifically, each round of audio data processing in the first N-1 rounds includes an analysis operation on the audio data for extracting its key features and attributes; but it is not required to perform a data screening operation after each round of analysis. As long as at least one screening operation is performed according to the analysis result in the entire first N-1 rounds of processing, the overall data processing efficiency and quality can be improved.
[0059] In addition, it should be noted that in the Nth round (i.e. the last round) of audio data processing, a final data screening operation is also required after the analysis operation. This screening step aims to ensure that the output target audio data meets the standard for training the speech synthesis model, and to eliminate the low-quality samples that may have been missed in the early processing, thereby further ensuring the accuracy and consistency of the training data.
[0060] The embodiment gradually screens the target audio data for training the speech synthesis model by sequentially performing N rounds of audio data processing on the obtained audio data to be processed, wherein each round of audio data processing includes an analysis operation on the audio data, and at least one round of performing a screening operation on the analyzed audio data is included in the first N-1 rounds of audio data processing. By introducing a data screening mechanism in the first N-1 rounds of audio data processing, a data processing mode of "processing and screening simultaneously" can be realized, which can effectively eliminate low-quality or abnormal data in the early stage of data processing, reduce the data load of subsequent processing steps, and gradually improve the overall data processing efficiency.
[0061] In the N rounds of audio data processing, each round of audio data processing includes an analysis operation on the audio data and a screening operation on the analyzed audio data; and the screening criteria of the screening mechanism corresponding to the N rounds of audio data processing gradually increase.
[0062] In the embodiment, each round of audio data processing in the N rounds of audio data processing includes an analysis operation on the audio data, and a data screening operation is performed immediately after the analysis operation is completed, so that the analysis and screening are alternately performed, and a total of N cycles are completed. Thus, it is ensured that the data is screened based on the latest analysis result at each processing stage, so as to reduce the data amount of subsequent analysis operations, and thus reduce the overall data processing efficiency.
[0063] It should be noted that after the Nth round, i.e. the last round of analysis operation is completed, a screening operation is also required to be performed, so as to ensure that all output data has undergone a complete quality control process, thereby improving the consistency and usability of the final output data.
[0064] The screening criteria of the screening mechanism of the N rounds of audio data processing is "gradually increased", which means that the standard used in each round of screening will be more stringent or more refined than the previous round, thereby constantly improving data quality while ensuring processing efficiency.
[0065] "Gradually increased screening criteria" means that in the N rounds of audio data processing, the quality judgment standard for each round of data screening operation is gradually improved. In this way, it is possible to dynamically adjust the screening strategy according to the data obtained by the current analysis at different processing stages, so as to retain as much potential available data as possible in the early stage, and to screen out the optimal audio data through a more refined evaluation mechanism in the later stage.
[0066] In some embodiments, the screening mechanism corresponding to the N rounds of audio data processing includes at least one of a screening mechanism based on audio duration, a screening mechanism based on audio quality, a screening mechanism based on the number of speaker transitions, a screening mechanism based on comparison between multiple text transcription results, and a screening mechanism based on the audio duration of the word segmentation in each audio segment.
[0067] In some examples, the N rounds of audio data processing can include first round of audio data processing, second round of audio data processing, third round of audio data processing, fourth round of audio data processing and fifth round of audio data processing. The screening mechanism corresponding to the N rounds of audio data processing includes, in turn: a screening mechanism based on audio duration, a screening mechanism based on audio quality, a screening mechanism based on the number of speaker transitions, a screening mechanism based on comparison between multiple text transcription results, and a screening mechanism based on the audio duration of the word segmentation in each audio segment
[0068] That is, the analysis in the first round of audio data processing includes audio segment segmentation, and the data screening mechanism includes a screening mechanism based on audio duration.
[0069] The analysis in the second round of audio data processing includes audio quality assessment, and the data screening mechanism includes a screening mechanism based on audio quality.
[0070] The analysis in the third round of audio data includes speaker detection, and the data screening mechanism includes a screening mechanism based on the number of speaker transitions.
[0071] The analysis in the fourth round of audio data processing includes audio transcription text, and the data screening mechanism includes a screening mechanism based on comparison between multiple text transcription results.
[0072] The analysis in the fifth round of audio data processing includes audio and text alignment processing, and the data screening mechanism includes a screening mechanism based on the audio duration of the word segmentation in each audio segment.
[0073] In the above five-round screening mechanism, the first-round screening mechanism performs preliminary screening based on the basic physical properties of the audio data to eliminate too short or too long speech segments, avoid invalid input during model training, and is relatively lenient. The second-round screening mechanism introduces subjective auditory quality and objective acoustic feature evaluation to remove segments with high background noise and unclear speech, thereby improving the robustness of the recognition model and being more stringent than the first round. The third-round screening mechanism focuses on the structural and semantic clarity of the speech content, eliminating segments with multiple conversations or frequent speaker switching to ensure the structural clarity of the training data and further improve the screening standard. The fourth round introduces a semantic level judgment standard, which excludes segments with inconsistent recognition results and semantic confusion through multi-model transcription comparison, and the screening standard is closer to the language understanding level. The fifth round screens from the perspective of maintaining high consistency of the output data in both speech and language dimensions to ensure accurate alignment of each word or phoneme with the audio and improve the accuracy and generalization ability of the model training. The specific implementation process of each round of audio data processing will be introduced in conjunction with the accompanying drawings as follows:
[0074] Figure 2 The flowchart of the first round of audio data processing in the N-round audio data processing provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the first round of audio data processing is performed on the audio data to be processed, including the following steps S201-S202: Figure 2
[0075] S201, in the first round of audio data processing, different segmentation granularities are used to perform multiple audio segmentations on the audio data to be processed, thereby obtaining an audio set after the first round of audio data processing.
[0076] Since the audio data to be processed has a large difference in length, some long audios can even reach several hours, and the speech synthesis model is usually more suitable for audio segments of 1 to 30 seconds. Therefore, in this embodiment, a VAD model is used to detect the speech segments of each piece of audio data to be processed, thereby obtaining a series of initial audio segments, and the initial audio segments are accurately segmented according to their starting and ending time points.
[0077] The VAD model has a certain granularity control ability when detecting speech segments, which can usually be realized by adjusting the detection probability threshold, setting the minimum speech or silence duration, etc. If the granularity is set too fine, the segmented sub-audio will be too short as a whole, which may lead to incomplete semantics and affect the language modeling effect of the subsequent speech synthesis model. Conversely, if the granularity is too coarse, the sub-audio will generally be longer, and a large number of segments exceeding 30 seconds will be generated, which will be eliminated in the subsequent first-round screening process, causing waste of data resources.
[0078] To balance the segmentation quality and data utilization, step S201 aims to adopt a multi-level VAD segmentation strategy, which can effectively control the length distribution of sub-audio while ensuring semantic coherence by performing segmentation operations on different granularity levels, thereby improving the quality and utilization efficiency of the final training data.
[0079] Specifically, in some embodiments, step S201 includes: performing audio segment segmentation on the audio data to be processed using a first segmentation granularity to obtain an initial audio set, the initial audio set including various audio segments; if the proportion of target audio segments in the initial audio set is greater than or equal to a preset proportion, performing segmentation on the target audio segments using a second segmentation granularity, and updating the initial audio set until the proportion of target audio segments in the updated initial audio set is less than the preset proportion, obtaining an audio set after the first round of audio data processing, the first segmentation granularity being greater than the second segmentation granularity; wherein the target audio segment is an audio segment with a duration greater than or equal to the second preset duration.
[0080] In this embodiment, the first segmentation granularity can be understood as a coarse-grained segmentation parameter configuration, including key parameters such as the first detection probability threshold, the first speech duration, and the first silence duration. The second segmentation granularity can be understood as a more refined segmentation strategy, and its parameter configuration includes key parameters such as the second detection probability threshold, the second speech duration, and the second silence duration.
[0081] Specifically, the first detection probability threshold is less than the second detection probability threshold, indicating that the identification of speech segments is more relaxed during preliminary segmentation; at the same time, the first speech duration is greater than the second speech duration, and the first silence duration is greater than the second silence duration, indicating that longer speech passages and silence intervals are allowed in the first round of segmentation, thereby forming a more extensive segmentation result.
[0082] After preliminary segmentation of the audio data using the first segmentation granularity, the initial audio set obtained generally has longer durations for each audio segment. Therefore, after completing the preliminary segmentation, it is necessary to detect the duration of each audio segment in the initial audio set to assess whether the current segmentation result meets the needs of subsequent modeling.
[0083] If the statistical results show that the proportion of longer audio segments is high, further segmentation of audio segments with a duration exceeding the second preset duration using a more refined second segmentation granularity is required, and the number of audio segments with a duration meeting the preset duration range requirement is required to meet the preset requirement.
[0084] For example, first, the first segmentation granularity is used to segment the audio data to be processed to obtain an initial audio set A, at this time, most of the audio segments in the initial audio set A have a long duration.
[0085] For the audio clips longer than 30 seconds (i.e., the target audio clips) in the initial audio set A, they are split twice using the second segmentation granularity to obtain a new audio set B. Then, a statistical analysis is performed on set B. If the total audio duration corresponding to the target audio clip accounts for less than 10% of the total audio duration of the entire set, the analysis ends; otherwise, the audio clips longer than 30 seconds need to be further segmented until the ratio requirement is met.
[0086] Through this multi-stage, coarse-to-fine progressive segmentation mechanism, the relationship between semantic integrity and segmentation accuracy can be effectively balanced, thereby improving the quality of the final audio dataset.
[0087] Continue reading Figure 2 In the case where the screening mechanism of the first round of audio data processing includes a screening mechanism based on audio duration, the first round of audio data processing further includes the following step S202:
[0088] S202: Delete, from the audio set after the first round of audio data processing, audio segments whose audio duration is less than a first preset duration and / or audio segments whose audio duration is greater than a second preset duration.
[0089] Specifically, the first preset duration can be 1 second, and the second preset duration can be 30 seconds. By deleting audio segments shorter than 1 second or longer than 30 seconds, it is possible to achieve an average retention rate of approximately 88% of the original audio data to be processed in each language after the first round of audio data processing.
[0090] Figure 3 This is a flowchart of the second round of audio data processing in the N rounds of audio data processing provided in the embodiment of this application. Figure 3 As shown, the second round of audio data processing is performed on the audio data to be processed, including the following steps S301-S302:
[0091] S301. During the second round of audio data processing, determine the quality index value corresponding to each audio segment in the audio set after the first round of audio data processing, and obtain the audio set after the second round of audio data processing.
[0092] In this step, the quality of each audio clip in the audio collection after the first round of audio data processing can be evaluated using two metrics: MOS (Mean Opinion Score) and Signal-to-Noise Ratio (SNR). These two metrics reflect the overall quality of the audio clip from the perspectives of subjective perception and background noise level, respectively.
[0093] MOS score is used to measure the overall perceptual quality of speech, usually ranging from 1 to 5, with higher values indicating better audio quality. In practical applications, an open-source algorithm NISQA (Non-Intrusive Speech Quality Assessment) can be used to automatically score audio segments. NISQA is a deep learning-based non-intrusive speech quality assessment model that can perform end-to-end quality prediction on input audio without reference to the original clean speech, with good generalization ability and evaluation accuracy.
[0094] On the other hand, the signal-to-noise ratio (SNR) is used to evaluate the proportion between the speech signal and the background noise in the audio, which is an important physical indicator of audio clarity. A higher signal-to-noise ratio means that the effective speech component in the audio has a higher proportion, and the background noise interference is smaller, which is beneficial to improve the training effect of the speech synthesis model and the clarity of the output speech.
[0095] By combining the quality evaluation results of MOS score and signal-to-noise ratio in two dimensions, the audio segments after the first round of processing can be finely selected, and samples with low quality, high noise, or unclear pronunciation can be removed, thereby further improving the overall quality and consistency of the audio dataset, and providing a more high-quality data basis for the subsequent training of the speech synthesis model.
[0096] Continuing to refer to Figure 3 In the case where the second-round audio data processing screening mechanism includes an audio quality-based screening mechanism, the second-round audio data processing process further includes the following step S302:
[0097] S302, deleting an audio segment in the second-round audio data processing audio set whose quality index value does not meet the preset quality requirement.
[0098] The preset quality requirement includes: the MOS score is greater than or equal to a first preset MOS score threshold; or, the MOS score is greater than or equal to a second preset MOS score threshold and less than the first MOS score threshold, and the signal-to-noise ratio is greater than or equal to a preset signal-to-noise ratio threshold.
[0099] In some examples, the first preset MOS score threshold can be 4.0, the second preset MOS score threshold can be 3.0, and the signal-to-noise ratio can be 17.
[0100] If the MOS score of an audio segment in the second-round audio data processing audio set is greater than or equal to 4.0, the audio segment is retained; or, if the MOS score of an audio segment in the second-round audio data processing audio set is greater than or equal to 3.0 and less than 4.0, and the signal-to-noise ratio is greater than or equal to 17, the audio segment is retained.
[0101] After the second round of audio data processing, the average retention rate of each language is about 57% of the original audio data to be processed.
[0102] Figure 4 The flowchart of the third round of audio data processing in the N rounds of audio data processing provided by the embodiment of the present application is shown in FIG. 4. As shown in FIG. 4, the third round of audio data processing is sequentially performed on the audio data to be processed, including the following steps S401-S402: Figure 4
[0103] S401, in the third round of audio data processing, speaker detection is performed on each audio segment in the audio set after the second round of audio data processing, to obtain an audio set after the third round of audio data processing.
[0104] Among them, the audio set after the second round of audio data processing includes each audio segment and its corresponding speaker change number and speaker change probability.
[0105] The audio data in this embodiment is mainly used for training of a speech synthesis model. Based on the requirement of the speech synthesis model that each audio segment contains a single speaker, a speaker endpoint detection model is used in this embodiment to perform speaker detection on each audio segment in the audio set after the second round of audio data processing, to obtain the speaker change number and the speaker change probability of each audio segment in the audio set after the second round of audio data processing.
[0106] The speaker endpoint detection model can be obtained by training a deep learning model using audio data of multiple languages. By inputting each audio segment in the audio set after the second round of audio data processing into the speaker endpoint detection model for speaker endpoint detection, the time point of speaker change and the corresponding probability value can be obtained.
[0107] Continuing to refer to Figure 4 In the case where the screening mechanism of the third round of audio data processing includes a screening mechanism based on the speaker change number, the third round of audio data processing process further includes the following step S402:
[0108] S402, deleting the audio segments in the audio set after the third round of audio data processing that do not meet the preset number of speaker change requirements, or do not meet the preset number of speaker change requirements and do not meet the preset probability threshold of the corresponding speaker change probability.
[0109] For example, after the third round of audio data processing is completed, a screening mechanism based on speaker change detection is further performed on each audio segment in the obtained audio set, to ensure the coherence and consistency of the retained audio in terms of semantics and speech features.
[0110] Specifically, the rules of the screening mechanism are as follows:
[0111] For an audio segment in which no speaker transition is detected, it is considered that the content is completed by a single speaker, the speech style is uniform, and the semantics are coherent, and thus it is retained.
[0112] For an audio segment in which one speaker transition is detected and the probability value of the transition is lower than 0.7, although there is a speaker change, the confidence of the change is low, and it can be a false detection or a boundary case in which the transition is not obvious, and thus it can be retained.
[0113] For an audio segment in which two speaker transitions are detected and the probability value of each transition is lower than 0.3, although there are two potential speaker switches, since the confidence is low, the overall speech style can still be relatively consistent, and thus it can be retained.
[0114] If an audio segment does not meet any of the above conditions, for example, a high-frequency or high-confidence speaker transition occurs, or a complex scene such as a multi-person conversation is mixed, it is determined that the data is not suitable for high-quality speech synthesis training, and thus it is deleted.
[0115] By setting such a hierarchical retention condition, the data quality boundary can be flexibly controlled on the premise of ensuring the naturalness of the speech and the integrity of the semantics, so as to construct a more robust and applicable training data set.
[0116] Through the third round of audio data processing, the average retention rate of each language is about 50% of the original audio data to be processed.
[0117] Figure 5 The flowchart of the fourth round of audio data processing in the N rounds of audio data processing provided by the embodiments of the present application is shown in FIG. 5. Figure 5 As shown in FIG. 5, the fourth round of audio data processing is performed on the audio data to be processed, including the following steps S501-S502:
[0118] S501, in the fourth round of audio data processing, at least two speech recognition models are used to perform speech recognition on the audio set after the third round of audio data processing, to obtain at least two sets of transcribed texts.
[0119] The training data of the speech synthesis model usually requires to provide high-quality text-audio pairs, and has a high requirement on the consistency of the word and sound, and generally requires that the alignment accuracy at the word level reaches at least 90%. In order to improve the consistency of the word and sound between the text transcribed result and the audio content, the embodiments of the present application use multiple speech recognition models to perform transcribing in parallel, and enhance the accuracy and robustness of the recognition result through the collaborative manner of multiple models.
[0120] In some embodiments, step S501 comprises: performing speech recognition on the audio set processed in the third round of audio data processing by using the first speech recognition model and the second speech recognition model respectively, to obtain a first transcription text set and a second transcription text set.
[0121] The first speech recognition model and the second speech recognition model can be WhisperX model and Iflytek ASR model respectively. These models each have different training data bases and language understanding capabilities, and show strong complementarity in various contexts and speech conditions. By fusing the recognition results of multiple models, the bias and misrecognition rate brought by a single model can be effectively reduced, thereby further improving the alignment quality between the final text and the audio, and meeting the demand of high-precision word-sound consistency for speech synthesis training.
[0122] It should be noted that the first speech recognition model and the second speech recognition model can perform transcription tasks in parallel, or can perform transcription in a non-parallel manner as needed. Those skilled in the art can flexibly configure according to actual application scenarios and system resource conditions, and the present embodiment does not make specific limitations.
[0123] Continuing to refer to Figure 5 In the case where the fourth round of audio data processing includes a filtering mechanism based on comparison between multiple text transcription results, the fourth round of audio data processing process further comprises the following step S502:
[0124] S502, based on comparison between the at least two transcription text sets, obtaining an audio set processed in the fourth round of audio data processing and a corresponding transcription text set.
[0125] In some embodiments, step S502 comprises: performing a first preprocessing operation on the first transcription text set and the second transcription text set respectively, to obtain a preprocessed first transcription text set and a preprocessed second transcription text set; in the case where the comparison result between the preprocessed first transcription text set and the preprocessed second transcription text set does not meet the preset requirement, performing a second preprocessing operation on the preprocessed first transcription text set and the preprocessed second transcription text set respectively, until the comparison result between the new preprocessed first transcription text set and the new preprocessed second transcription text set meets the preset requirement, to obtain the audio set processed in the fourth round of audio data processing and the corresponding transcription text set.
[0126] For example, it is assumed that after completing the third round of audio data processing, an audio set containing multiple audio segments is obtained. For each of the audio segments, the corresponding transcription result in the first transcription text set is transcription text A, and the corresponding transcription result in the second transcription text set is transcription text B.
[0127] To determine whether the two transcriptions are consistent, first, the transcription A and the transcription B are respectively pre-processed, including converting to lowercase and removing all punctuation, to obtain pre-processed transcription a1 and pre-processed transcription b1. Then, the pre-processed transcription a1 and the pre-processed transcription b1 are compared character by character, and if they are completely consistent, the transcription A and the transcription B are retained.
[0128] If the pre-processed transcription a1 and the pre-processed transcription b1 are not completely consistent, it is determined whether there are numbers in the pre-processed transcription a1 and the pre-processed transcription b1. If there are numbers, the numbers are normalized to convert them into a textual representation in the corresponding language, to obtain pre-processed transcription a2 and pre-processed transcription b2. Then, the pre-processed transcription a2 and the pre-processed transcription b2 are compared string by string, and if they are completely consistent, the transcription A and the transcription B are retained.
[0129] If the pre-processed transcription a2 and the pre-processed transcription b2 are still not completely consistent, it is determined whether there are spaces in the pre-processed transcription a2 and the pre-processed transcription b2. If there are spaces, the spaces in the pre-processed transcription a2 and the pre-processed transcription b2 are deleted to obtain pre-processed transcription a3 and pre-processed transcription b3. Then, the pre-processed transcription a3 and the pre-processed transcription b3 are compared character by character, and if they are completely consistent, the transcription A and the transcription B are retained.
[0130] If the pre-processed transcription a3 and the pre-processed transcription b3 are still not completely consistent, the edit distance between the pre-processed transcription a3 and the pre-processed transcription b3 is calculated. If the edit distance is less than or equal to a preset value (for example, 1), the transcription A and the transcription B are retained.
[0131] If the edit distance is greater than the preset value, the "same pronunciation different forms" are processed according to the language characteristics of different languages. For example, in Japanese, the same word can be written using Chinese characters or using hiragana; in Arabic, a word can have a vowel marker or can omit it. For the above-mentioned language characteristics, it is necessary to process them into a unified format to obtain pre-processed transcription a4 and pre-processed transcription b4, and then recalculate the edit distance between the pre-processed transcription a4 and the pre-processed transcription b4. If the edit distance is less than or equal to a preset value (for example, 1), the transcription A and the transcription B are retained; otherwise, they are deleted.
[0132] After the above screening process, the final remaining transcription text A (obtained by WhisperX transcription) and transcription text B (obtained by Iflytek ASR transcription) will enter the next stage of the decision-making process. For each language, the transcription text with higher comprehensive recognition accuracy in that language can be selected as the final transcription result. For example, if WhisperX has higher comprehensive recognition accuracy in that language, then A is retained as the final transcription result, otherwise B is retained as the final transcription result.
[0133] Through the fourth round of audio data processing, the average retention rate of each language is about 32% of the original audio quantity.
[0134] Figure 6 The flowchart of the fifth round of audio data processing in the N rounds of audio data processing provided by the embodiments of the present application is shown in FIG. 6. As shown in FIG. 6, the fifth round of audio data processing is performed on the audio data to be processed, including the following steps S601-S602: Figure 6
[0135] S601, in the fifth round of audio data processing, the alignment processing is performed on the audio set and the corresponding transcription text set after the fourth round of audio data processing, to obtain the alignment result between the audio set and the corresponding transcription text set after the fourth round of audio data processing.
[0136] The alignment result includes each word in each audio segment in the audio set after the fourth round of audio data processing and the corresponding audio duration.
[0137] The speech synthesis model needs duration information at the word level in the training process to achieve accurate modeling of pronunciation duration. For Chinese, "word" is usually taken as the basic language unit; while for English, "word" is taken as the basic unit.
[0138] In order to obtain duration information at the word level, the audio-text pair after the fourth round of audio data processing needs to be aligned in time by the alignment model. The alignment model can use the open source Montreal ForcedAligner (MFA), which is a speech alignment tool based on Hidden Markov Model (HMM) and Deep Neural Network (DNN), and can achieve high-precision duration alignment between audio signals and their corresponding text content. The MFA model can be trained with a small amount of multi-lingual supervised data, so as to adapt to the alignment needs in multiple language scenarios.
[0139] Continuing to refer to Figure 6 In the case where the fifth round of audio data processing includes a screening mechanism based on the audio duration of each word in the audio segment, the fifth round of audio data processing process further includes the following step S602:
[0140] S602, for each audio segment in the fourth round of audio data processed audio set, if the number of target words in each word in the audio segment is greater than the preset value, or the average audio duration of each word in the audio segment is greater than the first preset average or less than the second preset average, delete the audio segment, wherein the target word is a word with a duration greater than a preset duration.
[0141] For example, assuming that the text corresponding to the audio segment y is {x1, x2...xn} (where xi represents the i th word in the text), after time alignment processing, its duration sequence at the word level can be obtained {t1, t2...tn} (where ti represents the duration corresponding to the i th word in the text, unit: second).
[0142] If the duration of a word or phrase ti in the duration sequence is greater than the preset duration (for example, 2 seconds), it indicates that there may be an unrecognized non-human voice area (such as silence or background noise) in the audio segment, and the alignment model incorrectly assigns this part of the non-human voice to the context of the human voice part, resulting in the duration of the corresponding word or phrase being artificially lengthened. Therefore, the audio-text pair needs to be deleted.
[0143] In addition, the average duration a of the audio segment can also be calculated and compared with the average duration b of all audio segments in the language. If a>1.3b or a<0.8b, it indicates that the audio speed is too fast or too slow, which may affect the training quality of the speech synthesis model, so the audio-text pair needs to be deleted.
[0144] Through the fifth round of audio data processing, the average retention rate of each language is about 29% of the original audio data to be processed.
[0145] In summary, the embodiment has the following beneficial effects:
[0146] (1) The hierarchical processing scheme of "processing and screening at the same time" is adopted, which cross-combines data parsing and data screening, not only effectively improves the overall processing efficiency, but also takes into account the retention rate and quality of the finished data.
[0147] (2) By using multiple speech recognition models to simultaneously transcribe the audio data, and based on the multiple transcribed texts, cross-validation screening rules are used for data screening to improve the consistency of the audio-text pairs after screening in terms of words and pronunciation in each language, and to ensure the usability of the output data in the multi-language speech synthesis scene.
[0148] (3) By dividing the long audio multiple times with different granularities, the semantic integrity is ensured while the retention rate of the final output data is effectively improved.
[0149] Exemplary apparatus
[0150] Corresponding to the audio data processing method described above, the embodiments of the present application also provide an audio data processing device. Figure 7 is a structural schematic diagram of an audio data processing device provided by the embodiments of the present application. As shown in Figure 7 the audio data processing device provided by the embodiments of the present application includes an acquisition unit 701 and a processing unit 702; wherein the acquisition unit 701 is configured to acquire audio data to be processed; the processing unit 702 is configured to sequentially perform N rounds of audio data processing on the audio data to be processed to obtain target audio data; the target audio data is used to train a speech synthesis model, and N is a positive integer; wherein each round of audio data processing in the N rounds of audio data processing includes audio data analysis and filtering of the analyzed audio data; the filtering criteria of the filtering mechanism corresponding to the N rounds of audio data processing gradually increase.
[0151] In some embodiments, the audio data to be processed includes various audio segments; the filtering mechanism corresponding to the N rounds of audio data processing includes at least one of the following: an audio duration-based filtering mechanism, an audio quality-based filtering mechanism, a speaker transition number-based filtering mechanism, a comparison-based filtering mechanism among multiple text transcription results, and a filtering mechanism based on the audio duration of the word segmentation in each audio segment.
[0152] In some embodiments, the processing unit 702 sequentially performs N rounds of audio data processing on the audio data to be processed to obtain target audio data, including: in the first round of audio data processing, the audio data to be processed is segmented into multiple audio segments with different segmentation granularities to obtain an audio set after the first round of audio data processing; and the audio segments with audio duration less than a first preset duration and / or the audio segments with audio duration greater than a second preset duration in the audio set after the first round of audio data processing are deleted.
[0153] In some embodiments, the processing unit 702 performs multiple audio segment splitting on the audio data to be processed with different splitting granularities to obtain an audio set after first round of audio data processing, including: performing audio segment splitting on the audio data to be processed with a first splitting granularity to obtain an initial audio set, the initial audio set including various audio segments; if a proportion of a target audio segment in the initial audio set is greater than or equal to a preset proportion, performing splitting on the target audio segment with a second splitting granularity, and updating the initial audio set until the proportion of the target audio segment in the updated initial audio set is less than the preset proportion, to obtain an audio set after first round of audio data processing, the first splitting granularity being greater than the second splitting granularity; wherein the target audio segment is an audio segment with a time length greater than or equal to the second preset time length.
[0154] In some embodiments, the processing unit 702 sequentially performs N rounds of audio data processing on the audio data to be processed to obtain target audio data, and further includes: in the second round of audio data processing, determining a quality index value corresponding to each audio segment in the audio set after first round of audio data processing to obtain an audio set after second round of audio data processing; deleting an audio segment in the audio set after second round of audio data processing whose quality index value does not meet a preset quality requirement.
[0155] In some embodiments, the processing unit 702 sequentially performs N rounds of audio data processing on the audio data to be processed to obtain target audio data, and further includes: in the third round of audio data processing, performing speaker detection on each audio segment in the audio set after second round of audio data processing to obtain an audio set after third round of audio data processing, the audio set after second round of audio data processing including each audio segment and a corresponding number of speaker changes and a corresponding probability of speaker change; deleting an audio segment in the audio set after third round of audio data processing whose number of speaker changes does not meet a preset number requirement, or whose number of speaker changes does not meet the preset number requirement and whose corresponding probability of speaker change does not meet a preset probability threshold.
[0156] In some embodiments, the processing unit 702 sequentially performs N rounds of audio data processing on the audio data to be processed to obtain target audio data, and further includes: in the fourth round of audio data processing, performing speech recognition on the audio set after third round of audio data processing using at least two speech recognition models to obtain at least two sets of transcribed texts; based on comparison between the at least two sets of transcribed texts, obtaining an audio set after fourth round of audio data processing and a corresponding set of transcribed texts.
[0157] In some embodiments, the processing unit 702 employs at least two speech recognition models to perform speech recognition on the third round of audio data processed audio set to obtain at least two sets of transcribed texts, including: employing a first speech recognition model and a second speech recognition model to perform speech recognition on the third round of audio data processed audio set to obtain a first set of transcribed texts and a second set of transcribed texts; wherein, based on the comparison between the at least two sets of transcribed texts, the fourth round of audio data processed audio set and the corresponding set of transcribed texts are obtained, including: performing a first preprocessing operation on the first set of transcribed texts and the second set of transcribed texts respectively to obtain a preprocessed first set of transcribed texts and a preprocessed second set of transcribed texts; in the case that the comparison result between the preprocessed first set of transcribed texts and the preprocessed second set of transcribed texts does not meet the preset requirement, performing a second preprocessing operation on the preprocessed first set of transcribed texts and the preprocessed second set of transcribed texts respectively until the comparison result between the new preprocessed first set of transcribed texts and the new preprocessed second set of transcribed texts meets the preset requirement, obtaining the fourth round of audio data processed audio set and the corresponding set of transcribed texts.
[0158] In some embodiments, the processing unit 702 sequentially performs N rounds of audio data processing on the audio data to be processed to obtain target audio data, and further includes: in the fifth round of audio data processing, performing alignment processing on the fourth round of audio data processed audio set and the corresponding set of transcribed texts to obtain an alignment result between the fourth round of audio data processed audio set and the corresponding set of transcribed texts, the alignment result including each word piece in each audio segment in the fourth round of audio data processed audio set and the corresponding audio duration; for each audio segment in the fourth round of audio data processed audio set, if the number of target word pieces in each word piece of the audio segment is greater than a preset value, or the average audio duration of each word piece in the audio segment is greater than a first preset average value or less than a second preset average value, the audio segment corresponding to the word piece is deleted; wherein, the target word piece is a word piece with an audio duration greater than a preset duration.
[0159] The audio data processing apparatus provided by the present embodiment belongs to the same application concept as the audio data processing method provided by the above-mentioned embodiments of the present application, can execute the audio data processing method provided by any of the above-mentioned embodiments of the present application, and has the corresponding function modules and beneficial effects of executing the audio data processing method. Technical details not described in detail in the present embodiment can be referred to the specific processing content of the audio data processing method provided by the above-mentioned embodiments of the present application, which will not be described here.
[0160] The functions implemented by the obtaining unit 701 and the processing unit 702 above can be implemented by the same or different processors, and the embodiments of the present application are not limited.
[0161] Corresponding to the training method of the speech synthesis model described above, the embodiments of the present application also provide a training device of a speech synthesis model. Figure 8 is a structural schematic diagram of a training device of a speech synthesis model provided by the embodiments of the present application. As shown in Figure 8 The training device of the speech synthesis model provided by the embodiments of the present application includes an obtaining unit 801 and a training unit 802; the obtaining unit 801 is configured to obtain audio sample data, the audio sample data being obtained by using the audio data processing method described above; and the training unit 802 is configured to train a speech synthesis model based on the audio sample data.
[0162] It should be understood that the units in the above device can be implemented in the form of processor calling software. For example, the device includes a processor connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to implement any of the above methods or to realize the functions of the units of the device, wherein the processor can be a general processor, such as CPU or microprocessor, and the memory can be an internal memory or an external memory of the device. Alternatively, the units in the device can be implemented in the form of hardware circuit, and the functions of part or all of the units can be realized by designing the hardware circuit, which can be understood as one or more processors; for example, in one implementation, the hardware circuit is ASIC, and the functions of part or all of the units are realized by designing the logical relationship of elements in the circuit; for example, in another implementation, the hardware circuit can be realized by PLD, and taking FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuration file, so as to realize the functions of part or all of the units. All units of the above device can be implemented in the form of processor calling software, or implemented in the form of hardware circuit, or implemented in the form of processor calling software and hardware circuit.
[0163] In the embodiments of the present application, the processor is a circuit with signal processing capability. In one implementation, the processor can be a circuit with instruction reading and running capability, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can implement certain functions through a logic relationship of a hardware circuit, which is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA, etc. In the reconfigurable hardware circuit, the processor loads a configuration document to implement the hardware circuit configuration. It can be understood that the processor loads instructions to implement the functions of the above units.
[0164] It can be seen that each unit in the above apparatus can be one or more processors (or processing circuits) configured to implement the above methods, such as a CPU, a GPU, an NPU, a TPU, a DPU, a microprocessor, a DSP, an ASIC, an FPGA, or a combination of at least two of these processor forms.
[0165] In addition, each unit in the above apparatus can be integrated together or can be independently implemented. In one implementation, these units are integrated together to implement a SOC. The SOC can include at least one processor for implementing any of the above methods or the functions of the units of the apparatus. The at least one processor can be different, such as including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0166] Exemplary electronic device
[0167] The embodiments of the present application provide an electronic device, referring to Figure 9 The electronic device includes:
[0168] a memory 200 and a processor 210;
[0169] The memory 200 is connected with the processor 210, and is configured to store programs.
[0170] The processor 210 is configured to implement the audio data processing method disclosed in any of the above embodiments by running the programs stored in the memory 200.
[0171] Specifically, the above electronic device can further include a bus, a communication interface 220, an input device 230, and an output device 240.
[0172] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected to each other through a bus. Among them:
[0173] The bus can include a path for transmitting information between various components of the computer system.
[0174] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready-to-use programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0175] The processor 210 can include a main processor and can also include a baseband chip, a modem, etc.
[0176] The memory 200 stores programs for executing the technical solutions of the present application, and can also store operating systems and other key services. Specifically, the program can include program code, and the program code includes computer operation instructions. More specifically, the memory 200 can include read-only memory (ROM), other types of static storage devices that can store static information and instructions, random access memory (RAM), other types of dynamic storage devices that can store information and instructions, disk storage, flash, etc.
[0177] The input device 230 can include a device that receives data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer or a gravity sensor, etc.
[0178] The output device 240 can include a device that allows information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0179] The communication interface 220 can include a device using any transceiver to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0180] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any one of the audio data processing methods or the training methods of the speech synthesis model provided by the embodiments of the present application.
[0181] The embodiments of the present application also provide a chip, which comprises a processor and a data interface. The processor reads and runs a program stored on a memory through the data interface to execute the audio data processing method or the training method of the speech synthesis model introduced in any of the above embodiments. The specific processing process and its beneficial effects can be referred to the above embodiments of the audio data processing method or the training method of the speech synthesis model.
[0182] Exemplary computer program product and storage medium
[0183] In addition to the above method and device, the embodiments of the present application can also be a computer program product, which comprises computer program instructions. When the computer program instructions are run by a processor, the processor executes the steps of the audio data processing method or the training method of the speech synthesis model according to various embodiments of the present application described in any of the above embodiments of the present application.
[0184] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0185] In addition, the embodiments of the present application can also be a storage medium, which stores a computer program. The computer program is executed by a processor to perform the steps of the audio data processing method or the training method of the speech synthesis model according to various embodiments of the present application described in any of the above embodiments of the present application. The steps of the above audio data processing method or the training method of the speech synthesis model can be implemented.
[0186] For each of the above method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action order described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0187] It should be noted that each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between embodiments can be mutually referred to.
[0188] The steps in the method of each embodiment of the application can be adjusted, combined and deleted in sequence according to actual needs. The technical features described in each embodiment can be replaced or combined.
[0189] The modules and sub-modules in the device and terminal of each embodiment of the application can be combined, divided and deleted according to actual needs.
[0190] In several embodiments provided by the application, it should be understood that the disclosed terminal, device and method can be implemented by other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or sub-modules is only a logical function division, and other division manners can be adopted in actual implementation, for example, a plurality of sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed mutual elements can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.
[0191] The modules or sub-modules described as separate components can or can not be physically separated, and the components of the modules or sub-modules can or can not be physical modules or sub-modules, that is, they can be located in one place or distributed on a plurality of network modules or sub-modules. Part or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0192] In addition, each functional module or sub-module in each embodiment of the application can be integrated in one processing module, or each module or sub-module can exist physically, or two or more modules or sub-modules can be integrated in one module. The integrated module or sub-module can be realized in the form of hardware or in the form of software functional module or sub-module.
[0193] Those skilled in the art will further appreciate that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their functionality, their composition, and their manner of operation. Whether such functionality is implemented in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled persons can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0194] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0195] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are more especially used for the purpose of identification in claims. Moreover, the terms "comprise", "include" or "contain" or any other variant thereof, are intended to cover non-exclusive inclusions, such that processes, methods, articles, or apparatuses that comprise, include, or contain a list of elements, not only include those elements, but also other elements not expressly listed or otherwise inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0196] The above description of disclosed embodiments provides enabling disclosure sufficient for one of ordinary skill in the art to implement or use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing audio data, characterized in that: include: Get the audio data to be processed; Performing N rounds of audio data processing on the audio data to be processed in sequence to obtain target audio data; The target audio data is used to train a speech synthesis model, and N is a positive integer; Each of the N rounds of audio data processing includes parsing the audio data and filtering the parsed audio data. The screening criteria of the screening mechanism corresponding to the N rounds of audio data processing gradually increase.
2. The method according to claim 1, characterized in that The audio data to be processed includes various audio clips; The screening mechanisms corresponding to the N rounds of audio data processing include: a screening mechanism based on audio duration, a screening mechanism based on audio quality, a screening mechanism based on the number of speaker changes, a screening mechanism based on comparison between multiple text transcription results, and a screening mechanism based on the audio duration of the word segmentation in each audio segment.
3. The method according to claim 1, characterized in that Performing N rounds of audio data processing on the audio data to be processed in sequence to obtain target audio data, including: In the first round of audio data processing, the audio data to be processed is divided into multiple audio segments using different segmentation granularities to obtain an audio set after the first round of audio data processing; Delete the audio segments whose audio duration is less than the first preset duration and / or the audio segments whose audio duration is greater than the second preset duration in the audio set after the first round of audio data processing.
4. The method according to claim 3, characterized in that The method of performing multiple audio segmentation on the audio data to be processed using different segmentation granularities to obtain an audio set after the first round of audio data processing includes: Segmenting the audio data to be processed into audio segments using a first segmentation granularity to obtain an initial audio set, wherein the initial audio set includes each audio segment; If the proportion of the target audio segment in the initial audio set is greater than or equal to a preset proportion, segmenting the target audio segment using a second segmentation granularity and updating the initial audio set until the proportion of the target audio segment in the updated initial audio set is less than a preset ratio, thereby obtaining an audio set after the first round of audio data processing, where the first segmentation granularity is greater than the second segmentation granularity; The target audio segment is an audio segment whose audio duration is greater than or equal to the second preset duration.
5. The method according to claim 3, characterized in that Performing N rounds of audio data processing on the audio data to be processed in sequence to obtain target audio data, further comprising: During the second round of audio data processing, determining the quality index value corresponding to each audio segment in the audio set after the first round of audio data processing, to obtain the audio set after the second round of audio data processing; The audio segments whose quality index values do not meet the preset quality requirements are deleted from the audio set after the second round of audio data processing.
6. The method according to claim 5, characterized in that Performing N rounds of audio data processing on the audio data to be processed in sequence to obtain target audio data, further comprising: During the third round of audio data processing, speaker detection is performed on each audio segment in the audio set after the second round of audio data processing to obtain an audio set after the third round of audio data processing, wherein the audio set after the second round of audio data processing includes each audio segment and the corresponding number of speaker changes and speaker change probabilities; Delete the audio segments in the audio set after the third round of audio data processing, where the number of speaker changes does not meet the preset number requirement, or the number of speaker changes does not meet the preset number requirement and the corresponding speaker change probability does not meet the preset probability threshold.
7. The method according to claim 6, characterized in that Performing N rounds of audio data processing on the audio data to be processed in sequence to obtain target audio data, further comprising: In the fourth round of audio data processing, at least two speech recognition models are used to perform speech recognition on the audio set after the third round of audio data processing, to obtain at least two transcribed text sets; Based on the comparison between the at least two transcribed text sets, an audio set after the fourth round of audio data processing and a corresponding transcribed text set are obtained.
8. The method according to claim 7, characterized in that The at least two speech recognition models are used to perform speech recognition on the audio set after the third round of audio data processing, respectively, to obtain at least two transcribed text sets, including: Using a first speech recognition model and a second speech recognition model to perform speech recognition on the audio set after the third round of audio data processing, respectively, to obtain a first transcribed text set and a second transcribed text set; The method of obtaining an audio set and a corresponding transcribed text set after the fourth round of audio data processing based on the comparison between the at least two transcribed text sets includes: performing a first preprocessing operation on the first transcribed text set and the second transcribed text set respectively to obtain a preprocessed first transcribed text set and a preprocessed second transcribed text set; When the comparison result between the preprocessed first transcription text set and the preprocessed second transcription text set does not meet the preset requirements, a second preprocessing operation is performed on the preprocessed first transcription text set and the preprocessed second transcription text set respectively until the comparison result between the new preprocessed first transcription text set and the new preprocessed second transcription text set meets the preset requirements, thereby obtaining the audio set and the corresponding transcription text set after the fourth round of audio data processing.
9. The method according to claim 7, characterized in that Performing N rounds of audio data processing on the audio data to be processed in sequence to obtain target audio data, further comprising: During the fifth round of audio data processing, the audio set after the fourth round of audio data processing and the corresponding transcribed text set are aligned to obtain an alignment result between the audio set after the fourth round of audio data processing and the corresponding transcribed text set, the alignment result including each word segment and corresponding audio duration of each audio segment in the audio set after the fourth round of audio data processing; For each audio segment in the audio set after the fourth round of audio data processing, if the number of target segmented words in each segmented word in the audio segment is greater than a preset value, or the average audio duration of each segmented word in the audio segment is greater than a first preset average value or less than a second preset average value, then delete the audio segment; The target segmentation word is a segmentation word whose audio duration is greater than a preset duration.
10. A method for training a speech synthesis model, characterized in that: include: Acquire audio sample data, where the audio sample data is obtained using the method according to any one of claims 1 to 9; A speech synthesis model is obtained by training based on the audio sample data.
11. An audio data processing device, characterized in that: include: An acquisition unit, used for acquiring audio data to be processed; a processing unit, configured to sequentially perform N rounds of audio data processing on the audio data to be processed to obtain target audio data; the target audio data is used to train a speech synthesis model, where N is a positive integer; Each of the N rounds of audio data processing includes parsing the audio data and filtering the parsed audio data. The screening criteria of the screening mechanism corresponding to the N rounds of audio data processing gradually increase.
12. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method according to any one of claims 1 to 9 by running the program in the memory.
13. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 9.