Language data preprocessing method for multilingual and complex scene

By combining the AutoPrep framework and deep learning algorithms, end-to-end speech data preprocessing in complex multilingual scenarios is achieved, solving the problems of high phoneme mapping error rate and poor adaptability of segmentation technology, improving speech signal-to-noise ratio and speaker annotation efficiency, and making it suitable for cross-language application scenarios.

CN120913580AActive Publication Date: 2025-11-07SHENYI FUTURE TECHNOLOGY (GUANGDONG HENGQIN) CO LTD

Patent Information

Application Number
CN202511447963.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-07
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing speech data preprocessing technologies suffer from problems such as high phoneme mapping error rates in multilingual and complex scenarios, poor adaptability of traditional segmentation techniques, and low efficiency in engineering implementation due to lack of metadata annotation.

Method used

The AutoPrep framework is adopted, which combines cross-lingual pre-trained models and deep learning algorithms to achieve end-to-end language data preprocessing through speech enhancement, speech segmentation, speaker clustering and quality filtering modules, including dynamic block denoising, multilingual VAD model, speaker embedding extraction and language adaptive scoring.

Benefits of technology

It improves the signal-to-noise ratio and language independence of speech in less commonly spoken languages, enhances speech segmentation accuracy and speaker annotation efficiency, reduces manual annotation costs, and ensures the scalability and stability of the system in complex multilingual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913580A_ABST
    Figure CN120913580A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of language data processing, and discloses a language data preprocessing method in a multilingual and complex scene, which is a voice data preprocessing system in the multilingual and complex scene based on an AutoPrep framework, integrates five modules, namely a voice enhancement module, a voice segmentation module, a speaker clustering module, a target voice extraction module and a quality filtering module. According to the scheme, differential suppression of steady-state noise and transient-state noise in multilingual voice signals is realized, and particularly in a small language (such as Kazakh and Talanx) scene, the voice signal-to-noise ratio and the language independence of voice features are effectively improved, so that the voice data can be automatically and structurally processed, and the voice signal-to-noise ratio and the language independence of the voice features can be effectively improved. The problem that in the prior art, the phoneme mapping error rate is high due to the fact that small languages lack an exclusive phoneme system processing module is solved, and the availability and the processing effect of low-resource language data are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of language data processing, and in particular to a language data preprocessing method for multi-lingual and complex scenarios. BACKGROUND

[0002] The development of speech data preprocessing has been continuously iterating with the evolution of speech technology. In the early stage, it mainly focused on basic processing such as noise reduction and endpoint detection in a single language and pure environment. With the increasing demand for multi-lingual interaction and complex scenarios (such as noisy environment and multi-sound source interference), a preprocessing system integrating language adaptation, environment robustness enhancement, and cross-language feature normalization has been gradually developed. Meanwhile, deep learning is used to realize end-to-end noise suppression and feature optimization. The importance of this lies in that in multi-lingual and complex scenarios, high-quality preprocessing is the cornerstone of downstream tasks such as speech recognition and synthesis. It can effectively eliminate the interference of pronunciation differences of different languages, environmental noise, device distortion, etc., improve the purity and consistency of speech signals, avoid model training bias caused by original data quality problems, and lay a key foundation for accurate understanding and expression of multi-lingual speech systems in complex scenarios. It is the core prerequisite for solving the challenges of cross-language and strong interference in practical applications.

[0003] The existing speech data preprocessing technology has many limitations in practical applications.

[0004] In terms of language processing, the current scheme has formed a mature processing flow for high-resource languages (such as Chinese and English), covering pronunciation normalization based on phoneme dictionary, dialect accent adaptive model, and large-scale corpus-driven acoustic feature extraction. However, for small languages such as Kazakh and Tagalog, due to the lack of customized processing modules for their unique phonological rules (such as Kazakh vowel harmony and Tagalog stress pattern), they only rely on transfer learning or cross-language models, resulting in high phoneme set mapping error rate.

[0005] In terms of scene adaptability, endpoint detection and simple noise reduction techniques have been applied in near-field pure scenes. However, in complex scenarios, traditional multi-person overlapping speech separation algorithms often fail in single-channel conditions. Far-field recording has positioning deviation due to microphone array calibration errors, and real-time noise suppression lacks robust algorithms, making it difficult to cope with real environments such as conference room reverberation and street noise.

[0006] At the level of data labeling and meta-information processing, the mainstream scheme mainly adopts manual or semi-automatic labeling of audio start and end time, lacks key meta-information such as speaker ID, signal-to-noise ratio quality score and language type label, cannot construct speaker embedding due to un-labeled speaker attribute, and makes data difficult to directly connect with ASR / TTS model training requirements due to missing fine-grained labels such as environmental noise type, which requires additional investment of manpower for secondary labeling, and seriously restricts the efficiency of multi-lingual speech system engineering landing; therefore, the present application proposes a language data preprocessing method for multi-lingual and complex scenarios to solve the above problems. SUMMARY

[0007] In order to solve the above problems, the present application provides a language data preprocessing method for multi-lingual and complex scenarios.

[0008] The language data preprocessing method for multi-lingual and complex scenarios provided by the present application adopts the following technical scheme:

[0009] A language data preprocessing method for multi-lingual and complex scenarios, based on the AutoPrep framework, audio sequentially flows through a speech enhancement module, a speech segmentation module, a speaker clustering module, a target speech extraction module and a quality filtering module for processing, and realizes end-to-end processing from raw speech to structured data through the fusion of a cross-language pre-training model and a deep learning algorithm.

[0010] The speech enhancement module is composed of dynamic blocking, XLSR-53, BSRNN and pBSRNN core components, and realizes intelligent denoising processing of input audio.

[0011] The speech segmentation module frames the audio to be input into a pre-trained VAD model, which is based on a TDNN-Transformer architecture and can accurately identify whether each frame of audio belongs to “speech” or “non-speech”.

[0012] The speaker clustering module inputs the data processed by speech segmentation into the WeSpeaker-XL model after blocking to extract speaker embedding vectors, and constructs a similarity matrix by calculating the cosine similarity between each pair of embedding vectors.

[0013] Preferably, the speech enhancement module is mainly composed of dynamic blocking, a pre-training model with cross-lingual feature extraction capability and a speech enhancement model core component, and realizes intelligent denoising processing of input audio.

[0014] Preferably, in the speech enhancement module:

[0015] S11, the input audio is first divided into long audio and short audio according to the time length;

[0016] S12, the segmented audio is extracted by a pre-trained model with cross-lingual feature extraction capability to extract language-independent robust features, suppress environmental noise and capture long-distance dependencies of the audio sequence;

[0017] S13, the long audio segment and the short audio segment are processed by a speech enhancement model, the long audio segment captures long-time sequence dependency, and the short audio segment quickly suppresses burst noise;

[0018] S14, the denoised audio segment is integrated by time sequence alignment and overlap splicing technology, the long audio adopts a weighted overlap addition algorithm to eliminate boundary effects, and the short audio is post-processed to optimize spectral continuity, thereby outputting a high-fidelity denoised audio.

[0019] Preferably, in the speech segmentation module:

[0020] S21, the audio signal is subjected to amplitude normalization processing;

[0021] S22, the audio is subjected to frame processing;

[0022] S23, the framed audio is input into a trained voice activity detection model;

[0023] S24, according to the speech probability output by the model, an adaptive silence threshold is set, and continuous frames with high probability are determined as speech segments, and the rest are classified as silence;

[0024] S25, the effective speech is labeled with start and end time, and the corresponding waveform data is cut out from the original audio.

[0025] Preferably, in the speaker clustering module:

[0026] S31, the data processed by speech segmentation is segmented and input into a speaker embedding extraction network model to extract speaker embedding vectors, and a similarity matrix S is constructed by calculating the cosine similarity between each pair of embedding vectors;

[0027] S32, spectral clustering is performed thereon, the S is directly converted into an adjacency matrix W, a degree matrix D is calculated, and a symmetric normalized Laplacian matrix is constructed

[0028]

[0029] S33, eigenvalue decomposition is performed on L sym , the number of speakers k is automatically estimated by using the "eigenvalue gap", the first k eigenvectors are selected to form a matrix U, and each row vector is normalized, and then K-means is applied to divide them into k clusters;

[0030] S34, the cosine similarity between the cluster centers or representative embeddings is calculated, and if it exceeds a dynamic threshold, similar clusters are merged to further refine the speaker division;

[0031] S35. Assign a unique speaker ID to each cluster and map the position of each segment of the original audio embedded in the sliding window to the corresponding ID.

[0032] Preferably, for continuous streams of recordings longer than 2 hours, the spectral clustering state can be reset every 2 hours: retaining historical cluster centers, matching and updating cluster assignments with newly generated embedding vectors to maintain speaker ID consistency, while cleaning up old data.

[0033] Preferably, in the target speech extraction module:

[0034] S41. Merge adjacent speech segments from the same speaker in the speaker clustering results, and input the merged speech segments into STFT to extract spectral features and provide frequency domain representation for subsequent processing.

[0035] S42. Combine the spectral features obtained from STFT with the target speaker embedding, and use this as input to the advanced speech separation network model in target speech extraction to obtain M. tgt (t,f) and M noiset (t,f);

[0036] S43. For single-channel cases, use M. tgt (t,f) through The frequency domain representation of the separated target speaker is calculated and input into the RNN filter. The residual noise, reverberation, and frequency smoothing information are further modeled in the frequency domain, and then iSTFT is performed to transform the frequency-enhanced signal back to the time domain.

[0037] S44, For multi-channel cases, use M noiset The covariance and target direction estimation are calculated, and the MVDR beamforming formula is calculated as follows:

[0038]

[0039] in Let d(f) be the noise covariance matrix, d(f) be the target direction vector, and w(f) be the optimal weighting vector at frequency f. The multi-channel spectrum is weighted, and iSTFT is performed to transform the frequency-enhanced signal back to the time domain.

[0040] Preferably, the quality filtering module includes language adaptive scoring and anomaly detection technology:

[0041] In language adaptive scoring, the scoring model uses a differentiated threshold.

[0042] Anomaly detection detects anomalies by analyzing changes in audio energy, spectral characteristics, and time-domain waveforms.

[0043] Preferably, the quality filtering process comprises:

[0044] S51, receiving the speech segment from the target speech extraction module;

[0045] S52, applying a corresponding threshold value according to the language type to perform quality scoring;

[0046] S53, performing abnormality detection to eliminate or repair unqualified segments;

[0047] S54, outputting high-quality speech segments with quality score indicators, providing a basis for further screening and processing for downstream tasks.

[0048] In summary, the present application includes at least one of the following beneficial technical effects:

[0049] The present application introduces XLSR-53 cross-language pre-training model and BSRNN / pBSRNN dynamic block noise reduction algorithm, realizes the differential suppression of stable and transient noise in multilingual speech signals, especially in small language (such as Kazakh, Tagalog) scenarios, effectively improves the speech signal-to-noise ratio and language independence of speech features, overcomes the problem of high phoneme mapping error rate caused by the lack of exclusive phonetic processing module in small language in the prior art, and enhances the availability and processing effect of low-resource language data.

[0050] By constructing a multi-language voice activity detection (VAD) model based on the TDNN-Transformer hybrid architecture, the noise detection threshold is adaptively adjusted according to the signal-to-noise ratio, and the speech segment duration specification strategy is combined to realize accurate segmentation of continuous audio in complex scenes (such as street noise, far-field conference reverberation environment), improve the recognition accuracy of speech segment boundaries, significantly improve the speech separation effect in single-channel, multi-speaker overlapping speech environment, and solve the poor adaptability problem of traditional segmentation technology in complex conditions.

[0051] The present scheme uses WeSpeaker-XL model to extract multi-language universal speaker embedding, and combines incremental spectral clustering technology and language adaptive threshold to realize high-precision clustering and automatic speaker labeling of multi-speaker speech, solving the problem of lack of speaker meta-information in speech data in traditional preprocessing process and the need for a large amount of manual secondary labeling. This method effectively improves the speaker label generation efficiency, reduces the artificial labeling cost by 40%-60% while ensuring the clustering accuracy, and greatly shortens the engineering landing cycle of multilingual speech system.

[0052] By taking the speaker embedding as a prior input, combining the pBSRNN model and the simulated beamforming algorithm, precise speech extraction and reverberation suppression of the target speaker in single-channel and multi-channel scenarios are realized.

[0053] Finally, the application constructs a language-adaptive speech quality filtering system, sets DNSMOS and PDNSMOS score thresholds for high-resource and low-resource languages respectively, and combines an abnormality detection mechanism and a short-time noise repair strategy to effectively eliminate speech segments with low speech quality, incomplete content or non-speech interference, avoid the problem of over-filtering small language data due to low scores, and ensure the scalability and stability of the system when processing multi-lingual and large-scale speech data, which is suitable for cross-border customer service, multi-language conference and other diversified application scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 is a system flowchart of an embodiment of the application.

[0055] Figure 2 is an audio enhancement flowchart of an embodiment of the application.

[0056] Figure 3 is a VAD model structure diagram based on a TDNN-Transformer architecture of an embodiment of the application.

[0057] Figure 4 is a speaker clustering flowchart of an embodiment of the application.

[0058] Figure 5 is a SAM structure diagram of an embodiment of the application.

[0059] Figure 6 is a target speech extraction flowchart of an embodiment of the application. DETAILED DESCRIPTION

[0060] The following will be described in conjunction with the accompanying drawings Figure 1 - the accompanying drawings Figure 6 The application will be further described in detail.

[0061] Embodiment One: A language data preprocessing method for multi-lingual and complex scenarios

[0062] I. Overall framework design

[0063] A multi-lingual and complex scenario speech data preprocessing system is built based on the AutoPrep framework, which includes five core modules of speech enhancement, speech segmentation, speaker clustering, target speech extraction, and quality filtering. Through the fusion of cross-language pre-training models and deep learning algorithms, end-to-end processing from raw speech to structured high-quality data is realized. The system supports cross-language processing of 50+ languages, and outputs JSON structured data containing speaker labels and quality indicators.

[0064] The input audio needs to be formatted as WAV and the sampling rate needs to be adjusted to 16 kHz before preprocessing. Then, the audio sequentially passes through the speech enhancement, speech segmentation, speaker clustering, target speech extraction, and quality filtering modules. In the speaker clustering stage, the "merge" operation integrates adjacent speech segments belonging to the same speaker, and finally outputs the fully processed audio.

[0065] II. Core module design

[0066] 1. Speech enhancement

[0067] The speech enhancement module is mainly composed of dynamic blocking, XLSR-53, BSRNN, and pBSRNN core components, which can realize intelligent denoising processing of input audio. Input audio is divided into long audio (longer than 3 minutes) and short audio (shorter than 3 minutes) according to duration. During dynamic blocking, long audio uses a sliding blocking strategy with a "block window of 12 seconds and an offset of 4 seconds", ensuring the continuity of temporal information through overlapping windows, while short audio uses a "fixed window of 8 seconds" to balance computational efficiency and feature integrity.

[0068] The blocked audio extracts language-independent robust features with the help of XLSR-53. This model is trained on large-scale multi-lingual data and can break through language barriers, effectively suppress environmental noise, and capture long-distance dependencies of audio sequences.

[0069] BSRNN and pBSRNN are neural networks trained on mixed audio data in multi-lingual and multi-noise scenarios. Long audio blocks are processed by BSRNN, which has a bidirectional recursive structure suitable for capturing long temporal dependencies. Short audio blocks are processed by pBSRNN, which has a lightweight architecture that can quickly suppress burst noise for short audio.

[0070] Finally, each denoised audio segment is integrated through temporal alignment and overlapping splicing technology. Long audio uses a weighted overlapping addition algorithm to eliminate boundary effects, and short audio is post-processed to optimize spectral continuity, thus outputting high-fidelity denoised audio.

[0071] 2. Speech segmentation

[0072] To ensure the effectiveness of speech segmentation, amplitude normalization is performed on the audio signal after speech enhancement. The specific method is to divide the amplitude of each sampling point in the entire audio by the maximum absolute amplitude of the audio, so that the maximum amplitude of the waveform is compressed to the range of ±1. The mathematical expression is as follows:

[0073]

[0074] where, Yn represents the original amplitude of the nth sampling point, is the largest absolute amplitude in the whole audio, is the normalized amplitude value.

[0075] After normalization, the audio will be processed by frame blocking. The length of each frame is set to 25 milliseconds (corresponding to 400 sampling points), and the frame shift is 10 milliseconds (160 sampling points). If the tail of the audio is less than one frame, it will be padded with zeros. Then, each frame is multiplied by a Hamming window function to reduce spectral leakage at the frame boundaries and improve the accuracy of subsequent modeling.

[0076] After the above preprocessing, the framed audio will be input into the pre-trained VAD (Voice Activity Detection) model. This model is based on the TDNN-Transformer architecture and has been trained on multi-lingual silent and non-silent corpora. It can accurately identify the probability of each frame of audio belonging to "speech" or "non-speech".

[0077] Next, according to the speech probability output by the model, by setting an adaptive silence threshold (such as 0.8 for high SNR environment and 0.6 for low SNR environment), the continuous frames with high probability are determined as speech segments, and the rest are classified as silence. The preliminary detected speech segments are post-processed: first, merge adjacent speech segments with a silence interval of less than 1 second to preserve the integrity of the speech stream; second, for long speech segments exceeding 30 seconds, split them at the first silence point inside to prevent information blocks from being too long; finally, filter out short speech segments less than 1.5 seconds to improve the effectiveness and usability of the speech segments.

[0078] After the above processing, each valid speech segment will be labeled with start and end times, and the corresponding waveform data will be cut from the original audio. These structured and accurately bounded speech segments will serve as key inputs for subsequent speaker clustering, ensuring the robustness and high-quality output of the overall speech processing pipeline.

[0079] 3. Speaker clustering

[0080] After the speech segmentation processing data is blocked (every 1.5 seconds, 0.75 seconds sliding), it is input into the WeSpeaker-XL model to extract speaker embedding vectors. A similarity matrix is constructed by calculating the cosine similarity between each pair of embedding vectors. The formula for calculating the cosine similarity is as follows:

[0081]

[0082] where is the dot product, is the L2 norm.

[0083] Then spectral clustering is performed, converting S directly to an adjacency matrix W, computing a degree matrix D and constructing a symmetric normalized Laplacian matrix

[0084]

[0085] Eigenvalue decomposition is performed on L sym The number of speakers k is automatically estimated using the "Eigenvalue Gap", the top k eigenvectors are selected to form matrix U, and each row vector is normalized. Then K-means is applied to divide them into k clusters. After that, by calculating the cosine similarity between the cluster centers or representative embeddings, if it exceeds the dynamic threshold (usually set to 0.75, some language systems can be adjusted to 0.7 or 0.6), similar clusters are merged to further refine the speaker division.

[0086] Finally, a unique speaker ID (such as spk_001, spk_002…) is assigned to each cluster, and the position of the original audio embedding sliding window is mapped to the corresponding ID, the output format can be a timestamped JSON or CSV file, or directly embedded in the audio metadata.

[0087] For continuous streams or super-long recordings (> 2 hours), the process can reset the spectral clustering state every 2 hours: keep the historical cluster centers, match and update the cluster assignment with newly generated embedding vectors to maintain the consistency of speaker IDs, while cleaning up old data to ensure that the calculation and memory are controllable.

[0088] 4. Target speech extraction

[0089] Merge the speech segments of the same speaker in the speaker clustering result, and input the merged speech segments into STFT to extract spectral features, providing frequency domain representation for subsequent processing. Combine the spectral features obtained by STFT with the target speaker embedding to serve as the input of the pBSRNN model in target speech extraction, and obtain M tgt (t, f) and M noiset (t, f). The pBSRNN here is different from the pBSRNN in the speech processing part. Unlike the pBSRNN in the speech enhancement part, the pBSRNN in target speech extraction adds a sequence attention module (SAM) after the Band & Sequence Modeling Module. This module enhances the model's ability to model the time dependence of the speech sequence through multi-head attention mechanism, significantly improving the separation performance in complex scenarios. This pBSRNN model is trained using multi-lingual, mixed speech and single-person speaking audio datasets, optimizing the generalization ability to multiple languages and noise environments.

[0090] For single-channel cases, M tgt (t, f) is obtained by The frequency domain representation of the target speaker after separation is calculated and input to the RNN post-filter Further model the residual noise, reverberation, frequency domain smoothing information in the frequency domain, and then perform iSTFT to transform the frequency domain enhanced signal back to the time domain.

[0091] For the multi-channel case, the M noiset The covariance sum and target direction estimation are calculated, and the MVDR beamforming formula is calculated as follows:

[0092]

[0093] Where is the noise covariance matrix, d(f) is the target direction vector, and w(f) represents the optimal weighting vector at frequency f.

[0094] Then the multi-channel spectrum is weighted , and then iSTFT is performed to transform the frequency domain enhanced signal back to the time domain.

[0095] Through the above two sets of processes for single-channel and multi-channel, the target speech extraction module can achieve high-quality separation and dereverberation of the target speaker under different hardware and environmental conditions.

[0096] 5. Quality filtering

[0097] Quality filtering (Quality Filtering) is a key step in the speech processing process, aiming to eliminate low-quality speech segments and ensure that downstream tasks (such as speech recognition, speaker recognition, etc.) can receive high-quality and reliable input data, thereby improving the performance and robustness of the overall system. This module realizes effective control of speech quality through language adaptive scoring and anomaly detection technology.

[0098] In terms of language adaptive scoring, the system will use different quality thresholds based on different language resources. For high-resource languages (such as Chinese and English), due to the rich training data and mature model support, the system sets a higher quality standard, requiring the DNSMOS (a deep learning-based speech denoising quality evaluation index) of the speech segment to be greater than 3.2, and the PDNSMOS (a variant or related index of DNSMOS) to be greater than 3.8, to ensure the high quality of the input speech. For low-resource languages (such as Tajik), due to the limited training data and limited model evaluation accuracy, the system relaxes the threshold, requiring DNSMOS to be greater than 2.8 and PDNSMOS to be greater than 3.2, while combining speaker similarity compensation (similarity greater than 0.5) to balance between quality and data availability.

[0099] Abnormality detection is another core part of quality filtering, which aims to identify and remove speech segments containing non-speech burst noise (such as siren sound, impact sound) or device failure (such as microphone disconnection, audio distortion). The system detects abnormalities by analyzing the energy change, spectral features and time domain waveform of the audio. For short noise with a duration less than 0.5 seconds, the system will use interpolation technology based on adjacent frames to repair it to preserve valuable speech content as much as possible; while for abnormal segments that cannot be repaired, they are directly removed.

[0100] The overall process of quality filtering can be summarized as follows: first, receive the speech segments from the target speech extraction module; then, apply the corresponding DNSMOS and PDNSMOS threshold for quality scoring according to the language type; then, perform abnormality detection and remove or repair unqualified segments; finally, output high-quality speech segments with quality score indicators (such as DNSMOS, PDNSMOS, STOI) to provide further filtering and processing basis for downstream tasks. Through this series of technical means, the quality filtering module effectively guarantees the quality of the speech segments and provides a reliable data foundation for subsequent processing.

[0101] Embodiment two:

[0102] 1. Alternative of speech enhancement module: In the denoising process, in addition to BSRNN or pBSRNN, deep learning models such as Demucs, SEGAN, DCCRN, etc. can also be used for speech enhancement to improve signal-to-noise ratio and reduce reverberation. In addition, the XLSR-53 model can be replaced by wav2vec2.0, HuBERT, Whisper, etc. pre-trained models with cross-lingual feature extraction capabilities to adapt to different application requirements and computing resources.

[0103] 2. Alternative of speech segmentation method: In the multi-language VAD (Voice Activity Detection), deep learning models based on Conformer, Bi-GRU, Bi-LSTM, etc. can be used to replace the TDNN-Transformer architecture, or multi-modal information (such as video or voiceprint features) can be introduced to improve the accuracy of silence / speech boundary recognition.

[0104] 3. Speaker clustering alternative: In addition to the incremental spectral clustering method, end-to-end speaker separation methods (such as EEND), clustering methods based on PLDA (Probabilistic Linear Discriminant Analysis), or self-supervised clustering strategies can be used to separate and label multiple speaker voices; at the same time, WeSpeaker-XL can also be replaced by ECAPA-TDNN, ResNet-based embedding model, and other speaker embedding extraction networks.

[0105] 4. Target speech extraction method alternatives: For the separation of target speakers, not limited to using pBSRNN models, advanced speech separation networks such as Conv-TasNet, DPRNN, Sepformer, etc. can also be used; in terms of analog neural beamforming, traditional MVDR or GSS (Geometric Source Separation) methods can also be used to complete the directional extraction of multi-channel signals.

[0106] 5. Quality filtering mechanism alternatives: In addition to DNSMOS and PDNSMOS score models, PESQ, STOI, SI-SDR, etc. can also be used to combine subjective and objective evaluation indicators to perform multi-dimensional evaluation of speech quality; the anomaly detection part can also be replaced by a convolutional neural network (CNN) or a Transformer model to detect spectral anomaly patterns instead of energy / waveform statistics to enhance the intelligence and accuracy of detection.

[0107] System structure alternatives: The overall processing flow can be trimmed or the order adjusted according to the computing resources and real-time requirements. For example, in a high-performance computing environment, speech enhancement and target speech extraction can be combined into an integrated separation model; in an edge computing scenario, only the speech segmentation and quality filtering modules can be retained to quickly extract usable segments for real-time recognition.

[0108] It should be noted that in this text, terms such as "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device.

[0109] Although embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. A language data preprocessing method for multi-lingual and complex scenes, characterized in that: Based on the AutoPrep framework, the audio sequentially flows through the five modules of speech enhancement module, speech segmentation module, speaker clustering module, target speech extraction module and quality filtering module for processing, through the fusion of cross-language pre-training model and deep learning algorithm, the end-to-end processing from raw speech to structured data is realized; The speech enhancement module is composed of dynamic block, XLSR-53, BSRNN and pBSRNN core components, and realizes intelligent denoising processing of the input audio; The frame audio in the speech segmentation module will be input into the pre-trained VAD model, which is based on the TDNN-Transformer architecture and can accurately identify whether each frame of audio belongs to "speech" or "non-speech"; The data processed by the speech segmentation module is input into the WeSpeaker-XL model after blocking to extract the speaker embedding vector, and a similarity matrix is constructed by calculating the cosine similarity between each pair of embedding vectors.

2. The method according to claim 1, characterized in that: The speech enhancement module is mainly composed of dynamic block, pre-trained model with cross-language feature extraction capability and speech enhancement model core component, which realizes intelligent denoising processing of the input audio.

3. The method according to claim 2, characterized in that: In the speech enhancement module: S11, the input audio is first divided into long audio and short audio according to the time length; S12, the audio after blocking is extracted by the pre-trained model with cross-language feature extraction capability to extract language-independent robust features, suppress environmental noise and capture long-distance dependence of audio sequence; S13, both long audio blocks and short audio blocks are processed by the speech enhancement model, and long audio blocks capture long-time sequence dependence; Short audio blocks quickly suppress burst noise; S14, the denoised audio segments are integrated through time alignment and overlap splicing technology, long audio adopts weighted overlap addition algorithm to eliminate boundary effect, and short audio is post-processed to optimize spectral continuity, so as to output high-fidelity denoised audio.

4. The method of claim 1, wherein the method is used for multi-lingual and complex scene language data preprocessing. In the speech segmentation module: S21, the amplitude of the audio signal is normalized; S22, the audio is frame processed; S23, the frame audio will be input into the trained speech activity detection model; S24, according to the speech probability output by the model, set the adaptive silence threshold, judge the continuous frames with high probability as speech segments, and the rest as silence; S25, label the start and end time of the effective speech, and cut out the corresponding waveform data from the original audio.

5. The method for preprocessing language data in multiple languages and complex scenes according to claim 1, characterized in that: In the speaker clustering module: S31, the data processed by the speech segmentation module is input into the speaker embedding extraction network model after blocking to extract the speaker embedding vector, and a similarity matrix S is constructed by calculating the cosine similarity between each pair of embedding vectors; S32, spectral clustering is performed on it, the S is directly converted into an adjacency matrix W, the degree matrix D is calculated, and the symmetric normalized Laplacian matrix is constructed S33, to L sym Eigenvalue decomposition is performed, the number of speakers k is automatically estimated using the "eigenvalue gap", the first k eigenvectors are selected to form a matrix U, each row vector is normalized, and K-means is applied to divide them into k clusters. S34, by calculating the cosine similarity between the cluster centers or representative embeddings, if it exceeds the dynamic threshold, similar clusters are merged to further refine the speaker division; S35, assign a unique speaker ID to each cluster, and map the position of each embedded sliding window of the original audio to the corresponding ID.

6. The method of claim 5, wherein the method is for preprocessing language data in a multilingual and complex scenario. For continuous stream > 2 hours recording, the spectral clustering state can be reset every 2 hours: keep the history cluster centers, match and update the cluster assignment with newly generated embedding vectors to maintain the consistency of speaker ID, while cleaning up the old data.

7. The method for preprocessing language data in multiple languages and complex scenes according to claim 1, characterized in that: In the target speech extraction module: S41, the speech segments of adjacent and same speakers in the speaker clustering result are merged, and the merged speech segments are input into STFT to extract spectral features and provide frequency domain representation for subsequent processing; S42, combine the spectrum features obtained by STFT with the target speaker embedding as the input of the advanced speech separation network model in target speech extraction, to obtain M tgt (t, f) and M noiset (t, f); S43, M is used for single channel case tgt (t, f) is passed through The separated target speaker's frequency domain representation is calculated and input to the RNN post-filter The residual noise, reverberation, and frequency domain smoothing information are further modeled in the frequency domain, and then the iSTFT is performed to transform the frequency domain enhanced signal back to the time domain; S44, for the case of multiple channels by M noiset The covariance sum and the target direction estimate are computed and the MVDR beamforming formula is computed as follows: wherein is the noise covariance matrix, d(f) is the target direction vector, and w(f) represents the optimal weighting vector at frequency f, which is applied to the multi-channel spectrum, and the iSTFT is performed to transform the frequency-domain enhanced signal back to the time domain.

8. The method for preprocessing language data in multiple languages and complex scenes according to claim 1, characterized in that: The quality filtering module includes language adaptive scoring and anomaly detection technology: In language adaptive scoring, the scoring model uses differentiated thresholds; Anomaly detection detects anomalies by analyzing the energy change, spectral features and time domain waveform of the audio.

9. The method of claim 8, wherein the method is used for preprocessing language data in a multi-lingual and complex scenario. The quality filtering process includes: S51, receiving the speech segments from the target speech extraction module; S52, applying corresponding thresholds for quality scoring according to the language type; S53, performing anomaly detection to eliminate or repair unqualified segments; S54, outputting high-quality speech segments with quality score indicators, providing further screening and processing basis for downstream tasks.

Citation Information

Patent Citations

  • Voiceprint recognition method based on phoneme information and electronic equipment

    CN116403587A

  • Decoupling type voice self-supervision pre-training method

    CN118841029A

Cited By

  • Speech translation method and device of end cloud translation system, equipment and medium

    CN121438808A