A method and device for generating an ASR audio corpus based on a multi-modal large model

By generating and filtering audio corpora using a multimodal large model, the problems of high cost and insufficient diversity in audio corpus generation in existing technologies are solved, achieving the generation of high-quality corpora and the efficient adaptability of the ASR system in complex environments.

CN120340506BActive Publication Date: 2025-12-05ZHIMING RIXIN (NANJING) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510619927.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-12-05
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Existing methods for generating audio corpora are costly and struggle to cover all scenarios and speaker characteristics. Traditional text-to-speech technology produces speech that lacks the noise and diversity of real-world scenarios, resulting in poor adaptability of ASR systems in complex environments.

Method used

A multimodal large model is used to encode the target domain text into semantic vectors. Speech is generated by combining speech control parameters and speaker features. Noise annotation, text annotation, sentiment annotation and speaker annotation are performed by superimposing preset scene noise and adversarial noise. Word error rate threshold and semantic similarity threshold are set, and target corpus is selected from multimodal annotation files.

Benefits of technology

Generating high-quality ASR audio corpora that meet specific needs improves the matching degree and usability of the corpora with actual application scenarios, and enhances the robustness and recognition accuracy of the ASR system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340506B_ABST
    Figure CN120340506B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for generating ASR audio corpus based on a multimodal large model, and relates to the field of audio corpus. In the method, a semantic vector and a conditional vector are spliced into a joint vector to generate first speech; target noise is selected from a preset noise library according to a scene label, and the target noise is superimposed on the first speech to generate noise-bearing speech, and an adversarial noise is injected to generate second speech; the second speech is subjected to noise labeling, text labeling, emotion labeling and speaker labeling, and is aligned to generate a multimodal labeling file; according to the scene label, the noise type and the speaker information of the multimodal labeling file, a word error rate threshold and a semantic similarity threshold are set, and target corpus is selected from the multimodal labeling file according to the word error rate threshold and the semantic similarity threshold. By implementing the technical scheme provided in the application, high-quality audio corpus that meets specific requirements and is effectively screened can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio corpora, specifically to a method and apparatus for generating ASR audio corpora based on a multimodal large model. Background Technology

[0002] In the field of speech recognition technology, Automatic Speech Recognition (ASR) systems have made significant progress in recent years. With the continuous development of artificial intelligence technology, ASR systems have been widely used in many fields such as intelligent customer service, voice assistants, and smart homes, greatly improving the efficiency of people's information acquisition and interaction, and promoting the transformation of human-computer interaction methods towards a more natural and convenient direction.

[0003] In existing technologies, training high-quality ASR systems requires a large and diverse amount of audio corpus. Traditional audio corpus generation methods mainly fall into two categories. One is collecting speech data from real-world environments. This method requires significant human and material resources to record, organize, and annotate the speech data. For example, organizing a large-scale recording team to collect speech in different scenarios. The other is synthesizing speech through simple text-to-speech technology. This method mainly relies on rules or machine learning models to convert input text into speech output; however, the synthesized speech often lacks the noise and diversity of real-world scenarios.

[0004] However, existing methods for generating and annotating audio corpora have significant drawbacks. Collecting real-world speech data is costly and struggles to cover all scenarios and speaker characteristics, resulting in insufficient corpus diversity. Speech synthesized using simple text-to-speech techniques cannot simulate complex real-world scenarios and lacks noise and other interference factors, making the trained ASR system poorly adaptable to complex environments in practical applications. Summary of the Invention

[0005] This application provides a method and apparatus for generating ASR audio corpus based on a multimodal large model, which can generate high-quality ASR audio corpus that meets specific needs and has been effectively filtered, thereby improving the matching degree and usability of the corpus with actual application scenarios.

[0006] The first aspect of this application provides a method for generating ASR audio corpora based on a multimodal large model, applied to a corpus generation platform, the method comprising:

[0007] The target domain text is encoded into a semantic vector using a multimodal large model, the speech control parameters and speaker features are encoded into a conditional vector, the semantic vector and the conditional vector are concatenated into a joint vector, and the joint vector is decoded into the first speech.

[0008] Select target noise from a preset noise library based on scene tags, superimpose the target noise onto the first speech to generate noisy speech, and inject adversarial noise into the noisy speech to generate second speech;

[0009] The second speech is annotated with noise, text, emotion and speaker, and the annotated data is aligned to the same timeline to generate a multimodal annotation file;

[0010] Based on the scene labels, noise types, and speaker information of the multimodal annotation file, a word error rate threshold and a semantic similarity threshold are set, and target corpora are selected from the multimodal annotation file according to the word error rate threshold and the semantic similarity threshold.

[0011] Optionally, encoding the speech control parameters and speaker features into a conditional vector, concatenating the semantic vector and the conditional vector into a joint vector, and decoding the joint vector into the first speech includes:

[0012] Speech rate, pause interval, intonation and emotional intensity are mapped to a first vector, speaker identifier, dialect, timbre and age are mapped to a second vector, and the first vector and the second vector are concatenated to generate a conditional vector.

[0013] The semantic vector is assigned a first weight and the conditional vector a second weight through an attention mechanism. The semantic vector and the conditional vector are then concatenated along the channel dimension according to the first weight and the second weight to form a joint vector.

[0014] The joint vector is decoded frame by frame into a Mel spectrum, and the Mel spectrum is inversely transformed into a time-domain waveform to obtain the first speech.

[0015] Optionally, the step of selecting target noise from a preset noise library based on scene labels and superimposing the target noise onto the first speech to generate noisy speech includes:

[0016] Extract the semantic features of the scene labels, and match the target noise from the preset noise library based on the semantic features;

[0017] The energy weights of the first speech and the target noise are calculated based on the signal-to-noise ratio. The Mel spectra of the first speech and the target noise are aligned on the time axis and superimposed frame by frame according to the energy weights to obtain the first intermediate spectrum.

[0018] The first intermediate spectrum is converted into a time-domain signal by inverse Mel transform to obtain noisy speech.

[0019] Optionally, injecting adversarial noise into the noisy speech to generate the second speech includes:

[0020] Based on the gradient or decision boundary of the target ASR model, adversarial perturbations are generated, and the adversarial perturbations are converted into time-domain waveforms or spectral features to obtain adversarial noise.

[0021] The injection time window of the adversarial noise is located based on the speech content, and the adversarial noise and the noisy speech are superimposed in the Mel spectrum domain to obtain the second intermediate spectrum;

[0022] The second intermediate spectrum is subjected to inverse Mel transform to generate the second speech containing adversarial noise.

[0023] Optionally, the step of performing noise annotation, text annotation, sentiment annotation, and speaker annotation on the second speech, and aligning the annotation data to the same timeline to generate a multimodal annotation file includes:

[0024] The target frame of the adversarial noise in the second speech is located according to the injection control parameters of the adversarial noise, and noise annotation data is obtained according to the target frame;

[0025] An initial transcribed text is generated based on the speech content of the second speech. The first target speech segment, which is masked by scene noise or adversarial noise, is located based on the noise annotation data. The initial transcribed text of the first target speech segment is corrected by combining the multimodal large model with contextual semantics. The timestamp of the corrected initial transcribed text is aligned with the second speech using a dynamic time warping algorithm to obtain text annotation data.

[0026] An initial emotional result is generated based on the acoustic features of the second speech. The second target speech segment that is not masked by scene noise or adversarial noise is located based on the noise annotation data. The initial emotional result of the second target recording segment is corrected based on the corrected initial transcribed text. The timestamp of the corrected initial emotional result is aligned with the second speech using a dynamic time warping algorithm to obtain emotional annotation data.

[0027] Obtain the speaker embedding vector of the second speech, and perform similarity calculation and match the speaker identifier with a preset speaker database to obtain speaker annotation data;

[0028] A multimodal annotation file is generated based on the noise annotation data, the text annotation data, the sentiment annotation data, and the speaker annotation data.

[0029] Optionally, setting the word error rate threshold and semantic similarity threshold based on the scene labels, noise types, and speaker information of the multimodal annotation file includes:

[0030] Based on scene tags, match word error rate baseline thresholds and semantic similarity baseline thresholds from a preset mapping table;

[0031] The noise type in the multimodal annotation file is parsed, the signal-to-noise ratio value in the multimodal annotation file is mapped to the noise intensity level, and the matching word error rate baseline threshold and the semantic similarity baseline threshold are adjusted according to the noise type and the noise intensity level.

[0032] Based on the speaker identifier, the speaker attributes are queried from the preset speaker database. Based on the speaker attributes, the matching word error rate baseline threshold and the semantic similarity baseline threshold, which have undergone the first adjustment, are adjusted a second time to obtain the word error rate threshold and the semantic similarity threshold.

[0033] Optionally, the step of filtering target corpora from the multimodal annotation file based on the word error rate threshold and the semantic similarity threshold includes:

[0034] The transcribed text of each speech in the multimodal annotation file is time-aligned with a preset reference text, the number of errors is counted, and the word error rate is calculated based on the number of errors, including the number of substitution errors, the number of insertion errors, and the number of deletion errors.

[0035] Semantic encoding is performed on the transcribed text of each speech in the multimodal annotation file and a preset reference text, and cosine similarity is calculated;

[0036] The target corpus is determined from speech samples whose cosine similarity is greater than or equal to the semantic similarity threshold and whose word error rate is less than or equal to the word error rate threshold.

[0037] A second aspect of this application provides a system for generating ASR audio corpora based on a multimodal large model, including a vector module, a noise module, an annotation module, and an execution module, wherein:

[0038] The vector module is configured to encode target domain text into semantic vectors using a multimodal large model, encode speech control parameters and speaker features into conditional vectors, concatenate the semantic vectors and the conditional vectors into a joint vector, and decode the joint vector into first speech.

[0039] The noise module is configured to select target noise from a preset noise library based on scene labels, superimpose the target noise onto the first speech to generate noisy speech, and inject adversarial noise into the noisy speech to generate a second speech.

[0040] The annotation module is configured to perform noise annotation, text annotation, sentiment annotation, and speaker annotation on the second speech, and align the annotation data to the same timeline to generate a multimodal annotation file;

[0041] The execution module is configured to set a word error rate threshold and a semantic similarity threshold based on the scene labels, noise type, and speaker information of the multimodal annotation file, and to filter target corpora from the multimodal annotation file based on the word error rate threshold and the semantic similarity threshold.

[0042] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.

[0043] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.

[0044] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0045] 1. By encoding target domain text into semantic vectors using a multimodal large model, and combining speech control parameters and speaker features to generate speech, precise alignment of semantic and acoustic features is achieved. By concatenating semantic and conditional vectors, multi-dimensional control of speech is supported, enhancing the diversity of synthesized speech;

[0046] 2. By superimposing preset scene noise and adversarial noise (such as random perturbation signals), highly complex noisy speech is generated to simulate interference in real-world scenarios. Injecting adversarial noise forces the model to learn noise-insensitive features, enhancing its robustness to minor perturbations and reducing the false recognition rate in actual deployments.

[0047] 3. The synthesized speech is annotated with noise type, transcribed text, sentiment tags, speaker ID, etc., generating structured metadata to support multi-task learning. The annotated data is precisely aligned with the speech timeline to ensure a strict match between the transcribed text and the speech segments, improving data reliability;

[0048] 4. Based on scene labels, noise type, and speaker information, dynamically set word error rate thresholds and semantic similarity thresholds to select high-quality corpora from multimodal annotation files. Eliminate low-quality corpora to ensure the accuracy of training data. Prioritize retaining corpora strongly related to the target scene to improve model performance in specific domains. Reduce redundant data to lower model training costs. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating a method for generating ASR audio corpus based on a multimodal large model disclosed in an embodiment of this application;

[0050] Figure 2 This is a schematic diagram of a module of an ASR audio corpus generation system based on a multimodal large model disclosed in an embodiment of this application;

[0051] Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.

[0052] Explanation of reference numerals in the attached figures: 201, Vector module; 202, Noise module; 203, Labeling module; 204, Execution module; 301, Processor; 302, Communication bus; 303, User interface; 304, Network interface; 305, Memory. Detailed Implementation

[0053] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0054] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0055] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0056] This embodiment discloses a method for generating ASR audio corpora based on a multimodal large model, which is applied to a corpus generation platform. Figure 1 This is a flowchart illustrating a method for generating ASR audio corpus based on a multimodal large model disclosed in an embodiment of this application. Figure 1 As shown, the method includes the following steps:

[0057] S101. Use a multimodal large model to encode the target domain text into a semantic vector, encode the speech control parameters and speaker features into a conditional vector, concatenate the semantic vector and the conditional vector into a joint vector, and decode the joint vector into the first speech.

[0058] S102. Select target noise from a preset noise library according to scene tags, superimpose the target noise onto the first speech to generate noisy speech, and inject adversarial noise into the noisy speech to generate second speech;

[0059] S103. Perform noise annotation, text annotation, emotion annotation and speaker annotation on the second speech, and align the annotation data to the same timeline to generate a multimodal annotation file;

[0060] S104. Based on the scene labels, noise types, and speaker information of the multimodal annotation file, set a word error rate threshold and a semantic similarity threshold, and filter target corpus from the multimodal annotation file according to the word error rate threshold and the semantic similarity threshold.

[0061] The target domain text is encoded using a multimodal large model, transforming textual information into semantic vectors. This process captures the semantic meaning of the text, ensuring that the generated speech accurately reflects the text content. Speech control parameters (such as speech rate and intonation) and speaker features (such as timbre and accent) are encoded into conditional vectors. These conditional vectors provide additional control information for speech generation, making the generated speech more stylistically and stylistically appropriate for specific needs. The semantic and conditional vectors are concatenated into a joint vector, which integrates the text semantics and the conditional information of speech generation. The multimodal large model then decodes the joint vector to generate the first speech. Based on scene labels, target noise matching the specific scene is selected from a pre-defined noise library. For example, to simulate an outdoor environment, wind noise or traffic noise can be selected. The selected target noise is superimposed on the first speech to generate noisy speech. This step simulates noise interference in a real-world environment, making the generated speech closer to the actual application scenario. Finally, adversarial noise is injected into the noisy speech. The purpose of adversarial noise is to increase the difficulty of speech recognition and improve the robustness of the ASR model, enabling it to accurately recognize speech even in the face of noise interference. The second speech is annotated with noise (marking noise type and intensity), text (recording the corresponding text content), sentiment (identifying the emotional tendency in the speech), and speaker (determining the speaker's identity). These annotations provide rich information for subsequent corpus selection and application. The above annotated data are aligned to the same timeline, ensuring that each annotation accurately corresponds to a specific time period of the speech, generating a multimodal annotation file. This alignment facilitates subsequent comprehensive analysis and processing of the speech and annotation information. Based on the scene labels, noise types, and speaker information in the multimodal annotation file, word error rate thresholds and semantic similarity thresholds are set. The word error rate threshold measures the accuracy of the ASR system in speech recognition, while the semantic similarity threshold evaluates the semantic similarity between the recognition result and the original text. Based on the set word error rate and semantic similarity thresholds, target corpora that meet the requirements are selected from the multimodal annotation file. These target corpora can be used to train and evaluate the ASR model, improving its performance in complex scenarios.

[0062] Optionally, encoding the speech control parameters and speaker features into a conditional vector, concatenating the semantic vector and the conditional vector into a joint vector, and decoding the joint vector into the first speech includes:

[0063] Speech rate, pause interval, intonation and emotional intensity are mapped to a first vector, speaker identifier, dialect, timbre and age are mapped to a second vector, and the first vector and the second vector are concatenated to generate a conditional vector.

[0064] The semantic vector is assigned a first weight and the conditional vector a second weight through an attention mechanism. The semantic vector and the conditional vector are then concatenated along the channel dimension according to the first weight and the second weight to form a joint vector.

[0065] The joint vector is decoded frame by frame into a Mel spectrum, and the Mel spectrum is inversely transformed into a time-domain waveform to obtain the first speech.

[0066] Speech rate: This is usually measured by the number of syllables or words uttered per unit of time. For example, 120 words per minute is defined as a medium speech rate, mapped to a vector element value; while 80 words per minute is a slower speech rate, mapped to a different value. Pause interval: Based on punctuation marks (such as commas, periods, etc.) and natural semantic pauses in the text, the duration and location of pauses are statistically analyzed. For example, a pause after a period might be set to 1 second, mapped to a larger vector element value; a pause after a comma might be 0.5 seconds, mapped to a relatively smaller value. Intonation: Intonation reflects the fluctuations in speech, which can be quantified by analyzing the fundamental frequency of the speech signal. For example, a rise in fundamental frequency is represented as a positive number, and a fall as a negative number; different rises or falls correspond to different values, thus mapping to the first vector. Emotional intensity: Based on a pre-defined emotion classification model (such as categorizing emotions into happiness, sadness, anger, etc.), a numerical value is assigned to each emotion category according to the text content and the expected intensity of emotional expression. For example, strong feelings of happiness are mapped to higher values, while weaker feelings of happiness are mapped to lower values. After the above mapping, the values ​​corresponding to all parameters are combined into a vector, namely the first vector. This first vector contains the encoded information of the speech control parameters, providing a basis for controlling speech rate, pauses, intonation, etc. in subsequent speech generation. Speaker Identifier: A unique identifier is assigned to each speaker and directly mapped to a vector element. For example, speaker A is mapped to 1, speaker B to 2, etc. Dialect: A dialect encoding table is established according to different dialect types (such as Cantonese, Sichuan-Chongqing dialect, Wu dialect, etc.), mapping each dialect to a specific value or vector combination. For example, Cantonese is mapped to the vector [0, 1, 0], Sichuan-Chongqing dialect is mapped to [1, 0, 0], etc. Timbre: Timbre is the unique feature of a speaker's voice, which can be extracted by analyzing and extracting the spectral features of the speech signal (such as formant frequencies, spectral envelope, etc.), quantizing these features and mapping them to the second vector. For example, high-frequency formant prominence may correspond to higher vector element values, while low-frequency formant prominence corresponds to lower values. Age: Age is divided into different intervals (e.g., 0-10 years old, 11-20 years old, etc.), with each interval corresponding to a numerical value. For example, 0-10 years old is mapped to 1, 11-20 years old is mapped to 2, and so on. The numerical values ​​mapped from each speaker's features are combined into a second vector, which carries the encoded information of the speaker's features and is used to control the speaker's style and characteristics in the generated speech. The first and second vectors are concatenated to generate a conditional vector. The concatenation method is usually a simple vector join operation. For example, if the first vector is [a1, a2, ..., an] and the second vector is [b1, b2, ..., bm], then the concatenated conditional vector is [a1, a2, ..., an, b1, b2, ..., bm].This conditional vector integrates information from speech control parameters and speaker features, providing a foundation for subsequent joint vector generation.

[0067] By assigning weights to semantic and conditional vectors through an attention mechanism and concatenating them along the channel dimension to form a joint vector, the process achieves the organic integration of semantic and control information. The attention mechanism is a technique for automatically learning the importance of different information. During the generation of the joint vector, the attention mechanism analyzes the semantic and conditional vectors separately, assigning a first weight to the semantic vector and a second weight to the conditional vector based on their relevance to the final speech generation task. For the semantic and conditional vectors, query, key, and value vectors can be generated through linear transformations, respectively. For example, the semantic vector S undergoes linear transformations W1, W2, and W3 to obtain the query vector Qs=S×W1, the key vector Ks=S×W2, and the value vector Vs=S×W3; the conditional vector C undergoes linear transformations to obtain Qa=C×W1, Ka=C×W2, and Va=C×W3. The similarity between the query vector of the semantic vector and the key vector of the conditional vector, as well as the similarity between the query vector of the conditional vector and the key vector of the semantic vector, is calculated. Similarity can be calculated using methods such as dot product and cosine similarity. For example, the dot product of Qs and Ka yields the similarity score S1 = Qs·Ka, and the dot product of Qa and Ks yields the similarity score S2 = Qa·Ks. The similarity scores are normalized using the softmax function to obtain the first and second weights. For example, the first weight α = softmax(S1), and the second weight β = softmax(S2). Based on the assigned first weight α and second weight β, the semantic vector and conditional vector are concatenated along the channel dimension. During concatenation, the semantic vector and conditional vector can be combined in a certain proportion. For example, the joint vector U = α × S + β × C (the addition operation here is a weighted concatenation along the channel dimension). Other concatenation methods can also be used, such as directly connecting the semantic vector and conditional vector along the channel dimension and then performing a linear transformation according to the weights. Through this concatenation method, the joint vector integrates semantic and control information, providing a more comprehensive input for subsequent speech decoding.

[0068] The process of decoding the joint vector frame by frame into a Mel spectrum, and then inversely transforming the Mel spectrum into a time-domain waveform, ultimately yields the first speech. This process realizes the conversion from vector representation to audible speech. A neural network decoder (such as a decoder based on a recurrent neural network (RNN), long short-term memory (LSTM), or Transformer) is used to decode the joint vector frame by frame. The decoder receives the joint vector as input and generates each frame of the Mel spectrum progressively. During the decoding process, the decoder predicts the Mel spectrum features of the current frame based on the information in the joint vector and combined with previously generated frames. The Mel spectrum is a method of converting speech signals from the time domain to the frequency domain, simulating the human ear's perception of sound frequencies. Each frame of the Mel spectrum contains the frequency distribution information of the speech signal within a specific time period. The generated Mel spectrum is transformed back into a time-domain waveform through operations such as the inverse Mel filter bank and the inverse short-time Fourier transform (ISTFT). The inverse Mel filter bank transforms the Mel spectrum back into a linear spectrum, while ISTFT transforms the linear spectrum back into a time-domain signal. After the above inverse transformation, the speech signal in the time domain, i.e., the first speech, is obtained. This first speech combines the textual semantic information carried by the semantic vector with the speech control parameters and speaker feature information contained in the conditional vector, and has specific characteristics such as speech rate, intonation, emotional intensity, and speaker style.

[0069] By mapping speech control parameters such as speech rate, pause intervals, intonation, and emotional intensity to a first vector, and speaker features such as speaker identification, dialect, timbre, and age to a second vector, fine-grained encoding of core elements in the speech generation process is achieved. This provides rich and controllable feature inputs for subsequent generation of personalized and contextualized speech. Concatenating these two types of vectors to generate a conditional vector allows for flexible combination of different speech control parameters and speaker features, quickly adapting to diverse generation needs such as different speaker identities, emotional states, and speech styles, thus improving the flexibility and versatility of the speech generation system. Utilizing an attention mechanism to dynamically assign weights to the semantic and conditional vectors enables intelligent judgment of the relative importance of semantic and conditional information in the speech generation process based on the specific generation task and contextual information, giving the model adaptive capabilities when processing different types of text and scenarios. Based on the weight allocation results, the semantic and conditional vectors are concatenated along the channel dimension to form a joint vector, achieving deep fusion and joint modeling of semantic and conditional information at the model level. This contributes to generating semantically accurate, feature-rich, and expected speech. The joint vector is decoded frame by frame into a Mel spectrum. Utilizing this intermediate representation, the spectral features and details of the speech are effectively preserved, laying a solid foundation for subsequent waveform generation and ensuring high-quality spectral output. An inverse transform restores the Mel spectrum to a time-domain waveform, ultimately yielding the first speech. This process accurately restores the time-domain features of the speech, resulting in a natural, fluent, and distortion-free auditory experience, meeting the high speech quality requirements of practical applications.

[0070] Optionally, the step of selecting target noise from a preset noise library based on scene labels and superimposing the target noise onto the first speech to generate noisy speech includes:

[0071] Extract the semantic features of the scene labels, and match the target noise from the preset noise library based on the semantic features;

[0072] The energy weights of the first speech and the target noise are calculated based on the signal-to-noise ratio. The Mel spectra of the first speech and the target noise are aligned on the time axis and superimposed frame by frame according to the energy weights to obtain the first intermediate spectrum.

[0073] The first intermediate spectrum is converted into a time-domain signal by inverse Mel transform to obtain noisy speech.

[0074] The semantic feature extraction of scene labels utilizes natural language processing (NLP) technology to transform scene labels (such as "outdoor street," "factory workshop," and "noisy restaurant environment") into computer-understandable semantic vectors. These vectors contain core scene information, such as environment type and noise source characteristics, providing a semantic basis for subsequent noise matching. Based on the extracted semantic features, target noise is matched against a pre-set noise database. This database stores a large number of noise samples of different types and intensities, each with a semantic label. By calculating the similarity between the semantic features of the scene label and the semantic labels of the noise samples, the noise with the highest similarity is selected as the target noise, achieving precise matching between noise and scene and ensuring that the superimposed noise conforms to the characteristics of the actual scene. Signal-to-noise ratio (SNR) is an indicator that measures the energy relationship between speech and noise signals. Based on the set target SNR, the energy weights of the first speech and the target noise are calculated. The energy weights reflect the proportion of energy occupied by speech and noise in the mixed signal. By adjusting the energy weights, the relative intensity of speech and noise in the generated noisy speech can be controlled to meet the requirements of noise interference levels in different scenarios, such as simulating slight background noise or strong interference noise environments. Align the Mel spectra of the first speech and the target noise on the time axis. Mel spectrum is a method of converting speech signals to the frequency domain, which better matches the auditory perception characteristics of the human ear. Alignment ensures that the speech and noise are synchronized in the time dimension, facilitating subsequent frame-by-frame superposition. According to the calculated energy weights, the Mel spectra of the first speech and the target noise are added frame by frame to obtain the first intermediate spectrum. Frame-by-frame superposition allows precise control of the mixing ratio of speech and noise at each moment, ensuring the accuracy and continuity of the mixed spectrum in both time and frequency dimensions. The first intermediate spectrum is then converted into a time-domain signal using the inverse Mel transform. The inverse Mel transform is the reverse process of the Mel spectrum transform, restoring the frequency domain Mel spectrum information to a time-domain waveform signal. During the conversion, the energy distribution and frequency characteristics of the speech and noise after spectral superposition are preserved, ensuring that the restored time-domain signal accurately reflects the sound characteristics of the mixed speech and noise. The time-domain signal obtained after the inverse Mel transform is the noisy speech. Because the energy weights are precisely controlled during the spectrum superposition process, and accurate alignment and processing are performed in the time and frequency domains, the generated noisy speech sounds natural and realistic. The mixing effect of noise and speech meets the needs of real-world scenarios, providing high-quality training or testing data for subsequent speech recognition, speech enhancement and other tasks.

[0075] The generated noisy speech can simulate various real-world noise interference scenarios, providing rich noisy training data for speech processing models (such as ASR models). This exposes the models to more diverse noise environments during training, thereby improving their robustness and generalization ability in practical applications. Through semantic matching of scene labels and flexible adjustment of the signal-to-noise ratio, noisy speech with specific noise characteristics and interference levels can be customized to meet the diverse and realistic speech data requirements of different industries (such as intelligent customer service, security monitoring, and in-vehicle voice interaction). Based on a pre-set noise library and automated processing flow, large amounts of noisy speech data can be generated quickly and efficiently, avoiding the high cost and low efficiency of manually collecting and labeling noisy speech data in real-world scenarios. This provides a convenient data generation method for speech technology research and application development.

[0076] Optionally, injecting adversarial noise into the noisy speech to generate the second speech includes:

[0077] Based on the gradient or decision boundary of the target ASR model, adversarial perturbations are generated, and the adversarial perturbations are converted into time-domain waveforms or spectral features to obtain adversarial noise.

[0078] The injection time window of the adversarial noise is located based on the speech content, and the adversarial noise and the noisy speech are superimposed in the Mel spectrum domain to obtain the second intermediate spectrum;

[0079] The second intermediate spectrum is subjected to inverse Mel transform to generate the second speech containing adversarial noise.

[0080] The core of adversarial noise generation lies in leveraging the characteristics of the target ASR (Automatic Speech Recognition) model. Once an ASR model is trained to a certain stage, it possesses an internal mechanism for classifying or transforming input speech features. This mechanism can abstract decision boundaries (used to distinguish different speech categories) or reflect the impact of input changes on the output through gradient information. Based on gradients: During model training, the gradient represents the rate of change of the loss function with respect to the model parameters. When generating adversarial perturbations, the gradient can be seen as the direction in the input speech feature space that causes the model output to change most significantly. By calculating the gradient of the target ASR model with respect to the input speech and making small adjustments to the speech features along the gradient direction, adversarial perturbations that cause the model output to misrecognize can be generated. Based on decision boundaries: The decision boundary is the boundary by which the model distinguishes different speech categories. By analyzing the distribution of speech features near the decision boundary, small perturbations that can cross the boundary can be found. These perturbations, when added to the original speech, cause the speech to "slip" from the originally correctly recognized category into the incorrect category, thus forming adversarial perturbations. In practical implementation, iterative optimization algorithms such as Fast Signed Gradient Method (FGSM) and Basic Iterative Method (BIM) can be used. Taking FGSM as an example, it calculates the gradient of the loss function with respect to the input speech, signs the gradient, multiplies it by a small step size factor, and adds the result to the original speech features to obtain adversarial perturbations. These perturbations are usually small and imperceptible to humans, but can have a significant impact on ASR models. The generated adversarial perturbations may initially exist in the form of speech features (such as Mel frequency cepstral coefficients). To add them to noisy speech, they need to be converted into time-domain waveforms. This process is achieved through inverse transformations, such as using the inverse process of the Mel filter bank or the inverse discrete cosine transform, to map the perturbation from the feature space back to the time domain, obtaining the adversarial noise waveform in the time domain. Alternatively, the adversarial perturbation can be converted into spectral features, such as the Mel spectrum. Through specific transformation matrices or algorithms, the perturbation is mapped to the Mel spectral domain, making it adversarial noise in the form of spectral features that can be directly superimposed on the speech Mel spectrum. Determining the injection time window of the adversarial noise based on the speech content is crucial. This requires speech activity detection and semantic analysis of noisy speech. Speech activity detection identifies valid speech segments and silent segments in the speech, while semantic analysis further determines key semantic units (such as words and phrases) and their temporal locations. Injecting adversarial noise near key semantic units or during time periods that significantly impact ASR model recognition can more effectively interfere with the model's recognition process. For example, injecting noise during the pronunciation of important words in speech may cause the model to misidentify those words, leading to errors in the recognition of the entire sentence. The generated adversarial noise (whether in the spectral form after time-domain waveform transformation or directly generated spectral features) is then aligned with the Mel spectrum of the noisy speech on the time axis.To ensure synchronization between the two in the time dimension, subsequent superposition operations can accurately add noise to the corresponding positions in the speech. In the Mel spectrum domain, adversarial noise is added frame by frame to the Mel spectrum of the noisy speech according to certain rules (such as linear superposition) to obtain the second intermediate spectrum. Frame by frame superposition ensures the accuracy of the superposition process, making the mixing effect of noise and speech in the spectrum as expected, and avoiding anomalies caused by time or frequency misalignment. The second intermediate spectrum is then subjected to an inverse Mel transform to convert it from the Mel spectrum domain back to the time domain signal. The inverse Mel transform process is similar to the conversion from the Mel spectrum to the time domain waveform when generating noisy speech. Through inverse operations of the Mel filter bank, inverse discrete cosine transform, etc., the spectral information is restored to the time domain waveform signal. The time domain signal obtained after the inverse Mel transform is the second speech containing adversarial noise. Because the injection position and intensity of noise are precisely controlled during the spectrum superposition process, the generated second speech may not have obvious audible abnormalities, but it will produce incorrect recognition results in ASR model recognition due to the presence of adversarial noise, thus achieving the purpose of testing and enhancing the robustness of the ASR model.

[0081] Generated adversarial noise-infused speech can simulate potential malicious attacks in real-world scenarios, helping researchers and developers understand the weaknesses of ASR models when facing adversarial interference. This allows for targeted improvements and optimizations, enhancing the model's stability and reliability in practical applications. Introducing adversarial noise-infused speech data as training samples during ASR model training enables the model to learn more robust feature representations, strengthening its resistance to noise interference. Furthermore, using this type of data for testing during model evaluation allows for more accurate assessment of the model's performance in complex environments. Adversarial noise injection techniques provide an important tool for research in the field of speech security. Researchers can analyze the impact of different types of adversarial noise on ASR models, explore effective attack strategies, and further study corresponding defense methods, such as developing adversarial noise detection algorithms and improving model architecture to enhance robustness, thus driving the development of speech security technology. Generating adversarial noise-infused speech data expands the training dataset of ASR models, increasing data diversity and complexity. This helps the model learn more comprehensive speech features, reduces overfitting, and improves the model's generalization ability on unseen data.

[0082] Optionally, the step of performing noise annotation, text annotation, sentiment annotation, and speaker annotation on the second speech, and aligning the annotation data to the same timeline to generate a multimodal annotation file includes:

[0083] The target frame of the adversarial noise in the second speech is located according to the injection control parameters of the adversarial noise, and noise annotation data is obtained according to the target frame;

[0084] An initial transcribed text is generated based on the speech content of the second speech. The first target speech segment, which is masked by scene noise or adversarial noise, is located based on the noise annotation data. The initial transcribed text of the first target speech segment is corrected by combining the multimodal large model with contextual semantics. The timestamp of the corrected initial transcribed text is aligned with the second speech using a dynamic time warping algorithm to obtain text annotation data.

[0085] An initial emotional result is generated based on the acoustic features of the second speech. The second target speech segment that is not masked by scene noise or adversarial noise is located based on the noise annotation data. The initial emotional result of the second target recording segment is corrected based on the corrected initial transcribed text. The timestamp of the corrected initial emotional result is aligned with the second speech using a dynamic time warping algorithm to obtain emotional annotation data.

[0086] Obtain the speaker embedding vector of the second speech, and perform similarity calculation and match the speaker identifier with a preset speaker database to obtain speaker annotation data;

[0087] A multimodal annotation file is generated based on the noise annotation data, the text annotation data, the sentiment annotation data, and the speaker annotation data.

[0088] When generating the second speech, adversarial noise injection is performed based on specific control parameters, such as the injection time window and intensity. Using these injection control parameters, the target frames of the adversarial noise in the second speech can be accurately located. For example, if the control parameters specify that adversarial noise is injected starting at the 5th second of the speech and lasting for 2 seconds, then the frame number range corresponding to this noise can be calculated based on the speech sampling rate and frame length, thus determining the target frame. Based on the located target frame, noise annotation data is generated. The annotation data can include information such as noise type (scene noise, adversarial noise), the start and end frame numbers of the noise, and the noise intensity. This information provides a foundation for subsequent analysis of the impact of noise on tasks such as speech recognition and emotion recognition. Using speech recognition technology, the speech content of the second speech is converted into initial transcribed text. This is usually achieved through an automatic speech recognition model, which converts the speech signal into a corresponding text sequence. However, due to the presence of scene noise and adversarial noise in the second speech, the initial transcribed text may contain errors. Based on the noise annotation data, the first target speech segment masked by scene noise or adversarial noise is located. In these speech segments, the speech signal is severely interfered with by noise, making speech recognition difficult, and errors in the initial transcribed text often concentrate here. A multimodal large model is used, incorporating contextual semantic information from the first target speech segment, to correct the initial transcribed text. The multimodal large model can comprehensively consider semantic, syntactic, and contextual factors of speech, inferring the possible content of the masked speech segment by analyzing the content of surrounding normal speech segments, thereby optimizing the initial transcribed text. The Dynamic Time Warping (DTW) algorithm is used to align the timestamp of the corrected initial transcribed text with the second speech. The DTW algorithm can handle the scaling changes of the speech signal on the time axis, finding the optimal alignment path between text and speech, ensuring that each character in the text annotation data accurately corresponds to its corresponding position in the speech, resulting in accurate text annotation data. Based on the acoustic features of the second speech, such as pitch, speech rate, and energy, an emotion recognition model is used to generate initial emotion results. These acoustic features are closely related to human emotional states; for example, high pitch and rapid speech rate may indicate excitement or tension. Emotion recognition models learn from a large amount of emotion-labeled speech data to establish a mapping relationship between acoustic features and emotion categories, thereby classifying new speech into emotions. Based on noise-labeled data, a second target speech segment unmasked by scene noise or adversarial noise is located. In these segments, the speech signal is relatively clear, allowing the emotion recognition model to extract acoustic features and determine the emotional state more accurately. The initial emotion result for the second target speech segment is further refined based on the corrected initial transcribed text. Text content can provide additional semantic information for emotion judgment; for example, certain words or phrases themselves carry obvious emotional tendencies. By combining textual information, a more comprehensive understanding of the emotions expressed by the speech can be achieved, improving the accuracy of emotion labeling.Similarly, a dynamic time warping algorithm is used to align the timestamps of the corrected initial sentiment results with the second speech. This ensures that the sentiment annotation data accurately reflects the emotional changes in the speech over different time periods, generating sentiment annotation data consistent with the speech timeline. Speaker embedding vectors for the second speech are obtained using speaker recognition technology. A speaker embedding vector is an abstract representation of a speaker's acoustic features, containing unique information such as timbre and pronunciation habits. Deep learning models, such as speaker verification models, are typically used to extract these embedding vectors from the speech signal. The speaker embedding vectors are then compared to a pre-defined speaker database for similarity calculation. This database stores embedding vectors from multiple known speakers. By calculating the similarity between the two (e.g., cosine similarity), the speaker identifier most similar to the second speech's embedding vector is found, thus identifying the speaker of the second speech. The matched speaker identifier is used as speaker annotation data to annotate the second speech with speaker identity information, facilitating subsequent analysis and processing of different speakers' speech. Finally, information from different dimensions—noise annotation data, text annotation data, sentiment annotation data, and speaker annotation data—is integrated together. Each labeled data point includes timestamp information to ensure alignment and correlation along the same timeline. The integrated data is then used to construct a multimodal annotation file in a specific file format. This file records information such as noise type, text content, emotional state, and speaker identity at each time point in the speech, providing comprehensive and detailed annotation information for speech processing tasks. The multimodal annotation file can use structured data formats such as JSON and XML for easy subsequent reading and processing.

[0089] Multimodal annotation files provide rich annotation information for tasks such as speech recognition, emotion recognition, and speaker recognition. Information from different modalities can complement and verify each other, thereby improving the accuracy and robustness of these tasks. For example, in speech recognition, combining emotion annotation information can better understand the semantics and context of speech, reducing errors caused by noise interference. In emotion recognition, text annotation information can provide semantic support for emotion judgment, improving the accuracy of emotion recognition. It also provides detailed data support for academic research in the field of speech. Researchers can use multimodal annotation files to analyze the recognition performance, emotional expression characteristics, and speaker feature changes under different noise environments, promoting the development and innovation of speech technology. In practical voice application development, such as intelligent customer service, voice assistants, and voice security, multimodal annotation files can help developers better understand and process user voice input, achieving more intelligent and personalized voice interaction functions. For example, in intelligent customer service, emotion annotation information can be used to adjust customer service response strategies in a timely manner, improving user satisfaction.

[0090] Optionally, setting the word error rate threshold and semantic similarity threshold based on the scene labels, noise types, and speaker information of the multimodal annotation file includes:

[0091] Based on scene tags, match word error rate baseline thresholds and semantic similarity baseline thresholds from a preset mapping table;

[0092] The noise type in the multimodal annotation file is parsed, the signal-to-noise ratio value in the multimodal annotation file is mapped to the noise intensity level, and the matching word error rate baseline threshold and the semantic similarity baseline threshold are adjusted according to the noise type and the noise intensity level.

[0093] Based on the speaker identifier, the speaker attributes are queried from the preset speaker database. Based on the speaker attributes, the matching word error rate baseline threshold and the semantic similarity baseline threshold, which have undergone the first adjustment, are adjusted a second time to obtain the word error rate threshold and the semantic similarity threshold.

[0094] Scene labels describe the general application scenario of the speech, such as "quiet indoors," "noisy street," "conference room," and "in a car." The interference factors and noise levels experienced by the speech vary significantly across different scenarios, thus affecting the performance of the speech recognition system, such as the word error rate (WER) and the accuracy of semantic understanding (which can be measured by semantic similarity). A pre-defined mapping table stores the baseline thresholds for the word error rate and semantic similarity for different scene labels. These baseline thresholds reflect the minimum standards that a well-performing speech recognition and semantic understanding system should achieve in the corresponding scenario. For example, in a quiet indoor scenario, due to less noise interference and higher speech clarity, the baseline threshold for the word error rate can be set lower (e.g., 5%), while the baseline threshold for semantic similarity can be set higher (e.g., 0.9). In a noisy street scenario, due to greater background noise and increased speech recognition difficulty, the baseline threshold for the word error rate will increase accordingly (e.g., 15%), while the baseline threshold for semantic similarity will decrease (e.g., 0.7). Based on scene labels in the multimodal annotation file, the corresponding word error rate baseline threshold and semantic similarity baseline threshold are quickly matched from a preset mapping table, providing a basis for subsequent threshold adjustments. Signal-to-noise ratio (SNR) is the ratio of speech signal power to noise signal power, usually expressed in decibels (dB). It reflects the relative intensity of effective speech components and noise components in a speech signal; a higher SNR indicates better speech quality and less noise interference. Different SNR ranges are mapped to predefined noise intensity levels. For example, an SNR greater than 20dB corresponds to a "low noise level," an SNR between 10dB and 20dB corresponds to a "medium noise level," and an SNR less than 10dB corresponds to a "high noise level." This classification method more intuitively describes the degree of noise interference to speech. Different types of noise affect speech recognition and semantic understanding through different mechanisms. For example, white noise is a random noise, uniformly distributed across all frequencies, and will uniformly interfere with all frequency bands of the speech signal; while sudden noise (such as the sound of a door closing or a car horn) will cause strong interference to speech at a specific point in time, leading to local distortion of the speech signal. Based on the noise type and intensity level, the baseline thresholds for word error rate and semantic similarity are adjusted. Generally, the higher the noise intensity level or the more severe the noise interference with speech, the higher the baseline threshold for word error rate and the lower the baseline threshold for semantic similarity. For example, in a high-noise-level white noise scenario, the baseline threshold for word error rate might increase by 5% and the baseline threshold for semantic similarity might decrease by 0.1; while in a medium-noise-level scenario with sudden noise, the baseline threshold for word error rate might increase by 3% and the baseline threshold for semantic similarity might decrease by 0.05. This adjustment method more accurately reflects the difficulty of speech recognition and semantic understanding under different noise conditions. The preset speaker database stores relevant information for multiple known speakers, including attributes such as age, gender, accent, and speech rate.These attributes influence the characteristics of speech signals, thereby affecting the performance of speech recognition and semantic understanding. For example, people of different ages have different pronunciation characteristics; children and the elderly may have unclear pronunciation, slow or fast speech rates, etc.; people from different regions may have different accents, making it difficult for speech recognition systems to recognize certain pronunciations. Based on the speaker identifiers in the multimodal annotation file, the corresponding speaker attribute information is retrieved from a pre-set speaker database. Different speaker attributes affect the accuracy of speech recognition and semantic understanding; therefore, a second adjustment is needed to the word error rate baseline threshold and semantic similarity baseline threshold after the first adjustment, based on the speaker attributes. For example, for speakers with heavy accents, the speech recognition system will find it more difficult to recognize their speech, so the word error rate baseline threshold needs to be further increased; while for speakers with fast speech rates, the system may not be able to accurately recognize each word in time, which will also lead to an increase in the word error rate and the difficulty of semantic understanding, so the semantic similarity baseline threshold needs to be appropriately reduced. The specific adjustment range can be set based on extensive experimental data and experience. For example, for speakers with noticeable accents, the word error rate baseline threshold can be increased by 2% on top of the first adjustment, while the semantic similarity baseline threshold can be decreased by 0.03. After baseline threshold matching based on scene labels, the first adjustment combining noise type and intensity level, and the second adjustment based on speaker attributes, the final word error rate threshold and semantic similarity threshold are obtained. These two thresholds comprehensively consider the scene, noise level, and speaker characteristics, enabling a more accurate evaluation of the performance of speech recognition and semantic understanding systems. In practical applications, the word error rate output by the speech recognition system can be compared with the set word error rate threshold, and the semantic similarity output by the semantic understanding system can be compared with the set semantic similarity threshold to determine whether the system's performance meets the standards under the current speech input. If the word error rate is below the threshold and the semantic similarity is above the threshold, the system performance is good; otherwise, the system needs optimization and improvement. Setting different thresholds based on different speaker attributes can provide personalized voice services for different users. For example, for users with heavy accents, appropriately relaxing the word error rate threshold while lowering the semantic similarity threshold can improve the system's ability to recognize and understand the user's speech, thereby enhancing the user experience.

[0095] Optionally, the step of filtering target corpora from the multimodal annotation file based on the word error rate threshold and the semantic similarity threshold includes:

[0096] The transcribed text of each speech in the multimodal annotation file is time-aligned with a preset reference text, the number of errors is counted, and the word error rate is calculated based on the number of errors, including the number of substitution errors, the number of insertion errors, and the number of deletion errors.

[0097] Semantic encoding is performed on the transcribed text of each speech in the multimodal annotation file and a preset reference text, and cosine similarity is calculated;

[0098] The target corpus is determined from speech samples whose cosine similarity is greater than or equal to the semantic similarity threshold and whose word error rate is less than or equal to the word error rate threshold.

[0099] Time-aligning the transcribed text of each piece of speech in the multimodal annotation file with a preset reference text is to compare the content of the two under the same semantic and syntactic structure framework, so as to accurately evaluate the transcription accuracy of the speech recognition system for speech content. For example, a piece of speech may contain multiple sentences and words. Time alignment can ensure that each word and sentence in the transcribed text and the reference text corresponds to the corresponding time segment of the speech, avoiding inaccurate error statistics caused by time misalignment. The dynamic time warping algorithm can be used to achieve time alignment. Specifically, for speech signals, acoustic features such as Mel-frequency cepstral coefficients can be extracted first, and the transcribed text and the reference text are aligned in terms of feature sequences according to the pronunciation time points of the speech, laying a foundation for subsequent error statistics. The number of errors includes substitution errors, insertion errors, and deletion errors. A substitution error means that a certain word in the transcribed text is misrecognized as another word. For example, "apple" is recognized as "pingguo"; an insertion error means that there are words in the transcribed text that do not exist in the reference text. For example, the reference text is "I like to eat fruits", and the transcribed text is "I like to eat fruits and apples" (the extra "and apples"); a deletion error means that a certain word in the reference text is missing in the transcribed text. For example, the reference text is "I went for a walk in the park today", and the transcribed text is "I went to the park today" (missing "for a walk"). Based on time alignment, the transcribed text and the reference text are compared word by word. When a word mismatch is found, the error is counted according to the error type. For example, three counters are set to record the number of substitution errors, insertion errors, and deletion errors respectively. The entire text sequence is traversed, and for each error type encountered, the corresponding counter is incremented by 1. The ratio of the sum of the number of substitution errors, insertion errors, and deletion errors to the total number of words in the reference text is used as the word error rate. Semantic encoding of the transcribed text of each piece of speech in the multimodal annotation file with a preset reference text is to convert the text into a vector form that can represent its semantic information. Taking the BERT (Bidirectional Encoder Representations from Transformers) model as an example, the transcribed text and the reference text are respectively input into the BERT model to obtain their corresponding semantic vectors. Before inputting into the model, preprocessing operations such as word segmentation and adding special tokens need to be performed on the text. The semantic vectors output by the model can capture semantic relationships, context information, etc. in the text, providing a basis for subsequent similarity calculation. Cosine similarity is an index to measure the similarity between two vectors, and its value range is between [-1, 1]. For semantic vectors, the greater the cosine similarity, the more similar the semantics of the two texts. Calculating the cosine similarity between two vectors is achieved by calculating the cosine value of the angle between them. The word error rate threshold and the semantic similarity threshold are preset according to actual application requirements and system performance requirements.Word error rate (BER) thresholds measure the accuracy of speech recognition systems in transcribing speech content, while semantic similarity thresholds assess the semantic consistency between the transcribed text and the reference text. For example, in a voice interaction system with high accuracy requirements, the BER threshold might be set at 10%, and the semantic similarity threshold at 0.85. When selecting target corpora from multimodal annotation files, two conditions must be met simultaneously: cosine similarity greater than or equal to the semantic similarity threshold, and the BER less than or equal to the BER threshold. Only speech samples that meet both conditions are considered high-quality corpora. The selected target corpora can be used to further optimize speech recognition and semantic understanding models. Because these corpora possess high accuracy and semantic consistency, adding them to the training dataset improves the model's adaptability to different speech scenarios, speaker characteristics, and noise interference, thereby enhancing model performance. In different business scenarios, target corpora can serve as a basis for evaluating and selecting speech processing systems. For example, in intelligent customer service systems, analyzing target corpora can assess the system's understanding and response capabilities to user voice commands, providing a reference for system optimization and upgrades. Meanwhile, the target corpus can also be used to construct speech datasets for specific fields, meeting the speech processing needs of different industries.

[0100] This embodiment also discloses a system for generating ASR audio corpora based on a multimodal large model. Figure 2 This is a schematic diagram of a module of an ASR audio corpus generation system based on a multimodal large model disclosed in an embodiment of this application, such as... Figure 2 As shown, the system includes a vector module 201, a noise module 202, a labeling module 203, and an execution module 204, wherein:

[0101] Vector module 201 is configured to use a multimodal large model to encode target domain text into semantic vectors, encode speech control parameters and speaker features into conditional vectors, concatenate the semantic vectors and the conditional vectors into a joint vector, and decode the joint vector into first speech;

[0102] Noise module 202 is configured to select target noise from a preset noise library according to scene labels, superimpose the target noise onto the first speech to generate noisy speech, and inject adversarial noise into the noisy speech to generate a second speech;

[0103] The annotation module 203 is configured to perform noise annotation, text annotation, emotion annotation and speaker annotation on the second speech, and align the annotation data to the same timeline to generate a multimodal annotation file;

[0104] The execution module 204 is configured to set a word error rate threshold and a semantic similarity threshold based on the scene labels, noise types and speaker information of the multimodal annotation file, and to filter target corpus from the multimodal annotation file based on the word error rate threshold and the semantic similarity threshold.

[0105] Optionally, the vector module 201 is configured to:

[0106] Speech rate, pause interval, intonation and emotional intensity are mapped to a first vector, speaker identifier, dialect, timbre and age are mapped to a second vector, and the first vector and the second vector are concatenated to generate a conditional vector.

[0107] The semantic vector is assigned a first weight and the conditional vector a second weight through an attention mechanism. The semantic vector and the conditional vector are then concatenated along the channel dimension according to the first weight and the second weight to form a joint vector.

[0108] The joint vector is decoded frame by frame into a Mel spectrum, and the Mel spectrum is inversely transformed into a time-domain waveform to obtain the first speech.

[0109] Optionally, the noise module 202 is configured to:

[0110] Extract the semantic features of the scene labels, and match the target noise from the preset noise library based on the semantic features;

[0111] The energy weights of the first speech and the target noise are calculated based on the signal-to-noise ratio. The Mel spectra of the first speech and the target noise are aligned on the time axis and superimposed frame by frame according to the energy weights to obtain the first intermediate spectrum.

[0112] The first intermediate spectrum is converted into a time-domain signal by inverse Mel transform to obtain noisy speech.

[0113] Optionally, the noise module 202 is configured to:

[0114] Based on the gradient or decision boundary of the target ASR model, adversarial perturbations are generated, and the adversarial perturbations are converted into time-domain waveforms or spectral features to obtain adversarial noise.

[0115] The injection time window of the adversarial noise is located based on the speech content, and the adversarial noise and the noisy speech are superimposed in the Mel spectrum domain to obtain the second intermediate spectrum;

[0116] The second intermediate spectrum is subjected to inverse Mel transform to generate the second speech containing adversarial noise.

[0117] Optionally, the annotation module 203 is configured to:

[0118] The target frame of the adversarial noise in the second speech is located according to the injection control parameters of the adversarial noise, and noise annotation data is obtained according to the target frame;

[0119] An initial transcribed text is generated based on the speech content of the second speech. The first target speech segment, which is masked by scene noise or adversarial noise, is located based on the noise annotation data. The initial transcribed text of the first target speech segment is corrected by combining the multimodal large model with contextual semantics. The timestamp of the corrected initial transcribed text is aligned with the second speech using a dynamic time warping algorithm to obtain text annotation data.

[0120] An initial emotional result is generated based on the acoustic features of the second speech. The second target speech segment that is not masked by scene noise or adversarial noise is located based on the noise annotation data. The initial emotional result of the second target recording segment is corrected based on the corrected initial transcribed text. The timestamp of the corrected initial emotional result is aligned with the second speech using a dynamic time warping algorithm to obtain emotional annotation data.

[0121] Obtain the speaker embedding vector of the second speech, and perform similarity calculation and match the speaker identifier with a preset speaker database to obtain speaker annotation data;

[0122] A multimodal annotation file is generated based on the noise annotation data, the text annotation data, the sentiment annotation data, and the speaker annotation data.

[0123] Optionally, the execution module 204 is configured to:

[0124] Based on scene tags, match word error rate baseline thresholds and semantic similarity baseline thresholds from a preset mapping table;

[0125] The noise type in the multimodal annotation file is parsed, the signal-to-noise ratio value in the multimodal annotation file is mapped to the noise intensity level, and the matching word error rate baseline threshold and the semantic similarity baseline threshold are adjusted according to the noise type and the noise intensity level.

[0126] Based on the speaker identifier, the speaker attributes are queried from the preset speaker database. Based on the speaker attributes, the matching word error rate baseline threshold and the semantic similarity baseline threshold, which have undergone the first adjustment, are adjusted a second time to obtain the word error rate threshold and the semantic similarity threshold.

[0127] Optionally, the execution module 204 is configured to:

[0128] The transcribed text of each speech in the multimodal annotation file is time-aligned with a preset reference text, the number of errors is counted, and the word error rate is calculated based on the number of errors, including the number of substitution errors, the number of insertion errors, and the number of deletion errors.

[0129] Semantic encoding is performed on the transcribed text of each speech in the multimodal annotation file and a preset reference text, and cosine similarity is calculated;

[0130] The target corpus is determined from speech samples whose cosine similarity is greater than or equal to the semantic similarity threshold and whose word error rate is less than or equal to the word error rate threshold.

[0131] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided above belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0132] This embodiment also discloses an electronic device, referring to... Figure 3 The electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.

[0133] The communication bus 302 is used to enable communication between these components.

[0134] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0135] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0136] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0137] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. Figure 3 As shown, the memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for generating ASR audio corpus based on a multimodal large model.

[0138] exist Figure 3In the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 301 can be used to call the application stored in the memory 305 for a method of generating ASR audio corpus based on a multimodal large model. When executed by one or more processors 301, the electronic device executes one or more methods as described in the above embodiments.

[0139] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0140] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.

[0142] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0143] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0144] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0145] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for generating ASR audio corpus based on a multimodal large model, characterized in that, Applied to a corpus generation platform, the method includes: The target domain text is encoded into a semantic vector using a multimodal large model, the speech control parameters and speaker features are encoded into a conditional vector, the semantic vector and the conditional vector are concatenated into a joint vector, and the joint vector is decoded into the first speech. Select target noise from a preset noise library based on scene tags, superimpose the target noise onto the first speech to generate noisy speech, and inject adversarial noise into the noisy speech to generate second speech; The second speech is annotated with noise, text, emotion and speaker, and the annotated data is aligned to the same timeline to generate a multimodal annotation file; Based on the scene labels, noise types, and speaker information of the multimodal annotation files, a word error rate threshold and a semantic similarity threshold are set. Target corpora are then selected from the multimodal annotation files based on these thresholds. The step of encoding speech control parameters and speaker features into a conditional vector, concatenating the semantic vector and the conditional vector into a joint vector, and decoding the joint vector into the first speech includes: Speech rate, pause interval, intonation and emotional intensity are mapped to a first vector, speaker identifier, dialect, timbre and age are mapped to a second vector, and the first vector and the second vector are concatenated to generate a conditional vector. The semantic vector is assigned a first weight and the conditional vector a second weight through an attention mechanism. The semantic vector and the conditional vector are then concatenated along the channel dimension according to the first weight and the second weight to form a joint vector. The joint vector is decoded frame by frame into a Mel spectrum, and the Mel spectrum is inversely transformed into a time-domain waveform to obtain the first speech.

2. The method for generating ASR audio corpus based on a multimodal large model according to claim 1, characterized in that, The step of selecting target noise from a preset noise library based on scene labels and superimposing the target noise onto the first speech to generate noisy speech includes: Extract the semantic features of the scene labels, and match the target noise from the preset noise library based on the semantic features; The energy weights of the first speech and the target noise are calculated based on the signal-to-noise ratio. The Mel spectra of the first speech and the target noise are aligned on the time axis and superimposed frame by frame according to the energy weights to obtain the first intermediate spectrum. The first intermediate spectrum is converted into a time-domain signal by inverse Mel transform to obtain noisy speech.

3. The method for generating ASR audio corpus based on a multimodal large model according to claim 1, characterized in that, The step of injecting adversarial noise into the noisy speech to generate a second speech includes: Based on the gradient or decision boundary of the target ASR model, adversarial perturbations are generated, and the adversarial perturbations are converted into time-domain waveforms or spectral features to obtain adversarial noise. The injection time window of the adversarial noise is located based on the speech content, and the adversarial noise and the noisy speech are superimposed in the Mel spectrum domain to obtain the second intermediate spectrum; The second intermediate spectrum is subjected to inverse Mel transform to generate the second speech containing adversarial noise.

4. The method for generating ASR audio corpus based on a multimodal large model according to claim 3, characterized in that, The step of performing noise annotation, text annotation, sentiment annotation, and speaker annotation on the second speech, and aligning the annotation data to the same timeline to generate a multimodal annotation file includes: The target frame of the adversarial noise in the second speech is located according to the injection control parameters of the adversarial noise, and noise annotation data is obtained according to the target frame; An initial transcribed text is generated based on the speech content of the second speech. The first target speech segment, which is masked by scene noise or adversarial noise, is located based on the noise annotation data. The initial transcribed text of the first target speech segment is corrected by combining the multimodal large model with contextual semantics. The timestamp of the corrected initial transcribed text is aligned with the second speech using a dynamic time warping algorithm to obtain text annotation data. An initial emotional result is generated based on the acoustic features of the second speech. The second target speech segment that is not masked by scene noise or adversarial noise is located based on the noise annotation data. The initial emotional result of the second target recording segment is corrected based on the corrected initial transcribed text. The timestamp of the corrected initial emotional result is aligned with the second speech using a dynamic time warping algorithm to obtain emotional annotation data. Obtain the speaker embedding vector of the second speech, and perform similarity calculation and match the speaker identifier with a preset speaker database to obtain speaker annotation data; A multimodal annotation file is generated based on the noise annotation data, the text annotation data, the sentiment annotation data, and the speaker annotation data.

5. The method for generating ASR audio corpus based on a multimodal large model according to claim 1, characterized in that, The step of setting the word error rate threshold and semantic similarity threshold based on the scene labels, noise types, and speaker information of the multimodal annotation file includes: Based on scene tags, match word error rate baseline thresholds and semantic similarity baseline thresholds from a preset mapping table; The noise type in the multimodal annotation file is parsed, the signal-to-noise ratio value in the multimodal annotation file is mapped to the noise intensity level, and the matching word error rate baseline threshold and the semantic similarity baseline threshold are adjusted according to the noise type and the noise intensity level. Based on the speaker identifier, the speaker attributes are queried from the preset speaker database. Based on the speaker attributes, the matching word error rate baseline threshold and the semantic similarity baseline threshold, which have undergone the first adjustment, are adjusted a second time to obtain the word error rate threshold and the semantic similarity threshold.

6. The method for generating ASR audio corpus based on a multimodal large model according to claim 1, characterized in that, The step of filtering target corpora from the multimodal annotation file based on the word error rate threshold and the semantic similarity threshold includes: The transcribed text of each speech in the multimodal annotation file is time-aligned with a preset reference text, the number of errors is counted, and the word error rate is calculated based on the number of errors, including the number of substitution errors, the number of insertion errors, and the number of deletion errors. Semantic encoding is performed on the transcribed text of each speech in the multimodal annotation file and a preset reference text, and cosine similarity is calculated; The target corpus is determined from speech samples whose cosine similarity is greater than or equal to the semantic similarity threshold and whose word error rate is less than or equal to the word error rate threshold.

7. A system for generating ASR audio corpus based on a multimodal large model, characterized in that, It includes a vector module, a noise module, a labeling module, and an execution module, among which: The vector module is configured to encode target domain text into semantic vectors using a multimodal large model, encode speech control parameters and speaker features into conditional vectors, concatenate the semantic vectors and the conditional vectors into a joint vector, and decode the joint vector into first speech. The noise module is configured to select target noise from a preset noise library based on scene labels, superimpose the target noise onto the first speech to generate noisy speech, and inject adversarial noise into the noisy speech to generate a second speech. The annotation module is configured to perform noise annotation, text annotation, sentiment annotation, and speaker annotation on the second speech, and align the annotation data to the same timeline to generate a multimodal annotation file; The execution module is configured to set a word error rate threshold and a semantic similarity threshold based on the scene labels, noise type, and speaker information of the multimodal annotation file, and to filter target corpora from the multimodal annotation file according to the word error rate threshold and the semantic similarity threshold. The step of encoding speech control parameters and speaker features into a conditional vector, concatenating the semantic vector and the conditional vector into a joint vector, and decoding the joint vector into the first speech includes: Speech rate, pause interval, intonation and emotional intensity are mapped to a first vector, speaker identifier, dialect, timbre and age are mapped to a second vector, and the first vector and the second vector are concatenated to generate a conditional vector. The semantic vector is assigned a first weight and the conditional vector a second weight through an attention mechanism. The semantic vector and the conditional vector are then concatenated along the channel dimension according to the first weight and the second weight to form a joint vector. The joint vector is decoded frame by frame into a Mel spectrum, and the Mel spectrum is inversely transformed into a time-domain waveform to obtain the first speech.

8. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method and device for constructing corpus, computing equipment and storage medium

    CN111104546A

  • Speech recognition system and method for automatic training through speech synthesis method

    CN112669825A