Audio multi-scene noise adding processing method and device, equipment and medium

By using a multi-scenario audio noise processing method, a variety of noise audio samples are generated using a latent diffusion model. This solves the problems of resource consumption and scenario limitations in existing technologies and improves the training and robustness of the model in various acoustic scenarios.

CN118737119BActive Publication Date: 2026-05-05GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG UNIV OF TECH
Filing Date
2024-07-12
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing technologies, it is necessary to consume a lot of human resources and time to build a real-world audio noise reduction method, and the application scenarios are limited to known or foreseeable scenarios, making it difficult to effectively train models in various acoustic scenarios.

Method used

A multi-scenario audio noise processing method is adopted. Noise audio samples are generated through a latent diffusion model. Combined with Gaussian noise distribution and text embedding, audio samples of various noise types are gradually generated. The volume multiple threshold is randomly selected for synthesis to generate noise-added audio.

Benefits of technology

It improves the training accuracy and robustness of the model in various acoustic scenarios, simplifies the operation process, and significantly improves the processing efficiency of adding noise to audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118737119B_ABST
    Figure CN118737119B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, device, and medium for multi-scene audio noise enhancement. The method includes: an audio service system acquiring noise types in a target acoustic scene and the original audio requiring multi-scene noise enhancement; the audio service system transmitting each noise type as text embedding to a potential diffusion model in a noise generation system, using Gaussian noise distribution and text embedding as starting points in the potential diffusion model to progressively generate noise audio samples; copying each noise audio sample according to multiple preset volume multiple thresholds to determine the noise audio samples corresponding to the multiple preset volume multiple thresholds; randomly selecting one or more noise audio samples corresponding to each noise type's preset volume multiple thresholds and synthesizing them with the original audio requiring multi-scene noise enhancement to obtain the noise-enhanced audio. This application enables the model to better adapt to the actual environment and improves its robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and in particular to an audio multi-scene noise addition processing method, corresponding apparatus, electronic device and computer-readable storage medium. Background Technology

[0002] As the market continues to develop, various audio task models are emerging, providing services such as voice wake-up and speech recognition. However, most training datasets directly process audio signals, only serving as data augmentation and failing to guarantee effectiveness in specific application scenarios. Therefore, there is an urgent need for an audio contextualization method tailored to specific application scenarios. Audio contextualization requires adding scene noise to the audio signal to simulate a real acoustic environment; this is the audio noise-adding method in audio signal processing.

[0003] After production and market entry, most audio task models, influenced by scene factors, gradually shifted towards meeting the needs of different noise environments in practical applications. With the content remaining unchanged, the acoustic environment evolved, causing the actual performance of the model to deviate from the expected results, thus compromising the user experience.

[0004] With the rapid development of hardware computing power, scene-based noise reduction technology has also made great progress. Scene-based noise reduction methods are mainly divided into two categories: one is to record sound by actually building a scene, and the other is to simulate the acoustic scene through software simulation. Since actual construction requires a lot of human resources and time, more and more researchers tend to explore software simulation further. Because scene-based noise reduction methods are based on recorded actual noise and construct or simulate according to existing scene information, they are also called non-generative scene-based noise reduction methods. Although non-generative methods can specifically improve the model's performance in target acoustic scenes, their application scenarios are limited to known or predictable scenes. In practical applications, replicating acoustic scenes is time-consuming and labor-intensive, and researchers find it difficult to collect enough scene-based audio to train models.

[0005] In summary, existing technologies require significant human resources and time to build, and their application scenarios are limited to known or foreseeable scenarios. In practical applications, replicating acoustic scenarios is time-consuming and labor-intensive, and researchers find it difficult to collect sufficient scenario-based audio to train models. The applicant has made corresponding explorations to address these issues. Summary of the Invention

[0006] The purpose of this application is to solve the above-mentioned problems by providing an audio multi-scene noise addition processing method, corresponding device, electronic device and computer-readable storage medium.

[0007] To achieve the various objectives of this application, the following technical solution is adopted:

[0008] An audio multi-scene noise addition method proposed to meet one of the purposes of this application includes:

[0009] In response to the audio multi-scenario noise addition processing command, the audio service system obtains the noise type in the target acoustic scene and the original audio that needs to be processed for multi-scenario noise addition.

[0010] The audio service system transmits each type of noise as a text embedding to the potential diffusion model in the noise generation system. In the potential diffusion model, Gaussian noise distribution and the text embedding are used as the starting point to generate noise audio samples step by step.

[0011] The audio service system copies each noise audio sample according to multiple preset volume multiple thresholds to determine the noise audio sample corresponding to the multiple preset volume multiple thresholds;

[0012] The audio service system randomly selects one or more noise audio samples corresponding to preset volume multiple thresholds in each noise type, and synthesizes them with the original audio that needs to undergo multi-scene noise addition processing to obtain noise-added audio.

[0013] Optionally, the audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including:

[0014] In the latent diffusion model, an audio prior based on contrastive text-audio pre-training is generated;

[0015] A variational autoencoder is used as the decoder, and the Mel spectrogram is reconstructed based on the audio prior.

[0016] A pre-defined adversarial generative network is used as a vocoder to generate high-quality noise audio samples based on the Mel spectrogram.

[0017] Optionally, the audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including:

[0018] The potential diffusion model includes both the diffusion process and the reverse diffusion process;

[0019] During the diffusion process, the text embedding at each time step n∈[1,...,N] is given by the following formula:

[0020]

[0021] Where, β n It is a predefined noise scale, and satisfies 0 < β1 < ... < β n <...<β N <1, α n It is 1-β n Reparameterization, This indicates the noise level at each step. This represents the injected standard Gaussian noise distribution at the last time step N. It has a standard isotropic Gaussian distribution;

[0022] For model optimization, the training objective is estimated using reweighted noise:

[0023]

[0024] Where θ represents the current parameter status. This indicates the calculation of ∈ and ∈ θ (z n ,n,E y The similarity of ) , where ∈ is injected noise, ∈ θ (z n ,n,E y ) is the prediction noise, z n The prediction noise is a Gaussian distribution, where n is the time step and E is the value of E. x It is a comparison of the pre-trained audio encoder f in text-to-audio pre-training. audio (·) Embedding of the generated audio waveform x;

[0025] During the reverse diffusion process, from the Gaussian noise distribution and text embedding E y Begin by embedding the text in E y For the conditional denoising process, the audio prior z0 is gradually generated through the following steps:

[0026]

[0027] The mean and variance are parameterized as follows:

[0028]

[0029] Where, ∈ θ (z n ,n,E y ) is the predicted noise, During the training phase, based on the audio embedding E of the audio sample x x Learn to generate audio prior z0, and provide text embedding E during the prediction phase. y To predict noise ∈ θ (zn ,n,E y ).

[0030] Optionally, the audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including:

[0031] In the contrastive text-audio pre-training, noisy audio samples are represented as x, and text descriptions are represented as y, using a text encoder f. text (·) and audio encoder f audio (·) Extract the text embedding E respectively y and audio embedding E x .

[0032] Optionally, the audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including:

[0033] In a variational autoencoder, the variational autoencoder consists of an encoder and a decoder with stacked convolutional modules;

[0034] The encoder compresses the Mel spectrogram X into a latent space. Where r represents the compression ratio;

[0035] The audio prior representation generated by the decoder from the latent diffusion model Constructing Mel spectrograms Using a pre-defined generative adversarial network as a vocoder, from the Mel spectrogram Noise audio samples

[0036] Optionally, the audio service system randomly selects one or more noise audio samples corresponding to preset volume multiple thresholds for each noise type, and synthesizes them with the original audio that needs to undergo multi-scene noise addition processing, including:

[0037] The audio service system opens and reads each audio file, converting it into a NumPy array;

[0038] Determine the maximum length of all audio data and pad any data that is too short with zeros;

[0039] Stack all audio data column-wise into a two-dimensional array, then flatten it into a one-dimensional array;

[0040] Create a new WAV audio file and write the flattened audio data into it.

[0041] Optionally, the basic network architecture of the adversarial generative network is the HiFi-GAN adversarial generative network.

[0042] An audio multi-scene noise processing apparatus provided for another purpose of this application includes:

[0043] The audio acquisition module is configured to respond to audio multi-scenario noise addition processing commands, and the audio service system acquires the noise type in the target acoustic scene and the original audio that needs to be processed for multi-scenario noise addition.

[0044] The noise audio generation module is configured such that the audio service system transmits each noise type as a text embedding to the potential diffusion model of the noise generation system, and in the potential diffusion model, Gaussian noise distribution and the text embedding are used as the starting point to generate noise audio samples step by step.

[0045] The audio sample copying module is configured to copy each noise audio sample according to multiple preset volume multiple thresholds in the audio service system, so as to determine the noise audio sample corresponding to the multiple preset volume multiple thresholds;

[0046] The audio synthesis module is configured to randomly select one or more noise audio samples corresponding to a preset volume multiple threshold in each noise type, and synthesize them with the original audio that needs to undergo multi-scene noise addition processing to obtain noise-added audio.

[0047] An electronic device provided for another purpose of this application includes a central processing unit and a memory, the central processing unit being configured to invoke and run a computer program stored in the memory to perform the steps of the audio multi-scene noise addition processing method of this application.

[0048] A computer-readable storage medium is provided for another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the audio multi-scene noise addition processing method, which, when called by a computer, executes the steps included in the corresponding method.

[0049] Compared to existing technologies, this application addresses the problems of existing technologies requiring significant human and time investment for actual construction, and being limited to known or foreseeable scenarios. Furthermore, replicating acoustic scenarios is time-consuming and labor-intensive in practical applications, and researchers often struggle to collect sufficient contextualized audio for model training. This application offers the following benefits, including but not limited to:

[0050] The audio multi-scene noise addition method of this application can greatly improve the accuracy of model training and learning in various acoustic scenarios by adding noise to the audio multi-scene noise addition method. It can significantly improve the robustness of the model and evaluate the model performance. By adding various noises in the real world to the training data, the model can better adapt to the actual environment and improve its robustness. By adding noise to the test data, the performance of the model in practical applications can be evaluated.

[0051] Furthermore, the audio multi-scenario noise addition method of this application can add different types of noise simultaneously, is simple to operate, has a short processing time, and significantly improves the processing efficiency of adding noise to audio. Attached Figure Description

[0052] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0053] Figure 1 This is a flowchart illustrating the audio multi-scene noise reduction processing method in the embodiments of this application;

[0054] Figure 2 This is a schematic diagram illustrating the audio multi-scene noise addition processing performed by the noise generation system in the embodiments of this application;

[0055] Figure 3 This is a schematic diagram of the structure of the noise generation system based on the latent diffusion model in an embodiment of this application;

[0056] Figure 4 This is a flowchart of the audio classification model trained in the embodiments of this application;

[0057] Figure 5 This is a schematic block diagram of the audio multi-scene noise reduction processing device in the embodiments of this application;

[0058] Figure 6 This is a schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Implementation

[0059] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0060] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0061] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0062] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant) that may include a radio frequency receiver, pager, internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.

[0063] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.

[0064] It should be noted that the concept of "server" used in this application can also be extended to the case of server clusters. Based on the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can be independent of each other but accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method in this application.

[0065] One or more of the technical features of this application, unless explicitly specified herein, can be deployed on a server and accessed by a client remotely calling the online service interface provided by the server, or can be directly deployed and run on a client for access.

[0066] Unless otherwise specified, the neural network models referenced or potentially referenced in this application may be deployed on a remote server and invoked remotely on the client, or deployed on a client with the capability to invoke directly. In some embodiments, when running on the client, the corresponding intelligence may be acquired through transfer learning in order to reduce the requirements on the client's hardware resources and avoid excessive consumption of the client's hardware resources.

[0067] Unless otherwise specified, all data involved in this application may be stored remotely on a server or on a local terminal device, as long as it is suitable for use by the technical solution of this application.

[0068] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.

[0069] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.

[0070] Please see Figure 1 as well as Figure 2 In one embodiment of the audio multi-scene noise addition processing method of this application, the method includes:

[0071] Step S10: In response to the audio multi-scene noise addition processing instruction, the audio service system obtains the noise type in the target acoustic scene and the original audio that needs to be processed for audio multi-scene noise addition.

[0072] The audio service system can respond to audio multi-scene noise addition processing commands. The audio service system can obtain the noise type in the target acoustic scene and the original audio that needs to be processed for audio multi-scene noise addition. The noise type in the target acoustic scene can be the hawking of street vendors, the crying of babies, the chirping of birds outside the window, etc. All of these can be used as noise types in the target acoustic scene to generate noise audio samples. The noise audio samples can be used as the training set for the audio classification model to train the audio classification model.

[0073] In some embodiments, the audio classification model may be a convolutional neural network (CNN), a recurrent neural network (RNN) and its variants, a hybrid model (CNN+RNN), a self-attention model (Transformer), etc.

[0074] In some embodiments, Convolutional Neural Networks (CNNs) are widely used in audio classification. Typically, audio data is converted into spectrograms or mel-spectrograms and then fed into a CNN for processing and classification. CNNs effectively capture the temporal relationships between local features and the spectrum, making them suitable for many audio classification tasks, such as speech recognition and environmental sound classification.

[0075] In some embodiments, recurrent neural networks (RNNs) and their variants: RNNs and their variants (such as Long Short-Term Memory Networks (LSTM), Gated Recurrent Units (GRUs), etc.) are well adapted to processing sequential data (such as audio waveforms). They can model the dynamic changes of audio data over time, and therefore perform well in audio classification tasks that require consideration of temporal information, such as sentiment analysis and speech emotion recognition.

[0076] In some embodiments, hybrid models (CNN+RNN) are used, and in many practical applications, CNN and RNN are combined to leverage their respective strengths. For example, CNNs can be used to extract spectral features and generate high-level feature representations, which are then fed into an RNN for classification. This approach is commonly used for speech recognition and long audio segment classification.

[0077] Step S20: The audio service system transmits each type of noise as a text embedding to the potential diffusion model of the noise generation system. In the potential diffusion model, Gaussian noise distribution and the text embedding are used as the starting point to generate noise audio samples step by step.

[0078] After the audio service system obtains the noise type in the target acoustic scene and the original audio that needs to be subjected to multi-scene noise addition processing, the audio service system transmits each noise type as a text embedding to the potential diffusion model in the noise generation system. In the potential diffusion model, Gaussian noise distribution and the text embedding are used as the starting point to generate noise audio samples step by step. The text embedding is a prompt text.

[0079] Furthermore, the audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including:

[0080] Step S201: Generate audio priors based on contrastive text audio pre-training in the latent diffusion model;

[0081] Step S203: Use a variational autoencoder as a decoder and reconstruct the Mel spectrogram based on the audio prior.

[0082] Step S205: Using a preset adversarial generative network as a vocoder, high-quality noise audio samples are generated based on the Mel spectrogram.

[0083] Specifically, please refer to Figure 3 The basic network architecture of the adversarial generative network is the Hi Fi-GAN adversarial generative network, and the potential diffusion model includes the diffusion process and the de-diffusion process.

[0084] During the diffusion process, the text embedding at each time step n∈[1,...,N] is given by the following formula:

[0085]

[0086]

[0087] Where, β n It is a predefined noise scale, and satisfies 0 < β1 < ... < β n <...<β N <1, α n It is 1-β n Reparameterization, This indicates the noise level at each step. This represents the injected standard Gaussian noise distribution at the last time step N. It has a standard isotropic Gaussian distribution;

[0088] For model optimization, the training objective is estimated using reweighted noise:

[0089]

[0090] Where θ represents the current parameter configuration (i.e., various matrices and weights). This indicates the calculation of ∈ and ∈ θ (z n ,n,E y The similarity of ) , where ∈ is injected noise, ∈ θ (z n ,n,E y ) is the prediction noise, z n The prediction noise is a Gaussian distribution, where n is the time step and E is the value of E. x It is a comparison of the pre-trained audio encoder f in text-to-audio pre-training. audio (·) Embedding of the generated audio waveform x;

[0091] During the reverse diffusion process, from the Gaussian noise distribution and text embedding E y Begin by embedding the text in E y For the conditional denoising process, the audio prior z0 is gradually generated through the following steps:

[0092]

[0093] The mean and variance are parameterized as follows:

[0094]

[0095] Where, ∈ θ (z n ,n,E y ) is the predicted noise, During the training phase, based on the audio embedding E of the audio sample x x Learn to generate audio prior z0, and provide text embedding E during the prediction phase. y To predict noise ∈ θ (z n ,n,E y ).

[0096] In some embodiments, during contrastive text-audio pre-training, noisy audio samples are represented as x, and text descriptions are represented as y, which are encoded using a text encoder f. text (·) and audio encoder f audio (·) Extract the text embedding E respectively y and audio embedding E x .

[0097] Furthermore, the audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including:

[0098] Step S2001: In the variational autoencoder, the variational autoencoder consists of an encoder and a decoder with stacked convolutional modules;

[0099] Step S2003: The encoder compresses the Mel spectrogram X into the latent space. Where r represents the compression ratio;

[0100] Step S2005: The decoder generates an audio prior representation from the latent diffusion model. Constructing Mel spectrograms Using a pre-defined generative adversarial network as a vocoder, from the Mel spectrogram Noise audio samples

[0101] Step S30: The audio service system copies each noise audio sample according to multiple preset volume multiple thresholds to determine the noise audio samples corresponding to the multiple preset volume multiple thresholds;

[0102] In the potential diffusion model, Gaussian noise distribution and text embedding are used as the starting point. After generating noise audio samples step by step, the audio service system copies each noise audio sample according to multiple preset volume multiple thresholds to determine the noise audio samples corresponding to the multiple preset volume multiple thresholds. The multiple preset volume multiple thresholds can be 0%, 25%, 50%, 75%, and 100%, etc.

[0103] Step S40: The audio service system randomly selects one or more noise audio samples corresponding to preset volume multiple thresholds in each noise type, and synthesizes them with the original audio that needs to be subjected to multi-scene noise addition processing to obtain noise-added audio.

[0104] After determining the noise audio samples corresponding to the multiple preset volume multiple thresholds, the audio service system randomly selects one or more noise audio samples corresponding to the preset volume multiple thresholds in each noise type, and synthesizes them with the original audio that needs to undergo multi-scene noise addition processing to obtain the noise-added audio.

[0105] In some embodiments, the user inputs the noise type NoiseType = [nt] in the target acoustic scene into the audio service system. 1 ,nt 2 ,…,ntm ] and the original audio A;

[0106] The audio service system will include each noise type. i As prompt text E y Input noise generation system, from Gaussian noise distribution and text embedding E y Initially, generate noisy audio samples step by step.

[0107] The audio service system will copy each noisy audio file at five different volume levels: 0%, 25%, 50%, 75%, and 100%.

[0108] For each noise type nt i Randomly select a volume multiplier It is then combined with the original audio to form a multi-channel audio, i.e., a noise-added frequency A′.

[0109] In some embodiments, the audio service system randomly selects one or more noise audio samples corresponding to preset volume multiple thresholds for each noise type, and synthesizes them with the original audio that needs to undergo multi-scene noise addition processing, including:

[0110] Step S401: The audio service system opens and reads each audio file and converts it into a NumPy array;

[0111] Step S403: Determine the maximum length of all audio data and pad any data that is too short with zeros;

[0112] Step S405: Stack all audio data column by column into a two-dimensional array, and then flatten it into a one-dimensional array;

[0113] Step S407: Create a new WAV audio file and write the flattened audio data into it.

[0114] Specifically, the audio file is opened, and the `readframes` method from Python's standard audio processing library is used to read all the audio data in the file. This audio data is then converted into an array; that is, the audio file `noise_type_0.wav` is represented as `[a1, a2, ..., a...`. i The audio file noise_type_1.wav is represented as [b1,b2,…,b]. j The audio file noise_type_2.wav is represented as [c1,c2,…,c…]. k The audio file original_audio.wav is represented as [d1, d2, ..., d...]. l.

[0115] Find the maximum length among all audio data, and then pad the shorter audio data. That is, when i < j < l < k, noise_type_0.wav is represented as [a1, a2, …, a i , a i+1 , …, a k , where a i+1 to a k are all 0; noise_type_1.wav is represented as [b1, b2, …, b j , b j+1 , …, b k , where b j+1 to b k are all 0; original_audio.wav is represented as [d1, d2, …, d l , d<00000​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​Step S4003: Copy multiple volumes: Use methods such as torchaudio to sample the audio, and multiply the sampled values ​​by 0, 0.25, 0.5, 0.75 and 1 respectively to generate multiple sets of audio with different volumes, and save them as multiple audio records.

[0122] Step S4005: Synthesize the noisy dataset: Mix the data using methods such as random, traverse each original audio in the original dataset, randomly select three types of noise, extract the noise audio after adjusting the volume, and use methods such as wave to synthesize multi-channel audio to generate the noisy dataset.

[0123] Step S4007: Train the model: Use the noisy dataset as the training set to train a more robust audio task model.

[0124] In some embodiments, the audio multi-scenario noise addition method of this application is applicable to evaluating the performance of existing models in real-world application scenarios by adding noise to test data. For example, after a guide robot is deployed in a street environment and its voice wake-up and speech recognition models are trained, the street noise generated by the audio multi-scenario noise addition method of this application is added to the test dataset to form a street dataset, thereby testing the performance of the speech model on the street dataset to evaluate its performance in real-world applications.

[0125] As can be seen from the above embodiments, compared with the prior art, this application addresses the problems of the prior art requiring a large amount of human resources and time to actually build, and the application scenarios being limited to known or foreseeable scenarios. In practical applications, replicating acoustic scenarios is time-consuming and labor-intensive, and researchers find it difficult to collect enough scenario-based audio to train models. This application has, but is not limited to, the following beneficial effects:

[0126] Compared to existing technologies, this application addresses the problems of existing technologies requiring significant human and time investment for actual construction, and being limited to known or foreseeable scenarios. Furthermore, replicating acoustic scenarios is time-consuming and labor-intensive in practical applications, and researchers often struggle to collect sufficient contextualized audio for model training. This application offers the following benefits, including but not limited to:

[0127] The audio multi-scene noise addition method of this application can greatly improve the accuracy of model training and learning in various acoustic scenarios by adding noise to the audio multi-scene noise addition method. It can significantly improve the robustness of the model and evaluate the model performance. By adding various noises in the real world to the training data, the model can better adapt to the actual environment and improve its robustness. By adding noise to the test data, the performance of the model in practical applications can be evaluated.

[0128] Furthermore, the audio multi-scenario noise addition method of this application can add different types of noise simultaneously, is simple to operate, has a short processing time, and significantly improves the processing efficiency of adding noise to audio.

[0129] Please see Figure 5 An audio multi-scene noise processing device provided for one of the purposes of this application includes an audio acquisition module 1100, a noise audio generation module 1200, an audio sample copying module 1300, and an audio synthesis module 1400. The audio acquisition module 1100 is configured to respond to audio multi-scene noise addition processing instructions, and the audio service system acquires the noise type in the target acoustic scene and the original audio that needs to be processed for multi-scene noise addition; the noise audio generation module 1200 is configured to transmit each noise type as a text embedding to the potential diffusion model in the noise generation system, and use Gaussian noise distribution and the text embedding as the starting point in the potential diffusion model to gradually generate noise audio samples; the audio sample copying module 1300 is configured to copy each noise audio sample according to multiple preset volume multiple thresholds to determine the noise audio samples corresponding to the multiple preset volume multiple thresholds; the audio synthesis module 1400 is configured to randomly select one or more noise audio samples corresponding to preset volume multiple thresholds in each noise type, and synthesize them with the original audio that needs to be processed for multi-scene noise addition to obtain the noise-added audio.

[0130] Based on any embodiment of this application, please refer to Figure 6 Another embodiment of this application also provides an electronic device, which can be implemented by a computer device, such as... Figure 6 The diagram shows the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, they enable the processor to implement an audio multi-scene noise enhancement processing method. The processor of the computer device provides computing and control capabilities, supporting the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When these computer-readable instructions are executed by the processor, they enable the processor to execute the audio multi-scene noise enhancement processing method of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 6The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0131] In this embodiment, the processor is used to execute... Figure 5 The system contains the specific functions of each module and its sub-modules, and the memory stores the program code and various data required to execute the aforementioned modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the audio multi-scene noise enhancement processing device of this application, and the server can call the server's program code and data to execute the functions of all sub-modules.

[0132] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the audio multi-scene noise addition processing method described in any embodiment of this application.

[0133] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the audio multi-scene noise addition processing method described in any embodiment of this application.

[0134] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0135] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

[0136] In summary, the audio multi-scenario noise reduction method of this application can improve the accuracy of the model through model training and learning in various acoustic scenarios, significantly improve the robustness of the model and evaluate its performance. By adding various noises from the real world to the training data, the model can better adapt to the actual environment and improve its robustness. By adding noise to the test data, the model's performance in practical applications is greatly enhanced.

Claims

1. A method for multi-scene audio noise reduction processing, characterized in that, include: In response to the audio multi-scenario noise addition processing command, the audio service system obtains the noise type in the target acoustic scene and the original audio that needs to be processed for multi-scenario noise addition. The audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including: The potential diffusion model includes a diffusion process and a reverse diffusion process; in the diffusion process, the text is embedded at each time step. The transition probability is given by the following formula: , , in, It is a predefined noise scale and satisfies , express The reparameterized coefficients, , This indicates the noise level at each step. The standard Gaussian distribution of the injected noise is represented at the last time step. , It has a standard isotropic Gaussian distribution; For model optimization, the training objective is estimated using reweighted noise: , in, This refers to the current parameter status. It's injected noise. It is predictive noise. It is a time step. It is a comparison of the pre-trained audio encoder in text-to-audio pre-training. Generated noise audio samples Audio embedding; During the reverse diffusion process, from the Gaussian noise distribution and text embedding Begin with the text embedded The conditional denoising process gradually generates audio priors through the following steps. ,include: , , The mean parameter is: , The variance parameter is: , in, It predicts noise; during the training phase, it is based on the noise audio samples. audio embedding Learning to generate audio priors During the prediction phase, the text embedding is provided. To predict noise ; In contrastive text-audio pre-training Represents a noisy audio sample. This represents a text description, which uses a text encoder. and audio encoder Extract text embeddings separately and audio embedding ; In a variational autoencoder, the variational autoencoder consists of an encoder and a decoder with stacked convolutional modules; the encoder outputs a Mel-spectrum image. Compressed into potential space ,in, The compression ratio is represented by the audio prior representation generated by the decoder from the latent diffusion model. Constructing Mel spectrograms Using a pre-defined generative adversarial network as a vocoder, the vocoder is obtained from the Mel spectrogram. Generate noisy audio samples ; The audio service system copies each noise audio sample according to multiple preset volume multiple thresholds to determine the noise audio sample corresponding to the multiple preset volume multiple thresholds; The audio service system randomly selects one or more noise audio samples corresponding to preset volume multiple thresholds in each noise type, and synthesizes them with the original audio that needs to undergo multi-scene noise addition processing to obtain noise-added audio.

2. The audio multi-scene noise reduction processing method according to claim 1, characterized in that, The audio service system transmits each noise type as a text embedding to the latent diffusion model in the noise generation system. The latent diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including: In the latent diffusion model, an audio prior based on contrastive text audio pre-training is generated; A variational autoencoder is used as the decoder, and the Mel spectrogram is reconstructed based on the audio prior. A pre-defined adversarial generative network is used as a vocoder to generate high-quality noise audio samples based on the Mel spectrogram.

3. The audio multi-scene noise reduction processing method according to claim 1, characterized in that, The audio service system randomly selects one or more noise audio samples corresponding to preset volume multiple thresholds for each noise type, and synthesizes them with the original audio that needs to undergo multi-scene noise addition processing. This includes the following steps: The audio service system opens and reads each audio file, converting it into a NumPy array; Determine the maximum length of all audio data and pad any data that is too short with zeros; Stack all audio data column-wise into a two-dimensional array, then flatten it into a one-dimensional array; Create a new WAV audio file and write the flattened audio data into it.

4. The audio multi-scene noise reduction processing method according to any one of claims 1 to 3, characterized in that, The basic network architecture of the adversarial generative network is the HiFi-GAN adversarial generative network.

5. An audio multi-scene noise processing device, characterized in that, include: The audio acquisition module is configured to respond to audio multi-scenario noise addition processing commands, and the audio service system acquires the noise type in the target acoustic scene and the original audio that needs to be processed for multi-scenario noise addition. A noise audio generation module is configured such that the audio service system transmits each noise type as a text embedding to a potential diffusion model in the noise generation system. The potential diffusion model uses a Gaussian noise distribution and the text embedding as starting points to progressively generate noise audio samples, including: The potential diffusion model includes a diffusion process and a reverse diffusion process; in the diffusion process, the text is embedded at each time step. The transition probability is given by the following formula: , , in, It is a predefined noise scale and satisfies , express The reparameterized coefficients, , This indicates the noise level at each step. The standard Gaussian distribution of the injected noise is represented at the last time step. , It has a standard isotropic Gaussian distribution; For model optimization, the training objective is estimated using reweighted noise: , in, This refers to the current parameter status. It's injected noise. It is predictive noise. It is a time step. It is a comparison of the pre-trained audio encoder in text-to-audio pre-training. Generated noise audio samples Audio embedding; During the reverse diffusion process, from the Gaussian noise distribution and text embedding Begin with the text embedded The conditional denoising process gradually generates audio priors through the following steps. ,include: , , The mean parameter is: , The variance parameter is: , in, It predicts noise; during the training phase, it is based on the noise audio samples. audio embedding Learning to generate audio priors During the prediction phase, the text embedding is provided. To predict noise ; In contrastive text-audio pre-training Represents a noisy audio sample. This represents a text description, which uses a text encoder. and audio encoder Extract text embeddings separately and audio embedding ; In a variational autoencoder, the variational autoencoder consists of an encoder and a decoder with stacked convolutional modules; the encoder outputs a Mel-spectrum image. Compressed into potential space ,in, The compression ratio is represented by the audio prior representation generated by the decoder from the latent diffusion model. Constructing Mel spectrograms Using a pre-defined generative adversarial network as a vocoder, the vocoder is obtained from the Mel spectrogram. Generate noisy audio samples ; The audio sample copying module is configured to copy each noise audio sample according to multiple preset volume multiple thresholds in the audio service system, so as to determine the noise audio sample corresponding to the multiple preset volume multiple thresholds; The audio synthesis module is configured to randomly select one or more noise audio samples corresponding to a preset volume multiple threshold in each noise type, and synthesize them with the original audio that needs to undergo multi-scene noise addition processing to obtain noise-added audio.

6. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 4, which, when invoked by a computer, executes the steps included in the corresponding method.

Citation Information

Patent Citations

  • Voice noise method and system for data enhancement

    CN110211575A

  • Voice noise reduction method and device based on model fusion and storage medium

    CN117789744A