Audio noise reduction system, noise reduction method and vehicle-mounted audio system

By using a combination of deep neural networks and adaptive mask generators in the audio noise reduction system, the problem of poor performance in complex noise environments is solved, and high-efficiency and low-latency audio noise reduction is achieved, adapting to a variety of devices and formats, and reducing hardware costs.

CN119400197BActive Publication Date: 2025-05-13SHANGHAI LINGJING ACOUSTIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510006612.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-13
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The existing audio noise reduction technology is not effective when dealing with complex, dynamic and variable noise environments, and has problems such as negative impact on vocal quality, long delay, poor equipment compatibility and high hardware costs.

Method used

An audio noise reduction system based on deep neural network is adopted, including an input audio preprocessing module, a deep neural network separation module, an adaptive mask generator and an output reconstruction module. The system achieves fine separation of vocals and noise through feature extraction, encoder decoder architecture and adaptive mask generator, and operates on NPUs or GPUs through optimized model architecture, reducing latency and hardware costs.

Benefits of technology

It realizes efficient noise reduction in complex and dynamic noise environments, keeps the human voice clear and natural, reduces processing delays, adapts to audio signals of different devices and formats, and reduces hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119400197B_ABST
    Figure CN119400197B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio noise reduction system, a noise reduction method and a vehicle-mounted audio system. The noise reduction system comprises an input audio preprocessing module, which preprocesses an input audio signal to generate an audio spectrum; a deep neural network separation module, which is configured with: a feature extraction layer, which extracts audio features, namely, time features and frequency features, from the audio spectrum; an encoder, which is used to compress the audio features output by the feature extraction layer and generate a latent space representation according to the audio features; a decoder, which reconstructs and separates human voice and noise according to the latent space representation; an adaptive mask generator, which generates a time-series spectrum mask according to the frequency difference between noise and human voice and the time feature to separate the human voice audio; and an output reconstruction module, which is configured to inversely transform the frequency features of the separated human voice audio signal to reconstruct time-domain audio data, which can adapt to different noise types and avoid speech distortion or erroneous suppression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI, and in particular to an audio noise reduction system, a noise reduction method and a vehicle-mounted audio system. Background Art

[0002] In application scenarios such as in-vehicle systems, smart home devices, audio and video conferencing, and real-time communication devices, noise interference will have a negative impact on user experience. Audio noise reduction is a technology that aims to reduce or eliminate background noise in audio to improve audio quality, especially playing a very important role in improving speech clarity in noisy environments.

[0003] Traditional noise suppression methods are mainly implemented through filtering, noise threshold, frequency suppression and other algorithms. Although these methods can reduce noise to a certain extent, they also have at least the following problems:

[0004] First, while background noise is reduced or eliminated, noise reduction technology can also negatively affect the quality of the human voice in the audio, causing sound distortion or incorrect suppression of speech parts.

[0005] Secondly, current noise reduction technology constructs different noise models for different types of noise, such as wind noise, crowd noise, and machine noise. However, such noise models for specific types have poor adaptability to other types of noise. Therefore, such noise models will behave unstable in a changing noise environment, which is reflected in the incomplete removal of noise from multiple noise sources, or can only process noise in specific frequency bands, and cannot cope with complex, dynamic, and changing audio scenes.

[0006] Third, at present, when using deep learning models to process real-time audio, there are often problems such as long delays and insufficient processing efficiency, which makes it difficult to meet application scenarios with strict requirements for low latency, such as instant calls, live broadcasts, real-time meetings, in-car voice control, and smart AI voice assistants. Processing delays will lead to a poor user experience.

[0007] Fourth, the current noise reduction system has high requirements for input audio format or device compatibility in practical applications, and cannot flexibly adapt to different devices or audio signal sources of different formats, and its application scenarios are limited.

[0008] In addition, traditional audio processing models usually rely on GPUs for high-intensity computing, resulting in high hardware costs and high power consumption. They are not suitable for embedded devices or resource-constrained devices, nor for large-scale deployment.

[0009] The disclosure of the above background technology content is only used to assist in understanding the concept and technical solution of the present application. It does not necessarily belong to the prior art of the present application, nor does it necessarily provide technical guidance. In the absence of clear evidence that the above content has been disclosed before the filing date of the present application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the invention

[0010] The purpose of the present invention is to provide an intelligent noise reduction system that can adapt to dynamically changing noise types, optimize the model structure and calculation path based on NPU, and achieve low-latency high-definition voice output.

[0011] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0012] An audio noise reduction system comprises the following modules:

[0013] An input audio preprocessing module, which is configured to preprocess the input audio signal to generate an audio spectrum;

[0014] Deep neural network separation module, which is configured with:

[0015] A feature extraction layer, which extracts audio features from the audio spectrum, wherein the audio features include time features and frequency features;

[0016] An encoder-decoder architecture, wherein the encoder of the architecture is used to compress the audio features output by the feature extraction layer and generate a latent space representation based on the audio features, and the decoder of the architecture reconstructs and separates the human voice and noise based on the latent space representation;

[0017] An adaptive mask generator that generates a temporal spectral mask based on the frequency difference between noise and human voice and combines the temporal features to separate the human voice audio;

[0018] The output reconstruction module is configured to perform an inverse transformation on the frequency characteristics of the separated human voice audio signal to reconstruct the time domain audio data.

[0019] Furthermore, based on any one of the technical solutions or a combination of multiple technical solutions described above, the audio noise reduction system also includes an optimization module, which is configured to optimize the model architecture of the deep neural network separation module so that it can run on an NPU or a GPU.

[0020] Further, based on any one of the above-mentioned technical solutions or a combination of multiple technical solutions, the optimization module optimizes the model architecture by the following operations:

[0021] Reducing the model volume of the deep neural network separation module through pruning and / or quantization operations;

[0022] And / or, optimizing the model calculation path of the deep neural network separation module for the multi-core structure of the NPU or GPU.

[0023] Further, based on any one of the above-mentioned technical solutions or a combination of multiple technical solutions, the deep neural network separation module is trained by the following steps:

[0024] Collect multiple noise audios, generate corresponding noise audio spectra, and mark the noise type to which the noise audio spectra belong according to the spectrum of the noise audio spectra; and collect multiple human voice audios, and generate corresponding human voice audio spectra;

[0025] Mix the noise audio with the human voice audio to obtain a mixed audio, and generate a corresponding mixed audio spectrum;

[0026] Using the mixed audio and mixed audio spectrum as learning samples, and using the noise audio spectrum, the annotation of the noise type and the human voice audio spectrum as sample labels;

[0027] The convolutional neural network or the recursive neural network is trained using the collected multiple learning samples and the corresponding labels, and the convolutional neural network or the recursive neural network performs calculations based on an NPU (Neural Processing Unit) or a GPU (Graphics Processing Unit).

[0028] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, the NPU or GPU performs the following operations:

[0029] Split each learning sample into multiple audio frames, each audio frame has the same number of channels;

[0030] Generate a feature vector for each audio frame, wherein the feature vector includes values ​​of each channel of the audio frame;

[0031] Based on the feature vector, estimating the instantaneous function of the human voice audio spectrum and the noise audio spectrum in the audio frame;

[0032] Calculate the reference instantaneous function of the human voice audio spectrum and the noise audio spectrum according to the noise audio spectrum and the human voice audio spectrum in the label of the sample;

[0033] Determine discrete points of the reference instantaneous function according to the time length of each audio frame and the offset time length between adjacent audio frames;

[0034] Taking the reference instantaneous function at the discrete point as the target, the weight of the convolutional neural network or the recurrent neural network is optimized so that the gap between the estimated instantaneous function and the reference instantaneous function at the corresponding discrete point is reduced.

[0035] Furthermore, based on any one of the technical solutions or a combination of multiple technical solutions described above, the instantaneous function is based on any one of the estimation methods including silence segment estimation, minimum mean square error estimation, minimum tracking algorithm, histogram noise estimation algorithm, quantile noise estimation, recursive average noise algorithm, and A-likelihood ratio method.

[0036] Further, according to any one of the technical solutions or a combination of multiple technical solutions described above, before estimating the instantaneous function, a mean and variance normalization operation is performed on the values ​​of each channel in the feature vector.

[0037] Further, according to any one of the above-mentioned technical solutions or a combination of multiple technical solutions, learning samples corresponding to different noise types are collected, and the convolutional neural network or recursive neural network is trained using a binary cross entropy loss function;

[0038] The trained deep neural network separation module is configured with model calculation paths for different noise types;

[0039] The adaptive mask generator adjusts the frequency difference between noise and human voice involved in generating the spectrum mask according to different model calculation paths.

[0040] Further, based on any one of the above-mentioned technical solutions or a combination of multiple technical solutions, the system runs in real time based on an NPU or a GPU, and the NPU or the GPU divides each n frames of input audio into an audio group, and performs reasoning on each audio group to infer whether a dynamic change of noise type occurs in the audio group;

[0041] If so, when the adaptive mask generator generates a spectrum mask corresponding to the time feature of the audio group, the frequency difference between the corresponding noise and the human voice is adjusted to change the generation rule of the spectrum mask.

[0042] According to another aspect of the present invention, the present invention provides an audio noise reduction method, comprising the following steps:

[0043] Input the audio signal to be denoised;

[0044] Preprocessing the input audio signal to generate an audio spectrum;

[0045] A pre-built deep neural network separation module is configured to: extract audio features from the audio spectrum, the audio features including time features and frequency features; compress the extracted audio features and generate a latent space representation based on the audio features; reconstruct and separate human voice and noise based on the latent space representation;

[0046] Based on the frequency difference between noise and human voice, combined with the time characteristics, a temporal spectrum mask is generated to separate the human voice audio;

[0047] The frequency characteristics of the separated human voice audio signal are inversely transformed to reconstruct the time domain audio data.

[0048] Further, based on any one of the above-mentioned technical solutions or a combination of multiple technical solutions, the deep neural network separation module is trained by the following steps:

[0049] Collect multiple noise audios, generate corresponding noise audio spectra, and mark the noise type to which the noise audio spectra belong according to the spectrum of the noise audio spectra; and collect multiple human voice audios, and generate corresponding human voice audio spectra;

[0050] Mix the noise audio with the human voice audio to obtain a mixed audio, and generate a corresponding mixed audio spectrum;

[0051] Using the mixed audio and mixed audio spectrum as learning samples, and using the noise audio spectrum, the annotation of the noise type and the human voice audio spectrum as sample labels;

[0052] The convolutional neural network or the recursive neural network is trained using the collected multiple learning samples and the corresponding labels, and the convolutional neural network or the recursive neural network performs calculations based on an NPU (Neural Processing Unit) or a GPU (Graphics Processing Unit).

[0053] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, the NPU or GPU performs the following operations:

[0054] Split each learning sample into multiple audio frames, each audio frame has the same number of channels;

[0055] Generate a feature vector for each audio frame, wherein the feature vector includes values ​​of each channel of the audio frame;

[0056] Based on the feature vector, estimating the instantaneous function of the human voice audio spectrum and the noise audio spectrum in the audio frame;

[0057] Calculate the reference instantaneous function of the human voice audio spectrum and the noise audio spectrum according to the noise audio spectrum and the human voice audio spectrum in the label of the sample;

[0058] Determine discrete points of the reference instantaneous function according to the time length of each audio frame and the offset time length between adjacent audio frames;

[0059] Taking the reference instantaneous function at the discrete point as the target, the weight of the convolutional neural network or the recurrent neural network is optimized so that the gap between the estimated instantaneous function and the reference instantaneous function at the corresponding discrete point is reduced.

[0060] Further, according to any one of the above-mentioned technical solutions or a combination of multiple technical solutions, before estimating the instantaneous function, a mean and variance normalization operation is first performed on the values ​​of each channel in the feature vector;

[0061] The instantaneous function is based on any one of the estimation methods of silence segment estimation, minimum mean square error estimation, minimum tracking algorithm, histogram noise estimation algorithm, quantile noise estimation, recursive average noise algorithm, and A-likelihood ratio method.

[0062] Further, according to any one of the above-mentioned technical solutions or a combination of multiple technical solutions, learning samples corresponding to different noise types are collected, and the convolutional neural network or recursive neural network is trained using a binary cross entropy loss function;

[0063] The trained deep neural network separation module is configured with model calculation paths for different noise types;

[0064] Before generating the spectral mask, the frequency difference between the noise and the human voice involved in generating the spectral mask is adjusted according to different model calculation paths.

[0065] Further, based on any one of the above-mentioned technical solutions or a combination of multiple technical solutions, the deep neural network separation module is run in real time based on the NPU or GPU, and the NPU or GPU divides each n frames of input audio into an audio group, and performs reasoning on each audio group to infer whether a dynamic change of noise type occurs in the audio group;

[0066] If so, when generating a spectrum mask corresponding to the time feature of the audio group, the frequency difference between the corresponding noise and the human voice is adjusted to change the generation rule of the spectrum mask.

[0067] According to another aspect of the present invention, the present invention provides a vehicle audio system, comprising a speaker and the audio noise reduction system as described above, wherein the speaker is used to play the time domain audio reconstructed by the audio noise reduction system from the separated human voice audio signal.

[0068] The beneficial effects brought by the technical solution provided by the present invention are as follows:

[0069] a. Optimize the deep learning model to significantly improve processing efficiency, which is suitable for application scenarios with high real-time requirements;

[0070] b. Through the adaptive learning function, the system can automatically adjust the noise reduction strategy according to different noise environments to ensure the noise reduction effect in complex, dynamic and changing environments;

[0071] c. Build an improved deep learning audio separation model, using a multi-layer feature extraction mechanism to perform a more refined separation of human voice and noise, ensuring that the quality of human voice is preserved while noise is reduced, avoiding the problem of human voice distortion;

[0072] d. It can be used without relying on the high-performance computing capabilities of GPUs, such as using NPUs to complete computing processing, reducing hardware deployment costs and having broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0074] Figure 1 A schematic diagram of a framework of an audio noise reduction system provided for an exemplary embodiment of the present invention;

[0075] Figure 2 A schematic diagram of a training process of a deep neural network separation module provided for an exemplary embodiment of the present invention;

[0076] Figure 3 A schematic diagram of a flow chart of a neural network performing operations based on an NPU provided for an exemplary embodiment of the present invention;

[0077] Figure 4 A schematic diagram of a spectrum of an audio noise reduction system for separating human voice audio and noise audio provided by an exemplary embodiment of the present invention;

[0078] Figure 5 A flowchart of an audio noise reduction method provided for an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0079] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0080] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0081] Users are increasingly demanding audio quality. From a user experience perspective, audio separation and noise reduction technologies need to maintain high sound quality while having real-time processing capabilities and the ability to adapt to complex, dynamically changing noise environments. The present invention aims to provide an audio noise reduction method that maintains clear and natural human voice quality, reduces processing delays, and enhances adaptive capabilities, so as to achieve rapid adjustment of audio separation algorithms for different noise scenarios to adapt to noise changes in a variety of complex environments.

[0082] In one embodiment of the present invention, an audio noise reduction system is provided. Figure 1 As shown, the system includes the following modules:

[0083] The input audio preprocessing module is configured to preprocess the input audio signal, such as using short-time Fourier transform (STFT) to convert the audio signal into time domain and frequency domain to generate an audio spectrum. Figure 4 Spectrogram of the original sound shown in ;

[0084] Deep neural network separation module, which is configured with:

[0085] A feature extraction layer, which extracts audio features from the audio spectrum, wherein the audio features include time features and frequency features;

[0086] The encoder-decoder architecture is used to compress the audio features output by the feature extraction layer and generate a latent space representation based on the audio features. The decoder of the architecture reconstructs and separates the human voice and noise based on the latent space representation, such as Figure 4 The separated spectrogram of human voice and the spectrogram of noise shown in ;

[0087] An adaptive mask generator generates a temporal spectral mask based on the frequency difference between noise and human voice, combined with temporal features, to separate the human voice audio: the spectral mask indicates which frequency components correspond to noise, that is, they should be suppressed. Applying this mask to the signal can remove the noise signal while retaining the human voice signal.

[0088] The output reconstruction module is configured to perform an inverse transformation on the frequency characteristics of the separated human voice audio signal to reconstruct the time domain audio data.

[0089] In the field of deep learning, latent space denoising involves mapping noise data to a low-dimensional latent space and performing denoising on the data in this space. In this embodiment, a mask-based audio denoising technology is specifically used.

[0090] Continue to see Figure 1 The audio noise reduction system also includes an optimization module, which is configured to optimize the model architecture of the deep neural network separation module so that it can run on the NPU or GPU. The specific optimization operation is to reduce the model volume of the deep neural network separation module through pruning and / or quantization operations; and optimize the model calculation path of the deep neural network separation module for the multi-core structure of the NPU or GPU.

[0091] In this embodiment, the deep neural network separation module needs to be trained through specific learning samples, so that the trained deep neural network separation module configures different model calculation paths for different noise types;

[0092] The adaptive mask generator adjusts the frequency difference between the noise and the human voice involved in generating the spectrum mask according to different model calculation paths. The frequency difference is adjusted, that is, the generation rule of the spectrum mask is adjusted accordingly, so that when the noise type changes, a spectrum mask that is compatible with the type of noise can be adaptively generated, thereby improving the stability of the system in a changeable noise environment and being able to cope with complex, dynamic and changeable audio scenes.

[0093] like Figure 2 As shown, the deep neural network separation module is trained by the following steps:

[0094] Collect multiple noise audios, generate corresponding noise audio spectra, and mark the noise type to which the noise audio spectra belong according to the spectrum of the noise audio spectra; and collect multiple human voice audios, and generate corresponding human voice audio spectra;

[0095] Mix the noise audio with the human voice audio to obtain a mixed audio, and generate a corresponding mixed audio spectrum;

[0096] Using the mixed audio and mixed audio spectrum as learning samples, and using the noise audio spectrum, the annotation of the noise type and the human voice audio spectrum as sample labels;

[0097] The convolutional neural network or recurrent neural network is trained using the collected multiple learning samples and corresponding labels. It should be noted that in order to adapt to different types of noise, learning samples of different noise types are collected in this embodiment, and the convolutional neural network or recurrent neural network is trained using a binary cross entropy loss function. The convolutional neural network or recurrent neural network performs operations based on an NPU (Neural Processing Unit) or a GPU (Graphics Processing Unit). Specifically, the NPU or GPU performs operations such as Figure 3 The operation flow shown is:

[0098] Split each learning sample into multiple audio frames, each audio frame has the same number of channels;

[0099] Generate a feature vector for each audio frame, wherein the feature vector includes the values ​​of each channel of the audio frame. For example, if the audio has 5 channels, the feature vector is a one-dimensional vector with 5 values.

[0100] Based on the feature vector, an instantaneous function of the human voice audio spectrum and the noise audio spectrum in the audio frame is estimated, such as sampling silence segment estimation, minimum mean square error estimation, minimum tracking algorithm, histogram noise estimation algorithm, quantile noise estimation, recursive average noise algorithm or A-likelihood ratio method; in a specific embodiment, before estimating the instantaneous function, mean and variance normalization operations are performed on the values ​​of each channel in the feature vector;

[0101] Calculate the reference instantaneous function of the human voice audio spectrum and the noise audio spectrum according to the noise audio spectrum and the human voice audio spectrum in the label of the sample;

[0102] Determine discrete points of the reference instantaneous function according to the time length of each audio frame and the offset time length between adjacent audio frames;

[0103] Taking the benchmark instantaneous function at the discrete point as the target, the weights of the convolutional neural network or the recursive neural network are optimized so that the gap between the estimated instantaneous function and the benchmark instantaneous function at the corresponding discrete point is reduced until the neural network model completes convergence.

[0104] The instantaneous function of the human voice audio spectrum and the noise audio spectrum may be the instantaneous frequency obtained by estimating the spectrum envelope of the analysis signal through short-time Fourier transform (STFT), or may be analyzed and processed through Hilbert transform, which is not specifically limited in the present invention.

[0105] In one embodiment of the present invention, a processing method for saving computing power is provided: the system runs in real time based on an NPU or a GPU, and the NPU or the GPU divides each n frames of input audio into an audio group, and performs reasoning on each audio group to infer whether a dynamic change of noise type occurs in the audio group;

[0106] If so, when the adaptive mask generator generates a spectrum mask corresponding to the time feature of the audio group, the frequency difference between the corresponding noise and the human voice is adjusted to change the generation rule of the spectrum mask;

[0107] If it is inferred that no dynamic change of the noise type occurs in the audio group, the current generation rule of the spectrum mask is maintained.

[0108] On the other hand, this processing method can also provide a basis for the self-check algorithm: when it is inferred that a change in the noise type has occurred in the audio group but the model has not recognized it, the self-check can be initiated, such as running the model again to identify multiple frames of audio in the audio group. If the change in the noise type still cannot be identified, manual intervention is prompted to troubleshoot the problem. If the change in noise type is indeed found, the audio group is used as a learning sample to optimize the model training.

[0109] The AI ​​model of the present invention does not rely on the computing power of the GPU and can run on the NPU, effectively reducing hardware costs and power consumption. It can be applicable to embedded devices and can be promoted and applied in various in-vehicle audio and video systems.

[0110] In one embodiment of the present invention, an audio noise reduction method is provided. Figure 5 As shown, the following steps are included:

[0111] Input the audio signal to be denoised;

[0112] Preprocessing the input audio signal to generate an audio spectrum;

[0113] A pre-built deep neural network separation module is configured to: extract audio features from the audio spectrum, the audio features including time features and frequency features; compress the extracted audio features and generate a latent space representation based on the audio features; reconstruct and separate human voice and noise based on the latent space representation;

[0114] Based on the frequency difference between noise and human voice, combined with the time characteristics, a temporal spectrum mask is generated to separate the human voice audio;

[0115] The frequency characteristics of the separated human voice audio signal are inversely transformed to reconstruct the time domain audio data.

[0116] like Figure 2 As shown, the deep neural network separation module is trained by the following steps:

[0117] Collect multiple noise audios, generate corresponding noise audio spectra, and mark the noise type to which the noise audio spectra belong according to the spectrum of the noise audio spectra; and collect multiple human voice audios, and generate corresponding human voice audio spectra;

[0118] Mix the noise audio with the human voice audio to obtain a mixed audio, and generate a corresponding mixed audio spectrum;

[0119] Using the mixed audio and mixed audio spectrum as learning samples, and using the noise audio spectrum, the annotation of the noise type and the human voice audio spectrum as sample labels;

[0120] The convolutional neural network or the recursive neural network is trained using the collected multiple learning samples and the corresponding labels, and the convolutional neural network or the recursive neural network performs calculations based on the NPU or the GPU.

[0121] like Figure 3 As shown, the NPU or GPU performs the following operations:

[0122] Split each learning sample into multiple audio frames, each audio frame has the same number of channels;

[0123] Generate a feature vector for each audio frame, wherein the feature vector includes values ​​of each channel of the audio frame;

[0124] Based on the feature vector, the instantaneous function of the human voice audio spectrum and the noise audio spectrum in the audio frame is estimated; before estimating the instantaneous function, the mean and variance normalization operations are performed on the values ​​of each channel in the feature vector; the specific estimation algorithm used can be any one of the silent segment estimation, minimum mean square error estimation, minimum tracking algorithm, histogram noise estimation algorithm, quantile noise estimation, recursive average noise algorithm, and A-likelihood ratio method;

[0125] Calculate the reference instantaneous function of the human voice audio spectrum and the noise audio spectrum according to the noise audio spectrum and the human voice audio spectrum in the label of the sample;

[0126] Determine discrete points of the reference instantaneous function according to the time length of each audio frame and the offset time length between adjacent audio frames;

[0127] Taking the reference instantaneous function at the discrete point as the target, the weight of the convolutional neural network or the recurrent neural network is optimized so that the gap between the estimated instantaneous function and the reference instantaneous function at the corresponding discrete point is reduced.

[0128] In this embodiment, the noise type can be adapted to the dynamic change to ensure the effect of noise reduction. The learning samples collected by the deep neural network separation module for training need to have different noise types, and the convolutional neural network or recursive neural network is trained using a binary cross entropy loss function.

[0129] The trained deep neural network separation module is configured with model calculation paths for different noise types;

[0130] Before generating the spectral mask, the frequency difference between the noise and the human voice involved in generating the spectral mask is adjusted according to different model calculation paths.

[0131] Further, the deep neural network separation module is run in real time based on the NPU or GPU, and the NPU or GPU divides each n frames of input audio into an audio group, and performs reasoning on each audio group to infer whether a dynamic change of noise type occurs in the audio group;

[0132] If so, when generating a spectrum mask corresponding to the time feature of the audio group, the frequency difference between the corresponding noise and the human voice is adjusted to change the generation rule of the spectrum mask.

[0133] An embodiment of the present invention further provides a vehicle-mounted audio system, comprising a speaker and the audio noise reduction system as described above, wherein the speaker is used to play the time domain audio reconstructed by the audio noise reduction system from the separated human voice audio signal.

[0134] The audio noise reduction method and the audio noise reduction system provided in the embodiment of the present invention belong to the same inventive concept. The entire contents of the audio noise reduction system embodiment are incorporated into the audio noise reduction method embodiment by citing the entire embodiment, and the entire contents of the audio noise reduction system embodiment are also incorporated into the vehicle-mounted audio system embodiment, which will not be repeated here.

[0135] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0136] The above is only a specific implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An audio noise reduction system, characterized in that: Includes the following modules: An input audio preprocessing module, which is configured to preprocess the input audio signal to generate an audio spectrum; Deep neural network separation module, which is configured with: A feature extraction layer, which extracts audio features from the audio spectrum, wherein the audio features include time features and frequency features; An encoder-decoder architecture, wherein the encoder of the architecture is used to compress the audio features output by the feature extraction layer and generate a latent space representation based on the audio features, and the decoder of the architecture reconstructs and separates the human voice and noise based on the latent space representation; An adaptive mask generator that generates a temporal spectral mask based on the frequency difference between noise and human voice and combines the temporal features to separate the human voice audio; An output reconstruction module is configured to inversely transform the frequency characteristics of the separated human voice audio signal to reconstruct time domain audio data; The deep neural network separation module is trained by the following steps: collecting multiple noise audios, generating corresponding noise audio spectra, and marking the noise type to which the noise audio spectra belong according to the spectrogram of the noise audio spectra; and collecting multiple human voice audios, and generating corresponding human voice audio spectra; mixing the noise audio with the human voice audio to obtain mixed audio, and generating a corresponding mixed audio spectrum; using the mixed audio and the mixed audio spectrum as learning samples, and using the noise audio spectrum, the annotation of the noise type and the human voice audio spectrum as sample labels; using the collected multiple learning samples and the corresponding labels to train the convolutional neural network or the recursive neural network; The convolutional neural network or recurrent neural network performs the following operations: divide each learning sample into multiple audio frames, each audio frame has the same number of channels; generate a feature vector for each audio frame, the feature vector includes the values ​​of each channel of the audio frame; based on the feature vector, estimate the instantaneous function of the human voice audio spectrum and the noise audio spectrum in the audio frame; calculate the benchmark instantaneous function of the human voice audio spectrum and the noise audio spectrum according to the noise audio spectrum and the human voice audio spectrum in the label of the sample; determine the discrete points of the benchmark instantaneous function according to the time length of each audio frame and the offset time length between adjacent audio frames; and optimize the weights of the convolutional neural network or recurrent neural network with the benchmark instantaneous function at the discrete points as the target, so that the gap between the estimated instantaneous function and the benchmark instantaneous function at the corresponding discrete points is reduced.

2. The audio noise reduction system according to claim 1, characterized in that: It also includes an optimization module, which is configured to optimize the model architecture of the deep neural network separation module so that it can run on an NPU or a GPU.

3. The audio noise reduction system according to claim 2, characterized in that: The optimization module optimizes the model architecture by performing the following operations: Reducing the model volume of the deep neural network separation module through pruning and / or quantization operations; And / or, optimizing the model calculation path of the deep neural network separation module for the multi-core structure of the NPU or GPU.

4. The audio noise reduction system according to claim 1, characterized in that: The convolutional neural network or recurrent neural network performs operations based on an NPU or a GPU.

5. The audio noise reduction system according to claim 1, characterized in that: The instantaneous function is based on any one of the estimation methods of silence segment estimation, minimum mean square error estimation, minimum tracking algorithm, histogram noise estimation algorithm, quantile noise estimation, recursive average noise algorithm, and A-likelihood ratio method.

6. The audio noise reduction system according to claim 1, characterized in that: Before estimating the instantaneous function, mean and variance normalization operations are performed on the values ​​of each channel in the feature vector.

7. The audio noise reduction system according to claim 1, characterized in that: Collect learning samples corresponding to different noise types, and train the convolutional neural network or the recurrent neural network using a binary cross entropy loss function; The trained deep neural network separation module is configured with model calculation paths for different noise types; The adaptive mask generator adjusts the frequency difference between noise and human voice involved in generating the spectrum mask according to different model calculation paths.

8. The audio noise reduction system according to claim 7, characterized in that: The system runs in real time based on an NPU or a GPU, and the NPU or the GPU divides each n frames of input audio into an audio group, and performs reasoning on each audio group to infer whether a dynamic change of noise type occurs in the audio group; If so, when the adaptive mask generator generates a spectrum mask corresponding to the time feature of the audio group, the frequency difference between the corresponding noise and the human voice is adjusted to change the generation rule of the spectrum mask.

9. An audio noise reduction method, characterized in that: The following steps are involved: Input the audio signal to be denoised; Preprocessing the input audio signal to generate an audio spectrum; A pre-built deep neural network separation module is configured to: extract audio features from the audio spectrum, the audio features including time features and frequency features; Compressing the extracted audio features and generating a latent space representation based on the audio features; reconstructing and separating human voice and noise based on the latent space representation; Based on the frequency difference between noise and human voice, combined with the time characteristics, a temporal spectrum mask is generated to separate the human voice audio; Performing an inverse transformation on the frequency characteristics of the separated human voice audio signal to reconstruct the time domain audio data; The deep neural network separation module is trained by the following steps: collecting multiple noise audios, generating corresponding noise audio spectra, and marking the noise type to which the noise audio spectra belong according to the spectrogram of the noise audio spectra; and collecting multiple human voice audios, and generating corresponding human voice audio spectra; mixing the noise audio with the human voice audio to obtain mixed audio, and generating a corresponding mixed audio spectrum; using the mixed audio and the mixed audio spectrum as learning samples, and using the noise audio spectrum, the annotation of the noise type and the human voice audio spectrum as sample labels; using the collected multiple learning samples and the corresponding labels to train the convolutional neural network or the recursive neural network; The convolutional neural network or recurrent neural network performs the following operations: divide each learning sample into multiple audio frames, each audio frame has the same number of channels; generate a feature vector for each audio frame, the feature vector includes the values ​​of each channel of the audio frame; based on the feature vector, estimate the instantaneous function of the human voice audio spectrum and the noise audio spectrum in the audio frame; calculate the benchmark instantaneous function of the human voice audio spectrum and the noise audio spectrum according to the noise audio spectrum and the human voice audio spectrum in the label of the sample; determine the discrete points of the benchmark instantaneous function according to the time length of each audio frame and the offset time length between adjacent audio frames; and optimize the weights of the convolutional neural network or recurrent neural network with the benchmark instantaneous function at the discrete points as the target, so that the gap between the estimated instantaneous function and the benchmark instantaneous function at the corresponding discrete points is reduced.

10. The audio noise reduction method according to claim 9, characterized in that: The convolutional neural network or recurrent neural network performs operations based on an NPU or a GPU.

11. The audio noise reduction method according to claim 9, characterized in that: Before estimating the instantaneous function, firstly perform mean and variance normalization operations on the values ​​of each channel in the feature vector; The instantaneous function is based on any one of the estimation methods of silence segment estimation, minimum mean square error estimation, minimum tracking algorithm, histogram noise estimation algorithm, quantile noise estimation, recursive average noise algorithm, and A-likelihood ratio method.

12. The audio noise reduction method according to claim 9, characterized in that: Collect learning samples corresponding to different noise types, and train the convolutional neural network or the recurrent neural network using a binary cross entropy loss function; The trained deep neural network separation module is configured with model calculation paths for different noise types; Before generating the spectral mask, the frequency difference between the noise and the human voice involved in generating the spectral mask is adjusted according to different model calculation paths.

13. The audio noise reduction method according to claim 12, characterized in that: The deep neural network separation module is run in real time based on an NPU or a GPU, and the NPU or the GPU divides each n frames of input audio into an audio group, and performs reasoning on each audio group to infer whether a dynamic change of noise type occurs in the audio group; If so, when generating a spectrum mask corresponding to the time feature of the audio group, the frequency difference between the corresponding noise and the human voice is adjusted to change the generation rule of the spectrum mask.

14. A car audio system, characterized in that: It comprises a speaker and the audio noise reduction system as claimed in any one of claims 1 to 8, wherein the speaker is used to play the time domain audio reconstructed by the audio noise reduction system from the separated human voice audio signal.

Citation Information

Patent Citations

  • Microphone array speech denoising enhancement method with low signal-to-noise ratio effect

    CN110827847A

  • Voice signal recognition method and device, electronic equipment and storage medium

    CN113571063A