Wind noise removal in sound signals acquired by vehicle microphones

EP4804186A1Pending Publication Date: 2026-09-09VALEO TELEMATIK & AKUSTIK GMBH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
EP2025161883
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

In vehicle applications involving external microphones, wind noise significantly affects the quality of captured audio signals (i.e. lower signal to noise ratio (SNR)), especially at higher speeds.

Benefits of technology

[0011]It is also provided a computer-implemented method of machine-learning, for training the DNN. The method of machine-learning comprises providing a training dataset of training examples, each training example comprising a log-power spectrum of an example signal representing a sound acquired by a vehicle microphone and a log-power spectrum representing the example signal without wind noise. The method of machine-learning further comprises training the DNN based on the training dataset. The training comprises minimizing a loss. The loss penalizes, for each training example, a disparity. The disparity is between a cleaned log-power spectrum outputted by the DNN for an input log-power spectrum of the example signal of the training example and the log-power spectrum of the training example representing the example signal without wind noise. The method of machine-learning may be referred to as "the learning method".

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF0001
    Figure IMGF0001
  • Figure IMGF0002
    Figure IMGF0002
  • Figure IMGF0003
    Figure IMGF0003
Patent Text Reader

Abstract

The disclosure notably relates to a computer-implemented method for processing together two or more sound signals. Each signal represents a sound acquired at a same time frame by a respective microphone of a vehicle (i.e. all signals are acquired at a same time frame, each signal being acquired by a respective microphone of the vehicle). The vehicle comprises at least two microphones. The method comprises, for each microphone of the vehicle, applying a signal-cleaning function to the sound signal acquired by the microphone. The signal-cleaning function is a function configured to take as input a sound signal and to remove wind noise in the input sound signal. The method thereby outputs a clean sound signal for each respective sound signal acquired by each respective microphone.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to the field of computer programs and systems, and more specifically to methods, systems, programs and a neural network for wind noise removal in sound signals acquired by vehicle microphones.BACKGROUND

[0002] In vehicle applications involving external microphones, wind noise significantly affects the quality of captured audio signals (i.e. lower signal to noise ratio (SNR)), especially at higher speeds. This background noise can greatly impair the performance of sound event detection systems in vehicles. For example, systems designed to detect and locate siren sources may struggle to accurately identify siren signals amidst wind noise, particularly at higher speeds, resulting in compromised performance in detecting and localizing emergency vehicles.

[0003] Furthermore, mitigating wind noise poses a significant challenge when utilizing external microphones. Due to wind noise heavily masking most frequency bands, conventional signal processing methods aim at separating and removing wind frequency components and may thus inadvertently eliminate crucial frequency features of the target signals. Consequently, traditional noise reduction techniques often produce highly distorted output signals when attempting to reduce wind noise in exterior microphones.

[0004] Moreover, conventional methods for wind noise reduction in microphone arrays typically use multi-channel input and single-channel output (MISO) configurations. These methods often employ sound level comparison and / or beamforming techniques across multiple input channels to optimize the output signal and minimize wind noise. While MISO models may be effective for single-input applications such as siren detection systems, they are unsuitable as a preprocessing stage for multi-input applications like siren localization systems. In other words, MISO configurations eliminate the spatial information of channels by merging them into one, which prevents the use of this information for sound source localization. Existing multi-input siren detection and localization systems can suffer from wind noise interference at higher speeds.SUMMARY

[0005] It is therefore provided a computer-implemented method for processing together two or more sound signals. Each signal represents a sound acquired at a same time frame by a respective microphone of a vehicle (i.e. all signals are acquired at a same time frame, each signal being acquired by a respective microphone of the vehicle). The vehicle comprises at least two microphones. The method comprises, for each microphone of the vehicle, applying a signal-cleaning function to the sound signal acquired by the microphone. The signal-cleaning function is a function configured to take as input a sound signal and to remove wind noise in the input sound signal. The method thereby outputs a clean sound signal for each respective sound signal acquired by each respective microphone. The method may be referred to as "the processing method".

[0006] The function may comprise a Deep Neural Network (DNN). The DNN is configured and trained to take as input a log-power spectrum of a sound signal acquired by a vehicle microphone and to output a log-power spectrum of the sound signal without wind noise.

[0007] The function may further comprise a pre-processing module and a post-processing module. The pre-processing module is configured to take as input a sound signal acquired by a vehicle microphone and to output a log-power spectrum of the sound signal to be inputted to the DNN. The post-processing module is configured to take as input the log-power spectrum of the sound signal outputted by the DNN and to reconstruct a corresponding sound signal based on phase data extracted from the input sound signal.

[0008] The DNN may be configured to take as input a matrix. The matrix is formed by: a log-power spectrum of the sound signal acquired by a respective microphone at said same time frame, one or more log-power spectrum each of a sound signal acquired by the respective microphone at a previous time frame, and, one or more log-power spectrum each of a sound signal acquired by the respective microphone at a next time frame.

[0009] The one or more log-power spectrum each of a sound signal acquired by the respective microphone at a previous time frame may consist of three log-power spectra corresponding to three previous time frames. The one or more log-power spectrum each of a sound signal acquired by the respective microphone at a next time frame may consist of three log-power spectra corresponding to three next time frames.

[0010] The processing method may be performed in real-time during travel of the vehicle. The vehicle may travel at speed such that a signal to wind noise ratio is smaller than -20 decibels.

[0011] It is also provided a computer-implemented method of machine-learning, for training the DNN. The method of machine-learning comprises providing a training dataset of training examples, each training example comprising a log-power spectrum of an example signal representing a sound acquired by a vehicle microphone and a log-power spectrum representing the example signal without wind noise. The method of machine-learning further comprises training the DNN based on the training dataset. The training comprises minimizing a loss. The loss penalizes, for each training example, a disparity. The disparity is between a cleaned log-power spectrum outputted by the DNN for an input log-power spectrum of the example signal of the training example and the log-power spectrum of the training example representing the example signal without wind noise. The method of machine-learning may be referred to as "the learning method".

[0012] The training may comprise a residual learning.

[0013] It is also provided a DNN obtainable according to the learning method (i.e. having the same weights and architecture as a DNN trained by the learning method), for example a DNN obtained according to the learning method (i.e. the exact DNN that results from the learning method, with its weights / parameters set by the training according to the training method).

[0014] It is further provided a computer program comprising instructions for performing the processing method and / or the learning method.

[0015] It is further provided a computer data readable storage medium having recorded thereon the computer program and / or the DNN.

[0016] It is further provided a computer system comprising a processor coupled to a memory, the memory having recorded thereon the computer program and / or the DNN.

[0017] It is further provided a device comprising a data storage medium having recorded thereon the computer program and / or the neural network.

[0018] The device may form or serve as a non-transitory computer-readable medium, for example on a SaaS (Software as a service) or other server, or a cloud based platform, or the like. The device may alternatively comprise a processor coupled to the data storage medium. The device may thus form a computer system in whole or in part (e.g. the device is a subsystem of the overall system). The system may further comprise a graphical user interface coupled to the processor.

[0019] The computer system may be a vehicle on-board computer system coupled with microphones of the vehicle.

[0020] It is also provided a vehicle equipped with the computer system.BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Non-limiting examples will now be described in reference to the accompanying drawings, where: FIG.s 1 to 6 illustrate the methods; and FIG. 7 shows an example of the system. DETAILED DESCRIPTION

[0022] It is provided a computer-implemented method for processing together two or more sound signals. Each signal represents a sound acquired at a same time frame by a respective microphone of a vehicle (i.e. all signals are acquired during a same time frame, each signal being acquired by a respective microphone of the vehicle). The vehicle comprises at least two microphones. The method comprises, for each microphone of the vehicle, applying a signal-cleaning function to the sound signal acquired by the microphone. The signal-cleaning function is a function configured to take as input a sound signal and to remove wind noise in the input sound signal. The method thereby outputs a clean sound signal for each respective sound signal acquired by each respective microphone. As previously said, the method may be referred to as "the processing method".

[0023] The processing method forms an improved solution for wind noise reduction in sound signals acquired by vehicle microphones.

[0024] Indeed, as previously explained, conventional methods for noise reduction in sound signals acquired by vehicle microphones use a so-called "MISO" methodology (or configuration). The principles of this methodology are as follows. The vehicle comprises several microphones (for example organized as microphones arrays), and at each time frame (e.g. at regular time frames) all microphones acquire a sound signal (i.e. there are, for each time frame, one sound signal per microphone). The sound signal of a respective microphone may also be referred to as "input channel" or "microphone input channel". There is thus one input channel per microphone. MISO then works as follows: it takes as input all the input channels, processes them altogether to remove wind noise as previously described, and output a single clean output channel (i.e. a single output clean (free of wind noise) signal). Thus, MISO stands for "Multi-Input-Single-Output". This implies that MISO methods loose spatial information of channels by merging them into a single one. MISO may thus be effective for single-input applications, such as siren detection, but they are unsuitable for a preprocessing stage for multi-input applications, like siren localization models.

[0025] To the contrary, the present processing method implements a SISO (Single-Input-Single-Output) methodology or configuration: for each signal acquired by a respective one of the vehicle microphones, the method applies a signal-cleaning function (i.e. a function that reduces / removes wind noise in the signal) to that signal, and outputs a respective clean signal (i.e. without wind noise, or with substantially reduced wind noise) for that input signal. In other words, there are as many output cleaned signals (or output channels) as there are input signals (or input channels). In other words, the wind noise is removed from each individual channel, and there is no merging of the channels. The method may thus be used as a pre-processing step for multi-input applications like siren localization systems: the signals cleaned by the method are processable by any siren localization model to detect a siren based on these signals. Siren detection and localization systems implemented in vehicles indeed rely on the spatial information from exterior microphones to detect and pinpoint emergency vehicles. However, at higher speeds, wind noise severely degrades the siren signals captured by the microphones, hampering the ability of the system to accurately detect and locate the sources of sirens. The proposed methods effectively mitigate the detrimental impact of wind noise on target siren signals, ensuring that clean microphone signals are available for use by detection and localization systems.

[0026] FIG. 1 illustrates the differences between the MISO and SISO methodologies.

[0027] For each individual signal acquired by a respective, the signal function may comprise a DNN configured and trained to take as input a log-power spectrum of a sound signal acquired by the vehicle microphone and to output a log-power spectrum of the sound signal without wind noise. In other words, for each individual input channel, the method may use this DNN to clean the channel. As previously outlined, it is also proposed a computer-implemented method of machine-learning, for training the DNN.

[0028] Therefore, the present disclosure proposes an end-to-end deep learning-based approach for wind noise reduction in exterior vehicle microphones. Exterior microphones often capture target sound signals that are highly degraded by strong wind noise, especially at high speeds (e.g. such that signal to wind noise ratio is lower than -20 decibels). Conventional methods can lead to significant distortion in the enhanced signals, because, as previously said, for cleaning the signals, they may remove significant / crucial frequency components. The present disclosure solves this problem by providing and using this DNN that predicts clean signals from noisy ones without eliminating the spatial information necessary for sound event localization. The DNN is trained to extract optimal features from the spectra of noisy signals (highly degraded by wind noise), estimate optimal filters, and predict (i.e., reconstruct) the spectra of clean target signals. The proposed DNN-based methodology utilizes the log-power spectra of a microphone signal, effectively removing wind noise from the signal power and predicting the log-power spectra of the desired clean signal at the output. The log-power spectra accurately represent the formants of audio signals (e.g., siren signals) as peak regions in the power spectrum. The clean signal may then be reconstructed based on the predicted spectra and the phases extracted from the input signal.

[0029] The processing method is a method for processing together two or more sound signals. In specific, the method takes as input each sound signal and processes it individually, wherein by "processes" it is meant that the method removes wind noise in the sound signal. The method outputs the clean signals, i.e. the signals with the wind noise removed. All sound signals are relative to a same time frame, i.e. they each represent a respective sound acquired by a respective microphone of the vehicle, but each sound is acquired during the said same time frame. Each signal is thus acquired by one respective microphone of the vehicle (i.e. there is one signal per microphone of the vehicle). The vehicle comprises several microphones (at least two). The microphones may be mounted on the vehicle as a distribution arrays of microphones, as known in the art. Each microphone is a vehicle exterior microphone, and may be any exterior microphone model known in the art. A sound signal may also be referred to as "audio signal". The vehicle may be any type of land vehicle such as a car, a bus, a truck, or a motorbike. Any sound herein may be a siren sound. In particular the DNN discussed herein may be trained on training examples that stem from siren sound signals.

[0030] As outlined above, the method is performed for a time frame. In other words, the microphones acquire sound signals (one per microphone) during the same time frame. Each signal is thus a discrete time series of values, where each value corresponds to a specific time within the time frame and represents an amplitude of the sound acquired by the respective microphone at that time. The amplitude may be an air pressure or a voltage. To process the signals, the time frame may be divided into overlapping intervals T i = [t i , t i + T] , , where t i is the starting sample of the i-th time frame, and T is the duration of the frame. For a given microphone M, the signal may be mathematically represented as x M = (x T1 , x T2 ... , x Tn ), where T i denotes the i-th time interval. The signals may be further processed using a windowing function, applied to each time frame to mitigate spectral leakage during subsequent analysis. The overlapping of intervals ensures continuity and effective segmentation for analysis. The method processes the signals for all microphones for this time frame. It is to be understood that the method may be repeated for several (e.g. consecutive time frames). In particular the method may be performed in real-time during a travel of the vehicle, where continuously (regularly) the microphones provide sound signals corresponding to successive time frames. The vehicle may travel at speed such that a signal to wind noise ratio (i.e. for the signals considered in the methods) is smaller than -20 decibels, which corresponds to a high speed. Signal-to-noise ratio (also referred to as SNR or S / N) is a measure that compares the level of a desired signal to the level of background noise. SNR is defined as the ratio of signal power to noise power (here wind noise), often expressed in decibels. A ratio higher than 1:1 (greater than 0 dB) indicates more signal than noise. The method thus handles situations, typically high-speed situations, where wind noise is so significant that the signal to noise ratio is smaller than -20 decibels.

[0031] The processing method comprises, for each microphone of the vehicle, applying a signal-cleaning function to the sound signal acquired by the microphone. The signal-cleaning function is a function configured to take as input a sound signal and to remove wind noise in the input sound signal. By "remove wind noise" it is meant that the function performs wind noise reduction / removal, i.e. the function determines a signal that corresponds to the original signal but without (or, at least, with significantly less) wind noise. The processing method thereby outputs a clean sound signal (i.e. the sound signal without the wind noise) for each respective sound signal acquired by each respective microphone

[0032] The signal cleaning function may comprise a Deep Neural Network (DNN). The Deep Neural Network is configured and trained to (i.e. it has the architecture and training to) take as input a log-power spectrum of a sound signal acquired by a vehicle microphone and to output a log-power spectrum of the sound signal without wind noise.

[0033] As known per se, the log-power spectrum of a signal is a representation of the content of the signal power in the frequency domain, expressed on a logarithmic scale. The power spectrum of a signal describes how the power (or energy) of the signal is distributed across different frequencies. Mathematically, it is the squared magnitude of the signal's Fourier transform, which represents the amplitude of different frequency components. In other words, for a signal x(t), its power spectrum P(f) can be expressed as: P f = F x t f 2 where [x(t)}(f) is the Fourier transform of x(t), and P(f) is the power at frequence f. To make the power spectrum more manageable, especially when dealing with signals that span a large range of magnitudes, the power spectrum may be transformed to a logarithmic scale. This compression helps visualize data more effectively, particularly when there are very large and very small power values. The log-power spectrum is given by: L f = log P f = log F x t f 2 .

[0034] A base-10 logarithm may also be used as follows: L f = 10 log 10 P f .

[0035] The advantages of using the log-power spectrum of the microphone signals as input to the DNN-based model of the present disclosure for wind noise reduction are as follows: Improved Dynamic Range Representation: The log power spectrum compresses the dynamic range of the raw power spectrum, making it easier for the neural network to learn patterns. This compression ensures that both strong siren signals and weaker noise components (such as wind noise) are represented in a range that the network can process effectively; Enhanced Discrimination Between Noise and Signal: Siren signals have distinct harmonic structures and spectral energy patterns, while wind noise is typically broadband and lacks such structure. The log power spectrum accentuates these differences, helping the neural network to distinguish between the structured siren signal and unstructured wind noise more effectively; Robustness to Amplitude Variations: Logarithmic scaling makes the feature representation less sensitive to variations in amplitude, which can occur due to changes in the distance between the microphone and the siren or environmental conditions. This robustness improves the generalization of the DNN to different recording scenarios.

[0036] The signal cleaning function may further comprise a-preprocessing module and a post-processing module. The pre-processing module is configured to take as input a sound signal acquired by a vehicle microphone and to output a log-power spectrum (e.g. of the successive time frames of) the sound signal to be inputted to the DNN. The post-processing module is configured to take as input the log-power spectrum of (e.g. of the corresponding time frame of) the sound signal outputted by the DNN and to reconstruct a corresponding sound signal based on phase data extracted from the input sound signal. The pre-processing may for that extract (compute) the phase of the input signal (by any known method), which is then kept on memory and reuse to reconstruct the sound signal using the phase of the input signal and its cleaned log-power spectrum. Computing the log-power spectrum of the input signal by the pre-processing module may be performed by any known method for this purpose. Likewise, reconstruction of the signal using the input signal phase and the clean low-power spectrum by the post-processing method may be performed by any known method for this purpose.

[0037] The following should be noted, to understand while the proposed disclosure does not perform noise reduction in the phase as well. When a target signal is affected by an additive background noise (like wind noise) both magnitude and phase information may be degraded. However, there are some points about the wind noise reduction focusing on the magnitude of the signal: Preservation of Relative Phase for Localization Accuracy: Siren localization relies on inter-microphone phase differences to estimate the direction of arrival (DOA). Reusing the noisy phase ensures that these relative phase relationships are preserved, which is critical for accurate localization. Enhancing the phase could introduce distortions that compromise this directional information; Magnitude Enhancement Suffices for Wind Noise Reduction: A siren's perceptually significant features, such as its harmonic structure and spectral energy, are predominantly captured in the magnitude spectrum. Enhancing the magnitude improves the signal-to-noise ratio (SNR) effectively, while the noisy phase is typically adequate for reconstructing the signal without impairing localization performance; Wind Noise Characteristics and Impact on Phase: Wind noise is typically broadband and low-frequency, leading to a more random phase distribution in the noisy signal. However, the siren signal (often a modulated tone or distinct spectral signature) has a more structured phase component.

[0038] Even when the phase is noisy, the siren's distinctive harmonic structure is primarily represented in the magnitude spectrum, making magnitude enhancement sufficient for separating the siren signal from wind noise.

[0039] The DNN may be configured to take as input a matrix. The matrix is formed by: a log-power spectrum of the sound signal acquired by a respective microphone at (during) said same time frame; one or more log-power spectrum each of a sound signal acquired by the respective microphone at (during) a previous time frame; and and one or more log-power spectrum each of a sound signal acquired by the respective microphone at (during) a next time frame.

[0040] The one or more log-power spectrum each of a sound signal acquired by the respective microphone at a previous time frame may consist in three log-power spectra corresponding to three previous times (e.g. the three time frames before to the said same time frame (also referred to as "current time frame")). The one or more log-power spectrum each of a sound signal acquired by the respective microphone at a next time frame may consist in three log-power spectra corresponding to three next times (e.g. the three time frames after to the said same time frame). The method may comprise, for each sound signal frame to which the DNN is applied, applying first the pre-processing module to that frame and also to the next and previous frames (i.e. the pre-processing module is configured to compute the log-power spectrum for all the frames), to obtain their log-power spectra and form the matrix, and then inputting the matrix to the DNN.

[0041] In other words, the matrix may have 7 rows: 3 rows each encoding the log-power spectrum of the sound signal acquired by the respective microphone during one of the said previous time frames, 1 row encoding the log-power spectrum of the sound signal acquired by the respective microphone at said same time frame, and 3 rows each encoding the log-power spectrum of the sound signal acquired by the respective microphone during one of the said next time frames. Each row may have a size of 196 and comprise the values of the log-power spectrum it encodes for 196 frequency bins (e.g. distributed between 200Hz and 2000Hz).

[0042] The output of the DNN may however be a single vector encoding the cleaned (i.e. without or with reduced wind noise) log-power spectrum of the sound-signal acquired during the said same time frame. This vector may be of size 196 and comprise the log-power spectrum values for 196 frequency bins (e.g. distributed between 200Hz and 2000Hz).

[0043] The DNN may have a CNN architecture. CNN stands for Convolutional Neural Network, as known in the field of machine-learning. The CNN architecture may comprise 2D convolutional layers (for example two 2D convolutional layers), followed by residual blocks (for example five residual blocks each of three convolutional layers), followed by a 2D convolutional layer, followed by a flatten layer, followed by fully connected layers (for example three fully connected layers). The CNN may use a RELU activation function, except for the last layer which may use a linear activation function. In the first 2D convolutional layers, the convolutions may be applied on the Fourier transform (frequency bins) of each frame over all the microphones. In the subsequence 2D convolutional layers (after the residual blocks), the convolutions may be applied on both the Fourier transforms of the frames and on the frames themselves.

[0044] It is also provided a computer-implemented method of machine-learning, for training the DNN. As previously outlined, this method may be referred to as "the learning method".

[0045] The learning method is a method of machine-learning of a model, which is the DNN in the present disclosure. As known per se from the field of machine-learning, the processing of an input by a model includes applying operations to the input, the operations being defined by data including weight values or parameters. Learning a model (e.g. a neural network or a regressor) thus includes determining values of the weights / parameters based on a dataset configured for such learning (e.g. optimizing values of the weights / parameters by minimizing a loss function between the estimated and desired outputs, based on a dataset configured for such learning), such a dataset being possibly referred to as a learning dataset or a training dataset. For that, the dataset includes data pieces each forming a respective training sample or training example. The training samples / examples represent the diversity of the situations where the model is to be used after being learnt. Any training dataset herein may comprise a number of training samples / examples higher than 1000, 10000, 100000, or 1000000. In the context of the present disclosure, by "learning a model based on a dataset", it is meant that the dataset is a learning / training dataset of the model, based on which the values of the weights / parameters are set. In the present disclosure, the training dataset is the dataset of training examples, on which the DNN is trained.

[0046] As known per se from machine-learning, a neural network may be defined by its architecture, parameters, and hyperparameters. The architecture consists of layers, starting with the input layer whose neuron count may be determined by the dimensionality of the input data. This layer is followed by several hidden layers with a given number of neurons and activation functions. These layers and neurons define the network's depth and width, while the activation functions may introduce nonlinearity into the model. The output layer may have as many neurons as the variables in the output data. The interconnections between these layers defines the topology of the neural network. The parameters of the neural network are the learnable weights and biases, which are determined in the training process. In contrast, the hyperparameters are pre-defined settings that are not learned from the training data. These encompasses the number of hidden layers, neurons per layer and much more. To train a neural network, at least two settings may be defined. First, a loss function, which is a metric that measures the error between the training data and the model's prediction, such as the mean square error (MSE). Second, an optimizer, which modifies the model's weights and biases during the training process to minimize the loss function. Each optimizer has its own set of hyperparameters.

[0047] The learning method comprises providing a training dataset of training examples. Each training example comprises a log-power spectrum of an example signal representing a sound acquired by a vehicle microphone and a log-power spectrum representing the example signal without wind noise. The log-power-spectrum may of the example signal may for example be provided as a matrix like the one previously discussed, i.e. with 7 rows: 3 rows each encoding the log-power spectrum of the sound signal acquired by the microphone during one a previous time frame, 1 row encoding the log-power spectrum of the sound signal, and 3 rows each encoding the log-power spectrum of the sound signal acquired by the microphone during one next time frame. Each row may have a size of 196 and comprise the values of the log-power spectrum it encodes for 196 frequency bins (e.g. distributed between 200Hz and 2000Hz). The matrix has depth 1, i.e. the input matrix is a 3D input of dimensions (7, 196, 1). The log-power spectrum representing the example signal without wind noise may be of the same type as the output of the DNN i.e. a single vector encoding the clean log-power spectrum (without wind noise) of the sound-signal of the training example. As previously discussed, this vector may be of size 196 and comprise the log-power spectrum values for 196 frequency bins (e.g. distributed between 200Hz and 2000Hz). For each training example, the DNN thus takes as input this matrix and is trained to output a single vector similar (in its values) to the single vector of the clean log-power spectrum (without wind noise) of the sound-signal of the training example.

[0048] Thus, the training dataset comprises, for each training example, data (i.e. the matrix discussed above) encoding the non-clean sound signal of that training example and data (i.e. the single vector discussed above) encoding the corresponding clean signal. It is to be understood that, to increase robustness of the training, the training dataset may comprise (e.g. in a small proportion) training examples where the non-clean signal is actually already clean. In any case, the DNN has in the training examples references of clean signals, without wind noise, for example noisy sound signals, and thus learns to determine a clean signal for a corresponding input noisy one. In the training examples, and as known in the art of machine-learning, the examples sound signals may correspond to a sufficient variability of vehicle speed and / or wind strength.

[0049] Providing the training dataset may comprise creating the training dataset by creating the training examples (or at least a part thereof). This may include accessing any suitable database of sound signals acquired by vehicle microphones. Databases exist (or may be created based on existing publicly available sound data) which comprise sound signals (both clean signals and signals with wind noise) acquired by vehicle microphones at different directions and with different wind noises corresponding to different vehicle speeds. In such databases, for each noisy signal, the corresponding clean signal (without wind noise) is available. This allows to obtain the pairs of noisy and clean example signals, as well as the pairs of clean and clean example signals to obtain the training examples. In implementations, to train the DNN, the inventors used a dataset based on real-world data recorded in various environments using microphones mounted on different zones of test vehicles. Additionally, they incorporated sound data from publicly available sources, such as Freesound and the Kaggle database.

[0050] Creating the training dataset may comprise extracting a sufficient number and variability of all these pairs from the database and forming the training examples by obtaining their log-power spectra (for example by applying the pre-processing module) and forming the corresponding input matrices and output single vectors, which were previously discussed. Providing the training dataset may alternatively comprise retrieving (e.g. downloading) the already-created training examples (or at least a part thereof) from a (e.g. distant) memory or database or server where they have been stored further to their creation (e.g. with the creation methodology discussed above).

[0051] The learning method then comprises training the DNN based on the training dataset. The training comprises minimizing a loss. The loss penalizes, for each training example, a disparity. The disparity is a disparity between a cleaned log-power spectrum outputted by the DNN for an input log-power spectrum of the example signal of the training example and the log-power spectrum of the training example representing the example signal without wind noise. In other words: 1) each training example comprises the log-power spectrum of an example noisy (sometimes (but in a small proportion) already clean as previously explained) sound signal (e.g. in the form of the above-discussed matrix for example) and the corresponding clean log-power spectrum (i.e. the spectrum of that signal without noise) of that signal (i.e. in the form of the single vector previously discussed); 2) the DNN outputs, for the log-power spectrum of the noisy signal taken as input, a prediction of a clean log-power spectrum; and 3) the disparity is between that prediction and the clean log-power spectrum of the training example (also referred to as "target"). The disparity may be a distance, such that the loss has large value as long as the prediction is not sufficiently similar to the target. During training, parameters of the DNN are modified while the value of the loss remains to large on the training dataset. The loss may be a mean-square-error (MSE) and the training may be based on gradient backpropagation as known in the art of machine-learning.

[0052] The training may comprise a residual learning. This may comprise using residual blocks in the DNN architecture (between the convolutional layers as previously discussed) and applying batch normalization after each layer during the backpropagation process. This improves the training. Indeed, training very deep neural networks presents the challenge of enhancing gradient influence in deeper layers. Batch normalization applied after each layer, and residual learning implemented in bottleneck blocks, improve gradient flow during the backpropagation process. This technique ensures that the DNN trains effectively, even with a deep architecture, by maintaining the stability of gradient magnitudes and enabling the learning of complex features.

[0053] FIG. 2 shows an illustration of the DNN architecture and of its training according to implementations of the learning method. The learning rate may be 0.001 and the optimizer may be the well-known Adam optimizer in these examples. The proposed DNN in these implementations thus functions as a regression model, with the mean squared error (MSE) serving as the loss function between the predicted output and the desired output. It comprises several convolutional and fully connected layers, which are optimized during training to predict the log-power spectra of the corresponding clean signal. The parameters of the convolutional layers are optimized to estimate peak regions within highly noisy power spectral coefficients and subsequently reconstruct the clean signal based on the predicted power spectra.

[0054] In summary, the methods offer the following advantages. The conventional approach relies on traditional signal processing methods like cross-correlation and level comparison followed by beamforming for wind noise reduction, whereas the proposed model is an end-to-end DNN-based model capable of predicting a clean target signal from a severely noisy input signal degraded by wind noise. The conventional model is a MISO model and cannot serve as a preprocessing stage for models requiring information from multiple input channels (e.g., siren localization models) for wind noise reduction. Conversely, the proposed model is a SISO model, preserving individual microphone signal information, such as phase information, after enhancement, thereby serving as a preprocessing stage for multi-input approaches. The proposed model effectively removes severe wind noise from exterior microphones when driving at high speeds. The use of recurrent neural networks (specifically, long short-term memory or LSTM layers) in the prior model can increase system complexity. In contrast, the proposed DNN model is based on convolutional neural networks (CNNs), offering less complexity. CNNs utilize weight sharing to analyse input data, resulting in fewer trainable parameters compared to LSTM layers with similar neurons. Using highly processed data as input to DNNs may diminish performance in extracting optimal features from raw data during training. The proposed methods employ more raw data (e.g., log-power spectra) as DNN inputs, facilitating estimation and extraction of optimal features for wind noise reduction during training.

[0055] FIG. 3 illustrates the performance of the proposed methods in reducing wind noise from the target siren signal. The upper plot shows the spectrogram of the noisy siren signal (Tool siren) at a speed of 45 MPH and a distance of 100 meters. The bottom plot shows the spectrogram of the signal enhanced by the proposed SISO wind noise reduction model. The original and enhanced SNRs are, respectively, -15.30 dB and +9.92 dB, resulting in an SNR improvement of +25.23 dB.

[0056] FIG. 4 also illustrates the performance of the proposed methods in reducing wind noise from the target siren signal. Upper plot shows the spectrogram of the noisy siren signal (HiLo siren) at a speed of 75 MPH and a distance of 100 meters. Bottom plot shows the spectrogram of the signal enhanced by the proposed SISO wind noise reduction model. The original and enhanced SNRs are, respectively, -23.90 dB and +0.70 dB, resulting in an SNR improvement of +24.60 dB.

[0057] FIG. 5 illustrates validation results of the above-described training.

[0058] FIG. 6 illustrates a comparison between the performances of different siren detection models for difference types for a noisy original dataset of siren signals ("original dataset") and for the corresponding dataset of signals cleaned by the DNN and the proposed methods ("enhanced dataset").

[0059] It is also provided a DNN obtainable according to the learning method (i.e. having the same weights and architecture as a DNN trained by the learning method), for example a DNN obtained according to the learning method (i.e. the exact DNN that results from the learning method, with its weights / parameters set by the training according to the training method.

[0060] The methods may be integrated into a same computer-implemented method of sound signal processing which comprises the learning method (offline stage) and then the processing method (online stage).

[0061] The methods are computer-implemented. This means that steps (or substantially all the steps) of the methods are executed by at least one computer, or any system alike. Thus, steps of the methods are performed by the computer, possibly fully automatically, or, semi-automatically. In examples, the triggering of at least some of the steps of the methods may be performed through user-computer interaction. The level of user-computer interaction required may depend on the level of automatism foreseen and put in balance with the need to implement user's wishes. In examples, this level may be user-defined and / or pre-defined.

[0062] A typical example of computer-implementation of a method is to perform the method with a system adapted for this purpose. The system may comprise a processor coupled to a memory and a graphical user interface (GUI), the memory having recorded thereon a computer program comprising instructions for performing the method. The memory may also store a database. The memory is any hardware adapted for such storage, possibly comprising several physical distinct parts (e.g. one for the program, and possibly one for the database).

[0063] FIG. 7 shows an example of the system, wherein the system is a client computer system, e.g. a vehicle on-board computer.

[0064] The client computer of the example comprises a central processing unit (CPU) 1010 connected to an internal communication BUS 1000, a random access memory (RAM) 1070 also connected to the BUS. The client computer is further provided with a graphical processing unit (GPU) 1110 which is associated with a video random access memory 1100 connected to the BUS. Video RAM 1100 is also known in the art as frame buffer. A mass storage device controller 1020 manages accesses to a mass memory device, such as hard drive 1030. Mass memory devices suitable for tangibly embodying computer program instructions and data include all forms of nonvolatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks. Any of the foregoing may be supplemented by, or incorporated in, specially designed ASICs (application-specific integrated circuits). A network adapter 1050 manages accesses to a network 1060. The client computer may also include a haptic device 1090 such as cursor control device, a keyboard or the like. A cursor control device is used in the client computer to permit the user to selectively position a cursor at any desired location on display 1080. In addition, the cursor control device allows the user to select various commands, and input control signals. The cursor control device includes a number of signal generation devices for input control signals to system. Typically, a cursor control device may be a mouse, the button of the mouse being used to generate the signals. Alternatively or additionally, the client computer system may comprise a sensitive pad, and / or a sensitive screen. The computer may also be connected to the microphones of the vehicle and thus configured to receive these signals and process them according to the methods (possibly after suitable analog to digital signal conversion).

[0065] The computer program may comprise instructions executable by a computer, the instructions comprising means for causing the above system to perform the method. The program may be recordable on any data storage medium, including the memory of the system. The program may for example be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. The program may be implemented as an apparatus, for example a product tangibly embodied in a machine-readable storage device for execution by a programmable processor. Method steps may be performed by a programmable processor executing a program of instructions to perform functions of the method by operating on input data and generating output. The processor may thus be programmable and coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. The application program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language if desired. In any case, the language may be a compiled or interpreted language. The program may be a full installation program or an update program. Application of the program on the system results in any case in instructions for performing the method. The computer program may alternatively be stored and executed on a server of a cloud computing environment, the server being in communication across a network with one or more clients. In such a case a processing unit executes the instructions comprised by the program, thereby causing the method to be performed on the cloud computing environment.

Examples

Embodiment Construction

[0022]It is provided a computer-implemented method for processing together two or more sound signals. Each signal represents a sound acquired at a same time frame by a respective microphone of a vehicle (i.e. all signals are acquired during a same time frame, each signal being acquired by a respective microphone of the vehicle). The vehicle comprises at least two microphones. The method comprises, for each microphone of the vehicle, applying a signal-cleaning function to the sound signal acquired by the microphone. The signal-cleaning function is a function configured to take as input a sound signal and to remove wind noise in the input sound signal. The method thereby outputs a clean sound signal for each respective sound signal acquired by each respective microphone. As previously said, the method may be referred to as "the processing method".

[0023]The processing method forms an improved solution for wind noise reduction in sound signals acquired by vehicle microphones.

[0024]Indee...

Claims

1. A computer-implemented method for processing together two or more sound signals each representing a sound acquired at a same time frame by a respective microphone of a vehicle, the vehicle comprising at least two microphones, the method comprising: - for each microphone of the vehicle, applying a signal-cleaning function to the sound signal acquired by the microphone, the signal-cleaning function being a function configured to take as input a sound signal and to remove wind noise in the input sound signal, the method thereby outputting a clean sound signal for each respective sound signal acquired by each respective microphone.

2. The method of claim 1, wherein the function comprises a Deep Neural Network (DNN) configured and trained to take as input a log-power spectrum of a sound signal acquired by a vehicle microphone and to output a log-power spectrum of the sound signal without wind noise.

3. The method of claim 2, wherein the function further comprises: - a pre-processing module configured to take as input a sound signal acquired by a vehicle microphone and to output a log-power spectrum of the sound signal to be inputted to the DNN; and - a post-processing module configured to take as input the log-power spectrum of the sound signal outputted by the DNN and to reconstruct a corresponding sound signal based on phase data extracted from the input sound signal.

4. The method of claim 2 or 3, wherein the DNN is configured to take as input a matrix formed by: - a log-power spectrum of the sound signal acquired by a respective microphone at said same time frame; - one or more log-power spectrum each of a sound signal acquired by the respective microphone at a previous time frame; and - one or more log-power spectrum each of a sound signal acquired by the respective microphone at a next time frame.

5. The method of claim 4, wherein: - the one or more log-power spectrum each of a sound signal acquired by the respective microphone at a previous time frame consist in three log-power spectra corresponding to three previous times; and - the one or more log-power spectrum each of a sound signal acquired by the respective microphone at a next time frame consist in three log-power spectra corresponding to three next times.

6. The method of any one of claims 1 to 5, wherein the method is performed in real-time during travel of the vehicle.

7. The method of claim 6, wherein the vehicle travels at speed such that a signal to wind noise ratio is smaller than -20 decibels.

8. A computer-implemented method of machine-learning, for training a Deep Neural Network (DNN) comprised in the signal-cleaning function of any one of claims 1 to 7, the DNN being configured to take as input a log-power spectrum of a sound signal acquired by a vehicle microphone and to output a log-power spectrum of the sound signal without wind noise, the method comprising: - providing a training dataset of training examples, each training example comprising a log-power spectrum of an example signal representing a sound acquired by a vehicle microphone and a log-power spectrum representing the example signal without wind noise; and - training the DNN based on the training dataset, the training comprising minimizing a loss which penalizes, for each training example, a disparity between a cleaned log-power spectrum outputted by the DNN for an input log-power spectrum of the example signal of the training example and the log-power spectrum of the training example representing the example signal without wind noise.

9. The method of claim 8, wherein the training comprises a residual learning.

10. A Deep Neural Network obtainable according to the method of claim 8 or 9.

11. A computer program comprising instructions which, when executed by a computer-system, cause the computer system to perform the method of any one of claims 1 to 7 and / or the method of any one of claims 8 or 9.

12. A computer-readable data storage medium having recorded thereon the computer program of claim 11 and / or the Deep Neural Network of claim 10.

13. A computer-system comprising a processor coupled to a memory, the memory having recorded thereon the computer program of claim 11 and / or the Deep Neural Network of claim 10.

14. The computer system of claim 13, wherein the computer system is a vehicle on-board computer system coupled with microphones of the vehicle.

15. A vehicle equipped with the computer system of claim 14.

Citation Information

Patent Citations

  • Spatial audio wind noise detection

    US20220199100A1

  • Wind noise suppresor

    US20230063839A1

  • Emergency response vehicle detection for autonomous driving applications

    US20230410650A1

  • Switchable Noise Reduction Profiles

    US20240112690A1