Signal processing device, signal processing method, and signal processing program
A deep learning-based approach estimates virtual microphone signals without relying on physical models, effectively increasing the number of observation microphones and improving array signal processing performance.
Patent Information
- Application Number
- JP2022577952
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-01-29
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2041-01-29
AI Technical Summary
Existing array signal processing techniques face limitations in improving performance when the number of microphones is small, as previous methods rely on physical models that do not always hold true, making it difficult to estimate signals from virtual microphones accurately.
A signal processing device utilizing a deep learning model with a neural network to estimate signals from virtually placed microphones without explicit assumptions, enabling the virtual increase of observation microphones by directly processing acoustic signals.
The method allows for accurate estimation of virtual microphone signals, enhancing microphone array performance even with a limited number of microphones, and improves speech enhancement and signal processing outcomes.
Smart Images

Figure 0007789019000008 
Figure 0007789019000009 
Figure 0007789019000010
Abstract
Description
[Technical Field]
[0001] The present invention relates to a signal processing device, a signal processing method, a signal processing program, a learning device, a learning method, and a learning program. [Background technology]
[0002] Array signal processing techniques using microphone arrays (multiple microphones) are widely used in various applications such as speech enhancement, sound source separation, and sound source direction estimation.
[0003] The performance of array signal processing basically depends on the number of microphones, but in actual operation, many devices have limitations that make it difficult to increase the number of microphones. For this reason, there is a need to improve the performance of microphone array technology when the number of microphones is small.
[0004] In response to this, research has been conducted on methods that enable the virtual number of observation microphones to be increased by estimating the signals from virtual microphones placed in locations where no actual microphones are installed. For example, there is a method for estimating the phase components of virtual microphone signals based on a physical model. The physical model is a model that assumes plane waves, sparseness of speech, a microphone array with sufficiently close spacing, etc. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Hiroki Katahira, “Nonlinear speech enhancement by virtual increase of channels and maximum SNR beamformer”, [online], [Retrieved January 25, 2021], Internet<URL:https: / / asp-eurasipjournals.springeropen.com / track / pdf / 10.1186 / s13634-015-0301-3.pdf> Summary of the Invention [Problem to be solved by the invention]
[0006] Previous research has estimated the signal from a virtual microphone based on a physical model, but this physical model does not always hold true, making it difficult to estimate the signal (especially the phase) from the virtual microphone.
[0007] The present invention has been made in consideration of the above, and aims to provide a signal processing device, a signal processing method, a signal processing program, a learning device, a learning method, and a learning program that can estimate signals from virtually placed microphones without making explicit assumptions about the signals. [Means for solving the problem]
[0008] In order to solve the above-mentioned problems and achieve the object, the signal processing device of the present invention is a signal processing device that processes an acoustic signal, and is characterized by having an estimation unit that estimates an observation signal of a virtually placed virtual microphone from an observation signal of an input real microphone using a deep learning model having a neural network.
[0009] The learning device according to the present invention is characterized by having an input unit that receives, as learning data, an observation signal of a real microphone and an observation signal that is actually observed at the position of a virtually placed virtual microphone that is the target of estimation; an estimation unit that estimates the observation signal of the virtual microphone from the input observation signal of the real microphone using a deep learning model having a neural network; and an update unit that updates the parameters of the neural network so that the observation signal of the virtual microphone estimated by the estimation unit approaches the observation signal that was actually observed at the position of the virtual microphone. [Effects of the Invention]
[0010] According to the present invention, signals from virtually placed microphones can be estimated without making any explicit assumptions about the signals. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram schematically illustrating an example of an estimation device according to the first embodiment. [Figure 2] FIG. 2 is a flowchart illustrating the processing procedure of the estimation process according to the first embodiment. [Figure 3] FIG. 3 is a diagram schematically illustrating an example of a learning device according to the second embodiment. [Figure 4] FIG. 4 is a flowchart showing the processing procedure of the learning process according to the second embodiment. [Figure 5] FIG. 5 is a diagram schematically illustrating an example of a signal processing device according to the third embodiment. [Figure 6] Figure 6 shows the microphone array layout for the CHiME-4 corpus. [Figure 7] FIG. 7 is a diagram illustrating an example of a computer that implements an estimation device, a learning device, and a signal processing device by executing a program. DETAILED DESCRIPTION OF THE INVENTION
[0012] An embodiment of the present invention will be described in detail below with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are denoted by the same reference numerals. Note that, hereinafter, when "^A" is written for A, which is a vector, matrix, or scalar, it is considered to be equivalent to "a symbol with "^" written immediately above "A."
[0013] [Embodiment 1] In the first embodiment, an estimation device that estimates signals from virtually arranged virtual microphones for array signal processing using a microphone array will be described.
[0014] The estimation device according to the first embodiment estimates signals from virtually placed microphones (virtual microphones) without making any explicit assumptions about the signals. Fig. 1 is a diagram schematically illustrating an example of the estimation device according to the first embodiment.
[0015] The estimation device 10 (estimation unit) is realized by loading a predetermined program into a computer or the like including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The estimation device 10 also has a communication interface for transmitting and receiving various information to and from other devices connected via a wired connection or a network, etc.
[0016] As shown in Fig. 1, an estimation device 10 according to the first embodiment includes an NN 11. For the sake of simplicity, Fig. 1 illustrates an example in which two channels corresponding to actual observed real microphones are received and one channel corresponding to a virtual microphone is generated.
[0017] NN11 estimates the observed signal (amplitude and phase components) of a virtually placed virtual microphone from the observed signal observed by the input real microphone. The real microphones are actually installed microphones (microphones 1 and 3 in Figure 1). The observed signal r of the real microphone is the mixed acoustic signal observed by the real microphone (solid circle 1 and 3 in Figure 1). The virtual microphone is a microphone (microphone 2 in Figure 1) virtually placed at a position different from the position of the real microphone. NN11 estimates and outputs the observed signal ^v of the virtual microphone (dashed circle 2 in Figure 1).
[0018] The NN 11 is, for example, a time-domain deep learning model with high phase estimation performance. The NN 11 is a NN that operates directly in the time domain without being based on physical assumptions and can accurately estimate time-domain signals. The estimation device 10 uses the NN 11 to estimate time-domain signals that are observation signals of a virtual microphone from time-domain signals that are input observation signals of a real microphone. Hereinafter, in the first embodiment, a NN-based virtual microphone signal estimation (NN-VME: Neural Network-based Virtual Microphone Estimator) is proposed, which is a method for estimating observation signals of a virtual microphone directly from the time domain. Note that the NN 11 does not necessarily have to be a time-domain model and may be realized by a frequency-domain model. The NN 11 includes an encoder 111, a convolution block 112, and a decoder 113.
[0019] The encoder 111 is a neural network that maps an acoustic signal to a predetermined feature space, i.e., converts the acoustic signal into a feature vector. The convolution block 112 is a set of layers for performing one-dimensional convolution, etc. The decoder 113 is a neural network that maps features in a predetermined feature space to the space of acoustic signals, i.e., converts feature vectors into acoustic signals. The NN 11 outputs the observed signal converted by the decoder 113 as an estimated signal ^v of the virtual microphone.
[0020] The configurations of the convolution block, encoder, and decoder may be the same as those described in Reference 1 (Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,” IEEE / ACM Trans. ASLP, vol. 27, no. 8, pp. 1256-1266, 2019.). The time-domain acoustic signal may be obtained by the method described in Reference 1. In the following description, each feature is represented by a vector.
[0021] [Estimation process] Next, we will explain the case where the NN 11 simultaneously estimates one or more virtual microphones. c is the T long-term domain waveform of the cth real microphone, and ^v c´ denotes the estimated signal of the c'th virtual microphone. The real microphone signal r={r c=1 ,…,r c=Cr}, the NN-VME module NN11 generates a virtual microphone signal ^v={^v c´=1 ,…,^v c´=Cv} is estimated as shown in equation (1).
[0022]
number
[0023] where C r denotes the number of observation channels (i.e., real microphones), and C v denotes the number of virtual estimation channels (i.e., virtual microphones), and NN-VME(·) is the neural network.
[0024] [Estimation processing procedure] 2 is a flowchart showing the processing procedure of the estimation process according to embodiment 1. When an observed signal r of a real microphone is input, the estimation device 10 converts the input observed signal r of the real microphone in the time domain into a feature (step S1). The convolution block 112 performs one-dimensional convolution (step S2).
[0025] The decoder 113 converts the feature amount into an observation signal at the position of the virtual microphone (step S3). The NN 11 outputs the observation signal converted by the decoder 113 as an estimated signal ^v of the virtual microphone (step S4).
[0026] [Effects of the First Embodiment] In this way, the estimation device 10 uses a time-domain deep learning model with high phase estimation performance to directly estimate the observed signal of the virtual microphone from the observed signal observed by the input real microphone. In the tenth embodiment, such a data doblind framework makes it possible to directly estimate the signal (amplitude and phase components) of the virtual microphone without making any explicit assumptions about the signal (e.g., a physical model). Then, the estimation device 10 uses a time-domain deep learning model with high phase estimation performance to realize estimation of both the amplitude and phase of the virtual microphone signal.
[0027] Therefore, according to the first embodiment, it is possible to virtually increase the number of observation microphones, and even when the number of microphones is small, it is possible to improve the performance of the microphone array technology.
[0028] [Embodiment 2] Next, a second embodiment will be described. In the second embodiment, a learning device that performs learning of the NN 11 in the estimation device 10 will be described. In order to have the NN 11, which is an NN-VNE module, estimate the signal of the virtual microphone, the learning device 20 employs supervised learning, and uses, as learning data, observed signals of real microphones at the positions of the virtual microphones in addition to observed signals of real microphones that are actually placed during operation.
[0029] Fig. 3 is a diagram schematically illustrating an example of a learning device according to embodiment 2. Note that the same components as those in embodiment 1 are denoted by the same reference numerals and will not be described again. For simplicity of explanation, Fig. 3 illustrates an example in which learning device 20 receives two channels corresponding to real microphones and performs learning on NN 11 that generates one channel corresponding to a virtual microphone.
[0030] 3 is realized by, for example, loading a predetermined program into a computer or the like including a ROM, RAM, CPU, etc., and having the CPU execute the predetermined program. The learning device 20 also has a communication interface for transmitting and receiving various information to and from other devices connected via a wired connection or a network, etc. The learning device 20 has an NN 11, an input unit 21, and a parameter update unit 22.
[0031] The input unit 21 receives, as learning data, an observation signal (indicated by solid circle 1 and 3 in FIG. 3) of a real microphone (microphone 1 and 3) installed during operation and an observation signal (indicated by solid circle 2 in FIG. 3) actually observed at the position of a virtually placed virtual microphone (microphone 2) that is the estimation target. The input unit 21 inputs, to the NN, a time-domain observation signal r (indicated by solid circle 1 and 3 in FIG. 3) of a real microphone installed during operation. The input unit 21 inputs, to the parameter update unit 22, an observation signal t (indicated by solid circle 2 in FIG. 1) actually observed at the position of the virtual microphone.
[0032] NN11 (estimation unit) estimates the observed signal ^v (in Figure 3, the dashed circle 2) of a virtually placed virtual microphone (microphone 2) from the observed signal r observed by the input real microphones (microphones 1 and 3).
[0033] The parameter update unit 22 updates the parameters of the NN 11 so that the observed signal ^v of the virtual microphone estimated by the NN 11 approaches the observed signal t actually observed at the position of the virtual microphone.
[0034] [Learning process] Next, the learning process will be described. The learning device 20 employs supervised learning to have the NN 11, which is an NN-VME module, estimate the virtual microphone signal. Therefore, during learning, the observed signals of the real microphones at the positions of the virtual microphones are used as the learning target, along with the observed signals of the real microphones.
[0035] So, assume that a set of input and target signals {r, t} is available, where t={t c´=1 ,...,t c´=Cv} and t c´ denotes the target signal for the c'th virtual microphone. Figure 3 shows the case where a subset of microphones (e.g., channels 1 and 3) are assigned as network input values r, and another subset (e.g., channel 2) is used as network target values t.
[0036] The neural network (NN) 11 is trained based on the time-domain loss between the estimated signal and the real signal at the virtual microphone position. The parameter updater 22 employs a scale-dependent signal-to-noise ratio (SNR) as the loss, for example, as shown in Equation (2).
[0037]
number
[0038] Here, as explained in equation (1), ^v=NN-VME(r).
[0039] [Learning process procedure] Next, a description will be given of the learning process according to Embodiment 2. Fig. 4 is a flowchart showing the processing procedure of the learning process according to Embodiment 2.
[0040] 4, as learning data, an observed signal of a real microphone installed during operation and an observed signal actually observed at the position of a virtually placed virtual microphone, which is the estimation target, are input (step S11). An input unit 21 inputs a time-domain observed signal r of the real microphone installed during operation to the NN 11 (step S12).
[0041] The NN11 performs the same processing as steps S1 to S4 shown in FIG. 2 to estimate the observed signal ^v of the virtually placed virtual microphone from the observed signal r observed by the input real microphone (steps S13 to S16).
[0042] The parameter update unit 22 updates the parameters of the NN 11 so that the observed signal ^v of the virtual microphone estimated by the NN 11 approaches the observed signal t actually observed at the position of the virtual microphone (step S17). The parameter update unit 22 updates the parameters of the NN 11 so that the loss calculated by equation (2) is optimized.
[0043] Then, the parameter update unit 22 determines whether or not a termination condition has been reached (step S18). If the termination condition has been reached (step S18: Yes), the learning device 20 ends the process. If the termination condition has not been reached (step S18: No), the process returns to step S12. The termination condition may be, for example, that the number of parameter updates for the NN 11 has reached a predetermined number, that the loss value used for parameter update has become equal to or less than a predetermined threshold, or that the amount of parameter update (such as the derivative of the loss function value) has become equal to or less than a predetermined threshold.
[0044] [Effects of the second embodiment] Thus, unlike the training of speech enhancement methods, the learning device 20 according to the second embodiment does not require paired noisy and clean signals, but only requires observed signals from multiple real microphones as training data. In other words, the learning device 20 requires only observed signals (mixed acoustic signals) containing multi-channel noise as training data, so there are no restrictions on the form of the device, and mixed acoustic signals from multiple channels can be used as training data. In other words, the learning device 20 can use actual recordings taken with multiple microphones as training data, rather than simulated recordings.
[0045] Therefore, it is easy and inexpensive to prepare training data for the learning device 20. Furthermore, by using a large amount of training data, the learning device 20 can build a powerful NN 11, which enables precise modeling of real recordings.
[0046] [Embodiment 3] The estimation device 10 enables generation of virtual microphone signals, which can be used for various array processing. Therefore, in the third embodiment, a configuration in which the estimation device 10 is combined with a frequency domain beamformer will be described as an example.
[0047] [Signal processing device] Fig. 5 is a diagram schematically illustrating an example of a signal processing device according to embodiment 3. The signal processing device 100 illustrated in Fig. 5 is realized, for example, by loading a predetermined program into a computer or the like including a ROM, a RAM, a CPU, etc., and causing the CPU to execute the predetermined program. The signal processing device 100 also has a communication interface for transmitting and receiving various information to and from other devices connected via a wired connection or a network, etc. The signal processing device 100 includes an estimation device 10, a microphone signal processing unit 30, and an application unit 40 (signal processing unit).
[0048] The microphone signal processing unit 30 generates a speech enhancement signal from which noise components have been removed, based on the observed signal of the real microphone and the observed signal of the virtual microphone estimated by the estimation device 10. Note that the microphone signal processing unit 30 may also include sound source separation processing, sound source localization processing, etc.
[0049] The application unit 40 performs another task-dependent process using the speech enhancement signal. The application unit 40 performs, for example, speech recognition processing. Note that the processing order of the signal processing device 100 is an example, and there are cases where speech recognition processing is performed after sound source separation processing, and cases where speech enhancement processing or sound source separation processing is performed after sound source localization processing.
[0050] [Speech enhancement processing] [Basic Procedure] First, the estimation device 10 is used to estimate the real microphone signal r∈R as explained in equation (1). T×Cr As a virtual microphone signal ^v∈R T×Cv and estimate the extended microphone signal y=[r,^v]∈R T×C (C=C r +C v ) is obtained. Next, the microphone signal processor 30 uses a frequency domain beamformer in addition to the enhanced microphone signal in a frequency domain representation (i.e., a Short-Time Fourier Transform (STFT)) to obtain the enhanced speech signal. Finally, an inverse STFT is used to recover the enhanced time-domain waveform.
[0051] STFT area^X t,f The enhanced speech signal at ∈C is ^X t,f =w H f Y t,f where Y t,f ∈C C is a vector containing the C-channel STFT coefficients of the extended microphone signal at time-frequency bin (t,f), and w f ∈C C is a vector containing the beamforming filter coefficients, H denotes the conjugate transpose.
[0052] [MVDR format] The microphone signal processing unit 30 uses, for example, Minimum Variance Distortionless Response (MVDR) beamforming (Reference 2: Mehrez Souden, Jacob Benesty, and Sofiene Affes, “On optimal frequency-domain multichannel linear filtering for noise reduction,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 18, no. 2, pp. 260-276, 2009.) to generate time-invariant filter coefficients w fis calculated as shown in equation (3).
[0053]
number
[0054] where Φ S f ∈C C×C and Φ N f ∈C C×C are the spatial covariance (SC) matrices of the speech signal and noise signal, respectively. C is a one-hot vector representing the reference microphone.
[0055] Then, using the time-frequency mask, the SC matrix is estimated as shown in equation (4) (Reference 3: Jahn Heymann, Lukas Drude, and Reinhold Haeb-Umbach, “Neural network based spectral mask estimation for acoustic beamforming”, in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 196-200.).
[0056]
number
[0057] where ν∈{S,N}. m S t,f ∈[0,1] and m N t,f ∈[0,1] are the time-frequency masks of speech and noise, respectively.
[0058] [Virtual microphone loading] Experiments described below reveal that the use of virtual microphones in beamforming, while effective in increasing signal-to-distortion ratios (SDR), does not necessarily improve automatic speech recognition (ASR) performance because the virtual microphone estimation introduces processing artifacts.
[0059] To reduce the effect of this artifact, we use the virtual microphone loading term Z∈R in equation (5). C , the SC matrix Φ N f That is, in the microphone signal processing unit 30, ,of A loading term that reduces the weight of the virtual microphone channel is added to the spatial covariance matrix of the noise signal.
[0060]
number
[0061] where Z={z c,c´} C,C c=1,c´=1 is a matrix with zeros except for the diagonal elements corresponding to the virtual microphones. cv,cv = 1, and c v denotes the channel index corresponding to the virtual microphone, and ε is a loading hyperparameter that controls the contribution of the virtual microphone to the beamformer. For example, setting a large value for ε means that the virtual microphone contains significant noise that is uncorrelated with other microphones. Therefore, the estimated beamformer can improve ASR performance by reducing the channel weight of the virtual microphone.
[0062] [Effects of the Third Embodiment] The virtual microphone signals estimated by the estimation device 10 with the NN-VME module also allow for improved performance of speech enhancement and signal processing enhanced by the NN-VME.
[0063] [experiment] To evaluate NN-VME, we conducted the following two evaluations: Experiment 1, which evaluated the virtual microphone estimation performance using NN-VME, and Experiment 2, which evaluated the enhancement performance of a beamformer using estimated virtual microphones. Note that while the experiments reported the results of estimating one virtual microphone, it can naturally be expanded to estimate multiple virtual microphones.
[0064] Figure 6 shows the microphone array layout for the CHiME-4 corpus. All microphones in Figure 6 face forward except for microphone 2.
[0065] [Experimental conditions] We evaluated NN-VME on the CHiME-4 corpus (Reference 4: Jon Barker, Ricard Marxer, Emmanuel Vincent, and Shinji Watanabe, “The third CHiME speech separation and recognition challenge: Dataset, task, and baselines,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2015, pp. 504–511.) The CHiME-4 corpus contains speech recorded using a tablet device equipped with a six-channel rectangular microphone array, as shown in Figure 6. This corpus includes both simulated data and real-world recordings from noisy public environments.
[0066] The training set consisted of 3 hours of real speech data from 4 speakers and 15 hours of simulated speech data from 83 speakers. The evaluation set included 1,320 utterances from 4 speakers, including real speech data and simulated speech data containing noise. Of these utterances, utterances due to microphone malfunctions were removed, leaving 1,149 utterances for the evaluation set.
[0067] The evaluation metrics used were the SDR and word error rate (WER) of BSSEval (Reference 5: Emmanuel Vincent, Remi Gribonval, and Cedric Fevotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 14, no. 4, pp. 1462-1469, 2006.). To evaluate the performance of the virtual microphone estimation, we calculated the SDR between the estimated virtual microphone signal in the channel corresponding to the virtual microphone and the observed real microphone signal.
[0068] To evaluate the enhancement performance of the beamformer, we used a clean reverberant signal in the fourth channel as a reference signal. This evaluation is performed only on simulated data, as it requires access to a clean signal.
[0069] When evaluating ASR performance, we used Kaldi's CHiME-4 recipe (Reference 6: Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, “The Kaldi speech recognition toolkit,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2011. Reference 7: [online], [Retrieved January 25, 2021], Internet).<https: / / github.com / kaldi-asr / kaldi / tree / master / egs / chime4 / s5_6ch> ) was used.This is based on a deep neural network-hidden Markov model hybrid acoustic model (Reference 9: Herve Bourlard and Nelson Morgan, Connectionist speech recognition: A hybrid approach, 1994; Reference 10: Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdelrahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, and Brian Kingsbury, 2016) trained with the lattice-free maximum mutual information criterion (Reference 8: Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI”, in Interspeech, 2016, pp. 2751-2755.). (See "Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups," IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 8297, 2012.) A trigram language model was used for decoding.
[0070] [Experimental configuration] The network configuration of NN-VME was based on the Conv-TasNet network architecture. Following the description in Reference 1, the hyperparameters were set as N = 256, L = 20, B = 256, H = 512, P = 3, X = 8, and R = 4.
[0071] We trained the NN-VME using the Adam algorithm with gradient clipping (Diederik P Kingma and Jimmy Ba, “Adam: A method for stochastic optimization”, in International Conference on Learning Representations (ICLR), 2015.) with an initial learning rate of 0.0001. Training was terminated after 200 epochs.
[0072] The MVDR beamformer is available from the GitHub repository used in Kaldi's CHiME-4 recipe (Reference 12: [online], [searched January 25, 2021], Internet).<URL:https: / / github.com / fgnt / nn-gev,> We used the pre-trained mask estimation model provided by [Reference 3]. For the STFT calculation, we used a Blackman window with a set length and shift of 64 ms and 16 ms, respectively. For the ASR experiments, the loading hyperparameter ε in Eq. (5) was set to 0.05.
[0073] [Experimental Results] [Evaluation of virtual microphone estimation performance] Table 1 shows the SDR [dB] of the virtual microphone estimation using the noisy observed signal as the reference signal.
[0074] [Table 1]
[0075] In Table 1, RM represents the real microphone, and VM represents the virtual microphone estimated by NN-VME (NN11). Here, the reference signal for calculating the SDR is not a clean signal, but an observed signal containing noise from the channel corresponding to the virtual microphone. Therefore, the virtual microphone estimation performance can also be evaluated for real recordings.
[0076] In Table 1, the first column "eval ch" indicates the channel index of the virtual or real microphone signal used as the estimated signal in calculating the SDR. The second column "ref ch" indicates the channel index of the real microphone signal used as the reference signal. Here, the notation "5(4,6)" indicates that the virtual microphone signal in channel 5 was estimated using the real microphone signals in channels 4 and 6. As a reference, the score is compared with the SDR obtained with the closest real microphone (i.e., the one with the highest SDR). These results are shown in the first row (eval ch4, ref ch5) and the fourth row (eval ch5, ref ch6) of Table 1.
[0077] Table 1 shows that the signal estimated by the NN-VME module (e.g., "5(4,6)") has a higher SDR score than the observed signal recorded by a nearby microphone (e.g., "4"). These results demonstrate that even in real-world recordings, the NN-VME (NN11) can estimate virtual microphone signals that are not actually observed by the microphones by utilizing spatial information inferred from the few observed real microphone signals.
[0078] Table 1 shows the results of interpolation (i.e., virtual microphones positioned between real microphones) (e.g., "5(4,6)") and horizontal extrapolation (e.g., "6(4,5)"). In both cases, the NN-VME (NN11) can predict virtual microphone signals with low time waveform distortion, with an SDR of approximately 12 dB or higher.
[0079] [Evaluation of beamformer enhancement performance] Table 2 shows the SDR [dB] of the beamformer using the clean signal as the reference signal. Note that a higher SDR value indicates better performance, and a lower WER [%] value indicates better performance.
[0080] [Table 2]
[0081] In Table 2, VM BF indicates a beamformer using estimated virtual microphones (output of NN11), and RM BF indicates a beamformer using only real microphones. In Table 2, the columns "real" and "virtual" in the "used ch" column indicate the channel indexes corresponding to the real and virtual microphones used to form the beamformer, respectively. For example, "VM BF" in row (4) is formed using two real microphone signals (i.e., channels 4 and 6) and one virtual microphone signal (i.e., channel 5).
[0082] Table 2 shows that the VM BF proposed in the first embodiment (e.g., row (4)) achieves a higher SDR score than the RM BF formed by the same real microphone signal (e.g., row (2)). Here, another RM BF (e.g., row (3)) corresponds to the upper limit performance of the VM BF.
[0083] To evaluate the performance of the beamformer on real-world recordings, we performed an ASR evaluation in addition to the SDR-based evaluation described above. Table 2 also shows the WERs of the RM BF and VM BF evaluated on real data.
[0084] In actual recordings, the table also confirms that the VM BF proposed in the first embodiment (e.g., row (4)) reduced the WER by 0.9% compared to the corresponding RM BF (e.g., row (2)). A similar trend was observed when more microphones were used (rows (5) to (7)).
[0085] These results demonstrate that the estimated virtual microphone signals improve enhancement performance when combined with a beamformer.
[0086] Furthermore, Table 2 shows the results of VM BF with virtual microphone loading. The WER score of VM BF without loading is 15.1% under the same conditions as row (4), and 13.4% under the same conditions as row (7). This indicates that virtual microphone loading is effective in improving the ASR performance of VM BF.
[0087] Thus, it was shown that the virtual microphone signal estimated by NN-VME (NN11) improves the performance of speech enhancement and signal processing enhanced by NN-VME.
[0088] [System configuration of the embodiment] The components of the estimation device 10, the learning device 20, and the signal processing device 100 are conceptual functional components and do not necessarily need to be physically configured as shown in the figure. In other words, the specific forms of distribution and integration of the functions of the estimation device 10, the learning device 20, and the signal processing device 100 are not limited to those shown in the figures, and all or part of them can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.
[0089] Furthermore, all or any part of the processes performed in the estimation device 10, the learning device 20, and the signal processing device 100 may be realized by a CPU, a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU. Furthermore, each process performed in the estimation device 10, the learning device 20, and the signal processing device 100 may be realized as hardware using wired logic.
[0090] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.
[0091] [program] 7 is a diagram showing an example of a computer in which the estimation device 10, the learning device 20, and the signal processing device 100 are realized by executing a program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0092] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0093] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the programs that define the processes of the estimation device 10, the learning device 20, and the signal processing device 100 are implemented as program modules 1093 in which code executable by the computer 1000 is written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, the program modules 1093 for executing processes similar to those of the functional configurations of the estimation device 10, the learning device 20, and the signal processing device 100 are stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced with an SSD (Solid State Drive).
[0094] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.
[0095] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0096] Although the present invention has been described above as an embodiment, the present invention is not limited to the descriptions and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention. [Explanation of symbols]
[0097] 10 Estimation device 11 Neural Networks (NN) 111 Encoder 112 Convolution Blocks 113 Decoder 20 Learning Device 21 Input section 22 Parameter update section 30 Microphone signal processing section 40 Application Section 100 signal processing section
Claims
1. A signal processing device for processing an acoustic signal, an estimation unit that estimates a time domain signal corresponding to the amplitude and phase components of an observation signal of a virtually placed virtual microphone from a time domain signal corresponding to the amplitude and phase components of an observation signal of an input real microphone using a deep learning model having a neural network; a microphone signal processing unit that generates a speech enhancement signal from which a noise signal has been removed, based on the time domain signal of the real microphone and the time domain signal of the virtual microphone estimated by the estimation unit; an application unit that performs signal processing using the speech enhancement signal; and The signal processing device, wherein the microphone signal processing unit adds a loading term that reduces the weight of the virtual microphone channel to a spatial covariance matrix of a noise signal.
2. A signal processing method executed by a signal processing device, comprising: a step of estimating, using a deep learning model having a neural network, time-domain signals corresponding to amplitude and phase components of observation signals of a virtually placed virtual microphone from time-domain signals corresponding to amplitude and phase components of observation signals of an input real microphone; generating a speech-enhanced signal from which a noise signal has been removed, based on the time-domain signal of the real microphone and the time-domain signal of the virtual microphone estimated in the estimating step; performing signal processing using the speech enhancement signal; Including, A signal processing method, wherein the generating step includes adding a loading term that reduces the weight of the virtual microphone channel to a spatial covariance matrix of a noise signal.
3. A signal processing program for causing a computer to function as the signal processing device according to claim 1.
Citation Information
Patent Citations
Automobile interior noise control method
CN108806664A