Signal processing device, signal processing program, and signal processing method
The signal processing device uses a generative adversarial network to correct speech enhancement distortions unsupervisedly, addressing practical limitations of conventional methods by enhancing audio quality without paired data.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2026-03-17
AI Technical Summary
Conventional speech enhancement techniques suffer from artificial processing distortion and residual noise, particularly in time-frequency spectrograms, and require paired data of distorted and undistorted signals for supervised learning, making them impractical for many use cases.
A signal processing device utilizing a deep neural network within a generative adversarial network framework for unsupervised distortion correction, trained on samples of distorted and undistorted audio signals, to reduce distortion without requiring paired data.
The method effectively reduces distortion and residual noise in speech enhancement signals, improving audio quality without introducing additional distortions, and operates efficiently in real-world environments without relying on paired data.
Smart Images

Figure 0007831761000003 
Figure 0007831761000004 
Figure 0007831761000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to a signal processing device, a signal processing program, and a signal processing method, and can be applied, for example, to a process that reduces distortion from an audio signal that has been distorted by any signal processing. [Background technology]
[0002] Currently, speech enhancement technology, which emphasizes the target sound component from an observation signal mixed with interfering noise, has become an indispensable pre-processing technique in various audio processing applications. Ideally, the enhanced audio obtained here should not only have the interfering noise source removed, but also be free from unpleasant processing distortion.
[0003] Conventional audio enhancement techniques can be broadly classified into linear processing-based approaches and nonlinear processing-based approaches. Audio obtained through nonlinear audio enhancement processing such as time-frequency masking (see Non-Patent Literature 1) and DAE (Denoising Auto Encoder) (see Non-Patent Literature 2) contains artificial and unpleasant distortions that result mainly from the loss of spectral components of the target sound source, in addition to residual noise from unwanted sounds.
[0004] In response to this, conventional methods have been proposed to suppress nonlinear distortion by performing time smoothing in the cepstrum region (see Non-Patent Document 2).
[0005] Furthermore, attempts have been made to suppress interfering sound components while reducing processing distortion of the target sound source by integrating time-frequency masking and adversarial DAE (see Non-Patent Literature 3). In this adversarial learning-based method, by learning the mapping of the observed signal to the corresponding ground truth signal, it becomes possible to restore spectral components lost in time-frequency masking, thereby achieving speech enhancement for signals with severe processing distortion. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Ozgur Yilmaz, Scott Rickard, “Blind Separation of Speech Mixtures via Time-Frequency Masking”, IEEE Trans. on signal Proc, 1830-1847, 2004, [Retrieved February 11, 2022], [Online] INTERNET,<URL: https: / / www.ee.columbia.edu / ~dpwe / papers / YilR02-bsstfm.pdf > [Non-Patent Document 2] Xugang Lu, Yu Tsao, Shigeki Matsuda, Chiori Hori1, “Speech Enhancement Based on Deep Denoising Autoencoder”, INTERSPEECH, 436-440, 2013, [Retrieved February 11, 2022], [Online] INTERNET,<URL:https: / / www.citi.sinica.edu.tw / papers / yu.tsao / 3582-F.pdf> [Non-Patent Document 3] Naohiro Tawara, Tetsunori Kobayashi, Masaru Fujieda, Kazuhiro Katagiri, Takashi Yazu, Tetsuji Ogawa, “Adversarial autoencoder for reducing nonlinear distortion”, IPSJ, 2018, [Retrieved February 11, 2022], [Online] INTERNET,<URL: http: / / www.apsipa.org / proceedings / 2018 / pdfs / 0001669.pdf> [Overview of the project] [Problems that the invention aims to solve]
[0007] Incidentally, artificial processing distortion and residual noise occur locally on the time-frequency spectrogram. Therefore, conventional techniques designed to globally analyze the time-frequency spectrogram present different challenges for each.
[0008] Specifically, the technology described in Non-Patent Document 2 had the problem of producing another distortion similar to reverberation. Furthermore, the technology described in Non-Patent Document 3, being a supervised learning method, required the preparation of paired data of observed signals and ground truth signals in order to learn the mapping between the observed signal and the corresponding ground truth signal. As described above, conventional technologies have the problem of not being practical because it is not realistic to prepare paired data of observed signals and ground truth signals for all possible use cases.
[0009] In light of the above problems, there is a need for a signal processing device, a signal processing program, and a signal processing method that can reduce signal distortion caused by speech enhancement technology without requiring pair data of a distorted speech signal and a corresponding undistorted ground truth signal. [Means for solving the problem]
[0010] The first signal processing device of the present invention includes distortion correction means for correcting distortion of an input signal subjected to nonlinear audio enhancement processing using a deep neural network, wherein the deep neural network is Based on training data that includes samples of distorted audio signals and samples of undistorted audio signals, It has a learning model trained within the framework of a generative adversarial network using generators and discriminators, The generator uses the deep neural network, and the objective function of the deep neural network as the generator includes the input-output loss between the input signal and the output signal of the deep neural network. It is characterized by the following:
[0011] The second signal processing program of the present invention causes a computer to function as a distortion correction means that corrects the distortion of an input signal subjected to nonlinear audio enhancement processing using a deep neural network, wherein the deep neural network is Based on training data that includes samples of distorted audio signals and samples of undistorted audio signals, It has a learning model trained within the framework of a generative adversarial network using generators and discriminators, The generator uses the deep neural network, and the objective function of the deep neural network as the generator includes the input-output loss between the input signal and the output signal of the deep neural network. It is characterized by the following:
[0012] The third invention is a signal processing method performed by a signal processing apparatus. The signal processing apparatus includes distortion correction means. The distortion correction means corrects the distortion of an input signal subjected to non-linear voice enhancement processing using a deep neural network. The deep neural network has a learning model trained in the framework of an adversarial generation network using a generator and a discriminator, Based on training data that includes samples of distorted audio signals and samples of undistorted audio signals, and is characterized by the above. The generator uses the deep neural network, and the objective function of the deep neural network as the generator includes the input-output loss between the input signal and the output signal of the deep neural network.
Effect of the Invention
[0013] According to the present invention, it is possible to reduce the distortion of a signal generated by a voice enhancement technique without requiring paired data of a voice signal including distortion and a correct signal without distortion corresponding to the voice signal.
Brief Description of the Drawings
[0014] [Figure 1] It is a block diagram showing a functional configuration of a signal processing apparatus according to an embodiment. [Figure 2] It is a block diagram showing a hardware configuration of a signal processing apparatus according to an embodiment. [Figure 3] It is a block diagram showing a configuration when training a distortion correction DNN constituting a signal processing apparatus according to an embodiment. [Figure 4] It is a diagram (image diagram) showing an example of a model applied to a discriminator used when adversarially training a distortion correction DNN according to an embodiment. [Figure 5] It is a diagram (part 1) showing a model applied to a simulation of a sound collection apparatus according to an embodiment. [Figure 6] It is a diagram (part 2) showing a model applied to a simulation of a sound collection apparatus according to an embodiment. [Figure 7] It is a diagram (part 1) showing an evaluation result of a sound collection apparatus according to an embodiment. [[ID=四十]] [Figure 8] It is a diagram (part 2) showing an evaluation result of a sound collection apparatus according to an embodiment. <00001百零二> [Figure 9] This is Figure (3) showing the evaluation results of the sound collection device according to the embodiment. [Figure 10] This is a block diagram (part 1) showing the configuration when a framework of adversarial networks with cyclical consistency is applied during the training of the distortion correction means (distortion correction DNN) according to the embodiment. [Figure 11] This is a block diagram (part 2) showing the configuration when a framework of adversarial networks with cyclical consistency is applied during the training of the distortion correction means (distortion correction DNN) according to the embodiment. [Figure 12] This is a block diagram (part 3) showing the configuration when a framework of adversarial networks with cyclical consistency is applied during the training of the distortion correction means (distortion correction DNN) according to the embodiment. [Modes for carrying out the invention]
[0015] (A) Main embodiment Hereinafter, an embodiment of the signal processing device, signal processing program, and signal processing method according to the present invention will be described in detail with reference to the drawings.
[0016] (A-1) Configuration of the embodiment Figure 1 is a block diagram showing the overall configuration of the signal processing device 10 in this embodiment.
[0017] The signal processing device 10 includes a nonlinear audio enhancement means 11 and a distortion correction means 12.
[0018] The nonlinear speech enhancement means 11 is responsible for processing the "observation signal S1," which is an audio signal (acoustic signal) containing the speech component as the target sound, by nonlinear speech enhancement processing (hereinafter also referred to as "nonlinear speech enhancement processing") and outputting a "speech enhancement signal S2."
[0019] The distortion correction means 12 corrects the distortion contained in the audio enhancement signal S2 (distortion caused by nonlinear audio enhancement processing) to obtain a distortion-corrected audio signal (acoustic signal) called "distortion-corrected signal S3".
[0020] The distortion correction means 12 corrects the distortion using the distortion correction DNN 121. The distortion correction DNN 121 is a DNN (Deep Neural Network) that learns through a learning process described later.
[0021] The configuration and format of the observed signal S1 are not limited. As shown in Figure 1, in this embodiment, the observed signal S1 is assumed to be an audio signal (acoustic signal) observed (captured) by a microphone array unit 20 equipped with two microphone arrays MA (MA1, MA2). Microphone arrays MA1 and MA2 are assumed to be 2-channel microphone arrays each equipped with two microphones M (M1, M2). Each microphone array MA1 and MA2 is assumed to be placed at any location in the space where the target area where the target sound source (speaker) is located exists. The number and type of microphones / microphone arrays equipped in the microphone array unit 20 are not limited, and various configurations corresponding to the nonlinear speech enhancement means 11 can be applied.
[0022] Furthermore, the nonlinear sound enhancement technique used by the nonlinear sound enhancement means 11 is not limited, but in this embodiment, MUBASE (Multiple beam-forming area sound enhancement) (see Reference 1) will be used for explanation. In MUBASE processing, the common part of the fan-shaped spatial filter configured in the front direction of each microphone array MA is emphasized, thereby emphasizing only sound sources contained in a specific region (sound originating from the target area). In other words, MUBASE obtains a signal in which the speech, which is the sound of the target area, is emphasized by a process (hereinafter also called "area sound acquisition processing") that captures the sound of the target area, which originates from the target area (speakers within the target area), based on the beamformer output of multiple microphone array MAs. In this embodiment, an example of applying the above-described MUBASE as the nonlinear sound enhancement means 11 will be explained, but other nonlinear sound enhancement techniques may also be applied.
[0023] [Reference 1] Kazuhiro Katagiri, Tokuo Yamaguchi, Takashi Yazu, and Yoong Keok Lee, “Multiple beam-forming area sound enhancement (MUBASE) and stereophonic area sound reproduction (SASR) system”, SIGGRAPH Asia 2015 Emerging Technologies, 2015, [searched on February 11, 2022], [Online] INTERNET, <URL: https: / / dl.acm.org / doi / 10.1145 / 2818466.<2818493> Next, an example of the hardware configuration of the signal processing device 10 will be described.
[0024] The signal processing device 10 may be entirely configured by hardware (e.g., a dedicated chip, etc.), or may be configured partly or entirely as software (a program). The signal processing device 10 may be configured, for example, by installing a program (including the signal processing program of the embodiment) on a computer having a processor and a memory.
[0025] FIG. 2 is a block diagram showing an example of the hardware configuration of the signal processing device 10.
[0026] FIG. 2 shows an example of the hardware configuration when the signal processing device 10 is configured using software (a computer).
[0027] The signal processing device 10 shown in FIG. 2 has a computer 400 in which a program (including the sound collection program of the embodiment) is installed as a hardware component. Also, the computer 400 may be a computer dedicated to the sound collection program, or may be configured to be shared with programs of other functions. [[ID=The computer 400 shown in Figure 2 includes a processor 401, a primary storage unit 402, and a secondary storage unit 403. The primary storage unit 402 is a storage means that functions as the working memory of the processor 401, and can be a high-speed memory such as DRAM (Dynamic Random Access Memory). The secondary storage unit 403 is a storage means that records various data such as the OS (Operating System) and program data (including the sound acquisition program data according to the embodiment), and can be a non-volatile memory such as FLASH (trademark registered) memory, HDD, or SSD. In the computer 400 of this embodiment, when the processor 401 starts up, it reads the OS and program (including the sound acquisition program according to the embodiment) recorded in the secondary storage unit 403, and loads them onto the primary storage unit 402 for execution. Note that the specific configuration of the computer 400 is not limited to the configuration in Figure 2, and various configurations can be applied. For example, if the primary storage unit 402 is a non-volatile memory (e.g., FLASH memory), the secondary storage unit 403 may be omitted.
[0029] Next, we will describe the detailed configuration of the DNN121 distortion correction unit.
[0030] Figure 3 is a block diagram showing the configuration of the distortion correction DNN121 during training.
[0031] As shown in Figure 3, the distortion-corrected DNN121 can learn unsupervisedly using the framework of a GAN (Generative Adversarial Network) that performs adversarial learning.
[0032] In this case, within the GAN framework, the distortion-corrected DNN121 corresponds to the Generator. Also, in Figure 3, the discriminator 122 is positioned as an element of the Discriminator within the GAN framework.
[0033] In Figure 3, to distinguish it from the input / output signals (S2, S3) of the trained distortion correction DNN 121, the input signal of the distortion correction DNN 121 during training is shown as S4 and the output signal as S5. Also in Figure 3, the distortion-free audio signal input to the classifier 122 is shown as "S6," and the classification loss (identification loss) acquired by the classifier 122 is shown as "S7."
[0034] In this case, the discriminator 122 performs the process of distinguishing between the distortion-free audio signal S6 (true) and the output signal S5 (false) of the generator (distortion-correcting DNN 121). The distortion-correcting DNN 121 then learns to deceive the discriminator 122, which distinguishes between the distortion-free audio signal S6 (true) and the output signal S5 (false) of the generator (distortion-correcting DNN 121).
[0035] In this case, the input signal S4 may be a signal that has undergone audio enhancement processing and includes distortion. Specifically, for example, the signal output by the nonlinear audio enhancement means 11 may be used as the input signal S4. In this case, the distortion-free audio signal S6 does not need to be the correct signal (the distortion-free target sound signal included in the input signal S4) corresponding to the input signal S4 (the signal including distortion). In other words, the distortion-free audio signal S6 does not need to be the pair data (correct signal) corresponding to the input signal S4 (the signal including distortion) itself.
[0036] In the framework shown in Figure 3, an input / output loss calculation means 123 is arranged to calculate the input / output loss S8, which is the loss between the input signal S4 and the output signal S5 of the distortion correction DNN 121. Details of the input / output loss S8 will be described later.
[0037] As described above, in this embodiment of the distortion correction DNN121, adversarial learning is performed within the framework of a GAN as shown in Figure 3, eliminating the need for pairs of correct signals (pair data) corresponding to the input signal S4. This is an important requirement for constructing and operating the distortion correction means 12 using only signals obtained in a real environment.
[0038] Next, the parameters used in the GAN framework when training the distortion correction DNN 121 shown in FIG. 3 will be described.
[0039] Here, the parameter θ of the discriminator 122 D and the parameter θ of the distortion correction DNN 121 (generator) G are obtained by minimizing the objective function L shown in Equation (1) D L G is minimized.
[0040]
Equation
[0041] In Equation (1), "c" represents the distortion-free audio signal S6, "x" represents the audio enhancement signal S2 obtained by the non-linear audio enhancement means 11, and "λ" represents a coefficient for adjusting the balance between the discrimination loss S7 and the input-output loss S8.
[0042] In Equation (1), L BCE (c) is the discrimination loss (discrimination loss S7) in the discriminator 122. Here, as the loss function (the loss function applied to L BCE (c)), binary cross-entropy loss is used, but it is not limited to this. As the loss function used for the discrimination loss (discrimination loss S7) of the discriminator 122, for example, least squares loss or EMD (Earth Mover Distance) may be applied.
[0043] Also, in Equation (1), L L1 (x, G(x)) represents the input-output loss (input-output loss S8) between the input signal S4 and the output signal S5 in the distortion correction DNN 121 (generator). That is, L L1 (x, G(x)) represents the input-output loss S8 calculated by the input-output loss calculation means 123. Here, as the loss function (L L1The loss function applied to (x, G(x)) is assumed to be the L1 loss, but is not limited to this. For example, the input / output loss calculation means 123 (input / output loss S8) may use the L2 loss as the loss function.
[0044] Furthermore, in equation (1), G(x) is the output signal S5 for the input signal S4 of the distortion correction DNN121 (generator). In equation (1), L1 loss (L L1 (x, G(x)) was added as a constraint to maintain the waveform characteristics of the audio signal in the output signal S5.
[0045] Next, we will explain the specific model of the distortion correction DNN121.
[0046] This section describes the model structure when constructing the distortion-corrected DNN121 within the framework of GAN (Generative Adversarial Network). Here, the signals processed by the distortion-corrected DNN121 (input signal S4, output signal S5) are assumed to be time-frequency domain signals obtained by the short-time Fourier transform.
[0047] While any DNN model can be applied to Distortion Correction DNN121, it is preferable to apply the U-net type, an encoder-decoder type DNN that is widely used in GAN (Generative Adversarial Network) based speech enhancement. For example, the model described in Reference 2 can be applied to Distortion Correction DNN121. [Reference 2] Olaf Ronneberger, Philipp Fischer, Thomas Brox, “U-net: Convolutional Networks for Biomedical Image Segmentation”, MICCAI, 2015, [Retrieved February 11, 2022], [Online] INTERNET,<URL: https: / / arxiv.org / pdf / 1505.04597.pdf > Any model used within the GAN framework can be applied to the classifier 122. In this embodiment, the model applied to the classifier 122 is described as one of the following two types, but is not limited to these. Figure 4 is a diagram (conceptual) showing an example of a model applied to the classifier 122 in this embodiment.
[0048] In this embodiment, the first model applied to the classifier 122 is a model that performs two-dimensional convolution (2D convolution) on the entire input time-frequency spectrum and determines truth or falsity for the entire input (hereinafter referred to as the "two-dimensional convolutional model" or "2DConvGAN"). Examples of the two-dimensional convolutional model (2DConvGAN) applied to the classifier 122 include those described in references 3 and 4. [Reference 3] Santiago Pascual, Antonio Bonafonte, Joan Serra, “SEGAN: Speech Enhancement Generative Adversarial Network”, arXiv preprint arXiv:1703.09452, 2017, [Retrieved February 11, 2022], [Online] INTERNET,<URL: https: / / arxiv.org / pdf / 1703.09452.pdf> [Reference 4] Alec Radford, Luke Metz, Soumith Chintala, “UNSUPERVISED REPRESENTATION LEARNING WITH DEEP CONVOLUTIONAL GENERATIVE ADVERSARIAL NETWORKS”, CoRR abs / 1511. 06434, 2015, [Retrieved February 11, 2022], [Online] INTERNET,<URL: https: / / arxiv.org / pdf / 1511.06434.pdf > Furthermore, the second model applied to the discriminator 122 of this embodiment is a model that performs convolution up to the final layer and determines truth or falsity for each local patch of the input spectrum (hereinafter referred to as the "local patch type model" or "PatchGAN"). An example of a local patch type model (PatchGAN) applied to the discriminator 122 is the configuration shown in Reference 5. In the speech-enhanced signal S2 obtained by performing nonlinear speech enhancement processing on the observed signal S1, residual noise and artificial processing distortion occur locally on the time-frequency spectrum. Therefore, it is desirable to use a discriminator that determines truth or falsity (presence or absence of distortion) for each patch, and in this respect, the local patch type model (PatchGAN) is suitable. [Reference 5] Chuan Li, Michael Wand, “Precomputed Real-Time Texture Synthesis with Markovian Generative Adversarial Networks”, Proc. ECCV, 702-716, 2016, [Retrieved February 11, 2022], [Online] INTERNET,<URL: https: / / arxiv.org / pdf / 1604.04382.pdf > Figure 4(a) is an illustrative diagram showing an example where classifier 122 performs classification processing using a two-dimensional convolutional model, and Figure 4(b) is an illustrative diagram showing an example where classifier 122 performs classification processing using a local patch model.
[0049] In Figure 4, the matrix input to the classifier 122 as the signal to be classified (time-frequency spectrum) is shown as D101.
[0050] In Figure 4(a), the matrix D101a represents the result of a two-dimensional convolution of matrix D101 using a two-dimensional convolutional model. Also in Figure 4(a), R1 represents the numerical value of the classification result for D101 by the two-dimensional convolutional model.
[0051] In the classification process using a two-dimensional convolutional model, as shown in Figure 4(a), a single numerical value is output as the classification result R1 obtained by performing convolution on the entire input D101. Here, the classification result (truth / fake result) by the classifier 122 is assumed to be a numerical value in the range of 0.0 to 1.0.
[0052] In Figure 4(b), in the local patch model, a portion (patch) of the input matrix D101 is defined as D201. In Figure 4(b), the matrix representing the process of two-dimensional convolution of patch region D201 in the local patch model is defined as D201a. As shown in Figure 4(b), the local patch model outputs a single numerical value (a number in the range of 0.0 to 1.0) as the identification result R201 obtained as a result of convolution on patch region D201. In the local patch model shown in Figure 4(b), the entire input D101 is divided into 16 (4x4) patches (blocks), and two-dimensional convolution is performed to obtain 16 (4x4) numerical identification results (numerical values in the same format as R201). In Figure 4(b), the entire set of identification results (16 identification results) for each patch is referred to as the identification result group R2. In the model shown in Figure 4(b), 16 patches (4x4) are set on input D101 for the sake of simplicity. However, when applying a local patch type model in classifier 122, the number and location (range) of patches set on input D101 are not limited. Classifier 122 performs an evaluation of the entire input D101 based on the classification result group R2, and outputs a single numerical value (a number in the range of 0.0 to 1.0) as the final classification result. In this case, the method by which classifier 122 evaluates the classification result group R2 is not limited. For example, classifier 122 may output the average value of each numerical value constituting the classification result group R2 as the final classification result. Alternatively, classifier 122 may extract some numerical values (for example, a predetermined number of numerical values from the upper or lower end) from the numerical values constituting the classification result group R2 and output the average value of the extracted numerical values as the final classification result.
[0053] The distortion correction means 12 in this embodiment supports both an operation mode in which a learning process is performed on the distortion correction DNN 121 (hereinafter referred to as the "learning process mode") and an operation mode in which distortion correction processing of the audio enhancement signal S2 is performed on the distortion correction DNN 121 (hereinafter referred to as the "signal processing mode").
[0054] When the distortion correction means 12 operates in learning processing mode, it is supplied with training data that includes samples of audio signals containing distortion due to nonlinear audio enhancement processing (hereinafter referred to as "distortion-containing audio signals") (samples that become input signals S4) and samples of clean audio signals without distortion (samples that become undistorted audio signals S6). In this case, the distortion correction means 12 causes the distortion correction DNN 121 to perform adversarial learning using the training data within the framework of a GAN as shown in Figure 3. As a result, the distortion correction DNN 121 can acquire a learning model that has been learned (deep learning) based on the supplied training data.
[0055] As described above, in the signal processing device 10 of this embodiment, in order to acquire a learning model for converting a speech-enhanced signal S2, which includes distortion processed by a nonlinear speech enhancement technique, into a distortion-free speech signal, a distortion correction DNN 121 is learned by unsupervised learning based on adversarial learning (GAN). In the framework of adversarial learning (GAN), the distortion correction DNN 121 corresponds to a generator and is learned to deceive a discriminator 122 that distinguishes between a distortion-free speech signal S6 (true) and the output signal S5 (false) of the generator. Since the artificial processing distortion and residual noise caused by the speech enhancement technique occur locally on the time-frequency spectrogram, it is preferable in the signal processing device 10 of this embodiment to apply a local patch type model (PatchGAN) to the discriminator 122's determination of whether or not distortion is present. Furthermore, in the signal processing device 10 of this embodiment, the input / output loss calculation means 123 feeds back the input / output loss S8, which is the loss between the input signal S4 (a signal including signal distortion and residual noise) and the output signal S5 of the distortion correction DNN 121, to the distortion correction DNN 121. Moreover, in this embodiment, as shown in equation (1), the objective function of the distortion correction DNN 121 is configured to include the input / output loss S8. Furthermore, in the signal processing device 10 of this embodiment, the distortion correction DNN 121 is configured as a U-net type, which is an encoder-decoder type DNN.
[0056] (A-2) Operation of the embodiment Next, the operation of the signal processing device 10 of this embodiment, which has the configuration described above (signal processing method according to the embodiment), will be explained.
[0057] First, we will explain the processing when the distortion correction means 12 (distortion correction DNN 121) of the signal processing device 10 is operating in learning processing mode.
[0058] When training data is supplied to the distortion correction means 12 operating in learning processing mode, the distortion correction means 12 inputs the training data into the GAN framework shown in Figure 3, causing the distortion correction DNN 121 to perform training processing (learning the process of extracting target area sounds using a neural network). At this time, the training data includes samples of distortion-containing audio signals and samples of distortion-free audio signals.
[0059] In the GAN framework shown in Figure 3, a sample of a distorted audio signal included in the training data is supplied as an input signal S4 to the distortion correction DNN 121 and the input / output loss calculation means 123. In addition, an undistorted audio signal included in the training data is supplied as an undistorted audio signal S6 to the classifier 122. As a result, the input signal S4 is processed by the DNN in the distortion correction DNN 121, and the processing result is output as an output signal S5. The classifier 122 performs discrimination processing on the output signal S5, and the discrimination loss S7 is obtained as a result of the discrimination processing and fed back to the distortion correction DNN 121. Furthermore, the input / output loss calculation means 123 obtains the input / output loss (L1 loss) between the input signal S4 and the output signal S5 and feeds it back to the distortion correction DNN 121. Through the above processing, the distortion correction DNN 121 performs learning processing (learning of distortion correction processing by deep neural network).
[0060] Next, we will describe the operation of the distortion correction means 12 (distortion correction DNN121) of the signal processing device 10 when it is operating in signal processing mode.
[0061] The observed signal S1 is supplied to the nonlinear speech enhancement means 11, where nonlinear speech enhancement processing is performed on the observed signal and a speech enhancement signal S2 is output. Then, this speech enhancement signal S2 is supplied to the distortion correction means 12 (distortion correction DNN 121) which operates in signal processing mode, where the distortion correction DNN 121 performs distortion correction on the speech enhancement signal S2 using a trained DNN and outputs a distortion-corrected signal S3.
[0062] Next, the present inventor will explain the simulation (hereinafter referred to as "this simulation") that was performed to construct and evaluate the quality of the signal processing device 10.
[0063] First, let's explain the conditions for this simulation.
[0064] Figure 5 shows the model (conditions) for acquiring (observing) the observed signal S1 in this simulation.
[0065] In this simulation, as shown in Figure 5, it is assumed that the two microphone arrays MA1 and MA2 (2-channel microphone arrays), the target sound source, and the interfering sound source are all located on the same plane. Furthermore, in this simulation, the room that constitutes the sound field of the model environment shown in Figure 5 is assumed to be 7m x 7m x 3m (a room with a floor area of 7m x 7m and a height of 3m). Additionally, reverberation is excluded from the simulation conditions.
[0066] In Figure 5, the lines connecting the positions (center positions) of the two microphones M1 and M2 in microphone arrays MA1 and MA2 are denoted as L1 and L2, respectively. Also in Figure 5, the midpoint between the positions (center positions) of the two microphones M1 and M2 in microphone arrays MA1 and MA2 (the center point of the microphone array; the midpoint on lines L1 and L2) are indicated as P1 and P2, respectively. Furthermore, in Figure 5, the midpoint of the line L0 connecting positions P1 and P2 in microphone arrays MA1 and MA2 (the midpoint between microphone arrays MA1 and MA2) is indicated as P0. In addition, in Figure 5, the direction of microphone array MA2 (position P2) relative to P0 is defined as 0°, and the direction of microphone array MA1 (position P1) relative to P0 is defined as 180°, and the target sound source and interfering sound source are assumed to arrive from an angle between 0° and 180° relative to P0. In the following, the direction from P0 where the apparent sound source and interfering sound source are located will also be referred to as the "angle of arrival" or "direction of arrival." In Figure 5, the angle between line L0 and line L1, which indicates the orientation of the microphone array MA1, is θ. MA1Let θ be the angle between line L0 and line L2, which indicates the orientation of the microphone array MA2. MA2 That is what they say.
[0067] In this simulation, the distance between microphones M1 and M2 in each microphone array MA1 and MA2 was set to 3 cm. Furthermore, the distance between microphone arrays MA1 and MA2 (the distance between positions P1 and P2) was set to 40 cm. Additionally, in this simulation, θ MA1 , θ MA2 The angles were set to 25° for each. In other words, in this simulation, each microphone array MA1 and MA2 is positioned at an angle of 25° from the front.
[0068] Figure 6 shows the positions of each sound source within the environment shown in Figure 5 in this simulation.
[0069] As shown in Figure 6, the target sound source is located on a semicircle at a distance of 0.4m from P0, and the interfering sound source (sound source in the non-target area) is located on the line of a semicircle at a distance of 0.8m from P0. In this simulation, the direction of arrival of the target sound source was set to the front direction (90°), and the direction of arrival of the interfering sound source was set to one of the following directions: 15°, 45°, 135°, or 165°.
[0070] In this simulation, observation signals (acoustic signals) captured by microphone arrays MA1 and MA2 in a model environment as shown in Figures 5 and 6 were acquired through computer simulation, and the results of inputting the acquired observation signals into a signal processing device 10 were evaluated. Specifically, in this simulation, PyRoomAcoustics (see Reference 6 below) was used to set up a model environment as shown in Figures 5 and 6 and acquire the impulse response. The acquired impulse response was then convolved with the dry sources (dry sources of the target sound source and the interfering sound source) to obtain the observation signal S1 (observation signal of microphone arrays MA1 and MA2).
[0071] [Reference 6] Scheibler, E. Bezzam, I. Dokmani´c, “Pyroomacoustics: A Python package for audio room simulations and array processing algorithms”, Proc. IEEE ICASSP, 2018 In this simulation, 2310 utterances (speech data) from the TIMIT corpus (see Reference 7 below) were used as the sound source (target sound source and interference sound source) used as the dry source signal when acquiring the observed signal S1 (when acquired in the simulation environment shown in Figure 5), and as the sound source for the distortion-free speech signal S6 input to the classifier 122 (hereinafter referred to as "training distortion-free speech data").
[0072] [Reference 7] JS Garofolo, LF Lamel, WM Fisher, JGFiscus, DS Pallett, NL Dahlgren, V. Zue, “TIMIT acoustic phonetic continuous speech corpus,” Linguistic Data Consotrium, 1992. In this simulation, the U-net type DNN constituting the distortion correction DNN121 was configured with eight two-dimensional convolutional layers (Conv2D × 8 layers) on the encoder side (first half) and eight two-dimensional deconvolutional layers (Conv2DTrans × 8 layers) on the decoder side (second half). Furthermore, in this simulation, the input and output signals to the distortion correction DNN121 were 16kHz audio data. Additionally, the number of parameters in the U-net type DNN constituting the distortion correction DNN121 was set to 5,782,2337.
[0073] In this simulation, we evaluated both the case where a two-dimensional convolutional model (2DConvGAN) and the case where a local patch model (PatchGAN) was applied to the classifier 122. Furthermore, in this simulation, a five-layer two-dimensional convolutional layer (2DConv × 5 layers) was applied to the classifier 122. In addition, in this simulation, the structure of the two types of classifier 122 was adjusted so that the number of parameters was approximately the same, so that the difference in the number of parameters would not affect the evaluation results. Specifically, in this simulation, the number of parameters for classifier 122 when the two-dimensional convolutional model (2DConvGAN) was applied was set to 2,792,129, and the number of parameters for classifier 122 when the local patch model (PatchGAN) was applied was set to 2,764,481. Furthermore, in this simulation, when the local patch model (PatchGAN) was applied to classifier 122, 31 × 20 patches were set on the time-frequency spectrum of the output signal S5 for classification.
[0074] In this simulation, 11,000 mixed speeches obtained by superimposing the target sound source and interference sound source at levels from -5dB to 5dB were used as the observation signal S1 (hereinafter referred to as "training observation data") used during training (training processing mode). In addition, 1,000 mixed speeches obtained by superimposing the target sound source and interference sound source at levels of -3[dB], 0[dB], and 3[dB] were used as the observation signal S1 (hereinafter referred to as "evaluation observation data") used during evaluation (signal processing mode). Hereafter, the level at which the target sound source and interference sound source are superimposed on the observation signal S1 will be referred to as the "superposition level". Note that the sound sources (dry source signals) from which the training undistorted speech data, training observation data, and evaluation observation data are derived are different, and the speakers are also different.
[0075] In this simulation, MUBASE was used as the nonlinear speech enhancement processing applied to the nonlinear speech enhancement means 11, as described above. In this simulation, the training observation data was processed by MUBASE (area sound acquisition processing) and input as the input signal S4 to the distortion correction means 12 (distortion correction DNN 121).
[0076] In this simulation, Adam (see reference 8 below) was used as the optimization algorithm during the training of the distortion-corrected DNN121 (GAN framework shown in Figure 3). In this simulation, during the training of the distortion-corrected DNN121 (GAN framework shown in Figure 3), λ in equation (1) was set to 3.5, the mini-batch size to 100, the number of epochs to 250, and the learning rate to 0.001.
[0077] [Reference 8] D. Kingma, and J. Ba, “Adam: A method for stochastic optimization”, International Conference on Learning Representations (ICLR), 2015. Next, the results of this simulation will be explained using Figures 7 to 9.
[0078] Figures 7 to 9 show the results of evaluating the sound quality of the unprocessed observation signal S1 (hereinafter also referred to as "Observation"), the voice-enhanced signal S2 (a signal that has undergone voice-enhanced processing (area pickup) using the conventional MUBASE) (hereinafter also simply referred to as "MUBASE"), and the distortion-corrected signal S3 (a signal obtained by distortion correction processing of the voice-enhanced signal S2 using the distortion correction DNN121) in this simulation. Figures 7 to 9 show the sound quality evaluation results for the distortion-corrected signal S3, specifically for the signal that has undergone distortion correction processing using a learning model that applies a 2DConvGAN (two-dimensional convolutional model) (hereinafter also referred to as "MUBASE-2DConvGAN") and the signal that has undergone distortion correction processing using a learning model that applies a PatchGAN (local patch model) (hereinafter also referred to as "MUBASE-PatchGAN").
[0079] Figures 7 to 9 show the evaluation results of speech quality for Observation, MUBASE, MUBASE-2DConvGAN, and MUBASE-PatchGAN, respectively, when the superposition level of the evaluation observation data is varied by -3dB, 0dB, and 3dB. In Figures 7 to 9, the speech quality evaluation metrics PESQ (Perceptual Evaluation Of Speech Quality), STOI (Short-Time Objective Intelligibility), and SDR (Signal-to-Distortion Ratio) are used as measures to evaluate the distortion correction performance of the speech signal, respectively.
[0080] From the evaluation results in Figures 7 to 9, it can be seen that for all evaluation metrics (PESQ, STOI, and SDR), the output corrected with distortion correction DNN121 (MUBASE-2DConvGAN and MUBASE-PatchGAN) shows improved audio quality compared to the output from MUBASE. Furthermore, from the evaluation results in Figures 7 to 9, it can be seen that, among the outputs corrected with distortion correction DNN121 for all evaluation metrics (PESQ, STOI, and SDR), MUBASE-PatchGAN (distortion correction processing using a local patch model) has higher sound quality than MUBASE-2DConvGAN (distortion correction processing using a two-dimensional convolution model). Thus, it is clear that the sound quality of the MUBASE output is improved by distortion correction DNN121, and that MUBASE-PatchGAN (distortion correction processing using a local patch model) is superior.
[0081] (A-3) Effects of the Embodiment This embodiment can achieve the following effects.
[0082] In this embodiment, the signal processing device 10 corrects the distortion of the speech enhancement signal S2 using a distortion correction DNN 121 that performs adversarial learning using a GAN framework, as described above. As a result, the signal processing device 10 in this embodiment can perform distortion correction processing using a trained DNN without requiring paired data (input signal S4 and its corresponding ground truth signal). Furthermore, as shown in the simulation results above, in this embodiment, by performing distortion correction processing using the distortion correction DNN 121, distortion and residual noise of the signal generated by the nonlinear processing (speech enhancement processing) by the nonlinear speech enhancement means 11 can be reduced without generating any further distortion after processing, thereby obtaining a speech enhancement signal that is pleasant to listen to.
[0083] Furthermore, in this embodiment of the signal processing device 10, an example is shown in which a two-dimensional convolutional model (2DConvGAN) or a local patch model (PatchGAN) is applied as the model of the discriminator 122 used for training the distortion correction DNN 121. Since artificial processing distortion and residual noise caused by speech enhancement technology occur locally on the time-frequency spectrogram, it is preferable to apply a local patch model (PatchGAN) to the discriminator 122 for determining whether or not distortion is present. The suitability of applying a local patch model (PatchGAN) to the discriminator 122 is also supported by the simulation results described above.
[0084] Furthermore, in the signal processing device 10 of this embodiment, the input / output loss calculation means 123 processes the objective function of the distortion correction DNN 121 so that it includes the loss (input / output loss S8) between the input signal S4 (a signal including signal distortion and residual noise) and the output signal S5. If the signal processing device 10 were not equipped with the input / output loss calculation means 123, it would suffice for the determination by the classifier 122 to be true, which could lead to the DNN learning a distortion correction process where, for example, the volume of the output signal S5 fluctuates wildly regardless of the volume of the input signal S4. However, in the signal processing device 10 of this embodiment, by being equipped with the input / output loss calculation means 123, such learning can be suppressed, and an output signal S5 can be obtained in which distortion has been corrected to have characteristics similar to the input signal S4 in the output signal S5 of the distortion correction DNN 121.
[0085] (B) Other embodiments The present invention is not limited to the embodiments described above, and modified embodiments such as those exemplified below can also be cited.
[0086] (B-1) In the signal processing device 10 (distortion correction means 12) of the above embodiment, a configuration that does not support the learning processing mode is also possible (for example, a configuration in which a learning model has already been acquired or a learning model is acquired from an external source). Note that if the distortion correction means 12 also supports the learning processing mode (if it supports both the signal processing mode and the learning processing mode, it is necessary to provide a discriminator 122 and an input / output loss calculation means 123. On the other hand, if the distortion correction means 12 does not support the learning processing mode (if it only supports the signal processing mode), the discriminator 122 and the input / output loss calculation means 123 may be omitted.
[0087] (B-2) In the above embodiment, the signal processing device 10 was configured to include a nonlinear speech enhancement means 11, but it may also be configured to include only a distortion correction means 12 and to perform only the processing of correcting distortion from the supplied speech enhancement signal S2.
[0088] (B-3) In the above embodiment, L1 loss and L2 loss were given as examples of losses calculated by the input / output loss calculation means 123. However, in this case, the output signal S5 is made to resemble the input signal S4 which contains artificial processing distortion and residual noise, so there is a risk that the processing distortion and residual noise in the output signal S5 may not be fully corrected. For this reason, when training the distortion correction means 12, unsupervised learning may be performed using the framework of an adversarial network with cycle consistency. Examples of adversarial networks applicable to the distortion correction means 12 include the technology described in Reference 9. [Reference 9] Zhong Meng, Jinyu Li, Yifan Gong, Biing-Hwang (Fred) Juang, “S Cycle-Consistent Speech Enhancement”, arXiv:1809.02253v2 [eess.AS] 30 Apr 2019, [Retrieved February 15, 2022], [Online] INTERNET,<URL: https: / / arxiv.org / pdf / 1809.02253.pdf > Figures 10 to 12 are block diagrams showing the configuration when a framework of adversarial networks with cyclical consistency is applied during the learning of the distortion correction means 12.
[0089] In this case, the distortion correction means 12, as shown in Figure 10, further includes a distortion restoration DNN 124 that corresponds to the inverse transform of the distortion correction DNN 121 in the learning processing mode (during learning), a second classifier 125 (hereinafter also called "distortion classifier 125") which, contrary to the classifier 122 (hereinafter also called "distortion-free classifier 122A"), distinguishes signals containing processing distortion and residual noise as true values and distortion-free audio signals as false values, and a second input / output calculation means 126 (hereinafter also called "distortion restoration loss calculation means 126") which acquires the input / output losses of the distortion restoration DNN 124. In the following, the input / output loss calculation means 123 will also be referred to as "distortion correction loss calculation means 123A".
[0090] In this case, the distortion correction means 12 operating in learning processing mode will perform combined learning of the distortion correction DNN 121 and the distortion restoration DNN 124.
[0091] At this time, the objective functions of the distortion correction DNN121 and the distortion restoration DNN124 are: (a) the distortion-free identification loss S7(Ldc) obtained by inputting the output signal S5(Yo) obtained by passing the input signal S4(X) including processing distortion and residual noise through the distortion correction DNN121 to the distortion-free identification loss S7(Ldc) obtained by inputting the distortion-free identification loss S7(Ldc) to the distortion-free identification loss obtained by passing the input signal S4(X) and the output signal S5(Yo) through the distortion restoration DNN124 to the distortion restoration signal S9(Xr) The distortion correction loss S14(Lcc) includes (c) a distortion restoration loss S10(Lnn), (c) a distortion restoration signal S11(Xo) obtained by passing the distortion-free audio signal S6(Y) through the distortion restoration DNN124 and inputting the resulting distortion restoration signal S11(Xo) to the distortion discriminator 125, and (d) a distortion correction loss S14(Lcc) obtained by passing the distortion-free audio signal S6(Y) and the distortion restoration signal S11(Xo) through the distortion correction DNN121 to the distortion correction signal S13(Yr). Furthermore, the objective functions of the distortion correction DNN121 and distortion restoration DNN124 may also include, (e) as shown in Figure 11, the identity distortion loss S16(Lin) between the input signal S4(X) including processing distortion and residual noise and the identity distortion signal S15(Xi) obtained by passing X through the distortion restoration DNN124, and (f) as shown in Figure 12, the identity distortion-free loss S18(Lic) between the distortion-free audio signal S6(Y) and the identity distortion-free signal S17(Yi) obtained by passing Y through the distortion correction DNN121. Here, the parameters of the distortion correction DNN121 (generator) are obtained by minimizing the objective function L(F,G,Dv,Du) shown in equation (2).
[0092]
number
[0093] In equation (2), F is the strain correction DNN121 (generator), G is the strain restoration DNN124, Dv is the distortion-free discriminator 122A, and Du is the strain discriminator 125. Also, Lnn is the strain restoration loss S10, Lcc is the strain correction loss S14, Ldc is the distortion-free discriminator loss S7, Ldn is the strain discriminator loss S12, Lin is the identity strain loss S16, and Lic is the identity distortion-free loss S18. Furthermore, λ1, λ2, λ3, λ4, and λ5 represent coefficients that adjust the balance of multiple losses. [Explanation of Symbols]
[0094] 10...Signal processing device, 11...Nonlinear voice enhancement means, 12...Distortion correction means, 20...Microphone array unit, 122...Identifier, 123...Input / output loss calculation means, M, M1, M2...Microphones, MA, MA1, MA2...Microphone array, S1...Observation signal, S2...Voice enhancement signal, S3...Distortion-corrected signal, S4...Input signal, S5...Output signal, S6...Distortion-free voice signal, S7...Identification loss, S8...Input / output loss.
Claims
1. It includes a distortion correction means that corrects the distortion of an input signal that has undergone nonlinear audio enhancement processing using a deep neural network, The deep neural network has a learning model trained within the framework of a generative adversarial network using a generator and a discriminator, based on training data including samples of distorted audio signals and samples of undistorted audio signals. The generator uses the deep neural network described above. The objective function of the deep neural network, which acts as the generator, includes the input-output loss between the input and output signals of the deep neural network. A signal processing device characterized by the following:
2. The signal processing apparatus according to claim 1, characterized in that the loss function applied to the input / output loss is L1 loss.
3. The signal processing device according to claim 1, characterized in that the deep neural network has a learning model learned within the framework of an adversarial network having consistency through circulation.
4. The signal processing device according to any one of claims 1 to 3, characterized in that the discriminator performs identification of the presence or absence of distortion for each local patch.
5. The signal processing device according to any one of claims 1 to 4, characterized in that the input signal is an acoustic signal obtained by area sound pickup processing, which captures sound from a target area with the target area as the sound source, based on the beamformer output of a plurality of microphone arrays.
6. The signal processing device according to any one of claims 1 to 5, characterized in that the deep neural network is composed of a U-net type model.
7. Computers, This system functions as a distortion correction method that uses a deep neural network to correct distortion in input signals that have undergone nonlinear audio enhancement processing. The deep neural network has a learning model trained within the framework of a generative adversarial network using a generator and a discriminator, based on training data including samples of distorted audio signals and samples of undistorted audio signals. The generator uses the deep neural network described above. The objective function of the deep neural network, which acts as the generator, includes the input-output loss between the input and output signals of the deep neural network. A signal processing program characterized by the following:
8. In a signal processing method performed by a signal processing device, The signal processing device includes distortion correction means, The distortion correction means corrects the distortion of the input signal that has undergone nonlinear audio enhancement processing using a deep neural network. The deep neural network has a learning model trained within the framework of a generative adversarial network using a generator and a discriminator, based on training data including samples of distorted audio signals and samples of undistorted audio signals. The generator uses the deep neural network described above. The objective function of the deep neural network, which acts as the generator, includes the input-output loss between the input and output signals of the deep neural network. A signal processing method characterized by the following:
Citation Information
Patent Citations
Series data converter, learning apparatus, and program
JP2019101391A
Signal processing device, signal processing program, signal processing method, and sound collection device
JP2020012980A