Signal processing device, signal processing method, and signal processing program

The signal processing device enhances speech recognition by adding the original observation signal to the enhancement signal, reducing artifact components and improving performance in single-channel systems.

JP7722467B2Active Publication Date: 2025-08-13NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023564720
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-08-13
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

Single-channel speech enhancement technologies have limited effectiveness in improving speech recognition performance, and even enhanced signals can degrade recognition compared to noisy observed signals.

Method used

A signal processing device that generates an emphasis signal from an observation signal, adds the original observation signal to the enhancement signal, and performs speech recognition on the combined signal to reduce artifact components.

Benefits of technology

Improves speech recognition performance by reducing the impact of artifact elements in single-channel speech enhancement, demonstrated through increased Signal-to-Artifact Ratio (SAR) and decreased Word Error Rate (WER).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007722467000012
    Figure 0007722467000012
  • Figure 0007722467000013
    Figure 0007722467000013
  • Figure 0007722467000014
    Figure 0007722467000014
Patent Text Reader

Abstract

A signal processing device (10) comprises: a speech enhancement unit (11) that generates, from an observed signal, an enhanced signal obtained by enhancing the speech of a speaker; an original sound addition unit (12) that adds the observed signal to the enhanced signal; and a speech recognition unit (13) that performs speech recognition on the enhanced signal to which the observed signal has been added by the original sound addition unit (12).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a signal processing device, a signal processing method, and a signal processing program. [Background technology]

[0002] Building a speech recognition system that is robust against acoustic interference such as background noise and reverberation is a challenge in speech processing. Here, it has been confirmed that multi-channel speech enhancement technology (beamformer) using multiple microphones can significantly improve speech recognition performance. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Szu-Jui Chen, Aswin Shanmugam Subramanian, Hainan Xu, and Shinji Watanabe, "Building state-of-the-art distant speech recognition using the chime-4 challenge with a setup of speech enhancement baseline", in Interspeech, 2018, pp. 1571-1575. Summary of the Invention [Problem to be solved by the invention]

[0004] On the other hand, single-channel speech enhancement technology using a single microphone has limited effectiveness in improving speech recognition performance, as even when using an enhancement signal with noise removed, speech recognition performance can sometimes be worse than when using a noisy observed signal.

[0005] In reality, many devices only have a single microphone. Therefore, to realize a robust speech recognition system, it is important to develop speech enhancement techniques for single channels as well as multi-channels.

[0006] The present invention has been made in view of the above, and has an object to provide a signal processing device, a signal processing method, and a signal processing program that enable improvement of speech recognition performance through speech enhancement. [Means for solving the problem]

[0007] In order to solve the above-mentioned problems and achieve the object, a signal processing device according to the present invention is characterized by having a speech emphasis unit that generates an emphasis signal in which a speaker's speech is emphasized from an observation signal, an addition unit that adds the observation signal to the emphasis signal, and a speech recognition unit that performs speech recognition on the emphasis signal to which the observation signal has been added by the addition unit. [Effects of the Invention]

[0008] According to the present invention, it is possible to improve speech recognition performance by speech enhancement. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating signal decomposition of an enhancement signal by orthogonal projection. [Figure 2] FIG. 2 is a diagram showing the word error rate (WER) for the evaluation emphasis signal. [Figure 3] FIG. 3 is a diagram illustrating signal decomposition of a modified enhancement signal obtained by adding an observation signal to an enhancement signal. [Figure 4] FIG. 4 is a diagram illustrating an example of a configuration of a signal processing device according to an embodiment. [Figure 5] FIG. 5 is a flowchart showing the processing procedure of the signal processing method according to the embodiment. [Figure 6] FIG. 6 is a diagram showing the SDR, SNR, and SAR for the modified enhancement signal. [Figure 7] FIG. 7 shows the WER scores for the modified enhancement signals. [Figure 8] FIG. 8 is a diagram showing the WER scores of the signal processing device for the observed signals from the actual recording. [Figure 9] FIG. 9 is a diagram illustrating an example of a computer that implements a signal processing device by executing a program. DETAILED DESCRIPTION OF THE INVENTION

[0010] An embodiment of the present invention will be described in detail below with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the drawings, identical parts are denoted by the same reference numerals. Note that, hereinafter, when "^A" is written for A, which is a vector or matrix, it is assumed to be the same as "a symbol with "^" written immediately above "A". When " ̄A" is written for A, which is a vector or matrix, it is assumed to be the same as "a symbol with " ̄" written immediately above "A".

[0011] [Embodiment Mode] In this embodiment, as an example, a signal processing method for improving speech recognition performance is proposed based on the results of an analysis of factors that cause degradation of speech recognition performance by an enhanced signal obtained by single-channel speech enhancement (SE). Note that, in this embodiment, a signal processing method for a speech signal (observed signal) recorded by a single microphone (single channel) is described, but it is not limited to a single channel, and can also be applied to speech signals recorded by multiple microphones (multi-channel).

[0012] [Analysis of Enhanced Signals] First, we analyzed the factors that degrade speech recognition performance for signals enhanced by single-channel SE.

[0013] It is generally assumed that the processing distortion caused by single-channel SE is the cause of speech recognition performance degradation. However, there has been no systematic detailed analysis or elucidation of these distortions, particularly their impact on speech recognition. We believe that elucidating the impact of single-channel SE estimation errors on speech recognition is essential for improving SE front-end design.

[0014] Here we focus on the single-channel SE task, y ∈ R T denotes the T long-term domain waveform of the observed signal. The observed signal y is modeled as equation (1). s∈R T denotes the source signal. n∈R T denotes the background noise signal.

[0015]

number

[0016] The purpose of SE is to reduce the noise signal n from the observed signal y. When the observed signal y is input, the enhancement signal ^s∈R T is estimated as ^s = SE(y), where SE(·) denotes the SE processing performed, for example, by a neural network.

[0017] Next, to analyze the effect of the SE estimation error on speech recognition performance, we investigated the SE estimation error decomposition using orthogonal projection. Figure 1 illustrates the signal decomposition of the enhancement signal using orthogonal projection.

[0018] Since the enhancement signal ^s is obtained by performing estimation processing, it is inevitable that it will contain estimation errors. The enhancement signal ^s is decomposed using orthogonal projection as shown in Equation (2).

[0019]

number

[0020] In equation (2), s targetindicates the target sound source element, and e noise ∈R T indicates the noise factor (error), and e artif ∈R T indicates the artifact factor (error) (see Figure 1).

[0021] Specifically, orthogonal projection error decomposition decomposes the error in the SE into a noise component and an artifact component, which are obtained by projecting the SE error onto a speech / noise subspace spanning the speech / noise signal and a subspace orthogonal to the speech / noise subspace.

[0022] Noise element e noise Since the input signal consists of a linear combination of a speech signal and a noise signal, it is expected to be a naturally observable signal. These are called natural signals. Since similar noise elements naturally appear in training samples, the influence of these natural signals on speech recognition performance may be limited.

[0023] On the other hand, the artifact element e artif is composed of signals that cannot be represented by a linear combination of speech and noise signals (see Figure 1), and is an artificial / unnatural signal. This unnatural signal may be highly diverse and may rarely appear in training samples. Therefore, we hypothesize that speech recognition is more sensitive to artifact elements than noise elements.

[0024] The SE evaluation indices used are the signal to distortion ratio (SDR) (equation (3)), the signal to noise ratio (SNR) (equation (4)), and the signal to artifact ratio (SAR) (equation (5)).

[0025]

number

[0026]

number

[0027]

number

[0028] Next, the artifact element e artif We conducted an experiment to investigate the influence of noise factors on speech recognition performance. artif and noise element e noise In order to measure the effect of the error on speech recognition performance, the emphasis signal was modified by changing the magnitude of the error component, and speech recognition was performed using the modified emphasis signal as input.

[0029] Specifically, after decomposing the enhancement signal s using orthogonal projection, the artifact element e artif and noise element e noise By increasing or decreasing as in equation (6), the enhancement signal ^s ω ∈R T was synthesized.

[0030]

number

[0031] ω noise is the noise element e noise is a parameter that controls the amount of nartif is the artifact element e artif In this experiment, we used various enhancement signals ^s with different ratios of noise and artifact elements. ω To obtain ω noise and ω artif The values of and were changed. This allows us to control the SNR and SAR values while maintaining the same target sound source element s target By inputting such a modified emphasis signal as an evaluation emphasis signal into a speech recognition system, the influence of each error element on speech recognition performance was directly measured.

[0032] Figure 2 shows the WER for the evaluation-enhanced signal. Figure 2(a) is a 3D plot showing the speech recognition results for the evaluation-enhanced signal with the noise / artifact error ratio changed. Figure 2(b) shows the WER for the evaluation-enhanced signal with the noise / artifact error ratio changed. noise and ω artif 2(b) shows the corresponding 2D plots obtained by changing only one of the weights. In Fig. 2(b), baseline(obs.) represents the baseline WER score of the observed signal, and the square symbols represent the WER score of the original enhanced signal without any changes. The same applies to baseline(obs.) and square symbols in Figs. 7 and 8.

[0033] As shown in Figure 2, it can be seen that the original enhanced signal actually degrades the speech recognition performance compared to the observed signal. As shown in Figure 2, the artifact element e artif It has been observed that speech recognition performance can be significantly improved by reducing the noise factor e noise These results show that the noise factor e noise and artifact element e artif and artifact element e artif It was confirmed that the effect of the two methods on the degradation of speech recognition performance was greater than that of the two methods.

[0034] Based on this knowledge, this embodiment proposes a signal processing method for improving speech recognition performance. In this embodiment, as an approach to reducing the influence of artifact elements, a method for reducing the ratio of artifact components in a signal input to a speech recognition system is considered.

[0035] In this embodiment, the original sound (observed signal) is added to the emphasis signal to reduce the proportion of artifact elements in the signal input to the speech recognition system. Specifically, a signal obtained by adding the scaled observed signal y to the emphasis signal ^s is input to the speech recognition system as the modified emphasis signal _s. The modified emphasis signal _s∈R T is calculated as shown in equation (7).

[0036]

number

[0037] ω obs ≧0 is a parameter that controls the amount of observed signal y added to the enhanced signal ^s. Figure 3 is a diagram illustrating signal decomposition of a modified enhanced signal obtained by adding an observed signal to an enhanced signal. As shown in Figures 1 and 3, the artifact element e artif corresponds to the perpendicular line of the enhancement signal s to the Sn plane. Even when the observation signal y is added to the enhancement signal s, the observation signal y is parallel to the Sn plane, so the artifact element e artif The length of the vector of does not change between the modified emphasis signal ̂s and the emphasis signal ̂s.

[0038] In contrast, the modified emphasis signal s is obtained by adding the observed signal y to the emphasis signal s, so that the target sound source element s is target and the noise element  ̄e noise Therefore, the modified enhancement signal  ̄s increases the artifact component e artif Therefore, by using the modified enhancement signal s, the ratio of the artifact element e artif This reduces the impact of the original voice on speech recognition, which is expected to improve speech recognition performance. Below, we can mathematically prove that adding the original voice contributes to improving speech recognition performance.

[0039] The SAR improvement value SARi is calculated as shown in Equation (8). If SARi>0, the artifact element e artif The ratio of Ps∈R decreases. T×T is the source signal {s T} L-1T=0 (L-1 is the number of maximum allowable delays) denotes the orthogonal projection matrix on the subspace spanned by P s,n ∈R T×T is the difference between the source signal and the noise signal {s T , n T} L-1T=0denotes the orthogonal projection matrix on the subspace spanned by and.

[0040]

number

[0041] In the second column of equations in (8), P s,n y=y and  ̄e artif =^e artif As shown in the third column of equation (8), <P s,n When s,y>>0, SARi>0. Therefore, to improve the SAR of the original enhancement signal ^s=SE(y), <P s,n The sufficient condition is s,y>>0. This sufficient condition can also be rewritten as equation (9), and under this relaxed condition, it can be proven that adding the original sound reduces the proportion of artifact components in the modified emphasis signal s.

[0042]

number

[0043] [Signal processing device] A signal processing device to which original sound addition is applied in order to improve speech recognition performance will be described below. Fig. 4 is a diagram schematically illustrating an example of the configuration of a signal processing device according to an embodiment.

[0044] A signal processing device 10 according to the embodiment is realized, for example, by loading a predetermined program into a computer or the like including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and having the CPU execute the predetermined program. The signal processing device 10 also has a communication interface for transmitting and receiving various information to and from other devices connected via a network, etc. As shown in FIG. 1, the signal processing device 10 has a voice enhancement unit 11, an original voice addition unit 12 (addition unit), and a voice recognition unit 13. An observed signal y recorded in a single channel is input to the signal processing device 10, and the signal processing device 10 outputs a voice recognition result, for example, by converting the sound signal into text.

[0045] The speech enhancement unit 11 receives an input of an observed signal y recorded via a single channel. The speech enhancement unit 11 generates an enhanced signal ^s by enhancing the speaker's voice from the observed signal y, with the aim of reducing a noise signal n from the observed signal y. The speech enhancement unit 11 performs speech enhancement processing using, for example, a neural network.

[0046] The original sound adding unit 12 adds the observed signal y (original sound) to the emphasis signal ^s. The original sound adding unit 12 adds the weighted observed signal y to the emphasis signal ^s, and inputs the resulting signal to the speech recognition unit 13 as a modified emphasis signal s (see equation (7)).

[0047] The original sound adding unit 12 adjusts the weight ω of the observed signal y to be added to the emphasis signal ^s in accordance with the ratio of the noise signal contained in the observed signal y. obs For example, when the ratio of the noise signal included in the observed signal y is lower than a certain value, the original sound adding unit 12 adjusts the weight ω obs Alternatively, if the ratio of the noise signal included in the observed signal y is higher than a certain value, the original sound adding unit 12 may reduce the value of the weight ω obs The original sound adding unit 12 estimates the SNR of the observed signal y, and based on this estimation result, sets the weight ω obs You may also determine the value of

[0048] In addition, the original sound adding unit 12 may weight both the observed signal y and the observed signal to be added to the enhancement signal ^s, such that the sum of the weight of the observed signal y and the weight of the observed signal to be added to the enhancement signal ^s is 1, as shown in equation (10).

[0049]

number

[0050] Furthermore, the original sound adding unit 12 may appropriately set a weight α of the observed signal y and a weight β of the observed signal to be added to the emphasis signal ^s, as shown in equation (11).

[0051]

number

[0052] The speech recognition unit 13 performs speech recognition on the modified emphasis signal s. The speech recognition unit 13 outputs a speech recognition result obtained by converting the voice signal into text, for example. The speech recognition unit 13 performs speech emphasis processing using, for example, a trained deep learning model.

[0053] [Signal processing method] Next, a description will be given of a signal processing method executed by the signal processing device 10. Fig. 5 is a flowchart showing the processing procedure of the signal processing method according to the embodiment.

[0054] As shown in Fig. 5, when the signal processing device 10 receives an input of an observed signal y, the speech enhancement unit 11 performs speech enhancement processing to generate an enhancement signal ^s that emphasizes the speaker's speech from the observed signal y (step S1). The original sound addition unit 12 performs original sound addition processing to add the observed signal y to the enhancement signal ^s (step S2). The original sound addition unit 12 inputs the signal obtained by adding the observed signal y to the enhancement signal ^s as a modified enhancement signal _s to the speech recognition unit 13. The speech recognition unit 13 performs speech recognition processing on the modified enhancement signal _s (step S3) and outputs the speech recognition result.

[0055] [Evaluation experiment] The speech recognition accuracy of the signal processing device 10 was actually evaluated. A neural network-based time-domain denoising network (Denoising-TasNet) was adopted as the speech enhancement unit 11. A deep neural network-hidden Markov model (DNN-HMM) hybrid ASR (Automatic Speech Recognition) system based on Kaldi's standard method was adopted as the speech recognition unit 13. Data sets of reproduced reverberant speech signals were generated from the Wall Street Journal (WSJ0) corpus as a speech source and the CHiME-3 corpus as a noise source, and these were used as a training set, development set, and evaluation set.

[0056] Fig. 6 shows the SDR, SNR, and SAR for the modified enhancement signal s. Fig. 7 shows the WER score for the modified enhancement signal s. Figs. 6 and 7 show the ω obs The results were obtained by varying the value of between 0.0 and 1.5.

[0057] As shown in Figure 6, obs As ω increases, i.e., with each additional observation, the SDR and SNR decrease, while the SAR increases monotonically. obs As s increases, an improvement in SAR is observed, and the ratio of artifact components to the modified enhancement signal s decreases. Following this improvement in SAR, an improvement in WER is observed, as shown in Figure 7.

[0058] Therefore, by adding the original sound, the signal processing device 10 was able to improve the speech recognition performance compared to the reference observed signal and the original emphasized signal s. In other words, the signal processing device 10 was able to improve the speech recognition performance of the single-channel SE front-end by reducing the proportion of artifact elements in the modified emphasized signal s, i.e., by increasing the SAR.

[0059] Next, we performed an evaluation on real recordings. To confirm the results of real recordings, we used the actually recorded speech data (et05_real) from the CHiME-3 dataset. Figure 8 shows the WER scores of the signal processing device 10 for observed signals from real recordings.

[0060] 8, it was observed that the signal processing device 10 reduced the WER even when applied to actual recordings. In other words, it was proven that the effect of improving speech recognition performance by reducing artifact elements also applies to actual recordings.

[0061] [Effects of the embodiment] In this way, the signal processing device 10 according to the embodiment adds the observed signal y to the enhancement signal ^s to reduce the influence of artifact elements on speech recognition performance, and inputs the result to the speech recognition unit 13. This has proven that the signal processing device 10 can monotonically increase the SAR value and improve speech recognition performance. It has also been found that the signal processing device 10 effectively improves speech recognition performance even in actual recordings.

[0062] Up until now, it has been difficult to improve speech recognition performance, especially with single-channel speech enhancement. Furthermore, there has been no prior work on adding source audio as a front-end for speech recognition.

[0063] The signal processing device 10 according to the present embodiment has succeeded in improving speech recognition performance in single-channel speech enhancement by simply adding a simple process of adding an original sound (observed signal) to an enhancement signal before speech recognition.

[0064] [System configuration of the embodiment] Each component of the signal processing device 10 is a functional concept and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of the signal processing device 10 is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.

[0065] Furthermore, all or any part of the processes performed in the signal processing device 10 may be realized by a CPU, a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU. Furthermore, each process performed in the signal processing device 10 may be realized as hardware using wired logic.

[0066] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.

[0067] [program] 9 is a diagram showing an example of a computer in which a program is executed to realize the signal processing device 10. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0068] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0069] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the signal processing device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration of the signal processing device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced with an SSD (Solid State Drive).

[0070] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.

[0071] The program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090, but may also be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0072] Although the present invention has been described above as an embodiment, the present invention is not limited to the descriptions and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention. [Explanation of symbols]

[0073] 10. Signal Processing Device 11 Speech enhancement unit 12 Original sound addition section 13 Voice Recognition Unit

Claims

1. A speech enhancement unit that generates an enhancement signal that enhances the speaker's speech from an observed signal by performing estimation processing; an adding unit that adds the observation signal, which is the original sound, to the emphasis signal; a speech recognition unit that performs speech recognition on the emphasis signal to which the observation signal has been added by the adding unit; and the enhancement signal includes an estimation error defined using an orthogonal projection in addition to the target sound source signal; the estimation error includes a noise component, which is a naturally observable signal, and an artifact component, which is an unnatural signal that rarely appears in training samples and cannot be represented by a linear combination of a speech signal and a noise signal; the addition unit adds the observed signal, which is an original sound including the target sound source signal and the noise element, to the emphasis signal, thereby reducing a ratio of artifact elements in the signal input to the speech recognition unit.

2. 2. The signal processing apparatus according to claim 1, wherein the observed signal is an audio signal recorded by a single microphone.

3. 3. The signal processing device according to claim 1, wherein the adding unit adjusts a weight of the observed signal to be added to the emphasis signal in accordance with a ratio of a noise signal contained in the observed signal.

4. 4. The signal processing device according to claim 3, wherein the adding unit weights only the observed signal to be added to the enhancement signal, or weights both the observed signal and the observed signal to be added to the enhancement signal such that a sum of the weight of the observed signal and the weight of the observed signal to be added to the enhancement signal is 1.

5. A method performed by a signal processing device, comprising: generating an emphasis signal in which the speaker's voice is emphasized from the observed signal by performing estimation processing; adding the observation signal, which is the original sound, to the enhancement signal; performing speech recognition on the enhancement signal to which the observation signal has been added in the adding step; Including, the enhancement signal includes an estimation error defined using an orthogonal projection in addition to the target sound source signal; the estimation error includes a noise component, which is a naturally observable signal, and an artifact component, which is an unnatural signal that rarely appears in training samples and cannot be represented by a linear combination of a speech signal and a noise signal; a signal processing method, characterized in that the adding step adds the observed signal, which is an original sound including the target sound source signal and the noise element, to the emphasized signal, thereby reducing a ratio of artifact elements in the emphasized signal on which the speech recognition is performed.

6. A step of generating an emphasis signal in which the speaker's voice is emphasized from the observed signal by performing estimation processing; adding the observed signal, which is the original sound, to the enhancement signal; performing speech recognition on the enhancement signal to which the observation signal has been added in the adding step; on the computer, the enhancement signal includes an estimation error defined using an orthogonal projection in addition to the target sound source signal; the estimation error includes a noise component, which is a naturally observable signal, and an artifact component, which is an unnatural signal that rarely appears in training samples and cannot be represented by a linear combination of a speech signal and a noise signal; a signal processing program for reducing a ratio of artifact elements in the enhancement signal for speech recognition by adding the observation signal, which is an original sound including the target sound source signal and the noise element, to the enhancement signal in the adding step;

Citation Information

Patent Citations

  • Noise reduction processing method / Device and program storage medium

    JP2000082999A

  • System and method for reducing noise by using single microphone

    JP2001092491A

  • Portable electronic apparatus

    JP2012093641A