Signal processing device, signal processing method, and signal processing program
The signal processing device addresses the issue of delayed adaptive filter coefficients by estimating signal ratios with delayed signals to control coefficient updates, achieving rapid low-noise, low-distortion output.
Patent Information
- Application Number
- JP2022002604
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-01-11
AI Technical Summary
Existing noise cancellers fail to accurately account for the delay caused by the acoustic impulse response of adaptive filters, leading to insufficient control of adaptive filter coefficients and prolonged times to achieve low residual noise and low signal distortion.
A signal processing device that estimates the ratio of amplitudes or powers of mixed signals using delayed signals to control adaptive filter coefficients, incorporating a delay unit to account for acoustic impulse response and adjust the coefficient update step size based on this ratio.
Enables rapid achievement of an output signal with minimal residual noise and low distortion by accurately controlling adaptive filter coefficients, considering the impact of acoustic delays.
Smart Images

Figure 0007772601000004 
Figure 0007772601000005 
Figure 0007772601000006
Abstract
Description
[Technical Field]
[0001] The present invention relates to a signal processing technique for eliminating noise, interference signals, echoes, and the like that are mixed in a signal. [Background technology]
[0002] Speech signals input from a microphone, handset, or the like are often superimposed with background noise, posing a significant problem in speech coding and speech recognition. Patent Documents 1 and 2 disclose a two-input noise canceller using two adaptive filters as a signal processing device designed to cancel acoustically superimposed noise. The device inputs a noisy signal (a signal containing a mixture of a desired signal and noise) and a reference signal (mainly including a signal correlated with the noise), cancels some or all of the noise, and outputs an enhanced signal (a signal in which the desired signal is enhanced). A step-size calculation unit calculates the coefficient update step size of the second adaptive filter using the signal-to-noise ratio of the noisy signal estimated using the first adaptive filter. The first adaptive filter operates in the same way as the second adaptive filter, but the coefficient update step size of the first adaptive filter is set to a larger value than the coefficient update step size of the second adaptive filter. Therefore, the output of the first adaptive filter has high adaptability to environmental changes but is inferior to the second adaptive filter in noise estimation accuracy.
[0003] The step size calculation unit evaluates the signal-to-noise ratio of the noisy signal estimated using the first adaptive filter, and when the speech signal is louder than the noise, it determines that interference from the speech signal is large and provides a small coefficient update step size to the second adaptive filter. Conversely, when the speech signal is quieter than the noise, it determines that interference from the speech signal is small and provides a large coefficient update step size to the second adaptive filter. In this way, by controlling the second adaptive filter with the coefficient update step size provided by the step size calculation unit, sufficient adaptability to environmental changes and low distortion in the signal after noise cancellation are output.
[0004] Patent Document 3 discloses a configuration in which the first adaptive filter is omitted from the configurations of Patent Documents 1 and 2. The signal-to-noise ratio is approximated by the ratio between the desired signal (such as speech) estimated using a second adaptive filter and the output of the second adaptive filter, and the second adaptive filter itself is controlled using a step size calculated based on the signal-to-noise ratio. Furthermore, Patent Document 3 discloses a noise canceller configuration that extends the configurations of Patent Documents 1 and 2, and also cancels speech signals mixed into noise when crosstalk due to speech signals is present, which is a significant influence of speech signals mixed into noise at the input of a two-noise input device. In addition to the configurations of Patent Documents 1 and 2, Patent Document 3 includes a third adaptive filter that cancels crosstalk from a reference signal. To accurately cancel noise from the speech signal input, a second step size calculation unit calculates a coefficient update step size and controls the third adaptive filter.
[0005] That is, the noise cancellers in Patent Documents 1 to 3 control updating of the adaptive filter coefficients using a signal-to-noise ratio estimated using the signal after noise cancellation and the adaptive filter output. By using a small step size when the signal-to-noise ratio is high and a large step size when the signal-to-noise ratio is low, they achieve both fast convergence and low-distortion output signals.
[0006] However, in the noise cancellers of Patent Documents 1 to 3, the adaptive filter coefficients are not updated at all. This is because the initial values of the adaptive filter coefficients are usually set to zero. An adaptive filter with zero coefficients outputs zero. Because this is the denominator of the estimated signal-to-noise ratio, the estimated signal-to-noise ratio becomes an extremely large value, and the corresponding step size is set to zero. A step size of zero means that no coefficient update is performed. To avoid this, the step size must be forcibly set to a non-zero value immediately after the coefficient update begins. However, no clear design method is disclosed regarding what value the step size should actually be set to or how long it should remain non-zero. In other words, manual control of the step size is required to achieve both fast convergence and low-distortion output signals in a two-input noise canceller.
[0007] Patent Document 4 discloses the configuration of a two-input noise canceller that does not require manual control of the step size and achieves both fast convergence and low-distortion output signals. A new signal-to-noise ratio, defined as the ratio of the post-noise-cancellation signal to the reference signal, is used immediately after the device is activated to control the update of the adaptive filter coefficients, solving the problem of non-coefficient update. When the adaptive filter coefficients grow, they are switched to the signal-to-noise ratio disclosed in Patent Document 3. The coefficient growth is evaluated when the new signal-to-noise ratio becomes sufficiently close to the conventional signal-to-noise ratio. When crosstalk is present, a third adaptive filter is introduced to cancel the speech signal from the noise input signal, and the step size is controlled using the same principle as the second adaptive filter. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Japanese Patent Application Publication No. 10-215193 [Patent Document 2] Japanese Patent Application Laid-Open No. 2000-172299 [Patent Document 3] International Publication No. 2012 / 046582 [Patent Document 4] International Publication No. 2019 / 092798 Summary of the Invention [Problem to be solved by the invention]
[0009] However, the signal processing device described in Patent Document 4 does not provide a sufficiently accurate signal-to-noise ratio, defined as the ratio between a signal after noise cancellation and a reference signal. This is because the effect of delay caused by the impulse response of the acoustic system that the adaptive filter approximates is not taken into consideration. As a result, the adaptive filter coefficient update cannot be appropriately controlled, and an output signal with low residual noise and low signal distortion cannot be obtained in a short time.
[0010] An object of the present invention is to provide a technique for solving the above problems. [Means for solving the problem]
[0011] In order to achieve the above object, the device according to the present invention comprises: a first input means for inputting a first mixed signal in which the first signal and the second signal are mixed; a second input means for inputting a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed; a first adaptive filter that filters the second mixture signal to generate an estimate of the second signal; a first subtraction unit that generates an estimate of the first signal from the first mixed signal and an estimate of the second signal; a first delay unit that delays the second mixture signal using a coefficient of the first adaptive filter to generate a delayed second mixture signal; an estimation unit that estimates a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; Equipped with The first mixing ratio is used to control the first adaptive filter.
[0012] In order to achieve the above object, the method according to the present invention comprises: Input the first mixed signal, which is a mixture of the first and second signals, a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed is input; processing the second mixed signal with a first adaptive filter to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixture signal using a coefficient of the first adaptive filter to generate a delayed second mixture signal; estimating an amplitude or power ratio between the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; The first mixing ratio is used to control generation of an estimate of the second signal.
[0013] In order to achieve the above object, the program according to the present invention comprises: On the computer, inputting a first mixed signal in which a first signal and a second signal are mixed; inputting a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed; processing the second mixed signal with a first adaptive filter to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixture signal using a coefficient of the first adaptive filter to generate a delayed second mixture signal; a step of estimating a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; using the first mixing ratio to control generation of an estimate of the second signal; Execute the following. [Effects of the Invention]
[0014] According to the present invention, it is possible to obtain a signal processing device that can take into account the effect of delay caused by the acoustic impulse response that approximates the adaptive filter, appropriately control the update of the adaptive filter coefficients, and obtain an output signal with little residual noise and little signal distortion in a short time. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a block diagram showing a configuration of a signal processing device according to a first embodiment of the present invention. [Figure 2] FIG. 10 is a block diagram showing the configuration of a signal processing device according to a second embodiment of the present invention. [Figure 3] FIG. 10 is a block diagram showing a first configuration of an estimation unit according to a second embodiment of the present invention. [Figure 4] FIG. 10 is a diagram showing time transitions of values according to the second embodiment of the present invention. [Figure 5] FIG. 10 is a block diagram showing a second configuration of an estimation unit according to a second embodiment of the present invention. [Figure 6] FIG. 10 is a block diagram showing a configuration of a signal processing device according to a third embodiment of the present invention. [Figure 7] FIG. 10 is a block diagram showing a first configuration of an estimation unit according to a third embodiment of the present invention. [Figure 8] FIG. 11 is a block diagram showing a second configuration of an estimation unit according to the third embodiment of the present invention. [Figure 9] 1 is a block diagram showing the configuration of a computer according to a first embodiment of the present invention. [Figure 10] 10 is a flowchart showing an example of signal processing by a processor of the computer shown in FIG. 9. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the components described in the following embodiments are merely examples and are not intended to limit the technical scope of the present invention.
[0017] 1. First embodiment A signal processing device 100 according to a first embodiment of the present invention will be described with reference to Fig. 1. The signal processing device 100 in Fig. 1 is a device that obtains an estimated value e1(k) of a first signal from a first mixture signal xP(k) in which a first signal and a second signal are mixed.
[0018] As shown in FIG. 1, the signal processing device 100 includes a first input unit 101, a second input unit 102, an adaptive filter 103, a subtraction unit 104, an estimation unit 106, a coefficient update control unit 107, and a delay unit 110.
[0019] Of these, the first input unit 101 receives a first mixed signal xP(k) in which the first signal and the second signal are mixed. The second input unit 102 receives a second mixed signal xR(k) in which the third signal and the fourth signal are mixed. The first signal and the third signal originate from the same signal source A and are correlated with each other. The second signal and the fourth signal originate from the same signal source B and are correlated with each other.
[0020] The subtractor 104 receives the first mixture signal xP(k) and an estimate n1(k) of the second signal mixed in the first mixture signal xP(k), and outputs an estimate e1(k) of the first signal. Then, the adaptive filter 103 performs filtering on the second mixture signal xR(k) using a coefficient 141 that is updated based on the estimate e1(k) of the first signal, in order to obtain the estimate n1(k) of the second signal.
[0021] The delay unit 110 delays the second mixture signal xR(k) using the coefficient vector w1(k) of the adaptive filter 103 to obtain a delayed second mixture signal xRD(k).
[0022] The estimation unit 106 estimates the ratio of the amplitudes or powers of the first and second signals as a first mixture ratio R1(k) using the estimated value e1(k) of the first signal, the estimated value n1(k) of the second signal, the delayed second mixture signal xRD(k), and the coefficient vector w1(k) of the adaptive filter 103. When the value of the first mixture ratio R1(k) obtained by the estimation unit 106 is large, the coefficient update control unit 107 outputs a control signal μ1(k) to the adaptive filter 103 to reduce the amount of update of the coefficients 141 of the adaptive filter 103.
[0023] According to this embodiment having such a configuration, the first mixture ratio R1(k) is estimated using a delayed second mixture signal xRD(k) obtained by delaying the second mixture signal xR(k), and the coefficient update of the adaptive filter 103 is controlled. As a result, it is possible to take into consideration the influence of delay caused by the acoustic impulse response that approximates the adaptive filter 103, and it is possible to obtain an estimated value e1(k) of the first signal with less residual second signal and less distortion in a short time.
[0024] 2. Second Embodiment As a signal processing device according to a second embodiment of the present invention, a noise canceller will be described which receives a noisy signal (a signal obtained by mixing a desired signal and noise) and a reference signal (mainly including a signal correlated with noise), cancels part or all of the noise, and outputs an emphasized signal (a signal obtained by emphasizing the desired signal). Here, the noisy signal corresponds to a first mixed signal obtained by mixing a first signal and a second signal, the reference signal corresponds to a second mixed signal obtained by mixing a third signal and a fourth signal, and the emphasized signal (an estimated value of the first signal) corresponds to the desired signal.
[0025] [2.1. Basic noise cancellation technology] The following is a brief explanation of the basic technology of noise cancellation, which uses an adaptive filter to cancel noise, interference signals, echoes, etc. mixed in with a desired signal input from a microphone, handset, communication channel, etc., or to emphasize the desired signal.
[0026] As disclosed in Patent Documents 1 to 3, a two-input noise canceller generates, from a reference signal, a pseudo-noise (an estimated value of a second signal) corresponding to a noise component mixed into speech at the speech input terminal, using an adaptive filter that approximates the impulse response of the acoustic path from the noise source to the speech input terminal. The noise component is then suppressed by subtracting this pseudo-noise from a signal (first mixed signal) input to the speech input terminal. Here, the mixed signal refers to a signal in which a desired (speech) signal and noise are mixed, and is generally supplied to the speech input terminal from a microphone or handset. The reference signal is a signal correlated with the noise component in the noise source and is captured near the noise source. By capturing the reference signal near the noise source in this way, the reference signal can be considered to be approximately equal to the noise component in the noise source. The adaptive filter receives the reference signal supplied to the reference input terminal.
[0027] The coefficients of the adaptive filter are corrected by calculating the correlation between the error obtained by subtracting pseudo-noise from the noisy signal and the reference signal input to the reference input terminal. As coefficient correction algorithms for such adaptive filters, Patent Documents 1 to 3 disclose the "Least Mean-Square Algorithm (LMS)" and the "Learning Identification Method (LIM)." The LIM is also called the Normalized LMS (NLMS) algorithm.
[0028] Coefficient update using the normalized LMS algorithm is expressed by equation (1), where μ1(k) is the step size at time k. The coefficient vector w1(k) is defined by equation (2) using its elements. The reference signal vector xR(k) is expressed by equation (3) using its elements. Note that T represents the transpose of the vector, and the total number of coefficients is L.
number
[0029] The LMS algorithm and LIM are types of algorithms known as gradient methods, and the speed and accuracy of coefficient updates depend on a constant called the coefficient update step size. Filter coefficients are updated by multiplying the coefficient update step size by an error. To reduce interference with the coefficient update caused by the desired signal (the estimated value of the first signal) contained in the error, the coefficient update step size must be set to an extremely small value or zero. Constantly setting the coefficient update step size to a small value reduces the adaptive filter coefficients' ability to track environmental changes. Patent Documents 1 to 3 above disclose a method for solving the problem of increased output error or distortion of the desired signal. Furthermore, because the desired signal is generally speech, the term "speech" will be used hereinafter. However, the scope of the present invention is not limited to speech and encompasses any type of signal, including acoustic (audio) signals. Furthermore, the coefficient update algorithm is not limited to LMS or LIM.
[0030] [2.2. Configuration of noise canceller] 2 is a block diagram showing the overall configuration of a noise cancellation device 200 according to this embodiment. The noise cancellation device 200 may function as part of a device such as a digital camera, a laptop computer, a mobile phone, a hearing aid, a television, a smart speaker, or a robot, but the present invention is not limited to these and may be applied to any signal processing device that is required to cancel noise from an input signal.
[0031] As shown in FIG. 2, noise cancellation device 200 receives a noisy signal (first mixture signal) xP(k) containing a mixture of speech (first signal) and noise (second signal) from input terminal 201. It also receives a reference signal (second mixture signal) xR(k) containing a mixture of speech and noise from input terminal 202, and outputs an estimated value e1(k) of speech (enhanced signal) from output terminal 205. Adaptive filter 203 includes adaptive filter 103 and coefficient update control unit 107 shown in FIG. 1. Adaptive filter 203 receives first mixture ratio R1(k), calculates step size μ1(k), and updates coefficient vector w1(k) using the calculated step size μ1(k). The noise canceller 200 generates pseudo-noise n1(k) by modifying a reference signal xR(k) that is correlated with the noise to be cancelled using an adaptive filter 203, and then subtracts this from a noisy signal xP(k), which is a speech signal with noise superimposed thereon, thereby canceling the noise.
[0032] A noisy signal xP(k) is supplied to an input terminal 201 as a sequence of sample values. The noisy signal xP(k) is transmitted to a subtraction unit 204. A reference signal xR(k) is supplied to an input terminal 202 as a sequence of sample values. The reference signal xR(k) is transmitted to an adaptive filter 203 and an estimation unit 206.
[0033] The adaptive filter 203 performs a convolution operation on the reference signal xR(k) and the filter coefficients, and transmits the result as pseudo noise n1(k) to the subtraction unit 204 and the estimation unit 206. The adaptive filter 203 also supplies the coefficient vector w1(k) to the estimation unit 206.
[0034] The subtraction unit 204 receives the noisy signal xP(k) from the input terminal 201 and the pseudo-noise n1(k) from the adaptive filter 203. The subtraction unit 204 subtracts the pseudo-noise n1(k) from the noisy signal xP(k) and transmits the result as an estimated value e1(k) of the speech signal (estimated value of the first signal) to the output terminal 205, while simultaneously feeding it back to the adaptive filter 203.
[0035] Delay unit 210 receives reference signal xR(k) and coefficient vector w1(k) of adaptive filter 203, and delays reference signal xR(k) by d1(k) samples based on coefficient vector w1(k) of adaptive filter 203 to obtain delayed reference signal xRD(k). That is, xRD(k)=xR(k-d1(k)). d1(k) can be obtained as follows.
[0036] The natural number d1(k) is the delay in the impulse response of the acoustic system identified by the adaptive filter 103 and corresponds to the element w1(k, d1(k)) for which the absolute value of the coefficient vector w1(k) of the adaptive filter 203 is maximum. That is, d1(k) is the number (= position relative to the beginning) of the element with the largest absolute value among the elements of the coefficient vector w1(k). Various methods are known for finding the maximum value, including bubble sort. For example, in bubble sort, all elements are compared with a temporary maximum value determined initially, and the larger of the comparison results is replaced with the temporary maximum value each time. The last remaining value is the true maximum value. If the temporary maximum value and the element number are recorded simultaneously, the number of the last remaining maximum value becomes the desired d1(k). The initial value of d1(k), i.e., d1(0), can be set as d1(0) = 0.
[0037] In bubble sort, the comparison of the temporary maximum value with all elements is usually completed within one sampling period to determine d1(k). However, it is also possible to perform only a portion of the comparison, such as once or twice, within one sampling period, resulting in a single d1(k) being determined over multiple sampling periods. This reduces the number of comparison operations performed within one sampling period, leading to a reduction in the total amount of operations and power consumption. Conversely, because it takes time to update d1(k), the adaptive filter 203's ability to track fluctuations in the impulse response of the acoustic system it approximates deteriorates.
[0038] The elements of the coefficient vector w1(k) can be divided into multiple groups, and the sum of absolute values or sum of squares of the elements in each group can be used as coefficient information. The position corresponding to the group with the largest coefficient information can be used as d1(k). The position corresponding to a group can be the position of the first element, the position of the middle element, or the position of the last element in the group. Since the number of groups is smaller than the total number of elements, the number of comparison operations performed within one sampling period can be reduced, leading to a reduction in the total amount of operations and power consumption. In addition, the averaging effect of the elements within each group can absorb small fluctuations in element values caused by coefficient updates, resulting in a stable delay estimation result.
[0039] The elements of the coefficient vector w1(k) may be divided into multiple groups, and the position corresponding to the group with the largest coefficient information in each group may be set as d1(k). In this case, the number of comparisons of coefficient information in each group may be reduced and performed over multiple sampling periods. This further reduces the number of comparison operations performed within one sampling period, leading to a further reduction in the total amount of operations and power consumption.
[0040] The simplest setting for d1(k) is to make d1(k) a constant. A single predefined constant may be used, or multiple predefined constants may be used in sequence. For example, natural numbers M1 to M5 where M1 < M2 < … < M5 =< L and arbitrary natural numbers T1, T2…, T5 are predefined in advance. At the start of the operation of the noise canceling device 200, d1(k) is set to 0, and when the coefficient update is completed T1 times, d1(T1 + 1) = M1 is set. Subsequently, when the coefficient update is completed T2 times with d1(k) = M1 (T1 < k =< T1 + T2), d1(T2 + 1) = M2 is set. Thereafter, the same process can be repeated to sequentially set M3 to M5 as d1(k). When simulation is performed at the design stage of the device and the convergence characteristics of the coefficient vector w1(k) are known, setting d1(k) to a constant is particularly effective. The comparison operation for determining d1(k) becomes unnecessary, leading to a reduction in the total amount of computation and power consumption. At this time, the delay unit 210 receives the reference signal xR(k) and delays it by d1(k) samples to obtain the delayed reference signal xRD(k), and does not require the coefficient vector w1(k) of the adaptive filter 203.
[0041] The estimator 206 receives the estimated value e1(k) of the voice, the pseudo-noise n1(k) (the output of the adaptive filter 203), the delayed reference signal xRD(k), and the coefficient vector w1(k) of the adaptive filter 203, estimates the ratio of the amplitude or power of the voice and noise at the input terminal 201 as the first mixing ratio R1(k), and transmits it to the adaptive filter 203. The adaptive filter 203 updates the coefficient vector w1(k) using a small step size μ1(k) when the first mixing ratio R1(k) is large and a large step size μ1(k) when the first mixing ratio R1(k) is small. Methods for controlling the step size using the first mixing ratio R1(k), that is, the estimated value of the signal-to-noise ratio, are disclosed in detail in Patent Documents 1 to 3. Also, as disclosed in Patent Documents 1 to 3, the first mixing ratio R1(k) may be averaged and then used in the calculation of the step size μ1(k). The estimation accuracy for the ratio of the amplitude or power of the voice and noise is improved.
[0042] 2.3. First Configuration of Estimation Unit 206 FIG. 3 is a block diagram showing a first internal configuration of estimation unit 206. Estimation unit 206 includes signal ratio estimation unit 301, signal ratio estimation unit 302, and mixer 305. Signal ratio estimation unit 301 receives speech estimate e1(k) and pseudo-noise n1(k) and estimates the amplitude or power ratio of speech to noise at input terminal 201 as second mixture ratio R2(k). Second mixture ratio R2(k) may be the ratio of the amplitude or power of speech estimate e1(k) and pseudo-noise n1(k), or may be calculated by adding a small constant to the amplitude or power of the speech estimate e1(k) and pseudo-noise n1(k). Adding a small constant has the effect of preventing divergence of the quotient due to division. Alternatively, one or both of speech estimate e1(k) and pseudo-noise n1(k) may be averaged before use. Averaging can improve the accuracy of the ratio calculation.
[0043] Signal ratio estimation section 302 receives speech estimate e1(k) and delayed reference signal xRD(k), and estimates the ratio of the amplitude or power of speech to noise at input terminal 201 as third mixture ratio R3(k). Third mixture ratio R3(k) may be the ratio of the amplitude or power of speech estimate e1(k) to delayed reference signal xRD(k), or may be calculated by adding a small constant to these amplitudes or powers. Adding a small constant has the effect of preventing divergence of the quotient due to division. Alternatively, either or both of speech estimate e1(k) and delayed reference signal xRD(k) may be averaged before use.
[0044] Mixing unit 305 mixes second mixture ratio R2(k) and third mixture ratio R3(k) and outputs the mixture result as first mixture ratio R1(k). Second mixture ratio R2(k) and third mixture ratio R3(k) may be mixed by weighted addition, or may be mixed using a more complex higher-order polynomial. Prior to mixing, one or both of second mixture ratio R2(k) and third mixture ratio R3(k) may be averaged. Averaging can improve the calculation accuracy of first mixture ratio R1(k), i.e., the approximation accuracy of the amplitude or power of speech and noise at input terminal 201.
[0045] For simplicity, consider the case where the first mixture ratio R1(k) is calculated by mixing the second mixture ratio R2(k) and the third mixture ratio R3(k) using weighted addition. The weights for both are set to equal 1. The coefficients of the adaptive filter 203 are generally initialized to zero. Therefore, when coefficient updating begins, the pseudo-noise n1(k) is zero, and the second mixture ratio R2(k) has a zero denominator and is infinite. Therefore, when the step size μ1(k) of the adaptive filter 203 is calculated using the second mixture ratio R2(k), it becomes an extremely small value or zero, and the coefficient does not grow. If the coefficient does not grow, the pseudo-noise n1(k) does not increase, and the same problem persists.
[0046] On the other hand, the denominator of the third mixture ratio R3(k) is the delayed reference signal xRD(k), which is not necessarily zero at the start of coefficient updating. This is because the microphone input contains minute signals such as environmental noise. Even if the delayed reference signal xRD(k) is zero, it will not remain at zero. Therefore, the third mixture ratio R3(k) will not become infinite, and the corresponding step size μ1(k) will not become a minimum value. Therefore, the coefficients of the adaptive filter 203 grow as the coefficients are updated and converge to a value that represents the acoustic characteristics of the path from the noise signal source to the input terminal 201. When the delayed reference signal xRD(k) is zero, the reference signal xR(k) will also be zero, and the coefficients of the adaptive filter 203 will not be updated. Therefore, even if the third mixture ratio R3(k) is an extremely large value, it does not cause any problems. However, when the coefficients of adaptive filter 203 grow to a certain extent and the pseudo-noise n1(k) grows, the third mixture ratio R3(k) has a lower approximation accuracy to the ratio of the amplitude or power of speech to noise at input terminal 201 than the second mixture ratio R2(k).
[0047] Therefore, mixer 305 sets the weight of third mixture ratio R3(k) to a large value when coefficient update of adaptive filter 203 starts, and decreases it as the coefficient grows. The weight of second mixture ratio R2(k) is set to a small value when coefficient update of adaptive filter 203 starts, and increases it over time. This means that the proportion of third mixture ratio R3(k) in first mixture ratio R1(k) decreases in accordance with the number of coefficient updates.
[0048] For example, if the weight of the third mixture ratio R3(k) is set to 1 when the coefficient update of adaptive filter 203 begins, the weight of the second mixture ratio R2(k) becomes 0 because the sum of the weights is 1. The growth of the coefficients corresponds to the number of coefficient updates. Therefore, the weight of the third mixture ratio R3(k) is set to 1 when the coefficient update of adaptive filter 203 begins, and the weight decreases toward 0 corresponding to the number of coefficient updates of coefficient vector w1(k). Correspondingly, the weight of the second mixture ratio R2(k) increases from 0 to 1. If the initial value of the weight of the third mixture ratio R3(k) is set to 1, and the weight becomes 1 at some point or is set to 1, the third mixture ratio R3(k) is switched to the second mixture ratio R2(k). Similar effects can be obtained even if the sum of the weights is set to a value other than 1 or if the initial value of the weight for the third mixture ratio R3(k) is set to a value other than 1. The weight can be determined using the amount of change in the coefficient vector w1(k) of the adaptive filter 203.
[0049] FIG. 4 is a diagram schematically illustrating the time transition of a signal according to the second embodiment, showing the delayed reference signal xRD(k), the pseudo-noise n1(k) output by the adaptive filter 203, and the coefficient vector w1(k). The vertical axis in FIG. 4 represents the signal value in terms of power, and the coefficient is the norm of the coefficient vector. The horizontal axis represents the number of coefficient updates of the adaptive filter 203, expressed as the number of samples k. Comparing xRD(k) and n1(k) is equivalent to comparing the third mixture ratio R3(k) and the second mixture ratio R2(k), which have the same numerator and the denominators xRD(k) and n1(k), respectively. Because xRD(k) is unrelated to the coefficient updates, n1(k) determines the relationship between the third mixture ratio R3(k) and the second mixture ratio R2(k). n1(k) is determined by xRD(k) and the coefficient vector w1(k), and its amplitude or power increases with the coefficient update, excluding the change due to the increase or decrease of xRD(k). That is, xRD(k) is constant, and n1(k) increases with the increase of the coefficient vector w1(k) due to the coefficient update.
[0050] The relationship between the delayed reference signal xRD(k), the pseudo-noise n1(k) output by the adaptive filter 203, and the coefficient vector w1(k) described above is clear from FIG. 4. For simplicity, in FIG. 4, xRD(k), which is unrelated to the coefficient update, is shown as a constant. n1(k), which depends on xRD(k), increases smoothly, and when the coefficient vector w1(k) converges, n1(k) also saturates. Because the amount of change in the coefficient vector w1(k) of the adaptive filter 203 decreases with coefficient update, the amount of change in the coefficient vector w1(k) can be used as an indicator of how much n1(k) has grown from zero. That is, the mixer 305 determines the weight of the third mixing ratio R3(k) based on the time change in the coefficient vector w1(k).
[0051] By using xRD(k), which is obtained by delaying xR(k) by d1(k), it is possible to reduce the difference between the power of n1(k) and the power of xR(k), i.e., the time lag between them (the lag along the horizontal axis in FIG. 4). This means that the time lag between the increase or decrease in the power of n1(k) and the increase or decrease in the power of xRD(k) is smaller than the time lag between the increase or decrease in the power of n1(k) and the increase or decrease in the power of xR(k). This is particularly significant when xR(k) is a non-stationary signal, different from white noise. In other words, xRD(k) can approximate n1(k) with higher accuracy than xR(k). This leads to more accurate step size control through the high-accuracy third mixing ratio R3(k), which is desirable for controlling adaptive filter 203. Therefore, it is possible to obtain an estimated speech value with small residual noise and low distortion in a short time.
[0052] The time change of the coefficient vector w1(k) used to estimate the first mixture ratio R1(k) may be the time change of the sum of squares or the sum of absolute values of the elements of the coefficient vector w1(k), or the time change of the partial sum of squares or the partial sum of absolute values. When using the partial sum, the elements of the coefficient vector may be thinned out, or a portion of the coefficient vector may be cut out. By using the partial sum, it is possible to reduce the amount of calculation required for the time change of the coefficient while suppressing a decrease in the evaluation accuracy of the time change. Furthermore, the coefficient changes smoothly because it does not depend on the input signal. Therefore, it has smaller fluctuations than other indices that depend on the input signal, and can accurately detect a saturation state.
[0053] Equation (4) shows an example of mixing the second mixture ratio R2(k) and the third mixture ratio R3(k) by weighted addition. ζ1(k) expressed in equation (5) satisfies ζ1(0) = 0 and increases as the coefficients are updated. δ1(k), which corresponds to the time change in the coefficient vector w1(k) of the adaptive filter 203, can be calculated using equation (6). M is a predetermined natural number that determines the frequency at which the time change in the coefficient vector w1(k) is calculated. The larger M, the smaller the fluctuation in δ1(k) and the more accurate the detection of coefficient saturation. However, the larger the delay in the change in δ1(k), the longer the detection of coefficient saturation. Equation (5) represents an example in which, when the value of the time change δ1(k) becomes less than the threshold ε1, the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k) is set to 0, and the first mixture ratio R1(k) is switched from the third mixture ratio R3(k) to the second mixture ratio R2(k). This is an example of a method for determining whether the value of the time change δ1(k) has become sufficiently small. If ζ1(k) = δ1(k) / |δ1(M)| is used instead of equation (5), the content ratio of the third mixture ratio R3(k) is determined by the time change of the coefficient vector w1(k).
number
[0054] In this way, the mixer 305 can set the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k) to either 100% or 0% based on the comparison result between the time change δ1(k) and the threshold ε1. Alternatively, the mixer 305 may repeatedly compare the time change δ1(k) of the coefficient vector w1(k) with the threshold ε1, and set the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k) to 100% if the time change δ1(k) of the coefficient vector w1(k) is equal to or greater than the threshold ε1. Otherwise, the mixer 305 may repeatedly set the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k). This allows the first mixture ratio R1(k) to have high resilience to environmental changes, enabling more accurate coefficient update even when, for example, the amount of update of the coefficient vector w1(k) of the adaptive filter 203 increases rapidly.
[0055] The mixer 305 may set the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k) to 0% if the time change δ1(k) of the coefficient vector w1(k) is less than the threshold ε1 for a predetermined first constant L1 or more consecutive times. Otherwise, the mixer 305 may set the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k) to 100%. After the content ratio of the third mixture ratio R3(k) is set to 0% once, the comparison of δ1(k) and ε1 may be stopped by setting the threshold ε1 to a sufficiently large value. Here, L1 is an integer equal to or greater than 2.
[0056] The mixing unit 305 may set the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k) to 0% when δ1(k) is less than the threshold ε1 for a predetermined third constant L3 or more consecutive times in a section where the time change δ1(k) of the coefficient vector w1(k) has a number of samples equal to the predetermined second constant L2. Otherwise, the content ratio of the third mixture ratio R3(k) in the first mixture ratio R1(k) may be set to 100%. The second constant L2 and the third constant L3 are both integers greater than or equal to 2, satisfying L2 > L3. Furthermore, the continuous condition may be removed from the above evaluation, and the number of times δ1(k) is less than the threshold ε1 for a predetermined third constant L3 or more may be evaluated. After the content ratio of the third mixture ratio R3(k) is set to 0% once, the evaluation of δ1(k) may be stopped by setting L3 to a value greater than L2.
[0057] Note that mixer 305 may set the minimum value of the weights of second mixture ratio R2(k) and third mixture ratio R3(k) to a value greater than 0 rather than 0.
[0058] 2.4. Second Configuration of Estimation Unit 206 FIG. 5 is a block diagram showing a second internal configuration of the estimation unit 206. The estimation unit 206 includes a mixer 506 and a signal ratio estimator 503. The mixer 506 mixes a delayed reference signal xRD(k) (delayed second mixed signal) with a pseudo-noise n1(k) (estimated value of the second signal) based on a coefficient vector w1(k) to generate a first mixed signal n3(k). The signal ratio estimator 503 receives the speech estimate e1(k) and the first mixed signal n3(k) and estimates the ratio of the amplitude or power of the speech to the noise at the input terminal 201 as a first mixed ratio R1(k). The first mixed ratio R1(k) may be the ratio of the amplitude or power of the speech estimate e1(k) and the first mixed signal n3(k), or may be calculated by adding a small constant to the amplitude or power of the speech estimate e1(k) and the first mixed signal n3(k). Adding a small constant has the effect of preventing the quotient from diverging due to division. Alternatively, either or both of the estimated value e1(k) and the first mixed signal n3(k) may be averaged before use, which can improve the accuracy of calculating the ratio.
[0059] The second internal configuration of estimation unit 206 shown in Fig. 5 is equivalent to the first internal configuration of estimation unit 206 shown in Fig. 3. That is, in the first internal configuration shown in Fig. 3, two estimates for the amplitude or power ratio of speech to noise at input terminal 201 are generated in signal ratio estimation units 301 and 302, and these estimates are mixed to calculate first mixture ratio R1(k). The second internal configuration of estimation unit 206 shown in Fig. 5 mixes two types of noise estimates, i.e., delayed reference signal xRD(k) and pseudo-noise n1(k), to generate first mixed signal n3(k), determine the denominator, and calculate first mixture ratio R1(k) by combining this with the numerator, speech estimate e1(k). These two types of configurations are possible because the same numerator, i.e., speech estimate e1(k), is used when estimating the amplitude or power ratio of speech to noise at input terminal 201 in the first internal configuration shown in Fig. 3 and the second internal configuration shown in Fig. 5. The second internal configuration of the estimation unit 206 shown in FIG. 5 has fewer components and is simpler than the first internal configuration shown in FIG.
[0060] The operation of mixer 506 is the same as that of mixer 305, except that the input signals, second mixing ratio R2(k) and third mixing ratio R3(k), are replaced with pseudo-noise n1(k) and delayed reference signal xRD(k) (delayed second mixing signal), respectively, and therefore a detailed description thereof will be omitted.
[0061] With the above configuration, this embodiment can take into account the effect of delay caused by the acoustic impulse response that approximates the adaptive filter 203, and appropriately control the update of the coefficients of the adaptive filter 203, thereby making it possible to obtain an output signal with little residual noise and little signal distortion in a short time.
[0062] 3. Third Embodiment In the explanation so far, it has been assumed that the reference signal is noise itself, by capturing the reference signal near the noise source. However, in reality, there are cases where this condition cannot be met. In such cases, the reference signal is composed of noise and an audio signal mixed in with it. Such audio signal components mixed into the reference signal are called crosstalk. Patent Document 3 discloses the configuration of a noise cancellation device for use when crosstalk exists.
[0063] In this embodiment, a second adaptive filter is introduced to cancel crosstalk, similar to the noise cancellation. The second adaptive filter, which approximates the impulse response of the acoustic path (crosstalk path) from the audio signal source to the reference input terminal, is used to generate pseudo crosstalk corresponding to the audio signal components mixed in at the reference input terminal. Then, by subtracting this pseudo crosstalk from the signal (reference signal) input to the reference input terminal, the audio signal components (crosstalk) mixed in the reference input are canceled.
[0064] A noise cancellation device 800 according to a third embodiment of the present invention will be described with reference to Fig. 6. Compared to the second embodiment, the noise cancellation device according to this embodiment includes a subtraction unit 804, an adaptive filter 803, and a delay unit 810 in addition to the subtraction unit 204 and the adaptive filter 203, and the estimation unit 206 is replaced with an estimation unit 806. The other configurations and operations are the same as those of the second embodiment, so the same configurations are denoted by the same reference numerals and detailed description will be omitted.
[0065] Noise cancellation device 800 generates pseudo crosstalk n2(k) (an estimated value of the third signal) by modifying a signal (output at output terminal 205 = estimated speech signal or emphasis signal) correlated with the crosstalk to be canceled using adaptive filter 803. Then, this is subtracted from a reference signal xR(k) containing a mixture of speech and noise, thereby canceling the crosstalk. When updating the coefficients of adaptive filter 803, the step size is controlled using a fourth mixture ratio R4(k) that approximates the ratio of the amplitude or power of the fourth signal to the third signal. This allows smooth coefficient updating, and as a result, an output signal with low residual noise and little signal distortion can be obtained in a short time.
[0066] The noisy signal xP(k) is supplied to the input terminal 201 as a sample value sequence and transmitted to a subtraction unit 204. The reference signal xR(k) is supplied to the input terminal 202 as a sample value sequence and transmitted to a subtraction unit 804.
[0067] The subtraction unit 804 receives the reference signal xR(k) from the input terminal 202 and the pseudo crosstalk n2(k) from the adaptive filter 803. The subtraction unit 804 subtracts the pseudo crosstalk n2(k) from the reference signal xR(k) and transmits the result as a noise estimate e2(k) (estimate of the fourth signal) to the output terminal 805, while simultaneously feeding it back to the adaptive filter 803. The subtraction unit 804 also supplies the noise estimate e2(k) to the estimation unit 806.
[0068] The adaptive filter 803 performs a convolution operation on the speech estimate e1(k) (emphasis signal) and the filter coefficients, and transmits the result as pseudo crosstalk n2(k) (estimate of the third signal) to the subtraction unit 804 and the estimation unit 806. The adaptive filter 803 also supplies the coefficient vector w2(k) to the estimation unit 806.
[0069] Delay unit 810 receives speech estimate e1(k) and coefficient vector w2(k) of adaptive filter 803, and delays speech estimate e1(k) by d2(k) samples based on coefficient vector w1(k) of adaptive filter 203 to obtain delayed speech estimate e1D(k). That is, e1D(k)=e1(k-d2(k)). d2(k) can be obtained in the same way as d1(k) already described.
[0070] Estimation unit 806 receives speech estimate e1(k), pseudo-noise n1(k) (output of adaptive filter 203), noise delay estimate e2D(k) (delay estimate of the fourth signal), and coefficient vector w1(k) of adaptive filter 203, estimates the ratio of the amplitude or power of speech to noise at input terminal 201 as first mixture ratio R1(k), and transmits this to adaptive filter 203. Adaptive filter 203 updates the coefficients using a small step size μ1(k) when first mixture ratio R1(k) is large and a large step size μ1(k) when first mixture ratio R1(k) is small. Note that adaptive filter 203 shown in FIG. 6 differs from adaptive filter 203 shown in FIG. 2 in that noise estimate e2(k) is input instead of reference signal xR(k). Methods for controlling the step size using the first mixture ratio R1(k), i.e., an estimated value of the signal-to-noise ratio, are disclosed in detail in Patent Documents 1 to 3. As disclosed in Patent Documents 1 to 3, the first mixture ratio R1(k) may be averaged and then used to calculate the step size μ1(k). This improves the estimation accuracy of the ratio of the amplitude or power of speech to noise at input terminal 201.
[0071] Estimation unit 806 further receives noise estimate e2(k), pseudo crosstalk n2(k) (output of adaptive filter 803), speech delay estimate e1D(k), and coefficient vector w2(k) of adaptive filter 803, estimates the ratio of the amplitudes or powers of the noise and crosstalk (fourth signal and third signal) at input terminal 202 as fourth mixing ratio R4(k), and transmits this to adaptive filter 803. The size of coefficient vector w2(k) shown in equation (7) may be equal to or different from size L of coefficient vector w1(k).
number
[0072] Adaptive filter 803 updates the coefficients using a small step size μ2(k) when fourth mixing ratio R4(k) is large and a large step size μ2(k) when fourth mixing ratio R4(k) is small. Methods for controlling the step size μ2(k) using the estimated value of fourth mixing ratio R4(k), i.e., the noise-to-crosstalk ratio at input terminal 202, are disclosed in detail in Patent Documents 1 to 3. Furthermore, as disclosed in detail in Patent Documents 1 to 3, fourth mixing ratio R4(k) may be averaged before being used to calculate step size μ2(k). This improves the estimation accuracy of the ratio of the amplitude or power of noise to crosstalk at input terminal 202.
[0073] 3.1. First Configuration of Estimation Unit 806 Fig. 7 is a block diagram showing a first internal configuration of estimation section 806. In addition to the configuration of estimation section 206, estimation section 806 includes a signal ratio estimation section 901, a signal ratio estimation section 902, and a mixer 905. Signal ratio estimation section 302 shown in Fig. 7 is the same as signal ratio estimation section 302 shown in Fig. 3 except that it uses a delayed noise estimate value e2D(k) instead of delayed reference signal xRD(k).
[0074] The signal ratio estimator 901 receives the noise estimate e2(k) and the pseudo crosstalk n2(k) and estimates the ratio of the amplitude or power of the noise to the crosstalk as a fifth mixture ratio R5(k). The fifth mixture ratio R5(k) may be the ratio of the amplitude or power of the noise estimate e2(k) to the pseudo crosstalk n2(k), or may be calculated by adding a small constant to the amplitude or power. Adding a small constant has the effect of preventing divergence of the quotient due to division. Alternatively, one or both of the noise estimate e2(k) and the pseudo crosstalk n2(k) may be averaged before use. Averaging can improve the accuracy of the ratio calculation.
[0075] Signal ratio estimator 902 receives noise estimate e2(k) and speech delay estimate e1D(k) (first signal delay estimate) and estimates the ratio of the amplitude or power of the noise to the crosstalk at input terminal 202 as sixth mixture ratio R6(k). Sixth mixture ratio R6(k) may be the ratio of the amplitude or power of noise estimate e2(k) and speech delay estimate e1D(k), or may be calculated by adding a small constant to these amplitudes or powers. Adding a small constant has the effect of preventing divergence of the quotient due to division. Alternatively, either or both of noise estimate e2(k) and speech delay estimate e1D(k) may be averaged before use.
[0076] Mixing unit 905 mixes fifth mixture ratio R5(k) and sixth mixture ratio R6(k) using coefficient vector w2(k) of adaptive filter 803 and outputs the mixed result as fourth mixture ratio R4(k). The fifth mixture ratio R5(k) and sixth mixture ratio R6(k) may be mixed by weighted addition using coefficient vector w2(k) of adaptive filter 803, or may be mixed using a more complex higher-order polynomial. Prior to mixing, either or both of fifth mixture ratio R5(k) and sixth mixture ratio R6(k) may be averaged. Averaging can improve the calculation accuracy of fourth mixture ratio R4(k), i.e., the approximation accuracy of the amplitude or power of noise and crosstalk.
[0077] The operation of mixer 905 is the same as that of mixer 305 by changing the input and output signals of mixer 305 to those of mixer 905 as shown in FIG. 7, so a detailed description will be omitted.
[0078] 3.2. Second Configuration of Estimation Unit 806 Fig. 8 is a block diagram showing a second internal configuration of estimating section 806. In addition to mixing section 506 and signal ratio estimating section 503, estimating section 806 further includes mixing section 1106 and signal ratio estimating section 1103. Mixing section 506 in Fig. 8 is configured to receive a delayed noise estimate e2D(k) instead of the delayed reference signal xRD(k) (delayed second mixed signal) in mixing section 506 in Fig. 5. The operation of mixing section 506 in Fig. 8 is the same as the operation of mixing section 506 in Fig. 5, and therefore detailed description thereof will be omitted.
[0079] The mixer 1106 mixes the delay estimate e1D(k) of the speech (the delayed enhancement signal, i.e., the delay estimate of the first signal) and the pseudo crosstalk n2(k) (the estimate of the third signal) using the coefficient vector w2(k) of the adaptive filter 803 to generate a second mixed signal n4(k). The operation of the mixer 1106 is the same as that of the mixer 506 except that the input and output signals of the mixer 506 are changed to those of the mixer 1106 as shown in Fig. 8, and therefore a detailed description thereof will be omitted.
[0080] The signal ratio estimator 1103 receives the noise estimate e2(k) and the second mixed signal n4(k) and estimates the ratio of the amplitude or power of the noise to the crosstalk as a fourth mixed ratio R4(k). The fourth mixed ratio R4(k) may be the ratio of the amplitude or power of the noise estimate e2(k) to the second mixed signal n4(k), or may be calculated by adding a small constant to the amplitude or power of the noise estimate e2(k) and the second mixed signal n4(k). Adding a small constant has the effect of preventing divergence of the quotient due to division. Alternatively, one or both of the noise estimate e2(k) and the second mixed signal n4(k) may be averaged before use. Averaging can improve the accuracy of the ratio calculation.
[0081] The second internal configuration of estimator 806 shown in Fig. 8 is equivalent to the first internal configuration of estimator 806 shown in Fig. 7. That is, in the first internal configuration shown in Fig. 7, two estimates for the ratio of the amplitude or power of noise to crosstalk are generated by signal ratio estimators 901 and 902, and then these estimates are mixed to calculate fourth mixture ratio R4(k). In the second internal configuration shown in Fig. 8, two types of estimates, i.e., speech delay estimate e1D(k) and pseudo crosstalk n2(k), are mixed using coefficient vector w2(k) of adaptive filter 803 to generate second mixed signal n4(k), determine the denominator, and then calculate fourth mixture ratio R4(k) by combining this with noise estimate e2(k), which is the numerator. These two types of configurations are possible because the first internal configuration shown in Fig. 7 and the second internal configuration shown in Fig. 8 use the same numerator, i.e., noise estimate e2(k), when estimating the ratio of the amplitude or power of noise to crosstalk at input terminal 202. The second internal configuration of the estimating unit 806 shown in FIG. 8 has fewer components and is simpler than the first internal configuration shown in FIG.
[0082] With the above configuration, this embodiment can take into account the effect of delay caused by the acoustic impulse responses that approximate the adaptive filters 203 and 803, and appropriately control the coefficient update of the adaptive filters 203 and 803, thereby making it possible to obtain, in a short time, an output signal with little residual noise and little signal distortion.
[0083] 4. Other Embodiments Although multiple embodiments of the present invention have been described in detail above, any system or device that combines the separate features included in each embodiment also falls within the scope of the present invention.
[0084] The present invention may be applied to a system consisting of multiple devices, or to a single device. Furthermore, the present invention is also applicable when an information processing program (signal processing program) that realizes the functions of the above-described embodiments is supplied directly or remotely to a system or device. Such a program is executed by a processor such as a DSP (Digital Signal Processor) that constitutes a signal processing device or noise cancellation device. Furthermore, the scope of the present invention also includes a program installed on a computer to realize the functions of the present invention, a medium storing the program, and a WWW (World Wide Web) server from which the program can be downloaded.
[0085] 9 is a configuration diagram of a computer 1200 that executes a signal processing program when the first to third embodiments are configured by the signal processing program. The computer 1200 includes an input unit 1201, a processor 1203, an output unit 1202, and a memory 1204.
[0086] The processor 1203 controls the operation of the computer 1200 by reading a signal processing program stored in the memory 1204. The processor 1203 is, for example, a processor such as a DSP, a CPU (Central Processing Unit), or an MPU (Micro-Processing Unit). The memory 1204 includes one or more of a RAM (Random Access Memory), a ROM (Read-Only Memory), a flash memory, an EPROM (Erasable Programmable Read-Only Memory), and an EEPROM (Electrically Erasable Programmable Read-Only Memory).
[0087] Fig. 10 is a flowchart showing an example of signal processing by the processor of the computer shown in Fig. 9. The example shown in Fig. 10 is a flowchart in the case where the computer 1200 functions as the signal processing device 100 according to the first embodiment.
[0088] As shown in FIG. 10, in step S10, the processor 1203 executing the signal processing program first inputs, from the input unit 1201, a first mixture signal xP(k) in which a first signal and a second signal are mixed, and then inputs a second mixture signal xR(k) in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed. In step S11, the processor 1203 processes the second mixture signal xR(k) with a first adaptive filter (adaptive filter 103) to generate an estimate n1(k) of the second signal, and in step S12, generates an estimate e1(k) of the first signal from the first mixture signal xP(k) and the estimate n1(k) of the second signal. In step S13, the processor 1203 delays the second mixture signal xR(k) to generate a delayed second mixture signal xRD(k). In step S14, the processor 1203 estimates the ratio of the amplitudes or powers of the first signal and the second signal as a first mixture ratio R1(k) using the estimated value e1(k) of the first signal, the estimated value n1(k) of the second signal, the delayed second mixture signal xRD(k), and the coefficient vector w1(k) of the first adaptive filter (adaptive filter 103). In step S15, the processor 1203 controls the generation of the estimated value n1(k) of the second signal using the first mixture ratio R1(k), thereby achieving the same effects as in the first embodiment.
[0089] Furthermore, when the computer 1200 functions as the signal processing device according to the second embodiment, the processor 1203 receives the first mixture signal xP(k) and the second mixture signal xR(k) in step S10, and processes the second mixture signal xR(k) using a first adaptive filter (adaptive filter 203) to generate an estimated value n1(k) of the second signal in step S11. The processor 1203 generates an estimated value e1(k) of the first signal from the first mixture signal xP(k) and the estimated value n1(k) of the second signal in step S12. The processor 1203 delays the second mixture signal xR(k) to generate a delayed second mixture signal xRD(k) in step S13. In step S14, the processor 1203 generates a first mixture ratio R1(k) using the first signal estimate e1(k), the second signal estimate n1(k), the delayed second mixture signal xRD(k), the first mixture signal xP(k), and the coefficient 141 of the first adaptive filter (adaptive filter 203). In step S15, the processor 1203 controls the generation of the second signal estimate n1(k) using the first mixture ratio R1(k).
[0090] Furthermore, when the computer 1200 functions as the signal processing device according to the third embodiment, the processor 1203 delays the first signal estimate e1(k) to generate a delay estimate e1D(k) of the first signal and delays the fourth signal estimate e2(k) to generate a delay estimate e2D(k) of the fourth signal in step S14. The processor 1203 also generates a first mixture ratio R1(k) using the first signal estimate e1(k), the second signal estimate n1(k), the fourth signal delay estimate e2D(k), and the coefficient 141 of the first adaptive filter (adaptive filter 203). The processor 1203 also processes the first signal estimate e1(k) with the second adaptive filter (adaptive filter 803) to generate a third signal estimate n2(k), and generates a fourth signal estimate e2(k) by subtracting the third signal estimate n2(k) from the second mixture signal xR(k). Furthermore, processor 1203 processes fourth signal estimate e2(k) instead of reference signal xR(k) in the first adaptive filter (adaptive filter 203), and further estimates the ratio of the amplitudes or powers of noise and crosstalk as a fourth mixing ratio R4(k) using the fourth signal estimate e2(k), the third signal estimate n2(k), the first signal delay estimate e1D(k), and the coefficients of the second adaptive filter (adaptive filter 803). In step S15, processor 1203 further uses the fourth mixing ratio R4(k) to control generation of the third signal estimate n2(k).
[0091] [5. Other] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0092] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0093] [6. Effects] As described above, the signal processing device 100 according to the first embodiment includes a first input unit 101 (corresponding to an example of a first input means), a second input unit 102 (corresponding to an example of a second input means), an adaptive filter 103 (corresponding to an example of a first adaptive filter), a subtraction unit 104 (corresponding to an example of a first subtraction unit), a delay unit 110, and an estimation unit 106. The first input unit 101 receives a first mixture signal xP(k) in which a first signal and a second signal are mixed. The second input unit 102 receives a second mixture signal xR(k) in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed. The adaptive filter 103 filters the second mixture signal xR(k) to generate an estimate n1(k) of the second signal. The subtractor 104 generates an estimated value e1(k) of the first signal from the first mixture signal xP(k) and the estimated value n1(k) of the second signal. The delay unit 110 delays the second mixture signal xR(k) based on the coefficients of the adaptive filter 103 to generate a delayed second mixture signal xRD(k). The estimator 106 estimates the amplitude or power ratio of the first signal to the second signal at the input terminal 201 as a first mixture ratio R1(k) using the estimated value e1(k) of the first signal, the estimated value n1(k) of the second signal, the delayed second mixture signal xRD(k), and the coefficients 141 of the adaptive filter 103. The coefficient update controller 107 outputs a control signal μ1(k) to the adaptive filter 103 to reduce the amount of update of the coefficients 141 of the adaptive filter 103 when the value of the first mixture ratio R1(k) obtained by the estimator 106 is large. The signal processing device 100 uses the control signal μ1(k) to control the adaptive filter 103. This allows the signal processing device 100 to take into account the effect of delay caused by the acoustic impulse response that the adaptive filter approximates, and appropriately controls the update of the adaptive filter coefficients, thereby enabling an output signal with little residual noise and little signal distortion to be obtained in a short time.
[0094] The signal processing device according to the second embodiment includes an input terminal 201 (corresponding to an example of a first input means), an input terminal 202 (corresponding to an example of a second input means), an adaptive filter 203 (corresponding to an example of a first adaptive filter), a subtraction unit 204 (corresponding to an example of a first subtraction unit), a delay unit 210, and an estimation unit 206. The input terminal 201 receives a noisy signal xP(k) (corresponding to an example of a first mixture signal) in which a speech signal (corresponding to an example of a first signal) and noise (corresponding to an example of a second signal) are mixed. The input terminal 202 receives a reference signal xR(k) (corresponding to an example of a second mixture signal) in which a signal correlated with the speech signal (corresponding to an example of a third signal) and a signal correlated with the noise (corresponding to an example of a fourth signal) are mixed. The adaptive filter 203 filters the reference signal xR(k) to generate pseudo-noise n1(k) (corresponding to an example of an estimated value of the second signal). Subtraction unit 104 generates a speech estimate e1(k) (corresponding to an example of a first signal estimate e1(k)) from noisy signal xP(k) and pseudo-noise n1(k). Delay unit 210 delays reference signal xR(k) based on the coefficients of adaptive filter 203 to generate a delayed reference signal xRD(k). Estimation unit 206 estimates the ratio of the amplitude or power of the speech signal to noise at input terminal 201 as a first mixture ratio R1(k) using speech estimate e1(k), pseudo-noise n1(k), delayed reference signal xRD(k), and coefficient vector w1(k) of adaptive filter 203. Noise canceller 200 controls adaptive filter 203 using first mixture ratio R1(k). As a result, the signal processing device of the second embodiment can take into account the effect of delay caused by the acoustic impulse response that approximates the adaptive filter, and can appropriately control the update of the adaptive filter coefficients to obtain an output signal with little residual noise and little signal distortion in a short time.
[0095] The signal processing device according to the third embodiment includes input terminal 201, input terminal 202, adaptive filter 203, delay unit 210, subtraction unit 204, subtraction unit 804 (corresponding to an example of a second subtraction unit), adaptive filter 803 (corresponding to an example of a second adaptive filter), delay unit 810, and estimation unit 806. The adaptive filter 803 performs filtering on the estimated value e1(k) of the first signal to generate pseudo crosstalk n2(k) (corresponding to an example of an estimated value of the third signal). The subtraction unit 804 subtracts the pseudo crosstalk n2(k) from the reference signal xR(k) to generate a noise estimated value e2(k) (corresponding to an example of an estimated value of the fourth signal). The adaptive filter 203 receives the noise estimated value e2(k) instead of the reference signal xR(k). In addition to the functions of the estimation unit 206, the estimation unit 806 further uses the noise estimate e2(k), the noise delay estimate e2D(k), the pseudo crosstalk n2(k), the first signal delay estimate e1D(k), and the coefficients of the adaptive filter 803 to estimate the ratio of the amplitude or power of the noise and the crosstalk at the input terminal 202 as a fourth mixing ratio R4(k). The signal processing device according to the third embodiment controls the adaptive filter 803 using the fourth mixing ratio R4(k). This allows the signal processing device according to the third embodiment to take into account the influence of delay caused by the acoustic impulse response approximated by the adaptive filter, and to appropriately control the update of the adaptive filter coefficients to obtain an output signal with low residual noise and low signal distortion in a short time.
[0096] The above describes the embodiments of the present application in detail based on the drawings, but this is merely an example, and the present invention can be implemented in other forms that include the embodiments described in the Disclosure of the Invention section and that have been modified and improved in various ways based on the knowledge of those skilled in the art.
[0097] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, a subtraction section can be read as subtraction means or a subtraction circuit. (Appendix 1) a first input means for inputting a first mixed signal in which the first signal and the second signal are mixed; a second input means for inputting a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed; a first adaptive filter (adaptive filter 103, 203) for filtering the second mixture signal to generate an estimate of the second signal; a first subtraction unit (subtraction unit 104, 204) that generates an estimate of the first signal from the first mixed signal and an estimate of the second signal; a first delay unit (delay unit 210) that delays the second mixture signal using a coefficient of the first adaptive filter to generate a delayed second mixture signal; an estimation unit (estimation unit 106, 206, 806) that estimates a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; Equipped with The first adaptive filter is controlled using the first mixing ratio. Signal processing device. (Appendix 2) The estimation unit (estimation unit 106, 206) a first signal ratio estimator (signal ratio estimator 301) that estimates a ratio of amplitudes or powers of the first signal and the second signal as a second mixture ratio using the estimated value of the first signal and the estimated value of the second signal; a second signal ratio estimator (signal ratio estimator 302) that estimates a ratio of amplitude or power of the first signal to the second signal as a third mixture ratio using the estimated value of the first signal and the delayed second mixture signal; a first mixer (mixer 305) that mixes the second mixture ratio and the third mixture ratio based on a time change in a coefficient of the first adaptive filter to generate the first mixture ratio; 2. The signal processing device according to claim 1, comprising: (Appendix 3) The first mixing section (mixing section 305) When coefficient update of the first adaptive filter starts, the content ratio of the third mixture ratio in the first mixture ratio is set to 100%, and when a time change in the coefficient of the first adaptive filter becomes sufficiently small, the content ratio of the third mixture ratio in the first mixture ratio is set to 0%. 3. The signal processing device of claim 2. (Appendix 4) The estimation unit (estimation unit 106, 206) a second mixer (mixer 506) that mixes the delayed second mixed signal and the estimated value of the second signal based on time changes in the coefficients of the first adaptive filter to generate a first mixed signal; a third signal ratio estimator (signal ratio estimator 503) that estimates a ratio of amplitudes or powers of the first signal and the second signal as the first mixture ratio using the first mixed signal and an estimated value of the first signal; 2. The signal processing device according to claim 1, comprising: (Appendix 5) The second mixing section (mixing section 506) When the coefficient update of the first adaptive filter starts, the content ratio of the delayed second mixture signal in the first mixture signal is set to 100%, and when the time change of the coefficient of the first adaptive filter becomes sufficiently small, the content ratio of the delayed second mixture signal in the first mixture signal is set to 0%. 5. The signal processing device of claim 4. (Appendix 6) a second adaptive filter (adaptive filter 803) that filters the estimate of the first signal to generate an estimate of the third signal; a second subtraction unit (subtraction unit 804) that subtracts the estimated value of the third signal from the second mixed signal to generate an estimated value of the fourth signal; a second delay unit (delay unit 810) that delays the estimate of the first signal using a coefficient of the second adaptive filter to generate a delayed estimate of the first signal; The first adaptive filter is an estimate of the fourth signal is input instead of the second mixed signal; the first delay unit delays the estimated value of the fourth signal instead of the second mixed signal to generate a delayed estimated value of the fourth signal; The estimation unit a delay estimate value of the fourth signal is used as an input instead of the delayed second mixed signal; further estimating an amplitude or power ratio between the fourth signal and the third signal as a fourth mixture ratio using the fourth signal estimate, the third signal estimate, the delay estimate of the first signal, and a coefficient of the second adaptive filter; The second adaptive filter is controlled using the fourth mixing ratio. 2. The signal processing device of claim 1. (Appendix 7) The estimation unit (estimation unit 806) a first signal ratio estimator (signal ratio estimator 301) that estimates a ratio of amplitudes or powers of the first signal and the second signal as a second mixture ratio using the estimated value of the first signal and the estimated value of the second signal; a second signal ratio estimator (signal ratio estimator 302) that estimates a ratio of amplitudes or powers of the first signal and the second signal as a third mixture ratio using the estimated value of the first signal and the delayed estimated value of the fourth signal; a first mixer (mixer 305) that mixes the second mixture ratio and the third mixture ratio based on a time change in a coefficient of the first adaptive filter to generate the first mixture ratio; a fourth signal ratio estimator (signal ratio estimator 901) that estimates a ratio of amplitudes or powers of the fourth signal and the third signal as a fifth mixture ratio using the estimated value of the fourth signal and the estimated value of the third signal; a fifth signal ratio estimator (signal ratio estimator 902) that estimates a ratio of amplitudes or powers of the fourth signal and the third signal as a sixth mixture ratio using the estimated value of the fourth signal and the delayed estimated value of the first signal; a third mixer (mixer 905) that mixes the fifth mixture ratio and the sixth mixture ratio based on a time change in the coefficient of the second adaptive filter to generate the fourth mixture ratio; 7. The signal processing device according to claim 6, comprising: (Appendix 8) The first mixing section (mixing section 305) setting a content ratio of the third mixture ratio in the first mixture ratio to 100% when a coefficient update of the first adaptive filter starts, and setting a content ratio of the third mixture ratio in the first mixture ratio to 0% when a time change in the coefficient of the first adaptive filter becomes sufficiently small; The third mixing section (mixing section 905) The proportion of the sixth mixture ratio in the fourth mixture ratio is set to 100% when the coefficient update of the second adaptive filter starts, and the proportion of the sixth mixture ratio in the fourth mixture ratio is set to 0% when the change over time in the coefficient of the second adaptive filter becomes sufficiently small. 9. The signal processing device according to claim 7 or 8. (Appendix 9) The estimation unit (estimation unit 806) a second mixer (mixer 506) that mixes the delay estimate of the fourth signal and the estimate of the second signal based on time variations in the coefficients of the first adaptive filter to generate a first mixed signal; a third signal ratio estimator (signal ratio estimator 503) that estimates a ratio of amplitudes or powers of the first signal and the second signal as the first mixture ratio using estimated values of the first mixed signal and the first signal; a fourth mixer (mixer 1106) that mixes the delay estimate of the first signal and the estimate of the third signal based on time variations in the coefficients of the second adaptive filter to generate a second mixed signal; a sixth signal ratio estimator (signal ratio estimator 1103) that estimates a ratio of amplitudes or powers of the fourth signal and the third signal as the fourth mixture ratio using estimated values of the second mixed signal and the fourth signal; 7. The signal processing device according to claim 6, comprising: (Appendix 10) The second mixing section (mixing section 506) setting a content ratio of the delay estimate value of the fourth signal in the first mixed signal to 100% when a coefficient update of the first adaptive filter starts, and setting a content ratio of the corrected estimate value of the fourth signal in the first mixed signal to 0% when a time change in the coefficient of the first adaptive filter becomes sufficiently small; The fourth mixing unit (mixing unit 1106) When coefficient updating of the second adaptive filter starts, the content ratio of the delay estimate value of the first signal in the second mixed signal is set to 100%, and when time change in the coefficient of the second adaptive filter becomes sufficiently small, the content ratio of the corrected estimate value of the first signal in the second mixed signal is set to 0%. 10. The signal processing device of claim 9. (Appendix 11) The time change of the coefficient is The time change of the sum of squares or absolute values of the coefficients is 11. The signal processing device of claim 3, 5, 8, or 10. (Appendix 12) The time change of the coefficient is The time change of the squared partial sum or absolute partial sum of the coefficients is 11. The signal processing device of claim 3, 5, 8, or 10. (Appendix 13) Input the first mixed signal, which is a mixture of the first and second signals, a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed is input; processing the second mixed signal with a first adaptive filter (adaptive filter 103, 203) to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixture signal using a coefficient of the first adaptive filter to generate a delayed second mixture signal; estimating an amplitude or power ratio between the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; and controlling generation of an estimate of the second signal using the first mixing ratio. Signal processing methods. (Appendix 14) processing the estimate of the first signal with a second adaptive filter (adaptive filter 803) to generate an estimate of the third signal; subtracting the estimate of the third signal from the second mixed signal to generate an estimate of the fourth signal; delaying the estimate of the first signal using coefficients of the second adaptive filter to generate a delayed estimate of the first signal; the first adaptive filter processes an estimate of the fourth signal instead of the second mixed signal; delaying the estimate of the fourth signal instead of the second mixed signal to generate a delayed estimate of the fourth signal; further using the fourth signal estimate, the third signal estimate, the delay estimate of the first signal, and coefficients of the second adaptive filter, further estimating a ratio of amplitude or power of the fourth signal to the third signal as a fourth mixture ratio; and controlling generation of an estimate of the third signal using the fourth mixing ratio. 14. A signal processing method according to claim 13. (Appendix 15) On the computer, inputting a first mixed signal in which a first signal and a second signal are mixed; inputting a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed; processing the second mixed signal with a first adaptive filter (adaptive filter 103, 203) to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixture signal using a coefficient of the first adaptive filter to generate a delayed second mixture signal; a step of estimating a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; using the first mixing ratio to control generation of an estimate of the second signal; A signal processing program that executes the above. (Appendix 16) On the computer, processing the estimate of the first signal with a second adaptive filter (adaptive filter 803) to generate an estimate of the third signal; subtracting the estimate of the third signal from the second mixture signal to generate an estimate of the fourth signal; delaying the estimate of the first signal using coefficients of the second adaptive filter to generate a delayed estimate of the first signal; processing an estimate of the fourth signal instead of the second mixed signal with the first adaptive filter; further using the fourth signal estimate, the third signal estimate, the delay estimate of the first signal, and coefficients of the second adaptive filter, further estimating a ratio of amplitudes or powers of the fourth signal and the third signal as a fourth mixing ratio; using the fourth mixing ratio to control generation of an estimate of the third signal; 17. The signal processing program according to claim 16, (Appendix 17) Input the first mixed signal, which is a mixture of the first and second signals, a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed is input; processing the second mixed signal with a first adaptive filter (adaptive filter 103, 203) to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixed signal to generate a delayed second mixed signal; estimating an amplitude or power ratio between the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; and controlling generation of an estimate of the second signal using the first mixing ratio. Signal processing methods. [Explanation of symbols]
[0098] 100 Signal processing device 101 First input section 102 Second input section 103, 203 Adaptive filter (corresponding to an example of the first adaptive filter) 104, 204 subtraction unit (corresponding to an example of a first subtraction unit) 106, 206, 806 Estimation part 107 Coefficient update control section 141 coefficient 200, 800 Noise canceller (equivalent to an example of a signal processing device) 201 input terminal (corresponding to an example of the first input unit) 202 input terminal (corresponding to an example of the second input unit) 205, 805 output terminal 301 signal ratio estimation unit (corresponding to an example of a first signal ratio estimation unit) 302 signal ratio estimation unit (corresponding to an example of a second signal ratio estimation unit) 305 Mixing section (corresponding to an example of the first mixing section) 310 correction unit (corresponding to an example of the first correction unit) 506 Mixing section (equivalent to an example of the second mixing section) 503 signal ratio estimation unit (corresponding to an example of a third signal ratio estimation unit) 710 Correction unit (corresponding to an example of the second correction unit) 803 Adaptive filter (equivalent to an example of a second adaptive filter) 804 subtraction unit (corresponding to an example of the second subtraction unit) 901 signal ratio estimation unit (corresponding to an example of a fourth signal ratio estimation unit) 902 signal ratio estimation unit (corresponding to an example of a fifth signal ratio estimation unit) 905 Mixing section (equivalent to an example of the third mixing section) 1106 Mixing section (equivalent to an example of the 4th mixing section) 1103 signal ratio estimation unit (corresponding to an example of a sixth signal ratio estimation unit) A signal source B signal source xP(k) 1st mixed signal xR(k) 2nd mixed signal xRD(k) Delayed second mixed signal e1(k) Estimated value of the first signal, estimated value of the speech signal e2(k) Noise estimate (equivalent to an example of the fourth signal estimate) e1D(k) Delay estimate of the first signal, delay estimate of the voice signal e2D(k) Delayed noise estimate (equivalent to an example of the delayed fourth signal estimate) n1(k) is the estimated value of the second signal, pseudo-noise (corresponding to an example of the estimated value of the second signal) n2(k) pseudo crosstalk (corresponding to an example of an estimate of the third signal) n3(k) mixed signal (corresponding to an example of the first mixed signal) n4(k) mixed signal (corresponding to an example of the second mixed signal) R1(k) First mixture ratio R2(k) Second mixture ratio R3(k) Third mixture ratio R4(k) 4th mixture ratio R5(k) 5th mixture ratio R6(k) 6th mixture ratio
Claims
1. a first input means for inputting a first mixed signal in which the first signal and the second signal are mixed; a second input means for inputting a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed; a first adaptive filter for filtering the second mixed signal to generate an estimate of the second signal; a first subtraction unit that generates an estimate of the first signal from the first mixed signal and an estimate of the second signal; a first delay unit that delays the second mixture signal using a coefficient of the first adaptive filter to generate a delayed second mixture signal; an estimation unit that estimates a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; Equipped with The first mixing ratio is used to control the first adaptive filter. Signal processing device.
2. The estimation unit a first signal ratio estimator that estimates a ratio of amplitudes or powers of the first signal and the second signal as a second mixture ratio using the estimated value of the first signal and the estimated value of the second signal; a second signal ratio estimator that estimates a ratio of amplitudes or powers of the first signal and the second signal as a third mixture ratio using the estimated value of the first signal and the delayed second mixture signal; a first mixer that mixes the second mixture ratio and the third mixture ratio based on a time change in a coefficient of the first adaptive filter to generate the first mixture ratio; The signal processing device according to claim 1 , comprising:
3. The first mixing section When coefficient update of the first adaptive filter is started, the content ratio of the third mixture ratio in the first mixture ratio is set to 100%, and when a time change in the coefficient of the first adaptive filter becomes sufficiently small, the content ratio of the third mixture ratio in the first mixture ratio is set to 0%. The signal processing device according to claim 2 .
4. The estimation unit a second mixer that mixes the delayed second mixed signal and the estimated value of the second signal based on time changes in coefficients of the first adaptive filter to generate a first mixed signal; a third signal ratio estimator that estimates, as the first mixture ratio, a ratio of amplitudes or powers of the first signal and the second signal using the first mixed signal and an estimated value of the first signal; The signal processing device according to claim 1 , comprising:
5. The second mixing section is When the coefficient update of the first adaptive filter starts, the content ratio of the delayed second mixture signal in the first mixture signal is set to 100%, and when the time change of the coefficient of the first adaptive filter becomes sufficiently small, the content ratio of the delayed second mixture signal in the first mixture signal is set to 0%. The signal processing device according to claim 4 .
6. a second adaptive filter that filters the estimate of the first signal to generate an estimate of the third signal; a second subtraction unit that subtracts the estimated value of the third signal from the second mixed signal to generate an estimated value of the fourth signal; a second delay unit that delays the estimate of the first signal using a coefficient of the second adaptive filter to generate a delayed estimate of the first signal; The first adaptive filter is an estimate of the fourth signal is input instead of the second mixed signal; the first delay unit delays the estimate of the fourth signal instead of the second mixed signal to generate a delayed estimate of the fourth signal; The estimation unit a delay estimate value of the fourth signal is input instead of the delayed second mixed signal; further estimating, as a fourth mixture ratio, a ratio of amplitudes or powers of the fourth signal and the third signal using the estimated value of the fourth signal, the estimated value of the third signal, the delay estimated value of the first signal, and a coefficient of the second adaptive filter; The second adaptive filter is controlled using the fourth mixing ratio. The signal processing device according to claim 1 .
7. The estimation unit a first signal ratio estimator that estimates a ratio of amplitudes or powers of the first signal and the second signal as a second mixture ratio using the estimated value of the first signal and the estimated value of the second signal; a second signal ratio estimator that estimates a ratio of amplitudes or powers of the first signal and the second signal as a third mixture ratio using the estimated value of the first signal and the delayed estimated value of the fourth signal; a first mixer that mixes the second mixture ratio and the third mixture ratio based on a time change in a coefficient of the first adaptive filter to generate the first mixture ratio; a fourth signal ratio estimator that estimates a ratio of amplitudes or powers of the fourth signal and the third signal as a fifth mixture ratio using the estimated value of the fourth signal and the estimated value of the third signal; a fifth signal ratio estimator that estimates a ratio of amplitudes or powers of the fourth signal and the third signal as a sixth mixture ratio using the estimated value of the fourth signal and the delayed estimated value of the first signal; a third mixer that mixes the fifth mixture ratio and the sixth mixture ratio based on a time change in a coefficient of the second adaptive filter to generate the fourth mixture ratio; The signal processing device according to claim 6, comprising:
8. The first mixing section setting a content ratio of the third mixture ratio in the first mixture ratio to 100% when a coefficient update of the first adaptive filter starts, and setting a content ratio of the third mixture ratio in the first mixture ratio to 0% when a time change in the coefficient of the first adaptive filter becomes sufficiently small; The third mixing section is When coefficient update of the second adaptive filter is started, the content ratio of the sixth mixture ratio in the fourth mixture ratio is set to 100%, and when a time change in the coefficient of the second adaptive filter becomes sufficiently small, the content ratio of the sixth mixture ratio in the fourth mixture ratio is set to 0%. The signal processing device according to claim 7 .
9. The estimation unit a second mixer that mixes the delay estimate of the fourth signal and the estimate of the second signal based on time variations in the coefficients of the first adaptive filter to generate a first mixed signal; a third signal ratio estimator that estimates a ratio of amplitudes or powers of the first signal and the second signal as the first mixture ratio using estimated values of the first mixed signal and the first signal; a fourth mixer that mixes the delay estimate of the first signal and the estimate of the third signal based on time variations in coefficients of the second adaptive filter to generate a second mixed signal; a sixth signal ratio estimator that estimates a ratio of amplitudes or powers of the fourth signal and the third signal as the fourth mixture ratio using estimated values of the second mixed signal and the fourth signal; The signal processing device according to claim 6, comprising:
10. The second mixing section is setting a content ratio of the delay estimate value of the fourth signal in the first mixed signal to 100% when coefficient update of the first adaptive filter starts, and setting a content ratio of the corrected estimate value of the fourth signal in the first mixed signal to 0% when a time change in the coefficient of the first adaptive filter becomes sufficiently small; The fourth mixing section is When coefficient updating of the second adaptive filter starts, the content ratio of the delay estimate value of the first signal in the second mixed signal is set to 100%, and when time change in the coefficient of the second adaptive filter becomes sufficiently small, the content ratio of the corrected estimate value of the first signal in the second mixed signal is set to 0%. The signal processing device according to claim 9 .
11. The time change of the coefficient is The time change of the sum of squares or the sum of absolute values of the coefficients is 11. The signal processing device according to claim 3, 5, 8 or 10.
12. The time change of the coefficient is is the time change of the squared partial sum or absolute partial sum of the coefficients 11. The signal processing device according to claim 3, 5, 8 or 10.
13. A first mixed signal in which the first signal and the second signal are mixed is input; a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed is input; processing the second mixed signal with a first adaptive filter to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixture signal using coefficients of the first adaptive filter to generate a delayed second mixture signal; estimating a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; and controlling generation of an estimate of the second signal using the first mixing ratio. Signal processing methods.
14. processing the estimate of the first signal with a second adaptive filter to generate an estimate of the third signal; subtracting the estimate of the third signal from the second mixed signal to generate an estimate of the fourth signal; delaying the estimate of the first signal using coefficients of the second adaptive filter to generate a delayed estimate of the first signal; the first adaptive filter processes an estimate of the fourth signal instead of the second mixed signal; delaying the estimate of the fourth signal instead of the second mixed signal to generate a delayed estimate of the fourth signal; further using the fourth signal estimate, the third signal estimate, the delay estimate of the first signal, and coefficients of the second adaptive filter; further estimating a ratio of amplitude or power of the fourth signal to the third signal as a fourth mixture ratio; and controlling generation of an estimate of the third signal using the fourth mixing ratio.
14. The signal processing method according to claim 13.
15. On the computer, inputting a first mixed signal in which the first signal and the second signal are mixed; inputting a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed; processing the second mixed signal with a first adaptive filter to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixture signal using coefficients of the first adaptive filter to generate a delayed second mixture signal; a step of estimating a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; using the first mixing ratio to control generation of an estimate of the second signal; A signal processing program that executes the above.
16. On the computer, processing the estimate of the first signal with a second adaptive filter to generate an estimate of the third signal; subtracting the estimate of the third signal from the second mixture signal to generate an estimate of the fourth signal; delaying the estimate of the first signal using coefficients of the second adaptive filter to generate a delayed estimate of the first signal; processing an estimate of the fourth signal instead of the second mixed signal with the first adaptive filter; further using the fourth signal estimate, the third signal estimate, the delay estimate of the first signal, and coefficients of the second adaptive filter; further estimating a ratio of amplitude or power of the fourth signal to the third signal as a fourth mixing ratio; using the fourth mixing ratio to control generation of an estimate of the third signal; 16. The signal processing program according to claim 15, which causes the signal processing program to execute the following:
17. A first mixed signal in which the first signal and the second signal are mixed is input; a second mixed signal in which a third signal correlated with the first signal and a fourth signal correlated with the second signal are mixed is input; processing the second mixed signal with a first adaptive filter to generate an estimate of the second signal; generating an estimate of the first signal from the first mixed signal and an estimate of the second signal; delaying the second mixture signal by sequentially switching among a plurality of predetermined constants to generate a delayed second mixture signal; estimating a ratio of amplitudes or powers of the first signal and the second signal as a first mixture ratio using the estimated value of the first signal, the estimated value of the second signal, the delayed second mixture signal, and a coefficient of the first adaptive filter; and controlling generation of an estimate of the second signal using the first mixing ratio. Signal processing methods.
Citation Information
Patent Citations
Voice input device
JP1994075591A
Echo canceller device
JP1997055687A
Method and device for erasing noise
JP1998215193A
Noise elimination method and noise eliminating device using it
JP2000172299A
Signal processing method and apparatus
WO2005024787A1