Acoustic signal enhancement device, acoustic signal enhancement method, and program

The acoustic signal enhancement device addresses the issue of suppressing time-varying unwanted sounds by optimizing filter coefficients based on temporal and spatial sound classification, ensuring accurate suppression even with estimation errors in acoustic transfer characteristics.

JP7810178B2Active Publication Date: 2026-02-03NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023531342
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-30
Filing Date
2021-09-30
Publication Date
2026-02-03
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing acoustic signal enhancement devices fail to accurately suppress time-varying unwanted sounds when the estimated value of acoustic transfer characteristics contains errors or cannot be obtained.

Method used

An acoustic signal enhancement device that includes a beamformer unit, a switch unit, and a weighted spatial covariance estimator, which updates parameters to suppress unwanted sounds by classifying temporal and spatial states of recorded sound, optimizing filter coefficients based on a likelihood function that assumes a complex Gaussian distribution for the target sound.

Benefits of technology

The device effectively suppresses time-varying unwanted sounds with high precision even when the estimated acoustic transfer characteristics contain errors or are unavailable, enhancing audio signal quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007810178000037
    Figure 0007810178000037
  • Figure 0007810178000038
    Figure 0007810178000038
  • Figure 0007810178000039
    Figure 0007810178000039
Patent Text Reader

Abstract

This acoustic signal enhancement device updates a parameter using a frequency-divided recorded sound as input, and includes: a beam former unit that performs beam former processing on the basis of an updated weighted spatial covariance matrix with a switch weight being defined as a weight indicating, in the classification of the time-varying spatial state of the recorded sound, a ratio at which the recorded sound belongs to respective classes at respective times, and updates an auxiliary estimated value of a target sound; a switch unit that, on the basis of the updated auxiliary estimated value, updates the switch weight and the power of the target sound, and outputs an estimated value of the target sound; and a weighted spatial covariance estimation unit that, on the basis of the updated switch weight and power, updates the weighted spatial covariance matrix.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an audio signal enhancement device, an audio signal enhancement method, and a program for suppressing noise and reverberation from recorded sound and separating and estimating each target sound. [Background technology]

[0002] Non-Patent Document 1 discloses an acoustic signal enhancement device that estimates a target sound by switching over time between multiple outputs obtained by applying recorded sound to a beamformer (see FIG. 1). According to the acoustic signal enhancement device 8 in Non-Patent Document 1, under the condition that estimated values ​​of the acoustic transfer characteristics (hereinafter simply referred to as acoustic transfer characteristics) related to the direct sound and early reflected sound of the target sound are given, it is determined which of multiple beamformer outputs to use based on the criterion of minimizing the power of the processed sound, and the filter coefficients of each beamformer are optimized, thereby enhancing the acoustic signal.

[0003] Non-Patent Document 2 discloses an audio signal enhancement device that achieves audio signal enhancement even in reverberant environments by sequentially applying a dereverberation process that suppresses reverberation in recorded sound and a beamformer (see Fig. 2). According to the audio signal enhancement device 9 in Non-Patent Document 2, audio signal enhancement is performed by simultaneously optimizing dereverberation and each filter coefficient of the beamformer under the condition that an estimated value of the acoustic transfer characteristics of the target sound is given, based on the criterion that the target sound follows a Gaussian distribution whose power changes over time. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Kouei Yamaoka, Nobutaka Ono, Shoji Makino, and Takeshi Yamada, TIME-FREQUENCY-BIN-WISE SWITCHING OF MINIMUM VARIANCE DISTORTIONLESS RESPONSE BEAMFORMER FOR UNDERDETERMINED SITUATIONS, Proc. IEEE ICASSP, pp.7908-7912, 2019. [Non-patent document 2] Tomohiro Nakatani, Christoph Boeddeker, Keisuke Kinoshita, Rintaro Ikeshita, Marc Delcroix, Reinhold Haeb-Umbach, Jointly optimal denoising, dereverberation, and source separation, IEEE / ACM Trans. Audio, Speech, and Language Processing, vol. 28, pp. 2267 - 2282, 2020. Summary of the Invention [Problem to be solved by the invention]

[0005] According to Non-Patent Document 1, the filter coefficients of the beamformer are optimized without taking into consideration the statistical properties of the target sound, so if the estimated value of the acoustic transfer characteristic contains an estimation error or if the acoustic transfer characteristic cannot be obtained, the accuracy of the acoustic signal enhancement deteriorates.

[0006] Therefore, an object of the present invention is to provide an acoustic signal enhancement device that can accurately suppress time-varying unwanted sounds even when the estimated value of the acoustic transfer characteristic contains an estimation error or when the acoustic transfer characteristic cannot be obtained. [Means for solving the problem]

[0007] The acoustic signal enhancement device of the present invention is a device that receives frequency-divided recorded sound as input and updates parameters, and includes a beamformer unit, a switch unit, and a weighted spatial covariance estimator. The switch weights are weights that indicate the proportion of the recorded sound at each time point in a classification of temporally changing spatial states of the recorded sound to which category it belongs. The beamformer unit performs beamforming processing based on the updated weighted spatial covariance matrix and updates an auxiliary estimate of the target sound. The switch unit updates the switch weights and the power of the target sound based on the updated auxiliary estimate and outputs an estimate of the target sound. The weighted spatial covariance estimator updates the weighted spatial covariance matrix based on the updated switch weights and power. [Effects of the Invention]

[0008] According to the acoustic signal enhancing device of the present invention, even when the estimated value of the acoustic transfer characteristic contains an estimation error or when the acoustic transfer characteristic cannot be obtained, it is possible to suppress time-varying unwanted sounds with high precision. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing the configuration of an acoustic signal enhancement device disclosed in Non-Patent Document 1. [Figure 2] FIG. 1 is a block diagram showing the configuration of an acoustic signal enhancement device of Non-Patent Document 2. [Figure 3] FIG. 1 is a block diagram showing the configuration of an acoustic signal enhancement device according to a first embodiment. [Figure 4] 3 is a flowchart showing the operation of the acoustic signal enhancing device according to the first embodiment. [Figure 5] FIG. 2 is a block diagram showing the configuration of a switching beam former unit according to the first embodiment. [Figure 6] 10 is a flowchart showing the operation of a switching beam former unit according to the first embodiment. [Figure 7] FIG. 10 is a block diagram showing the configuration of an acoustic signal enhancing device according to a second embodiment. [Figure 8] 10 is a flowchart showing the operation of the acoustic signal enhancing device according to the second embodiment. [Figure 9]FIG. 10 is a block diagram showing the configuration of an acoustic signal enhancing device according to a third embodiment. [Figure 10] 10 is a first flowchart showing the operation of the acoustic signal enhancing device according to the third embodiment. [Figure 11] 10 is a second flowchart showing the operation of the acoustic signal enhancing device according to the third embodiment. [Figure 12] FIG. 10 is a block diagram showing the configuration of an acoustic signal enhancing device according to a fourth embodiment. [Figure 13] 10 is a flowchart showing the operation of the acoustic signal enhancing device according to the fourth embodiment. [Figure 14] FIG. 2 is a diagram showing an example of the functional configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described in detail. Components having the same functions are given the same numbers, and duplicated explanations will be omitted. [Example]

[0011] Hereinafter, signals to be suppressed by the audio signal enhancing device (noise, reverberation, and other target sounds in each target sound estimation) will be collectively referred to as unwanted sounds.

[0012] The functional configuration of the target sound enhancement device of Example 1 will be described below with reference to Fig. 3. As shown in the figure, the target sound enhancement device 1 of this example includes a reverberation suppression unit 11, a second switch unit 12, a switching beamformer unit 13, and a weighted spatiotemporal covariance estimation unit 14, and is a device that receives as input a recorded sound that has been frequency-divided using a short-time Fourier transform or the like and an estimated value of the acoustic transfer characteristics of the target sound, and repeatedly updates parameters until a predetermined stopping condition is met.

[0013] In the following description, the same processing is performed individually for each frequency, so the frequency number f of all symbols will be omitted.

[0014] <Filter configuration> The dereverberation unit 11

number

number

[0015] However, x t (x is bold, t is italic) is the recorded sound vector at time t (t is italic), x - t (x is in bold, t is in italics) is a time series vector of past recorded sounds from time t-L+1 to time tD (L is the filter order, D is the predicted delay of the dereverberation process), G t ∈C M(L-D)×M are dereverberation filters (G is bold, t is italic, C M(L-D)×M is the set of M(LD) × M-dimensional complex matrices, where M is the number of microphones), W t ∈C M×N is the noise suppression filter (W is bold, t is italic, C M×N is the set of all M×N-dimensional complex matrices), G t and W t Combined, the current recorded sound vector x t (x is bold, t is italic) and the vector x of the past recordings - t Convolutional BeamFormer (CBF) applied to the time series (x in bold), (·) H represents the conjugate transpose of a matrix.

[0016] The filter coefficients of equations (1) and (2) are further realized by the weighted sum of a plurality of coefficients as in equation (3).

number

[0017] <Optimization criteria> Estimated target sound y n,t As shown in equation (4), the mean is 0 and the variance is λ n,t Assume that the distribution follows a complex Gaussian distribution.

number

number

number

[0018] That is, the parameters (all filter coefficients, switch weights, power of each target sound (= variance of complex Gaussian distribution)) that maximize this likelihood function are found.

[0019] <Optimization method> Since there is no known method for finding the parameters that maximize equation (7) in a closed form, optimization is performed by repeating the process of updating individual parameters in turn (while fixing other parameters).

[0020] <Process flow: Initialization> Power of each target sound λ n,t The recorded sound is dereverberated using the conventional weighted prediction error minimization (WPE) method (Reference Non-Patent Document 1), and is initialized with the power of each target sound calculated using a minimum power distortionless response beamformer (Reference Non-Patent Document 2). Note that the method for initializing the power of each target sound is not limited to the above, and any method can be used.

[0021] (Reference non-patent document 1: Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, Biing-Hwang, Speech dereverberation based on variance-normalized delayed linear prediction, IEEE Trans. Audio, Speech, and Language Processing, vol. 18, no. 7, pp. 1717-1731, 2010.) (Reference Non-Patent Document 2: Livnat Ehrenberg, Sharon Gannot, Amir Leshem, Ephraim Zehavi, Sensitivity analysis of MVDR and MPDR beamformers, Proc. IEEE Convention of Electrical and Electronics Engineers in Israel, 2010) Furthermore, all switch weights are initialized with random numbers.

[0022] <Processing flow: Repeated processing> The following process is repeated until convergence (or a certain number of times).

[0023] [Weighted spatiotemporal covariance estimator 14] The weighted spatio-temporal covariance estimator 14 updates the weighted spatio-temporal covariance matrix based on the first switch weight, the second switch weight, and the power (S14). More specifically, the weighted spatio-temporal covariance estimator 14 calculates the weighted spatio-temporal covariance matrix R n,i,j , P n,i,j Update (R,P are in bold, n,i,j are in italics).

number

[0024] [Dereverberation section 11] The dereverberation unit 11 performs dereverberation processing on the recorded sound, executes beamforming processing based on the updated weighted spatiotemporal covariance matrix, and updates the auxiliary dereverberated sound of the target sound (S11). More specifically, the dereverberation unit 11 calculates each filter coefficient G i Update (1≦i≦I).

number

number

number

number

[0025] The switching beamformer unit 13 outputs the updated dereverberated sound z t (z is bold, t is italic) and repeat the following process a certain number of times for each target sound n.

[0026] [Weighted spatial covariance estimator 133] The weighted spatial covariance estimator 133 calculates the spatial covariance matrix Σ for each output (1≦j≦J) of the beamformer using equation (16). n,j (n, j are in italics) is updated (S133).

number

[0027] By feeding back the switch weight and the power of the target sound to weighted spatial covariance update unit 133, it is possible to simultaneously consider and optimize the viewpoint of whether the sound is background sound or target sound (effect of the audio model) and the viewpoint of how the background sound is spatially distributed (effect of the first switch), and since it is possible to classify the spatial distribution of the background sound centered on the background sound section, it is possible to accurately suppress unnecessary sounds that change over time without being affected much by errors even if the estimated value of the acoustic transfer characteristics of the target sound contains errors.

[0028] A speech model consisting of time-varying power is used to distinguish whether or not each time frame contains a target sound. Specifically, a spatial covariance matrix is ​​calculated based on the maximum likelihood method, weighted by the inverse of the speech power, to obtain a spatial covariance matrix that primarily emphasizes noise sections. By estimating a beamformer using this spatial covariance matrix, it is possible to accurately minimize the noise power (even if the estimated acoustic transfer characteristics of the target sound contain errors).

[0029] Furthermore, the larger the eigenvalue of Σ in equation (16), the more the beamformer is optimized to weaken the corresponding direction, and if the spatial covariance with respect to the estimated value of the power of the target sound has a large value, it is considered to be noise and is updated to weaken it.

[0030] [Beamformer 131] The beam former 131 calculates each filter coefficient w n,j (1≦j≦J) is updated (S131).

number

number

[0031] [Modification of the beam former unit 131] In Reference Non-Patent Document 3, the beamformer estimation in the form of Equation (17) is performed by using the acoustic transfer characteristic h n It is disclosed that the above can be transformed into the following form, which does not require the above.

number

[0032] Methods for determining the spatial covariance matrix Φn of the target audio from the recorded audio are disclosed in, for example, Non-Patent Documents 3, 4, and 5.

[0033] (Reference Non-Patent Document 3: M. Souden, J. Benesty, S. Affes, “On optimal frequency-domain multichannel linear filtering for noise reduction,” IEEE Transactions on Audio, Speech, and Language Processing, 18 (2), pp. 260-276, 2010.) (Reference non-patent document 4: J. Heymann, L. Drude, C. Boeddeker, P. Hanebrink, R. Haeb-Umbach, “BEAMNET: END-TO-END TRAINING OF A BEAMFORMER-SUPPORTED MULTI-CHANNEL ASR SYSTEM,” Proc. ICASSP, pp. 5325-5329, 2017.) (Reference non-patent document 5: Takuya Yoshioka, Nobutaka Ito, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Masakiyo Fujimoto, Chengzhu Yu, Wojciech J Fabian, Miquel Espi, Takuya Higuchi, Shoko Araki, Tomohiro Nakatani, “The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices,” Proc. 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), pp. 436-443, 2015.) When the modified example of the beam former unit 131 is used, the target sound emphasis device does not need to receive an estimated value of the acoustic transfer characteristic as an input.

[0034] [First switch unit 132] The first switch unit 132 uses the first switch weight δ of each output (1≦j≦J) of the beamformer in equation (19). n,j,t The first switch unit 132 is used to classify the background sound in each time frame into several spatial states (such as from which direction louder noise is heard) and to estimate a different beamformer for each state.

number

number

number

[0035] The functional configuration of the target sound enhancement device of Example 2 will be described below with reference to Fig. 7. As shown in the figure, the target sound enhancement device 2 of this example includes a beam former 21, a first switch 22, and a weighted spatial covariance estimation unit 23, and has the same configuration as the switching beam former 13 in Example 1. The target sound enhancement device 2 receives as input a recorded sound that has been frequency-divided using a short-time Fourier transform or the like, and an estimated value of the acoustic transfer characteristic of the target sound, and repeatedly updates parameters until a predetermined stopping condition is met.

[0036] <Filter configuration> The beam former 21 calculates the dereverberation signal z t Recorded sound x t The filter coefficients of equation (2) are further realized by the weighted sum of multiple coefficients as shown in equation (3).

[0037] w in equation (3) n,j (w is bold, n,j are italic) and δ n,j,t (Italics) are the filter coefficients of the jth beamformer for the nth target sound and the first switch weights at time t.

[0038] <Optimization criteria> The estimated target sound has a mean of 0 and a variance of λ as shown in Equation (4). n,tIt is assumed that the filter follows a complex Gaussian distribution. Under the assumptions of Equation (4), Equation (5), and Equation (6), the likelihood function of Equation (7) is the criterion for optimizing the audio signal enhancement process. h in Equation (7) n is an estimate of the acoustic transfer characteristic of the nth target sound. In other words, we find the parameters (all filter coefficients, switch weights, and power of each target sound (= variance of complex Gaussian distribution)) that maximize this likelihood function.

[0039] <Optimization method> Since there is no known method for finding the parameters that maximize equation (7) in a closed form, optimization is performed by repeating the process of updating individual parameters in turn (while fixing other parameters).

[0040] <Process flow: Initialization> Power of each target sound λ n,t The recorded sounds are initialized with the power of each target sound calculated using a conventional minimum power distortionless response beamformer (Reference Non-Patent Document 2). Furthermore, all switch weights are initialized with random numbers.

[0041] <Processing flow: Repeated processing> The following process is repeated until convergence (or a certain number of times).

[0042] [Weighted spatial covariance estimator 23] The weighted spatial covariance estimator 23 updates the weighted spatial covariance matrix based on the updated switch weights and powers (S23). More specifically, the weighted spatial covariance estimator 23 calculates the spatial covariance matrix Σ for each output (1≦j≦J) of the beamformer in equation (16). n,j Update.

[0043] [Beamformer 21] The beamformer unit 21 performs beamforming processing based on the updated weighted spatial covariance matrix, and updates the auxiliary estimated value of the target sound (S21). More specifically, the beamformer unit 21 updates each filter coefficient w n,jThe beamformer unit 21 updates each auxiliary estimated value y j,t is updated using equation (18).

[0044] [First switch unit 22] The first switch unit 22 updates the switch weights and the power of the target sound based on the updated auxiliary estimate, and outputs the target sound estimate (S22). More specifically, the first switch unit 22 updates the first switch weights δ n,j,t Update.

[0045] The first switch unit 22 calculates the estimated value y n,t Update.

[0046] The first switch unit 22 uses the power λ of the target sound in equation (21). n,t The first switch unit 22 updates the estimated value y n,t Output. [Example]

[0047] <Symbol conversion> In the following examples, δ t,f (j) is the first switch weight for the jth "output of the separation matrix" at (time, frequency) = (t, f). Also, β t,f (i,j) is β t,f (i,j) =γ t,f (i) δ t,f (j) The fusion switch weights are set to satisfy the following.

[0048] <Features of the Acoustic Signal Enhancement Device of the Third Embodiment> The acoustic signal enhancing device of this embodiment is capable of highly accurate estimation even when an estimated value of the acoustic transfer characteristic cannot be obtained in advance (=blind processing).

[0049] Furthermore, in order to realize blind processing, an optimization criterion different from that of the previous embodiments is used.

[0050] The acoustic signal enhancement device of this embodiment estimates N target sounds and MN noise components simultaneously. That is, it processes the problem as dereverberation + sound source separation. Accordingly, the beamformer has the following configuration.

[0051] The estimation target is a separation matrix consisting of N beamformers that estimate the target sound and MN beamformers that estimate the noise components.

[0052] All beamformers included in the separation matrix are switched simultaneously. In the first and second embodiments, the beamformer is switched independently for each target sound.

[0053] <Filter configuration> Dereverberation processing is performed according to equation (22).

number

[0054] Although equation (22) is substantially the same as equation (1), in this embodiment, the frequency f needs to be expressed individually, so the above expression is used. The same applies to the following equations.

[0055] Beamformer processing for sound source separation is performed according to equation (23).

number

[0056] The filter coefficients of equations (22) and (23) are further realized by the weighted sum of a plurality of coefficients as in equation (24) (similar to the first embodiment).

number

[0057] β in Equation (25) t,f (i,j) (=γ t,f (i) δ t,f (j) ) is the switch weight for the i-th dereverberation filter and the j-th separation matrix at time t and frequency f. t,f (i,j) are all γ t,f (i) δ t,f (j) may be substituted for the calculation.

[0058] Using equation (24), y obtained from equations (22) and (23) t,f can be calculated as follows:

number

[0059] <Optimization criteria> The estimated sound sources are mutually independent as shown in equation (26),

number

number

number

number

[0060] <Optimization method> Since there is no known method for finding the parameters that maximize Equation (28) in a closed form, optimization is performed by repeating the process of updating each parameter in turn (while fixing other parameters).

[0061] The functional configuration of the target sound enhancement device 3 of this embodiment will be described below with reference to Fig. 9. As shown in the figure, the target sound enhancement device 3 of this embodiment includes a reverberation suppression unit 11, a beamformer unit 32, a switch unit 33, a weighted spatial covariance estimator 34, and a weighted spatio-temporal covariance estimator 35. The operation of the target sound enhancement device 3 (first flowchart) will be described below with reference to Fig. 10.

[0062] <Process flow: Initialization> The target sound emphasis device 3 calculates the power λ of each target sound. n,t,f , filter coefficient G f (i) , W f (j) is initialized with the power and filter coefficients (common to all switches) of each separated sound calculated using a conventional blind convolution beamformer (Reference Non-Patent Document 6) for the recorded sound, and all switch weights are initialized with random numbers (S30).

[0063] (Reference non-patent document 6: Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Hiroshi Sawada, Computationally efficient and versatile framework for blind speech separation and dereverberation, Proc. Interspeech, pp. 91-95, 2020.) <Processing flow: Repeated processing until convergence condition is reached> The target voice emphasis device 3 repeats the following processes (S35, S11, execution of the second flowchart) until convergence occurs.

[0064] <Processing flow: Weighted spatiotemporal covariance estimation> The weighted spatiotemporal covariance estimator 35 calculates a weighted spatiotemporal covariance matrix R n,f (i,j) ,P n,f (i,j) Update (S35).

number

number

[0065] <Processing flow: Weighted spatial covariance estimation> The weighted spatial covariance estimator 34 calculates the weighted spatial covariance matrix Σ for each sound source included in the output (1≦j≦J) of each separation matrix using equation (36). n,f (j) Update (S34).

number

number

number

[0066] In source separation, the signal power λ n,t,fIt has been shown that by making σ take a common value across all frequencies, the order of sound sources separated at different frequencies can be aligned (see, for example, Non-Patent Document 7).

[0067] (Reference Non-Patent Document 7: Nobutaka Ono and Shigeki Miyabe, Auxiliary-function-based independent component analysis for super-Gaussian sources, in LVA / ICA. Springer, pp. 165-172, 2010.) In the present invention, this method can be used in the following procedure.

[0068] The weighted spatial covariance estimator calculates the frequency average λ of the power of each signal based on equation (43). n,t Ask for.

number

[0069] In the third embodiment, the first and second switch weights are updated simultaneously after updating the filter coefficients for both dereverberation and sound source separation. However, the switch weights do not necessarily have to be updated at this timing, and it is not necessary to update the two switches simultaneously. For example, the following configuration is also possible.

[0070] After updating the dereverberation filter coefficients, update both switch weights or only the second switch weight.

[0071] After updating the filter coefficients for source separation, update both switch weights or only the first switch weight.

[0072] At any timing, the switch weights may be updated based on the criterion of maximizing the likelihood function, with other parameters being fixed.

[0073] <Functional configuration of the target sound enhancement device 4 according to the fourth embodiment> As shown in FIG. 12, the target sound enhancement device 4 of this embodiment includes a beam former unit 32, a switch unit 43, and a weighted spatial covariance estimation unit .

[0074] <Changes from Example 3> Skip dereverberation processing and perform blind source separation.

[0075] Dereverberation Filter G f (i) and the second switch weight γ t,f (i) has been deleted.

[0076] The dereverberation unit 11 and the weighted spatiotemporal covariance estimation unit 35 are omitted.

[0077] The beamformer 32 and the weighted spatial covariance estimator 34 receive the auxiliary dereverberated sound z t,f (i) Not the recorded sound x t Enter.

[0078] The switch unit 43 skips the estimation process of the second switch weight.

[0079] <Optimization criteria> The filter configuration is the same as in Example 3. However, the likelihood function in equations (28) and (29) is G f (i) and γ t,f (i) For example,

number

number

[0080] <Optimization method> Other than adopting the above filter configuration, this is the same as Example 3.

[0081] Hereinafter, the operation of the target voice emphasis device 4 will be described with reference to FIG.

[0082] <Process flow: Initialization> The target sound emphasis device 4 calculates the power λ of each target sound. n,t,f , filter coefficient W f (j) is initialized with the power and filter coefficients (common to all switches) of each separated sound obtained using a conventional blind source separation method (reference non-patent document 7) for the recorded sound, and all switch weights are initialized with random numbers (S40).

[0083] <Processing flow: Repeated processing until convergence condition (or a certain number of times) is met> The target voice emphasis device 4 repeats the following processes (S34, S32, S43) until convergence occurs (or a certain number of times).

[0084] <Processing flow: Weighted spatial covariance estimation> The weighted spatial covariance estimator 34 calculates the weighted spatial covariance matrix Σ for each sound source included in the output (1≦j≦J) of each separation matrix using equation (36). n,f (j) Update (S34).

[0085] <Processing flow: Beamformer processing> The beam former 32 calculates each filter coefficient w n,f (j) (1≦n≦M, 1≦j≦J) and update the auxiliary estimates y t,f (i,j) is updated using equation (39) (S32).

[0086] <Processing flow: Switch processing> The switch unit 43 calculates the estimated value yt,f After updating, the power λ of each sound source is calculated using equation (40). n,t,f (1≦n≦M) is updated, and the first switch weight is updated by equation (41) (more specifically, equation (44) below) (S43).

number

[0087] <Experiment> When audio signal enhancement processing was applied to the audio recorded with three microphones of two people speaking simultaneously in a noisy and reverberant environment, the following experimental results were obtained. It can be seen that the audio signal enhancement devices of Examples 1 and 3 have higher accuracy than the conventional method (Non-Patent Document 2). [Table 1] <Effects> According to the acoustic signal enhancement device 1 of the first embodiment, the switch weights, the power of the target sound, the coefficients of the dereverberation processing, and the coefficients of the beamformer are optimized by iterative processing based on the criterion that the power of the target sound follows a Gaussian distribution that changes over time. Therefore, even if the acoustic transfer characteristics of the target sound contain errors or the recorded sound contains reverberation, it is possible to accurately suppress unnecessary sounds that change over time.

[0088] According to the acoustic signal enhancement device 2 of the second embodiment, the switch weights, the power of the target sound, and the coefficients of each beamformer are optimized by iterative processing based on the criterion that the power of the target sound follows a Gaussian distribution that changes over time. Therefore, even if the estimated value of the acoustic transfer characteristic contains an estimation error, the time-varying unwanted sound can be suppressed with high accuracy.

[0089] In addition, optimization can be performed by simultaneously considering the perspective of whether the sound is background sound or target sound (effectiveness of the audio model) and the perspective of how the background sound is spatially distributed (effectiveness of the first switch).

[0090] As a result, the spatial distribution of the background sound can be classified around the background sound section, so that even if the acoustic transfer characteristics of the target sound contain errors, the time-varying unnecessary sounds can be accurately suppressed without being affected much by the errors.

[0091] <Additional Notes> The device of the present invention may, for example, be a single hardware entity having an input unit to which a keyboard or the like can be connected, an output unit to which an LCD display or the like can be connected, a communication unit to which a communication device (e.g., a communication cable) capable of communicating with an external device can be connected, a CPU (which may also include a central processing unit, cache memory, registers, etc.), memories such as RAM and ROM, an external storage device such as a hard disk, and buses connecting these input unit, output unit, communication unit, CPU, RAM, ROM, and external storage device so that data can be exchanged between them. If necessary, the hardware entity may also be provided with a device (drive) capable of reading and writing to a recording medium such as a CD-ROM. A physical entity equipped with such hardware resources includes a general-purpose computer.

[0092] The external storage device of the hardware entity stores the programs required to realize the above-mentioned functions and the data required for processing these programs (the programs may be stored in a ROM, which is a read-only storage device, for example, instead of an external storage device). Data obtained by processing these programs is stored in RAM, the external storage device, etc. as appropriate.

[0093] In a hardware entity, each program stored in an external storage device (or ROM, etc.) and the data required to process each program are loaded into memory as needed, and interpreted, executed, and processed by the CPU as appropriate, resulting in the CPU realizing a predetermined function (each component represented as a unit, means, etc., above).

[0094] The present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of the present invention. Furthermore, the processes described in the above embodiments may not only be executed in chronological order according to the order described, but may also be executed in parallel or individually depending on the processing capacity of the device that executes the processes or as needed.

[0095] As described above, when the processing functions of the hardware entities (apparatuses of the present invention) described in the above embodiments are realized by a computer, the processing contents of the functions that the hardware entities should have are described by a program. Then, by executing this program on a computer, the processing functions of the hardware entities are realized on the computer.

[0096] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 10020 of the computer shown in Figure 14 and operating the control unit 10010, input unit 10030, output unit 10040, etc.

[0097] The program describing the processing contents can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include magnetic recording devices, optical disks, magneto-optical recording media, and semiconductor memories. Specifically, examples of magnetic recording devices include hard disk drives, flexible disks, and magnetic tapes; optical disks include DVDs (Digital Versatile Discs), DVD-RAMs (Random Access Memory), CD-ROMs (Compact Disc Read Only Memory), and CD-Rs (Recordable) / RWs (Rewritable); magneto-optical recording media include MOs (Magneto-Optical discs), and semiconductor memories include EEP-ROMs (Electrically Erasable and Programmable-Read Only Memory).

[0098] The program may be distributed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to another computer via a network, thereby distributing the program.

[0099] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the received program each time a program is transferred from a server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. In this embodiment, the program includes information used for processing by a computer that is equivalent to a program (such as data that is not a direct instruction to the computer but has properties that define computer processing).

[0100] In addition, in this embodiment, a hardware entity is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.

Claims

1. An audio signal enhancement device that receives frequency-divided recorded sound as an input and updates parameters, The switch weight is a weight that indicates the proportion of the recorded sound at each time point in the classification of the spatial state of the recorded sound that changes over time, to which category the recorded sound belongs, a beamformer unit that receives an estimate of the acoustic transfer characteristics of the target sound as an input, performs beamforming processing based on the updated weighted spatial covariance matrix, and updates an auxiliary estimate of the target sound; a switch unit that updates the switch weight and the power of the target sound based on the updated auxiliary estimate, and outputs an estimate of the target sound; a weighted spatial covariance estimator that updates the weighted spatial covariance matrix based on the updated switch weights and the powers; Acoustic signal enhancement device.

2. An audio signal enhancement device that receives frequency-divided recorded sound as an input and updates parameters, The first switch weight is a weight indicating the proportion of the recorded sound at each time point in the classification of the spatial state of the recorded sound that changes over time to which category, The second switch weight is a weight indicating the proportion of the recorded sound at each time point in the classification of the time-space state of the recorded sound that changes over time to which category, a dereverberation unit that performs dereverberation processing on the recorded sound based on the updated weighted spatio-temporal covariance matrix and updates an auxiliary dereverberated sound of the target sound; a switch unit that updates and outputs the second switch weights and the dereverberated sound based on the auxiliary dereverberated sound, the updated power of the target sound, and the updated beamformer coefficients; a switching beamformer unit that receives an estimate of acoustic transfer characteristics of a target sound as input, updates the estimate of the target sound, the beamformer coefficients, the power of the target sound, and the first switch weight of the target sound based on the updated dereverberation-reduced sound, and outputs the estimate of the target sound; a weighted space-time covariance estimator that updates the weighted space-time covariance matrix based on the first switch weight, the second switch weight, and the power; Acoustic signal enhancement device.

3. 3. The acoustic signal enhancement device according to claim 2, The switching beam former unit a beamformer unit that receives an estimate of the acoustic transfer characteristics of the target sound as an input, performs beamforming processing based on the updated weighted spatial covariance matrix, and updates an auxiliary estimate of the target sound; a first switch unit that updates the first switch weight and the power of the target sound based on the updated auxiliary estimate value and outputs an estimate value of the target sound; a weighted spatial covariance estimator that updates the weighted spatial covariance matrix based on the updated first switch weights and the power; Acoustic signal enhancement device.

4. An audio signal enhancement device that receives input of sounds recorded by a plurality of microphones, the acoustic signal enhancement device acquires an initial value of the power of each sound source based on the recorded sound; The first switch weight is a weight indicating the proportion of the recorded sound at each time point in the classification of the spatial state of the recorded sound that changes over time to which category, a weighted spatial covariance estimator that updates a weighted spatial covariance matrix for estimating a coefficient for determining a target sound of a beamformer based on the first switch weight, the power of each sound source or an initial value of the power of each sound source, and a recorded sound; A beamformer unit that updates beamformer coefficients that estimate separated sounds of a separation matrix based on the weighted spatial covariance matrix, and updates auxiliary estimates of each sound source based on the updated beamformer coefficients and the recorded sounds; a switch unit that acquires the beamformer coefficients, updates an estimate of all sound sources based on the auxiliary estimate of each sound source and the first switch weight, updates the power of each sound source based on the estimate of all sound sources, updates the first switch weight based on the power of each sound source, and outputs an estimate of each sound source. Acoustic signal enhancement device.

5. An audio signal enhancement device that receives input of sounds recorded by a plurality of microphones, The first switch weight is a weight indicating the proportion of the recorded sound at each time point in the classification of the spatial state of the recorded sound that changes over time to which category, The second switch weight is a weight indicating the proportion of the recorded sound at each time point in the classification of the time-space state of the recorded sound that changes over time to which category, a weighted spatial covariance estimator that updates a weighted spatial covariance matrix for estimating coefficients for determining a target sound of a beamformer based on the first and second switch weights, the power of each sound source, and an auxiliary dereverberation-reduced sound; a beamformer unit that updates beamformer coefficients that estimate separated sounds of a separation matrix based on the weighted spatial covariance matrix, and updates auxiliary estimates of each sound source based on the updated beamformer coefficients and the auxiliary dereverberation sound; a switch unit that acquires the beamformer coefficients, updates an estimate of all sound sources based on the auxiliary estimate of each sound source and the first and second switch weights, updates a power of each sound source based on the estimate of all sound sources, updates the first switch weight and the second switch weight based on the power of each sound source, and outputs an estimate of each sound source; a weighted spatio-temporal covariance estimator that updates a weighted spatio-temporal covariance matrix for estimating filter coefficients for dereverberation processing based on the first and second switch weights and the power of each of the sound sources; a dereverberation unit that updates a filter coefficient for dereverberation processing based on the beamformer coefficients and the weighted spatio-temporal covariance matrix, and updates the auxiliary dereverberated sound. Acoustic signal enhancement device.

6. An audio signal enhancement method executed by an audio signal enhancement device that receives frequency-divided recorded sound as input and updates parameters, comprising: The switch weight is a weight that indicates the proportion of the recorded sound at each time point in the classification of the spatial state of the recorded sound that changes over time, to which category the recorded sound belongs, a beamforming step of receiving an estimate of the acoustic transfer characteristics of the target sound, performing beamforming processing based on the updated weighted spatial covariance matrix, and updating an auxiliary estimate of the target sound; a switching step of updating the switch weights and the power of the target sound based on the updated auxiliary estimate and outputting an estimate of the target sound; a weighted spatial covariance estimation step of updating the weighted spatial covariance matrix based on the updated switch weights and the powers. Acoustic signal enhancement method.

7. An audio signal enhancement method executed by an audio signal enhancement device that receives frequency-divided recorded sound as input and updates parameters, comprising: The first switch weight is a weight indicating the proportion of the recorded sound at each time point in the classification of the spatial state of the recorded sound that changes over time to which category, The second switch weight is a weight indicating the proportion of the recorded sound at each time point in the classification of the time-space state of the recorded sound that changes over time to which category, a dereverberation step of performing a dereverberation process on the recorded sound, executing a beamforming process based on the updated weighted spatio-temporal covariance matrix, and updating an auxiliary dereverberated sound of the target sound; a switching step of updating and outputting the second switch weights and the dereverberated sound based on the auxiliary dereverberated sound, the updated power of the target sound, and the updated beamformer coefficients; a switching beamformer step of receiving an estimate of acoustic transfer characteristics of a target sound as an input, updating the estimate of the target sound, the beamformer coefficients, the power of the target sound, and the first switch weight of the target sound based on the updated dereverberation-reduced sound, and outputting the estimate of the target sound; a weighted space-time covariance estimation step of updating the weighted space-time covariance matrix based on the first switch weight, the second switch weight, and the power. Acoustic signal enhancement method.

8. A program that causes a computer to function as the acoustic signal enhancement device according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice signal enhancement system and method

    CN102938254A

  • Driving unit of el display device

    JP1977027393A

  • Method and device for noise component suppression processing method

    JP2001100800A

  • Signal processing device, signal processing method and program

    JP2008134298A

  • Signal processing system, signal processing method and program

    JP2018036526A