Signal processing device, signal processing method, and program
The MaxSNR criterion in the signal processing device optimizes dereverberation and beamforming by fully utilizing spatial information, addressing the limitations of MVDR-based CBF in sound source extraction.
Patent Information
- Application Number
- JP2024541328
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2025-11-26
- Estimated Expiration
- 2042-08-17
AI Technical Summary
Existing signal source extraction techniques, such as the convolutional beamformer (CBF) based on the Minimum-Variance Distortionless Response (MVDR) criterion, fail to utilize all spatial information of the target sound source due to the compression of spatial information into a steering vector.
Introduce the MaxSNR criterion to design the CBF, utilizing a signal processing device with units for estimating spatial and spatiotemporal covariance matrices to fully leverage the spatial information of the target sound source, enabling joint optimization of dereverberation and beamforming.
The MaxSNR criterion allows for the complete utilization of spatial information, improving sound source extraction performance by enhancing dereverberation and beamforming processes.
Smart Images

Figure 0007776016000040 
Figure 0007776016000041 
Figure 0007776016000042
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for estimating, with high quality, an audio signal contained in a signal recorded using a microphone. [Background technology]
[0002] When recording a speech signal using a microphone in a noisy reverberant environment, the quality of the speech signal contained in the recorded signal is low because the microphone contains unwanted components such as noise, reverberation, and interfering sounds in addition to the desired speech components. Therefore, signal source extraction techniques have been actively researched to estimate the speech signals contained in the recorded signal with high quality. A known method for signal source extraction using multiple sensors is the convolutional beamformer (CBF, see Non-Patent Document 1). The Minimum-Variance Distortionless Response (MVDR) criterion has been used to optimize the CBF (see Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] T. Nakatani, C. Boeddeker, K. Kinoshita, R. Ikeshita, M. Delcroix and R. Haeb-Umbach, "Jointly Optimal Denoising, Dereverberation, and Source Separation", IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2267-2282, 2020. Summary of the Invention [Problem to be solved by the invention]
[0004] However, when designing CBF based on the MVDR standard, the spatial information (spatial covariance matrix) of the target sound source to be extracted is compressed into a steering vector to design the CBF, which poses a problem in that it is not possible to use all of the spatial information possessed by the target sound source.
[0005] An object of the present invention is to provide a signal processing device, a signal processing method, and a program that can use all of the spatial information of a target sound source by introducing a MaxSNR criterion instead of the MVDR criterion. [Means for solving the problem]
[0006] In order to solve the above problems, according to one aspect of the present invention, a signal processing device includes: a second spatial covariance matrix estimation unit that estimates a spatial covariance matrix of a non-target sound source using an estimated value of the spatiotemporal covariance matrix of the non-target sound source; a dereverberation filter estimation unit that estimates a dereverberation filter using the estimated value of the spatiotemporal covariance matrix of the non-target sound source; a beamformer estimation unit that estimates a convolution beamformer using the estimated value of the spatial covariance matrix of an observed signal or the target sound source, the estimated value of the spatial covariance matrix of the non-target sound source, and the estimated dereverberation filter; and a sound source extraction unit that performs beamforming processing using the observed signal and the estimated convolution beamformer to estimate a sound source signal. [Effects of the Invention]
[0007] According to the present invention, by introducing the MaxSNR criterion, it is possible to obtain an effect that all spatial information of the target sound source can be used. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a functional block diagram of a signal processing device according to a first embodiment. [Figure 2] FIG. 2 is a diagram showing an example of a processing flow of the signal processing device according to the first embodiment. [Figure 3] FIG. 10 is a functional block diagram of a signal processing device according to a second embodiment. [Figure 4]FIG. 10 is a diagram showing an example of a processing flow of a signal processing device according to a second embodiment. [Figure 5] FIG. 10 is a functional block diagram of a signal processing device according to a third embodiment. [Figure 6] FIG. 11 is a diagram showing an example of a processing flow of a signal processing device according to a third embodiment. [Figure 7] FIG. 1 is a diagram showing an example of the configuration of a computer to which the present technique is applied. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described. In the drawings used in the following description, components having the same functions and steps performing the same processes will be denoted by the same reference numerals, and duplicated explanations will be omitted. In the following description, the symbols "^" and " - " etc. should normally be written directly above the character immediately following it, but due to limitations in text notation, they are written immediately before the character in question. In formulas, these symbols are written in their original positions. Furthermore, unless otherwise specified, processing performed on each element of a vector or matrix is assumed to apply to all elements of that vector or matrix.
[0010] <Sound source extraction problem> The problem addressed in this embodiment is a sound source extraction problem, in which a signal x observed by a microphone is f,t From the sound source signal s f,t Alternatively, the source signal s f,t The spatial image s with the reverberation removed f,t image =a f s f,t The problem is to estimate a frepresents the acoustic transfer function of the sound source. Note that the sound source signal is a signal based on the sound emitted by a sound source (target sound source) to be recorded by a microphone, and in this embodiment, the target sound source is a speaker (hereinafter also referred to as "target speaker"), the target sound is the voice spoken by the target speaker (hereinafter also referred to as "target voice"), and the target signal is a signal corresponding to the target voice. However, the target sound source is not limited to these, and may be any sound source other than a speaker, such as a sound source such as a musical instrument or a playback device, and the target sound is not limited to voice, but may also be a sound other than voice. Sound sources other than the target sound source are also called non-target sound sources.
[0011] <Key Points of the First Embodiment> The steering vector used by MVDR CBF is the spatial covariance matrix V S , and the MVDR CBF corresponds to the spatial covariance matrix V S In this embodiment, the MaxSNR criterion is introduced as a new criterion for designing the CBF. When designing the CBF using the MaxSNR criterion, the spatial information of the target sound source (the spatial covariance matrix V S ) has the advantage of being able to make full use of it.
[0012] First, we will explain the MaxSNR CBF. Let M be any integer equal to or greater than 2 that represents the number of microphones, L+1 be the number of taps of the CBF, and S + Let be the set of all non-negative definite matrices, and let A B Let be a square matrix with B rows and B columns, and let matrix A B×C Let be a matrix with B rows and C columns, and ^R N ∈S M+ML + Let be the spatiotemporal covariance matrix of the non-target sound source, and V S ∈S M + Let be the spatial covariance matrix of the target sound source, and O A×B Let be a zero matrix with A rows and B columns,
number
number
[0013] Furthermore, the MaxSNR CBF ^w of this embodiment has the feature that it can be decomposed into the product of the dereverberation filter ^G and the MaxSNR beamformer w for the instantaneous mixture, as shown in the following equation.
number
[0014] To explain that equation (2) can be decomposed as equation (3), ^w, ^R N is written as follows:
number
number
number
[0015] Here, the optimal solution of MaxSNR CBF ^w is opt of
number
number
number
number
[0016] Note that equation (7) can be solved as the optimal eigenvector of the generalized eigenvalue decomposition.
[0017] V S w opt = λ max V N w opt where λ max is the largest eigenvalue.
[0018] ^G in equation (8) is a multi-channel linear prediction (MCLP) based dereverberation filter used in dereverberation. Also, V in equation (9) N ^R N and can be considered as the spatial covariance matrix of the dereverberated non-target sound sources.
[0019] First Embodiment FIG. 1 is a functional block diagram of a signal processing device according to the first embodiment, and FIG. 2 shows the processing flow thereof.
[0020] The signal processing device 100 includes a first spatial covariance matrix estimator 110, a spatiotemporal covariance matrix estimator 120, a second spatial covariance matrix estimator 140, a dereverberation filter estimator 130, a beamformer estimator 150, a sound source extractor 160, and a spatial image estimator 170.
[0021] The signal processing device 100 receives an observed signal x observed by a microphone. f,t is input, and the sound source signal s f,t Alternatively, the source signal s f,t The (dereverberated) spatial image s f,t image =a f s f,t The observed signal is, for example, an acoustic signal observed by a microphone array consisting of multiple microphones. The output signal of the microphone may be input as is, or an output signal stored in some kind of storage device may be read and input, or the output signal of the microphone may be input after undergoing some processing. Note that f (f=1,...,F) indicates frequency, t (t=1,...,T) indicates frame number, and the observed signal x f,t , source signal s f,t is a frequency domain signal. However, the observed signal in the time domain is input and the observed signal in the frequency domain x f,t and the source signal s f,t The estimated value of may be converted into a time domain sound source signal in a time domain conversion unit (not shown) and output. The frequency domain conversion and the time domain conversion may be performed by any method, such as Fourier transform, inverse Fourier transform, etc.
[0022] The signal processing device 100 is a special device configured by loading a special program into a publicly known or dedicated computer having, for example, a central processing unit (CPU), a main memory (RAM), etc. The signal processing device 100 executes each process under the control of, for example, the central processing unit. Data input to the signal processing device 100 and data obtained by each process are stored in, for example, the main memory, and the data stored in the main memory is read out to the central processing unit as needed and used for other processes. At least a part of each processing unit of the signal processing device 100 may be configured by hardware such as an integrated circuit. Each storage unit included in the signal processing device 100 may be configured by, for example, a main storage unit such as a RAM (Random Access Memory), or middleware such as a relational database or a key-value store. However, each storage unit does not necessarily have to be included inside the signal processing device 100. It may be configured by an auxiliary storage unit including a hard disk, an optical disk, or a semiconductor memory element such as a flash memory, and be configured to be external to the signal processing device 100.
[0023] Each part will be explained below.
[0024] <First spatial covariance matrix estimation unit 110> The first spatial covariance matrix estimation unit 110 estimates the spatial covariance matrix of the target sound source (S110) and calculates the estimated value V S ∈S M + As a method for estimating the spatial covariance matrix of the target sound source, various methods can be used. For example, the first spatial covariance matrix estimation unit 110 outputs the observed signal x f,t is input, and the observed signal x f,t The section including the sound emitted by the target sound source (hereinafter also referred to as the target signal) is estimated from the target signal, and the spatial covariance matrix of the target sound source is estimated using the estimated target signal. In addition, if the direction of the target sound source is known, the spatial covariance matrix of the target sound source is approximated in advance by experiments or simulations, and the approximate value is used as the estimated value VS ∈S M + It may also be used as. <Spatio-temporal covariance matrix estimation unit 120> The space-time covariance matrix estimation unit 120 estimates the space-time covariance matrix of the non-target sound source (S120) and calculates the estimated value ^R N ∈S M+ML + As a method for estimating the spatiotemporal covariance matrix of the non-target sound source, various methods can be used. For example, the spatiotemporal covariance matrix estimation unit 120 outputs the observed signal x f,t is input, and the observed signal x f,t A section that does not include a sound emitted by the target sound source (hereinafter also referred to as a non-target signal) is estimated from the non-target signal, and the space-time covariance matrix of the non-target sound source is estimated using the estimated non-target signal. <Dereverberation filter estimation unit 130> The dereverberation filter estimation unit 130 estimates the spatiotemporal covariance matrix ^R N is used as input, and the estimated value ^R N The block matrices contained in - P N , - R N (S130), and outputs the estimated dereverberation filter ^G. For example, the dereverberation filter is estimated by equation (8).
number
number
number
[0025] <Second spatial covariance matrix estimation unit 140> The second spatial covariance matrix estimator 140 estimates the spatiotemporal covariance matrix ^R N is used as input, and the estimated value ^R N The block matrix R contained in N , - P N , - R N The spatial covariance matrix of the non-target sound source is estimated from (S140), and the estimated value V N ∈S M+ML + For example, the spatial covariance matrix of the non-target sound source is estimated by equation (9).
number
[0026] The second spatial covariance matrix estimator 140 estimates the spatiotemporal covariance matrix ^R N and the dereverberation filter ^G estimated by the dereverberation filter estimation unit 130 are input, and the estimated value ^R N The spatial covariance matrix of the non-target sound source may be estimated from the dereverberation filter ^G using equation (9).
[0027] <Beamformer estimation unit 150> The beamformer estimation unit 150 estimates the spatial covariance matrix V of the target sound source. S and the estimated spatial covariance matrix of the non-target sound source, V N and the estimated dereverberation filter ^G are input. The beamformer estimation unit 150 estimates the spatial covariance matrix V Sand the estimated spatial covariance matrix of the non-target sound source, V N From this, the MaxSNR beamformer w for the instantaneous mixture is opt Ask for.
number
[0028] V S w opt = λ max V N w opt where λ max is the largest eigenvalue.
[0029] The beamformer estimation unit 150 calculates the MaxSNR beamformer w for the instantaneous mixture. opt and the estimated dereverberation filter ^G, a convolution beamformer is estimated by equation (3) (S150), and the estimated convolution beamformer ^w is output.
number
[0030] y f,t =^w f H ^x f,t ∈ C ^w ∈ C M+ML ^w=[^w1| … | ^w F ] ^x f,t =[x f,t T |x f,t-D-1 T |…|xf,t-D-L T ] T ∈ C M+ML A H denotes the Hermitian transpose of A, and A T denotes the transpose of A, and Y=(y t ) t=1 T is an estimate of the source signal S and D is the prediction delay.
[0031] <Spatial image estimation unit 170> Convolutional beamformer ^w for each frequency bin f f Although the scale of is indefinite, the spatial image s f,t image A vector u that approximates f can be restored by estimating
[0032] s f,t image =a f s f,t ≒u f y f,t =(u f w f H )(^G f H ^x f,t )∈C M ^G=[^G1| … | ^G F ], where the vector u f is required to meet the following conditions:
[0033] (i) w f H u f =1 (no distortion constraint) (ii) u f ∝V N,f w f (Ideally a f ∝V N,f w f holds) V N =[V N,1 | … | V N,F ], and the two constraints make the vector u fis uniquely determined as follows:
number
[0034] s f,t image ≒u f y f,t <Effects> With the above configuration, by introducing the MaxSNR criterion, it is possible to use all the spatial information of the target sound source.
[0035] <Key Points of the Second Embodiment> The MVDR CBF (a method for estimating CBF based on the MVDR standard) requires a steering vector of the target sound source to be estimated separately in advance, which causes problems such as the sound source extraction performance of the MVDR CBF being highly dependent on the steering vector estimation performance and poor usability. This embodiment solves these problems.
[0036] To estimate the MaxSNR CBF, the estimated spatial covariance matrix V of the target sound source is used as shown in equations (1) and (2). S and the estimated spatiotemporal covariance matrix of the non-target sound source, ^R N must be requested in advance.
number
number
[0037] The Blind MaxSNR CBF of this embodiment is a method for estimating MaxSNR CBF by repeatedly performing calculations similar to the MaxSNR CBF given by equation (2) or equation (7).
[0038] The Blind MaxSNR CBF of this embodiment is a function of an arbitrary super-Gaussian function φ:R ≧0 →R and the following matrix ^R X Schur complement matrix V X Using the above, the blind MaxSNR CBF is defined as the following local optimum solution (Equations (20a) and (20b)).
number
number
number
[0039] More specifically, the estimate of the spatiotemporal covariance matrix of the non-target sound source, ^R N,f The space-time covariance matrix ^R, interpreted as Z,f The MaxSNR CBF is optimized without prior knowledge through iterative optimization that alternately repeats a process of calculating MaxSNR CBF ^w based on the following equations (21) and (22) and a process of estimating MaxSNR CBF ^w based on the following equations (23) to (26).
number
number
number
number
number
[0040] In addition, in each iteration of the above iterative optimization, MaxSNR CBF ^w is calculated for each frequency f=1,...,F based on the following equation (27). f The feature of this method is that the scales are aligned.
[0041] w f ←(u f,m ) * w f =(e m T u f ) * w f (27) where m (1≦m≦M) is the index of the reference microphone, * denotes the complex conjugate, and u f is expressed by equation (11) (where V N,f V instead of Z,f (using u f,m =e m T u f ∈C is u f is the mth element of
[0042] Second Embodiment The following description will focus on the differences from the first embodiment.
[0043] FIG. 3 is a functional block diagram of the signal processing device according to the first embodiment, and FIG. 4 shows the processing flow thereof.
[0044] The signal processing device 200 includes an initialization unit 201, a first spatial covariance matrix estimation unit 210, a space-time covariance matrix estimation unit 220, a second spatial covariance matrix estimation unit 240, a dereverberation filter estimation unit 230, a beamformer estimation unit 250, a sound source extraction unit 160, and a determination unit 280.
[0045] The signal processing device 200 receives an observed signal x observed by a microphone. f,t and the index m of the reference microphone are input, and the source signal s f,t where f is the frequency, t is the frame number, and the observed signal x f,t , source signal s f,t is a frequency domain signal. However, the observed signal in the time domain is input and the observed signal in the frequency domain x f,t and the source signal s f,tmay be converted into a time domain sound source signal in a time domain conversion unit (not shown) and output. The frequency domain conversion and the time domain conversion may be performed by any method, such as Fourier transform or inverse Fourier transform.
[0046] <Initialization section 201> The initialization unit 201 receives the index m of the reference microphone and calculates the initial value ^w of the convolution beamformer ^w to be estimated. 0 =[^w1 0 ,…,^w F 0 ] is set by the following formula (S201) and output.
number
number
number
number
number
[0047] <Dereverberation filter estimation unit 230> The dereverberation filter estimation unit 230 estimates the spatial-temporal covariance matrix ̂R Z,f is used as input, and the estimated value ^R Z,f Included in - P Z,f , - R Z,f (S230) and outputs the estimated dereverberation filter ^G. For example, the dereverberation filter is estimated by equation (25).
number
number
number
[0048] <Second spatial covariance matrix estimation unit 240> The second spatial covariance matrix estimation unit 240 estimates the spatial-temporal covariance matrix ^R Z,f is used as input, and the space-time covariance matrix ^R Z,f R included in Z,f , - P Z,f , - R Z,f The spatial covariance matrix of the non-target sound source is estimated from (S240), and the estimated value V Z,f ∈S M+ML + For example, the spatial covariance matrix of the non-target sound source is estimated by Equation (31).
number
[0049] The second spatial covariance matrix estimation unit 240 estimates the spatial-temporal covariance matrix ^R Z,f and the dereverberation filter ^G estimated by the dereverberation filter estimation unit 230 are input, and the spatiotemporal covariance matrix ^R Z,f The spatial covariance matrix of the non-target sound source may be estimated from the dereverberation filter ^G using equation (31).
[0050] <Beamformer estimation unit 250> The beamformer estimator 250 estimates the observed signal x f,t The estimated spatial covariance matrix V X =[V X,1 ,…,V X,F ] and the estimated spatial covariance matrix of the non-target sound source V Z =[V Z,1 ,…,V Z,F ] and the estimated dereverberation filter ^G=[^G ,1 ,…,^G F ] is input. The beamformer estimation unit 250 receives the observed signal x f,t The estimated spatial covariance matrix V Xand the estimated spatial covariance matrix of the non-target sound source, V Z From this, by equation (24), w f k+1 Ask for.
number
number
number
[0051] ^w k+1 ←(u f,m ) * ^w k+1 =(e m T u f ) * ^w k+1 (29) <Sound source extraction section 160> The sound source extraction unit 160 extracts the observed signal x f,tThe convolution beamformer ^w estimated as k+1 Using the inputs y and y, beamforming processing is performed using the following equation to estimate the sound source signal (S160), and the estimated value y f,t Output.
[0052] y f,t =(^w f k+1 ) H ^x f,t ∈ C ^w k+1 ∈C M+ML ^w k+1 =[^w1 k+1 | … | ^w F k+1 ] <Judgment section 280> The determination unit 280 determines whether the convergence condition is satisfied (S280), and if the convergence condition is satisfied (YES in S280), the estimated value y f,t is output as the output of the signal processing device, and the process ends. If the convergence condition is not satisfied (NO in S280), the determination unit 280 sends a control signal to each unit to repeat S220 to S160, and controls the processing of each unit. Note that the estimated value y f,t is used in the space-time covariance matrix estimation unit 220, and the calculation of equation (22) can be omitted. Note that the convergence conditions include whether the learning has been repeated a certain number of times (for example, several times), and whether the convolution beamformer ^w before and after estimation k+1 Conditions such as "Is the difference below a certain threshold?" can be used.
[0053] <Effects> With this configuration, it is possible to obtain the same effects as in the first embodiment. Furthermore, the Blind MaxSNR CBF of this embodiment is an ultra-fast method that can estimate the MaxSNR CBF with high accuracy in just a few iterations.
[0054] In this embodiment, the sound source signal s f,t The estimated value of y f,t However, the spatial image estimation unit 170 is provided, and the estimated value yf,t Using the spatial image s f,t image Approximation of u f y f,t may be calculated and output.
[0055] <Key Points of the Third Embodiment> In this embodiment, as a by-product of the Blind MaxSNR CBF of the second embodiment, the spatial covariance matrix V S is known (= estimated in advance), while the spatiotemporal covariance matrix of the unwanted sound, ^R N We develop "Iteratively Reweighted MaxSNR CBF (IR-MaxSNR CBF)," a method for estimating MaxSNR CBF with high accuracy under the condition that is unknown (i.e., not estimated in advance).
[0056] The spatial covariance matrix V of the target sound source S can be estimated with high accuracy, the information can be used to estimate the MaxSNR CBF with higher accuracy than the Blind MaxSNR CBF of the second embodiment.
[0057] Third Embodiment The following description will focus on the differences from the second embodiment.
[0058] FIG. 5 is a functional block diagram of a signal processing device according to the third embodiment, and FIG. 6 shows the processing flow thereof.
[0059] The signal processing device 300 includes an initialization unit 201, a first spatial covariance matrix estimator 110, a space-time covariance matrix estimator 220, a second spatial covariance matrix estimator 240, a dereverberation filter estimator 230, a beamformer estimator 250, a sound source extractor 160, and a determination unit 280.
[0060] This embodiment differs from the second embodiment in that it includes a first spatial covariance matrix estimator 110 instead of the first spatial covariance matrix estimator 210. The first spatial covariance matrix estimator 110 is as described in the first embodiment. Also, the beamformer estimator 250 estimates the observed signal xf,t The estimated spatial covariance matrix V X Instead, the estimated spatial covariance matrix V of the target sound source is S The other processing is the same as in the second embodiment.
[0061] <Other variations> The present invention is not limited to the above-described embodiments and modifications. For example, the various processes described above may not only be executed in chronological order as described, but may also be executed in parallel or individually depending on the processing capabilities of the devices that execute the processes or as needed. Other modifications are possible within the scope of the present invention.
[0062] <Programs and recording media> The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 7, and operating the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0063] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0064] The program may be distributed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to another computer via a network, thereby distributing the program.
[0065] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with each program transferred from the server computer. Alternatively, the server computer may not transfer the program to the computer, but may execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. In this embodiment, the program includes information used for processing by a computer that is equivalent to a program (such as data that is not a direct instruction to the computer but has properties that define computer processing).
[0066] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
Claims
1. a second spatial covariance matrix estimation unit that estimates a spatial covariance matrix of the non-target sound source using an estimated value of the space-time covariance matrix of the non-target sound source; a dereverberation filter estimator that estimates a dereverberation filter using the estimated value of the spatiotemporal covariance matrix of the non-target sound source; a beamformer estimation unit that estimates a convolution beamformer using an estimate of a spatial covariance matrix of an observed signal or a target sound source, an estimate of a spatial covariance matrix of the non-target sound source, and the estimated dereverberation filter; a sound source extraction unit that performs beamforming processing using the observed signal and the estimated convolution beamformer to estimate a sound source signal, Signal processing device.
2. 2. The signal processing device of claim 1, a first spatial covariance matrix estimation unit that estimates a section including a sound emitted from a target sound source (hereinafter also referred to as a target signal) from the observed signal and estimates a spatial covariance matrix of the target sound source using the estimated target signal; a space-time covariance matrix estimation unit that estimates a section that does not include a sound emitted by a target sound source (hereinafter also referred to as a non-target signal) from the observed signal, and estimates a space-time covariance matrix of the non-target sound source using the estimated non-target signal, the beamformer estimation unit estimates a convolution beamformer using an estimate of a spatial covariance matrix of the target sound source, an estimate of a spatial covariance matrix of the non-target sound source, and the estimated dereverberation filter. Signal processing device.
3. 2. The signal processing device of claim 1, a first spatial covariance matrix estimator that estimates a spatial covariance matrix of the observed signal using the observed signal; a space-time covariance matrix estimator that estimates a space-time covariance matrix of the non-target sound source by using the observed signal and the estimated convolution beamformer, the beamformer estimation unit estimates a convolution beamformer using an estimate of a spatial covariance matrix of the observed signals, an estimate of a spatial covariance matrix of the non-target sound source, and the estimated dereverberation filter; repeating the processes in the spatiotemporal covariance matrix estimator, the second spatial covariance matrix estimator, the dereverberation filter estimator, the beamformer estimator, and the sound source extractor until a convergence condition is satisfied; Signal processing device.
4. 2. The signal processing device of claim 1, a first spatial covariance matrix estimation unit that estimates a section including a sound emitted from a target sound source (hereinafter also referred to as a target signal) from the observed signal and estimates a spatial covariance matrix of the target sound source using the estimated target signal; a space-time covariance matrix estimator that estimates a space-time covariance matrix of the non-target sound source by using the observed signal and the estimated convolution beamformer, the beamformer estimation unit estimates a convolution beamformer using an estimate of a spatial covariance matrix of the target sound source, an estimate of a spatial covariance matrix of the non-target sound source, and the estimated dereverberation filter; repeating the processes in the spatiotemporal covariance matrix estimator, the second spatial covariance matrix estimator, the dereverberation filter estimator, the beamformer estimator, and the sound source extractor until a convergence condition is satisfied; Signal processing device.
5. A second spatial covariance matrix estimation step in which a computer estimates a spatial covariance matrix of a non-target sound source using an estimate of the spatiotemporal covariance matrix of the non-target sound source; a dereverberation filter estimation step in which a computer estimates a dereverberation filter using the estimated spatiotemporal covariance matrix of the non-target sound source; a beamformer estimation step in which a computer estimates a convolution beamformer using an estimated value of a spatial covariance matrix of an observed signal or a target sound source, an estimated value of a spatial covariance matrix of the non-target sound source, and the estimated dereverberation filter; a sound source extraction step of performing beamforming processing by a computer using the observed signal and the estimated convolution beamformer to estimate a sound source signal, Signal processing methods.
6. A program for causing a computer to function as the signal processing device according to any one of claims 1 to 4.
Citation Information
Patent Citations
Signal processing device, signal processing method, and program
WO2020121545A1