Signal processing device, signal processing method, and program

By selecting reference channels and sensor subsets based on SDR or pSDR, the signal processing device optimizes beamforming output quality for speech and acoustic signals, addressing the limitations of oSNR-based methods and reducing computational requirements.

WO2026078827A1PCT designated stage Publication Date: 2026-04-16NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/036179
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Conventional methods for selecting reference channels and sensor subsets in beamforming based on output SNR (oSNR) fail to accurately reflect the quality of the beamforming output signal, particularly for speech and acoustic signals, leading to suboptimal performance.

Method used

The proposed signal processing device selects reference channels and sensor subsets based on signal-to-distortion ratio (SDR) or predicted SDR (pSDR) to improve the quality of beamforming output signals.

Benefits of technology

This approach enhances the performance of beamforming by optimizing the selection process to better match the quality of speech and acoustic signals, improving the signal-to-distortion ratio and reducing computational load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024036179_16042026_PF_FP_ABST
    Figure JP2024036179_16042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is, inter alia, a signal processing device that performs reference channel selection or sensor subset selection based on SDR. A signal processing device selects at least one of a sensor subset and a reference channel on the basis of maximization of a value that corresponds to a signal-to-distortion ratio of an output signal in beamforming.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing device, signal processing method, and program

[0001] This invention relates to a signal separation technique (or sound source separation technique) for separating the signal before mixing from a mixed signal observed using two or more sensors (or microphones). The mixed signal includes the target sound and other interfering sounds, noise, etc.

[0002] Beamforming is being actively researched as a technique to separate the pre-mixed signal from the mixed signal observed by a sensor.

[0003] In a sensor network where sensors (microphones) are distributed in space, some sensors are located relatively close to the target sound source, while others are located far away. Generally, the signal-to-noise ratio (SNR) of signals observed by sensors close to the target sound source is high, while the SNR of signals observed by sensors far away is low. Therefore, the quality of the beamforming output signal (signal separation performance) strongly depends on the result of "reference channel selection," which is the sensor to which the reference for the beamforming output signal is set, that is, which sensor is selected as the reference channel for beamforming. For example, it is thought that selecting a sensor with a high SNR as the reference channel can improve the quality of the beamforming output signal. Non-patent document 1 proposes a method for selecting a reference channel for beamforming in which the reference channel with the highest SNR of the beamforming output signal is selected. This method is called reference channel selection based on output SNR (oSNR).

[0004] Furthermore, limiting the number of sensors used by beamforming can reduce the computational cost of the sensor network and lower its power consumption. For this reason, "sensor subset selection," which involves selecting which sensors to use for beamforming from among the sensors in the sensor network, is being actively researched. The aim is to reduce the number of sensors used while maintaining the quality of the beamforming output signal by selecting a sensor subset that is effective for beamforming. Non-patent document 2 proposes a method of selecting a sensor subset that maximizes the signal-to-noise ratio (SNR) of the beamforming output signal. This method is called sensor subset selection based on output SNR (oSNR).

[0005] TC Lawin-Ore and S. Doclo, "Reference Microphone Selection for MWF-based Noise Reduction Using Distributed Microphone Arrays", Speech Communication; 10. ITG Symposium, Braunschweig, Germany, 2012, pp. 1-4. J. Szurley, A. Bertrand, M. Moonen, P. Ruckebusch and I. Moerman, "Energy aware greedy subset selection for "Speech enhancement in wireless acoustic sensor networks", 2012 Proceedings of the 20th European Signal Processing Conference (EUSIPCO), Bucharest, Romania, 2012, pp. 789-793.

[0006] However, the output SNR (oSNR) of the beamforming output signal has a low correlation with the signal quality of speech and acoustic signals, such as the signal-to-distortion ratio (SDR) and PESQ (Perceptual Evaluation of Speech Quality) described in Reference 1. Therefore, there is a problem that the quality of the beamforming output signal may be low when reference channel selection and sensor subset selection are performed based on oSNR.

[0007] (Reference 1) E. Vincent, R. Gribonval and C. Fevotte, "Performance measurement in blind audio source separation", in IEEE Transactions on Audio, Speech, and Language Processing, vol. 14, no. 4, pp. 1462-1469, July 2006. The present invention aims to provide a signal processing device, a signal processing method, and a program that perform reference channel selection or sensor subset selection based on SDR or predicted SDR (pSDR) rather than oSNR.

[0008] To solve the above problems, according to one aspect of the present invention, the signal processing device selects at least one of a sensor subset and a reference channel based on maximizing a value corresponding to the signal-to-distortion ratio of the beamforming output signal.

[0009] The present invention has the effect of improving the performance of conventional oSNR-based methods.

[0010] Functional block diagram of the signal processing device according to the first embodiment. Diagram showing an example of the processing flow of the signal processing device according to the first embodiment. Functional block diagram of the signal processing device according to the second embodiment. Diagram showing an example of the processing flow of the signal processing device according to the second embodiment. Functional block diagram of the signal processing device according to the third embodiment. Diagram showing an example of the processing flow of the signal processing device according to the third embodiment. Diagram for explaining an algorithm showing an example of how to find a sensor subset. Functional block diagram of the signal processing device according to the fourth embodiment. Diagram showing an example of the processing flow of the signal processing device according to the fourth embodiment. Functional block diagram of the signal processing device according to the fifth embodiment. Diagram showing an example of the processing flow of the signal processing device according to the fifth embodiment. Diagram for explaining an algorithm showing an example of how to find a combination of a reference channel and a sensor subset. Functional block diagram of the signal processing device according to the sixth embodiment. Diagram showing an example of the processing flow of the signal processing device according to the sixth embodiment. Diagram showing an example of the configuration of a computer to which this method is applied.

[0011] Embodiments of the present invention will be described below. In the drawings used in the following description, components with the same function or steps that perform the same processing will be denoted by the same reference numerals, and redundant explanations will be omitted. In the following description, symbols such as "^" and "~" used in the text should ideally be placed directly above the following character, but due to the limitations of text notation, they are placed immediately before the character. In formulas, these symbols are written in their original positions. Furthermore, unless otherwise specified, the processing performed on each element of a vector or matrix will be applied to all elements of that vector or matrix.

[0012] <First Embodiment> Figure 1 shows a functional block diagram of the signal processing device 100 according to the first embodiment, and Figure 2 shows its processing flow.

[0013] The signal processing device 100 includes a target sound source covariance matrix estimation unit 120, a noise source covariance matrix estimation unit 110, a cross-correlation matrix estimation unit 130, a beamformer design unit 140, and a signal-to-distortion ratio calculation unit 150.

[0014] The signal processing device 100 takes as input the observation signal x(ω,t) observed by M spatially distributed sensors, selects the optimal sensor as the reference channel for beamforming, obtains and outputs the beamformer corresponding to the selected sensor. However, ω is the index of the frequency bin, where ω = 1, …, Ω, and t is the index of the time frame, where t = 1, …, T. In the following description, the frequency bin ω and the time frame t may be omitted on the premise that the signal in the short-time Fourier transform domain is processed. Let the target sound component contained in the observation signal x(ω,t) be x s (ω,t), the interference sound component be x i (ω,t), the noise component be x n (ω,t), and the noise source including interference sound and noise be x i+n = x i + x n . Let the target sound component x s (ω,t), the interference sound component x i , and the noise component x n be statistically independent components and mutually uncorrelated. x(ω,t) = x s (ω,t) + x i (ω,t) + x n ∈ C M , where C represents the set of all complex numbers.

[0015] The signal processing device 100 is a special device configured by loading a special program into a known or dedicated computer having, for example, a central processing unit (CPU) and a main memory (RAM). The signal processing device 100 executes each process under the control of, for example, the central processing unit. Data input to the signal processing device 100 and data obtained in each process are stored in, for example, the main memory, and the data stored in the main memory is read to the central processing unit as needed and used for other processes. Each processing unit of the signal processing device 100 may be composed of hardware such as an integrated circuit, at least in part. Each storage unit of the signal processing device 100 can be composed of, for example, a main memory such as RAM (Random Access Memory), or middleware such as a relational database or key-value store. However, each storage unit does not necessarily have to be located inside the signal processing device 100; it may be composed of an auxiliary storage device made of semiconductor memory elements such as a hard disk, optical disk, or flash memory, and may be located outside the signal processing device 100.

[0016] The following describes each part.

[0017] <Noise Source Covariance Matrix Estimation Unit 110> The noise source covariance matrix estimation unit 110 takes the observed signal x(ω,t) observed by M sensors as input and estimates the noise source x for each frequency bin ω=1,...,Ω. i+n The covariance matrix ^R i+n (ω) is estimated (S110), and the estimated value R i+n Outputs (ω). Noise source x i+n The covariance matrix ^R i+n (ω) is defined by the following equation. Here, Any estimation method can be used to estimate the covariance matrix of the noise source. For example, a time-frequency mask q for the noise source. i+n Using (ω,t), the covariance matrix ^R is given by the following equation. i+n (ω) can be estimated.

[0018] q i+n (ω,t)=1-q s (ω,t) (2) <Target Sound Source Covariance Matrix Estimation Unit 120> The target sound source covariance matrix estimation unit 120 takes the observed signal x(ω,t) observed by M sensors as input, and for each frequency bin ω=1,...,Ω, it estimates the target sound x s The covariance matrix ^R s (ω) is estimated (S120), and the estimated value R s Outputs (ω). Target sound x s The covariance matrix R s (ω) is defined by the following equation. Any estimation method can be used to estimate the covariance matrix of the target sound. For example, the time-frequency mask q for the target sound source, expressed by equation (3), can be used. s Using (ω,t), the covariance matrix ^R is given by the following equation. s (ω) can be estimated. R s (ω) ← P(R x (ω)-R i+n (ω)) (7) Here, R i+n (ω) is noise source x i+n =x i +x n The covariance matrix ^R i+n This is an estimated value of (ω) and is the output value of the noise source covariance matrix estimation unit 110. P(R) is R s This operator replaces the negative eigenvalues ​​of R with 0 to ensure that R is positive semi-definite. s When determining (ω), the target sound source covariance matrix estimation unit 120 uses the output value R of the noise source covariance matrix estimation unit 110. i+n Receive (ω).

[0019] <Cross-correlation matrix estimation unit 130> The cross-correlation matrix estimation unit 130 takes the observed signal x(ω,t) observed by M sensors as input and calculates the cross-correlation matrix ^C ​​between the target sound and the noise source for each frequency bin ω=1,...,Ω. s,i+n (ω) is estimated (S130), and the estimated value C s,i+n Outputs (ω). Cross-correlation matrix ^C ​​between target sound and noise source.s,i+n (ω) is defined by the following equation. Here, Cross-correlation matrix C s,i+n Any estimation method can be used to estimate (ω).

[0020] <Beamformer Design Unit 140> The beamformer design unit 140 estimates R s (ω) and estimated value R i+n With (ω) as the input, for each frequency bin ω=1,...,Ω, each channel r(r=1,2,...,M) is used as the reference channel for the beamformer w r (ω) is calculated (S140) and output. Any design method can be used for the beamformer. For example, a Rank-1 multichannel Winer filter defined by the following equation may be used. e r =[0,…,0,1,0,…0] is a unit vector where the r-th element is 1, and μ is a parameter that adjusts the noise suppression amount and is a non-negative real number. Tr(A) represents the trace of matrix A.

[0021] <Signal-to-distortion ratio calculation unit 150> The signal-to-distortion ratio calculation unit 150 calculates the estimated value R s (ω), estimated value R i+n (ω), estimated value C s,i+n (ω), and beamformer w r (ω) is used as input, and a reference channel is selected based on maximizing the value corresponding to the signal-to-distortion ratio (SDR) of the beamforming output signal (in this embodiment, the SDR itself).

[0022] For example, the signal-to-distortion ratio calculation unit 150 calculates the signal-to-distortion ratio (SDR) of the beamforming output signal for each frequency bin ω=1,...Ω (S150).

[0023] Here, the beamformer w when the observed signal x(ω,t) is input. r The separation signal of (ω) (the output signal of beamforming) is expressed by the following equation: A Η represents the complex conjugate transpose of A. Τx represents the transpose of A, and the signal of the reference channel is x s,r Let (ω,t)∈C, and ~s r (ω)=[x s,r (ω,1),…,x s,r (ω,T)] Τ ∈C T , ~y r (ω)=[y r (ω,1),…,y r (ω,T)] Τ ∈C T Let's assume that.

[0024] For example, the signal-to-distortion ratio calculation unit 150 calculates the narrowband SDR using the following formula. Here, Re(z) is an operator that returns the real part of the complex number z.

[0025] Furthermore, for example, the signal-to-distortion ratio calculation unit 150 calculates the broadband SDR using the following formula. Next, the signal-to-distortion ratio calculation unit 150 selects the channel r that maximizes the calculated SDR (narrowband SDR or broadband SDR) as the beamforming reference channel (S155), and the selected reference channel r select Beam Pharma W r_select Get (ω) and output it. Note that the subscript A_B is A B It means...

[0026] <Effects> With the above configuration, reference channel selection based on SDR can be performed, improving the performance of conventional oSNR-based methods.

[0027] <Second Embodiment> This section will focus on the differences from the first embodiment.

[0028] Figure 3 shows a functional block diagram of the signal processing device 200 according to the second embodiment, and Figure 4 shows its processing flow.

[0029] The signal processing device 200 includes a target sound source covariance matrix estimation unit 120, a noise source covariance matrix estimation unit 110, a beamformer design unit 140, and a signal-to-distortion ratio calculation unit 250. It does not include a cross-correlation matrix estimation unit 130, and the processing in the signal-to-distortion ratio calculation unit 250 differs from that of the signal processing device 100.

[0030] <Signal-to-distortion ratio calculation unit 250> The signal-to-distortion ratio calculation unit 250 calculates the estimated value R s (ω), estimated value R i+n (ω), and beamformer w r (ω) is used as input, and a reference channel is selected based on maximizing the value corresponding to the signal-to-distortion ratio (SDR) of the beamforming output signal (in this embodiment, the predicted signal-to-distortion ratio (pSDR: predicted SDR)).

[0031] For example, the signal-to-distortion ratio calculation unit 250 calculates a predicted value (pSDR: predicted SDR) of the signal-to-distortion ratio of the beamforming output signal for each frequency bin ω=1,...Ω (S250).

[0032] For example, the signal-to-distortion ratio calculation unit 250 calculates the narrowband pSDR using the following formula. Furthermore, for example, the signal-to-distortion ratio calculation unit 250 calculates the broadband pSDR using the following formula. Next, the signal-to-distortion ratio calculation unit 250 selects the channel r that maximizes the calculated pSDR (narrowband pSDR or broadband pSDR) as the beamforming reference channel (S255), and the selected reference channel r select Beam Pharma W r_select Get (ω) and output it.

[0033] <Effects> By adopting this configuration, the same effects as in the first embodiment can be obtained. Furthermore, by using pSDR instead of SDR, the computational load can be reduced.

[0034] <Third Embodiment> This section will focus on explaining the differences from the first embodiment.

[0035] Figure 5 shows a functional block diagram of the signal processing device 300 according to the third embodiment, and Figure 6 shows its processing flow.

[0036] The signal processing device 300 includes a target sound source covariance matrix estimation unit 120, a noise source covariance matrix estimation unit 110, a cross-correlation matrix estimation unit 130, a beamformer design unit 340, and a signal-to-distortion ratio calculation unit 350.

[0037] The signal processing device 300 takes as input observation signals x(ω,t) observed by M spatially distributed sensors and, given a beamforming reference channel, selects the optimal sensor subset, obtains the beamformer corresponding to the selected sensor subset, and outputs it.

[0038] In this embodiment, the reference channel r select It is assumed that the following are selected in advance. In this embodiment, any method may be used to select the reference channel. For example, the method used in the first embodiment, the second embodiment, Non-Patent Document 1, etc., may be used.

[0039] In this embodiment, a sensor subset is selected instead of a reference channel.

[0040] Consider a subset I of the set of M sensors {1, …, M}. Let I be called the sensor subset. The total number of sensors |I| (=K) included in the sensor subset I is predetermined. Of the M observation signals x(ω,t), let x be the observation signal corresponding to the sensor subset I. I We will represent this as (ω,t). Note that the reference channel is always included in the sensor subset I.

[0041] For each frequency bin ω=1, ..., Ω, the reference channel is r select Let w(I,r_select)(ω) be the beamformer used when beamforming a sensor subset I.

[0042] <Beamformer Design Unit 340> The beamformer design unit 340 estimates R s (ω) and estimated value R i+n (ω) and a subset I of multiple sensors are inputs, and for each frequency bin ω=1,...Ω, the reference channel is rselect and calculates the beamformer wI,r_select(ω) for each sensor subset I (S340), and outputs it. As a method for designing the beamformer, any design method may be used. For example, a Rank-1 multichannel Winer filter defined by the following equation may be used. Here, er_select,I ∈ C |I| is the unit vector e specified by the sensor subset I r_select ∈ C M obtained by extracting the rows of. μ is a parameter for adjusting the noise suppression amount and is a real number of 0 or more. R s,I (ω) and R i+n,I (ω) are the estimated values R s (ω) and the estimated value R i+n (ω) are submatrices.

[0043] <Signal-to-distortion ratio calculation unit 350> The signal-to-distortion ratio calculation unit 350 selects a sensor subset based on the maximization of a value corresponding to the signal-to-distortion ratio (SDR) of the output signal of beamforming (in this embodiment, the signal-to-distortion ratio itself), with the estimated value R s (ω), the estimated value R i+n (ω), the estimated value C s,i+n (ω), and the beamformer wI,r_select(ω) as inputs.

[0044] For example, the signal-to-distortion ratio calculation unit 350 obtains the signal-to-distortion ratio (SDR) of the output signal of beamforming of the beamformer wI,r_select(ω) for each sensor subset I for each frequency bin ω = 1,..., Ω (S350).

[0045] Here, the separated signal (output signal of beamforming) when the observation signal x I (ω,t) (observation signal specified by the sensor subset I) is input and the beamformer wI,r_select(ω) uses each sensor subset I is represented by the following equation. Let the signal of the reference channel be xs,r_select(ω,t) ∈ C, and ~s r_select\(\mathbf{r}(\omega)=[x_{s,r\_select(\omega,1)},\ldots,x_{s,r\_select(\omega,T)}]\) Τ \(\in\mathcal{C}\) T \(\text{,}\tilde{\mathbf{y}}\) r_select \(\mathbf{y}(\omega)=[\mathbf{y}(\omega,1),\ldots,\mathbf{y}(\omega,T)]\) r_select \(\text{,}\mathbf{y}(\omega,1),\ldots,\mathbf{y}(\omega,T)\) r_select \(\text{,}\mathbf{y}(\omega,T)]\) Τ \(\in\mathcal{C}\) T \(\text{be set as}\)

[0046] For example, the signal - to - distortion ratio calculation unit 350 obtains the SDR (narrowband SDR or broadband SDR) using Equation (11) or (15). However, instead of the estimated values \(\mathbf{R}(\omega)\) and \(\mathbf{R}(\omega)\) respectively, \(\hat{\mathbf{R}}(\omega)\) and \(\hat{\mathbf{R}}(\omega)\) are used. Also, in Equations (12) and (14), instead of the estimated value \(\mathbf{C}(\omega)\), \(\hat{\mathbf{C}}(\omega)\) is used. \(\hat{\mathbf{C}}(\omega)\) is a sub - matrix of the estimated value \(\mathbf{C}(\omega)\) consisting of the rows and columns specified by the sensor subset \(I\). s \(\mathbf{R}(\omega)\text{ and the estimated value}\mathbf{R}(\omega)\) i+n \(\text{, replace them with}\mathbf{R}(\omega)\text{ and}\mathbf{R}(\omega)\text{ respectively}\) s,I \(\text{,}\mathbf{R}(\omega)\text{ and}\mathbf{R}(\omega)\) i+n,I \(\text{In addition, in Equations (12) and (14), instead of the estimated value}\mathbf{C}(\omega)\) s,i+n \(\text{, replace it with}\mathbf{C}(\omega)\) I,s,i+n \(\mathbf{C}(\omega)\text{ is used.}\mathbf{C}(\omega)\) I,s,i+n \(\text{is a sub - matrix of the estimated value}\mathbf{C}(\omega)\text{ consisting of the rows and columns specified by the sensor subset}I\). s,i+n \(\text{

[0047] The signal - to - distortion ratio calculation unit 350 selects the sensor subset \(I\) for which the obtained SDR (narrowband SDR or broadband SDR) is maximized (S355), obtains the beamformer \(\mathbf{w}_{I,r\_select(\omega)}\) corresponding to the selected sensor subset \(I\), and outputs it.

[0048] <Effect> With the above configuration, sensor subset selection based on SDR can be performed, and the performance of the conventional method based on oSNR can be improved.

[0049] As a method for obtaining the sensor subset \(I\) that maximizes the SDR, all sensor subsets may be exhaustively searched, or the greedy method of Algorithm 1 in FIG. 7 may be used.

[0050] In Algorithm 1, first, for example, the reference channel \(r_0\) is set by the method of the first embodiment, the second embodiment, Non - Patent Document 1, etc. (S1 in FIG. 7), and the reference channel \(r_0\) is added to the temporary sensor subset \(I\) (S2 in FIG. 7).

[0051] One sensor not included in the provisional sensor subset I is added to the provisional sensor subset I, and a beamformer is designed for this set (hereinafter also referred to as the provisional sensor subset candidate) (corresponding to S340 in Figure 6), and the signal-to-distortion ratio of the beamforming output signal is calculated (corresponding to S350 in Figure 6). The signal-to-distortion ratio is calculated for all sensor subset candidate corresponding to the sensor not included in the provisional sensor subset I, and the sensor with the maximum signal-to-distortion ratio is added to the provisional sensor subset I to form a new provisional sensor subset I (corresponding to S4, S5 in Figure 7 and S355 in Figure 6). Processes S4 and S5 are repeated until the number of sensors included in the provisional sensor subset I reaches a predetermined number K (S3, S6 in Figure 7), and once the number of sensors included in the provisional sensor subset I reaches a predetermined number K, the provisional sensor subset I at that point is selected as the final sensor subset I (S7 in Figure 7).

[0052] <Fourth Embodiment> This section will focus on explaining the differences from the third embodiment.

[0053] In this embodiment, the second and third embodiments are combined.

[0054] Figure 8 shows a functional block diagram of the signal processing device 400 according to the fourth embodiment, and Figure 9 shows its processing flow.

[0055] The signal processing device 400 includes a target sound source covariance matrix estimation unit 120, a noise source covariance matrix estimation unit 110, a beamformer design unit 340, and a signal-to-distortion ratio calculation unit 450. It does not include a cross-correlation matrix estimation unit 130, and the processing in the signal-to-distortion ratio calculation unit 450 differs from that of the signal processing device 300.

[0056] <Signal-to-distortion ratio calculation unit 450> The signal-to-distortion ratio calculation unit 450 calculates the estimated value R s (ω), estimated value R i+n (ω) and beamformer wI,r_select(ω) are inputs, and a subset of sensors is selected based on maximizing the value corresponding to the signal-to-distortion ratio (SDR) of the beamforming output signal (in this embodiment, the predicted signal-to-distortion ratio (pSDR: predicted SDR)).

[0057] For example, the signal-to-distortion ratio calculation unit 450 calculates a predicted value (pSDR: predicted SDR) of the signal-to-distortion ratio of the beamforming output signals of beamformers wI,r_select(ω) for each sensor subset I for each frequency bin ω=1,...,Ω (S450).

[0058] For example, the signal-to-distortion ratio calculation unit 450 calculates the pSDR (narrowband pSDR or broadband pSDR) using equation (21) or (25). However, the estimated value R s (ω) and estimated value R i+n Replace (ω) with R s,I (ω) and R i+n,I Use (ω).

[0059] Next, the signal-to-distortion ratio calculation unit 450 selects the sensor subset I that maximizes the calculated pSDR (narrowband pSDR or broadband pSDR) (S455), obtains the beam pharma wI,r_select(ω) corresponding to the selected sensor subset I, and outputs it.

[0060] <Fifth Embodiment> This section will focus on explaining the differences from the third embodiment.

[0061] When selecting a sensor subset, whether the selected sensor subset is effective for beamforming strongly depends on the reference channel used by beamforming. Conventional methods and the third and fourth embodiments employ a two-step approach: first, the reference channel is identified, and then the sensor subset selection is solved using that reference channel. However, since the optimal reference channel differs for each sensor subset, the two-step method may not be able to select the optimal sensor subset. Here, "optimal" means that the quality of the beamforming output signal is high.

[0062] In this embodiment, a sensor subset and a reference channel pair are selected simultaneously (simultaneous optimization).

[0063] Figure 10 shows a functional block diagram of the signal processing device 500 according to the fifth embodiment, and Figure 11 shows its processing flow.

[0064] The signal processing device 500 includes a target sound source covariance matrix estimation unit 120, a noise source covariance matrix estimation unit 110, a cross-correlation matrix estimation unit 130, a beamformer design unit 540, and a signal-to-distortion ratio calculation unit 550.

[0065] The signal processing device 500 takes the observed signal x(ω,t) observed by M spatially distributed sensors as input, selects the optimal combination of reference channel and sensor subset for beamforming, acquires the beamform corresponding to the selected combination of reference channel and sensor subset, and outputs it.

[0066] Unlike the third embodiment, in this embodiment, the reference channel is not pre-selected.

[0067] In this embodiment, a combination of a reference channel and a sensor subset is selected, rather than a sensor subset.

[0068] <Beamformer Design Unit 540> The beamformer design unit 540 estimates R s (ω) and estimated value R i+n (ω) and a subset of multiple sensors I⊆[M] are inputs, and for each frequency bin ω=1,...,Ω, the reference channel is r * Let ∈ I, and beamformer w for each sensor subset I. I,r^* Calculate (ω) (S540) and output it. However, the subscript A^B is A B This means that any design method can be used for the beamformer. For example, a Rank-1 multichannel Winer filter defined by equation (31) may be used. However, the determined reference channel r select Instead, any channel r included in the sensor subset I * Using this, the beamformer w for each combination of reference channel and sensor subset I,r^* Calculate (ω) and output it.

[0069] <Signal-to-distortion ratio calculation unit 550> The signal-to-distortion ratio calculation unit 550 calculates the estimated value R s (ω), estimated value R i+n (ω), estimated value C s,i+n(ω), and beamformer w for each combination of reference channel and sensor subset I,r^* (ω) is used as input, and a combination of reference channel and sensor subset is selected based on maximizing the value corresponding to the signal-to-distortion ratio (SDR) of the beamforming output signal (in this embodiment, the signal-to-distortion ratio itself).

[0070] For example, the signal-to-distortion ratio calculation unit 550 calculates the beamformer w for each combination of the reference channel and the sensor subset for each frequency bin ω=1,...Ω. I,r^* The signal-to-distortion ratio (SDR) of the beamforming output signal of (ω) is calculated (S550). The method for calculating the SDR is the same as that of the signal-to-distortion ratio calculation unit 350.

[0071] The signal-to-distortion ratio calculation unit 550 calculates the combination of the reference channel and sensor subset that maximizes the calculated SDR (narrowband SDR or broadband SDR) (I,r * Select (S555), and select the combination (I,r * ) corresponds to Beam Pharma w I,r^* Get (ω) and output it.

[0072] <Effects> The above configuration provides the same effects as the first and third embodiments. Furthermore, by simultaneously selecting (simultaneously optimizing) pairs of sensor subsets and reference channels, the performance of conventional sensor subset selection can be improved.

[0073] Note that the combination of sensor subsets that maximizes SDR (I,r * As for how to find ), one can either exhaustively search all combinations, or use the greedy algorithm shown in Algorithm 2 in Figure 12.

[0074] Algorithm 2 first initializes the provisional sensor subset I (S1 in Figure 12).

[0075] One sensor not included in the provisional sensor subset I is added to the provisional sensor subset I, and one of the sensors in that set (hereinafter also called the provisional sensor subset candidate) is used as the reference channel (this combination of provisional sensor subset candidate and reference channel is also called the provisional combination), and a beamformer is designed. The reference channel of the provisional combination is changed, and a beamformer is designed for all provisional combinations corresponding to the provisional sensor subset candidate. Furthermore, the sensor to be added is changed, and a beamformer is designed for all provisional combinations corresponding to all provisional sensor subset candidates (corresponding to S540 in Figure 11). Next, the signal-to-distortion ratio of the beamforming output signal is calculated (corresponding to S550 in Figure 11). Among all provisional combinations of all provisional sensor subset candidates, the sensor of the provisional combination with the maximum signal-to-distortion ratio (more specifically, the sensor added in the provisional sensor subset candidate corresponding to the provisional combination with the maximum signal-to-distortion ratio) is added to the provisional sensor subset I, and a new provisional sensor subset I is created (corresponding to S3, S4 in Figure 12, and S555 in Figure 11). Processes S3 and S4 are repeated until the number of sensors included in the provisional sensor subset I reaches a predetermined number K (S2 and S5 in Figure 12). Once the number of sensors included in the provisional sensor subset I reaches the predetermined number K, the provisional combination at that point is set to the final sensor subset I and reference channel r. * Select this as a combination (S6 in Figure 12).

[0076] <Sixth Embodiment> This section will focus on explaining the differences from the fifth embodiment.

[0077] In this embodiment, the second embodiment and the fifth embodiment are combined.

[0078] Figure 13 shows a functional block diagram of the signal processing device 600 according to the sixth embodiment, and Figure 14 shows its processing flow.

[0079] The signal processing device 600 includes a target sound source covariance matrix estimation unit 120, a noise source covariance matrix estimation unit 110, a beamformer design unit 340, and a signal-to-distortion ratio calculation unit 650. It does not include a cross-correlation matrix estimation unit 130, and the processing in the signal-to-distortion ratio calculation unit 650 differs from that of the signal processing device 500.

[0080] <Signal-to-distortion ratio calculation unit 650> The signal-to-distortion ratio calculation unit 650 calculates the estimated value R s (ω), estimated value R i+n (ω), and beamformer w r (ω) is used as input, and a combination of reference channel and sensor subset is selected based on maximizing the value corresponding to the signal-to-distortion ratio (SDR) of the beamforming output signal (in this embodiment, the predicted signal-to-distortion ratio (pSDR: predicted SDR)).

[0081] For example, the signal-to-distortion ratio calculation unit 650 calculates the beamformer w for each combination of the reference channel and the sensor subset for each frequency bin ω=1,...Ω. I,r^* The predicted signal-to-distortion ratio (pSDR) of the beamforming output signal of (ω) is calculated (S650). The method for calculating pSDR is the same as that of the signal-to-distortion ratio calculation unit 450.

[0082] The signal-to-distortion ratio calculation unit 650 calculates the combination of the reference channel and the sensor subset that maximizes the calculated pSDR (narrowband pSDR or broadband pSDR) (I,r * ) select (S655), and select (I,r * ) corresponds to Beam Pharma w I,r^* Get (ω) and output it.

[0083] <Effects> With the above configuration, the same effects as those of the second, fourth, and fifth embodiments can be obtained.

[0084] <Modification> Furthermore, the present invention may also include a device (terminal) for using the apparatus, system, or method of the present invention via a network (telecommunication line). The "device (terminal) for use" may be equipped with functions necessary to obtain the effects of implementing the apparatus, system, or method of the present invention (for example, control functions, decoding functions, restoration functions, input / output functions, etc.). Note that a configuration including a device (terminal) for using the apparatus or method of the present invention via a network (telecommunication line) is also called a signal processing system.

[0085] <Hardware, Programs and Recording Media> The functions realized by the components described herein may be implemented in a circuitry or processing circuitry, including a general-purpose processor, an application-specific processor, an integrated circuit, an ASIC (Application Specific Integrated Circuit), a CPU (a Central Processing Unit), conventional circuits, and / or a combination thereof, programmed to realize the functions described herein. A processor includes transistors and other circuits and is considered a circuitry or processing circuitry. A processor may be a programmed processor that executes a program stored in memory.

[0086] In this specification, circuitry, unit, and means are hardware programmed to perform or execute the functions described herein. Such hardware may be any hardware disclosed herein, or any hardware known to be programmed to perform or execute the functions described herein.

[0087] If the hardware is a processor that is considered to be a type of circuitry, then the circuitry, means, or unit is a combination of hardware and software used to constitute the hardware and / or processor.

[0088] The various processes described above can be carried out by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 15, and then causing the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc. to operate.

[0089] The program describing this process can be recorded on a computer-readable recording medium. Any computer-readable recording medium can be used, such as a magnetic recording device, optical disc, magneto-optical recording medium, or semiconductor memory.

[0090] Furthermore, this program may be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs or CD-ROMs on which the program is recorded. Alternatively, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.

[0091] A computer executing such a program may, for example, first store the program recorded on a portable storage medium or a program transferred from a server computer in its own storage device. Then, when processing is to be executed, the computer reads the program stored on its own storage medium and executes the processing according to the read program. Alternatively, the computer may directly read the program from the portable storage medium and execute the processing according to that program, or it may sequentially execute the processing according to the received program each time a program is transferred to it from a server computer. Furthermore, the processing may be executed using a so-called ASP (Application Service Provider) type service, where the processing function is realized only by issuing execution instructions and obtaining results, without transferring the program from the server computer to this computer.In addition, the processing may be executed using a so-called SaaS (Software as a Service) type service, where a part of the server computer is made available to the user along with the program. Furthermore, the term "program" in this form includes information used for processing by an electronic computer that is equivalent to a program (data, etc., that is not a direct instruction to the computer but has the property of defining the processing of the computer).

[0092] Furthermore, in this configuration, the device is configured by executing a predetermined program on a computer, but at least a part of these processes may be implemented in hardware.

[0093] <Other Modifications> The present invention is not limited to the embodiments and modifications described above. For example, the various processes described above may not only be performed sequentially according to the description, but may also be performed in parallel or individually as needed, depending on the processing capacity of the device performing the processes. Other modifications can be made as appropriate without departing from the spirit of the present invention.

Claims

1. A signal processing device that selects at least one of a sensor subset and a reference channel based on maximizing a value corresponding to the signal-to-distortion ratio of the beamforming output signal.

2. A signal processing device according to claim 1, comprising simultaneously selecting the sensor subset and the reference channel.

3. A signal processing method in which a signal processing device selects at least one of a sensor subset and a reference channel based on maximizing a value corresponding to the signal-to-distortion ratio of the beamforming output signal.

4. A program for causing a computer to function as a signal processing device according to claim 1.