Signal processing device, signal processing method, and program
The ISS2 algorithm addresses the balance of fast convergence and low computational complexity in signal source separation by iteratively optimizing separation and mixing matrix components, achieving superior performance in signal processing tasks.
Patent Information
- Application Number
- JP2023568930
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2041-12-23
AI Technical Summary
Existing signal processing algorithms for signal source separation, such as IP2 and ISS1, either require significant computational resources per iteration or suffer from slow convergence, failing to balance fast convergence with low computational complexity effectively.
A signal processing device that employs a novel algorithm, ISS2, which updates the separation matrix by dividing the mixing matrix into sub-matrices and iteratively optimizing pairs of separation and mixing matrix components, using a multiplicative update method to achieve both fast convergence and reduced computational complexity.
The ISS2 algorithm significantly enhances convergence speed while maintaining low computational requirements, outperforming conventional methods in signal separation tasks.
Smart Images

Figure 0007740378000022 
Figure 0007740378000023 
Figure 0007740378000024
Abstract
Description
[Technical Field]
[0001] The present invention relates to a signal processing device, a signal processing method, and a program. [Background technology]
[0002] Signal source separation technology (or sound source separation technology), which estimates unmixed source signals from observed mixed signals, is a technology widely used for preprocessing of speech recognition, etc. Known methods for performing signal source separation using multiple sensors include Independent Component Analysis (ICA, Non-Patent Document 1) and Independent Vector Analysis (IVA, Non-Patent Document 2).
[0003] To date, an algorithm called Iterative Projection (IP) has been developed as an optimization algorithm for ICA and IVA. Two IPs have been developed so far: IP1 (Non-Patent Document 2) and IP2 (Non-Patent Document 3).
[0004] Another optimization algorithm for ICA and IVA that has been developed is called Iterative Source Steering (ISS, Non-Patent Document 4). In this specification, this ISS is referred to as ISS1. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] P. Common, "Independent component analysis, a new concept?" Signal processing 36.3 (1994), 287-314. [Non-patent document 2] N. Ono, "Stable and fast update rules for independent vector analysis based on auxiliary function technique," IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2011, pp. 189-192. [Non-patent document 3] Nobutaka Ono: Fast solutions for independent component analysis, independent vector analysis, and independent low-rank matrix analysis for three or more sound sources. Proceedings of the Acoustical Society of Japan, March 2018. [Non-patent document 4] R. Scheibler and N. Ono, "Fast and Stable Blind Source Separation with Rank-1 Updates," IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 236-240. Summary of the Invention [Problem to be solved by the invention]
[0006] IP2, an extension of IP1, has fast convergence but requires a large amount of calculation per iteration, while ISS1 requires a small amount of calculation per iteration but has slow convergence.
[0007] Therefore, an object of the present invention is to provide a signal processing device that achieves both fast convergence of IP2 and small computational complexity of ISS1. [Means for solving the problem]
[0008] The signal processing device of the present invention includes a separated signal updating unit.
[0009] The separation signal update unit solves the minimization problem of the upper bound function with respect to the separation matrix W in the upper bound minimization algorithm of the signal source separation technique IVA (independent vector analysis) by dividing the mixing matrix A into sub-matrices A l ,..., A L into d columns (d is an integer greater than or equal to 2), and updating one by one the pairs (W, A 1 ,..., A L )(l = 1,..., L), and updates the separation signal Y as the separation matrix W is updated.
Advantages of the Invention
Brief Description of the Drawings
Modes for Carrying Out the Invention
[0014] To estimate the source signal S, we use the separation matrix W (= A), which is the inverse matrix of A, instead of the mixing matrix A. -1 ) can be estimated. The separation result is Y = WX. The separation matrix W [k] ∈GL(m),k=1,...,K is defined by equation (2).
number
number
[0015] The model of the signal source separation technique IVA used in the present invention is defined as follows: In IVA, it is assumed that a multivariate vector of length K given by equation (3) follows a probability density function with quadratic or higher degree correlation.
number
number
number
[0016] The conventional algorithms IP1, IP2, and ISS1 and the algorithm ISS2 according to the present invention are algorithms that belong to a framework called the Majorization-Minimization Algorithm (MM algorithm). The MM algorithm for IVA is as follows:
[0017] <IVAのMMアルゴリズム> The MM algorithm for ICA has been proposed in Non-Patent Documents 1 to 3.
[0018] (Reference Non-Patent Document 1: N. Ono and S. Miyabe, “Auxiliary-function-based independent component analysis for super-Gaussian sources,” in Proc. LVA / ICA, 2010, pp. 165-172.) (Reference non-patent document 2: P. Ablin, A. Gramfort, J.-F. Cardoso, and F. Bach, “Stochastic algorithms with descent guarantees for ICA,” in Proc. AISTATS, 2019, pp. 1564-1573.) (Reference Non-Patent Document 3: N. Ono, “Stable and fast update rules for independent vector analysis based on auxiliary function technique,” in Proc. WASPAA, 2011, pp. 189-192.) where p(y) is a symmetric probability density function,
number
number
[0019] G'(r) / r is r∈(0,∞)=R >0 When p(y) is monotonically decreasing in G, p(y) is said to be a super-Gaussian distribution, where G' is the first derivative of G (see Non-Patent Documents 1, 2, 4 (pp. 60-61), and 5).
[0020] (Reference non-patent document 4: A. Benveniste, M. Metivier, and P. Priouret, Adaptive algorithms and stochastic approximations, 1st ed. Springer Science, 1990, vol. 22.) (Reference non-patent document 5: J. Palmer, D. Wipf, K. Kreutz-Delgado, and B. Rao, “Variational EM algorithms for non-Gaussian latent variable models,” in Proc. NIPS, vol. 18, 2005, pp. 1059-1066.) For example, the generalized Gaussian distribution (GGD) given by equation (5) is a super-Gaussian distribution.
number
number
number
number
number
[0021] (Reference Non-Patent Document 6: “Determinant maximization of a nonsymmetric matrix with quadratic constraints,” SIAM J. Optim., vol. 17, no. 4, pp. 997-1014, 2007.) (Reference Non-Patent Document 7: N. Ono, “Fast stereo independent vector analysis and its implementation on mobile phone,” in Proc. IWAENC, 2012, pp. 1-4.) However, when m≧3, no algorithm has been found to obtain the globally optimal solution of equation (12). Therefore, conventional block coordinate descent (BCD) algorithms IP1, IP2, and ISS1 have been developed to solve equation (12). These algorithms are called MM+BCD. In this invention, a new MM+BCD algorithm, ISS, has been developed. d This discloses the following.
[0022] Hereinafter, for simplicity of notation, the upper right index [k] will be omitted when explaining equation (12).
[0023] <MM+BCD Algorithm Disclosed in This Specification> Conventional algorithms IP1, IP2, and ISS1, and algorithm ISS disclosed in this specification d The difference lies in the way the MM algorithm solves the "optimization problem (Equation (12)) regarding the separation matrix W of the upper bound function (Equation (9))."
[0024] Conventional ISS (Reference Non-Patent Document 8, ISS1) is an algorithm that updates A column by column in each iteration.
[0025] (Reference Non-Patent Document 8: R. Scheibler and N. Ono, “Fast and stable blind source separation with rank-1 updates,” in Proc. ICASSP, 2020, pp. 236-240.)
[0026] The ISS2 algorithm disclosed in this specification is an algorithm that updates A by two columns at each iteration. To extend ISS1 to ISS2, we use the ISS algorithm that updates A by d columns for any d≧1. d This paper discloses a unified method for developing
[0027] <ISS d Definition of > Let d be a divisor of m. Let A be an L-submatrix A with d columns. 1 ,...,A LConsider dividing it into
number
number
number
number
number
number
Equation
Equation
Example
[0028] In the following Example 1, an optimization problem (Equation (7)) regarding the separation matrix W is solved by the algorithm ISS d defined in the method described in <Definition of ISS d (d is an arbitrary natural number), and a signal processing device 1 is disclosed. As described above, ISS d is an extension of the conventional method ISS1.
[0029] Specifically, it is an algorithm that updates the pair of (W, A l ) one by one according to the optimization problem (Equation (16)). Since the update rule (Equation (16)) is the same as the update rules (Equations (17) and (18)), (W, A l ) is updated according to the update rules (Equations (17) and (18)). The characteristic of ISS is the policy of updating the separation matrix W by updating a part of the mixing matrix A.
[0030] As described above, in ISS d , the algorithm with d = 1 coincides with the conventional ISS1, and in ISS d , the algorithm with d = 2 corresponds to ISS2 disclosed this time.
[0031] <Signal processing device 1> The functional configuration of a signal processing device 1 of this embodiment will be described with reference to Fig. 1. As shown in the figure, the signal processing device 1 of this embodiment includes an initial value setting unit 11, an auxiliary variable updating unit 12, a separated signal updating unit 13, and a control unit 14. The operation of the signal processing device 1 will be described below with reference to Fig. 2.
[0032] <Initial value setting section 11> The initial value setting unit 11 sets an appropriate initial value to the separation matrix W, and calculates the initial value Y of the separated signal by Y=WX (S11).
[0033] <Auxiliary variable update unit 12> The auxiliary variable update unit 12 repeatedly updates the auxiliary variable Λ under the control of the control unit 14 (S12).
[0034] <Separated signal update unit 13> The separated signal update unit 13 repeatedly updates the separated signal Y under the control of the control unit 14 (S13). Specifically, in the optimization problem (Equation (12)) regarding the separating matrix W of the upper bound function in the upper bound minimization algorithm of the signal source separation technique IVA (Independent Vector Analysis), the separated signal update unit 13 updates the mixing matrix A into a submatrix A having d columns (d is an integer of 2 or more). 1 ,...,A L (Equation (15)), and the separation matrix W and the submatrix A 1 ,...,A L The set (W,A l ) (l=1,...,L) one by one in accordance with the minimization problem for the upper bound function W (equation (16)), and the separated signal Y (=WX) is repeatedly updated (S13).
[0035] Since updating the separation matrix W is equivalent to updating the separation signals Y, it is sufficient to update only the separation signals Y without updating the separation matrix W.
[0036] <Control unit 14> The control unit 14 controls the auxiliary variable update unit 12 and the separated signal update unit 13 to alternately and repeatedly execute the operations until a predetermined condition is met.
[0037] The predetermined condition may be until a predetermined number of repetitions is reached, or until the update amount of each parameter becomes equal to or less than a predetermined threshold value, or the like.
[0038] <Experimental Results> Figure 3 shows the SDR improvement obtained with each method. We can see that the convergence of the proposed ISS2 is much faster than that of ISS1 and IP1, and comparable to that of IP2 (note that the SDR curves of IP2 and ISS2 nearly overlap). This is clear evidence of the effectiveness of our approach.
[0039] <Additional Notes> The device of the present invention may, for example, be a single hardware entity having an input section to which a keyboard or the like can be connected, an output section to which an LCD display or the like can be connected, a communication section to which a communication device (e.g., a communication cable) capable of communicating with an external device can be connected, a CPU (which may also have a central processing unit, cache memory, registers, etc.), memories such as RAM and ROM, an external storage device such as a hard disk, and buses connecting these input section, output section, communication section, CPU, RAM, ROM, and external storage device so that data can be exchanged between them. If necessary, the hardware entity may also be provided with a device (drive) capable of reading and writing recording media such as a CD-ROM. Examples of physical entities equipped with such hardware resources include general-purpose computers.
[0040] The external storage device of the hardware entity stores the programs required to realize the above-mentioned functions and the data required in the processing of these programs (not limited to the external storage device, but the programs may also be stored in a ROM, which is a read-only storage device). Furthermore, the data obtained by the processing of these programs is stored appropriately in the RAM or the external storage device.
[0041] In a hardware entity, each program stored in an external storage device (or ROM, etc.) and the data required to process each program are loaded into memory as needed, and interpreted, executed, and processed by the CPU as appropriate, resulting in the CPU realizing a predetermined function (each component represented as a unit, means, etc., above).
[0042] The present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of the present invention. Furthermore, the processes described in the above embodiments may not only be executed in chronological order according to the order described, but may also be executed in parallel or individually depending on the processing capacity of the device that executes the processes or as needed.
[0043] As mentioned above, when the processing functions of the hardware entities (devices of the present invention) described in the above embodiments are realized by a computer, the processing contents of the functions that the hardware entities should have are described by a program. Then, by executing this program on a computer, the processing functions of the hardware entities are realized on the computer.
[0044] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 10020 of the computer 10000 shown in Figure 4 and operating the control unit 10010, input unit 10030, output unit 10040, etc.
[0045] The program describing this processing content can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include magnetic recording devices, optical disks, magneto-optical recording media, and semiconductor memories. Specifically, for example, hard disk drives, flexible disks, and magnetic tapes can be used as magnetic recording devices; DVDs (Digital Versatile Discs), DVD-RAMs (Random Access Memory), CD-ROMs (Compact Disc Read Only Memory), and CD-Rs (Recordable) / RWs (Rewritable) can be used as optical disks; MOs (Magneto-Optical discs) can be used as magneto-optical recording media; and EEP-ROMs (Electrically Erasable and Programmable-Read Only Memory) can be used as semiconductor memories.
[0046] This program may be distributed, for example, by selling, transferring, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded. Furthermore, this program may be distributed by storing it in a storage device of a server computer and transferring it from the server computer to other computers via a network.
[0047] A computer that executes such a program, for example, first stores the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes processing in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute processing in accordance with the program. Furthermore, each time a program is transferred from a server computer to the computer, the computer may execute processing in accordance with the received program. Alternatively, the server computer may not transfer the program to the computer, but may instead implement the processing function by issuing an execution instruction and obtaining the results, thereby implementing the above-described processing through a so-called ASP (Application Service Provider) type service. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has the properties of determining the processing of the computer).
[0048] In addition, in this embodiment, a hardware entity is configured by executing a predetermined program on a computer, but at least a part of these processing contents may also be realized by hardware.
Claims
1. In the upper bound minimization algorithm of the signal source separation technique IVA (Independent Vector Analysis), the problem of minimizing a proxy function, which is a component of the upper bound function for the separating matrix W used to minimize the negative log-likelihood of the observed signal, is solved by dividing the mixing matrix A into a submatrix A with d columns (d is an integer equal to or greater than 2). 1 ,...,A L and divide the separation matrix W into the submatrix A 1 ,...,A L The set (W,A l ) (l=1,...,L) by updating them one by one under the constraint that the product of the separation matrix W and the mixing matrix A becomes an identity matrix when updating each set, and a separation signal update unit updates the separation signal Y in accordance with the update of the separation matrix W. Signal processing device.
2. 2. The signal processing device according to claim 1, d=2 Signal processing device.
3. A signal processing method executed by a signal processing device, comprising: In the upper bound minimization algorithm of the signal source separation technique IVA (Independent Vector Analysis), the problem of minimizing a proxy function, which is a component of the upper bound function for the separating matrix W used to minimize the negative log-likelihood of the observed signal, is solved by dividing the mixing matrix A into a submatrix A with d columns (d is an integer equal to or greater than 2). 1 ,...,A L and divide the separation matrix W into the submatrix A 1 ,...,A L The set (W,A l ) (l=1,...,L) by updating them one by one under the constraint that the product of the separation matrix W and the mixing matrix A becomes an identity matrix when updating each set, and updating the separated signal Y in accordance with the update of the separation matrix W. Signal processing methods.
4. A program that causes a computer to function as the signal processing device according to claim 1 or 2.