Sound source separation system

By generating virtual microphone signals and using the ICA method to separate sound source signals, the high cost problem caused by the need for the same number of microphones in existing technologies is solved, and efficient multi-source signal separation is achieved.

CN121237115APending Publication Date: 2025-12-30ALPS ALPINE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510849408.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-06-24
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing ICA methods require the same number of microphones as the number of sound sources when separating multi-source signals, resulting in high costs.

Method used

By using n microphones to generate m virtual microphone signals, and using the ICA method to separate L sound source signals, where the m virtual microphone signals are directional in different directions and L≥n, the separation of sound source signals is achieved.

Benefits of technology

It effectively separates sound source signals that outnumber the number of microphones, reducing the cost of sound source separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237115A_ABST
    Figure CN121237115A_ABST
Patent Text Reader

Abstract

The present invention addresses the problem of providing a "sound source separation system" for separating sound source signals of sound sources having a larger number than the number of microphones by an ICA method. The solution is that five beamformers (21) of a virtual microphone sound generation unit (2) generate five virtual microphone signals, which are outputs of five virtual single-directivity microphones having directivity in different directions, on the basis of outputs of four omnidirectional microphones (11) of a microphone group (1). An ICA processing unit (3) uses five virtual microphone signals inputted from a virtual microphone sound generation unit (2) as inputs from five different microphones (11), and separates and outputs five sound source signals, which are signals of five different sound sources, by an ICA method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technique for blind source separation of each sound source signal from a mixed signal composed of multiple sound source signals. Background Technology

[0002] As a technique for blind source separation of each sound source signal from a mixed signal composed of multiple sound source signals, the following technique is known: multiple microphones are arranged in an acoustic space containing multiple sound sources, and the sound source signals are separated from the mixed signal composed of each sound source signal picked up by each microphone by using the ICA (Independent Component Analysis) method, which separates each sound source signal by utilizing the statistical independence of each sound source signal (for example, Patent Document 1).

[0003] In addition, as a technology associated with the present invention, a technique for microphone arrays is known in which the outputs of multiple omnidirectional microphones are combined to generate the output of a virtual unidirectional microphone, and the direction of the directionality of the virtual unidirectional microphone is controlled (e.g., Patent Document 2).

[0004] Prior technical documents:

[0005] Patent documents:

[0006] Patent Document 1: Japanese Patent Application Publication No. 2007-215163

[0007] Patent Document 2: Japanese Patent Application Publication No. 2016-032260 Summary of the Invention

[0008] The problem that the invention aims to solve:

[0009] The aforementioned ICA method cannot be applied to the separation of sound source signals from a number of sound sources that is greater than the number of microphones. Therefore, it is necessary to set up a number of microphones that are equal to or greater than the number of sound sources whose signals need to be separated, which increases the cost of multi-source sound source separation.

[0010] Therefore, the objective of this invention is to use the ICA method to effectively separate sound source signals from a larger number of sound sources than the number of microphones.

[0011] Methods used to solve problems:

[0012] To achieve the aforementioned goal, the present invention provides a sound source separation system that performs blind source separation on multiple mixed signals composed of two or more sound source signals using the ICA (Independent Component Analysis) method. The system comprises: n microphones (where n > 2); a virtual microphone signal generation mechanism that generates m virtual microphone signals based on the output signals of the n microphones, wherein the m virtual microphone signals are the output signals of m (where m > n) virtual unidirectional microphones, each with directional characteristics in a different direction; and an ICA processing mechanism that uses the m virtual microphone signals as the multiple mixed signals and separates the L sound source signals, which are L different sound sources (where L > n and L ≤ m), using the ICA method.

[0013] Furthermore, in this sound source separation system, the relationship between L and m can also be set to satisfy L = m.

[0014] In the above sound source separation system, the directional direction of the m virtual unidirectional microphones can also be set at equal angular intervals.

[0015] Alternatively, in the above sound source separation system, the directional direction of L of the m virtual unidirectional microphones can be set to the direction of the L sound sources.

[0016] Alternatively, in the above sound source separation system, the virtual microphone signal generation mechanism can be configured to change the directionality of the m virtual microphone signals. In this sound source separation system, a sound source detection mechanism is provided to detect the direction of the L sound sources; and a directionality control mechanism is provided to control the virtual microphone signal generation mechanism so that the directionality of the L virtual unidirectional microphones among the m virtual unidirectional microphones becomes the direction of the L sound sources detected by the sound source detection mechanism.

[0017] In the above sound source separation system, the n microphones can also be configured in a car to pick up sound inside the car's cabin.

[0018] Based on the sound source separation system described above, m virtual microphone signals are generated from the outputs of n (where n > 2) microphones, which are the output signals of m (where m > n) virtual unidirectional microphones with directivity in different directions. Then, the ICA method is used to separate L sound source signals, which are signals from L different sound sources (where L > n and L ≤ m). Here, the outputs of the m virtual unidirectional microphones are equivalent to the outputs of m real unidirectional microphones. Therefore, the separation of the L sound source signals using the ICA method is equivalent to source separation of sound sources with fewer than the number of unidirectional microphones from the outputs of the m (where L ≤ m) real unidirectional microphones.

[0019] Therefore, the ICA method can effectively separate sound source signals from a larger number of sound sources than the number of microphones.

[0020] Invention effects:

[0021] As described above, according to the present invention, the ICA method can be used to effectively separate sound source signals from a number of sound sources that are greater than the number of microphones. Attached Figure Description

[0022] Figure 1 This is a block diagram illustrating the configuration of a sound source separation system according to an embodiment of the present invention.

[0023] Figure 2 This is a diagram showing the direction (pickup axis) of the unidirectional tone generated in an embodiment of the present invention.

[0024] Figure 3 This is a diagram illustrating an example of the configuration of a beamformer according to an embodiment of the present invention.

[0025] Figure 4 This is a diagram illustrating an applicable example of a sound source separation system according to an embodiment of the present invention.

[0026] Figure 5 This is a diagram illustrating other configuration examples of the sound source separation system according to embodiments of the present invention.

[0027] Figure 6 This is a block diagram illustrating other configuration examples of the sound source separation system according to embodiments of the present invention.

[0028] Explanation of reference numerals in the attached figures:

[0029] 1…microphone group, 2…virtual microphone sound generation unit, 3…ICA processing unit, 4…directional control unit, 11…microphone, 21…beamformer, 22…variable beamformer, 211…delay unit, 212…addition unit. Detailed Implementation

[0030] The embodiments of the present invention will be described below.

[0031] First, the first embodiment will be described.

[0032] Figure 1 This indicates the configuration of the sound source separation system involved in this embodiment.

[0033] As shown in the figure, the sound source separation system includes a microphone group 1, a virtual microphone sound generation unit 2, and an ICA processing unit 3.

[0034] Microphone group 1 has four omnidirectional microphones 11, namely M1, M2, M3 and M4, which are configured separately.

[0035] The virtual microphone sound generation unit 2 has five beamformers 21: a 0° beamformer, a 45° beamformer, a 90° beamformer, a 135° beamformer, and a 180° beamformer. The outputs of all four microphones 11 are input to each beamformer 21.

[0036] like Figure 2 As shown in Figure a, taking the predetermined direction of observation from the reference point RP of the pre-determined microphone group 1 as the 90° direction, the 0° beamformer 21 generates a virtual microphone signal Sig_0° as the output of a virtual monodirectional microphone with directionality in the 0° direction based on the output of the four microphones 11, and outputs it to the ICA processing unit 3. The 45° beamformer 21 generates a virtual microphone signal Sig_45° as the output of a virtual monodirectional microphone with directionality in the 45° direction based on the output of the four microphones 11, and outputs it to the ICA processing unit 3. The 90° beamformer 21 generates... The virtual microphone signal Sig_90°, which is the output of a virtual monodirectional microphone with a directional orientation in the 90° direction, is output to the ICA processing unit 3. The 135° beamformer 21 generates a virtual microphone signal Sig_135°, which is the output of a virtual monodirectional microphone with a directional orientation in the 135° direction, based on the output of the four microphones 11, and outputs it to the ICA processing unit 3. The 180° beamformer 21 generates a virtual microphone signal Sig_180°, which is the output of a virtual monodirectional microphone with a directional orientation in the 180° direction, based on the output of the four microphones 11, and outputs it to the ICA processing unit 3.

[0037] Furthermore, when using a microphone group 1 consisting of four microphones 11 arranged in one dimension, such as Figure 2As shown in b, the reference point RP can be set as the center of the microphone group 1, the 90° direction can be set as the direction perpendicular to the arrangement direction of the four microphones 11, i.e., the front direction, the 0° direction can be set as the right direction of the microphone group 1, and the 180° direction can be set as the left direction of the microphone group 1.

[0038] Here, taking the case where a delayed-addition type microphone array is formed by microphone group 1 and beamformer 21 as an example, the configuration of each beamformer 21 will be explained.

[0039] The five beamformers 21 have the same configuration. Figure 3 The configuration of the 45° beamformer 21 is represented by 'a'.

[0040] In this case, as shown in the figure, the 45° beamformer 21 includes: four delay units 211 corresponding to the four microphones 11 respectively, and an adder 212 that adds the outputs of each delay unit 211 and outputs it as a virtual microphone signal Sig_45° to the ICA processing unit 3.

[0041] exist Figure 3 In b, as shown in the case of using a microphone group 1 in which four microphones 11 are equally arranged in one dimension, in the sound arriving at each microphone 11 from the 45° direction, a delay difference is generated corresponding to the position of each microphone 11. When the interval between adjacent microphones 11 is d, the delay difference is: the sound arriving at microphone M2 is delayed by D2 = d × sin45° relative to the sound arriving at microphone M1, the sound arriving at microphone M3 is delayed by D3 = 2d × sin45° relative to the sound arriving at microphone M1, and the sound arriving at microphone M4 is delayed by D4 = 3d × sin45° relative to the sound arriving at microphone M1.

[0042] Delay times are set for the four delay units 211 to match the delays of the outputs of the four microphones 11, and also to match the delays of the virtual microphone signals output by the other beamformers 21. For example, when the delay time used to match the delays of the virtual microphone signals output by the other beamformers 21 is set to Dx, the delay time of the delay unit 211 corresponding to M1 is set to D4+Dx, the delay time of the delay unit 211 corresponding to M2 is set to (D4-D2)+Dx, the delay time of the delay unit 211 corresponding to M3 is set to (D4-D3)+Dx, and the delay time of the delay unit 211 corresponding to M4 is set to Dx.

[0043] As a result, the virtual microphone signal Sig_45° output by the summing unit 212 becomes a sound obtained by adding the sounds arriving at each microphone 11 from the 45° direction with the same phase, and becomes a directional signal in the 45° direction.

[0044] return Figure 1 The ICA processing unit 3 takes the five virtual microphone signals Sig_0°, Sig_45°, Sig_90°, Sig_135°, and Sig_180° input from the virtual microphone sound generation unit 2 as inputs from five different microphones 11, and uses the ICA method to separate and output the sound source signals Sig_AS1, Sig_AS2, Sig_AS3, Sig_AS4, and Sig_AS5 from the five different sound sources.

[0045] In this embodiment, based on the outputs of the four microphones 11, the outputs of five virtual unidirectional microphones, each directional in a different direction, are generated. For the outputs of these five virtual unidirectional microphones, sound source separation of five sound sources is performed using the ICA method. Here, the outputs of the five virtual unidirectional microphones are equivalent to the outputs of the five real unidirectional microphones; therefore, the processing of the ICA processing unit 3 is equivalent to performing sound source separation of five sound sources from the outputs of the five real unidirectional microphones using the ICA method, the same number as the number of unidirectional microphones.

[0046] Therefore, the ICA processing unit 3 can effectively separate sound source signals from a greater number of sound sources than the number of microphones 11 in the microphone group 1.

[0047] Next, an applicable example of the sound source separation system involved in this embodiment will be presented.

[0048] The sound source separation system can be applied, for example, to separate multiple sound sources inside a car. In this case, microphone group 1 is positioned to pick up a wide range of sounds inside the car. Such a position for wide sound pickup can be, for example, set to... Figure 4 The front part of the roof, the dashboard, or the center of the roof, as shown in a and b.

[0049] In addition, as sound sources and sound source signals inside a car, there are in-vehicle speakers and their speaker output sounds, in-vehicle equipment and their alarm sounds, and occupants and their spoken voices.

[0050] Here, in the case of applying a sound source separation system to the sound source separation inside a car, the sound source signal separated by the sound source separation system can be used as a noise source signal for active noise control, etc. In this active noise control, control is performed to make the sound of a sound source that is useless to each occupant of the car as noise and cannot be heard.

[0051] The embodiments of the present invention have been described above.

[0052] Here, the directional direction of the virtual microphone signal generated by each beamformer 21 of the virtual microphone sound generation unit 2 is set to 0°, 45°, 90°, 135°, and 180°. However, the directional direction of the virtual microphone signal generated by each beamformer 21 can be set to any direction as long as they are different from each other.

[0053] For example, if the direction of the sound source from which the sound source signal is separated is known in advance, the direction of the directional of the virtual microphone signal generated by each beamformer 21 can also be set to the direction of the known sound source.

[0054] That is, for example, such as Figure 5 As shown in a1, when microphone group 1 is positioned at the front end of the ceiling, and the vehicle-mounted speakers SP1, SP2, SP3, SP4, and SP5 in the figure are used as sound sources, and the speaker output sound is separated as the sound source signal, as follows: Figure 5 As shown in a2, the direction of each vehicle speaker SP1, SP2, SP3, SP4, SP5 when viewed from the microphone group 1 can also be set to the directional direction of the virtual microphone signal generated by each beamformer 21.

[0055] In addition, similarly, such as Figure 5 As shown in b1, with microphone group 1 positioned at the front end of the ceiling, and the occupants Hm1, Hm2, Hm3, Hm4, and Hm5 seated in the seats shown in the figure used as sound sources, and their speech signals separated as sound source signals, as follows: Figure 5 As shown in b2, the direction of the standard head position of the human body sitting in each seat when viewed from the microphone group 1 can also be set as the directional direction of the virtual microphone signal generated by each beamformer 21.

[0056] In addition, the number of microphones 11 in microphone group 1 is set to 4, and the number of virtual microphone signals generated by virtual microphone sound generation unit 2 and sound source signals separated by ICA processing unit 3 is set to 5. However, the number of microphones 11 in microphone group 1 can also be set to any number of 2 or more.

[0057] Furthermore, the number of virtual microphone signals generated by the virtual microphone sound generation unit 2 and the number of sound source signals separated by the ICA processing unit 3 can be set to any number greater than the number of microphones 11. Alternatively, in this case, the number of sound source signals separated by the ICA processing unit 3 can be less than the number of virtual microphone signals generated by the virtual microphone sound generation unit 2.

[0058] Furthermore, in this case, the directional direction of the virtual microphone signals generated by each beamformer 21 can be any direction as long as they are different from each other.

[0059] By setting the direction of the virtual microphone signal to the direction of the sound source, better sound source separation or faster convergence of the parameters (separation matrix) used in the sound source separation process can be expected.

[0060] Furthermore, in the above embodiments, the direction of the directional properties of the virtual microphone signals generated by each beamformer 21 of the virtual microphone sound generation unit 2 is fixed, but the direction of the directional properties of each virtual microphone signal can also be varied according to the location of the sound source.

[0061] That is, in this case, such as Figure 6 As shown, instead of the five beamformers 21 of the virtual microphone sound generation unit 2, there are five variable beamformers 22 that are variable directional beamformers, and a directional control unit 4 that detects the direction of the sound source and adjusts the directional direction of each variable beamformer 22 to the direction of the sound source. In addition, the ICA processing unit 3 takes the five virtual microphone signals Sig_Dir1, Sig_Dir2, Sig_Dir3, Sig_Dir4, and Sig_Dir5 input from each variable beamformer 22 of the virtual microphone sound generation unit 2 as inputs from five different microphones 11, and separates and outputs the five sound source signals Sig_AS1, Sig_AS2, Sig_AS3, Sig_AS4, and Sig_AS5, which are signals from five different sound sources, using the ICA method.

[0062] Here, for example, through Figure 3 In the configuration of beamformer 21 shown in figure a, by replacing the delay portion 211 with a variable delay portion with a variable delay time, a variable beamformer 22 with variable directivity can be constructed. In addition, in this case, the directivity control unit 4 can change the direction of directivity of the variable beamformer 22 by changing the delay time of each variable delay portion.

[0063] Furthermore, based on the difference in arrival time of the cross-correlation components included in the outputs of the four microphones 11, the direction of the sound source in the directional control unit 4 can be detected. Alternatively, based on the outputs of the four microphones 11, the output of a virtual microphone 11 with a single directionality for exploring the direction of the sound source can be generated, and the directionality of the virtual microphone 11 for exploring the direction of the sound source can be changed in a way that scans an acoustic space containing multiple sound sources, and the direction of the output sound of the virtual microphone 11 for exploring the direction of the sound source is detected as the direction of the sound source, thereby enabling the detection of the direction of the sound source in the directional control unit 4.

[0064] In addition, when only specific objects such as people are used as sound sources, the objects can be detected by using pattern matching of sensors such as cameras or LiDAR. In the directional control unit 4, the direction of the detected objects relative to the microphone group 1 is detected as the direction of each sound source.

[0065] In addition, the above uses Figure 3 The 'a' indicates an example of the configuration of the beamformer 21 in the case where the microphone array 1 and the beamformer 21 form a delay-addition type microphone array. However, as a configuration of the beamformer 21, other configurations such as a delay-subtraction type beamformer or a filter-sum type beamformer may also be used.

Claims

1. A sound source separation system, comprising performing blind source separation of each sound source signal from multiple mixed signals composed of two or more sound source signals using the ICA method, i.e., independent component analysis, characterized in that, having: n microphones, where n > 2; a virtual microphone signal generation mechanism that generates m virtual microphone signals from output signals of the n microphones, where m > n, the m virtual microphone signals being output signals of m virtual unidirectional microphones having directivity in respective different directions; and an ICA processing mechanism that separates L sound source signals, which are signals of different L sound sources, from the m virtual microphone signals as the plurality of mixed signals by an ICA method, where L > n and L ≤ m.

2. The sound source separation system according to claim 1, wherein a relationship between the L and the m satisfies L = m.

3. The sound source separation system according to claim 1, wherein directions of directivity of the m virtual unidirectional microphones are set at equiangular intervals.

4. The sound source separation system according to claim 1, wherein directions of directivity of the L virtual unidirectional microphones among the m virtual unidirectional microphones are set to be directions of the L sound sources.

5. The sound source separation system according to claim 1, wherein the virtual microphone signal generation mechanism is capable of changing the directions of directivity of the m virtual microphone signals, the sound source separation system has: a sound source detection mechanism that detects the directions of the L sound sources; and a directivity control mechanism that controls the virtual microphone signal generation mechanism so that the directions of directivity of the L virtual unidirectional microphones among the m virtual unidirectional microphones become the directions of the L sound sources detected by the sound source detection mechanism.

6. The sound source separation system according to any one of claims 1 to 5, wherein the n microphones are disposed in a vehicle to pick up sound in a passenger compartment of the vehicle.

Citation Information

Patent Citations

  • Sound source separation apparatus, program for sound source separation apparatus and sound source separation method

    JP2007215163A

  • Failure detection system and failure detection method

    JP2016032260A