Sound source separation system

The sound source separation system generates virtual microphone signals with different directivity to overcome the ICA method's limitations, enabling effective separation of more sound sources than real microphones, thus reducing costs and improving efficiency.

JP2026005536APending Publication Date: 2026-01-16ALPS ALPINE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024103963
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

The ICA method is limited in separating sound source signals from a number of sound sources greater than the number of microphones, leading to high costs due to the need for an equal or greater number of microphones.

Method used

A sound source separation system using n microphones generates m virtual microphone signals with different directivity directions, allowing the ICA method to separate L sound source signals where L>n, effectively utilizing fewer real microphones.

Benefits of technology

The system successfully separates sound source signals from a greater number of sound sources than the number of microphones, reducing costs and enhancing separation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026005536000001_ABST
    Figure 2026005536000001_ABST
Patent Text Reader

Abstract

To provide a "sound source separation system" for separating sound source signals of sound sources more than the number of microphones by an ICA method.SOLUTION: Five beam formers 21 of a virtual microphone sound generation part 2 generate five virtual microphone signals which are outputs of five virtual unidirectional microphones having directivities in different directions from outputs of four omnidirectional microphones 11 of a microphone set 1. The ICA processing section 3 receives the five virtual microphone signals input from the virtual microphone sound generation section 2 as inputs from the five different microphones 11, and separates and outputs five sound source signals which are signals of five different sound sources by the ICA method.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for performing blind source separation of each sound source signal from a mixed signal in which a plurality of sound source signals are mixed. [Background technology]

[0002] As a technique for performing blind source separation of each sound source signal from a mixed signal in which multiple sound source signals are mixed, a technique is known in which multiple microphones are placed in an acoustic space in which multiple sound sources exist, and each sound source signal is separated from the mixed signal in which the sound source signals picked up by each microphone are mixed, by using the ICA (Independent Component Analysis) method, which separates each sound source signal by taking advantage of the fact that each sound source signal is statistically independent (for example, Patent Document 1).

[0003] Also, a known technology related to the present invention is a microphone array technology that combines the outputs of multiple omnidirectional microphones to generate the output of a virtual unidirectional microphone and controls the direction of the directivity of this virtual unidirectional microphone (for example, Patent Document 2). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-215163 [Patent Document 2] Japanese Patent Application Laid-Open No. 2016-032260 Summary of the Invention [Problem to be solved by the invention]

[0005] The ICA method described above cannot be applied to separating sound source signals from a number of sound sources greater than the number of microphones, so it is necessary to provide microphones in a number equal to or greater than the number of sound sources whose sound source signals are to be separated, which results in high costs for separating multiple sound sources. Therefore, an object of the present invention is to use the ICA method to effectively separate sound source signals from sound sources whose number is greater than the number of microphones. [Means for solving the problem]

[0006] In order to achieve the above object, the present invention provides a sound source separation system that performs blind source separation of each sound source signal from a plurality of mixed signals in which two or more sound source signals are mixed by the ICA method (Independent Component Analysis), and the system includes n (where n>2) microphones, virtual microphone signal generation means that generates m virtual microphone signals, which are output signals of m (where m>n) virtual unidirectional microphones each having directivity in a different direction, from the output signals of the n microphones, and ICA processing means that separates L sound source signals, which are signals of L different sound sources (where L>n and L≦m), by the ICA method (Independent Component Analysis), using the m virtual microphone signals as the plurality of mixed signals.

[0007] In this sound source separation system, the relationship between L and m may be set so as to satisfy L=m. Here, in the above sound source separation system, the directivity directions of the m virtual unidirectional microphones may be set at equal angular intervals. Alternatively, in the above sound source separation system, the directivity directions of the L virtual unidirectional microphones among the m virtual unidirectional microphones may be set to the directions of the L sound sources. Alternatively, in the above sound source separation system, the virtual microphone signal generation means may be configured to be able to change the directivity directions of the m virtual microphone signals, and the sound source separation system may be provided with sound source detection means that detects the directions of the L sound sources, and directivity control means that controls the virtual microphone signal generation means so that the directivity directions of the L virtual unidirectional microphones among the m virtual unidirectional microphones are the directions of the L sound sources detected by the sound source detection means.

[0008] In the above sound source separation system, the n microphones may be installed in a vehicle so as to collect sounds inside the vehicle's cabin. According to the sound source separation system described above, m virtual microphone signals, which are output signals of m (where m>n) virtual unidirectional microphones having directivity in different directions, are generated from the outputs of n (where n>2) microphones, and L sound source signals, which are signals of L different sound sources (where L>n and L≦m), are separated from these m virtual microphone signals using the ICA method. Here, the outputs of the m virtual unidirectional microphones are equivalent to the outputs of m real unidirectional microphones, so the separation of L sound source signals using the ICA method is equivalent to separating the number of sound sources from the outputs of m (where L≦m) real unidirectional microphones into a number equal to or less than the number of unidirectional microphones.

[0009] Therefore, the ICA method can effectively separate the source signals of sound sources that are greater than the number of microphones. [Effects of the Invention]

[0010] As described above, according to the present invention, sound source signals of sound sources whose number is greater than the number of microphones can be satisfactorily separated using the ICA method. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing a configuration of a sound source separation system according to an embodiment of the present invention. [Figure 2]1 is a diagram showing the direction (sound collection axis) of unidirectional sound generated in an embodiment of the present invention. FIG. [Figure 3] FIG. 2 is a diagram illustrating an example of the configuration of a beamformer according to an embodiment of the present invention. [Figure 4] 1 is a diagram illustrating an application example of a sound source separation system according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating another example of the configuration of the sound source separation system according to the embodiment of the present invention. [Figure 6] FIG. 10 is a block diagram showing another example configuration of the sound source separation system according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, an embodiment of the present invention will be described. First, the first embodiment will be described. FIG. 1 shows the configuration of a sound source separation system according to this embodiment. As shown in the figure, the sound source separation system includes a microphone set 1, a virtual microphone sound generator 2, and an ICA processor 3. The microphone set 1 includes four omnidirectional microphones 11, M1, M2, M3, and M4, and the four microphones 11 are arranged at a distance from each other. The virtual microphone sound generation unit 2 has five beamformers 21: a 0° beamformer, a 45° beamformer, a 90° beamformer, a 135° beamformer, and a 180° beamformer, and each beamformer 21 receives the outputs of all four microphones 11.

[0013] As shown in FIG. 2a, a predetermined direction viewed from a predetermined reference point RP of the microphone set 1 is set to the 90° direction. The 0° beamformer 21 generates a virtual microphone signal Sig_0°, which is the output of a virtual unidirectional microphone having directivity in the 0° direction, from the outputs of the four microphones 11, and outputs it to the ICA processing unit 3. The 45° beamformer 21 generates a virtual microphone signal Sig_45°, which is the output of a virtual unidirectional microphone having directivity in the 45° direction, from the outputs of the four microphones 11, and outputs it to the ICA processing unit 3. The 90° beamformer 21 generates a virtual microphone signal Sig_45°, which is the output of a virtual unidirectional microphone having directivity in the 45° direction, from the outputs of the four microphones 11, and outputs it to the ICA processing unit 3. The 135° beamformer 21 generates a virtual microphone signal Sig_90°, which is the output of a virtual unidirectional microphone having directivity in the 90° direction from the output of the four microphones 11, and outputs this to the ICA processing unit 3. The 135° beamformer 21 generates a virtual microphone signal Sig_135°, which is the output of a virtual unidirectional microphone having directivity in the 135° direction from the outputs of the four microphones 11, and outputs this to the ICA processing unit 3. The 180° beamformer 21 generates a virtual microphone signal Sig_180°, which is the output of a virtual unidirectional microphone having directivity in the 180° direction from the outputs of the four microphones 11, and outputs this to the ICA processing unit 3.

[0014] When a microphone set 1 in which four microphones 11 are arranged in a one-dimensional manner is used as the microphone set 1, as shown in FIG. 2b, the reference point RP can be set to the center of the microphone set 1, the 90° direction can be set to the forward direction which is perpendicular to the arrangement direction of the four microphones 11, the 0° direction can be set to the right direction of the microphone set 1, and the 180° direction can be set to the left direction of the microphone set 1. Here, the configuration of each beam former 21 will be described by taking as an example a case where the microphone set 1 and the beam former 21 form a delay and sum type microphone array.

[0015] The five beamformers 21 have the same configuration, and FIG. 3a shows the configuration of the 45° beamformer 21 as a representative example. In this case, as shown in the figure, the 45° beamformer 21 has four delay units 211 that correspond one-to-one to each of the four microphones 11, and an adder unit 212 that adds the outputs of each delay unit 211 and outputs the result to the ICA processing unit 3 as a virtual microphone signal Sig_45°. Figure 3b shows the case where microphone set 1 is used, in which four microphones 11 are evenly arranged in one dimension. Sound arriving at each microphone 11 from a 45° direction has a delay difference depending on the position of each microphone 11, and this delay difference is such that, where d is the distance between adjacent microphones 11, sound arriving at microphone M2 is delayed by D2 = d × sin45° relative to the sound arriving at microphone M1, sound arriving at microphone M3 is delayed by D3 = 2d × sin45° relative to the sound arriving at microphone M1, and sound arriving at microphone M4 is delayed by D4 = 3d × sin45° relative to the sound arriving at microphone M1.

[0016] The four delay units 211 have delay times set so that the delays of the outputs of the four microphones 11 are aligned and so that the delays are aligned with the virtual microphone signals output by the other beamformers 21. For example, assuming that the delay time for aligning the delays with the virtual microphone signals output by the other beamformers 21 is Dx, the delay time of the delay unit 211 corresponding to M1 is set to D4+Dx, the delay time of the delay unit 211 corresponding to M2 is set to (D4-D2)+Dx, the delay time of the delay unit 211 corresponding to M1 is set to (D4-D3)+Dx, and the delay time of the delay unit 211 corresponding to M4 is set to Dx.

[0017] As a result, the virtual microphone signal Sig_45° output by the addition unit 212 is a sound obtained by adding together sounds arriving at each microphone 11 from a 45° direction in phase, and is a signal having directivity in the 45° direction. Returning to Figure 1, the ICA processing unit 3 receives five virtual microphone signals, Sig_0°, Sig_45°, Sig_90°, Sig_135°, and Sig_180°, from the virtual microphone sound generation unit 2 as inputs from five different microphones 11, and separates and outputs sound source signals Sig_AS1, Sig_AS2, Sig_AS3, Sig_AS4, and Sig_AS5 of five different sound sources using the ICA method.

[0018] As described above, according to this embodiment, outputs from five virtual unidirectional microphones with directivity in different directions are generated from the outputs of the four microphones 11, and five sound sources are separated from the outputs of these five virtual unidirectional microphones using the ICA method. Here, the outputs of the five virtual unidirectional microphones are equivalent to the outputs of five real unidirectional microphones, so the processing of the ICA processing unit 3 corresponds to separating five sound sources, the same number as the number of unidirectional microphones, from the outputs of the five real unidirectional microphones using the ICA method.

[0019] Therefore, the ICA processing unit 3 can effectively separate sound source signals from a greater number of sound sources than the number of microphones 11 in the microphone set 1. Next, an application example of the sound source separation system according to this embodiment will be described. The sound source separation system can be applied to, for example, separating multiple sound sources inside an automobile. In this case, the microphone set 1 is placed in a position where it can collect a wide range of sounds inside the automobile. Such a position where it can collect a wide range of sounds can be, for example, the front end of the ceiling inside the automobile, on the dashboard, or in the center of the ceiling, as shown in Figures 4a and 4b.

[0020] Furthermore, sound sources and sound source signals inside an automobile include in-vehicle speakers and their output sounds, in-vehicle devices and their alarm sounds, passengers and their speech sounds, and so on. Here, when the sound source separation system is applied to sound source separation inside an automobile, the sound source signals separated by the sound source separation system can be used as noise source signals for active noise control, which controls the sound of sound sources that are unnecessary for each passenger in the automobile so that they do not hear it as noise.

[0021] The embodiments of the present invention have been described above. Here, in the above, the directional directions of the virtual microphone signals generated by each Beaformer 21 of the virtual microphone sound generation unit 2 are set to 0°, 45°, 90°, 135°, and 180°, but the directional directions of the virtual microphone signals generated by each Beaformer 21 may be set to any direction as long as they are different from each other.

[0022] For example, if the direction of the sound source from which the sound source signal is to be separated is known in advance, the direction of the directivity of the virtual microphone signal generated by each Beaformer 21 may be set to the direction of the known sound source. That is, for example, as shown in Figure 5a1, when microphone set 1 is placed at the front end of the ceiling, each of the in-vehicle speakers SP1, SP2, SP3, SP4, and SP5 in the figure is used as a sound source, and the speaker output sound is separated as a sound source signal, as shown in Figure 5a2, the direction of each of the in-vehicle speakers SP1, SP2, SP3, SP4, and SP5 as seen from microphone set 1 may be set as the directional direction of the virtual microphone signal generated by each beamformer 21.

[0023] Similarly, as shown in Figure 5b1, when microphone set 1 is placed at the front end of the ceiling, passengers Hm1, Hm2, Hm3, Hm4, and Hm5 seated in each seat in the figure are used as sound sources, and the passengers' speech sounds are separated as sound source signals, the direction of the standard head position of the human body seated in each seat as seen from microphone set 1 may be set as the directional direction of the virtual microphone signal generated by each Beformer 21, as shown in Figure 5b2. In the above, the number of microphones 11 in microphone set 1 is set to four, and the number of virtual microphone signals generated by virtual microphone sound generation unit 2 and sound source signals separated by ICA processing unit 3 is set to five, but the number of microphones 11 in microphone set 1 may be any number greater than or equal to two.

[0024] Furthermore, the number of virtual microphone signals generated by the virtual microphone sound generation unit 2 and the number of sound source signals separated by the ICA processing unit 3 may be any number greater than the number of microphones 11. In this case, the number of sound source signals separated by the ICA processing unit 3 may be less than the number of virtual microphone signals generated by the virtual microphone sound generation unit 2.

[0025] In this case, the directivity directions of the virtual microphone signals generated by each beamformer 21 may be any directions as long as they are mutually different. By setting the direction of the directivity of the virtual microphone signal to the direction of the sound source in this way, better sound source separation and quick convergence of the parameters (separation matrix) used in the sound source separation process can be expected. In addition, in the above embodiment, the direction of the directivity of the virtual microphone signal generated by each beamformer 21 of the virtual microphone sound generation unit 2 is fixed, but the direction of the directivity of each virtual microphone signal may be variable depending on the position of the sound source. 6, in this case, five variable beamformers 22, which are beamformers with variable directivity, are provided instead of the five beamformers 21 of the virtual microphone sound generation unit 2, and a directivity direction control unit 4 is provided that detects the direction of a sound source and adjusts the direction of the directivity of each variable beamformer 22 to the direction of the sound source. The ICA processing unit 3 then uses the five virtual microphone signals Sig_Dir1, Sig_Dir2, Sig_Dir3, Sig_Dir4, and Sig_Dir5 input from each variable beamformer 22 of the virtual microphone sound generation unit 2 as inputs from five different microphones 11, and separates and outputs five sound source signals Sig_AS1, Sig_AS2, Sig_AS3, Sig_AS4, and Sig_AS5, which are signals of the five different sound sources, by the ICA method.

[0026] Here, variable beamformer 22 with variable directivity can be configured, for example, by replacing delay unit 211 in the configuration of beamformer 21 shown in Fig. 3a with a variable delay unit with a variable delay time. In this case, directivity direction control unit 4 can change the direction of the directivity of variable beamformer 22 by changing the delay time of each variable delay unit.

[0027] Furthermore, the detection of the direction of a sound source in the directivity direction control unit 4 can be performed based on the difference in arrival time of components that are highly correlated and are included in the outputs of the four microphones 11 at each microphone 11. Alternatively, the detection can be performed by generating the output of a unidirectional virtual microphone 11 for sound source direction searching from the outputs of the four microphones 11, and detecting the direction of the largest output sound level of the virtual microphone 11 for sound source direction searching as the direction of the sound source while changing the directivity direction of this virtual microphone 11 for sound source direction searching so as to scan an acoustic space containing multiple sound sources.

[0028] In addition, when only specific objects such as humans are considered to be sound sources, the objects may be detected by pattern matching using sensors such as cameras or LiDAR, and the directional direction control unit 4 may detect the direction of the detected objects relative to the microphone set 1 as the direction of each sound source. In addition, in the above, an example configuration of the beamformer 21 when a delay-addition type microphone array is formed using the microphone set 1 and the beamformer 21 has been shown using Figure 3a, but other configurations may be adopted for the beamformer 21, such as a delay-subtraction type beamformer or a filter-and-sum type beamformer. [Explanation of symbols]

[0029] 1...microphone set, 2...virtual microphone sound generation unit, 3...ICA processing unit, 4...directivity direction control unit, 11...microphone, 21...beamformer, 22...variable beamformer, 211...delay unit, 212...adder unit.

Claims

1. A sound source separation system that performs blind source separation on each sound source signal from a plurality of mixed signals in which two or more sound source signals are mixed, using an ICA method (Independent Component Analysis method), n microphones (where n>2); a virtual microphone signal generating means for generating m virtual microphone signals, which are output signals from m (where m>n) virtual unidirectional microphones each having directivity in a different direction, from the output signals from the n microphones; and an ICA processing means for separating L sound source signals, which are signals of L different sound sources (where L>n and L≦m), by an ICA method (Independent Component Analysis) using the m virtual microphone signals as the plurality of mixed signals.

2. 2. The sound source separation system according to claim 1, A sound source separation system characterized in that the relationship between L and m satisfies L=m.

3. 2. The sound source separation system according to claim 1, A sound source separation system, characterized in that the directional directions of the m virtual unidirectional microphones are set at equal angular intervals.

4. 2. The sound source separation system according to claim 1, A sound source separation system, characterized in that the directivity directions of the L virtual unidirectional microphones among the m virtual unidirectional microphones are set to the directions of the L sound sources.

5. 2. The sound source separation system according to claim 1, the virtual microphone signal generation means is capable of changing the direction of directivity of the m virtual microphone signals; The sound source separation system a sound source detection means for detecting the directions of the L sound sources; and a directivity control means for controlling the virtual microphone signal generation means so that the directivity directions of the L virtual unidirectional microphones among the m virtual unidirectional microphones are the directions of the L sound sources detected by the sound source detection means.

6. 6. A sound source separation system according to claim 1, 2, 3, 4 or 5, A sound source separation system, characterized in that the n microphones are installed in a vehicle so as to collect sounds inside the vehicle's cabin.

Citation Information

Patent Citations

  • Sound source separation apparatus, program for sound source separation apparatus and sound source separation method

    JP2007215163A

  • Failure detection system and failure detection method

    JP2016032260A