Sound acquisition system, sound acquisition method, and storage medium

By introducing a multi-microphone array and dynamic beamformer switching technology into a traditional beamforming processing unit, the problem of not being able to capture the voices of multiple speakers in traditional technology is solved, and efficient voice acquisition of multiple speakers is achieved.

CN116490924BActive Publication Date: 2026-02-10AUDIO TECHNICA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180068862.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-11
Filing Date
2021-10-12
Publication Date
2026-02-10
Estimated Expiration
2041-10-12

AI Technical Summary

Technical Problem

Traditional beamforming processing units cannot effectively capture the speech of other speakers when the target is pointing in the direction of the speaker.

Method used

Employing multiple microphone arrays and signal processing equipment, and combining first and second beamformers with sound source direction detection and directional control, the beamformer is dynamically switched to accommodate multiple speakers, enabling flexible acquisition of different sound sources.

Benefits of technology

It enables continuous and efficient acquisition of each speaker's voice when switching between multiple speakers, avoiding acquisition interruptions and changes in sound level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116490924B_ABST
    Figure CN116490924B_ABST
Patent Text Reader

Abstract

The sound collection system (S) has: a microphone array (1) including a plurality of microphones (2); a first beamformer (152) that outputs a first signal in which, among a plurality of sound signals based on sound arriving at the plurality of microphones (2), sound signals based on sound arriving from a direction within a first range are emphasized more than sound signals based on sound arriving from other directions; a second beamformer (153) that outputs a second signal in which, among the plurality of sound signals, sound signals based on sound arriving from a direction within a second range are emphasized more than sound signals based on sound arriving from other directions; a sound source direction detection unit (151) that detects a direction of a sound source that generates sound arriving at the plurality of microphones (2); and a directivity control unit (155) that, during the first beamformer (152) is outputting the first signal, if a change angle per unit time of the direction of the sound source detected by the sound source direction detection unit (151) is determined to be above a threshold value, causes the second beamformer (153) to output the second signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a sound acquisition system, a sound acquisition method, and a storage medium. Background Technology

[0002] A beamforming processing unit is known that uses the phase difference in audio signals observed by multiple microphones to perform beamforming processing in order to acquire sound when the target of sound acquisition is pointing towards the sound source (see, for example, Patent Document 1).

[0003] Existing technology

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2013-201525 Summary of the Invention

[0006] The problem the invention aims to solve

[0007] In traditional beamforming processing units, the sound source is assumed to be a single source. Therefore, in traditional beamforming processing units, if another speaker speaks while the sound acquisition target is pointing in the direction of the speaker, there is a problem that the other speaker's speech cannot be acquired.

[0008] Therefore, the present invention is made in view of these points, and its purpose is to enable the acquisition of speech from multiple speakers.

[0009] Solution for solving the problem

[0010] A sound acquisition system according to a first aspect of the present invention includes: a microphone array including a plurality of microphones; a first beamformer for outputting a first signal, the first signal being obtained by emphasizing a sound signal based on a direction from a first range more than a sound signal based on sound from other directions among a plurality of sound signals based on sound arriving at the plurality of microphones; a second beamformer for outputting a second signal, the second signal being obtained by emphasizing a sound signal based on a direction from a second range more than a sound signal based on sound from other directions among the plurality of sound signals; a sound source direction detection unit for detecting the direction of a sound source that generates sound arriving at the plurality of microphones; and a directivity control unit for causing the second beamformer to output the second signal when, during the period when the first beamformer is outputting the first signal, the change angle of the direction of the sound source detected by the sound source direction detection unit per unit time is determined to be equal to or greater than a threshold.

[0011] While the first beamformer is outputting the first signal, if the change angle of the direction of the sound source per unit time is determined to be less than a threshold, the directivity control unit can cause the first beamformer to continue outputting the first signal even though the first range has been changed.

[0012] While the first beamformer is outputting the first signal, if the change angle is determined to be equal to or greater than a threshold, the directivity control unit may reduce the output level of the first signal.

[0013] The directional control unit can reduce the output level of the first signal by using an attenuation factor based on the elapsed time after the change angle is determined to be equal to or greater than a threshold.

[0014] The directional control unit can increase the output level of the second signal while decreasing the output level of the first signal.

[0015] The directional control unit can increase the output level of the second signal at a rate greater than the rate of change used to decrease the output level of the first signal.

[0016] If it is determined that the direction of the sound source is not included in the first range, the directivity control unit can cause the second beamformer to output the second signal.

[0017] The directional control unit can determine a second range such that the second range includes the direction of the sound source before the second beamformer outputs the second signal.

[0018] While the second beamformer is outputting the second signal, if the change angle of the direction of the sound source detected by the sound source direction detection unit per unit time is determined to be equal to or greater than a threshold, the directivity control unit can cause the first beamformer to output the first signal.

[0019] The sound acquisition system may further include a storage unit for storing beamforming coefficients and the direction of the sound source detected by the sound source direction detection unit in association with each other. The directivity control unit may use the beamforming coefficients stored in the storage unit in association with the direction of the sound source detected by the sound source direction detection unit to cause the first beamformer or the second beamformer to output the first signal or the second signal.

[0020] The storage unit can store the directions of sound sources previously detected by the sound source direction detection unit and the beamformer coefficients previously calculated by the directivity control unit based on those directions in association with each other. If it is determined that the direction of a newly detected sound source by the sound source direction detection unit is the same as the direction of a sound source previously detected and stored in the storage unit, the directivity control unit can use the beamformer coefficients stored in association with the direction of the previously detected sound source.

[0021] A sound acquisition method according to a second aspect of the present invention includes the following steps: outputting a first signal, the first signal being obtained by emphasizing a sound signal based on a direction from a first range compared to a sound signal based on sound from other directions among a plurality of sound signals based on sound arriving at a plurality of microphones; detecting the direction of a sound source that generates sound arriving at the plurality of microphones; and, while outputting the first signal, if it is determined that the angle of change of the direction of the sound source per unit time is equal to or greater than a threshold, outputting a second signal, the second signal being obtained by emphasizing a sound signal based on a direction from a second range compared to a sound signal based on sound from other directions among the plurality of sound signals.

[0022] According to a storage medium of a stored program of a third aspect of the present invention, the program causes a computer to function as: a first beamformer for outputting a first signal, the first signal being obtained by emphasizing a sound signal based on a direction from a first range compared to a sound signal based on sound from other directions among a plurality of sound signals based on sound arriving at a plurality of microphones; a second beamformer for outputting a second signal, the second signal being obtained by emphasizing a sound signal based on a direction from a second range compared to a sound signal based on sound from other directions among the plurality of sound signals; a sound source direction detection unit for detecting the direction of a sound source that generates sound arriving at the plurality of microphones; and a directivity control unit for causing the second beamformer to output the second signal when, during the period when the first beamformer is outputting the first signal, the change angle of the direction of the sound source detected by the sound source direction detection unit per unit time is determined to be equal to or greater than a threshold.

[0023] The effects of the invention

[0024] According to the present invention, the voices of multiple speakers can be collected. Attached Figure Description

[0025] Figure 1 This is a diagram illustrating the general structure of the sound acquisition system S according to this embodiment.

[0026] Figure 2 The diagram illustrates the operation of the sound acquisition system S in acquiring multiple voices generated by multiple speakers, presented as a time series.

[0027] Figure 3 This is a diagram used to illustrate the structure of the sound acquisition system S.

[0028] Figure 4 This is a diagram illustrating the structure of the first beamformer 152.

[0029] Figure 5 This is a flowchart showing the processing flow performed by the beamforming processing unit 15 to determine whether a new sound source has been detected.

[0030] Figure 6 This is a flowchart showing the processing flow performed by the beamforming processing unit 15 for controlling the beamformer based on the detection of a new sound source. Detailed Implementation

[0031] <Summary of the sound acquisition system according to this embodiment>

[0032] Figure 1 This is a diagram illustrating the general structure of the sound acquisition system S according to this embodiment. Figure 1 This is a side view showing the interior of space R. For example, space R is a room in a building, but it is not limited to this; it can also be a corridor, lounge, stairwell, etc. Figure 1 As shown, the sound acquisition system S is installed on the inner top surface of space R, and speakers A1, A2 and A3 remain in space R. Figure 1 Speech words B1, B2, and B3 in the text are generated by speakers A1, A2, and A3, respectively. Figure 1 In this configuration, the sound acquisition system S is installed on the inner top surface of the space R. It should be noted that the sound acquisition system S can also be installed on the inner side or inner bottom surface of the space R.

[0033] The sound acquisition system S includes a microphone array and a signal processing device. The microphone array includes multiple microphones. The signal processing device includes multiple beamformers that process the sound arriving at the microphone array. The sound acquisition system S uses beamforming coefficients corresponding to the directions of the sound sources detected by the multiple beamformers to perform beamforming, thereby simulating the formation of multiple directional microphones. The beamformer coefficients will be described later.

[0034] Figure 2 The diagram illustrates the operation of the sound acquisition system S in acquiring multiple voices generated by multiple speakers, presented as a time series. Figure 2 The horizontal axis in the graph represents time. Figure 2The vertical axis shows “Speaker A1”, “Speaker A2” and “Speaker A3”, which indicate the duration of speech B1, B2 and B3 generated by speakers A1, A2 and A3, respectively. Figure 2 The "first beamformer" and "second beamformer" shown on the vertical axis indicate the duration of beamforming processing performed by the first and second beamformers included in the sound acquisition system S, and the speech with the direction of the sound source identified in the beamforming processing. "Output sound" indicates the speech acquired by the sound acquisition system S and output to an external device. The external device is, for example, a computer with a router or storage medium connected to a communication network.

[0035] like Figure 2 As shown, speaker A1 generates speech B1 from time T1 to time T3, speaker A2 generates speech B2 from time T2 to time T5, and speaker A3 generates speech B3 from time T4 to time T6. At time T1, the sound acquisition system S detects speech B1 to begin beamforming processing using the first beamformer and identifies the direction of the sound source of speech B1. At time T2, the sound acquisition system S detects speech B2 from a different direction than speech B1 to begin beamforming processing using the second beamformer, thereby identifying the direction of the sound source of speech B2. At time T3, the sound acquisition system S stops beamforming processing using the first beamformer.

[0036] At time T4, the sound acquisition system S detects the direction of the sound source of speech B3 and begins beamforming processing using the first beamformer. At time T5, the sound acquisition system S stops beamforming processing using the second beamformer. As a result, the sound acquisition system S acquires speech B1 from time T1 to time T2, and acquires both speech B1 and speech B2 from time T2 to time T3. The sound acquisition system S acquires speech B2 from time T3 to time T4, and acquires both speech B2 and speech B3 from time T4 to time T5. From time T5 to time T6, the sound acquisition system S acquires speech B3.

[0037] Because the sound acquisition system S has multiple beamformers as described above, it simulates the same situation as multiple narrow directional microphones pointing towards each sound source and acquires sound. Furthermore, even when the number of speakers exceeds the number of beamformers, and the speaker generating the speech is switched, the sound acquisition system S can acquire the speech of multiple speakers without interruption by switching multiple beamformers.

[0038] although Figure 2The sound acquisition system S stops beamforming processing along with the cessation of the speaker's generated speech, but beamforming processing can continue even after the cessation of the speaker's generated speech. For example, the sound acquisition system S may stop beamforming processing using the first beamformer that started at time T1, not at time T3, but at a time after a predetermined time interval from time T3. Furthermore, the sound acquisition system S may continue beamforming processing at time T3 without stopping the beamforming processing using the first beamformer. In this case, when the source direction of speech B3 is detected at time T4, the sound acquisition system S switches the beamforming direction using the first beamformer to the source direction of speech B3.

[0039] <Structure of Sound Acquisition System S>

[0040] Figure 3 This is a diagram illustrating the structure of a sound acquisition system S. The sound acquisition system S includes a microphone array 1 and a signal processing device 10. The microphone array 1 includes multiple microphones 2 (microphones 2a, 2b, 2c, and 2d). The multiple microphones 2 output electrical signals based on the arriving sound. The signal processing device 10 processes the electrical signals output from the multiple microphones 2 to increase the directivity towards the sound source, thereby emphasizing and outputting the sound generated from the sound source.

[0041] The signal processing device 10 includes an input unit 11, a first attenuation unit 12, a second attenuation unit 13, an output unit 14, and a beamforming processing unit 15. The input unit 11 includes, for example, a preamplifier and an analog-to-digital (A / D) converter. The input unit 11 converts multiple analog electrical signals input from each of the multiple microphones 2 into multiple digital signals to generate multiple sound signals. The input unit 11 generates, for example, multiple amplified signals obtained by amplifying the analog electrical signals input from the respective multiple microphones 2. The input unit 11 converts the multiple amplified signals into multiple digital signals to generate multiple sound signals. The input unit 11 outputs the generated multiple sound signals to the beamforming processing unit 15.

[0042] The first attenuation unit 12 and the second attenuation unit 13 reduce or increase the level of the signal input from the beamforming processing unit 15. The first attenuation unit 12 and the second attenuation unit 13 reduce or increase the level of the signal output from the beamforming processing unit 15 based on the attenuator gain obtained from the beamforming processing unit 15. The attenuator gain corresponds to an attenuation factor, which is the amount by which the signal level is reduced or increased relative to the signal level before it is reduced or increased in the first attenuation unit 12 and the second attenuation unit 13. The first attenuation unit 12 and the second attenuation unit 13 output the signal obtained by reducing or increasing the signal level to the output unit 14.

[0043] The output unit 14 outputs signals input from the first attenuation unit 12 and the second attenuation unit 13. The output unit 14 generates an output sound signal obtained by adding the signal output from the first attenuation unit 12 to the signal output from the second attenuation unit 13, and outputs the generated output sound signal. The output unit 14 includes, for example, a digital-to-analog (D / A) converter, and converts the digital output sound signal into an analog signal to output the converted analog signal.

[0044] The beamforming processing unit 15 includes a sound source direction detection unit 151, a first beamformer 152, a second beamformer 153, a storage unit 154, and a directivity control unit 155. The beamforming processing unit 15 may be configured, for example, by a processor for digital signal processing.

[0045] The sound source direction detection unit 151 detects the direction of the sound source that generates the sound arriving at the plurality of microphones 2. For example, if the microphone array 1 is mounted on the inner top surface of the space, the direction of the sound source is represented by the angle between a) a straight line extending vertically from the center position of the microphone array 1 and b) a straight line connecting the position of the microphone 2 and the position of the sound source. The sound source direction detection unit 151 detects the direction of the sound source, for example, based on the difference in the time when the sound arrives at each of the plurality of microphones, by using a delay-sum array method. The sound source direction detection unit 151 notifies the directional control unit 155 of the detected sound source direction.

[0046] Among multiple sound signals collected by multiple microphones 2, the first beamformer 152 outputs a first signal, which is obtained by emphasizing sound signals based on directions within a first range compared to sound signals based on sound from other directions. The first range is a range defined around the direction of a first sound source notified by the sound source direction detection unit 151. The size of the first range is determined, for example, by the number of microphones 2 and the beamforming coefficient set for the first beamformer 152.

[0047] The first beamformer 152 generates a first signal by combining multiple sound signals input from the input unit 11. Using beamformer coefficients input from the directivity control unit 155, the first beamformer 152 generates multiple sound signals such that the level of the sound signal based on sound from a first range direction is higher than the level of the sound signal based on sound from other directions. The first beamformer 152 generates the first signal by combining the generated multiple sound signals. The first beamformer 152 outputs the generated first signal to the first attenuation unit 12.

[0048] Figure 4This is a diagram illustrating the structure of the first beamformer 152. The first beamformer 152 includes a plurality of variable delay sections 161 (variable delay sections 161a, 161b, 161c and 161d), a plurality of gain adjustment sections 162 (gain adjustment sections 162a, 162b, 162c and 162d) and an adder section 163.

[0049] The variable delay unit 161 delays multiple sound signals acquired from the input unit 11 based on a delay amount input from the directivity control unit 155. The beamforming coefficient corresponds to the delay amount, which is a time interval corresponding to the difference in distance from the sound source to each of the multiple microphones 2 (hereinafter referred to as "propagation distance"), and the variable delay unit 161 delays the sound signals, for example, based on the delay amount of the beamforming coefficient. By causing the variable delay unit 161 to delay the sound signals by a time interval corresponding to the difference in propagation distance, the timing difference of the multiple sounds that have reached the multiple microphones 2 is corrected, thereby bringing the multiple sound signals from the direction with the strongest directivity of the first beamformer 152 into the same phase.

[0050] The gain adjustment unit 162 adjusts the signal gain after the variable delay unit 161 has caused a delay. The beamformer coefficients correspond to the gain, and the gain adjustment unit 162 amplifies or attenuates the signal delayed by the variable delay unit 161, for example, based on the gain corresponding to the beamformer coefficients. The gain of each of the plurality of gain adjustment units 162 is determined according to the beamformer coefficients.

[0051] The adder 163 adds multiple signals generated by multiple gain adjustment units 162. The signal output from the gain adjustment unit 162 corresponding to the direction within the first range is greater than the signal output from the other gain adjustment units 162. Therefore, the adder 163 adds multiple signals to generate a first signal, which is obtained by emphasizing the sound signal based on the direction within the first range more than the sound signal based on the sound from other directions.

[0052] Return to reference Figure 3 Among the multiple sound signals input from the input unit 11, the second beamformer 153 outputs a second signal, which is obtained by emphasizing sound signals based on directions within a second range compared to sound signals based on sounds from other directions. The second range is a range defined around the direction of a second sound source notified by the sound source direction detection unit 151. The size of the second range is determined, for example, by the number of the multiple microphones 2 and the beamforming coefficient set for the second beamformer 153.

[0053] The second beamformer 153 generates a second signal by combining multiple sound signals input from the input unit 11. The second beamformer 153 uses beamformer coefficients input from the directivity control unit 155 to generate multiple sound signals, such that the level of the sound signal based on sound from a direction within a second range is greater than the level of the sound signals based on sound from other directions. The second beamformer 153 generates the second signal by combining the generated multiple sound signals. The second beamformer 153 outputs the generated second signal to the second attenuation unit 13. The structure of the second beamformer 153 is similar to... Figure 4 The structure of the first beamformer 152 shown is the same.

[0054] Storage unit 154 includes storage media such as random access memory (RAM) and solid-state drive (SSD). Storage unit 154 stores attenuation coefficients used to calculate the attenuator gain used by the first attenuator unit 12 and the second attenuator unit 13. Storage unit 154 stores beamformer coefficients associated with the direction of the sound source.

[0055] The storage unit 154 can store the direction of the sound source detected by the sound source direction detection unit 151 and the beamformer coefficients in association with each other. For example, the storage unit 154 can store a) the direction of the sound source previously detected by the sound source direction detection unit 151 and b) the beamformer coefficients previously calculated by the directivity control unit 155 based on these directions in association with each other.

[0056] In addition, the storage unit 154 stores programs for enabling the processor to function as the sound source direction detection unit 151, the first beamformer 152, the second beamformer 153, and the directivity control unit 155.

[0057] The directivity control unit 155 determines the beamforming coefficients of the first beamformer 152 and the second beamformer 153 based on the direction of the sound source notified by the sound source direction detection unit 151, and controls the first beamformer 152 and the second beamformer 153. For example, the directivity control unit 155 uses the beamforming coefficients to cause the first beamformer 152 or the second beamformer 153 to output a first signal or a second signal, and these beamforming coefficients are stored in the storage unit 154 in association with the direction of the sound source detected by the sound source direction detection unit 151. Furthermore, the directivity control unit 155 controls the attenuation factors of the first attenuation unit 12 and the second attenuation unit 13.

[0058] If, based on the direction of the sound source notified by the sound source direction detection unit 151, it is determined that the sound source generating the sound has changed, the directivity control unit 155 changes the beamforming coefficients set for the first beamformer 152 and the second beamformer 153, as well as the attenuation factors of the first attenuation unit 12 and the second attenuation unit 13. To detect that the sound source has changed or moved, the directivity control unit 155 stores angle information indicating the direction of the sound source notified by the sound source direction detection unit 151 in the storage unit 154. The directivity control unit 155 calculates the change angle, which is the difference between the angle detected by the sound source direction detection unit 151 at the current moment and the angle indicated by the angle information stored in the storage unit 154 a unit of time prior (hereinafter referred to as the "previous angle").

[0059] If the angle of change per unit time (which is the difference between the current moment and the immediate preceding moment) is equal to or greater than a threshold, the directivity control unit 155 determines that the sound source generating the sound has changed. On the other hand, if the angle of change is less than the threshold, the directivity control unit 155 determines that the sound source generating the sound has moved. For example, the unit time is 0.1 seconds. The threshold is a value set based on the minimum directional difference between multiple sound sources, and is, for example, 10 degrees.

[0060] If a new sound source is detected, the directivity control unit 155 uses an unused beamformer from among multiple beamformers to perform signal processing within the range including the new sound source. Specifically, if, while the first beamformer 152 is outputting a first signal, it is determined that the angle of change per unit time of the direction of the sound source detected by the sound source direction detection unit 151 is equal to or greater than a threshold, the directivity control unit 155 causes the second beamformer 153 to output a second signal. That is, if it is determined that the direction of the sound source detected by the sound source direction detection unit 151 is the direction of a new sound source not included in the first range, the directivity control unit 155 causes the second beamformer 153 to output a second signal.

[0061] The directivity control unit 155 determines a second range such that it includes the direction of the newly detected sound source before the second beamformer 153 outputs the second signal. The directivity control unit 155 calculates beamformer coefficients corresponding to the determined second range and sets the calculated beamformer coefficients for the plurality of gain adjustment units 162, thereby causing the second beamformer 153 to output the second signal. By operating the directivity control unit 155 in this manner, when a new sound source begins to generate sound, the signal processing device 10 can acquire sound in a directivity state with a direction toward the new sound source.

[0062] On the other hand, if, during the period when the first beamformer 152 is outputting the first signal, it is determined that the angle of change of the direction of the sound source per unit time is less than a threshold, the directivity control unit 155 causes the first beamformer 152 to continue outputting the first signal even though the first range has changed. In other words, the directivity control unit 155 determines that the same sound source as the one detected at the current moment has been detected at the immediate previous moment, and continues to use the beamformer that acquires sound in a state of directivity towards a range including the detected sound source.

[0063] As described above, even if the detected sound source is determined to be in a different position than immediately before, if the angle of change of the direction of the sound source per unit time is less than a threshold, the directivity control unit 155 will not switch the currently operating beamformer. In other words, even if the position of the sound source has changed, if the angle of change of the direction of the sound source per unit time is less than the threshold, the directivity control unit 155 determines that the same sound source as immediately before has been detected. Then, the directivity control unit 155 changes the direction of directivity by changing the beamformer coefficient to be set for the operating beamformer based on the angle of change. The directivity control unit 155 operating in this way allows the signal processing device to acquire sound without switching the beamformer when generating speech, for example, while the speaker is moving, thus preventing changes in the level of the acquired sound.

[0064] If another new sound source (a sound source in a third direction) is detected while the second beamformer 153 is outputting a second signal, the directivity control unit 155 uses the first beamformer 152 to acquire the sound generated by the detected new sound source. If it is determined that the change angle of the direction of the sound source detected by the sound source direction detection unit 151 per unit time is equal to or greater than a threshold while the second beamformer 153 is outputting a second signal, the directivity control unit 155 causes the first beamformer 152 to output a first signal.

[0065] If the direction of a newly detected sound source is the same as the direction of a previously detected sound source, the directivity control unit 155 can use beamformer coefficients associated with the direction of the previously detected sound source. Specifically, if it is determined that the direction (third direction) of a newly detected sound source by the sound source direction detection unit 151 is the same as the previously detected first direction, the directivity control unit 155 uses beamformer coefficients associated with the first direction stored in the storage unit 154 to cause the first beamformer 152 to output a first signal. Since the directivity control unit 155 uses beamformer coefficients stored in the storage unit 154, the time required for the beamformer to start operation can be reduced.

[0066] As described above, whenever a new sound source is detected, the directivity control unit 155 alternately uses the first beamformer 152 and the second beamformer 153. As a result, even if there is a certain amount of time during which sound is generated from multiple sound sources simultaneously, the signal processing device 10 can acquire sound generated from multiple sound sources when the sound source is switched.

[0067] Next, the operation of the directional control unit 155 controlling the first attenuation unit 12 and the second attenuation unit 13 will be described. The directional control unit 155 calculates the attenuator gain of the first attenuation unit 12 and the second attenuation unit 13 based on the elapsed time after the moment a new sound source is detected. The directional control unit 155 adjusts the level of the signal output from the first attenuation unit 12 and the second attenuation unit 13 by setting the calculated attenuator gain for the first attenuation unit 12 and the second attenuation unit 13.

[0068] If a new sound source is detected, the directivity control unit 155 increases the output level of the attenuation unit downstream of the beamformer corresponding to the range including the new sound source. Conversely, the directivity control unit 155 decreases the output level of the attenuation unit downstream of the beamformer corresponding to the range not including the new sound source. The following describes a situation where a first range corresponding to a first signal output by a first beamformer stops including a sound source over time, and a second range corresponding to a second signal output by a second beamformer gradually changes over time to include a new sound source. In this case, the attenuation unit downstream of the first beamformer for reducing the signal level is the first attenuation unit 12, while the attenuation unit downstream of the second beamformer for increasing the signal level is the second attenuation unit 13.

[0069] If, during the period when the first beamformer 153 is outputting the first signal, it is determined that the change angle is equal to or greater than a threshold, the directivity control unit 155 reduces the output level of the first signal. When reducing the output level of the first signal, the directivity control unit 155 uses an attenuation factor based on the elapsed time after the determination that the change angle is equal to or greater than the threshold to reduce the output level of the first signal. The directivity control unit 155 operates the first attenuation unit with an attenuation factor corresponding to the attenuator gain determined based on the attenuation coefficient and the elapsed time.

[0070] For example, the attenuator gain is determined by multiplying the attenuation coefficient C by the elapsed time T. For example, the attenuation coefficient C is a negative fixed value. In this way, the attenuator gain calculated based on the elapsed time is set for the first attenuation unit 12. This allows the directivity control unit 155 to gradually attenuate the first signal, thus preventing the sound generated from the sound source from suddenly disappearing.

[0071] Furthermore, the directivity control unit 155 increases the output level of the second signal output from the second beamformer 153. For example, the directivity control unit 155 increases the output level of the second signal at a rate greater than the rate of change used to decrease the output level of the first signal. The rate of change is determined by the amount of change in output level per unit time. As described above, since the directivity control unit 155 increases the output level of the second signal at a rate greater than the rate of change used to decrease the output level of the first signal, the output level of the second signal increases in a short time. Therefore, the signal processing device 10 can output the voice of a person who has started speaking at a sufficient volume from the beginning. The directivity control unit 155 can increase the output level of the second signal while decreasing the output level of the first signal. Because the directivity control unit 155 operates in this way, a silent period between the first and second signals can be prevented when the signal processing device 10 switches outputs between the first and second signals.

[0072] <Process for Detecting and Processing New Sound Sources>

[0073] Figure 5 This is a flowchart illustrating the processing flow performed by the beamforming processing unit 15 for determining whether a new sound source has been detected. The sound source direction detection unit 151 acquires multiple sound signals amplified by the input unit 11 (S11). The sound source direction detection unit 151 detects the sound source direction based on the acquired multiple sound signals (S12).

[0074] The directional control unit 155 calculates the difference between the sound source direction detected by the sound source direction detection unit 151 at the current moment and the sound source direction at the immediate preceding moment (S13). If the calculated difference between the sound source directions is equal to or greater than a threshold ("Yes" in S14), the directional control unit 155 determines that a new sound source has been detected (S15). If the calculated difference between the sound source directions is less than the threshold ("No" in S14), the directional control unit 155 determines that the same sound source as the one at the immediate preceding moment has been detected (S16).

[0075] If the operation to terminate the detection process of a new sound source has not yet been performed ("No" in S17), the beamforming processing unit 15 repeats the processing from S11 to S17. If the operation to terminate the detection process of a new sound source has been performed ("Yes" in S17), the beamforming processing unit 15 terminates the detection process of the new sound source.

[0076] <Beamformer Control Process>

[0077] Figure 6 This is a flowchart showing the processing flow performed by the beamforming processing unit 15 for controlling the beamformer based on the detection of a new sound source. Figure 6The diagram illustrates the processing flow when the directivity control unit 155 controls one of the multiple beamformers included in the signal processing device. When the first beamformer 152 outputs a first signal in a directivity state towards the first sound source, the process begins... Figure 6 The flowchart shown.

[0078] The first beamformer 152 operates with beamformer coefficients for the first sound source (S21). If the second sound source has not been detected ("No" in S22), the directivity control unit 155 repeats the process for detecting the second sound source. If the second sound source is detected ("Yes" in S22), the directivity control unit 155 begins measuring the elapsed time (S23). The directivity control unit 155 reduces the attenuator gain for the first sound source by calculating the attenuator gain for the first sound source based on the measured elapsed time (S24).

[0079] If the directional control unit 155 detects a sound source other than the second sound source (e.g., a third sound source) during the period when the first beamformer 152 is not operating ("Yes" in S25), the directional control unit 155 applies the beamformer coefficients calculated for the third sound source to the first beamformer 152 (S26). The directional control unit 155 can obtain the beamformer coefficients for the third sound source through the reference storage unit 154. The first beamformer 152 starts operating based on the beamformer coefficients for the third sound source applied by the directional control unit 155 (S27). The directional control unit 155 increases the attenuator gain for the third sound source (S28).

[0080] If the directivity control unit 155 has not detected a third sound source during the period when the first beamformer 152 is not operating ("No" in S25), the directivity control unit 155 repeats the process for detecting the third sound source. If the operation to terminate the control process of the beamformer has not been performed ("No" in S29), the beamforming processing unit repeats the process from S21 to S28. If the operation to terminate the control process of the beamformer has been performed ("Yes" in S29), the beamforming processing unit 15 terminates the control process of the beamformer.

[0081] <Effects of Sound Acquisition System S>

[0082] As described above, the sound acquisition system S includes: a first beamformer 152, whose output is a first signal obtained by emphasizing a sound signal based on a direction from a first range among sound signals arriving at a plurality of microphones 2; and a second beamformer 153, whose output is a second signal obtained by emphasizing a sound signal based on a direction from a second range among a plurality of sound signals. Then, the directivity control unit 155 causes the beamformer to perform beamforming processing based on the direction switching of the sound source.

[0083] Even if the speaker generating the speech switches between multiple speakers, the sound acquisition system S can acquire multiple voices without interrupting the speech generated by multiple speakers.

[0084] It should be noted that, although Figure 1 The description covers the case with three speakers, but the sound acquisition system S can also be used in cases with four or more speakers. Although the sound acquisition system S is provided with two beamformers in the above description, by providing the sound acquisition system S with three or more beamformers, the sound acquisition system S can acquire sound with directivity toward each of the three or more sound source directions.

[0085] The present invention has been described based on exemplary embodiments. The scope of the present invention is not limited to the scope described in the above embodiments, and various changes and modifications can be made within the scope of the present invention. For example, all or part of the device can be configured using any functionally or physically distributed or integrated units. Furthermore, new exemplary embodiments generated by any combination of exemplary embodiments are included in the exemplary embodiments. Moreover, the effects of the new exemplary embodiments resulting from the combination also have the effects of the original exemplary embodiments.

[0086] [Description of reference numerals in the attached figures]

[0087] 1 microphone array

[0088] 2 microphones

[0089] 10 signal processing devices

[0090] 11 Input Section

[0091] 12 First Attenuation Section

[0092] 13 Second Attenuation Section

[0093] 14 Output Section

[0094] 15 Beamforming Processing Unit

[0095] 151 Sound Source Direction Detection Department

[0096] 152 First Beamformer

[0097] 153 Second Beamformer

[0098] 154 Storage Department

[0099] 155 Directional Control Unit

[0100] 161 Variable Delay Unit

[0101] 162 Gain Adjustment Section

[0102] 163 Addition Department

Claims

1. A sound acquisition system, comprising: A microphone array, which includes multiple microphones; A first beamformer is used to output a first signal, which is obtained by emphasizing a sound signal based on a direction from a first range compared with a sound signal based on sound from other directions among a plurality of sound signals based on sound arriving at a plurality of microphones; A second beamformer is used to output a second signal, which is obtained by emphasizing a sound signal based on a direction from a second range more than a sound signal based on sound from other directions among the plurality of sound signals; A sound source direction detection unit is used to detect the direction of the sound source that generates the sound that reaches the plurality of microphones; as well as The directivity control unit is configured to, during the period when the first beamformer is outputting the first signal, cause the second beamformer to output the second signal if the change angle of the direction of the sound source detected by the sound source direction detection unit per unit time is determined to be equal to or greater than a threshold. Wherein, the first range is a range defined around the direction of the first sound source notified by the sound source direction detection unit. The second range is a range defined around the direction of the second sound source notified by the sound source direction detection unit.

2. The sound acquisition system according to claim 1, wherein, While the first beamformer is outputting the first signal, if the change angle of the direction of the sound source per unit time is determined to be less than a threshold, the directivity control unit causes the first beamformer to continue outputting the first signal even though the first range has been changed.

3. The sound acquisition system according to claim 1 or 2, wherein, While the first beamformer is outputting the first signal, if the change angle is determined to be equal to or greater than a threshold, the directivity control unit reduces the output level of the first signal.

4. The sound acquisition system according to claim 3, wherein, The directional control unit reduces the output level of the first signal by using an attenuation factor based on the elapsed time after the change angle is determined to be equal to or greater than a threshold.

5. The sound acquisition system according to claim 3, wherein, The directional control unit increases the output level of the second signal while decreasing the output level of the first signal.

6. The sound acquisition system according to claim 3, wherein, The directional control unit increases the output level of the second signal at a rate greater than the rate of change used to decrease the output level of the first signal.

7. The sound acquisition system according to claim 1 or 2, wherein, If it is determined that the direction of the sound source is not included in the first range, the directivity control unit causes the second beamformer to output the second signal.

8. The sound acquisition system according to claim 1 or 2, wherein, Before the second beamformer outputs the second signal, the directional control unit determines a second range such that the second range includes the direction of the sound source.

9. The sound acquisition system according to claim 1 or 2, wherein, While the second beamformer is outputting the second signal, if the change angle of the direction of the sound source detected by the sound source direction detection unit per unit time is determined to be equal to or greater than a threshold, the directivity control unit causes the first beamformer to output the first signal.

10. The sound acquisition system according to claim 1 or 2, further comprising a storage unit, the storage unit being used to store beamforming coefficients and the direction of the sound source detected by the sound source direction detection unit in association with each other. in, The directivity control unit uses beamformer coefficients stored in the storage unit in association with the direction of the sound source detected by the sound source direction detection unit to cause the first beamformer or the second beamformer to output the first signal or the second signal.

11. The sound acquisition system according to claim 10, wherein, The storage unit stores the directions of the sound sources previously detected by the sound source direction detection unit and the beamformer coefficients previously calculated by the directivity control unit based on those directions, in association with each other. If the direction of a newly detected sound source determined by the sound source direction detection unit is the same as the direction of a previously detected sound source stored in the storage unit, the directivity control unit uses beamforming coefficients stored in association with the direction of the previously detected sound source.

12. A sound acquisition method, comprising the following steps: A first signal is output, which is obtained by emphasizing a sound signal based on a direction from a first range more than a sound signal based on sound from other directions among multiple sound signals based on sound arriving at multiple microphones; The direction of the sound source that generates the sound arriving at the plurality of microphones is detected using a sound source direction detection unit; as well as During the output of the first signal, if it is determined that the change angle of the direction of the sound source per unit time is equal to or greater than a threshold, a second signal is output. The second signal is obtained by selecting a sound signal from among the plurality of sound signals that emphasizes sound from a direction within a second range more than sound signals based on sound from other directions. Wherein, the first range is a range defined around the direction of the first sound source notified by the sound source direction detection unit. The second range is a range defined around the direction of the second sound source notified by the sound source direction detection unit.

13. A storage medium containing a stored program, the program being used to cause a computer to function as: A first beamformer is used to output a first signal, which is obtained by emphasizing a sound signal based on a direction from a first range compared with a sound signal based on sound from other directions among a plurality of sound signals based on sound arriving at a plurality of microphones; A second beamformer is used to output a second signal, which is obtained by emphasizing a sound signal based on a direction from a second range more than a sound signal based on sound from other directions among the plurality of sound signals; A sound source direction detection unit is used to detect the direction of the sound source that generates the sound that reaches the plurality of microphones; as well as The directivity control unit is configured to, during the period when the first beamformer is outputting the first signal, cause the second beamformer to output the second signal if the change angle of the direction of the sound source detected by the sound source direction detection unit per unit time is determined to be equal to or greater than a threshold. Wherein, the first range is a range defined around the direction of the first sound source notified by the sound source direction detection unit. The second range is a range defined around the direction of the second sound source notified by the sound source direction detection unit.

14. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method of claim 12.

Citation Information

Patent Citations

  • Beam forming processing unit

    JP2013201525A

  • Audio processing apparatus, audio processing system, and audio processing method

    CN105474666A

  • An audio frequency acquisition method and device based on a microphone array

    CN106098075A