Audio signal processing method, audio signal processing apparatus, audio signal processing system, and audio signal processing program

WO2026204557A1PCT designated stage Publication Date: 2026-10-01YAMAHA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/010331
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-03-17
Publication Date
2026-10-01

Smart Images

  • Figure JP2026010331_01102026_PF_FP_ABST
    Figure JP2026010331_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This audio signal processing method is used in an audio signal processing system provided with a plurality of microphones and a loudspeaker. The audio signal processing method comprises performing setting, for each of the plurality of microphones, so that a sound collection gain increases as a sound collection level increases, and further obtaining, for each of the plurality of microphones, a loop gain to be fed back via the loudspeaker and adjusting the sound collection gain on the basis of the loop gain.
Need to check novelty before this filing date? Find Prior Art

Description

Audio signal processing method, audio signal processing apparatus, audio signal processing system, and audio signal processing program

[0001] One embodiment of the present invention relates to an audio signal processing method, an audio signal processing apparatus, an audio signal processing system, and an audio signal processing program.

[0002] Patent Document 1 discloses an audio signal processing method that changes a first transmission amount transmitted to a first sound emitting device closest to a first sound collecting device that has detected a speaker to the minimum.

[0003] International Publication No. 2023 / 042699

[0004] Patent Document 1 does not disclose a gain adjustment method for a plurality of sound collecting devices when a speaker is switched. When a speaker is switched, if there is a time difference before increasing the gain of a microphone for a new speaker, the beginning of the utterance of the new speaker may not be collected. Further, if gains of a plurality of microphones are increased simultaneously when a speaker is switched, howling may occur.

[0005] An object of one embodiment of the present invention is to provide an audio signal processing method capable of preventing howling when a speaker is switched.

[0006] The audio signal processing method according to one embodiment of the present invention is used in an audio signal processing system including a plurality of microphones and a speaker. In the audio signal processing method, for each of the plurality of microphones, the higher the sound collection level is, the higher the sound collection gain is set; further, for each of the plurality of microphones, a loop gain that is fed back via the speaker is obtained, and the sound collection gain is adjusted based on the loop gain.

[0007] The audio signal processing method according to one embodiment of the present invention can prevent howling when a speaker is switched.

[0008] It is a schematic block diagram showing the configuration of the audio signal processing system 1. It is a block diagram showing the configuration of the audio signal processing apparatus 20. It is a diagram schematically showing an audio transmission path. It is a block diagram schematically showing a mixing gain Gij in the audio signal processing apparatus 20. It is a schematic block diagram showing the configuration of an audio signal processing system 1A according to Modification 1.

[0009] (First Embodiment) Figure 1 is a schematic block diagram showing the configuration of the sound signal processing system 1 according to the first embodiment. The sound signal processing system 1 comprises i microphones 10i (i=1 to m), a sound signal processing device 20, and j speakers 30j (j=1 to n).

[0010] The microphone 10i and speaker 30j are installed in the ceiling of the conference room 300, as an example. The microphone 10i and speaker 30j are connected to the sound signal processing device 20 by audio cables, communication cables, or wireless communication. For the sake of explanation, Figure 1 shows the sound signal processing device 20 as being installed in the ceiling, but in reality it is installed inside the conference room 300 or in a monitor room adjacent to the conference room 300. Furthermore, in this invention, the microphone 10i and speaker 30j do not need to be installed in the ceiling.

[0011] The sound signal processing system 1 is a system that amplifies the voice of speaker S and delivers it to listener L. Such a sound signal processing system 1 is used, for example, in meetings, seminars, or presentations. Figure 1 shows one speaker S and one listener L, but the number of speaker S and listener L is not limited to one each.

[0012] Microphone 10i captures the voice of speaker S. The sound signal processing device 20 receives the sound signal captured by microphone 10i. The sound signal processing device 20 performs signal processing such as mixing, gain adjustment, and equalization on the input sound signal. The sound signal processing device 20 outputs the processed sound signal to speaker 30j. Speaker 30j emits sound based on the input sound signal.

[0013] Figure 2 is a block diagram showing the configuration of the sound signal processing device 20. The sound signal processing device 20 consists of a general-purpose information processing device such as a personal computer or a smartphone, as an example. The general-purpose information processing device may be a device used by the speaker, or a device used by the operator running the conference or the installer setting up the conference system.

[0014] The sound signal processing device 20 includes a display 251, a user interface 252, a flash memory 253, a CPU 254, a RAM 255, and a communication interface 256.

[0015] The display unit 251 consists of, for example, an LCD or OLED, and displays various information. The user interface 252 is, for example, a touch panel stacked on the LCD or OLED of the display unit 251. Alternatively, the user interface 252 may be a keyboard or mouse. If the user interface 252 is a touch panel, the user interface 252, together with the display unit 251, constitutes a GUI (Graphical User Interface).

[0016] The communication interface 256 includes an audio interface and communication means such as wired LAN, wireless LAN, or Bluetooth®. The communication interface 256 is connected to the microphone 10i and the speaker 30j.

[0017] The CPU 254 is an example of a processor and is a control unit that controls the operation of the sound signal processing device 20. The CPU 254 performs various operations such as sound signal processing by reading a predetermined program, such as an application program, stored in the flash memory 253 (a storage medium) into the RAM 255 and executing it. The program may also be stored in a server (not shown). The CPU 254 may also download and execute a program from the server via a network.

[0018] Figure 3 is a schematic diagram showing the audio transmission path. Figure 4 is a schematic block diagram showing the gain of microphone 10i, the mixing gain in the sound signal processing device 20 (gain of the sound signal distributed to each speaker), and the gain of speaker 30j.

[0019] The speaker S's voice reaches the listener L via a sound amplification transmission path that passes through the microphone 10i, the sound signal processing device 20, and the speaker 30j. In addition, the speaker S's voice also reaches the listener L via a direct transmission path that passes through the space of the conference room 300 without passing through the sound amplification transmission path.

[0020] The transfer function of the direct transmission path from the speaker S to the listener L is H1, the transfer function from the speaker S to the microphone 10i is H2i, the gain of the microphone 10 is MGi, the gain of the sound signal distributed to each speaker for each microphone in the sound signal processing device 20 is Gij, the gain of speaker 30j is SGj, the transfer function from speaker 30j to the listening position is H3j, and the transfer function from speaker 30j to microphone 10i is H4ji.

[0021] The loop gain LGi of each microphone is expressed as the sum of the loop gains of each speaker, as shown in equation (1).

[0022] The transfer function from the speaker to the microphone increases as the speaker is closer to the microphone. On the other hand, the closer the speaker is to the microphone, the closer it is to the speaker's position, and the greater the sound pressure of the speaker's voice that reaches it through the direct transmission path. Therefore, the closer the speaker is to the microphone, the lower the speaker's gain can be.

[0023] Therefore, the sound signal processing device 20 sets the gain Gij of the sound signal to be distributed to the multiple speakers based on the distance between the microphone and each of the multiple speakers. The gain Gij is made proportional to the distance d between the speaker and the microphone, as shown in equation (2), for example.

[0024] The coefficient w in equation (2) is a weighting coefficient and is not mandatory. The weighting coefficient w is a value of 0 or greater. As a result, the speaker gain Gij increases as it is farther from the microphone. Equation (2) above can also be used as an equation that takes into account the statistical reflection component. In this case as well, the gain Gij is proportional to the distance d between the speaker and the microphone.

[0025] The distance between each microphone and each speaker may be manually entered by the user, or it may be determined by means of imaging with a camera or measurement using LiDAR (Light Detection and Ranging). Alternatively, it may be determined by measuring the impulse response when a test sound is output from each speaker.

[0026] Furthermore, the sound signal processing device 20 sets the sound pickup gain MGi based on gain sharing so that the sound pickup gain MGi increases as the sound pickup level increases.

[0027] Gain sharing is a setting method that shares the loop gain among multiple microphones in a given loop system. Loop gain can be determined, for example, by outputting a test sound from a speaker and measuring the impulse response. Gain sharing assumes that the transmission characteristics of the speaker and each microphone are equal, and that the gain is shared based on the ratio of the input energy to the microphones, so that the loop gain remains constant even if the gain of each microphone is changed. As shown in Figure 4, when signals from multiple microphone systems are distributed to multiple speakers independently of each other, the loop gain is also independent for each microphone system, and it is not possible to share a single gain. However, if there is acoustic coupling between the microphone systems, that is, if the input signal of one microphone system is amplified and then input to another microphone system and re-radiated, the concept of gain sharing can be applied to reduce the acoustic coupling between systems by increasing the pickup gain of the microphone system with a high pickup level and decreasing the pickup gain of the microphone system with a low pickup level. The sound signal processing device 20 can increase the system gain by setting up this gain sharing method.

[0028] However, since the gain is not shared (distributed) across a single system, certain combinations of microphone gains can easily lead to feedback. Furthermore, the degree of acoustic coupling varies depending on the relative positions of multiple speakers and microphones, meaning that different combinations of microphone gains are prone to feedback.

[0029] The sound signal processing device 20 of this embodiment determines the loop gain for each of the multiple microphones 10i while acoustic coupling is in place, or, while observing the loop gain, adjusts the pickup gain MGi according to the pickup level of each microphone so that the loop gain of each microphone system does not exceed a predetermined value (e.g., -6 dB) and the overall gain MGi, Gij, and SGj are maximized. Alternatively, the sound signal processing device 20 may detect the occurrence of howling instead of the loop gain and determine the gain MGi that does not cause howling by lowering the gain from the pickup gain MGi at the time howling is detected. In other words, after adjusting the gain for each of the multiple microphone systems, the sound signal processing device 20 gradually increases the gain while acoustic coupling is in place and sets the gain to a predetermined value (e.g., -6 dB) lower than the gain at which howling actually occurred.

[0030] The sound signal processing device 20 may further adjust the sound pickup gain MGi based on the distance of each of the multiple microphones to the speaker. For example, the sound signal processing device 20 sets the sound pickup gain MGi to be proportional to the distance, at or below the maximum gain that does not cause feedback. In this case, the sound signal processing device 20 can control the sound at or below the gain that does not cause feedback while keeping the amplification volume approximately constant regardless of the distance from the microphone.

[0031] The sound signal processing device 20 may determine the loop gain not only for the system of one speaker, but also for each system of multiple speakers, and adjust the sound pickup gain MGi based on the loop gain of all systems. In this case as well, the sound signal processing device 20 determines the loop gain for each speaker for each of the multiple microphones 10i, and adjusts the sound pickup gain MGi according to the sound pickup level of each microphone so that the loop gain of each system does not exceed a predetermined value (for example, -6 dB) and the overall gain MGi, Gij, and SGj is maximized.

[0032] The number of loop gains increases with the number of microphones and speakers. Therefore, loop gains can be determined not only by measuring the impulse response but also by simulation. For example, the transfer function H2i from the speaker S's position to the microphone 10i can be determined by simulation by inputting information such as the speaker S's position, the speaker S's sound pressure, the microphone 10i's position, the position of the conference room 300's walls, and the wall's reflectivity. The transfer function H1 of the direct transmission path can also be determined by simulation by inputting information such as the speaker S's position, the speaker S's sound pressure, the listener L's position, the conference room 300's walls, and the wall's reflectivity. Similarly, the transfer function H3j from speaker 30j to the listening position and the transfer function H4ji from speaker 30j to microphone 10i can also be determined by simulation by inputting information such as the speaker 30j's position, the conference room 300's walls, and the wall's reflectivity. Alternatively, the transfer functions H1, H2i, H3j, and H4ji can be statistically calculated from the volume and surface area of ​​the wall, as well as the sound absorption coefficient (or reflectivity) of the wall surface. They can also be determined solely from the distances to the speaker S, microphone 10i, speaker 30j, and listener L.

[0033] As described above, the sound signal processing device 20 of this embodiment is configured such that the sound pickup gain increases as the sound pickup level increases for each of the multiple microphones, and further, the loop gain is determined for each of the multiple microphones and the sound pickup gain is adjusted based on the loop gain, thereby preventing sound interruptions when switching microphones and eliminating the risk of howling.

[0034] (Modification 1) Figure 5 is a schematic block diagram showing the configuration of the sound signal processing system 1A according to Modification 1. The microphone according to Modification 1 is an array microphone 100 in which a plurality of microphone units are arranged. The sound signal processing device 20 or the array microphone 100 sets up multiple sound pickup beams by combining the sound signals picked up by the plurality of microphone units. Beamforming is a delayed summing process that, for example, adds a delay to the sound signals of each microphone unit and combines them to form a sensitivity peak (focal point) at the speaker's position. i sound pickup beams (i = 1 to k) are set. In the example in Figure 5, three sound pickup beams b1, b2, and b3 are set.

[0035] The sound signal processing device 20 adjusts the sound pickup gain MGi according to the sound pickup level for each of the multiple sound pickup beams bi, rather than for each individual microphone system of the array microphone 100. It sets the gain Gij of the sound signal to be distributed to the multiple speakers. That is, the gain MGi shown in the above embodiment corresponds to the sound pickup gain for each of the multiple sound pickup beams in Modification 1. For example, the sound signal processing device 20 adjusts the sound pickup gain MGi according to the sound pickup level for each direction of the multiple sound pickup beams. In this case, the sound signal processing device 20 may multiply the gain Gij by a correction value Gadjk (where k indicates the direction) for each direction.

[0036] In this case, the gain Gij increases with increasing speaker distance from the focal point of the sound-collecting beam and decreases with decreasing speaker distance. Therefore, the sound signal processing device 20 of Modification 1 can also ensure sufficient sound pressure for amplified sound while suppressing howling, and automatically provide amplified sound with better sound quality.

[0037] (Modification 2) The sound signal processing device 20 according to Modification 2 sets the amount of sound signal to be sent to the multiple speakers based on the loop gain for each of the multiple microphones. The sound signal processing device 20 determines the speaker gain Gij by setting a target value LG for the feedback (loop gain) from each speaker to the microphone. The target value LG for the loop gain from each speaker is made proportional to the distance d between the speaker and the microphone, as shown in equation (3), for example.

[0038] The coefficient w in equation (3) is a weighting coefficient and is not mandatory. The weighting coefficient w is a value of 0 or greater. When w = 0 in equation (3), the target value LG of the loop gain from each speaker is constant. However, even if the target value LG of the loop gain is constant, the speaker gain Gij increases the further it is from the microphone.

[0039] The sound signal processing device 20 determines the maximum loop gain level (dB) for each speaker with respect to microphone i, based on the distance between each microphone and each speaker, for example, as shown in the following formula (4).

[0040] l represents the speaker number. In this example, there is feedback from multiple speakers to a single microphone system, but the target loop gain from each speaker is set according to the distance, and the overall target loop gain for the microphone system is -6 dB. If the loop gain exceeds 0 dB, the feedback system may diverge and cause howling. Equation (4) means that the loop gain from all speakers to the microphone system is set to -6 dB, and the energy of the microphone system is distributed to each speaker with distance weighting. The coefficient w in equation (4) is a weighting coefficient and is not essential. The weighting coefficient w is a value of 0 or greater. The larger the value of the weighting coefficient w, the greater the gain of the speaker farther from the microphone.

[0041] Subsequently, the sound signal processing device 20 calculates the difference between the loop gain for each speaker and the maximum loop gain obtained above. Based on the calculated difference, the sound signal processing device 20 sets the gain Gij for each speaker. Note that the difference from the loop gain may be adjusted using the total gain of MGi, Gij, and SGj.

[0042] Thus, the sound signal processing device 20 of the modified example 2 determines the loop gain to be fed back to the microphone for each of the multiple speakers, determines a target value for the loop gain for each of the multiple speakers, and sets the gain Gij of the sound signal to be distributed to the multiple speakers based on the determined loop gain and the target value of the loop gain.

[0043] Through the above processing, as shown in Fig. 3, the sound pressure of the voice from speaker S at the position of listener L is expressed as the sum of the sound amplification transmission path represented by gain H2i·MGi·Gij·SGj·H3j and the direct transmission path represented by gain H1. To suppress howling, gain Gij is set to be smaller for speakers closer to the microphone, which results in a higher sound pressure in the direct transmission path. At positions far from the speaker, the sound pressure in the direct transmission path becomes smaller, so gain Gij is set to be larger for speakers farther from the microphone. Accordingly, the audio signal processing device 20 can ensure amplified sound with sufficient sound pressure while suppressing howling, and automatically provide amplified sound with better sound quality.

[0044] (Modification 3) The audio signal processing device 20 according to Modification 3 uses MG'i that further adjusts the sound collection gain of each microphone. In a state where the acoustic coupling exists, for each of the plurality of microphones 10i, while detecting howling occurrence in advance for a plurality of combinations of the sound collection gain MGi of the microphone in each microphone system, the audio signal processing device 20 obtains MG'i that maximizes the overall gain MG'i·MGi·Gij·SGj at which howling does not occur. The audio signal processing device 20 may set MG'i such that the loop gain of each microphone system does not exceed a predetermined value (for example, -6 dB) while measuring the loop gain.

[0045] Furthermore, the audio signal processing device 20 may calculate MG'i based on the measured transfer functions of each microphone and each speaker.

[0046] (Modification 4) In a state where the acoustic coupling exists, for each of the plurality of microphones 10i, while detecting howling occurrence in advance for a plurality of combinations of the microphone gain MGi of the microphone in each microphone system, the audio signal processing device 20 according to Modification 4 obtains a gain curve with respect to the input level of MGi that maximizes the overall gain MGi·Gij·SGj at which howling does not occur. The audio signal processing device 20 may set the gain curve such that the loop gain of each microphone system does not exceed a predetermined value (for example, -6 dB) while measuring the loop gain.

[0047] Furthermore, the audio signal processing device 20 may calculate the gain curve from the measured transfer functions of each microphone and each speaker.

[0048] (Modification 5) The sound signal processing device 20 in Modification 5 sets the sound pickup gain of the microphone with the highest sound pickup level to the maximum and the sound pickup gain of the other microphones to the minimum, and sets the gain to switch microphones according to the speaker's position or changes in the speaker himself. When switching the microphone to be muted (gain to maximum) from the first microphone to the second microphone in Modification 5, the sound signal processing device 20 gradually decreases the sound pickup gain of the first microphone while gradually increasing the sound pickup gain of the second microphone.

[0049] The sound signal processing device 20 in modified example 5 also ensures that there is no interruption in sound when switching microphones, and there is no risk of generating feedback during the switching process.

[0050] The above description of embodiments should be considered in all respects to be illustrative and not restrictive. The scope of the present invention is indicated by the claims, not by the above embodiments. Furthermore, the scope of the present invention is intended to include all modifications within the meaning and scope equivalent to the claims.

[0051] 1, 1A: Sound signal processing system, 10i: Microphone, 20: Sound signal processing device, 30j: Speaker, 100: Array microphone, 251: Display, 252: User I / F, 253: Flash memory, 254: CPU, 255: RAM, 256: Communication I / F, 300: Conference room

Claims

1. An audio signal processing method used in an audio signal processing system comprising multiple microphones and a speaker, wherein each of the multiple microphones is set such that the sound pickup gain increases as the sound pickup level increases, and further, for each of the multiple microphones, a loop gain is determined to be fed back through the speaker, and the sound pickup gain is adjusted based on the loop gain.

2. The sound signal processing method according to claim 1, wherein the speaker includes a plurality of speakers, and for each of the plurality of microphones, the loop gain to be fed back in each system via the plurality of speakers is determined, and the sound pickup gain is adjusted based on the loop gain of the entire system.

3. The sound signal processing method according to claim 2, wherein for each of the plurality of microphones, the amount of sound signal to be sent to the plurality of speakers is set based on the loop gain.

4. The sound signal processing method according to claim 1 or 2, wherein the sound pickup gain is adjusted based on the distance of each of the plurality of microphones to the speaker.

5. The sound signal processing method according to claim 1 or 2, wherein the loop gain is determined based on the impulse response of the transmission system from the speaker to the microphone.

6. The sound signal processing method according to claim 1 or claim 2, wherein the plurality of microphones are array microphones in which a plurality of microphone units are arranged, a plurality of sound pickup beams are set by synthesizing the sound signals picked up by the plurality of microphone units, and the sound pickup gain is adjusted for each of the plurality of sound pickup beams.

7. An audio signal processing device used in an audio signal processing system comprising a plurality of microphones and a speaker, comprising a processor that sets each of the plurality of microphones such that the sound pickup gain increases as the sound pickup level increases, wherein the processor determines the loop gain to be fed back through the speaker for each of the plurality of microphones, and adjusts the sound pickup gain based on the loop gain.

8. An audio signal processing system comprising: a plurality of microphones; a speaker; and an audio signal processing device that sets each of the plurality of microphones such that the sound pickup gain increases as the sound pickup level increases, wherein the audio signal processing device determines the loop gain to be fed back through the speaker for each of the plurality of microphones, and adjusts the sound pickup gain based on the loop gain.

9. An audio signal processing program for an audio signal processing device used in an audio signal processing system comprising multiple microphones and speakers, which causes each of the multiple microphones to be set such that the sound pickup gain increases as the sound pickup level increases, and further causes each of the multiple microphones to determine the loop gain to be fed back through the speaker, and to adjust the sound pickup gain based on the loop gain.