Sound signal processing method, sound signal processing device, sound signal processing system, and sound signal processing program
Patent Information
- Application Number
- PCT/JP2026/010328
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-17
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026010328_01102026_PF_FP_ABST
Abstract
Description
Audio signal processing method, audio signal processing apparatus, audio signal processing system, and audio signal processing program
[0001] One embodiment of the present invention relates to an audio signal processing method, an audio signal processing apparatus, an audio signal processing system, and an audio signal processing program.
[0002] Patent Literature 1 discloses a gain setting apparatus that dynamically changes the gain for each speaker with respect to a master volume in consideration of the loop gain of each system.
[0003] Japanese Patent Laying-Open No. 2018-137532
[0004] The technology of Patent Literature 1 does not consider sound transmitted from a speaker to a listener in a room.
[0005] An object of an embodiment of the present invention is to provide an audio signal processing method that achieves easy and appropriate volume control while also considering sound transmitted from a speaker to a listener in a room.
[0006] An audio signal processing method according to an embodiment of the present invention is an audio signal processing method used in an audio signal processing system including a microphone that collects sound of a speaker and a plurality of speakers that amplify and output the speaker's sound to a listener, the method comprising: distributing an audio signal collected by the microphone to the plurality of speakers; for each of the plurality of speakers, obtaining a loop gain that is fed back to the microphone; receiving a volume setting; and setting the gain of the audio signal distributed to the plurality of speakers in accordance with the received volume such that the loop gain does not exceed a predetermined value and based on distance attenuation of sound directly arriving from the speaker to the listener and transfer characteristics of sound arriving at the listener from the microphone via the plurality of speakers.
[0007] The audio signal processing method according to an embodiment of the present invention can achieve easy and appropriate volume control while also considering sound transmitted from a speaker to a listener in a room.
[0008] This is a schematic block diagram showing the configuration of the sound signal processing system 1. This is a block diagram showing the configuration of the sound signal processing device 20. This is a diagram schematically showing the audio transmission path. This is a schematic block diagram schematically showing the mixing gain Gij in the sound signal processing device 20. This is an example of the GUI displayed on the display unit 251. This is a schematic block diagram showing the configuration of the sound signal processing system 1A according to Modification 1. This is a schematic block diagram showing the configuration of the sound signal processing system 1B according to Modification 2. This is a schematic block diagram showing the configuration of the sound signal processing system 1C according to Modification 3. This is a schematic block diagram showing the configuration of the sound signal processing system 1D according to Modification 4.
[0009] (First Embodiment) Figure 1 is a schematic block diagram showing the configuration of the sound signal processing system 1 according to the first embodiment. The sound signal processing system 1 comprises i microphones 10i (i=1 to m), a sound signal processing device 20, and j speakers 30j (j=1 to n).
[0010] The microphone 10i and speaker 30j are installed in the ceiling of the conference room 300, as an example. The microphone 10i and speaker 30j are connected to the sound signal processing device 20 by audio cables, communication cables, or wireless communication. For the sake of explanation, Figure 1 shows the sound signal processing device 20 as being installed in the ceiling, but in reality it is installed inside the conference room 300 or in a monitor room adjacent to the conference room 300. Furthermore, in this invention, the microphone 10i and speaker 30j do not need to be installed in the ceiling.
[0011] The sound signal processing system 1 is a system that amplifies the voice of speaker S and delivers it to listener L. Such a sound signal processing system 1 is used, for example, in meetings, seminars, or presentations. Figure 1 shows one speaker S and one listener L, but the number of speaker S and listener L is not limited to one each.
[0012] Microphone 10i captures the voice of speaker S. The sound signal processing device 20 receives the sound signal captured by microphone 10i. The sound signal processing device 20 performs signal processing such as mixing, gain adjustment, and equalization on the input sound signal. The sound signal processing device 20 outputs the processed sound signal to speaker 30j. Speaker 30j emits sound based on the input sound signal.
[0013] Figure 2 is a block diagram showing the configuration of the sound signal processing device 20. The sound signal processing device 20 consists of a general-purpose information processing device such as a personal computer or a smartphone, as an example. The general-purpose information processing device may be a device used by the speaker, or a device used by the operator running the conference or the installer setting up the conference system.
[0014] The sound signal processing device 20 includes a display 251, a user interface 252, a flash memory 253, a CPU 254, a RAM 255, and a communication interface 256.
[0015] The display unit 251 consists of, for example, an LCD or OLED, and displays various information. The user interface 252 is, for example, a touch panel stacked on the LCD or OLED of the display unit 251. Alternatively, the user interface 252 may be a keyboard or mouse. If the user interface 252 is a touch panel, the user interface 252, together with the display unit 251, constitutes a GUI (Graphical User Interface).
[0016] The communication interface 256 includes an audio interface and communication means such as wired LAN, wireless LAN, or Bluetooth®. The communication interface 256 is connected to the microphone 10i and the speaker 30j.
[0017] The CPU 254 is an example of a processor and is a control unit that controls the operation of the sound signal processing device 20. The CPU 254 performs various operations such as sound signal processing by reading a predetermined program, such as an application program, stored in the flash memory 253 (a storage medium) into the RAM 255 and executing it. The program may also be stored in a server (not shown). The CPU 254 may also download and execute a program from the server via a network.
[0018] Figure 3 is a schematic diagram showing the audio transmission path. Figure 4 is a schematic block diagram showing the gain of microphone 10i, the mixing gain in the sound signal processing device 20 (gain of the sound signal distributed to each speaker), and the gain of speaker 30j.
[0019] The speaker S's voice reaches the listener L via a sound amplification transmission path that passes through the microphone 10i, the sound signal processing device 20, and the speaker 30j. In addition, the speaker S's voice also reaches the listener L via a direct transmission path that passes through the space of the conference room 300 without passing through the sound amplification transmission path.
[0020] The transfer function of the direct transmission path from the speaker S to the listener L is H1, the transfer function from the speaker S to the microphone 10i is H2i, the gain of the microphone 10 is MGi, the gain of the sound signal distributed to each speaker for each microphone in the sound signal processing device 20 is Gij, the gain of speaker 30j is SGj, the transfer function from speaker 30j to the listening position is H3j, and the transfer function from speaker 30j to microphone 10i is H4ji.
[0021] The loop gain LGi of each microphone is expressed as the sum of the loop gains of each speaker, as shown in equation (1).
[0022] The transfer function from the speaker to the microphone increases as the speaker is closer to the microphone. On the other hand, the closer the speaker is to the microphone, the closer it is to the speaker's position, and the greater the sound pressure of the speaker's voice that reaches it through the direct transmission path. Therefore, the closer the speaker is to the microphone, the lower the speaker's gain can be.
[0023] The sound signal processing device 20 sets the gain Gij of the sound signal to be distributed to each of the multiple speakers, based on the distance attenuation of sound traveling directly from the speaker to the listener and the transmission characteristics of sound traveling from the microphone to the listener via multiple speakers, so that the loop gain LGi does not exceed a predetermined value. As described above, the closer the speaker is to the microphone, the greater the sound pressure of the speaker's voice that reaches it via the direct transmission path, and the further the speaker is from the microphone, the smaller the sound pressure of the speaker's voice that reaches it via the direct transmission path. Therefore, the gain Gij reflects the distance attenuation of sound from the speaker's position to the listener's position.
[0024] More specifically, the target value of the loop gain LG is made proportional to the distance d between the speaker and the microphone, as shown in equation (2), for example.
[0025] The coefficient w in equation (2) is a weighting coefficient and is not mandatory. The weighting coefficient w is a value of 0 or greater. When w = 0 in equation (2), the target value LG of the loop gain from each speaker is constant. However, even if the target value LG of the loop gain is constant, the speaker gain Gij increases the further it is from the microphone.
[0026] The sound signal processing device 20 determines the maximum loop gain level (dB) for each speaker with respect to microphone i, based on the distance between each microphone and each speaker, for example, as shown in the following equation (3).
[0027] l represents the speaker number. In this example, there is feedback from multiple speakers to a single microphone system, but the target loop gain from each speaker is set according to the distance, and the overall target loop gain for the microphone system is -6 dB. If the loop gain exceeds 0 dB, the feedback system may diverge and cause howling. Equation (3) means that the loop gain from all speakers to the microphone system is set to -6 dB, and the energy of the microphone system is distributed to each speaker with distance weighting. The coefficient w in equation (3) is a weighting coefficient and is not essential. The weighting coefficient w is a value of 0 or greater. The larger the value of the weighting coefficient w, the greater the gain of the speaker farther from the microphone.
[0028] Subsequently, the sound signal processing device 20 calculates the difference between the loop gain for each speaker and the maximum loop gain obtained above. Based on the calculated difference, the sound signal processing device 20 sets the gain Gij for each speaker. Note that the difference from the loop gain may be adjusted using the total gain of MGi, Gij, and SGj.
[0029] Thus, the sound signal processing device 20 of the modified example 1 determines the loop gain to be fed back to the microphone for each of the multiple speakers, determines a target value for the loop gain for each of the multiple speakers, and sets the gain Gij of the sound signal to be distributed to the multiple speakers based on the determined loop gain and the target value of the loop gain.
[0030] As a result of the above processing, as shown in Figure 3, the sound pressure of the speaker S's voice at the listener L's position is expressed as the sum of the amplification transmission path represented by gains H2i, MGi, Gij, SGj, and H3j, and the direct transmission path represented by gain H1. To suppress howling, the gain Gij is made smaller for speakers closer to the microphone, but the sound pressure of the direct transmission path increases. At positions farther from the speaker, the sound pressure of the direct transmission path decreases, but the gain Gij is made larger for speakers further from the microphone. Therefore, the sound signal processing device 20 can ensure sufficient sound pressure for amplified sound while suppressing howling, and can automatically provide amplified sound with better sound quality.
[0031] The sound signal processing unit 20 determines the loop gain to be fed back to each of the multiple speakers for each microphone. The loop gain for each speaker is measured, for example, by outputting a test sound from speaker 30j.
[0032] The sound pressure of the speaker S's voice at the listener L's position is expressed as the sum of the amplification transmission path, represented by gains H2i, MGi, Gij, SGj, and H3j, and the direct transmission path, represented by gain H1. The sound signal processing device 20 sets the gain Gij so that the sound pressure of the speaker S's voice at the listener L's position becomes the required sound pressure. The transfer function H1 of the direct transmission path and the transfer function H2i from the speaker S's position to the microphone 10i can be determined by outputting a test sound from the speaker S's position and measuring it at the listener L's position. The transfer function H3j from the speaker 30j to the listening position can also be measured, for example, by outputting a test sound from the speaker 30j.
[0033] Alternatively, the transfer function H2i from the speaker S's position to the microphone 10i can be determined by simulation by inputting information such as the speaker S's position, the speaker S's sound pressure, the microphone 10i's position, the position of the conference room 300's walls, and the wall's reflectivity. The transfer function H1 of the direct transmission path can also be determined by simulation by inputting information such as the speaker S's position, the speaker S's sound pressure, the listener L's position, the conference room 300's walls, and the wall's reflectivity. Similarly, the transfer function H3j from speaker 30j to the listening position and the transfer function H4ji from speaker 30j to microphone 10i can also be determined by simulation by inputting information such as the speaker 30j's position, the conference room 300's walls, and the wall's reflectivity. Alternatively, the transfer functions H1, H2i, H3j, and H4ji can be statistically calculated from the volume, surface area, and sound absorption coefficient (or reflectivity) of the wall, or they can be determined solely from the distances to the speaker S, microphone 10i, speaker 30j, and listener L.
[0034] The positions of the speaker S, each microphone, each speaker, and the listener L may be manually entered by the user, or they may be determined by means of imaging with a camera or measurement using LiDAR (Light Detection and Ranging).
[0035] As described above, the sound signal processing device 20 provides the listener L with a sufficiently loud amplified sound while suppressing howling. In this embodiment, the sound signal processing device 20 further changes the obtained gain Gij through user operation.
[0036] Figure 5 shows an example of the GUI displayed on the display unit 251. The sound signal processing device 20 displays a slider as an operator in its GUI. The sound signal processing device 20 accepts volume settings via the slider. The slider is movable between gain 0 and 1 (0 dB to -∞ dB). When the slider is at its maximum value of 0 dB, the gain Gij is set to the maximum required sound pressure level to prevent feedback, as determined as described above. In other words, the user can adjust the volume from the maximum required sound pressure level to prevent feedback to mute by moving the slider. This allows the sound signal processing device 20 to easily and appropriately control the volume while also considering the sound transmitted from the speaker to the listener within the room.
[0037] (Modification 1) Figure 6 is a schematic block diagram showing the configuration of the sound signal processing system 1A according to Modification 1. The microphone according to Modification 1 is an array microphone 100 in which a plurality of microphone units are arranged. The sound signal processing device 20 or the array microphone 100 sets up a plurality of sound pickup beams by combining the sound signals picked up by the plurality of microphone units. Beamforming is a delayed summing process that, for example, adds a delay to the sound signals of each microphone unit and combines them to form a sensitivity peak (focal point) at the speaker's position. i sound pickup beams (i = 1 to k) are set. In the example in Figure 6, three sound pickup beams b1, b2, and b3 are set.
[0038] The sound signal processing device 20 sets the gain Gij of the sound signal to be distributed to the multiple speakers for each of the multiple sound-collecting beams bi, rather than for each individual microphone system of the array microphone 100. That is, the gain Gij shown in the above embodiment corresponds to the gain for each of the multiple sound-collecting beams in Modification 1. The sound signal processing device 20 sets the gain Gij of the sound signal to be distributed to the multiple speakers for each of the multiple sound-collecting beams. In this case, the sound signal processing device 20 may multiply the gain Gij by a correction value Gadjk (where k indicates the direction) for each direction.
[0039] In this case, the gain Gij increases with increasing speaker distance from the focal point of the sound-collecting beam and decreases with decreasing speaker distance. Therefore, the sound signal processing device 20 of Modification 1 can also ensure sufficient sound pressure for amplified sound while suppressing howling, and automatically provide amplified sound with better sound quality.
[0040] (Modification 2) Figure 7 is a schematic block diagram showing the configuration of the sound signal processing system 1B according to Modification 2. The sound signal processing system 1B of Modification 2 includes multiple information processing terminals 50 for each listener. In Modification 2, a GUI is displayed on each of the multiple information processing terminals 50 for each listener, and volume settings are accepted for each listener.
[0041] The sound signal processing device 20 receives volume settings for each listener via the information processing terminal 50 for each listener. The sound signal processing device 20 sets a gain G_ij according to the volume setting for each listener. The sound signal processing device 20 adjusts the gain G_ij, for example, in the following manner. For example, when increasing the sound pressure for a certain listener, the sound signal processing device 20 increases the gain with a smooth distribution of one first speaker corresponding to the listener, or a certain range of first speakers around the center of the one speaker, thereby obtaining a desired increase in sound pressure. On the other hand, since the loop gain increases as a whole, the sound signal processing device 20 adjusts the loop gain for each microphone to be equal to or less than a predetermined value by adjusting to decrease the gain of second speakers outside a certain range. In this way, the sound signal processing device 20 increases or decreases the gain of the first speakers corresponding to the listener who has received the volume setting, and decreases or increases the gain of second speakers other than the first speakers in the opposite direction to the gain of the first speakers. Accordingly, the sound signal processing device 20 distributes gain changes across a plurality of speakers, thereby enabling necessary adjustment while reducing the influence on other listeners. Note that the sound signal processing device 20 may adjust both the gain G_ij and the gain SG_j according to the volume setting for each listener.
[0042] The sound signal processing device 20 sets the gain G_ij such that the loop gain LG_i does not exceed a predetermined value, and the sound pressure of the voice of the speaker S at the position of the listener L becomes 0 to 1 times the required sound pressure (the gain value received via the slider). However, since it is also possible to increase the gain of one speaker that covers a listener by decreasing the gain of speakers outside a certain range, the upper limit is not limited to 1 time, and may be a value larger than 1.
[0043] Accordingly, when there are a plurality of listeners, the sound signal processing device 20 of Modification 2 can set G_ij such that a desired sound pressure is obtained for each listener through a simple operation.
[0044] (Modification 3) FIG. 8 is a schematic block diagram showing the configuration of a sound signal processing system 1C according to Modification 3. The sound signal processing system 1C of Modification 3 includes an information processing terminal 50 for each of a plurality of speakers. In Modification 3, a GUI is displayed on the information processing terminal 50 for each of the plurality of speakers, and volume setting is accepted for each speaker.
[0045] The sound signal processing device 20 accepts volume settings for each speaker via the information processing terminal 50 for each speaker. The sound signal processing device 20 sets the gain Gij according to the volume setting for each speaker.
[0046] The sound signal processing device 20 sets the gain Gij for each speaker such that the loop gain LGi does not exceed a predetermined value, and the sound pressure of the speaker's voice at the position of listener L is 0 to 1 times the required sound pressure (the gain value accepted via the slider).
[0047] Thereby, when there are a plurality of speakers, the sound signal processing device 20 of Modification 3 can set Gij to achieve a desired sound pressure through a simple operation for each utterance.
[0048] (Modification 4) FIG. 9 is a schematic block diagram showing the configuration of a sound signal processing system 1C according to Modification 4. The sound signal processing system 1D of Modification 4 includes an information processing terminal 50 used by listener L. In Modification 4, the information processing terminal 50 for one listener accepts volume settings for each of a plurality of speakers.
[0049] The sound signal processing device 20 receives volume settings for each speaker via the information processing terminal 50. The sound signal processing device 20 sets the gain Gij according to the volume settings for each speaker. The sound signal processing device 20 adjusts the gain Gij as follows, for example. When the sound signal processing device 20 adjusts the gain for a particular speaker, for example, it adjusts the input gain to the first microphone responsible for that speaker to be lowered. Furthermore, if there is one or more beamforming microphones as in the modified example 1, the sound signal processing device 20 can adjust the sound pressure of a speaker by setting a gain for each beam and assigning a beam to each speaker, or by changing the gain for each beam depending on the direction, etc. Note that the sound signal processing device 20 may adjust not only the gain Gij, but also the gain GMi, or a combination of gain Gij and gain GMi.
[0050] The sound signal processing device 20 sets the gain Gij for each speaker such that the loop gain LGi does not exceed a predetermined value, and the sound pressure of the speaker's voice at the listener L's position is 0 to 1 times the required sound pressure (the gain value received by the slider).
[0051] As a result, in the modified example 4, the sound signal processing device 20 allows the listener to set Gij to the desired sound pressure with a simple operation for each utterance, even when there are multiple speakers.
[0052] The above description of the embodiments should be considered in all respects to be illustrative and not restrictive. The first embodiment and modifications 1 to 4 described above are not mutually exclusive and can be applied in appropriate combination. The scope of the present invention is indicated by the claims, not by the embodiments described above. Furthermore, the scope of the present invention is intended to include all modifications within the meaning and scope equivalent to the claims.
[0053] 1, 1A: Sound signal processing system, 10i: Microphone, 20: Sound signal processing device, 30j: Speaker, 100: Array microphone, 251: Display, 252: User I / F, 253: Flash memory, 254: CPU, 255: RAM, 256: Communication I / F, 300: Conference room
Claims
1. A sound signal processing method used in a sound signal processing system comprising a microphone for picking up the voice of a speaker and a plurality of speakers for amplifying and outputting the voice of the speaker to a listener, the method comprising: distributing the sound signal picked up by the microphone to the plurality of speakers; determining the loop gain to be fed back to the microphone for each of the plurality of speakers; accepting a volume setting; and, according to the accepted volume, setting the gain of the sound signal to be distributed to the plurality of speakers so that the loop gain does not exceed a predetermined value, and based on the distance attenuation of the voice that travels directly from the speaker to the listener and the transmission characteristics of the voice that travels from the microphone to the listener via the plurality of speakers.
2. The sound signal processing method according to claim 1, wherein the listener includes multiple listeners, and the volume setting is received for each of the multiple listeners via a plurality of volume controls corresponding to the multiple listeners, and the gain is set according to the volume setting for each listener.
3. The sound signal processing method according to claim 2, wherein the gain of the first speaker corresponding to the listener who has received the volume setting is increased or decreased, and the gain of the second speakers other than the first speaker is decreased or increased in the opposite direction to the gain of the first speaker.
4. The sound signal processing method according to claim 1 or 2, wherein the speaker includes multiple speakers, and the volume setting is received for each of the multiple speakers via a plurality of volume controls corresponding to the multiple speakers, and the gain is set according to the volume setting for each speaker.
5. The sound signal processing method according to claim 1 or claim 2, wherein the microphone includes a plurality of microphones, and the loop gain is determined for each of the plurality of microphones.
6. The sound signal processing method according to claim 1, wherein the microphone is an array microphone in which a plurality of microphone units are arranged, a plurality of sound pickup beams are set by combining the sound signals picked up by the plurality of microphone units, and the loop gain is determined for each of the plurality of sound pickup beams.
7. The sound signal processing method according to claim 6, wherein the listener includes a plurality of listeners, the volume setting is received for each of the plurality of listeners via a plurality of volume controls corresponding to the plurality of listeners, and the gain is adjusted for each of the plurality of sound pickup beams according to the volume setting for each listener.
8. An audio signal processing device used in an audio signal processing system comprising a microphone for picking up a speaker's voice and a plurality of speakers for amplifying and outputting the speaker's voice to a listener, the device comprising a processor that distributes the audio signal picked up by the microphone to the plurality of speakers, determines the loop gain to be fed back to the microphone for each of the plurality of speakers, accepts a volume setting, and sets the gain of the audio signal to be distributed to the plurality of speakers according to the accepted volume, such that the loop gain does not exceed a predetermined value, and based on the distance attenuation of the voice that travels directly from the speaker to the listener and the transmission characteristics of the voice that travels from the microphone to the listener via the plurality of speakers.
9. A sound signal processing system comprising a microphone for picking up the voice of a speaker and a plurality of speakers for amplifying and outputting the voice of the speaker to a listener, the system comprising: a sound signal processing device that distributes the sound signal picked up by the microphone to the plurality of speakers; for each of the plurality of speakers, determines the loop gain to be fed back to the microphone; accepts a volume setting; and, according to the accepted volume, sets the gain of the sound signal to be distributed to the plurality of speakers so that the loop gain does not exceed a predetermined value, and based on the distance attenuation of the voice that travels directly from the speaker to the listener and the transmission characteristics of the voice that travels from the microphone to the listener via the plurality of speakers.
10. An audio signal processing program used in an audio signal processing system comprising a microphone for picking up a speaker's voice and a plurality of speakers for amplifying and outputting the speaker's voice to a listener, which causes the audio signal processing device to perform the following processes: distribute the audio signal picked up by the microphone to the plurality of speakers; determine the loop gain to be fed back to the microphone for each of the plurality of speakers; accept a volume setting; and, according to the accepted volume, set the gain of the audio signal to be distributed to the plurality of speakers so that the loop gain does not exceed a predetermined value, and based on the distance attenuation of the voice that travels directly from the speaker to the listener and the transmission characteristics of the voice that travels from the microphone to the listener via the plurality of speakers.