Audio signal processing method and audio signal processing system

The audio signal processing method addresses howling and complex processing issues by adjusting feed amounts to specific speakers, ensuring clear audio output and suppression of howling in multi-speaker environments.

JP7708193B2Active Publication Date: 2025-07-15YAMAHA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023548414
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-16
Filing Date
2022-09-05
Publication Date
2025-07-15
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

Existing audio signal processing systems face issues with howling suppression and complex signal processing when multiple speakers speak simultaneously, leading to unnecessary reduction of voice levels and difficulty in clearly hearing other speakers.

Method used

An audio signal processing method that adjusts feed amounts of audio signals collected by multiple microphones to minimize output to specific speakers while maintaining clear audio for others, using a system with speaker voice detection and feed amount adjustment units to suppress howling without complex processing.

Benefits of technology

Effectively suppresses howling and ensures clear audio output from multiple speakers without complex signal processing, allowing each speaker to hear others clearly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708193000001
    Figure 0007708193000001
  • Figure 0007708193000002
    Figure 0007708193000002
  • Figure 0007708193000003
    Figure 0007708193000003
Patent Text Reader

Abstract

This sound signal processing method is used by a sound signal processing system comprising a plurality of sound emission devices corresponding to a plurality of sound collection devices. Feed amount adjustment processing is carried out, in which the levels of sound signals picked up by the plurality of sound collection devices are set to a preset feed amount, and the sound signals are transmitted to the plurality of sound emission devices. The feed amount adjustment processing involves making a change to minimize a first feed amount transmitted from, among the plurality of sound collection devices, a first sound collection device which detected a speaker's voice to a first sound emission device closest to the first sound collection device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing apparatus performed when sound collected by a microphone is emitted from a speaker.

Background Art

[0002] Patent Document 1 and Patent Document 2 describe devices for suppressing howling.

[0003] Patent Document 1 describes a public address system including a plurality of microphones and a plurality of speakers. The public address system according to Patent Document 1 detects a speaker based on input signals from the plurality of microphones and selects a microphone corresponding to the position of the speaker. The public address system according to Patent Document 1 reduces the output level of a speaker arranged near the selected microphone.

[0004] Patent Document 2 describes a filter setting device including a plurality of microphones, a plurality of speakers, a mixer, an input filter, and an output filter. The filter setting device described in Patent Document 2 adjusts the frequency characteristics of the output filter so as to change the loop gain of an integrated system with respect to a sound signal mixed by the mixer and output to the plurality of speakers.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, in the configuration described in Patent Document 1, for example, when multiple speakers at different positions speak simultaneously, the level of other speaker voices output from the same speaker decreases along with the speaker voice to be suppressed in the speaker output. That is, the voice amplification system of Patent Document 1 unnecessarily reduces the level of the speaker output.

[0007] Also, the configuration of Patent Document 2 requires complex signal processing.

[0008] In view of the above circumstances, one aspect of the present disclosure aims to provide an audio signal processing method that does not require complex signal processing, suppresses howling, and can appropriately output voices other than the voice to be suppressed by howling.

Means for Solving the Problems

[0009] The audio signal processing method is an audio signal processing method used in an audio signal processing system including a plurality of sound collection devices and a corresponding plurality of sound playback devices. The method performs a feed amount adjustment process of setting the levels of the audio signals collected by the plurality of sound collection devices to a preset feed amount and transmitting the audio signals to the plurality of sound playback devices. In the feed amount adjustment process, the first feed amount transmitted from the first sound collection device that has detected the speaker voice among the plurality of sound collection devices to the first sound playback device closest to the first sound collection device is changed to be minimized.

Effects of the Invention

[0010] The audio signal processing method does not require complex signal processing, suppresses howling, and can appropriately output voices other than the voice to be suppressed by howling.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0012] A sound signal processing method and a sound signal processing system according to embodiments of the present invention will be described with reference to the drawings.

[0013] [Embodiment 1] [Audio conference environment (microphone (microphone), speaker, positional relationship of speakers)] FIG. 1 is a diagram showing an example of the arrangement of a plurality of microphones and a plurality of speakers in a space where an audio conference is held. As shown in FIG. 1, in a space (for example, a conference room) where an audio conference is held, there are a plurality of microphones MIC11 - MIC25 (MIC11, MIC12, MIC13, MIC14, MIC15, MIC21, MIC22, MIC23, MIC24, MIC25), and a plurality of speakers SP11 - SP43 (SP11, SP12, SP13, SP14, SP15, SP21, SP22, SP23, SP24, SP25, SP31, SP32, SP33, SP41, SP42, SP43). These microphones correspond to the "sound collection device" of the present invention, and these speakers correspond to the "sound playback device" of the present invention.

[0014] The plurality of microphones MIC11 - MIC15 are arranged side by side along the first side in the vicinity of the first side on the table TBL. The plurality of microphones MIC21 - MIC25 are arranged side by side along the second side (the side opposite to the first side) in the vicinity of the second side on the table TBL. The column of the plurality of microphones MIC11 - MIC15 and the column of the plurality of microphones MIC21 - MIC25 are separated with the table TBL in between.

[0015] The plurality of speakers SP11 - SP15 are arranged side by side along the first side of the table TBL near the first side. The plurality of speakers SP21 - SP25 are arranged side by side along the second side of the table TBL near the second side. The speaker SP11 is the speaker closest to the microphone MIC11. Similarly, as shown in FIG. 1, the speakers SP12 - SP15, SP21 - SP25 are the speakers closest to the microphones MIC12 - MIC15, MIC21 - MIC25 respectively. As shown in FIG. 1, the plurality of speakers SP31 - SP33 are arranged side by side along the third side (the side orthogonal to the first side and the second side) of the table TBL near the third side. The plurality of speakers SP41 - SP43 are arranged side by side along the fourth side (the side opposite to the third side) of the table TBL near the fourth side. The plurality of speakers SP11 - SP43 mainly emit sound in the direction of the speaker closest to each of them. Note that in FIG. 1, the plurality of speakers SP11 - SP43 are not arranged on the table TBL, but they may be arranged on the table TBL.

[0016] The arrangement of these plurality of microphones and plurality of speakers is not limited to the example of FIG. 1. The arrangement of the plurality of microphones and the plurality of speakers may be any arrangement as long as the speaker closest to each microphone can be determined. Also, the number of arrangements of the plurality of microphones and the number of arrangements of the plurality of speakers are not limited to the above numbers, and can be set according to the number of speakers conducting the meeting, etc.

[0017] The plurality of speakers 911 - 915, 921 - 925 are each near the plurality of microphones MIC11 - MIC15, MIC21 - MIC25 and conduct a meeting (conversation). More specifically, for example, the speaker 911 is near the microphone MIC11 and conducts a meeting using the microphone MIC11. At this time, it is preferable that the plurality of microphones MIC11 - MIC25 are set with directivity etc. so as to pick up the voices of the nearby speakers and hardly pick up the voices of other speakers.

[0018] [Configuration and Processing of Audio Signal Processing Device] FIG. 2 is a functional block diagram showing an example of the configuration of an audio signal processing system according to the first embodiment of the present invention. FIG. 3 is a flowchart showing an example of an audio signal processing method according to the first embodiment of the present invention.

[0019] As shown in FIG. 2, the audio signal processing system 10 includes a speaker voice detection unit 11, a feed amount adjustment unit 12, a plurality of microphones MIC11 - MIC25, a plurality of speakers SP11 - SP43, and a plurality of output amplifiers A11 - A43. Physically, the audio signal processing system 10 includes a main body device, a plurality of microphone housings, and a plurality of speaker housings.

[0020] The main body device is an information processing device including a DSP, a CPU, etc., a storage medium, and a predetermined electronic circuit. The main body device realizes the functions of the speaker voice detection unit 11 and the feed amount adjustment unit 12 by executing an audio signal processing program stored in the storage medium using a DSP, a CPU, etc. Also, the main body device realizes a plurality of output amplifiers A11 - A43 by a plurality of amplification devices.

[0021] The plurality of microphones MIC11 - MIC25 each have an individual microphone housing separately from the main body device. As shown in FIG. 1, the plurality of microphones MIC11 - MIC25 are arranged at predetermined positions in the space where the meeting is held. The plurality of microphones MIC11 - MIC25 are physically and electrically connected to the main body device.

[0022] The plurality of speakers SP11 - SP43 each have an individual speaker housing separately from the main body device. As shown in FIG. 1, the plurality of speakers SP11 - SP43 are arranged at predetermined positions in the space where the meeting is held. The plurality of speakers SP11 - SP43 are physically and electrically connected to the main body device.

[0023] More specifically, the audio signal processing system 10 has the following circuit configuration. A plurality of microphones MIC11 - MIC25 are connected to the speaker voice detection unit 11 and the feed volume adjustment unit 12. The plurality of microphones MIC11 - MIC25 acquire the sound collection signals and output the sound collection signals Sm11 - Sm25 to the speaker voice detection unit 11 and the feed volume adjustment unit 12 (S11). Specifically, the microphone MIC11 outputs the sound collection signal Sm11 to the speaker voice detection unit 11 and the feed volume adjustment unit 12. Similarly, the plurality of microphones MIC12 - MIC15, MIC21 - MIC25 output the sound collection signals Sm12 - Sm15, Sm21 - Sm25 to the speaker voice detection unit 11 and the feed volume adjustment unit 12.

[0024] [Detection of Speaker Voice] The speaker voice detection unit 11 detects the speaker voice using the plurality of sound collection signals Sm11 - Sm25 (S12). That is, the speaker voice detection unit 11 detects the microphone that picks up the voice of the speaker during pronunciation (speech). This microphone corresponds to the "first sound collection device" of the present invention. And this microphone is a specific microphone that is the target for adjusting the feed volume, and is hereinafter referred to as the "specific microphone".

[0025] FIG. 4 is a functional block diagram showing an example of the configuration of the speaker voice detection unit. As shown in FIG. 4, the speaker voice detection unit 11 includes a signal level detection unit 111 and a specific microphone detection unit 112.

[0026] The sound collection signals Sm11 - Sm25 of the plurality of microphones MIC11 - MIC25 are input to the signal level detection unit 111. The signal level detection unit 111 detects the signal levels (amplitudes) of the plurality of sound collection signals Sm11 - Sm25. The signal level detection unit 111 outputs the signal levels (amplitudes) of the plurality of sound collection signals Sm11 - Sm25 to the specific microphone detection unit 112.

[0027] The specific microphone detection unit 112 stores in advance a threshold value for speaker detection (hereinafter simply referred to as the threshold value). The threshold value is set to a value higher than, for example, the noise level of the space where the meeting is held and lower than the level of the sound collection signal when the speech in the meeting is picked up by the microphone.

[0028] The specific microphone detection unit 112 compares the signal levels of the plurality of sound collection signals Sm11 - Sm25 with a threshold value. The specific microphone detection unit 112 detects a sound collection signal that is equal to or greater than the threshold value. In other words, the specific microphone detection unit 112 detects, as a specific microphone, the microphone that has output a sound collection signal equal to or greater than the threshold value (S13). The specific microphone detection unit 112 outputs information on the specific microphone to the feed amount adjustment unit 12.

[0029] Generally speaking, the feed amount adjustment unit 12 adjusts the feed amounts of the plurality of sound collection signals Sm11 - Sm25 for the plurality of speakers SP11 - SP43 based on the specific microphone.

[0030] FIG. 5 is a functional block diagram showing an example of the configuration of the feed amount adjustment unit. As shown in FIG. 5, the feed amount adjustment unit 12 includes a speaker determination unit 121, a feed amount setting unit 122, a mixer 123, and a speaker DB 120. The feed amount setting unit 122 includes a suppression amount setting unit 1220.

[0031] The relationship between the specific microphone and the speakers to be adjusted for the feed amount is stored in the speaker DB120. FIG. 6 is a table showing the relationship between the specific microphone and the speakers stored in the speaker DB 120 As shown in FIG. 6, for each specific microphone, the first speaker Ka and The second speaker Ka and is associated. The first speaker is the speaker closest to the specific microphone. The second speaker is the speaker next closest to the specific microphone after the first speaker. This first speaker corresponds to the "first sound emitting device" of the present invention, and the second speaker corresponds to the "second sound emitting device" of the present invention.

[0032] The speaker determination unit 121 refers to the speaker DB 120 based on the information of the specific microphone, and determines the speaker to be adjusted for the feed amount (S14). Specifically, the speaker determination unit 121 refers to the speaker DB 120 based on the information of the specific microphone, and determines the first speaker and the second speaker for the specific microphone. For example, if the specific microphone is the microphone MIC13, the speaker determination unit 121 determines the speaker SP13 as the first speaker, and determines the speakers SP12 and SP14 as the second speakers. The speaker determination unit 121 outputs the information of the first speaker and the information of the second speaker to the feed amount setting unit 122.

[0033] The feed amount setting unit 122 sets the feed amounts for the sound pickup signals of the plurality of microphones MIC11 - MIC25 with respect to the plurality of speakers SP11 - SP43 based on the combination of the specific microphone, the first speaker, and the second speaker (S15). The feed amount for the first speaker corresponds to the "first feed amount" of the present invention, and the feed amount for the second speaker corresponds to the "second feed amount" of the present invention.

[0034] If the information indicating the combination of the specific microphone, the first speaker, and the second speaker is not input, the feed amount setting unit 122 sets the feed amount to a preset value for all combinations of the plurality of microphones MIC11 - MIC25 and the plurality of speakers SP11 - SP43. The preset feed amount is basically the same feed amount in all combinations of the plurality of microphones MIC11 - MIC25 and the plurality of speakers SP11 - SP43, but is not limited to exact coincidence, and may include an error or an adjustment difference for each speaker. Such a feed amount is referred to as a reference feed amount. Note that the feed amount is the adjustment amount of the signal level (amplitude) of the sound signal sent from the microphone to the speaker. The feed amount can be set, for example, by the gain for the sound pickup signal.

[0035] If the information indicating the combination of the specific microphone, the first speaker, and the second speaker is input, the feed amount setting unit 122 sets the feed amounts from the specific microphone to the first speaker and the second speaker to the reference feed quantity It is also adjusted to be small. At this time, the feed amount setting unit 122 uses the suppression amount set by the suppression amount setting unit 1220, which will be described in detail later. More specifically, the feed amount setting unit 122 adjusts the feed amount from the specific microphone to the first speaker to a minimum value smaller than the reference feed amount. Note that the feed amount from the specific microphone to the first speaker may be 0. That is, it is also possible not to send the sound collection signal from the specific microphone to the first speaker. Also, the feed amount setting unit 122 adjusts the feed amount from the specific microphone to the second speaker to be smaller than the reference feed amount and larger than the feed amount to the first speaker. Then, the feed amount setting unit 122 sets the feed amounts other than the feed amounts from the specific microphone to the first and second speakers to the reference feed amount.

[0036] The feed amount setting unit 122 outputs the feed amount for each combination of the plurality of microphones MIC11 - MIC25 and the plurality of speakers SP11 - SP43 to the mixer 123.

[0037] The mixer 123 is realized by a so-called matrix mixer. The mixer 123 uses the feed amount from the feed amount setting unit 122 to adjust the signal levels of the sound collection signals Sm11 - Sm25 of the plurality of microphones MIC11 - MIC25 and mix them for each of the plurality of speakers SP11 - SP43 (S16).

[0038] For example, when not all of the sound collection signals Sm11 - Sm25 are detected as speaker voices, the mixer 123 adjusts the levels of all the sound collection signals Sm11 - Sm25 with the reference feed amount. The mixer 123 mixes these level-adjusted sound collection signals for each of the plurality of speakers SP11 - SP43.

[0039] On the one hand, when a specific sound pickup signal is detected as the speaker's voice, the mixer 123 adjusts the level of the sound pickup signal sent from the specific microphone to the first speaker and the second speaker to be smaller than the reference sending amount. The mixer 123 adjusts the sending amounts of signals other than the sound pickup signal from the specific microphone to the first speaker and the second speaker at the reference sending amount. The mixer 123 mixes these level-adjusted sound pickup signals Sm11 - Sm25 for each of the plurality of speakers SP11 - SP43.

[0040] By performing such mixing processing, the mixer 123 generates speaker supply signals Ss11 - Ss43 for each of the plurality of speakers SP11 - SP43.

[0041] As a result, the level at which the sound pickup signal of the specific microphone is output from the first speaker is significantly reduced. Also, the level at which the sound pickup signal of the specific microphone is output from the second speaker is reduced. Therefore, howling via the specific microphone can be suppressed. On the other hand, the levels at which the sound pickup signal of the specific microphone is output from speakers other than the first speaker and the second speaker can be ensured at levels that can be heard by other speakers.

[0042] The sending amount adjustment unit 12 transmits (outputs) the speaker supply signals Ss11 - Ss43 to the output amplifiers A11 - A43 respectively. The output amplifiers A11 - A43 amplify the speaker supply signals Ss11 - Ss43 respectively to generate speaker drive signals So11 - So43. The output amplifiers A11 - A43 transmit (output) the speaker drive signals So11 - So43 to the plurality of speakers SP11 - SP43 respectively (S17).

[0043] Note that the processing of the sending amount adjustment unit 12 described above is processing as one means for realizing the following concept. FIGS. 7 and 8 are diagrams showing the concept of adjusting the sending amount according to an embodiment of the present invention. FIG. 7 shows the case where there is one speaker during speech. FIG. 8 shows the case where there are two speakers speaking simultaneously. The sizes of the black circles in FIGS. 7 and 8 and the sizes of the white circles in FIG. 8 indicate the magnitudes of the sending amounts.

[0044] (When there is only one speaker speaking) As shown in FIG. 7, when only the speaker 913 is speaking, among the plurality of sound collection signals Sm11 - Sm25, the signal level of the sound collection signal Sm13 becomes the maximum value equal to or greater than the threshold value, and the signal levels of the other sound collection signals are less than the threshold value.

[0045] The feed amount adjustment unit 12 sets the feed amount to all speakers to the reference feed amount with respect to the sound collection signal Sm13 caused by the speech of the speaker 913. That is, as shown by the black circles in FIG. 7, the feed amount adjustment unit 12 sets the same feed amount for all the speakers SP11 - SP43 with respect to the sound collection signal Sm13.

[0046] The feed amount adjustment unit 12 further performs the following processing. The feed amount adjustment unit 12 determines the speaker SP13 as the first speaker and the speakers SP12 and SP14 as the second speakers based on the microphone MIC13 detected as the specific microphone. The feed amount adjustment unit 12 sets the suppression amount for the speakers SP13, SP12, and SP14 of the sound collection signal Sm13, and does not set the suppression amount for the other speakers.

[0047] More specifically, the feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 (specific microphone) and the speaker SP13 (first speaker) to be approximately the same as the reference feed amount, as shown in FIG. 7. The feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 (specific microphone) and the speakers SP12 and SP14 (second speakers) to be smaller than the reference feed amount and a non - zero value, as shown in FIG. 7. Here, the suppression amount is set as a reference for the concept of the present invention. However, the feed amount adjustment unit 12 may simplify the processing by not setting the suppression amount and simply reducing the feed amount for the combination of the specific microphone and the first and second speakers.

[0048] At this time, it is more preferable that the feed rate adjustment unit 12 performs the following processing. Specifically, the feed rate adjustment unit 12 sets the suppression amount according to the following concept. The feed rate adjustment unit 12 sets the gain corresponding to the reference feed rate as gain A, and sets the degree of suppression as suppression index α. Also, the feed rate adjustment unit 12 sets the distance between the microphone MIC13 and the speaker SP13 as distance r1, and sets the distance between the microphone MIC13 and the speakers SP12 and SP14 as distance r2. Further, the feed rate adjustment unit 12 sets the influence degree of the distance on the suppression amount as influence degree n.

[0049] In this case, the suppression amount for the speaker SP13 is set to A / (r1 n ·α). Also, the suppression amount for the speakers SP12 and SP14 is set to A / (r2 n ·α). Thereby, the feed rate adjustment unit 12 can set the suppression amount as described above. These suppression amounts are values corresponding to the positional relationship between a plurality of speakers and a plurality of microphones.

[0050] The feed rate adjustment unit 12 sets the feed rate for all the speakers SP11 - SP43 by subtracting the suppression amount from the reference feed rate. Thereby, the feed rate adjustment unit 12 makes the feed rate of the sound collection signal Sm13 for the speaker SP13 smaller than the reference feed rate and adjusts it to the minimum value. Also, the feed rate adjustment unit 12 adjusts the feed rate of the sound collection signal Sm13 for the speakers SP12 and SP14 to be smaller than the reference feed rate. Further, the feed rate adjustment unit 12 sets the feed rate of the sound collection signal Sm13 for the speakers other than SP12, SP13, and SP14 to the reference feed rate.

[0051] (When there are two speakers speaking simultaneously) As shown in FIG. 8, when the speakers 913 and 924 are speaking simultaneously and no other speakers are speaking, the signal levels is the threshold of the sound collection signals Sm13 and Sm24 are is the threshold equal to or higher than a certain value, and the signal levels

[0052] The feed amount adjustment unit 12 sets the feed amount to all speakers to the reference feed amount for the sound collection signals Sm13 and Sm24. That is, as shown by the black circles in FIG. 8, the feed amount adjustment unit 12 sets the same feed amount for all speakers SP11 - SP43 for the sound collection signal Sm13. Also, as shown by the white circles in FIG. 8, the feed amount adjustment unit 12 sets the same feed amount for all speakers SP11 - SP43 for the sound collection signal Sm24.

[0053] The feed amount adjustment unit 12 further performs the following processing. Based on the microphone MIC13 detected as the specific microphone, the feed amount adjustment unit 12 determines the speaker SP13 as the first speaker and determines the speakers SP12 and SP14 as the second speakers. The feed amount adjustment unit 12 sets the suppression amounts for the speakers SP13, SP12, and SP14 of the sound collection signal Sm13. The feed amount adjustment unit 12 does not set the suppression amount for other speakers.

[0054] More specifically, the feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 and the speaker SP13 (the first speaker) to be approximately the same as the reference feed amount, as shown in FIG. 8. The feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 and the speakers SP12 and SP14 (the second speakers) to be smaller than the reference feed amount and a non - zero value, as shown in FIG. 8. Here too, the suppression amount is set with reference to the concept of the present invention. However, similar to the above - mentioned case, the feed amount adjustment unit 12 may simplify the process by not setting the suppression amount and simply reducing the feed amount.

[0055] Based on the microphone MIC24 detected as the specific microphone, the feed amount adjustment unit 12 determines the speaker SP24 as the first speaker and determines the speakers SP23 and SP25 as the second speakers. The feed amount adjustment unit 12 sets the suppression amounts for the speakers SP24, SP23, and SP25 of the sound collection signal Sm24. The feed amount adjustment unit 12 does not set the suppression amount for other speakers.

[0056] More specifically, as shown in FIG. 8, the feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC24 and the speaker SP24 (first speaker) to be approximately the same as the reference feed amount. The feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC24 and the speakers SP23, SP25 (second speakers) to be less than the reference feed amount and a non-zero value, as shown in FIG. 8. Here too, the suppression amount is set with reference to the concept of the present invention. However, as in the above case, the feed amount adjustment unit 12 may simplify the process by simply reducing the feed amount without setting the suppression amount.

[0057] The feed amount adjustment unit 12 sets the feed amount for all the speakers SP11 - SP43 by subtracting the suppression amount from the reference feed amount. As a result, as shown in FIG. 8, the feed amount adjustment unit 12 reduces the feed amount of the sound collection signal Sm13 for the speaker SP13 to be less than the reference feed amount and adjusts it to the minimum value. Also, the feed amount adjustment unit 12 adjusts the feed amount of the sound collection signal Sm13 for the speakers SP12, SP14 to be less than the reference feed amount. Further, the feed amount adjustment unit 12 sets the feed amount of the sound collection signal Sm13 for the speakers other than SP12, SP13, SP14 to the reference feed amount. Also, the feed amount adjustment unit 12 reduces the feed amount of the sound collection signal Sm24 for the speaker SP24 to be less than the reference feed amount and adjusts it to the minimum value. Also, the feed amount adjustment unit 12 adjusts the feed amount of the sound collection signal Sm24 for the speakers SP23, SP25 to be less than the reference feed amount. Further, the feed amount adjustment unit 12 sets the feed amount of the sound collection signal Sm24 for the speakers other than SP23, SP24, SP25 to the reference feed amount.

[0058] Therefore, in the speaker supply signal Ss13, the component of the sound collection signal Sm13 is significantly suppressed. On the other hand, in the speaker supply signal Ss13, the component of the sound collection signal Sm24 is not suppressed. Also, in the speaker supply signals Ss12, Ss14, the component of the sound collection signal Sm13 is suppressed to some extent. On the other hand, in the speaker supply signals Ss12, Ss14, the component of the sound collection signal Sm24 is not suppressed.

[0059] This suppresses howling caused by the microphone MIC13 and the plurality of speakers SP13, SP12, and SP14. Furthermore, the speaker 913 can clearly hear the speech of the speaker 924 even while speaking himself / herself.

[0060] Similarly, in the speaker supply signal Ss24, the component of the sound collection signal Sm24 is significantly suppressed. On the other hand, in the speaker supply signal Ss24, the component of the sound collection signal Sm13 is not suppressed. Also, in the speaker supply signals Ss23 and Ss25 to the speakers SP23 and SP25, the component of the sound collection signal Sm24 is suppressed to some extent. On the other hand, in the speaker supply signals Ss23 and Ss25, the component of the sound collection signal Sm13 is not suppressed.

[0061] This suppresses howling caused by the microphone MIC24 and the plurality of speakers SP24, SP23, and SP25. Furthermore, the speaker 924 can clearly hear the speech of the speaker 913 even while speaking himself / herself.

[0062] When the processes shown in FIGS. 7 and 8 are realized by a plurality of functional blocks, for example, the feed amount setting unit 122 shown in FIG. 5 includes a suppression amount setting unit 1220. The suppression amount setting unit 1220 sets the suppression amount as described above using the combination of the specific microphone, the first speaker, and the second speaker determined by the speaker determination unit 121. The feed amount setting unit 122 sets the feed amount for all combinations of the microphones and the speakers by subtracting the suppression amount set by the suppression amount setting unit 1220 from the reference feed amount.

[0063] With the configuration and processes as described above, the sound signal processing system 10 can emit the speech of the speaker from the plurality of speakers SP11 - SP43 as shown in FIGS. 9(A) and 9(B).

[0064] FIG. 9(A) and FIG. 9(B) are diagrams showing an example of the magnitude of the output sound from a plurality of speakers. FIGS. 9(A) and 9(B) have the same arrangement of a plurality of microphones, a plurality of speakers, and a plurality of speakers as in FIG. 1. FIG. 9(A) shows the case where only speaker 913 is speaking, and FIG. 9(B) shows the case where speaker 913 and speaker 924 are speaking simultaneously. In FIGS. 9(A) and 9(B), the hatched circles indicate the volume of each speaker when the voice of speaker 913 is emitted from the plurality of speakers SP11 - SP43, and the larger the radius of the circle, the larger the volume. In FIG. 9(B), the non-filled circles indicate the volume of each speaker when the voice of speaker 924 is emitted from the plurality of speakers SP11 - SP43, and the larger the radius of the circle, the larger the volume.

[0065] As shown in FIG. 9(A), when speaker 913 is speaking, the volume of speaker SP13 is the smallest, and the volumes of speakers SP12 and SP14 are the next smallest. And the volumes of the other plurality of speakers SP11, SP15, SP21 - SP43 become a predetermined value larger than the volumes of speakers SP12 - SP14. Thereby, the volume of the voice of speaker 913 propagating from speaker SP13 to microphone MIC13 can be suppressed, and howling can be suppressed. Furthermore, the volume of the voice of speaker 913 propagating from speakers SP12 and SP14 to microphone MIC13 can be suppressed, and howling can be further suppressed.

[0066] At this time, the volumes of speakers SP12 and SP14 are larger than the volume of speaker SP13. Therefore, speaker 912 and speaker 914 can hear the direct sound from speaker 913 and the voice of speaker 913 output from speakers SP12 and SP14. Therefore, according to the present embodiment, it is possible to suppress the difficulty of hearing the voice of speaker 913.

[0067] As shown in FIG. 9(B), when speaker 913 and speaker 924 are speaking simultaneously, various processes related to the voice of speaker 913 are as described above, and howling caused by the voice of speaker 913 is suppressed. On the other hand, as shown in FIG. 9(B), the speaker SP24 near speaker 924, who is speaking simultaneously with speaker 913, emits sound without suppressing the voice of speaker 913. Thus, speaker 924 can hear the voice of speaker 913 from speaker SP24 even while speaking.

[0068] Also, similar to the case of speaker 913 described above, as shown by the unshaded circles in FIG. 9(B), the volume of speaker SP24 is the smallest for the sound signal from microphone MIC24, and the volumes of speakers SP23 and SP25 are the next smallest. Then, the volumes of the other multiple speakers SP11 - SP15, SP21, SP22, SP31 - SP43 are set to a predetermined value larger than the volumes of speakers SP23 - SP25. Thereby, the volume of the voice of speaker 924 propagating from speaker SP24 to microphone MIC24 is suppressed, and howling can be suppressed. Furthermore, the volume of the voice of speaker 924 propagating from speakers SP23 and SP25 to microphone MIC24 is suppressed, and howling can be further suppressed.

[0069] Also, as shown in FIG. 9(B), speaker SP13 emits sound without suppressing the volume of the sound signal including the voice of speaker 924. Thus, speaker 913 can hear the voice of speaker 924 from speaker SP13 even while speaking.

[0070] In this way, the sound signal processing system 10 can suppress howling for each speaker while enabling each speaker to clearly hear the voices of other speakers when multiple speakers are speaking simultaneously. In the above description, the case where two people are speaking simultaneously is shown, but the same processing can be applied when three or more people are speaking simultaneously.

[0071] In the prior art, when suppressing howling as described above, the output to a plurality of speakers SP11 - SP43 is suppressed. That is, the configuration of the prior art adjusts the volume of the speaker drive signals So11 - So43. In this case, when multiple speakers are speaking simultaneously, the sound output from the speakers near one speaker becomes a sound in which the voices of other speakers who are speaking are also suppressed. Therefore, it becomes difficult for a speaker to clearly hear the voices of other speakers who are speaking simultaneously.

[0072] As described above, the sound signal processing system 10 can suppress howling and appropriately output voices other than the target of howling suppression. Further, the sound signal processing system 10 does not perform complex filter coefficient settings for howling suppression and does not require complex signal processing for howling suppression.

[0073] In the above description, speaker when there is one speaker or when there is one in each of a plurality of columns speaker when they are separated). However, when a plurality of speaker (simultaneously speaker ) are close to each other, the sound signal processing system 10 can suppress howling and appropriately output voices other than the target of howling suppression by using, for example, the following method.

[0074] For example, when the waveforms of a plurality of sound collection signals are equal to or greater than a threshold value, the specific microphone detection unit 112 compares the waveforms of these plurality of sound collection signals and detects whether the plurality of sound collection signals are the same speaker voice. If the plurality of sound collection signals are different speaker voices, the specific microphone detection unit 112 detects a plurality of specific microphones corresponding to each of the plurality of sound collection signals. The feed amount adjustment unit 12 performs the above-described feed amount adjustment process based on the plurality of specific microphones. At this time, if the plurality of sound collection signals are the same speaker voice, the specific microphone detection unit 112 detects the microphone corresponding to the sound collection signal with the maximum value as the specific microphone. Then, the feed amount adjustment unit 12 performs the above-described feed amount adjustment process based on this specific microphone.

[0075] [Embodiment 2] FIG. 10 is a functional block diagram showing an example of the configuration of a sound signal processing system according to a second embodiment of the present invention. FIG. 11 is the configuration of a feed amount adjustment unit according to the second embodiment of It is a functional block diagram showing an example. FIG. 12 is a flowchart showing an example of a sound signal processing method according to the second embodiment of the present invention.

[0076] The sound signal processing system 10A according to the second embodiment is different in that it includes a feed amount adjustment unit 12A instead of the feed amount adjustment unit 12 of the sound signal processing system 10 according to the first embodiment. Other configurations of the sound signal processing system 10A are the same as those of the sound signal processing system 10, and descriptions of the same parts are omitted.

[0077] The feed amount adjustment unit 12A is different in that it includes a mixer 123A instead of the mixer 123 of the feed amount adjustment unit 12 according to the first embodiment. Other configurations of the feed amount adjustment unit 12A are the same as those of the feed amount adjustment unit 12, and descriptions of the same parts are omitted.

[0078] As shown in FIG. 11, the mixer 123A is realized by a matrix mixer in the same manner as the mixer 123 of the first embodiment. However, the mixer 123A includes a filter that applies reverberation for each combination of a plurality of microphones MIC11 - MIC25 and a plurality of speakers SP11 - SP43. The filter that applies reverberation is, for example, an FIR filter or the like.

[0079] Specifically, a plurality of sound collection signals Sm11 - Sm25 are input to the mixer 123A. The mixer 123A performs convolution operation processing on the plurality of sound collection signals Sm11 - Sm25 with filter coefficients corresponding to each of the plurality of speakers SP11 - SP43. Thereby, the mixer 123A imparts a plurality of reverberant sounds for each combination of the plurality of microphones MIC11 - MIC25 and the plurality of speakers SP11 - SP43 (S21). At this time, the mixer 123A sets the filter coefficients for the convolution operation processing based on the positional relationship between each of the plurality of microphones MIC11 - MIC25 and the plurality of speakers SP11 - SP43, and the environment of the conference room (the size of the conference room, the structure of the walls and ceiling, etc.).

[0080] The filter coefficients for the convolution operation processing are set as follows, for example. The reverberant sound is composed of an early reflection sound component and a reverberation sound component. The early reflection sound is the sound that arrives at the sound reception point at an early time after the sound generated at the generation position (speaker position) is reflected by the wall, floor, or ceiling. Therefore, the filter coefficients for the convolution operation processing for the early reflection sound component are set based on the position of the speaker (the position of the microphone), the position of the speaker, and the environment of the conference room (the size of the conference room, the structure of the walls and ceiling, etc.). The reverberation sound is the sound that arrives at the sound reception point after the sound generated at the generation position is multiply reflected and arrives at the sound reception point following the early reflection sound. Therefore, the filter coefficients for the convolution operation processing for the reverberation sound component are set based on the environment of the conference room (the size of the conference room, the structure of the walls and ceiling, etc.). Note that the setting methods for these early reflection sound components and reverberation sound components are just examples, and other setting methods can also be used. Also, these early reflection sounds and reverberation sounds respectively correspond to the "indirect sound" of the present invention.

[0081] When such reverberant sounds are imparted, the mixer 123A adjusts the level of the sound collection signal with the reverberant sound imparted using the feed amount set by the feed amount setting unit 122 (S25). For example, the mixer 123A sets a coefficient for adjusting the amplitude level based on the feed amount and multiplies it by the sound collection signal with the reverberant sound imparted. Thereby, the mixer 123A can generate a sound collection signal with the amplitude level adjusted by the feed amount and the reverberant sound imparted.

[0082] The mixer 123A adjusts the amplitude level according to the feed rate and mixes the sound pickup signal with reverberation added thereto for each of the plurality of speakers SP11 - SP43 (S26). Thereby, the mixer 123A generates speaker supply signals Ss11r - Ss43r for each of the plurality of speakers SP11 - SP43 and outputs them to the plurality of output amplifiers A11 - A43. The plurality of output amplifiers A11 - A43 transmit the speaker supply signals Ss11r - Ss43r to the plurality of speakers SP11 - SP43 (S27).

[0083] With such a configuration, the sound signal processing system 10A exhibits the same operational effects as the sound signal processing system 10, and can provide each speaker with a realistic sound according to the environment of the conference room, the position of the speaker, and the position of each speaker from the plurality of speakers SP11 - SP43.

[0084] The reverberation may include at least one of an early reflection component and a reverberation component. At this time, by using the reverberation component, the sound signal processing system 10 A can realize a realistic conversation according to the shape of the conference room, the state of the wall surface, etc. Also, by using the early reflection component, the sound signal processing system 10 A can realize a realistic conversation according to the position of the speaker in the conference room and the position of the listener (another speaker listening to the voice of the speaker speaking).

[0085] In the above configuration, the sound signal processing system 10A applies reverberation for each combination of the plurality of microphones and the plurality of speakers. However, the sound signal processing system 10A can also apply reverberation, for example, as follows. The sound signal processing system 10A divides the plurality of microphones and the plurality of speakers into a plurality of groups according to their respective positions. The sound signal processing system 10A applies reverberation for each combination of the plurality of microphone groups and the plurality of speaker groups. Thereby, the sound signal processing system 10A can reduce the load of signal processing while imparting a sense of presence.

[0086] [Embodiment 3] FIG. 13 is a functional block diagram showing an example of the configuration of a sound signal processing system according to a third embodiment of the present invention. As shown in FIG. 13, a sound signal processing system 10B according to the third embodiment is different from the sound signal processing system 10 according to the first embodiment in that a plurality of equalizers are additionally provided. Other configurations of the sound signal processing system 10B are the same as those of the sound signal processing system 10, and descriptions of the same parts are omitted.

[0087] The sound signal processing system 10B includes a plurality of equalizers EQ11 - EQ43. A plurality of speaker supply signals Ss11 - Ss43 are input from the feed amount adjustment unit 12 to the plurality of equalizers EQ11 - EQ43.

[0088] The plurality of equalizers EQ11 - EQ43 generate tone adjustment signals Sq11 - Sq43 by performing predetermined signal processing on each of the plurality of speaker supply signals Ss11 - Ss43. The plurality of equalizers EQ11 - EQ43 output the plurality of tone adjustment signals Sq11 - Sq43 to the plurality of output amplifiers A11 - A43, respectively.

[0089] With such a configuration and processing, the sound signal processing system 10B exhibits the same operational effects as the sound signal processing system 10, and can emit sound with a desired tone color for each of the plurality of speakers.

[0090] Note that the parameter settings of the plurality of equalizers EQ11 - EQ43 may be input by the conference administrator or each speaker. Furthermore, the parameters of the plurality of equalizers EQ11 - EQ43 may be set based on the environment of the conference room or the like.

[0091] [Embodiment 4] FIG. 14 is a functional block diagram showing an example of the configuration of a sound signal processing system according to a fourth embodiment of the present invention. FIG. 15 is a functional block diagram showing an example of the configuration of a conference device used in the sound signal processing system according to the fourth embodiment.

[0092] In the above-described Embodiments 1 to 3, the case where all speakers gather in one conference room (physical space) to hold a meeting was shown. Embodiment 4 shows the case where speakers are in different conference rooms (different physical spaces) to hold a meeting. Note that the total number of speakers and the number of speakers in one conference room shown in this embodiment are examples and are not limited thereto.

[0093] As shown in FIG. 14, the audio signal processing system 10C includes a plurality of conference devices 81 to 83, a plurality of speakers SP81 to SP83, a plurality of transmission microphones MICn81 to MICn83, a plurality of cancellation microphones MICw81 to MICw83, and a server 80. These transmission microphones correspond to the "transmission sound collection device" of the present invention, and the cancellation microphones correspond to the "cancellation sound collection device" of the present invention. Also, these speakers correspond to the "sound playback device" of the present invention. Further, the server 80 corresponds to the "transmission signal generation unit" of the present invention.

[0094] The plurality of conference devices 81 to 83 and the server 80 are connected to a communication network 800. Thereby, the plurality of conference devices 81 to 83 and the server 80 perform data communication with each other through the network 800.

[0095] (Configuration in Conference Room ROOMa) The conference devices 81 and 82 are arranged in the conference room ROOMa. The speakers SP81 and SP82, the transmission microphones MICn81 and MICn82, and the plurality of cancellation microphones MICw81 and MICw82 are arranged in the conference room ROOMa. The speaker SP81, the transmission microphone MICn81, and the cancellation microphone MICw81 are connected to the conference device 81. The speaker SP82, the transmission microphone MICn82, and the cancellation microphone MICw82 are connected to the conference device 82.

[0096] Speaker SP81, transmission microphone MICn81, and cancellation microphone MICw81 are arranged near speaker 90a. Conference device 81 includes a transmission volume adjustment unit 812 and an IF819. The transmission volume adjustment unit 812 and the IF819 are connected to each other and are configured by an information processing device in the same manner as in the above-described embodiments. The IF819 is connected to the server 80 through the network 800. The cancellation microphone MICw81 and the speaker SP81 are connected to the transmission volume adjustment unit 812. The transmission microphone MICn81 is connected to the IF819. Speaker 90a conducts a meeting using the speaker SP81, the transmission microphone MICn81, the cancellation microphone MICw81, and the conference device 81.

[0097] Speaker SP82, transmission microphone MICn82, and cancellation microphone MICw82 are arranged near speaker 90b. Conference device 82 includes a transmission volume adjustment unit 822 and an IF829. The transmission volume adjustment unit 822 and the IF829 are connected to each other and are configured by an information processing device in the same manner as in the above-described embodiments. The IF829 is connected to the server 80 through the network 800. The cancellation microphone MICw82 and the speaker SP82 are connected to the transmission volume adjustment unit 822. The transmission microphone MICn8 2 is , is connected to the IF829. Speaker 90b conducts a meeting using the speaker SP82, the transmission microphone MICn82, the cancellation microphone MICw82, and the conference device 82.

[0098] (Configuration in Conference Room ROOMb) Conference device 83 is arranged in conference room ROOMb. Speaker SP83, transmission microphone MICn83, and cancellation microphone MICw83 are arranged in conference room ROOMb. The speaker SP83, the transmission microphone MICn83, and the cancellation microphone MICw83 are connected to the conference device 83.

[0099] Speaker SP83, transmission microphone MICn83, and cancellation microphone MICw83 are arranged near the speaker 90c. The conference device 83 includes a transmission volume adjustment unit 832 and an IF839. The transmission volume adjustment unit 832 and the IF839 are connected to each other and are configured by an information processing device in the same manner as in each of the above embodiments. The IF839 is connected to the server 80 through the network 800. The cancellation microphone MICw83 and the speaker SP83 are connected to the transmission volume adjustment unit 832. 3 is The transmission microphone MICn8 is connected to the IF839. The speaker 90c conducts a conference using the speaker SP83, the transmission microphone MICn83, the cancellation microphone MICw83, and the conference device 83.

[0100] (Features of the transmission microphone and the cancellation microphone) The sound pickup directivity of the plurality of transmission microphones MICn81, MICn82, MICn83 is narrow. Accordingly, the transmission microphone MICn81 picks up the voice of the speaker 90a and hardly picks up other voices. The transmission microphone MICn82 picks up the voice of the speaker 90b and hardly picks up other voices. The transmission microphone MICn83 picks up the voice of the speaker 90c and hardly picks up other voices.

[0101] The sound pickup directivity of the plurality of cancellation microphones MICw81 - MICw83 is wider than the sound pickup directivity of the plurality of transmission microphones MICn81 - MICn83. Accordingly, the cancellation microphone MICw81 picks up not only the voice of the speaker 90a but also the voice of the speaker 90b. The cancellation microphone MICw82 picks up not only the voice of the speaker 90b but also the voice of the speaker 90a. In other words, the cancellation microphones MICw81, MICw82 pick up the sound of the positions where the speakers 90a and 90b are located and the surrounding space. Similarly, the cancellation microphone MICw83 picks up the sound of the position where the speaker 90c is located in the conference room ROOMb and the surrounding space.

[0102] (Specific content of sound signal processing) The above configuration toTherefore, the audio signal processing system 10C executes the processing of the audio signal in the meeting as follows.

[0103] (Microphone pickup for speaker 90a) The transmission microphone MICn81 picks up the voice of speaker 90a and generates a pickup signal Sa. The transmission microphone MICn81 outputs the pickup signal Sa to the IF819 of the conference device 81. The IF819 transmits the pickup signal Sa to the server 80 through the network 800. The cancellation microphone MICw81 picks up the voices in the positions where speakers 90a and 90b are located in the conference room ROOMa and the surrounding space, and generates a cancellation pickup signal Sxa. Since the cancellation microphone MICw81 is omnidirectional, the cancellation pickup signal Sxa is substantially a pickup signal obtained by adding the pickup signal Sa of the voice of speaker 90a and the pickup signal Sb of the voice of speaker 90b (Sxa = Sa + Sb).

[0104] (Microphone pickup for speaker 90b) The transmission microphone MICn82 picks up the voice of speaker 90b and generates a pickup signal Sb. The transmission microphone MICn82 outputs the pickup signal Sb to the IF829 of the conference device 82. The IF829 transmits the pickup signal Sb to the server 80 through the network 800. The cancellation microphone MICw82 picks up the voices in the positions where speakers 90a and 90b are located in the conference room ROOMa and the surrounding space, and generates a cancellation pickup signal Sxb. Since the cancellation microphone MICw82 is omnidirectional, the cancellation pickup signal Sxb is substantially a pickup signal obtained by adding the pickup signal Sb of the voice of speaker 90b and the pickup signal Sa of the voice of speaker 90a (Sxb = Sb + Sa).

[0105] (Microphone pickup for speaker 90c) The transmission microphone MICn83 picks up the voice of speaker 90c and generates a pickup signal Sc. The transmission microphone MICn83 outputs the pickup signal Sc to the IF839 of the conference device 83. The IF839 transmits the pickup signal Sc to the server 80 through the network 800. The cancellation microphone MICw83 is for speakers 90 in the conference room ROOMb c isThe voice of the position where the user is located and the surrounding space is picked up to generate a cancellation pickup signal Sxc. Although the cancellation microphone MICw83 has an omnidirectional characteristic, since there is only the speaker 90c in the conference room ROOMb, the cancellation pickup signal Sxc substantially becomes the pickup signal Sc of the voice of the speaker 90c (Sxc = Sc).

[0106] (Processing at the server) The server 80 adds the pickup signal Sa, the pickup signal Sb, and the pickup signal Sc to generate an audio addition signal Sall (= Sa + Sb + Sc). The server 80 transmits the audio addition signal Sall to a plurality of conference devices 81-83 through the network 800. This audio addition signal Sall corresponds to the "transmission signal" of the present invention.

[0107] (Sound emission for the speaker 90a) The IF819 of the conference device 81 receives the audio addition signal Sall and outputs it to the feed amount adjustment unit 812.

[0108] As shown in FIG. 15, the feed amount adjustment unit 812 includes a delay processing unit 8121 and a subtractor 8122. The delay processing unit 8121 performs a delay process with a delay amount Δt on the cancellation pickup signal Sxa. The delay amount Δt in the conference device 81 is set by, for example, the transmission time of the pickup signal Sa from the conference device 81 to the server 80, the transmission time of the audio addition signal Sall from the server 80 to the conference device 81, and the generation time of the audio addition signal Sall in the server 80 added together.

[0109] The delay processing unit 8121 outputs the delayed cancellation pickup signal Δt(Sxa) to the subtractor 8122. Note that Δt(Sxa) shown in FIG. 15 is the signal Sxa(t) on the time axis with a delay amount Δ tIt means the signal Sxa(t + Δt) that has undergone delay processing. The subtractor 8122 receives the voice addition signal Sall from the IF819. When the voice addition signal Sall is input to the subtractor 8122, it is delayed with respect to the pickup signal Sa and the pickup signal for cancellation Sxa. The delay amount of this voice addition signal Sall is the same as the above-mentioned delay amount Δt, and the subtractor 8122 receives the voice addition signal Δt(Sall) with delay. Note that Δt(Sall) shown in FIG. 15 means the signal Sall(t + Δt) on the time axis. The subtractor 8122 subtracts the delayed pickup signal for cancellation Δt(Sxa) from the voice addition signal Δt(Sall) with delay. The voice addition signal Sall is the addition signal of the pickup signal Sa, the pickup signal Sb, and the pickup signal Sc, and the pickup signal for cancellation Sxa is the addition signal of the pickup signal Sa and the pickup signal Sb. Δt(Sxa)=Δt(Sa + Sb) shown in FIG. 15 means the signal Sxa(t + Δt)=Sa(t + Δt)+Sb(t + Δt) on the time axis. Δt(Sall)=Δt(Sa + Sb + Sc) shown in FIG. 15 means the signal Sall(t + Δt)=Sa(t + Δt)+Sb(t + Δt)+Sc(t + Δt) on the time axis.

[0110] Therefore, the output signal from the subtractor 8122 becomes a signal in which the pickup signal Sa and the pickup signal Sb are suppressed, and becomes the playback signal Δt(Sc). Δ t (Sc) means the signal Sc(t + Δt) on the time axis. The subtractor 8122 transmits (outputs) the playback signal Δt(Sc) to the speaker SP81. As a result, the playback signal Δt(Sc) transmitted from the feed amount adjustment unit 812 to the speaker SP81 is a signal adjusted so that the feed amounts of the pickup signal Sa and the pickup signal Sb are minimized.

[0111] The speaker SP81 plays the playback signal Δt(Sc), and the speaker 90a listens to this sound. As a result, the speaker 90a can listen to only the voice of the speaker 90c in another conference room ROOMb from the speaker SP81.

[0112] (Playback for speaker 90b) The IF 829 of the conference device 82 receives the voice addition signal Sall and outputs it to the feed amount adjustment unit 822. The feed amount adjustment unit 822 has the same configuration as the feed amount adjustment unit 812. As a result, the speaker 90b can only hear the voice of the speaker 90c in another conference room ROOMb from the speaker SP82.

[0113] (Effect on speakers in the same room in the audio signal processing system 10C) Here, since the speakers 90a and 90b are sitting in the conference room ROOMa, they directly hear each other's voices. For example, as shown in FIG. 14, the speaker 90a directly hears the voice Sbd of the speaker 90b. Similarly, the speaker 90b directly hears the voice Sad of the speaker 90a. On the other hand, due to the above-described processing, the speakers 90a and 90b sitting in the conference room ROOMa do not hear each other's voices from the speakers SP81 and SP82, respectively.

[0114] Therefore, the situation where the speakers 90a and 90b hear each other's voices with a double time difference is suppressed. As a result, the audio signal processing system 10C can provide a conference environment without discomfort for multiple speakers in the same room during a web conference.

[0115] Also, similar to the audio signal processing system 10, the audio signal processing system 10C does not emit the sound pickup signal of this microphone from the speaker closest to the microphone in the same room, so howling can be suppressed. As a result, in the conference room ROOMa where there are a plurality of speakers 90a and 90b and a plurality of microphones MICn81, MICn82 and a plurality of speakers SP81, SP82 are arranged, the audio signal processing system 10C can suppress howling due to the voice of the speaker 90a and howling due to the voice of the speaker 90b.

[0116] At this time, since the feed amount adjustment unit is composed of a delay processing unit and a subtractor, howling can be suppressed with a simple configuration and processing. Therefore, similar to the audio signal processing system 10, the audio signal processing system 10C does not require complex signal processing, can suppress howling, and can appropriately output voices other than the howling suppression target.

[0117] (Sound emission to speaker 90c) The IF 839 of the conference device 83 receives the voice addition signal Sall and outputs it to the feed volume adjustment unit 832. The feed volume adjustment unit 832 has the same configuration as the feed volume adjustment unit 812.

[0118] Therefore, the output signal from the feed volume adjustment unit 832 is a signal obtained by subtracting the delayed cancellation pickup signal Sxc from the voice addition signal Δt(Sall) with a delay. As a result, the output signal from the feed volume adjustment unit 832 becomes a signal in which the pickup signal Sc is suppressed, and becomes the sound emission signal Δt(Sa + Sb). The feed volume adjustment 832 unit transmits (outputs) the sound emission signal Δt(Sa + Sb) to the speaker SP83. Thus, the sound emission signal Δt(Sa + Sb) transmitted from the feed volume adjustment unit 832 to the speaker SP83 is a signal adjusted so that the feed volume of the pickup signal Sc is minimized.

[0119] The speaker SP83 emits the sound emission signal Δt(Sa + Sb), and the speaker 90c listens to this sound. Thus, the speaker 90c can listen to the voices of the speakers 90a and 90b in another conference room ROOMa from the speaker SP83.

[0120] Also, similar to the conference room ROOMa, in a conference room ROOMb where there are a plurality of speakers 90c and the microphone MICn83 and the speaker SP83 are arranged, the sound signal processing system 10C can suppress howling caused by the voices of the speakers 90c. And the sound signal processing system 10C does not require complicated signal processing even in the conference room ROOMb, can suppress howling, and can appropriately output voices other than the howling suppression target.

[0121] In addition, in the sound signal processing systems of the above-described embodiments, the mode of setting the second speaker was shown. However, the sound signal processing system can also set only the first speaker and omit the setting of the second speaker. Further, in the sound signal processing systems of the above-described embodiments, the number of second speakers to be set is not limited to two, and may be one or three or more depending on the arrangement of a plurality of speakers and the like. Further, the sound signal processing system may adjust the amount to be subtracted from the reference feed amount for each of the plurality of speakers according to the distance from the microphone that has picked up the voice of the speaker.

[0122] In addition, in the above-described embodiments, the case where the microphone is of the installed type was shown. In other words, in the above-described embodiments, the case where the microphone does not move was shown. However, even if the microphone (the microphone for generating the sound collection signal) moves with the speaker like a lapel microphone, the above-described configuration can be applied. In this case, the sound signal processing system may have, for example, the following configuration.

[0123] The sound signal processing system includes a configuration for detecting the position of the speaker (the microphone for generating the sound collection signal). Further, the sound signal processing system stores the positions of a plurality of speakers. The sound signal processing system calculates the distances between the microphone for generating the sound collection signal and the plurality of speakers from the detected position of the speaker (the microphone for generating the sound collection signal). The sound signal processing system determines the speaker closest to the microphone for generating the sound collection signal as the first speaker, and the next closest speaker as the second speaker. Hereinafter, the sound signal processing system adjusts the feed amount as described above. Thereby, the sound signal processing system can suppress howling and appropriately output voices other than the howling suppression target without requiring complicated signal processing.

[0124] In addition, in each of the above-described embodiments, the case where a plurality of microphones and a plurality of speakers are separately arranged has been shown. However, each of the plurality of microphones and the speaker closest to each microphone (first speaker) may be integrally arranged. For example, each of the plurality of microphones and the speaker closest to each microphone (first speaker) may be housed in one housing. Thereby, the positional relationship between each of the plurality of microphones and the speaker closest to each microphone (first speaker) is fixed. Therefore, the sound signal processing system can more reliably suppress howling.

[0125] The description of this embodiment is illustrative in all respects and not restrictive. The scope of the present invention is indicated not by the above-described embodiments but by the claims. Further, the scope of the present invention is intended to include all modifications within the meaning and scope equivalent to the claims.

Explanation of Reference Numerals

[0126] 10, 10A, 10B, 10C: Sound signal processing system 11: Speaker voice detection unit 12, 12A: Feed amount adjustment unit 80: Server 81, 82, 83: Conference device 111: Signal level detection unit 112: Specific microphone detection unit 120: Speaker DB 121: Speaker determination unit 122: Feed amount setting unit 123, 123A: Mixer 800: Network 812, 822, 832: Feed amount adjustment unit 819, 829, 839: IF 8121: Delay processing unit 8122: Subtractor

Claims

1. A sound signal processing method used in a sound signal processing system including a plurality of sound collection devices and a plurality of corresponding sound playback devices, wherein a feed amount adjustment process is performed to set the level of a sound signal collected by the plurality of sound collection devices to a preset reference feed amount and transmit the sound signal to the plurality of sound playback devices, and in the feed amount adjustment process, the feed amount transmitted from a first sound collection device that has detected a first speaker's voice among the plurality of sound collection devices to a first sound playback device closest to the first sound collection device is adjusted to a first feed amount smaller than the reference feed amount, and the sound signal is transmitted from the first sound collection device to the plurality of sound playback devices, and the feed amount transmitted from a second sound collection device that has detected a second speaker's voice among the plurality of sound collection devices to a second sound playback device closest to the second sound collection device is adjusted to a second feed amount smaller than the reference feed amount. A sound signal processing method.

2. In the feed amount adjustment process, the third feed amount of the sound signal transmitted to a third sound playback device that is closer to the first sound collection device than the first sound playback device is decreased, and the first feed amount is smaller than the third feed amount. The sound signal processing method according to claim 1.

3. The method includes a process of adding ambient sound to each of the plurality of sound signals collected by the plurality of sound collection devices. The sound signal processing method according to claim 1 or claim 2.

4. In the process of adding ambient sound, different ambient sounds are added to the respective sound signals collected by the plurality of sound collection devices. The sound signal processing method according to claim 3.

5. The plurality of sound collection devices and the plurality of sound playback devices each have an integrated shape including a pair of a sound collection device and a sound playback device. The sound signal processing method according to claim 1 or claim 2.

6. The first speaker's voice or the second speaker's voice is detected when a sound signal with a level equal to or higher than a predetermined level is input. The sound signal processing method according to claim 1 or claim 2.

7. A sound signal processing system including a plurality of sound collection devices and a plurality of corresponding sound playback devices, comprising a feed amount adjustment unit configured to set the level of a sound signal collected by the plurality of sound collection devices to a preset reference feed amount and transmit the sound signal to the plurality of sound playback devices, wherein the feed amount adjustment unit adjusts the feed amount transmitted from a first sound collection device that has detected a first speaker's voice among the plurality of sound collection devices to a first sound playback device closest to the first sound collection device to a first feed amount smaller than the reference feed amount, and transmits the sound signal from the first sound collection device to the plurality of sound playback devices, Adjust the transmission amount, which is transmitted from the second sound collection device that has detected the second speaker's voice among the plurality of sound collection devices, to a second transmission amount that is smaller than the reference transmission amount, to the second sound playback device closest to the second sound collection device. Audio signal processing system. **Claim 8** The transmission amount adjustment unit reduces the third transmission amount of the audio signal transmitted to the third sound playback device that is closer to the first sound collection device than the first sound playback device. The first transmission amount is smaller than the third transmission amount. The audio signal processing system according to claim 7. **Claim 9** The transmission amount adjustment unit adds ambient sound to each of the plurality of audio signals collected by the plurality of sound collection devices. The audio signal processing system according to claim 7 or claim 8. **Claim 10** The transmission amount adjustment unit adds different ambient sounds to each of the audio signals collected by the plurality of sound collection devices. The audio signal processing system according to claim 9. **Claim 11** The plurality of sound collection devices and the plurality of sound playback devices are each in an integrated shape including a pair of sound collection devices and sound playback devices. The audio signal processing system according to claim 7 or claim 8. **Claim 12** It includes a speaker voice detection unit that detects the first speaker's voice or the second speaker's voice when an audio signal above a predetermined level is input. The transmission amount adjustment unit performs a transmission amount adjustment process on the first speaker's voice or the second speaker's voice detected by the speaker voice detection unit. The audio signal processing system according to claim 7 or claim 8.

Citation Information

Patent Citations

  • Presentation system

    JP1992312098A

  • Multi-input echo canceller

    JP1994343196A

  • Loudspeaker system

    JP2006238254A

  • Filter setting device, filter setting method, and filter setting program

    JP2018142886A