Method for processing sound signal and system for processing sound signal

The sound signal processing system addresses the issue of unnecessary speaker output suppression and complex processing by using a system with specialized microphones and speakers to adjust feed amounts, effectively suppressing feedback and ensuring clear audio output for multiple speakers.

JP2025123436AInactive Publication Date: 2025-08-22YAMAHA CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025102993
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-16
Filing Date
2025-06-19
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing sound signal processing systems, such as those described in Patent Documents 1 and 2, face issues with unnecessary suppression of speaker output levels when multiple speakers speak simultaneously and require complex signal processing for howling suppression.

Method used

A sound signal processing method using a system with multiple sound collection devices, including a transmission sound collection device and a cancellation sound collection device with wider directivity, where collected signals are processed to adjust feed amounts to specific microphones and speakers, suppressing feedback without complex processing.

Benefits of technology

The method effectively suppresses feedback while allowing clear audio output from multiple speakers, ensuring each speaker can hear others without unnecessary volume suppression, and does not require complex signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025123436000001_ABST
    Figure 2025123436000001_ABST
Patent Text Reader

Abstract

To suppress howling without requiring complicated signal processing and to appropriately output sound other than howling suppression targets.SOLUTION: The method for processing a sound signal according to the present invention is a method for processing a sound signal used in a sound signal processing system including a plurality of sound collection devices and a plurality of sound emission devices. The plurality of sound collection devices include a transmission sound collection device and a cancellation sound collection device with wider directivity than the transmission sound collection device. A sound signal collected by the transmission sound collection device is transmitted to the sound emission devices, a cancellation sound collection signal collected by the cancellation sound collection device is cancelled from sound collection signals received by each sound emission device, and a sound signal after the cancellation is output to each sound emission device.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to a sound signal processing method and a sound signal processing device that are performed when sound picked up by a microphone is output from a speaker. [Background technology]

[0002] Patent Documents 1 and 2 describe devices for suppressing howling.

[0003] Patent Document 1 describes a public address system equipped with multiple microphones and multiple speakers. The public address system of Patent Document 1 detects a speaker based on input signals from the multiple microphones and selects a microphone corresponding to the speaker's position. The public address system of Patent Document 1 lowers the output level of a speaker located near the selected microphone.

[0004] Patent Document 2 describes a filter setting device that includes multiple microphones, multiple speakers, a mixer, an input filter, and an output filter. The filter setting device described in Patent Document 2 adjusts the frequency characteristics of the output filter so as to change the loop gain of the integrated system for sound signals that are mixed by the mixer and output to the multiple speakers. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-238254 [Patent Document 2] Japanese Patent Application Publication No. 2018-142886 Summary of the Invention [Problem to be solved by the invention]

[0006] However, with the configuration described in Patent Document 1, for example, when multiple speakers in different positions speak simultaneously, the level of the speaker's voice output from the speaker to be suppressed drops along with the level of the other speakers' voices output from the same speaker. In other words, the loudspeaker system of Patent Document 1 unnecessarily lowers the speaker output level.

[0007] Furthermore, the configuration of Patent Document 2 requires complex signal processing.

[0008] In consideration of the above circumstances, one aspect of the present disclosure aims to provide a sound signal processing method that does not require complex signal processing, suppresses howling, and can appropriately output audio other than the audio targeted for howling suppression. [Means for solving the problem]

[0009] The sound signal processing method is a sound signal processing method used in a sound signal processing system including a plurality of sound collection devices and a plurality of sound emitting devices, wherein the plurality of sound collection devices include a transmission sound collection device and a cancellation sound collection device having a wider directivity than the transmission sound collection device, and the sound collection signals collected by the transmission sound collection device are transmitted to the plurality of sound emitting devices, the cancellation sound collection signals collected by the cancellation sound collection device are cancelled from the collection signals received by each sound emitting device of the plurality of sound emitting devices, and the cancelled sound signals are output to each sound emitting device. [Effects of the Invention]

[0010] The sound signal processing method does not require complex signal processing, and can suppress feedback and appropriately output sounds other than those targeted by the feedback suppression. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram showing an example of the arrangement of a plurality of microphones and a plurality of speakers in a space where an audio conference is held. [Figure 2] FIG. 2 is a functional block diagram showing an example of the configuration of a sound signal processing system according to the first embodiment of the present invention. [Figure 3] FIG. 3 is a flowchart showing an example of a sound signal processing method according to the first embodiment of the present invention. [Figure 4] FIG. 4 is a functional block diagram showing an example of the configuration of the speaker voice detection unit. [Figure 5] FIG. 5 is a functional block diagram showing an example of the configuration of the feed amount adjustment unit. [Figure 6] FIG. 6 is a table showing the relationship between microphones and speakers stored in the speaker DB. [Figure 7] FIG. 7 is a diagram showing a concept of adjusting the feed amount according to an embodiment of the present invention. [Figure 8] FIG. 8 is a diagram showing a concept of adjusting the feed amount according to an embodiment of the present invention. [Figure 9] 9(A) and 9(B) are diagrams showing an example of the volume of sounds output from a plurality of speakers. [Figure 10] FIG. 10 is a functional block diagram showing an example of the configuration of a sound signal processing system according to the second embodiment of the present invention. [Figure 11] FIG. 11 is a functional block diagram showing an example of the configuration of a feed amount adjustment unit according to the second embodiment. [Figure 12] FIG. 12 is a flowchart showing an example of a sound signal processing method according to the second embodiment of the present invention. [Figure 13] FIG. 13 is a functional block diagram showing an example of the configuration of a sound signal processing system according to the third embodiment of the present invention. [Figure 14] FIG. 14 is a functional block diagram showing an example of the configuration of a sound signal processing system according to the fourth embodiment of the present invention. [Figure 15] FIG. 15 is a functional block diagram showing an example of the configuration of a conference device used in the sound signal processing system according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0012] A sound signal processing method and a sound signal processing system according to an embodiment of the present invention will be described with reference to the drawings.

[0013] [Embodiment 1] [Audio conference environment (position of microphone, speaker, and speaker)] Fig. 1 is a diagram showing an example of the arrangement of multiple microphones and multiple speakers in a space where an audio conference is held. As shown in Fig. 1, multiple microphones MIC11-MIC25 (MIC11, MIC12, MIC13, MIC14, MIC15, MIC21, MIC22, MIC23, MIC24, MIC25) and multiple speakers SP11-SP43 (SP11, SP12, SP13, SP14, SP15, SP21, SP22, SP23, SP24, SP25, SP31, SP32, SP33, SP41, SP42, SP43) are arranged in the space where the audio conference is held (e.g., a conference room). These microphones correspond to the "sound collection device" of the present invention, and these speakers correspond to the "sound emission device" of the present invention.

[0014] The multiple microphones MIC11-MIC15 are arranged in a row near a first side of the table TBL and along the first side. The multiple microphones MIC21-MIC25 are arranged in a row near a second side (the side opposite the first side) of the table TBL and along the second side. The row of the multiple microphones MIC11-MIC15 and the row of the multiple microphones MIC21-MIC25 are separated by the table TBL.

[0015] The multiple speakers SP11-SP15 are arranged in a row near a first side of the table TBL and along the first side. The multiple speakers SP21-SP25 are arranged in a row near a second side of the table TBL and along the second side. The speaker SP11 is the speaker closest to the microphone MIC11. Similarly, as shown in FIG. 1, the speakers SP12-SP15 and SP21-SP25 are the speakers closest to the microphones MIC12-MIC15 and MIC21-MIC25, respectively. As shown in FIG. 1, the multiple speakers SP31-SP33 are arranged in a row near a third side of the table TBL (a side perpendicular to the first and second sides) and along the third side. The multiple speakers SP41-SP43 are arranged in a row near a fourth side of the table TBL (a side opposite the third side) and along the fourth side. The multiple speakers SP11-SP43 each emit sound mainly in the direction of the speaker closest to them. Although the speakers SP11-SP43 are not arranged on the table TBL in FIG. 1, they may be arranged on the table TBL.

[0016] The arrangement of these multiple microphones and multiple speakers is not limited to the example shown in Figure 1. The arrangement of the multiple microphones and multiple speakers may be any arrangement that allows the speaker closest to each microphone to be determined. Furthermore, the number of multiple microphones and the number of multiple speakers are not limited to the above numbers and can be set depending on the number of speakers in the conference, etc.

[0017] Multiple speakers 911-915, 921-925 are respectively near multiple microphones MIC11-MIC15, MIC21-MIC25 and hold a conference (conversation). More specifically, for example, speaker 911 is near microphone MIC11 and holds a conference using microphone MIC11. In this case, it is preferable that the directivity, etc. of the multiple microphones MIC11-MIC25 is set so that they pick up the voices of nearby speakers and hardly pick up the voices of other speakers.

[0018] [Configuration and processing of sound signal processing device] Fig. 2 is a functional block diagram showing an example of the configuration of a sound signal processing system according to the first embodiment of the present invention, and Fig. 3 is a flowchart showing an example of a sound signal processing method according to the first embodiment of the present invention.

[0019] 2, the sound signal processing system 10 includes a speaker voice detection unit 11, a feed amount adjustment unit 12, multiple microphones MIC11-MIC25, multiple speakers SP11-SP43, and multiple output amplifiers A11-A43. Physically, the sound signal processing system 10 includes a main unit, multiple microphone housings, and multiple speaker housings.

[0020] The main device is an information processing device equipped with a DSP, a CPU, etc., a storage medium, and predetermined electronic circuits. The main device realizes the functions of the speaker voice detection unit 11 and the feed amount adjustment unit 12 by executing a sound signal processing program stored in the storage medium using the DSP, CPU, etc. The main device also realizes multiple output amplifiers A11-A43 using multiple amplification devices.

[0021] The multiple microphones MIC11-MIC25 each have their own microphone housing separate from the main device. As shown in Figure 1, the multiple microphones MIC11-MIC25 are placed in predetermined positions in the space where the conference will be held. The multiple microphones MIC11-MIC25 are physically and electrically connected to the main device.

[0022] The multiple speakers SP11-SP43 each have their own speaker housing separate from the main device. As shown in Figure 1, the multiple speakers SP11-SP43 are arranged at predetermined positions in the space where the conference will be held. The multiple speakers SP11-SP43 are physically and electrically connected to the main device.

[0023] More specifically, the sound signal processing system 10 has the following circuit configuration. The multiple microphones MIC11-MIC25 are connected to the speaker voice detection unit 11 and the feed amount adjustment unit 12. The multiple microphones MIC11-MIC25 acquire picked-up signals and output the picked-up signals Sm11-Sm25 to the speaker voice detection unit 11 and the feed amount adjustment unit 12 (S11). Specifically, the microphone MIC11 outputs the picked-up signal Sm11 to the speaker voice detection unit 11 and the feed amount adjustment unit 12. Similarly, the multiple microphones MIC12-MIC15 and MIC21-MIC25 output the picked-up signals Sm12-Sm15 and Sm21-Sm25 to the speaker voice detection unit 11 and the feed amount adjustment unit 12.

[0024] [Speaker voice detection] The talker voice detection unit 11 detects the talker voice using the plurality of collected sound signals Sm11-Sm25 (S12). That is, the talker voice detection unit 11 detects the microphone that collects the talker's voice during pronunciation (utterance). This microphone corresponds to the "first sound collection device" of the present invention. This microphone is the specific microphone for which the feed amount is to be adjusted, and is hereinafter referred to as the "specific microphone."

[0025] 4 is a functional block diagram showing an example of the configuration of the speaker voice detection unit 11. As shown in FIG.

[0026] The signal level detection unit 111 receives input of pickup signals Sm11-Sm25 from multiple microphones MIC11-MIC25. The signal level detection unit 111 detects the signal levels (amplitudes) of the multiple pickup signals Sm11-Sm25. The signal level detection unit 111 outputs the signal levels (amplitudes) of the multiple pickup signals Sm11-Sm25 to the specific microphone detection unit 112.

[0027] The specific microphone detection unit 112 stores in advance a threshold value for speaker detection (hereinafter simply referred to as the threshold value). The threshold value is set to, for example, a value higher than the noise level in the space where the conference is held and lower than the level of the picked-up signal when speech at the conference is picked up by the microphone.

[0028] The specific microphone detection unit 112 compares the signal levels of the multiple picked-up signals Sm11-Sm25 with a threshold. The specific microphone detection unit 112 detects picked-up signals that are equal to or greater than the threshold. In other words, the specific microphone detection unit 112 detects a microphone that outputs a picked-up signal that is equal to or greater than the threshold as a specific microphone (S13). The specific microphone detection unit 112 outputs information about the specific microphone to the transmission amount adjustment unit 12.

[0029] The send amount adjustment unit 12 generally adjusts the send amounts of the plurality of picked-up signals Sm11-Sm25 to the plurality of speakers SP11-SP43 based on the specific microphone.

[0030] Fig. 5 is a functional block diagram showing an example of the configuration of the feed amount adjustment unit 12. As shown in Fig. 5, the feed amount adjustment unit 12 includes a speaker determination unit 121, a feed amount setting unit 122, a mixer 123, and a speaker DB 120. The feed amount setting unit 122 includes a suppression amount setting unit 1220.

[0031] The speaker DB 120 stores the relationship between the specific microphone and the speaker whose feed amount is to be adjusted. FIG. 6 is a table showing the relationship between the specific microphone and the speaker stored in the speaker DB 120. As shown in FIG. 6, a first speaker and a second speaker are associated with each specific microphone. The first speaker is the speaker closest to the specific microphone. The second speaker is the speaker next closest to the specific microphone after the first speaker. This first speaker corresponds to the "first sound emitting device" of the present invention, and the second speaker corresponds to the "second sound emitting device" of the present invention.

[0032] The speaker determination unit 121 refers to the speaker DB 120 based on the information about the specific microphone and determines the speaker for which the feed amount is to be adjusted (S14). Specifically, the speaker determination unit 121 refers to the speaker DB 120 based on the information about the specific microphone and determines a first speaker and a second speaker for the specific microphone. For example, if the specific microphone is microphone MIC13, the speaker determination unit 121 determines speaker SP13 as the first speaker and speakers SP12 and SP14 as the second speakers. The speaker determination unit 121 outputs the information about the first speaker and the information about the second speaker to the feed amount setting unit 122.

[0033] The feed amount setting unit 122 sets the feed amounts of the signals picked up by the microphones MIC11-MIC25 for the speakers SP11-SP43 based on the combination of the specific microphone, the first speaker, and the second speaker (S15). The feed amount for the first speaker corresponds to the "first feed amount" of the present invention, and the feed amount for the second speaker corresponds to the "second feed amount" of the present invention.

[0034] If information indicating a combination of a specific microphone, a first speaker, and a second speaker is not input, the feed amount setting unit 122 sets a preset feed amount for all combinations of the multiple microphones MIC11-MIC25 and the multiple speakers SP11-SP43. The preset feed amount is basically the same for all combinations of the multiple microphones MIC11-MIC25 and the multiple speakers SP11-SP43, but it is not limited to a perfect match and may include errors or adjustment differences for each speaker. Such a feed amount is referred to as a reference feed amount. Note that the feed amount is the amount of adjustment for the signal level (amplitude) of the sound signal sent from the microphone to the speaker. The feed amount can be set, for example, by the gain for the picked-up signal.

[0035] When information indicating a combination of a specific microphone, a first speaker, and a second speaker is input, the feed amount setting unit 122 adjusts the feed amount from the specific microphone to the first speaker and the second speaker to be smaller than the reference feed amount. At this time, the feed amount setting unit 122 uses the suppression amount set by the suppression amount setting unit 1220, the details of which will be described later. More specifically, the feed amount setting unit 122 adjusts the feed amount from the specific microphone to the first speaker to a minimum value smaller than the reference feed amount. The feed amount from the specific microphone to the first speaker may be zero. In other words, it is also possible not to send a picked-up signal from the specific microphone to the first speaker. Furthermore, the feed amount setting unit 122 adjusts the feed amount from the specific microphone to the second speaker to be smaller than the reference feed amount but larger than the feed amount to the first speaker. The feed amount setting unit 122 then sets the feed amounts other than the feed amounts from the specific microphone to the first speaker and the second speaker to the reference feed amount.

[0036] The feed amount setting unit 122 outputs to the mixer 123 the feed amount for each combination of the plurality of microphones MIC11-MIC25 and the plurality of speakers SP11-SP43.

[0037] The mixer 123 is realized by a so-called matrix mixer. The mixer 123 adjusts the signal levels of the pickup signals Sm11-Sm25 of the multiple microphones MIC11-MIC25 using the feed amounts from the feed amount setting unit 122, and mixes the signals for each of the multiple speakers SP11-SP43 (S16).

[0038] For example, if none of the collected signals Sm11-Sm25 are detected as speaker voices, the mixer 123 adjusts the levels of all the collected signals Sm11-Sm25 by the reference feed amount, and mixes these level-adjusted collected signals for each of the multiple speakers SP11-SP43.

[0039] On the other hand, when a specific picked-up signal is detected as a speaker's voice, the mixer 123 adjusts the level of the picked-up signal sent from the specific microphone to the first and second speakers so that the amount is smaller than the reference amount. The mixer 123 adjusts the level of the sent signals other than the picked-up signal from the specific microphone to the first and second speakers by the reference amount. The mixer 123 mixes these level-adjusted picked-up signals Sm11-Sm25 for each of the multiple speakers SP11-SP43.

[0040] By performing such mixing processing, the mixer 123 generates speaker supply signals Ss11-Ss43 for the multiple speakers SP11-SP43, respectively.

[0041] This significantly reduces the level at which the signal picked up by the specific microphone is output from the first speaker. Also, the level at which the signal picked up by the specific microphone is output from the second speaker is reduced. Therefore, feedback via the specific microphone can be suppressed. Meanwhile, the level at which the signal picked up by the specific microphone is output from speakers other than the first and second speakers can be ensured to be audible to other speakers.

[0042] The feed amount adjuster 12 transmits (outputs) the speaker supply signals Ss11-Ss43 to the output amplifiers A11-A43, respectively. The output amplifiers A11-A43 amplify the speaker supply signals Ss11-Ss43 to generate speaker drive signals So11-So43. The output amplifiers A11-A43 transmit (output) the speaker drive signals So11-So43 to the multiple speakers SP11-SP43, respectively (S17).

[0043] The processing of the above-described forwarding amount adjustment unit 12 is processing as one means for realizing the following concept. Figs. 7 and 8 are diagrams showing the concept of forwarding amount adjustment according to an embodiment of the present invention. Fig. 7 shows the case where one speaker is speaking. Fig. 8 shows the case where two speakers are speaking simultaneously. The size of the black circles in Figs. 7 and 8 and the size of the white circles in Fig. 8 indicate the size of the forwarding amount.

[0044] (When there is only one speaker) As shown in FIG. 7, when only speaker 913 is speaking, among the plurality of picked-up signals Sm11-Sm25, the signal level of picked-up signal Sm13 is equal to or greater than the threshold and is the maximum value, and the signal levels of the other picked-up signals are below the threshold.

[0045] The feed amount adjustment unit 12 sets the feed amount to the reference feed amount for all speakers for the picked-up signal Sm13 resulting from the speech of the speaker 913. That is, as indicated by the black circles in Fig. 7, the feed amount adjustment unit 12 sets the same feed amount for all speakers SP11-SP43 for the picked-up signal Sm13.

[0046] The feed amount adjustment unit 12 further performs the following process. Based on the microphone MIC13 detected as the specific microphone, the feed amount adjustment unit 12 determines the speaker SP13 as the first speaker and the speakers SP12 and SP14 as the second speakers. The feed amount adjustment unit 12 sets the suppression amount for the picked-up signal Sm13 to the speakers SP13, SP12, and SP14, but does not set the suppression amount for the other speakers.

[0047] More specifically, the feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 (specific microphone) and the speaker SP13 (first speaker) to, for example, approximately the same as the reference feed amount, as shown in Fig. 7. The feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 (specific microphone) and the speakers SP12 and SP14 (second speakers) to a value that is smaller than the reference feed amount and is not 0, as shown in Fig. 7. Note that the suppression amount is set here as a reference for the concept of the present invention. However, the feed amount adjustment unit 12 may simplify the processing by simply reducing the feed amount for the combination of the specific microphone and the first and second speakers without setting a suppression amount.

[0048] At this time, it is more preferable that the feed amount adjustment unit 12 performs the following process. Specifically, the feed amount adjustment unit 12 sets the suppression amount based on the following concept. The feed amount adjustment unit 12 sets the gain corresponding to the reference feed amount as gain A, and the degree of suppression as suppression index α. Furthermore, the feed amount adjustment unit 12 sets the distance between the microphone MIC13 and the speaker SP13 as distance r1, and the distance between the microphone MIC13 and the speakers SP12 and SP14 as distance r2. Furthermore, the feed amount adjustment unit 12 sets the influence of the distance on the suppression amount as influence n.

[0049] In this case, the amount of suppression for the speaker SP13 is A / (r1 n ·α). The amount of suppression for the speakers SP12 and SP14 is set to A / (r2 n α). This allows the feed amount adjustment unit 12 to set the suppression amounts as described above. These suppression amounts are values ​​that correspond to the positional relationship between the multiple speakers and the multiple microphones.

[0050] The feed amount adjustment unit 12 sets the feed amounts for all speakers SP11-SP43 by subtracting the suppression amount from the reference feed amount. As a result, the feed amount adjustment unit 12 reduces the feed amount of the picked-up signal Sm13 for the speaker SP13 to less than the reference feed amount, adjusting it to the minimum value. The feed amount adjustment unit 12 also adjusts the feed amount of the picked-up signal Sm13 for the speakers SP12 and SP14 to less than the reference feed amount. Furthermore, the feed amount adjustment unit 12 sets the feed amount of the picked-up signal Sm13 for speakers other than the speakers SP12, SP13, and SP14 to the reference feed amount.

[0051] (When two speakers are speaking at the same time) As shown in FIG. 8, when speakers 913 and 924 are speaking simultaneously and no other speakers are speaking, the signal levels of picked-up signals Sm13 and Sm24 are equal to or greater than the threshold, and the signal levels of the other picked-up signals are less than the threshold.

[0052] The feed amount adjustment unit 12 sets the feed amounts for all speakers for the collected signal Sm13 and the collected signal Sm24 to the reference feed amount. That is, as indicated by the black circles in Fig. 8, the feed amount adjustment unit 12 sets the same feed amount for all speakers SP11-SP43 for the collected signal Sm13. Also, as indicated by the white circles in Fig. 8, the feed amount adjustment unit 12 sets the same feed amount for all speakers SP11-SP43 for the collected signal Sm24.

[0053] The feed amount adjustment unit 12 further performs the following process. Based on the microphone MIC13 detected as the specific microphone, the feed amount adjustment unit 12 determines the speaker SP13 as the first speaker and the speakers SP12 and SP14 as the second speakers. The feed amount adjustment unit 12 sets the amount of suppression of the picked-up signal Sm13 for the speakers SP13, SP12, and SP14. The feed amount adjustment unit 12 does not set the amount of suppression for the other speakers.

[0054] More specifically, the feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 and the speaker SP13 (first speaker) to, for example, approximately the same as the reference feed amount, as shown in Fig. 8. The feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC13 and the speakers SP12 and SP14 (second speakers) to a value that is smaller than the reference feed amount and is not 0, as shown in Fig. 8. Note that here too, the suppression amount is set as a reference for the concept of the present invention. However, as in the above case, the feed amount adjustment unit 12 may simplify the processing by simply reducing the feed amount without setting the suppression amount.

[0055] The feed amount adjustment unit 12 determines the speaker SP24 as the first speaker and the speakers SP23 and SP25 as the second speakers based on the microphone MIC24 detected as the specific microphone. The feed amount adjustment unit 12 sets the amount of suppression of the picked-up signal Sm24 to the speakers SP24, SP23, and SP25. The feed amount adjustment unit 12 does not set the amount of suppression for the other speakers.

[0056] More specifically, the feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC24 and the speaker SP24 (first speaker) to, for example, approximately the same as the reference feed amount, as shown in Fig. 8. The feed amount adjustment unit 12 sets the suppression amount for the combination of the microphone MIC24 and the speakers SP23 and SP25 (second speakers) to a value that is smaller than the reference feed amount and is not 0, as shown in Fig. 8. Note that here too, the suppression amount is set as a reference for the concept of the present invention. However, as in the above case, the feed amount adjustment unit 12 may simplify the processing by simply reducing the feed amount without setting the suppression amount.

[0057] The feed amount adjustment unit 12 sets the feed amounts for all speakers SP11-SP43 by subtracting the suppression amount from the reference feed amount. As a result, as shown in FIG. 8, the feed amount adjustment unit 12 adjusts the feed amount of the picked-up signal Sm13 for the speaker SP13 to be smaller than the reference feed amount, to the minimum value. The feed amount adjustment unit 12 also adjusts the feed amount of the picked-up signal Sm13 for the speakers SP12 and SP14 to be smaller than the reference feed amount. Furthermore, the feed amount adjustment unit 12 sets the feed amount of the picked-up signal Sm13 for speakers other than the speakers SP12, SP13, and SP14 to the reference feed amount. The feed amount adjustment unit 12 also adjusts the feed amount of the picked-up signal Sm24 for the speaker SP24 to be smaller than the reference feed amount, to the minimum value. The feed amount adjustment unit 12 also adjusts the feed amount of the picked-up signal Sm24 for the speakers SP23 and SP25 to be smaller than the reference feed amount. Furthermore, the feed amount adjustment unit 12 sets the feed amount of the picked-up signal Sm24 for the speakers other than the speakers SP23, SP24, and SP25 to the reference feed amount.

[0058] Therefore, in the speaker supply signal Ss13, the component of the picked-up signal Sm13 is significantly suppressed. On the other hand, in the speaker supply signal Ss13, the component of the picked-up signal Sm24 is not suppressed. Also, in the speaker supply signals Ss12 and Ss14, the component of the picked-up signal Sm13 is suppressed to some extent. On the other hand, in the speaker supply signals Ss12 and Ss14, the component of the picked-up signal Sm24 is not suppressed.

[0059] This suppresses feedback between the microphone MIC13 and the multiple speakers SP13, SP12, and SP14. Furthermore, speaker 913 can clearly hear what speaker 924 is saying even while he or she is speaking.

[0060] Similarly, in the speaker supply signal Ss24, the component of the picked-up signal Sm24 is significantly suppressed. On the other hand, in the speaker supply signal Ss24, the component of the picked-up signal Sm13 is not suppressed. Also, in the speaker supply signals Ss23 and Ss25 to the speakers SP23 and SP25, the component of the picked-up signal Sm24 is suppressed to some extent. On the other hand, in the speaker supply signals Ss23 and Ss25, the component of the picked-up signal Sm13 is not suppressed.

[0061] This suppresses feedback between the microphone MIC24 and the multiple speakers SP24, SP23, and SP25. Furthermore, speaker 924 can clearly hear what speaker 913 is saying even while he or she is speaking.

[0062] 7 and 8 are realized by a plurality of functional blocks, for example, the feed amount setting unit 122 shown in Fig. 5 includes a suppression amount setting unit 1220. The suppression amount setting unit 1220 sets the suppression amount as described above using the combination of the specific microphone, the first speaker, and the second speaker determined by the speaker determination unit 121. The feed amount setting unit 122 sets the feed amount for all combinations of microphone and speaker by subtracting the suppression amount set by the suppression amount setting unit 1220 from the reference feed amount.

[0063] With the above-described configuration and processing, the sound signal processing system 10 can output the speaker's speech from a plurality of speakers SP11-SP43, as shown in FIGS. 9(A) and 9(B).

[0064] 9(A) and 9(B) are diagrams showing an example of the volume of sound output from multiple speakers. In FIGS. 9(A) and 9(B), the arrangement of multiple microphones, multiple speakers, and multiple speakers is the same as in FIG. 1. In FIG. 9(A) and 9(B), the arrangement of multiple microphones, multiple speakers, and multiple speakers is the same as in FIG. 1. In FIG. 9(A) and 9(B), the hatched circles indicate the volume of each speaker when the voice of speaker 913 is emitted from multiple speakers SP11-SP43. The larger the radius of the circle, the higher the volume. In FIG. 9(B), the unfilled circles indicate the volume of each speaker when the voice of speaker 924 is emitted from multiple speakers SP11-SP43. The larger the radius of the circle, the higher the volume.

[0065] As shown in Figure 9(A), when speaker 913 is speaking, the volume of speaker SP13 is the lowest, followed by speakers SP12 and SP14. The volumes of the other speakers SP11, SP15, SP21-SP43 are set to predetermined values ​​that are higher than the volumes of speakers SP12-SP14. This reduces the volume of speaker 913's voice propagating from speaker SP13 to microphone MIC13, thereby suppressing feedback. Furthermore, the volume of speaker 913's voice propagating from speakers SP12 and SP14 to microphone MIC13 is reduced, thereby further suppressing feedback.

[0066] At this time, the volume of speakers SP12 and SP14 is greater than the volume of speaker SP13. Therefore, speakers 912 and 914 can hear the direct sound from speaker 913 and the voice of speaker 913 output from speakers SP12 and SP14. Therefore, according to this embodiment, it is possible to prevent the voice of speaker 913 from being difficult to hear.

[0067] As shown in Fig. 9(B), when speaker 913 and speaker 924 are speaking at the same time, the various processes related to the voice of speaker 913 are as described above, and feedback caused by the voice of speaker 913 is suppressed. On the other hand, as shown in Fig. 9(B), speaker SP24 near speaker 924, who is speaking at the same time as speaker 913, emits the voice of speaker 913 without suppressing it. This allows speaker 924 to hear the voice of speaker 913 from speaker SP24 even while speaker 924 is speaking.

[0068] Also, as in the case of speaker 913 described above, as shown by the open circles in FIG. 9(B), the volume of speaker SP24 is the lowest for the sound signal from microphone MIC24, followed by speakers SP23 and SP25. The volumes of the other speakers SP11-SP15, SP21, SP22, and SP31-SP43 are set to predetermined values ​​that are higher than the volumes of speakers SP23-SP25. This reduces the volume of the speaker 924's voice propagating from speaker SP24 to microphone MIC24, thereby suppressing feedback. Furthermore, the volume of the speaker 924's voice propagating from speakers SP23 and SP25 to microphone MIC24 is reduced, thereby further suppressing feedback.

[0069] 9(B), the speaker SP13 emits sound without suppressing the volume of the sound signal including the voice of the speaker 924. This allows the speaker 913 to hear the voice of the speaker 924 from the speaker SP13 even while the speaker 913 is speaking.

[0070] In this way, when multiple speakers are speaking simultaneously, the sound signal processing system 10 can suppress feedback for each speaker while allowing each speaker to clearly hear the voices of the other speakers. Note that although the above explanation has been given for a case where two people speak simultaneously, similar processing can also be applied to a case where three or more people speak simultaneously.

[0071] In the prior art, the above-described feedback is suppressed by suppressing the output to the multiple speakers SP11-SP43. That is, the configuration of the prior art adjusts the volume of the speaker drive signals So11-So43. In this case, when multiple speakers are speaking simultaneously, the sound output from the speaker near one speaker is a sound in which the voices of the other speakers are also suppressed. Therefore, it becomes difficult for a speaker to clearly hear the voices of the other speakers who are speaking simultaneously.

[0072] As described above, the sound signal processing system 10 can suppress feedback and appropriately output sounds other than those targeted for feedback suppression. Furthermore, the sound signal processing system 10 does not require complex filter coefficient settings or other complex signal processing for feedback suppression.

[0073] In the above description, the case where there is one speaker or one speaker in each of multiple rows (multiple speakers are far apart) has been described. However, when multiple speakers (simultaneous speakers) are close to each other, the sound signal processing system 10 can suppress feedback and appropriately output sounds other than those targeted for feedback suppression by using, for example, the following method.

[0074] For example, when multiple picked-up signals are equal to or greater than a threshold, the specific microphone detection unit 112 compares the waveforms of these multiple picked-up signals to detect whether the multiple picked-up signals represent the voices of the same speaker. If the multiple picked-up signals represent the voices of different speakers, the specific microphone detection unit 112 detects multiple specific microphones corresponding to each of the multiple picked-up signals. The feed amount adjustment unit 12 performs the above-mentioned feed amount adjustment process based on the multiple specific microphones. Note that, at this time, if the multiple picked-up signals represent the voices of the same speaker, the specific microphone detection unit 112 detects the microphone corresponding to the picked-up signal with the maximum value as the specific microphone. Then, the feed amount adjustment unit 12 performs the above-mentioned feed amount adjustment process based on this specific microphone.

[0075] [Embodiment 2] Fig. 10 is a functional block diagram showing an example of the configuration of a sound signal processing system according to a second embodiment of the present invention. Fig. 11 is a functional block diagram showing an example of the configuration of a feed amount adjustment unit according to the second embodiment. Fig. 12 is a flowchart showing an example of a sound signal processing method according to the second embodiment of the present invention.

[0076] The sound signal processing system 10A according to the second embodiment differs from the sound signal processing system 10 according to the first embodiment in that it includes a feed amount adjustment unit 12A instead of the feed amount adjustment unit 12 of the sound signal processing system 10. Other configurations of the sound signal processing system 10A are the same as those of the sound signal processing system 10, and descriptions of similar parts will be omitted.

[0077] The feed amount adjustment unit 12A differs from the feed amount adjustment unit 12 according to the first embodiment in that it includes a mixer 123A instead of the mixer 123 of the feed amount adjustment unit 12. Other configurations of the feed amount adjustment unit 12A are the same as those of the feed amount adjustment unit 12, and a description of similar parts will be omitted.

[0078] 11, the mixer 123A is realized by a matrix mixer, similar to the mixer 123 of the first embodiment. However, the mixer 123A includes a filter that adds a reflected sound to each combination of the microphones MIC11-MIC25 and the speakers SP11-SP43. The filter that adds the reflected sound is, for example, an FIR filter.

[0079] Specifically, a plurality of picked-up sound signals Sm11-Sm25 are input to the mixer 123A. The mixer 123A performs convolution calculation processing on the plurality of picked-up sound signals Sm11-Sm25 using filter coefficients corresponding to each of the plurality of speakers SP11-SP43. As a result, the mixer 123A adds a plurality of reflected sounds to each combination of the plurality of microphones MIC11-MIC25 and the plurality of speakers SP11-SP43 (S21). At this time, the mixer 123A sets the filter coefficients for the convolution calculation processing based on the relative positions of the plurality of microphones MIC11-MIC25 and the plurality of speakers SP11-SP43, and the environment of the conference room (such as the size of the conference room, the structure of the walls and ceiling, etc.).

[0080] The filter coefficients for the convolution calculation process are set, for example, as follows: Reflected sound is composed of an early reflected sound component and a reverberant sound component. Early reflected sound is sound generated at the generation position (speaker position) that arrives at the sound receiving point early after reflecting off the walls, floor, and ceiling. Therefore, the filter coefficients for the convolution calculation process for the early reflected sound component are set based on the position of the speaker (microphone position), the position of the speaker, and the environment of the conference room (the size of the conference room, the structure of the walls and ceiling, etc.). Reverberant sound is sound generated at the generation position that arrives at the sound receiving point after multiple reflections, and arrives at the sound receiving point after the early reflected sound. Therefore, the filter coefficients for the convolution calculation process for the reverberant sound component are set based on the environment of the conference room (the size of the conference room, the structure of the walls and ceiling, etc.). Note that the method for setting these early reflected sound components and reverberant sound components is just an example, and other setting methods may also be used. Furthermore, these early reflected sound and reverberant sound each correspond to the "indirect sound" of the present invention.

[0081] When adding such reflected sound, mixer 123A adjusts the level of the picked-up signal to which the reflected sound has been added, using the feed amount set by feed amount setting unit 122 (S25). For example, mixer 123A sets a coefficient for adjusting the amplitude level based on the feed amount, and multiplies the picked-up signal to which the reflected sound has been added by this coefficient. In this way, mixer 123A can generate a picked-up signal to which the reflected sound has been added, with the amplitude level adjusted by the feed amount.

[0082] The mixer 123A mixes the picked-up signals, the amplitude levels of which have been adjusted by the feed amount and to which the reflected sounds have been added, for each of the speakers SP11-SP43 (S26). As a result, the mixer 123A generates speaker supply signals Ss11r-Ss43r for each of the speakers SP11-SP43 and outputs them to the output amplifiers A11-A43. The output amplifiers A11-A43 transmit the speaker supply signals Ss11r-Ss43r to the speakers SP11-SP43 (S27).

[0083] With this configuration, the sound signal processing system 10A can achieve the same effects as the sound signal processing system 10, and can also provide each speaker with realistic sound from multiple speakers SP11-SP43 that is appropriate for the conference room environment, the speaker's position, and the position of each speaker.

[0084] The reflected sound may contain at least one of an early reflected sound component and a reverberant sound component. In this case, by using the reverberant sound component, the sound signal processing system 10A can realize a conversation with a sense of realism that corresponds to the shape of the conference room, the condition of the walls, etc. Also, by using the early reflected sound component, the sound signal processing system 10A can realize a conversation with a sense of realism that corresponds to the position of the speaker in the conference room and the position of the listener (another speaker listening to the voice of the speaker currently speaking).

[0085] In the above-described configuration, the sound signal processing system 10A adds reflected sound to each combination of multiple microphones and multiple speakers. However, the sound signal processing system 10A can also add reflected sound, for example, as follows. The sound signal processing system 10A divides multiple microphones and multiple speakers into multiple groups according to their respective positions. The sound signal processing system 10A adds reflected sound to each combination of multiple microphone groups and multiple speaker groups. In this way, the sound signal processing system 10A can reduce the load of signal processing while adding a sense of realism.

[0086] [Embodiment 3] Fig. 13 is a functional block diagram showing an example of the configuration of a sound signal processing system according to a third embodiment of the present invention. As shown in Fig. 13, the sound signal processing system 10B according to the third embodiment differs from the sound signal processing system 10 according to the first embodiment in that it additionally includes a plurality of equalizers. The other configuration of the sound signal processing system 10B is the same as that of the sound signal processing system 10, and a description of similar parts will be omitted.

[0087] The sound signal processing system 10B includes a plurality of equalizers EQ11 to EQ43, to which a plurality of speaker supply signals Ss11 to Ss43 are input from the feed amount adjustment unit 12.

[0088] The equalizers EQ11-EQ43 generate sound adjustment signals Sq11-Sq43 by performing predetermined signal processing on the speaker supply signals Ss11-Ss43, respectively, and output the sound adjustment signals Sq11-Sq43 to the output amplifiers A11-A43, respectively.

[0089] With such a configuration and processing, the sound signal processing system 10B can achieve the same effects as the sound signal processing system 10, and can also emit sounds with desired tones from each of the multiple speakers.

[0090] The parameters of the equalizers EQ11-EQ43 may be set by the conference manager or by each speaker. Furthermore, the parameters of the equalizers EQ11-EQ43 may be set based on the environment of the conference room, etc.

[0091] [Embodiment 4] Fig. 14 is a functional block diagram showing an example of the configuration of a sound signal processing system according to a fourth embodiment of the present invention. Fig. 15 is a functional block diagram showing an example of the configuration of a conference device used in the sound signal processing system according to the fourth embodiment.

[0092] In the above-described first to third embodiments, all speakers gather in one conference room (physical space) to hold a conference. In the fourth embodiment, speakers are present in different conference rooms (different physical spaces) to hold a conference. Note that the total number of speakers and the number of speakers in one conference room shown in the present embodiment are merely examples, and are not limited to these.

[0093] 14, the sound signal processing system 10C includes a plurality of conferencing devices 81-83, a plurality of speakers SP81-SP83, a plurality of transmitting microphones MICn81-MICn83, a plurality of canceling microphones MICw81-MICw83, and a server 80. These transmitting microphones correspond to the "transmitting sound collecting device" of the present invention, and the canceling microphone corresponds to the "canceling sound collecting device" of the present invention. Furthermore, these speakers correspond to the "sound emitting device" of the present invention. Furthermore, the server 80 corresponds to the "transmitting signal generating unit" of the present invention.

[0094] The plurality of conferencing devices 81-83 and the server 80 are connected to a communication network 800. As a result, the plurality of conferencing devices 81-83 and the server 80 perform data communication with each other through the network 800.

[0095] (Configuration of the ROOMa conference room) Conferencing devices 81 and 82 are placed in a conference room, ROOMa. Speakers SP81 and SP82, transmitting microphones MICn81 and MICn82, and multiple cancellation microphones MICw81 and MICw82 are placed in the conference room, ROOMa. The speaker SP81, transmitting microphone MICn81, and cancellation microphone MICw81 are connected to the conference device 81. The speaker SP82, transmitting microphone MICn82, and cancellation microphone MICw82 are connected to the conference device 82.

[0096] The speaker SP81, the transmitting microphone MICn81, and the cancellation microphone MICw81 are arranged near the speaker 90a. The conferencing device 81 includes a feed amount adjustment unit 812 and an IF 819. The feed amount adjustment unit 812 and the IF 819 are connected to each other and are configured by information processing devices as in the above-described embodiments. The IF 819 is connected to the server 80 via the network 800. The cancellation microphone MICw81 and the speaker SP81 are connected to the feed amount adjustment unit 812. The transmitting microphone MICn81 is connected to the IF 819. The speaker 90a holds a conference using the speaker SP81, the transmitting microphone MICn81, the cancellation microphone MICw81, and the conferencing device 81.

[0097] The speaker SP82, the transmitting microphone MICn82, and the cancellation microphone MICw82 are arranged near the speaker 90b. The conferencing device 82 includes a feed amount adjustment unit 822 and an IF 829. The feed amount adjustment unit 822 and the IF 829 are connected to each other and are configured by an information processing device as in the above-described embodiments. The IF 829 is connected to the server 80 via the network 800. The cancellation microphone MICw82 and the speaker SP82 are connected to the feed amount adjustment unit 822. The transmitting microphone MICn82 is connected to the IF 829. The speaker 90b holds a conference using the speaker SP82, the transmitting microphone MICn82, the cancellation microphone MICw82, and the conferencing device 82.

[0098] (Configuration of the conference room ROOMb) The conference device 83 is placed in the conference room ROOMb. The speaker SP83, the transmitting microphone MICn83, and the canceling microphone MICw83 are placed in the conference room ROOMb. The speaker SP83, the transmitting microphone MICn83, and the canceling microphone MICw83 are connected to the conference device 83.

[0099] The speaker SP83, the transmitting microphone MICn83, and the cancellation microphone MICw83 are arranged near the speaker 90c. The conferencing device 83 includes a feed amount adjustment unit 832 and an IF 839. The feed amount adjustment unit 832 and the IF 839 are connected to each other and are configured by an information processing device as in the above-described embodiments. The IF 839 is connected to the server 80 via the network 800. The cancellation microphone MICw83 and the speaker SP83 are connected to the feed amount adjustment unit 832. The transmitting microphone MICn83 is connected to the IF 839. The speaker 90c holds a conference using the speaker SP83, the transmitting microphone MICn83, the cancellation microphone MICw83, and the conferencing device 83.

[0100] (Features of the transmitting microphone and the canceling microphone) The sound pickup directivity of the multiple transmitting microphones MICn81, MICn82, and MICn83 is narrow. As a result, the transmitting microphone MICn81 picks up the voice of speaker 90a and hardly picks up other voices. The transmitting microphone MICn82 picks up the voice of speaker 90b and hardly picks up other voices. The transmitting microphone MICn83 picks up the voice of speaker 90c and hardly picks up other voices.

[0101] The sound pickup directivity of the multiple cancellation microphones MICw81-MICw83 is wider than the sound pickup directivity of the multiple transmission microphones MICn81-MICn83. As a result, the cancellation microphone MICw81 picks up not only the voice of speaker 90a but also the voice of speaker 90b. The cancellation microphone MICw82 picks up not only the voice of speaker 90b but also the voice of speaker 90a. In other words, the cancellation microphones MICw81 and MICw82 pick up sounds from the positions of speakers 90a and 90b and the space around them. Similarly, the cancellation microphone MICw83 picks up sounds from the position of speaker 90c in conference room ROOMb and the space around it.

[0102] (Specific details of sound signal processing) With the above-described configuration, the sound signal processing system 10C processes sound signals in a conference as follows.

[0103] (Speaker 90a) The transmitting microphone MICn81 picks up the voice of the speaker 90a and generates a picked-up voice signal Sa. The transmitting microphone MICn81 outputs the picked-up voice signal Sa to the IF 819 of the conferencing device 81. The IF 819 transmits the picked-up voice signal Sa to the server 80 via the network 800. The canceling microphone MICw81 picks up voices from the positions of the speakers 90a and 90b in the conference room ROOMa and the space around them and generates a canceling picked-up voice signal Sxa. Because the canceling microphone MICw81 has wide directivity, the canceling picked-up voice signal Sxa is essentially a picked-up voice signal obtained by adding together the picked-up signal Sa of the voice of the speaker 90a and the picked-up signal Sb of the voice of the speaker 90b (Sxa = Sa + Sb).

[0104] (Recording of speaker 90b) The transmitting microphone MICn82 picks up the voice of the speaker 90b and generates a picked-up voice signal Sb. The transmitting microphone MICn82 outputs the picked-up voice signal Sb to the IF 829 of the conferencing device 82. The IF 829 transmits the picked-up voice signal Sb to the server 80 via the network 800. The canceling microphone MICw82 picks up voices from the positions of the speakers 90a and 90b in the conference room ROOMa and the space around them and generates a canceling picked-up voice signal Sxb. Because the canceling microphone MICw82 has wide directivity, the canceling picked-up voice signal Sxb is essentially a picked-up voice signal obtained by adding together the picked-up signal Sb of the voice of the speaker 90b and the picked-up signal Sa of the voice of the speaker 90a (Sxb = Sb + Sa).

[0105] (Recording of speaker 90c) The transmitting microphone MICn83 picks up the voice of the speaker 90c and generates a picked-up voice signal Sc. The transmitting microphone MICn83 outputs the picked-up voice signal Sc to the IF 839 of the conferencing device 83. The IF 839 transmits the picked-up voice signal Sc to the server 80 via the network 800. The canceling microphone MICw83 picks up voice from the position of the speaker 90c in the conference room ROOMb and the space around it and generates a canceling picked-up voice signal Sxc. Although the canceling microphone MICw83 has wide directivity, since only the speaker 90c is present in the conference room ROOMb, the canceling picked-up voice signal Sxc is essentially the picked-up voice signal Sc of the voice of the speaker 90c (Sxc = Sc).

[0106] (Processing on the server) The server 80 adds the collected sound signals Sa, Sb, and Sc to generate an audio sum signal Sall (=Sa+Sb+Sc). The server 80 transmits the audio sum signal Sall to a plurality of conference devices 81-83 via a network 800. This audio sum signal Sall corresponds to the "transmission signal" of the present invention.

[0107] (Speech to Speaker 90a) The IF 819 of the conference device 81 receives the audio addition signal Sall and outputs it to the transmission amount adjustment unit 812 .

[0108] 15, the transmission amount adjustment unit 812 includes a delay processing unit 8121 and a subtractor 8122. The delay processing unit 8121 applies a delay amount Δt to the cancellation picked-up signal Sxa. The delay amount Δt in the conference device 81 is set, for example, by adding together the transmission time of the picked-up signal Sa from the conference device 81 to the server 80, the transmission time of the audio added signal Sall from the server 80 to the conference device 81, and the generation time of the audio added signal Sall in the server 80.

[0109] The delay processing unit 8121 outputs the delayed cancellation collected sound signal Δt(Sxa) to the subtractor 8122. Note that Δt(Sxa) shown in FIG. 15 means a signal Sxa(t+Δt) obtained by applying delay processing of a delay amount Δt to the time axis signal Sxa(t). The subtractor 8122 receives an audio added signal Sall from the IF 819. When input to the subtractor 8122, the audio added signal Sall is delayed with respect to the collected sound signal Sa and the cancellation collected sound signal Sxa. The delay amount of this audio added signal Sall is the same as the above-mentioned delay amount Δt, and the subtractor 8122 receives an audio added signal Δt(Sall) with a delay. Note that Δt(Sall) shown in FIG. 15 means the time axis signal Sall(t+Δt). The subtractor 8122 subtracts the delayed cancellation collected sound signal Δt(Sxa) from the delayed audio added signal Δt(Sall). The audio added signal Sall is an added signal of the collected sound signals Sa, Sb, and Sc, and the cancellation collected sound signal Sxa is an added signal of the collected sound signals Sa and Sb. Δt(Sxa)=Δt(Sa+Sb) shown in FIG. 15 means the time axis signal Sxa(t+Δt)=Sa(t+Δt)+Sb(t+Δt). Δt(Sall)=Δt(Sa+Sb+Sc) shown in FIG. 15 means the time axis signal Sall(t+Δt)=Sa(t+Δt)+Sb(t+Δt)+Sc(t+Δt).

[0110] Therefore, the output signal from the subtractor 8122 is a signal in which the sound collection signal Sa and the sound collection signal Sb are suppressed, and becomes the sound emission signal Δt(Sc). Δt(Sc) shown in FIG. 15 means the signal Sc(t+Δt) on the time axis. The subtractor 8122 transmits (outputs) the sound emission signal Δt(Sc) to the speaker SP81. As a result, the sound emission signal Δt(Sc) transmitted from the send amount adjustment unit 812 to the speaker SP81 is a signal adjusted so that the send amount of the sound collection signal Sa and the sound collection signal Sb is minimized.

[0111] The speaker SP81 emits the sound emission signal Δt(Sc), and the speaker 90a hears this sound. As a result, the speaker 90a can hear only the voice of the speaker 90c in the other conference room ROOMb from the speaker SP81.

[0112] (Speaker 90b) The IF 829 of the conferencing device 82 receives the audio addition signal Sall and outputs it to the audio forwarding amount adjustment unit 822. The audio forwarding amount adjustment unit 822 has the same configuration as the audio forwarding amount adjustment unit 812. This allows the speaker 90b to hear only the audio of the speaker 90c in another conference room ROOMb from the speaker SP82.

[0113] (Effect of sound signal processing system 10C on speakers in the same room) Here, speaker 90a and speaker 90b are present in the same conference room ROOMa, and therefore hear each other's voices directly. For example, as shown in FIG. 14, speaker 90a directly hears speaker 90b's voice Sbd. Similarly, speaker 90b directly hears speaker 90a's voice Sad. Meanwhile, due to the above-described processing, speaker 90a and speaker 90b, who are present in the same conference room ROOMa, do not hear each other's voices from speaker SP81 and speaker SP82, respectively.

[0114] Therefore, the speaker 90a and the speaker 90b are prevented from hearing each other's voices in double with a time lag. As a result, the sound signal processing system 10C can provide a natural conference environment for multiple speakers in the same room during a Web conference.

[0115] Furthermore, the sound signal processing system 10C can suppress feedback because it does not emit the signal picked up by a microphone from the speaker closest to that microphone in the same room, as the sound signal processing system 10. As a result, in a conference room ROOMa where multiple speakers 90a, 90b are present and where multiple microphones MICn81, MICn82 and multiple speakers SP81, SP82 are arranged, the sound signal processing system 10C can suppress feedback caused by the voice of speaker 90a and howling caused by the voice of speaker 90b.

[0116] In this case, since the feed amount adjustment unit is composed of a delay processing unit and a subtractor, feedback can be suppressed with a simple configuration and processing. Therefore, similar to the sound signal processing system 10, the sound signal processing system 10C does not require complex signal processing, and can suppress feedback and appropriately output sounds other than those targeted for feedback suppression.

[0117] (Speech to Speaker 90c) The IF 839 of the conference device 83 receives the audio addition signal Sall and outputs it to the forwarding amount adjustment unit 832. The forwarding amount adjustment unit 832 has the same configuration as the forwarding amount adjustment unit 812.

[0118] Therefore, the output signal from the feed amount adjustment unit 832 is a signal obtained by subtracting the delayed cancellation sound collection signal Sxc from the delayed sound addition signal Δt(Sall). As a result, the output signal from the feed amount adjustment unit 832 becomes a signal in which the sound collection signal Sc is suppressed, and becomes the sound emission signal Δt(Sa+Sb). The feed amount adjustment unit 832 transmits (outputs) the sound emission signal Δt(Sa+Sb) to the speaker SP83. As a result, the sound emission signal Δt(Sa+Sb) transmitted from the feed amount adjustment unit 832 to the speaker SP83 is a signal adjusted so that the feed amount of the sound collection signal Sc is minimized.

[0119] The speaker SP83 emits the sound signal Δt(Sa+Sb), and the speaker 90c hears this sound. This allows the speaker 90c to hear the voices of the speakers 90a and 90b in the other conference room ROOMa from the speaker SP83.

[0120] Similarly to conference room ROOMa, in conference room ROOMb, which has multiple speakers 90c and is equipped with microphones MICn83 and speakers SP83, the sound signal processing system 10C can suppress feedback caused by the voices of the speakers 90c. The sound signal processing system 10C can also suppress feedback in conference room ROOMb without requiring complex signal processing, and can appropriately output voices other than those targeted for feedback suppression.

[0121] In the sound signal processing system of each of the above-described embodiments, a second speaker is set. However, the sound signal processing system may set only the first speaker and omit setting the second speaker. In addition, in the sound signal processing system of each of the above-described embodiments, the number of second speakers set is not limited to two, and may be one or three or more depending on the arrangement of multiple speakers. In addition, the sound signal processing system may adjust the feed amount to be subtracted from the reference feed amount for each of multiple speakers depending on the distance from the microphone that picked up the speaker's voice.

[0122] Furthermore, in each of the above-described embodiments, the microphone is a stationary type. In other words, in each of the above-described embodiments, the microphone does not move. However, the above-described configuration can be applied even if the microphone (the microphone for generating the collected sound signal) moves with the speaker, such as a pin microphone. In this case, the sound signal processing system may have, for example, the following configuration.

[0123] The sound signal processing system has a configuration for detecting the position of a speaker (a microphone for generating a sound collection signal). The sound signal processing system also stores the positions of multiple speakers. From the detected position of the speaker (a microphone for generating a sound collection signal), the sound signal processing system calculates the distance between the microphone for generating a sound collection signal and the multiple speakers. The sound signal processing system determines the speaker closest to the microphone for generating a sound collection signal to be the first speaker, and determines the next closest speaker to be the second speaker. Thereafter, the sound signal processing system adjusts the send amount as described above. This allows the sound signal processing system to suppress feedback and appropriately output sounds other than those targeted for feedback suppression without requiring complex signal processing.

[0124] Furthermore, in each of the above-described embodiments, the cases where the multiple microphones and the multiple speakers are arranged separately have been described. However, each of the multiple microphones and the speaker (first speaker) closest to each microphone may be arranged integrally. For example, each of the multiple microphones and the speaker (first speaker) closest to each microphone may be housed in a single housing. This fixes the positional relationship between each of the multiple microphones and the speaker (first speaker) closest to each microphone. Therefore, the sound signal processing system can more reliably suppress howling.

[0125] The description of the present embodiment is illustrative in all respects and is not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention is intended to include all modifications that are equivalent to the claims and fall within the scope thereof. [Explanation of symbols]

[0126] 10, 10A, 10B, 10C: Sound signal processing system 11: Speaker voice detection unit 12, 12A: Feed rate adjustment section 80: Server 81, 82, 83: Conference equipment 111: Signal level detection unit 112: Specific microphone detection unit 120: Speaker DB 121: Speaker determination unit 122: Feed amount setting section 123, 123A: Mixer 800:Network 812, 822, 832: Feed rate adjustment unit 819, 829, 839:IF 8121: Delay processing unit 8122: Subtractor

Claims

1. A sound signal processing method used in a sound signal processing system including a plurality of sound collection devices and a plurality of sound emission devices, The plurality of sound collecting devices include a transmitting sound collecting device and a canceling sound collecting device having a wider directivity than the transmitting sound collecting device, a sound collection signal collected by the sound collection device for transmission is transmitted to the plurality of sound emission devices; canceling the cancellation sound collection signal collected by the cancellation sound collection device from the sound collection signals received by each of the plurality of sound emitting devices, and outputting the cancelled sound signal to each of the sound emitting devices; Sound signal processing method.

2. the transmitting sound collecting device includes a first transmitting sound collecting device and a second transmitting sound collecting device, the cancellation sound collecting device includes a first cancellation sound collecting device and a second cancellation sound collecting device, the plurality of sound emitting devices include a first sound emitting device and a second sound emitting device, the first transmitting sound collecting device, the first canceling sound collecting device, and the first sound emitting device are installed in a first conference room; the second transmitting sound collecting device, the second canceling sound collecting device, and the second sound emitting device are installed in a second conference room, the first transmitting sound collecting device collects the voice of a first speaker in the first conference room and transmits a first collected sound signal; the second transmitting sound collecting device collects the voice of a second speaker in the second conference room and transmits a second collected sound signal; the first cancellation sound collecting device collects a first cancellation sound collecting signal, the second cancellation sound collecting device collects a second cancellation sound collecting signal, a sound signal obtained by canceling the first cancellation sound collection signal from the second sound collection signal is output to the first sound emission device; A sound signal obtained by canceling the second cancellation sound collection signal from the first sound collection signal is output to the second sound emission device. The sound signal processing method according to claim 1 .

3. The transmitting sound collecting device further includes a third transmitting sound collecting device installed in the first conference room, The cancellation sound collecting device further includes a third cancellation sound collecting device installed in the first conference room, the plurality of sound emitting devices further include a third sound emitting device installed in the first conference room, the third transmitting sound collecting device collects the voice of a third speaker in the first conference room and transmits a third collected sound signal; the first cancellation sound collection device collects the voices of the first speaker and the third speaker as the first cancellation sound collection signal; the third cancellation sound collection device collects the voices of the first speaker and the third speaker to generate a third cancellation sound collection signal; a sound signal obtained by canceling the first cancellation sound collection signal from the second sound collection signal is output to the first sound emission device; A sound signal obtained by canceling the third cancellation sound collection signal from the second sound collection signal is output to the third sound emission device. The sound signal processing method according to claim 2 .

4. a server receives and adds the first collected signal, the second collected signal, and the third collected signal to generate an audio sum signal; a sound signal obtained by canceling the first cancellation sound collection signal from the audio added signal is output to the first sound emitting device; a sound signal obtained by canceling the second cancellation sound collection signal from the audio addition signal is output to the second sound output device; A sound signal obtained by canceling the third cancellation sound collection signal from the audio added signal is output to the third sound emitting device. The sound signal processing method according to claim 3 .

5. The first cancellation sound collection signal is subjected to a first delay processing, and a sound signal obtained by canceling the first cancellation sound collection signal after the first delay processing from the audio added signal is output to the first sound emitting device, The second cancellation sound collection signal is subjected to second delay processing, and a sound signal obtained by canceling the second cancellation sound collection signal after the second delay processing from the audio added signal is output to the second sound emitting device, The third cancellation sound collection signal is subjected to third delay processing, and a sound signal obtained by canceling the second cancellation sound collection signal after the third delay processing from the audio added signal is output to the third sound emitting device. The sound signal processing method according to claim 4.

6. A sound signal processing system including a plurality of sound collection devices and a plurality of sound emission devices corresponding to the plurality of sound collection devices, The plurality of sound collecting devices include a transmitting sound collecting device and a canceling sound collecting device having a wider directivity than the transmitting sound collecting device, the transmitting sound collection device transmits collected sound signals to the plurality of sound emission devices; The sound signal processing system includes a transmission amount adjustment unit that cancels a cancellation sound collection signal collected by the cancellation sound collection device from a sound collection signal received by each sound emitting device of the plurality of sound emitting devices, and outputs the canceled sound signal to each sound emitting device. Sound signal processing system.

7. the transmitting sound collecting device includes a first transmitting sound collecting device and a second transmitting sound collecting device, the cancellation sound collecting device includes a first cancellation sound collecting device and a second cancellation sound collecting device, the plurality of sound emitting devices include a first sound emitting device and a second sound emitting device, the first transmitting sound collecting device, the first canceling sound collecting device, and the first sound emitting device are installed in a first conference room; the second transmitting sound collecting device, the second canceling sound collecting device, and the second sound emitting device are installed in a second conference room, the first transmitting sound collecting device collects the voice of a first speaker in the first conference room and transmits a first collected sound signal; the second transmitting sound collecting device collects the voice of a second speaker in the second conference room and transmits a second collected sound signal; the first cancellation sound collecting device collects a first cancellation sound collecting signal, the second cancellation sound collecting device collects a second cancellation sound collecting signal, a sound signal obtained by canceling the first cancellation sound collection signal from the second sound collection signal is output to the first sound emission device; A sound signal obtained by canceling the second cancellation sound collection signal from the first sound collection signal is output to the second sound emission device. The sound signal processing system according to claim 6 .

8. The transmitting sound collecting device further includes a third transmitting sound collecting device installed in the first conference room, The cancellation sound collecting device further includes a third cancellation sound collecting device installed in the first conference room, the plurality of sound emitting devices further include a third sound emitting device installed in the first conference room, the third transmitting sound collecting device collects the voice of a third speaker in the first conference room and transmits a third collected sound signal; the first cancellation sound collection device collects the voices of the first speaker and the third speaker as the first cancellation sound collection signal; the third cancellation sound collection device collects the voices of the first speaker and the third speaker to generate a third cancellation sound collection signal; a sound signal obtained by canceling the first cancellation sound collection signal from the second sound collection signal is output to the first sound emission device; A sound signal obtained by canceling the third cancellation sound collection signal from the second sound collection signal is output to the third sound emission device. The sound signal processing system according to claim 7 .

9. a server receives and adds the first collected signal, the second collected signal, and the third collected signal to generate an audio sum signal; a sound signal obtained by canceling the first cancellation sound collection signal from the audio added signal is output to the first sound emitting device; a sound signal obtained by canceling the second cancellation sound collection signal from the audio addition signal is output to the second sound output device; A sound signal obtained by canceling the third cancellation sound collection signal from the audio added signal is output to the third sound emitting device. The sound signal processing system according to claim 8 .

10. The first cancellation sound collection signal is subjected to a first delay processing, and a sound signal obtained by canceling the first cancellation sound collection signal after the first delay processing from the audio added signal is output to the first sound emitting device, The second cancellation sound collection signal is subjected to second delay processing, and a sound signal obtained by canceling the second cancellation sound collection signal after the second delay processing from the audio added signal is output to the second sound emitting device, The third cancellation sound collection signal is subjected to third delay processing, and a sound signal obtained by canceling the second cancellation sound collection signal after the third delay processing from the audio added signal is output to the third sound emitting device. The sound signal processing system according to claim 9.

Citation Information

Patent Citations

  • Echo control system

    JP2002204187A

  • Headphone, and noise canceling circuit and method

    JP2009141679A

  • Portable telephone set, and recording method

    JP2009218750A

  • Noise reduction circuit and noise reduction method

    JP2013110579A

  • Head set

    JP2017028351A