Sound pickup device
By integrating a microphone, vibration sensor, echo canceller, and adaptive control unit, the sound collection device enhances audio quality by minimizing residual echoes through adaptive learning speed adjustments, addressing the issue of degraded audio quality from residual echoes.
Patent Information
- Application Number
- JP2022006136
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-01-19
AI Technical Summary
Existing sound collection devices that convert vibration signals into audio signals face issues with residual echo components that degrade audio quality, particularly during two-way communication, as conventional echo cancellers may erroneously update filter coefficients.
Incorporating a microphone for air vibrations, a vibration sensor, an echo canceller, an adaptive filter, and an adaptive control unit that adjusts learning speeds based on echo levels and voice activity to minimize residual echoes, ensuring high-quality audio signal conversion.
The solution effectively improves the quality of audio signals generated by vibration sensors by adaptively controlling echo cancellation, reducing residual echoes and maintaining audio clarity during communication.
Smart Images

Figure 0007790161000001 
Figure 0007790161000002 
Figure 0007790161000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound collection device. [Background technology]
[0002] Patent Document 1 describes a sound collection device that can capture clear sound even in a noisy environment by including a microphone that generates a sound signal based on air vibrations and a vibration sensor that generates a vibration signal corresponding to the sound signal based on bone vibrations. The former microphone is sometimes called an air conduction microphone, and the latter vibration sensor is sometimes called a bone conduction microphone.
[0003] The sound collection device described in Patent Document 1 includes a filtering unit that converts a vibration signal generated by a vibration sensor into an audio signal, and outputs an audio signal based on the vibration signal generated by the vibration sensor. The sound collection device described in Patent Document 1 is configured to update the filter coefficients of the filtering unit so that an error signal, which is the difference between the audio signal output from the filtering unit and the audio signal generated by the microphone, becomes smaller. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-251354 Summary of the Invention [Problem to be solved by the invention]
[0005] During two-way communication such as internet calls, a sound collector may pick up the user's own voice and transmit it to the other party, and the voice emitted by the other party and transmitted via the line may be output from a speaker. When the sound collector picks up the voice output from the speaker, the other party's voice contains an echo component. It is common to use an echo canceller to cancel the echo component superimposed on the user's voice signal. However, in a sound collector that converts a vibration signal into an audio signal, if there is a residual echo component that could not be canceled, the filtering unit may erroneously update the filter coefficients, degrading the quality of the audio signal based on the vibration signal.
[0006] The present invention aims to provide a sound collection device that can further improve the quality of an audio signal based on a vibration signal generated by a vibration sensor in an environment where the echo component of the voice of the communication partner may be superimposed on the user's audio signal. [Means for solving the problem]
[0007] The present invention includes a microphone that generates a first audio signal based on air vibrations, a vibration sensor that generates a vibration signal based on vibrations transmitted to a human body by speaking, an echo canceller that suppresses echo components superimposed on a first audio signal when the microphone picks up a second audio signal transmitted from a communication partner and received via a line, the second audio signal being reproduced by a speaker, an adaptive filter that uses an audio signal output from the echo canceller as a target signal and multiplies the vibration signal by a coefficient to generate a converted audio signal so as to bring the vibration signal closer to the target signal, a subtractor that generates a residual signal that is the difference between the target signal and the converted audio signal, and an adaptive control unit that controls the coefficient by which the adaptive filter multiplies the vibration signal so as to reduce the residual signal. The adaptive control unit includes a residual echo level estimating unit that estimates a residual echo level remaining in the target signal based on a sound pressure level of the target signal and a sound pressure level of the second audio signal, and an adaptive filter learning speed setting unit that controls the adaptive filter to update the coefficients at a first speed if the vibration signal indicates a voice section and the residual echo level is equal to or less than a predetermined threshold, and controls the adaptive filter to update the coefficients at a second speed slower than the first speed or not to update the coefficients if the conditions are not satisfied. To provide a sound collection device. The present invention includes a microphone that generates a first audio signal based on air vibrations, a vibration sensor that generates a vibration signal based on vibrations transmitted to the human body by speaking, an echo canceller that suppresses echo components superimposed on a first audio signal when the microphone picks up a second audio signal transmitted from a communication partner and received via a line and played back by a speaker, an adaptive filter that uses an audio signal output from the echo canceller as a target signal and multiplies the vibration signal by a coefficient to generate a converted audio signal so as to bring the vibration signal closer to the target signal, a subtractor that generates a residual signal that is the difference between the target signal and the converted audio signal, and a subtractor that reduces the coefficient by which the adaptive filter multiplies the vibration signal so as to reduce the residual signal. and an adaptive control unit that controls the adaptive filter to update the coefficients at a first speed, wherein the adaptive control unit comprises: a residual echo level estimation unit that estimates a residual echo level remaining in the target signal based on the sound pressure level of the target signal and the sound pressure level of the second audio signal; a level ratio calculation unit that calculates a level ratio between the residual echo level and a vibration signal level that indicates the sound pressure level of the vibration signal; and an adaptive filter learning speed setting unit that controls the adaptive filter to update the coefficients at a first speed if the conditions that the vibration signal indicates a voice section and the level ratio exceeds a predetermined threshold are satisfied, and controls the adaptive filter to update the coefficients at a second speed that is slower than the first speed or controls the adaptive filter not to update the coefficients if the conditions are not satisfied. [Effects of the Invention]
[0008] According to the sound collection device of the present invention, the quality of the sound signal based on the vibration signal generated by the vibration sensor can be further improved in an environment where the echo component of the voice of the communication partner may be superimposed on the user's sound signal. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a sound collection device according to an embodiment; [Figure 2] 1 is a block diagram showing an example of the configuration of an echo canceller included in a sound collection device according to an embodiment; [Figure 3A] FIG. 2 is a waveform diagram illustrating an audio signal generated by a microphone. [Figure 3B] FIG. 4 is a waveform diagram showing a vibration signal generated by a vibration sensor. [Figure 4] FIG. 4 is a characteristic diagram showing frequency characteristics of an audio signal and a vibration signal. [Figure 5] 3A to 3C are waveform diagrams showing examples of a voice signal generated by a microphone, a voice of the other party output from a speaker, and a vibration signal generated by a vibration sensor. [Figure 6] 3 is a block diagram showing a specific example of the configuration of the adaptive control unit 12 in FIG. 2. FIG. [Figure 7] FIG. 3 is a block diagram showing a specific example of the configuration of the adaptive filter 13 in FIG. 2. [Figure 8] FIG. 2 is a block diagram showing a first specific example configuration of the adaptive control unit 5 of FIG. [Figure 9] FIG. 2 is a block diagram showing a second specific example configuration of the adaptive control unit 5 of FIG. [Figure 10] FIG. 2 is a block diagram showing a specific example of the configuration of the adaptive filter 6 in FIG. [Figure 11A] 1 is a partial flowchart showing the operation of a sound collection device according to an embodiment. [Figure 11B] 11B is a partial flowchart continuing from FIG. 11A, showing the operation of the sound collection device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] A sound collection device according to an embodiment will be described below with reference to the accompanying drawings. FIG. 1 shows a sound collection device 200 according to an embodiment. In FIG. 1, a microphone 1 generates an audio signal (first audio signal) based on air vibrations. An A / D converter 2 performs A / D conversion on the analog audio signal supplied from the microphone 1, and supplies the digital audio signal to an echo canceller 20. Although the first audio signal is close to the audio perceived by humans through the ears, the first audio signal may contain echo components. Therefore, it is desirable to use the audio signal output from the echo canceller 20 as a target signal when converting a vibration signal, which will be described later, into an audio signal.
[0011] A digital voice signal (second voice signal), which is voice transmitted from the other party and received via the server and line 11 (hereinafter referred to as the other party voice), is supplied to the echo canceller 20 and the D / A converter 15. The second voice signal is sometimes referred to as the other party voice signal. The D / A converter 15 D / A converts the input digital voice signal and supplies an analog voice signal to the speaker 16. The speaker 16 reproduces the input voice signal and outputs the other party voice. At this time, when the microphone 1 picks up the other party voice output from the speaker 16, the voice spoken by the other party may be superimposed as an echo component on the voice spoken by the user.
[0012] The echo canceller 20 suppresses the echo component superimposed on the audio signal output from the A / D converter 2 using the audio signal received via the line 11. The echo canceller 20 supplies the audio signal with the echo component suppressed to the adaptive control unit 5 and the subtractor 7. Although the echo canceller 20 may not be able to completely cancel the echo component superimposed on the audio signal picked up by the microphone 1, the audio signal output from the echo canceller 20 will be referred to as an echo-canceled audio signal.
[0013] As an example, the echo canceller 20 can be configured as shown in Fig. 2. As shown in Fig. 2, the echo canceller 20 includes an adaptive control unit 12, an adaptive filter 13, and a subtractor 14. The adaptive control unit 12 generates an adaptive filter control signal for controlling the adaptive filter 13 and supplies it to the adaptive filter 13. The adaptive filter 13 multiplies the other party's voice signal by a coefficient in accordance with the adaptive filter control signal to generate a cancellation voice signal for canceling the echo component from the voice signal on which the echo component is superimposed, and supplies this cancellation voice signal to the subtractor 14. A specific configuration example of the adaptive filter 13 will be described later.
[0014] The echo canceller 20 is not limited to the configuration including the adaptive filter 13 as shown in Fig. 2, and other echo suppression methods may be used. The specific configuration of the echo canceller 20 is not limited.
[0015] Returning to FIG. 1 , the vibration sensor 3 generates a vibration signal based on vibrations transmitted to the human body (the body of the user of the sound collection device 200). The vibration sensor 3 is arranged so as to come into contact with the surface of the human body. The vibration sensor includes a vibration receiving device embedded in the body, a microphone arranged so as to come into direct contact with the human body, a camera that captures the vibrations transmitted to the surface of the human body as images, and a rangefinder that captures the vibrations transmitted to the surface of the human body as position information. The A / D converter 4 A / D converts the analog vibration signal supplied from the vibration sensor 3 and supplies the digital vibration signal to the adaptive control unit 5 and the adaptive filter 6.
[0016] FIG. 3A shows the audio signal generated by the microphone 1 when the user speaks, and FIG. 3B shows the vibration signal generated by the vibration sensor 3 during the same period as the audio signal in FIG. 3A. Comparing FIG. 3A and FIG. 3B reveals that the audio signal and the vibration signal have different sound pressure levels. FIG. 4 shows the frequency characteristics of the audio signal and the vibration signal. In some frequency bands, the sound pressure level of the vibration signal, shown by the dashed line, is lower than the sound pressure level of the audio signal, shown by the solid line. When the vibration signal is supplied to a speaker and output as audio, the resulting sound is muffled compared to when the audio signal generated by the microphone 1 is supplied to a speaker and output as audio, and sounds different from the original audio.
[0017] As will be described later, the adaptive filter 6 uses the echo-canceled voice signal output from the echo canceller 20 as a target signal, corrects the vibration signal so that it approaches the target signal, and generates a converted voice signal, which is supplied to the line 11. The line 11 is, for example, an Internet line. The converted voice signal is transmitted to the other party via the line 11 and an Internet communication server (not shown).
[0018] In Fig. 5, (a) shows the voice signal generated by the microphone 1, (b) shows the other party's voice output from the speaker 16, and (c) shows the vibration signal generated by the vibration sensor 3. In Fig. 5(b), sections b1, b2, and b3 are voice sections (speech sections) where the voice of the communication partner is present, and sections other than b1, b2, and b3 are non-voice sections (non-speech sections) where the other party's voice is not present. In Fig. 5(c), sections c1 and c2 are voice sections where the voice of the user is present, and sections other than c1 and c2 are non-voice sections where the user's voice is not present.
[0019] Section b3 overlaps with section c2 for the most part, and since both the other party's voice and the user's voice have high sound pressure levels, echo components are likely to remain even after echo cancellation by an echo canceller. Section b1 overlaps with section c1, and although the sound pressure level of the other party's voice is low, echo components may remain. Section b2 is located in a non-speech section of the user's voice, and it is expected that echo cancellation by an echo canceller will sufficiently cancel the echo components.
[0020] Fig. 6 shows a specific example of the configuration of the adaptive control unit 12 shown in Fig. 2. The adaptive control unit 12 includes a voice activity detection unit 121 and an adaptive filter learning speed setting unit 122. The voice activity detection unit 121 detects the voice activity of the other party's voice using a technique called VAD (Voice Activity Detection), and supplies other party's voice activity information to the adaptive filter learning speed setting unit 122. The voice activity detection unit 121 detects the voice activity based on at least whether or not the sound pressure level exceeds a predetermined level.
[0021] In general, the adaptive control unit 12 generates an adaptive filter control signal for changing the operation of the adaptive filter 13 depending on whether it is a voice section where the other party's voice is present or a non-voice section where the other party's voice is not present. Specifically, if the other party's voice section information indicates a voice section of the other party's voice, the adaptive filter learning speed setting unit 122 generates an adaptive filter control signal for setting the learning speed to active and supplies it to the adaptive filter 13. If the other party's voice section information indicates a non-voice section of the other party's voice, the adaptive filter learning speed setting unit 122 generates an adaptive filter control signal for setting the learning speed to save and supplies it to the adaptive filter 13.
[0022] The learning speed being active means that the adaptive operation in the adaptive filter 13 is actively promoted, and the learning speed being saved means that the adaptive operation in the adaptive filter 13 is suppressed or stopped.
[0023] Specifically, actively promoting the adaptive operation of the adaptive filter 13 means controlling the adaptive filter 13 to update the coefficients (described later) so that the adaptive filter 13 generates a cancellation signal for canceling the echo component at a first speed in a short time. Suppressing the adaptive operation of the adaptive filter 13 means controlling the adaptive filter 13 to update the coefficients at a second speed slower than the first speed over a long time. Stopping the adaptive operation of the adaptive filter 13 means controlling the adaptive filter 13 not to update the coefficients.
[0024] 7 shows a specific example of the configuration of adaptive filter 13 using an FIR filter. Adaptive filter 13 includes adaptive coefficient update unit 131, delay units 1321-132n, multipliers 1330-133n, and adders 1341-134n, where n is a number ranging from several tens to several hundreds. Delay units 1321-132n delay each sample of the input digital remote party voice signal by one clock and output the delayed signal. Multipliers 1330-133n multiply the sample input to delay unit 1321 and each sample output from delay units 1321-132n by a coefficient, respectively, and output the results.
[0025] The adders 1341 to 134n respectively add the outputs of the multipliers 1330 and 1331, the outputs of the adder 1341 and multiplier 1332, the outputs of the adder 1342 and multiplier 1333, ..., the outputs of the adder 134(n-1) (not shown) and multiplier 133n. As a result, the adder 134n outputs a cancellation audio signal for canceling the echo component from the audio signal on which the echo component is superimposed.
[0026] The subtractor 14 subtracts the cancellation audio signal from the audio signal on which the echo component is superimposed and output from the A / D converter 2, and outputs an echo-canceled audio signal. The adaptive coefficient update unit 131 updates the coefficients by which the multipliers 1330 to 133n multiply the input samples so as to generate a cancellation audio signal in which as little echo component remains as possible.
[0027] At this time, when the adaptive filter control signal is high indicating active, adaptive coefficient update unit 131 updates the coefficients supplied to multipliers 1330 to 133n in a short time.When the adaptive filter control signal is low indicating save, adaptive coefficient update unit 131 updates the coefficients supplied to multipliers 1330 to 133n over a long time, or does not update the coefficients.
[0028] 8 shows a first specific example of the configuration of the adaptive control unit 5. As shown in FIGS. 1 and 8, the adaptive control unit 5 receives the voice signal and vibration signal output from the echo canceller 20 as well as the other party's voice signal supplied from the line 11. The adaptive control unit 5 includes a voice activity detection unit 51, a residual echo level estimation unit 52, and an adaptive filter learning speed setting unit 55.
[0029] The voice activity detection unit 51 detects the voice activity of the vibration signal using a technique called VAD and supplies the voice activity information to the adaptive filter learning speed setting unit 55. The voice activity detection unit 51 detects the voice activity based on whether or not the sound pressure level exceeds a predetermined level. The voice signal output from the echo canceller 20 and the other party's voice signal are input to the residual echo level estimation unit 52. The residual echo level estimation unit 52 estimates the residual echo level remaining in the target signal by calculating the relative sound pressure level ratio per predetermined unit time between the sound pressure level of the other party's voice signal and the sound pressure level of the voice signal output from the echo canceller 20. The predetermined unit time is, for example, several milliseconds or several tens of milliseconds. The residual echo level estimation unit 52 supplies the residual echo level to the adaptive filter learning speed setting unit 55.
[0030] If the voice activity information indicates a user voice activity and the residual echo level is equal to or lower than a predetermined threshold, which is a first condition, the adaptive filter learning rate setting unit 55 generates an adaptive filter control signal for setting the learning rate to active and supplies it to the adaptive filter 6. If the first condition is not satisfied, the adaptive filter learning rate setting unit 55 generates an adaptive filter control signal for setting the learning rate to save and supplies it to the adaptive filter 6.
[0031] The learning speed being active means that the adaptive operation in the adaptive filter 6 is actively promoted, and the learning speed being saved means that the adaptive operation in the adaptive filter 6 is suppressed or stopped.
[0032] Specifically, actively promoting the adaptive operation of the adaptive filter 6 means controlling the adaptive filter 6 to update the coefficient (described later) that is multiplied by the vibration signal at a third speed in a short time. Suppressing the adaptive operation of the adaptive filter 6 means controlling the adaptive filter 6 to update the coefficient over a long time at a fourth speed that is slower than the third speed. Stopping the adaptive operation of the adaptive filter 6 means controlling the adaptive filter 6 not to update the coefficient. The third speed may be the same as or different from the first speed, and the fourth speed may be the same as or different from the second speed.
[0033] If the voice interval information does not indicate the user's voice interval, there is no voice signal that can serve as the target signal, so the learning speed should be set to save. If the voice interval information indicates the user's voice interval but the residual echo level exceeds the threshold, the presence of the residual echo component may degrade the quality of the converted voice signal, so the learning speed should be set to save. A threshold to be compared with the residual echo level that does not degrade the quality of the converted voice signal by the adaptive filter 6 may be measured in advance and stored in the memory unit.
[0034] Fig. 9 shows a second specific configuration example of the adaptive control unit 5. The adaptive control unit 5 includes a voice activity detection unit 51, a residual echo level estimation unit 52, a vibration signal level correction unit 53, a level ratio calculation unit 54, and an adaptive filter learning speed setting unit 55. In Fig. 9, the same parts as in Fig. 8 are designated by the same reference numerals, and their description may be omitted.
[0035] The vibration signal level correction unit 53 receives as input the voice section information of the vibration signal output from the voice section detection unit 51, the vibration signal, and the voice signal output from the echo canceller 20. The vibration signal level correction unit 53 calculates the relative sound pressure level ratio per predetermined unit time between the vibration signal and the voice signal output from the echo canceller 20 during the voice section of the vibration signal. The vibration signal level correction unit 53 also outputs a corrected sound pressure level obtained by correcting the sound pressure level of the vibration signal to a sound pressure level equivalent to the sound pressure level of the voice signal based on the relative sound pressure level ratio. The predetermined unit time is, for example, several milliseconds or several tens of milliseconds.
[0036] The audio signal picked up by the microphone 1 may contain echo components or environmental noise. If the sound pressure level of the vibration signal is corrected to a sound pressure level equivalent to the sound pressure level of the audio signal, a relatively accurate sound pressure level of the audio signal can be obtained that is not affected by the echo components or environmental noise.
[0037] 9 receives as input the voice signal output from the echo canceller 20, the other party's voice signal, and voice section information of the vibration signal. Similar to the voice section detection unit 121, the residual echo level estimation unit 52 detects the voice section of the other party's voice signal by a technique called VAD to generate other party's voice section information, and detects the sound pressure level of the other party's voice signal to generate other party's sound pressure information.
[0038] If the voice section information of the vibration signal indicates the user's non-voice section and the other party's voice section information indicates the voice section of the other party's voice signal, the microphone 1 does not pick up the voice emitted by the user but only the echo, and therefore the voice signal output from the echo canceller 20 contains only the echo component.
[0039] Therefore, when the voice section information of the vibration signal indicates a non-voice section of the user and the other party's voice section information indicates a voice section of the other party's voice signal, the residual echo level estimation unit 52 calculates the relative sound pressure level ratio per predetermined unit time between the other party's sound pressure information and the voice signal output from the echo canceller 20. The predetermined unit time here is also, for example, several milliseconds or several tens of milliseconds. The relative sound pressure level ratio calculated by the residual echo level estimation unit 52 corresponds to the estimated residual echo level. In this way, the residual echo level estimation unit 52 estimates the residual echo level.
[0040] The level ratio calculation unit 54 receives the residual echo level output from the residual echo level estimation unit 52 and the corrected sound pressure level output from the vibration signal level correction unit 53. The level ratio calculation unit 54 divides the corrected sound pressure level by the residual echo level to calculate the relative sound pressure level ratio between the corrected sound pressure level and the residual echo level. The residual echo level estimation unit 52 estimates in advance the residual echo level contained in the audio signal picked up by the microphone 1. The vibration signal level correction unit 53 obtains a corrected sound pressure level corresponding to the sound pressure level of the audio signal based on the vibration signal.
[0041] Therefore, the relative sound pressure level ratio calculated by the level ratio calculation unit 54 is an accurate sound pressure level ratio even when the microphone 1 is picking up environmental noise or when the voice of the user and the voice of the other party overlap. If the relative sound pressure level ratio calculated by the level ratio calculation unit 54 exceeds a predetermined threshold, the voice signal output from the echo canceller 20 contains almost no echo component, which means that the echo component has been canceled by the echo canceller 20. If the relative sound pressure level ratio calculated by the level ratio calculation unit 54 is equal to or less than the predetermined threshold, the voice signal output from the echo canceller 20 contains an echo component, which means that the echo component has not been canceled by the echo canceller 20.
[0042] The adaptive filter learning speed setting unit 55 receives as input the voice segment information output from the voice segment detection unit 51 and the relative sound pressure level ratio output from the level ratio calculation unit 54. If the voice segment information indicates a user's voice segment and the relative sound pressure level ratio output from the level ratio calculation unit 54 exceeds a threshold, the adaptive filter learning speed setting unit 55 generates an adaptive filter control signal for setting the learning speed to active and supplies it to the adaptive filter 6. If the second condition is not satisfied, the adaptive filter learning speed setting unit 55 generates an adaptive filter control signal for setting the learning speed to save and supplies it to the adaptive filter 6.
[0043] If the voice section information does not indicate the user's voice section, there is no voice signal that can serve as a target signal, so the learning speed should be set to save. If the voice section information indicates the user's voice section but the relative sound pressure level ratio is below the threshold, the presence of residual echo components may deteriorate the quality of the converted voice signal, so the learning speed should be set to save.
[0044] 9, as a third specific configuration example of the adaptive control unit 5, the other party's voice activity information generated by the residual echo level estimation unit 52 may be input to the adaptive filter learning rate setting unit 55. In this case, if the third condition is satisfied that the other party's voice activity information indicates a non-voice activity of the other party's voice signal and the voice activity information indicates a user's voice activity, the adaptive filter learning rate setting unit 55 generates an adaptive filter control signal for setting the learning rate to active and supplies the signal to the adaptive filter 6.
[0045] If a fourth condition is satisfied, that is, the other party voice section information indicates the voice section of the other party voice signal, the relative sound pressure level ratio output from the level ratio calculation unit 54 exceeds the threshold value, and the voice section information indicates the user's voice section, the adaptive filter learning speed setting unit 55 generates an adaptive filter control signal for setting the learning speed to active and supplies it to the adaptive filter 6.
[0046] If neither the third nor the fourth condition is satisfied, the adaptive filter learning rate setting unit 55 generates an adaptive filter control signal for setting the learning rate to save and supplies it to the adaptive filter 6 .
[0047] As a more preferable configuration, the adaptive control unit 5 shown in FIG. 9 includes a vibration signal level correction unit 53, and the level ratio calculation unit 54 calculates the relative sound pressure level ratio between the vibration signal level and the residual echo level, using the sound pressure level (corrected sound pressure level) of the vibration signal corrected by the vibration signal level correction unit 53 as the vibration signal level. For simplification, as a fourth specific configuration example of the adaptive control unit 5, the vibration signal level correction unit 53 may be omitted. In this case, the level ratio calculation unit 54 may calculate the level ratio between the vibration signal level indicating the sound pressure level of the vibration signal and the residual echo level. Furthermore, a threshold value for the level ratio between the vibration signal level and the residual echo level, at which the sound pressure level of the vibration signal is estimated to be sufficiently high and the quality of the converted voice signal by the adaptive filter 6 is maintained, may be measured in advance and stored in the storage unit.
[0048] If a fifth condition is met, that is, the voice segment information indicates a user's voice segment and the level ratio calculated by the level ratio calculation unit 54 exceeds a predetermined threshold, the adaptive filter learning speed setting unit 55 generates an adaptive filter control signal for setting the learning speed to active and supplies it to the adaptive filter 6. If the fifth condition is not met, the adaptive filter learning speed setting unit 55 generates an adaptive filter control signal for setting the learning speed to save and supplies it to the adaptive filter 6.
[0049] In FIG. 1, a subtractor 7 supplies the difference between the converted voice signal output from the adaptive filter 6 and the voice signal output from the echo canceller 20 to the adaptive filter 6 as a residual signal.
[0050] 10 shows a specific example of the configuration of adaptive filter 6 using an FIR filter. Adaptive filter 6 includes adaptive coefficient update unit 61, delay devices 621 to 62n, multipliers 630 to 63n, and adders 641 to 64n, where n is a number ranging from several tens to several hundreds. Delay devices 621 to 62n delay each sample of the input digital vibration signal by one clock and output the delayed signal. Multipliers 630 to 63n multiply the sample input to delay device 621 and each sample output from delay devices 621 to 62n by a coefficient, respectively, and output the results.
[0051] The adders 641 to 64n respectively add the outputs of the multipliers 630 and 631, the output of the adder 641 and the multiplier 632, the output of the adder 642 and the multiplier 63, ..., the output of the adder 64(n-1) (not shown) and the multiplier 63n. As a result, the adder 64n outputs a converted audio signal that has been corrected so as to bring the vibration signal output from the A / D converter 4 closer to the audio signal output from the echo canceller 20.
[0052] The subtractor 7 outputs a residual signal which is the difference between the converted voice signal output from the adder 64n and the voice signal output from the echo canceller 20. The adaptive coefficient update unit 61 updates the coefficients by which the multipliers 630 to 63n multiply the input samples so as to reduce the residual signal.
[0053] At this time, when the adaptive filter control signal is high indicating active, adaptive coefficient update unit 61 updates the coefficients supplied to multipliers 630 to 63n in a short time so as to reduce the residual signal.When the adaptive filter control signal is low indicating save, adaptive coefficient update unit 61 either updates the coefficients supplied to multipliers 630 to 63n in a long time so as to reduce the residual signal, or does not update the coefficients.
[0054] When an adaptive filter control signal for setting the learning speed to active is input, the adaptive filter 6 updates the coefficients supplied to the multipliers 630 to 63n in a short time to correct the vibration signal so that it approaches the audio signal. This allows the sound collection device 200 to immediately supply a converted audio signal with good audio quality to the line 11.
[0055] When an adaptive filter control signal for setting the learning speed to save is input, the adaptive filter 6 does not update the coefficients supplied to the multipliers 630 to 63n, or if it updates them, it does not update them immediately but updates them gradually over a long period of time. This allows the sound collection device 200 to supply to the line 11 a converted audio signal whose audio quality is maintained without causing any significant degradation in the audio quality of the converted audio signal.
[0056] The adaptive filter 6 obtains a coefficient that brings the vibration signal closer to the audio signal through learning when any of the first to fifth conditions is satisfied, and outputs a converted audio signal with good audio quality. Therefore, even if a state occurs in which none of the first to fifth conditions is satisfied, the adaptive filter 6 generates a converted audio signal using a coefficient that brings the already obtained vibration signal closer to the audio signal, so it can continue to output a converted audio signal with good audio quality.
[0057] 11A and 11B, a series of operations performed by the sound collection device 200 will be described. The flowcharts shown in Fig. 11A and 11B show operations when the adaptive control unit 5 has the second configuration example shown in Fig. 9.
[0058] 11A, when the power of the sound collection device 200 is turned on and processing starts, the adaptive control unit 12 generates other party voice section information and other party sound pressure information in step S1. In step S2, the adaptive control unit 12 determines whether or not it is a other party voice section based on the other party voice section information. If it is a other party voice section (YES), in step S3, the adaptive control unit 12 supplies an adaptive filter control signal indicating active to the adaptive filter 13. If it is not a other party voice section (NO), in step S4, the adaptive control unit 12 supplies an adaptive filter control signal indicating save to the adaptive filter 13.
[0059] Following step S3, in step S5, the adaptive filter 13 updates the coefficients to be supplied to the multipliers 1330 to 133n in a short time. Following step S4, in step S6, the adaptive filter 13 updates the coefficients to be supplied to the multipliers 1330 to 133n in a long time or does not update them.
[0060] In step S7, the adaptive control unit 5 determines a voice section based on the vibration signal, and in step S8, corrects the sound pressure level of the vibration signal. In parallel with steps S7 and S8, the adaptive control unit 5 estimates a residual echo level in step S9. Subsequently, in step S10, the adaptive control unit 5 calculates a relative sound pressure level ratio between the corrected sound pressure level and the residual echo level.
[0061] In step S11 of Fig. 11B, the adaptive control unit 5 determines whether or not it is a voice section based on the voice section information of the vibration signal. If it is a voice section (YES), the adaptive control unit 5 shifts the process to step S12. If it is not a voice section (NO), the adaptive control unit 5 shifts the process to step S14. In step S12, the adaptive control unit 5 determines whether or not the relative sound pressure level ratio between the corrected sound pressure level and the residual echo level exceeds a threshold. If the relative sound pressure level ratio exceeds the threshold (YES), the adaptive control unit 5 shifts the process to step S13. If the relative sound pressure level ratio does not exceed the threshold (NO), the adaptive control unit 5 shifts the process to step S14.
[0062] In step S13, the adaptive control unit 5 supplies an adaptive filter control signal indicating active to the adaptive filter 6. In step S14, the adaptive control unit 5 supplies an adaptive filter control signal indicating save to the adaptive filter 6. Following step S13, in step S15, the adaptive filter 6 updates the coefficients supplied to the multipliers 630 to 63n in a short time. Following step S14, in step S16, the adaptive filter 6 updates the coefficients supplied to the multipliers 630 to 63n over a long time or does not update them.
[0063] Following step S15 or S16, in step S17, the sound collection device 200 determines whether or not a power-off operation has been performed. If a power-off operation has not been performed (NO), the sound collection device 200 returns to step S1 in Fig. 11A and repeats the processes of steps S1 to S17. If a power-off operation has been performed (YES), the sound collection device 200 ends the process.
[0064] As described above, the sound collection device 200 does not always update the coefficient by which the converted audio signal is multiplied in the adaptive filter 6 so as to reduce the residual signal in a short time. The sound collection device 200 is configured to update the coefficient over a long period of time or not update the coefficient at all when there is a possibility that the quality of the converted audio signal will be deteriorated due to the presence of residual echo components. Therefore, the sound collection device 200 can improve the quality of the audio signal (converted audio signal) based on the vibration signal generated by the vibration sensor 3.
[0065] The present invention is not limited to the above-described embodiment, and various modifications are possible without departing from the spirit of the present invention. In the second and third configuration examples of the adaptive control unit 5, the residual echo level estimating unit 52 generates the other party's voice activity information. The other party's voice activity information used by the adaptive control unit 5 may be generated outside the adaptive control unit 5. The other party's voice activity information generated by the voice activity detecting unit 121 included in the adaptive control unit 12 shown in FIG. 6 may be input to the adaptive control unit 5. Furthermore, although the residual echo level estimating unit 52 generates the other party's sound pressure information, it may be generated outside the adaptive control unit 5. A sound pressure information detecting unit that detects the sound pressure level of the other party's voice signal may be provided within the adaptive control unit 12, and the other party's sound pressure information generated by the sound pressure information detecting unit may be input to the adaptive control unit 5.
[0066] 1, a selector may be provided that selects between the voice signal output from the echo canceller 20 and the converted voice signal output from the adaptive filter 6 and supplies the selected signal to the line 11. An environmental noise analysis unit may be provided that analyzes whether or not environmental noise is superimposed on the voice signal generated by the microphone 1, and the selector may select the voice signal output from the echo canceller 20 if no environmental noise is superimposed, or select the converted voice signal if environmental noise is superimposed.
[0067] 1, the components excluding the microphone 1, the vibration sensor 3, the line 11, and the speaker 16 may be configured by a microcomputer. In this case, a computer program stored in a non-transitory storage medium causes the central processing unit of the microcomputer to execute the above-described processes in the sound collection device 200. The components excluding the microphone 1, the vibration sensor 3, the line 11, and the speaker 16 may be configured by hardware and implemented as an integrated circuit. [Explanation of symbols]
[0068] 1 microphone 2,4 A / D converter 3. Vibration Sensor 5,12 Adaptive control unit 6,13 Adaptive Filter 7,14 Subtractor 15 D / A converter 16 speakers 20 Echo Canceller 52 Residual echo level estimation section 53 Vibration signal level correction section 54 Level ratio calculation section 55 Adaptive filter learning speed setting section 200 Sound collection device
Claims
1. a microphone that generates a first audio signal based on air vibrations; a vibration sensor that generates a vibration signal based on vibrations transmitted to the human body by speech; an echo canceller that suppresses echo components superimposed on the first audio signal by the microphone picking up a sound of a second audio signal transmitted from a communication partner and received via a line, the second audio signal being reproduced by a speaker; an adaptive filter that generates a converted voice signal by multiplying the vibration signal by a coefficient so that the voice signal output from the echo canceller is used as a target signal and the vibration signal approaches the target signal; a subtractor for generating a residual signal that is the difference between the target signal and the converted audio signal; an adaptive control unit that controls the coefficient by which the adaptive filter multiplies the vibration signal so as to update the coefficient so that the residual signal becomes smaller; Equipped with The adaptive control unit a residual echo level estimating unit that estimates a residual echo level remaining in the target signal based on a sound pressure level of the target signal and a sound pressure level of the second audio signal; an adaptive filter learning speed setting unit that controls the adaptive filter to update the coefficients at a first speed if the vibration signal indicates a voice section and the residual echo level is equal to or lower than a predetermined threshold, and that controls the adaptive filter to update the coefficients at a second speed slower than the first speed or not to update the coefficients if the conditions are not satisfied; Equipped with Sound pickup device.
2. a microphone that generates a first audio signal based on air vibrations; a vibration sensor that generates a vibration signal based on vibrations transmitted to the human body by speech; an echo canceller that suppresses echo components superimposed on the first audio signal by the microphone picking up a sound of a second audio signal transmitted from a communication partner and received via a line, the second audio signal being reproduced by a speaker; an adaptive filter that generates a converted voice signal by multiplying the vibration signal by a coefficient so that the voice signal output from the echo canceller is used as a target signal and the vibration signal approaches the target signal; a subtractor for generating a residual signal that is the difference between the target signal and the converted audio signal; an adaptive control unit that controls the coefficient by which the adaptive filter multiplies the vibration signal so as to update the coefficient so that the residual signal becomes smaller; Equipped with The adaptive control unit a residual echo level estimating unit that estimates a residual echo level remaining in the target signal based on a sound pressure level of the target signal and a sound pressure level of the second audio signal; a level ratio calculation unit that calculates a level ratio between a vibration signal level indicating a sound pressure level of the vibration signal and the residual echo level; an adaptive filter learning speed setting unit that controls the adaptive filter to update the coefficients at a first speed if the vibration signal indicates a voice section and the level ratio exceeds a predetermined threshold, and controls the adaptive filter to update the coefficients at a second speed slower than the first speed or not to update the coefficients if the conditions are not satisfied; Equipped with Sound pickup device.
3. the adaptive control unit further includes a vibration signal level correction unit that calculates a relative sound pressure level ratio between the vibration signal and the target signal in a sound section of the vibration signal, and corrects the sound pressure level of the vibration signal to a sound pressure level corresponding to the sound pressure level of the first sound signal based on the relative sound pressure level ratio, The level ratio calculation unit calculates a relative sound pressure level ratio between the vibration signal level and the residual echo level, using the sound pressure level of the vibration signal corrected by the vibration signal level correction unit as the vibration signal level. The sound pickup device according to claim 2 .
Citation Information
Patent Citations
Voice amplifying phone
JP2007060429A
Microphone and sound generation method
JP2007251354A
Speech signal processor
JP2010016564A
Audio signal generation system and method
JP2014502468A