Sound pick-up device, sound pick-up method, and sound pick-up program
The sound collection device enhances audio clarity by using adaptive filtering techniques to distinguish target speech from interfering sounds through electromyogram analysis, ensuring high-quality audio output.
Patent Information
- Application Number
- JP2023216760
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
AI Technical Summary
Existing sound collection devices struggle to accurately distinguish target speech from interfering voices or noise components, leading to erroneous inclusion of non-target sounds in the audio output.
A sound collection device utilizing an adaptive filter that updates filter coefficients based on electromyogram signals and voice signals to differentiate speaking and non-speaking sections, thereby suppressing non-target sounds.
Effectively suppresses the influence of non-target voices and noise components, improving audio quality by accurately controlling the approximation of vibration signals to voice signals during speaking intervals.
Smart Images

Figure 2025099816000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sound collection technology, and particularly to a sound collection device, a sound collection method, and a sound collection program that use a microphone and a vibration sensor.
Background Art
[0002] A sound collection device includes a microphone that generates an audio signal based on air vibration and a vibration sensor that generates a vibration signal corresponding to the audio signal based on bone vibration, thereby obtaining clear audio in a noisy environment (for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The sound collection device converts the vibration signal generated by the vibration sensor into an audio signal in a filtering unit, approximates the vibration signal to the audio signal that is originally transmitted through air, and provides an easy-to-hear audio. However, when the target speaker speaks, if the conversation audio of a person other than the speaker is covered, it will also be erroneously adapted to components other than the target sound. Therefore, it is necessary to accurately grasp the case where the audio of a person other than the speaker is mixed into the microphone. If it does not have a function of detecting the voice components and noise components of a person other than the speaker, when increasing the approximation degree to the microphone audio in the speaking section, there is a risk that the influence of the voice or noise components other than the target sound will be erroneously added to the converted audio.
[0005] The present invention has been made in view of such a situation, and its purpose is to provide a technology for suppressing the influence of voices or noise components other than the target sound.
Means for Solving the Problems
[0006] In order to solve the above problems, a sound collection device according to an aspect of the present invention includes an adaptive filter that performs an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal, and a coefficient update unit that updates the filter coefficient based on a residual signal that is a difference between the voice signal acquired by the microphone and the converted voice signal, and an adaptive control unit that adjusts the frequency of updating the filter coefficient in the coefficient update unit based on an electromyogram signal and a voice signal that reflect whether it is a user's speaking section or a non-speaking section.
[0007] Another aspect of the present invention is a sound collection method. This sound collection method includes a step of performing an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal, a step of updating the filter coefficient based on a residual signal that is a difference between the voice signal acquired by the microphone and the converted voice signal, and a step of adjusting the frequency of updating the filter coefficient based on an electromyogram signal and a voice signal that reflect whether it is a user's speaking section or a non-speaking section.
[0008] Another aspect of the present invention is a sound collection program. This sound collection program causes a computer to perform a step of performing an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal, a step of updating the filter coefficient based on a residual signal that is a difference between the voice signal acquired by the microphone and the converted voice signal, and a step of adjusting the frequency of updating the filter coefficient based on an electromyogram signal and a voice signal that reflect whether it is a user's speaking section or a non-speaking section.
[0009] Note that any combination of the above components, and those obtained by converting the expression of the present invention among a method, an apparatus, a system, a recording medium, a computer program, etc. are also effective as aspects of the present invention.
Advantages of the Invention
[0010] According to the present invention, it is possible to suppress the influence of voices or noise components other than the target sound.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Best Mode for Carrying Out the Invention
[0012] (Example 1) Before specifically describing the present invention, an overview will be first described. This example relates to a sound collection device. The sound collection device includes a microphone that generates an audio signal based on air vibration, a vibration sensor that generates a vibration signal based on vibration transmitted to the human body, and an adaptive filter that multiplies a coefficient to the vibration signal to generate a converted audio signal in order to correct the vibration signal so as to approximate it to the audio signal. Further, when the sound collection device detects the utterance situation of the target sound by electromyogram analysis and detects the mixing of audio components other than the target sound, the sound collection device stops the update (learning) of the filter coefficient in the adaptive filter and suppresses the converted audio signal from being approximated to the audio components other than the target sound.
[0013] The mixing of audio components other than the target sound is determined using a threshold value that serves as a reference for the correlation value between the vibration signal and the audio signal. The reference threshold value is made variable in conjunction with the correlation between the vibration signal and the audio signal in a non-audio section, that is, under background noise. Thereby, the audio quality of the converted audio signal is improved by detecting a higher-precision correlation. Note that the converted audio signal is a digital signal related to audio obtained by performing analog-to-digital conversion on the vibration signal acquired by the vibration sensor and further performing an operation using a predetermined filter coefficient, that is, an audio signal converted based on the vibration signal.
[0014] FIG. 1 shows the configuration of a sound collection device 100 according to Example 1. The sound collection device 100 includes a microphone 10, a first AD converter (Analog to Digital converter) 12, a vibration sensor 14, a second AD converter 16, an adaptive filter 18, a subtractor 20, a DA converter (Digital to Analog converter) 22, a correlation determination unit 32, an adaptive control unit 34, an electromyogram sensor 90, and an electromyogram analysis unit 92.
[0015] The microphone 10 converts air vibrations into an audio signal 200. Since the microphone 10 can accurately acquire the audio signal 200, the audio signal 200 acquired by the microphone 10 is relatively close to the audio signal 200 that a person perceives through the ear. Therefore, by using the audio signal 200 acquired by the microphone 10 as the target value of the vibration signal 202 described below, it is possible to maintain the audio quality of the converted audio signal 212 output by the adaptive filter 18 at a high level.
[0016] The first AD converter 12 outputs the audio signal 200 converted into a digital signal (hereinafter, this is also referred to as "audio signal 200") to the subtracter 20 and the correlation determination unit 32 by performing analog-digital conversion on the audio signal 200 acquired by the microphone 10.
[0017] The vibration sensor 14 acquires the speech signal emitted by a person mainly as a vibration signal 202 transmitted to the human body. The vibration sensor 14 may be a microphone that detects sound or vibration, or may be a camera (visual microphone) that detects vibration from pixel variations or color fluctuations between frames of an image and acquires an audio signal. Instead of the vibration sensor 14, a receiving device embedded in the body or a microphone in direct contact with the human body may be used, or a device such as a laser displacement meter that measures the vibration of the human body surface without contact with the human body may be used.
[0018] The second AD converter 16 outputs the vibration signal 202 converted into a digital signal (hereinafter, this is also referred to as "vibration signal 202") to the adaptive filter 18 and the correlation determination unit 32 by performing analog-digital conversion on the vibration signal 202 acquired by the vibration sensor 14.
[0019] The vibration signal 202 has the characteristics that, compared with the audio signal 200, its frequency band is narrow and it has a large bias in the frequency band. Due to this characteristic, the vibration signal 202 becomes a very muffled voice and sounds significantly different from the original audio signal 200. To improve this, in this embodiment, the audio signal 200 acquired at the same time is used as a reference signal, and the adaptive filter 18 is used to make the vibration signal 202 approach the audio signal 200.
[0020] The adaptive filter 18 receives the vibration signal 202 from the second AD converter 16 and the residual signal 210 from the subtractor 20. Based on the residual signal 210, the adaptive filter 18 performs an operation on the vibration signal 202 using the filter coefficient and outputs the operation result as a converted audio signal 212.
[0021] Figure 2 shows the configuration of the adaptive filter 18 according to Embodiment 1. The adaptive filter 18 includes a first delay element 50a to an (N - 1)th delay element 50n - 1 collectively referred to as a delay element 50, a first multiplier 52a to an Nth multiplier 52n collectively referred to as a multiplier 52, a first adder 54a to an (N - 1)th adder 54n - 1 collectively referred to as an adder 54, and a coefficient update unit 60.
[0022] The delay element 50 delays the vibration signal 202 and sequentially outputs it. The multiplier 52 multiplies the vibration signal 202 by the filter coefficient. The adder 54 sequentially adds the multiplication results at the multiplier 52. The addition result of the (N - 1)th adder 54n - 1 corresponds to the aforementioned operation result. The subtractor 20 calculates a residual signal 210, which is the difference value between the audio signal 200, which is the reference signal, and the converted audio signal 212, which is the operation result, and outputs it to the adaptive filter 18.
[0023] The coefficient update unit 60 updates the filter coefficient based on the audio signal 200 and the operation result by, for example, executing the LMS (Least Mean Square) algorithm as follows.
Equation
[0024] The DA converter 22 outputs a converted voice signal 212 (hereinafter referred to as the "output signal 214") converted into an analog signal by performing digital-to-analog conversion on the converted voice signal 212 output from the adaptive filter 18.
[0025] Here, when a voice component other than the target voice is mixed in the voice signal 200 acquired by the microphone 10, in the adaptive filter 18, an action of approximating to the voice of others may erroneously occur. The voice component other than the target voice is, for example, the voice of a person other than the target person, the voice reproduced by a handset speaker (not shown) at the call destination, and the like. Therefore, in order to improve the accuracy of the process for bringing the vibration signal 202 in the adaptive filter 18 closer to the voice signal 200, it is required to update (learn) the filter coefficient when no voice component other than the target voice is mixed in the voice signal 200 in the speaking section. As a preferable condition for updating (learning) the filter coefficient, it is desirable that it is the speaking section of the user. Note that the speaking section is a time zone in which the speech of the target person is detected.
[0026] In this embodiment, in order to detect whether it is the speaking interval or the non-speaking interval of the target person (user), the electromyogram sensor 90 attached near the face (mouth) of the target person measures the electromyogram generated by the muscle activity associated with the movement near the mouth of the target person. Since the electromyogram is generated by muscle contraction along with the movement of the mouth and throat, it can be measured regardless of whether it is a vowel or a consonant. The electromyogram sensor 90 outputs the measured electromyogram as an electromyogram signal 250 to the electromyogram analysis unit 92. Note that the electromyogram sensor 90 is a measuring device for electromyogram signals that has an electromyogram signal measuring unit that is attached to the muscle to be measured and measures an electromyogram signal indicating the muscle activity waveform, and an amplifying unit that amplifies the electromyogram signal.
[0027] The electromyogram analysis unit 92 identifies the speaking interval based on the electromyogram signal 250. For example, the electromyogram analysis unit 92 detects the presence or absence of muscle contraction in the mouth and throat based on the displacement amount of the electromyogram per unit time, and identifies the speaking interval when the displacement amount exceeds a predetermined threshold value. That is, when the target person speaks, an electromyogram associated with the muscle contraction of the mouth movement and the throat movement is generated, so the electromyogram analysis unit 92 measures the electromyogram signal 250 generated by the mouth movement and the throat movement. As a method for detecting the speaking interval from the electromyogram signal 250, for example, a method such as the literature: "Investigation of Recognition Parameters for Consonant Recognition Using Electromyogram Signals" (The 73rd National Conference of the Information Processing Society of Japan, 6P-4) may be utilized. The electromyogram analysis unit 92 outputs the electromyogram analysis result determining whether it is the speaking interval or the non-speaking interval of the target person as the speaking interval information 204 to the correlation determination unit 32. In order to improve the determination accuracy of the speaking interval, the electromyogram analysis unit 92 may use the voice signal 200 obtained by the vibration sensor 14 as a substitute for the voice signal 200, and supplementarily use the voice interval determination result by the voice signal 200. The effect of preventing false detection due to mouth movement without speech such as yawning can be obtained.
[0028] The human vocal cords generate the types of voice signals 200, and various voice signals 200 are generated by passing through the throat and the mouth, which are vocal tract tubes. However, there is a correlation between the vibration signal 202 associated with speech acquired at a predetermined human body part and the voice signal 200 obtained by the microphone 10 acquiring air vibrations at a predetermined position, which involves a spatial delay during propagation in space and transmission characteristics due to human body components (bones, muscles, skin, etc.).
Number
[0029] When compensating for a predetermined spatial delay (z1) and the gain correction amount with respect to the reference vibration signal V(t) to align the time axis and sound pressure level of the vibration signal 202 and the voice signal 200, it is expected that the signal values will be similar, although it also depends on the type of speech. If it is only the speech signal of the subject, by performing a cross-correlation calculation between the vibration signal 202 and the voice signal 200, a point with a high correlation is observed in the vicinity of the time difference z1 + z2. The general formula for cross-correlation is shown as follows.
Number
[0030] In Equation (3), it is observed that the point with the highest correlation has the maximum value Cmax of the correlation value C. Here, τ corresponds to the time difference (z1 + z2) between the speech signal sequence S(t) and the vibration signal sequence V(t) with the same signal source. By calculating Equation (3), the maximum value Cmax, which is the correlation value C of the point with the highest correlation between the vibration signal 202 and the speech signal 200, determines the number of samples T according to the sampling frequency from a predetermined analysis time width, gives a shift amount of about several samples to more than a dozen samples before and after centered on τ, and performs a convolution operation, and is observed as the numerical value of the maximum value of the sum of products.
[0031] The cross-correlation calculation may be performed as follows.
Number
[0032] On the other hand, when the voices of people other than the target speaker or noise components are mixed in, the correlation between the vibration signal 202 and the speech signal 200 decreases, so conversely, a numerical value with low correlation is observed for the correlation value. Therefore, the correlation determination unit 32 determines whether there is a multiple speech state in the time period when the target speaker speaks based on whether the correlation value is equal to or greater than a threshold value.
[0033] The correlation determination unit 32 detects a multiple talk state or a situation where the noise environment has deteriorated based on at least one of the maximum value Cmax of the correlation value C calculated by the above-described formula (3) or the minimum value Cmin of the correlation value C calculated by the formula (4). FIG. 3 shows the configuration of the correlation determination unit 32 according to the first embodiment. The correlation determination unit 32 includes a threshold setting unit 70, a correlation value calculation unit 72, and a state determination unit 74. The threshold setting unit 70 receives the utterance section information 204 indicating whether it is an utterance section or a non-utterance section from the myoelectric potential analysis unit 92. The threshold setting unit 70 sets, as the first threshold 216, the value that is larger than a predetermined value from the reference correlation value, based on the correlation value in the case of ambient noise without an external sound source and in a non-utterance section, and sets, as the second threshold 218, the value that is smaller than a predetermined value from the reference correlation value.
[0034] The correlation value calculation unit 72 calculates the correlation value 220 between the voice signal 200 and the vibration signal 202 by executing the above-described formula (3) or formula (4). At this time, in order to improve the accuracy of the correlation value 220, the sound pressure levels of the vibration signal 202 and the voice signal 200 may be made uniform in advance. For example, a gain adjustment unit may be provided after the first AD converter 12 in FIG. 2 so that the sound pressure levels of the voice signal 200 during speech are approximately the same, or adjustment may be made in the correlation value calculation unit 72.
[0035] The state determination unit 74 receives the first threshold 216 or the second threshold 218 from the threshold setting unit 70 and receives the correlation value 220 from the correlation value calculation unit 72. The state determination unit 74 determines the states of the subject's speech, multiple talk, other person's speech, and silence by comparing the first threshold 216 or the second threshold 218 with the correlation value 220. Further, the state determination unit 74 determines the control content of the coefficient update unit 60 according to the determined state. The subject's speech is a state where only the target person is speaking, the multiple talk is a double talk state or a state where external noise is mixed in, the other person's speech is a state where only a person other than the target person is speaking, and the silence is a state where there is no speech.
[0036] Figures 4(a) and 4(b) show the data structure of the table held in the state determination unit 74 according to the first embodiment. When the correlation value 220 is greater than or equal to the first threshold value 216 during the speech section, the state determination unit 74 determines that it is the target person's speech and determines the activation of the adaptation of the coefficient update unit 60. When the correlation value 220 is less than the first threshold value 216 during the speech section, the state determination unit 74 determines that it is a multiple speech and determines the stop of the adaptation of the coefficient update unit 60. When the correlation value 220 is greater than or equal to the second threshold value 218 during the non-speech section, the state determination unit 74 determines that it is silence and determines the save of the adaptation of the coefficient update unit 60. When the correlation value 220 is less than the second threshold value 218 during the non-speech section, the state determination unit 74 determines that it is the speech of others and determines the stop of the adaptation of the coefficient update unit 60.
[0037] Here, in the case of adaptation activation, "α" in Equation (1) is set to the first value. The first value is determined in advance. In the case of adaptation save, "α" in Equation (1) is set to the second value. The second value is smaller than the first value. In the case of adaptation stop, the update of the filter coefficient in the coefficient update unit 60 is stopped. Return to FIG. 3.
[0038] The vibration signal 202 indicates the vibration of the human body during speech, but compared with the audio signal 200 picked up by the microphone 10, the noise level tends to be high during ambient noise. When determining the correlation between the vibration signal 202 and the audio signal 200, the first threshold value 216 or the second threshold value 218 may be set based on the correlation value 220 during ambient noise without an external sound source and during the non-speech section. The transfer characteristics shown in Equation (2) have a varying range depending on the type of speech, and even when only the speaker is speaking, the correlation value 220 may fluctuate accordingly. In particular, due to the positional relationship between the part of the human body used as the measurement point to pick up the vibration signal 202 during speech and the vocal cords, throat, and mouth where the sound wave is generated, the correlation value 220 during speech may decrease.
[0039] However, since the correlation value 220 obtained during the ambient noise and non-speaking interval is calculated by a noise component that has no correlation between the vibration signal 202 and the voice signal 200, it will not fall below this correlation value 220 during speech. Therefore, the threshold setting unit 70 detects that it is a non-speaking interval based on the signal indicating whether it is a speaking interval or a non-speaking interval, and further determines, for the ambient noise state, that the correlation value 220 changes within a predetermined time constant width as a determination material.
[0040] The threshold setting unit 70 acquires, from the correlation value calculation unit 72, the correlation values 224 for a predetermined number of times calculated during the ambient noise and non-speaking interval. When the fluctuation range of the acquired correlation values 224 for the predetermined number of times is within the allowable range (within several times a predetermined value), it calculates a reference correlation value by averaging the correlation values 224 for the predetermined number of times acquired from the correlation value calculation unit 72, and determines the larger of the reference correlation value plus a predetermined value as the first threshold 216, and the smaller of the reference correlation value minus a predetermined value as the second threshold 218.
[0041] FIG. 5(a), FIG. 5(b), and FIG. 5(c) show the time changes of the sound pressure level, the electromyogram analysis result, and the correlation value 220 in the microphone 10 and the vibration sensor 14 according to the first embodiment. FIG. 5(a) shows the time changes of the vibration signal 202 and the voice signal 200. The horizontal axis in FIG. 5(a) represents time, and the vertical axis represents the sound pressure level. FIG. 5(b) shows the time change of the speech interval information 204, which is the electromyogram analysis result in the electromyogram analysis unit 92. The horizontal axis in FIG. 5(b) represents time, and the vertical axis represents the value of the speech interval information 204. High-level speech interval information 204 indicates "speech", and low-level speech interval information 204 indicates "silence (non-speaking)". FIG. 5(c) shows the time change of the correlation value 220 calculated in the correlation value calculation unit 72. The horizontal axis in FIG. 5(c) represents time, and the vertical axis represents the correlation value 220.
[0042] Since period L1 is the state of the subject's speech, in period L1, both the vibration signal 202 and the audio signal 200 have a drastically fluctuating sound pressure level. Since the vibration signal 202 and the audio signal 200 have a similar tendency in the state of the subject's speech, the correlation value 220 also increases. In period L2, since it is a silent state, in period L2, the average ambient noise level of the vibration signal 202 and the average ambient noise level of the audio signal 200 are obtained. Although these are within a certain correlation, since they are noise components that are uncorrelated with each other, the correlation value 220 becomes small.
[0043] Period L3 is the state of the subject's speech, similar to period L1. Therefore, the vibration signal 202, the audio signal 200, and the correlation value 220 in period L3 have the same tendency as those in period L1. Period L4 is a silent state, similar to period L2. Therefore, the vibration signal 202, the audio signal 200, and the correlation value 220 in period L4 have the same tendency as those in period L2. Since period L5 is the state of another person's speech, in period L5, the fluctuation of the sound pressure level of the vibration signal 202 is small, but the sound pressure level of the audio signal 200 fluctuates drastically. Therefore, the correlation value 220 in period L5 significantly decreases. The state of another person's speech can also be said to be a state where external noise is mixed in.
[0044] Period L6 is the state of multiple speeches. Since the vibration signal 202 in period L6 only contains the speech information of the subject himself / herself, the part with intense movement on the vertical axis becomes the speech signal of the target speaker. On the other hand, the audio signal 200 in period L6 contains the speech information of the subject himself / herself and the speech information of others. Therefore, the correlation value 220 in period L6 decreases. Since the speech information of the subject is included during multiple speeches, the correlation value 220 is higher than during another person's speech. Period L7 is the state of another person's speech, similar to period L5. Therefore, the vibration signal 202, the audio signal 200, and the correlation value 220 in period L7 have the same tendency as those in period L5.
[0045] The speaking intervals are period L1, period L3, and period L6. In the speaking intervals, in order to distinguish the state of the target person's speech and the state of multiple speeches, the first threshold value 216 shown in FIG. 5(b) is set. On the other hand, the non-speaking intervals are period L2, period L4, period L5, and period L7. In the non-speaking intervals, in order to distinguish the state of silence and the state of others' speech, the second threshold value 218 shown in FIG. 5(b) is set. Return to FIG. 1.
[0046] The adaptation control unit 34 receives the determination result (correlation information 206) in the correlation determination unit 32. If the determination result is adaptation active, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to set "α" to the first value. If the determination result is adaptation save, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to set "α" to the second value. If the determination result is stop, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to stop updating the filter coefficient. That is, the adaptation control unit 34 adjusts the degree of updating of the filter coefficient in the coefficient update unit 60 based on the vibration signal 202 and the audio signal 200.
[0047] The above configuration can be realized hardware-wise by the CPU, memory, and other LSIs of any computer, and software-wise by a program loaded into the memory, etc. Here, however, the functional blocks realized by their cooperation are depicted. Therefore, it is understood by those skilled in the art that these functional blocks can be realized in various forms by hardware only, software only, or a combination thereof.
[0048] The operation of the sound collection device 100 configured as described above will be described. FIG. 6 is a flowchart showing an update procedure for a first threshold value 216 and a second threshold value 218 by the sound collection device 100 according to the first embodiment. This shows an operation of setting the first threshold value 216 and the second threshold value 218 with respect to a correlation value 220 between a vibration signal 202 and an audio signal 200, which is a criterion for determining whether it is a state of multiple utterances. This operation is performed in a state where there is no utterance and no environmental noise, and the first threshold value 216 and the second threshold value 218 may be set by executing a series of operations in advance before the utterance, or the above state may be determined during the utterance to set the first threshold value 216 and the second threshold value 218.
[0049] The electromyogram sensor 90 acquires an electromyogram as an electromyogram signal 250 (S10). The measurement period of the electromyogram is appropriately several tens of msec in accordance with the processing unit of the sound. This is because, in order to determine the sound section by electromyogram analysis, it is necessary to have the same time resolution as the sound processing in view of the transition of the amount of change per unit time of the electromyogram.
[0050] The electromyogram analysis unit 92 determines whether it is a non-utterance section by performing an electromyogram analysis process on the electromyogram signal 250 acquired by the electromyogram sensor 90 (S12). The electromyogram analysis unit 92 calculates, for example, the average value of the electromyogram at a predetermined time, and estimates that it is an utterance section because the muscle contraction movement is active when the amount of displacement over time exceeds a predetermined threshold value. If it is a non-utterance section (Y in S12), the correlation determination unit 32 acquires the audio signal 200 by the microphone 10 and the audio signal 200 after digital-analog conversion by the first AD converter 12, and the vibration sensor 14 acquires the vibration signal 202 and the vibration signal 202 after analog-digital conversion by the second AD converter 16 (S14). The correlation value calculation unit 72 of the correlation determination unit 32 calculates the correlation value 220 based on the audio signal 200 and the vibration signal 202 a predetermined number of times according to a predetermined analysis time width and sampling frequency (S16).
[0051] The threshold setting unit 70 of the correlation determination unit 32 acquires the correlation values 224 for a predetermined number of times calculated in step 16 from the correlation value calculation unit 72, and determines whether the variation range of the correlation values is within the allowable range (S18). When the variation range of the correlation values is within the allowable range (Y in S18), that is, when the background noise continues, the threshold setting unit 70 updates the first threshold value 216 and the second threshold value 218 based on the correlation values 224 for a predetermined number of times acquired from the correlation value calculation unit 72 (S20). When the variation range of the correlation values is not within the allowable range (N in S18), step 20 is skipped. When it is not a non-speaking interval (N in S12), steps 14 to 20 are skipped. When the operation has not ended (N in S22), the process returns to step 10. When the operation has ended (Y in S22), the process is terminated.
[0052] FIG. 7 is a flowchart showing a processing procedure by the sound collection device 100 according to the first embodiment. Step 50 is the same as the operation shown in FIG. 6, and here shows a case where an interval in which background noise continues during use is detected, and the first threshold value 216 and the second threshold value 218 are always updated. When setting the first threshold value 216 and the second threshold value 218 in advance in a situation where background noise continues before an actual conversation starts, step 50 may be omitted.
[0053] The myoelectric potential sensor 90 acquires a myoelectric potential signal 250 (S52). The myoelectric potential analysis unit 92 determines whether it is a speaking interval by performing a myoelectric potential analysis process on the myoelectric potential signal 250 acquired by the myoelectric potential sensor 90 (S54). When it is a speaking interval (Y in S54), the correlation determination unit 32 acquires the audio signal 200 acquired by the microphone 10 and the audio signal 200 digitally-analog converted by the first AD converter 12, and the vibration signal 202 acquired by the vibration sensor 14 and the vibration signal 202 analog-digital converted by the second AD converter 16 (S56), and the correlation value calculation unit 72 of the correlation determination unit 32 calculates a correlation value 220 based on the audio signal 200 and the vibration signal 202 (S58).
[0054] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the first threshold value 216 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the first threshold value 216 (S60). If the correlation value 220 is not greater than or equal to the first threshold value 216 (N in S60), the state determination unit 74 of the correlation determination unit 32 estimates multiple speech (S62) and determines an adaptive stop (S64). If the correlation value 220 is greater than or equal to the first threshold value 216 (Y in S60), the state determination unit 74 of the correlation determination unit 32 estimates the target person's speech (S66) and determines an adaptive active (S68).
[0055] When it is not the speech section (N in S54), the correlation determination unit 32 acquires the audio signal 200 by the microphone 10, and acquires the audio signal 200 after digital - analog conversion by the first AD converter 12, and the vibration signal 202 acquired by the vibration sensor 14 and the vibration signal 202 after analog - digital conversion by the second AD converter 16 (S69). The correlation value calculation unit 72 of the correlation determination unit 32 calculates a correlation value 220 based on the audio signal 200 and the vibration signal 202 (S70).
[0056] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the second threshold value 218 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the second threshold value 218 (S72). If the correlation value 220 is not greater than or equal to the second threshold value 218 (N in S72), the state determination unit 74 of the correlation determination unit 32 estimates other person's speech (S74) and determines an adaptive stop (S76). If the correlation value 220 is greater than or equal to the second threshold value 218 (Y in S72), the state determination unit 74 of the correlation determination unit 32 estimates silence (S78) and determines an adaptive save (S80).
[0057] The adaptive filter 18 executes adaptive filter processing according to the control content determined by the state determination unit 74 of the correlation determination unit 32 (S82). If the operation has not ended (N in S84), it returns to step 50. If the operation has ended (Y in S84), the process is terminated.
[0058] According to this embodiment, by approximating the vibration signal to the characteristics of the voice signal, the influence of ambient noise components can be suppressed. Further, the presence of voices other than the speaker's voice and sudden noises is sequentially detected from the correlation between the vibration signal and the voice signal, and the degree of approximation to the voice signal is controlled based on the detection result, so that the influence of voices or noise components other than the target voice can be suppressed. Further, since the influence of voices or noise components other than the target voice is suppressed, the mixing of voice components of people other than the speaker is avoided and the voice quality of the speaker can be improved. Further, by detecting the movement of the face (mouth) of the subject by analyzing the amount of change in the myoelectric potential measured by the myoelectric potential sensor installed near the face (mouth) of the subject, the speaking interval of the subject can be determined with high accuracy. Further, since the speaking interval of the subject is determined with high accuracy, the control accuracy of the filter coefficient can be improved.
[0059] (Example 2) Next, Example 2 will be described. Example 2 relates to the sound collection device 100 in the same manner as Example 1. The sound collection device 100 according to Example 1 adjusts any one of adaptive active, adaptive stop, and adaptive save, that is, the degree of update of the filter coefficient, based on the myoelectric potential analysis result and the correlation value 220 between the vibration signal 202 and the voice signal 200. On the other hand, the sound collection device 100 according to Example 2 adjusts the frequency of update of the filter coefficient based on the myoelectric potential analysis result and the correlation value 220 between the vibration signal 202 and the voice signal 200. The sound collection device 100, the adaptive filter 18, and the correlation determination unit 32 according to Example 2 are of the same type as those in FIGS. 1 to 3. Here, the description will focus on the differences from Example 1.
[0060] The state determination unit 74 receives the first threshold value 216 or the second threshold value 218 from the threshold setting unit 70 and receives the correlation value 220 from the correlation value calculation unit 72. The state determination unit 74 determines the state of the subject speaking, multiple speaking, other person speaking, or silence by comparing the first threshold value 216 or the second threshold value 218 with the correlation value 220. Further, the state determination unit 74 determines the control content of the coefficient update unit 60 according to the determined state.
[0061] Figs. 8(a) and 8(b) show the data structure of the table held in the state determination unit 74 according to the second embodiment. When the correlation value 220 is greater than or equal to the first threshold value 216 in the speech section, the state determination unit 74 determines that it is the target person's speech and determines the update based on the first frequency of the filter coefficient in the coefficient update unit 60. When the correlation value 220 is less than the first threshold value 216 in the speech section, the state determination unit 74 determines that it is a multiple speech and determines the update based on the second frequency of the filter coefficient in the coefficient update unit 60. When the correlation value 220 is greater than or equal to the second threshold value 218 in the non-speech section, the state determination unit 74 determines that it is silent and determines the update based on the third frequency of the filter coefficient in the coefficient update unit 60. When the correlation value 220 is less than the second threshold value 218 in the non-speech section, the state determination unit 74 determines that it is the speech of another person and determines the update based on the second frequency of the filter coefficient in the coefficient update unit 60.
[0062] Here, the second frequency is made smaller than the first frequency, and the third frequency is made smaller than the first frequency and larger than the second frequency. For example, the first frequency is every time, the second frequency is once every 128 times, and the third frequency is between once every 2 times and once every 64 times. The values of each frequency are not limited to these. Return to Fig. 3.
[0063] The adaptation control unit 34 receives the determination result (correlation information 206) in the correlation determination unit 32. If the determination result is an update based on the first frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient at the first frequency. If the determination result is an update based on the second frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient at the second frequency. If the determination result is an update based on the third frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient at the third frequency. That is, the adaptation control unit 34 adjusts the update frequency of the filter coefficient in the coefficient update unit 60 based on the vibration signal 202 and the audio signal 200.
[0064] FIG. 9 is a flowchart showing a processing procedure by the sound collection device 100 according to the second embodiment. Step 100 is the same as the operation shown in FIG. 6, and here it shows a case where a section in which ambient noise continues during use is detected and the first threshold value 216 and the second threshold value 218 are always updated. If the first threshold value 216 and the second threshold value 218 are set in advance in a situation where ambient noise continues before an actual conversation starts, step 100 may be omitted.
[0065] The myoelectric potential sensor 90 acquires a myoelectric potential signal 250 (S102). The myoelectric potential analysis unit 92 determines whether it is a speech section by performing a myoelectric potential analysis process on the myoelectric potential signal 250 acquired by the myoelectric potential sensor 90 (S104). If it is a speech section (Y in S104), the correlation determination unit 32 acquires the audio signal 200 acquired by the microphone 10 and the audio signal 200 after digital - analog conversion by the first AD converter 12, and the vibration signal 202 acquired by the vibration sensor 14 and the vibration signal 202 after analog - digital conversion by the second AD converter 16 (S106), and the correlation value calculation unit 72 of the correlation determination unit 32 calculates a correlation value 220 based on the audio signal 200 and the vibration signal 202 (S108).
[0066] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the first threshold value 216 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the first threshold value 216 (S110). If the correlation value 220 is not greater than or equal to the first threshold value 216 (N in S110), the state determination unit 74 of the correlation determination unit 32 estimates multiple speech (S112) and determines an update at the second frequency (S114). If the correlation value 220 is greater than or equal to the first threshold value 216 (Y in S110), the state determination unit 74 of the correlation determination unit 32 estimates the target person's speech (S116) and determines an update at the first frequency (S118).
[0067] When it is not the speaking period (N in S104), the correlation determination unit 32 acquires the audio signal 200 by the microphone 10, and acquires the audio signal 200 that has been digitally-analog converted by the first AD converter 12 and the vibration signal 202 that has been analog-digital converted by the second AD converter 16 (S119). The correlation value calculation unit 72 of the correlation determination unit 32 calculates the correlation value 220 (S120).
[0068] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the second threshold value 218 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the second threshold value 218 (S122). When the correlation value 220 is not greater than or equal to the second threshold value 218 (N in S122), the state determination unit 74 of the correlation determination unit 32 estimates that someone else is speaking (S124) and determines an update at the second frequency (S126). When the correlation value 220 is greater than or equal to the second threshold value 218 (Y in S122), the state determination unit 74 of the correlation determination unit 32 estimates silence (S128) and determines an update at the third frequency (S130).
[0069] The adaptive filter 18 executes adaptive filter processing according to the control content determined by the state determination unit 74 of the correlation determination unit 32 (S132). When the operation has not ended (N in S134), the process returns to step 100. When the operation has ended (Y in S134), the process is terminated.
[0070] According to this embodiment, since the update frequency of the filter coefficient is adjusted based on the correlation value between the vibration signal and the voice signal, the influence of voices or noise components other than the target sound can be suppressed. Further, when the correlation value is equal to or greater than the first threshold value in the speaking interval, the filter coefficient is updated at the first frequency equal to or higher than the second frequency, so that the filter coefficient can be updated to be closer to the voice signal. Further, when the correlation value is smaller than the first threshold value in the speaking interval, the filter coefficient is updated at the second frequency smaller than the first frequency, so that the influence of voices or noise components other than the target sound can be suppressed. Further, when the correlation value is equal to or greater than the second threshold value in the non-speaking interval, the filter coefficient is updated at the third frequency smaller than the first frequency, so that the influence of the noise component can be suppressed. Further, when the correlation value is smaller than the second threshold value in the non-speaking interval, the filter coefficient is updated at the second frequency smaller than the first frequency, so that the influence of voices or noise components other than the target sound can be suppressed. Further, by imaging the face of the subject and detecting the movement of the face (mouth) of the subject by electromyogram analysis, the speaking interval of the subject can be determined with high accuracy. Further, since the speaking interval of the subject is determined with high accuracy, the control accuracy of the filter coefficient can be improved.
[0071] (Example 3) Next, Example 3 will be described. Example 3 relates to the sound collection device 100 as before. The sound collection device 100 according to Example 2 adjusts the update frequency of the filter coefficient based on the electromyogram analysis result and the correlation value 220 between the vibration signal 202 and the voice signal 200. On the other hand, the sound collection device 100 according to Example 3 adjusts the update frequency of the filter coefficient based on the electromyogram analysis result and the phoneme analysis result (voice recognition result) of the voice signal 200. Here, the description will focus on the differences from before.
[0072] FIG. 10 shows the configuration of the sound collection device 100 according to Example 3. The sound collection device 100 includes a microphone 10, a first AD converter 12, a vibration sensor 14, a second AD converter 16, an adaptive filter 18, a subtractor 20, a DA converter 22, an adaptive control unit 34, an electromyogram sensor 90, an electromyogram analysis unit 92, and a determination unit 94.
[0073] The microphone 10, the first AD converter 12, the vibration sensor 14, the second AD converter 16, the adaptive filter 18, the subtractor 20, the DA converter 22, and the adaptive control unit 34 are the same as before. Further, a determination unit 94 is included instead of the correlation determination unit 32 described above. The first AD converter 12 outputs the audio signal 200 to the subtractor 20 and the determination unit 94, and the second AD converter 16 outputs the vibration signal 202 to the adaptive filter 18.
[0074] The electromyogram sensor 90 outputs the measured electromyogram as an electromyogram signal 250 to the electromyogram analysis unit 92. The electromyogram analysis unit 92 executes electromyogram analysis processing on the electromyogram signal 250 to detect the movement of the subject's mouth from the electromyogram signal 250 and identify the speaking section. Further, the electromyogram analysis unit 92 identifies the characters pronounced by the subject from the electromyogram waveform based on the detected movement of the subject's mouth and throat. The characters pronounced by the subject are, for example, "a", "i",... The electromyogram analysis unit 92 outputs the speaking section information 204 indicating whether it is the speaking section or the non-speaking section of the subject, and the detected character information 205 in which the characters pronounced by the subject are detected from the electromyogram waveform based on the movement of the subject's mouth and throat to the determination unit 94.
[0075] The determination unit 94 receives the audio signal 200 from the first AD converter 12 and also receives the speaking section information 204 and the detected character information 205 from the electromyogram analysis unit 92. The determination unit 94 classifies the current state into any one of the states of the subject speaking, the silent state, the state of others speaking, and the state of multiple speaking based on the audio signal 200, the speaking section information 204, and the detected character information 205, and determines the update frequency according to the classified state. The determination unit 94 outputs the determined update frequency as a determination signal 252 to the adaptive control unit 34. The adaptive control unit 34 receives the determination result (determination signal 252) in the determination unit 94. The adaptive control unit 34 adjusts the update frequency of the filter coefficient in the coefficient update unit 60 of the adaptive filter 18 based on the determination result.
[0076] FIG. 11 shows the configuration of the determination unit 94 according to Embodiment 3. The determination unit 94 includes a state determination unit 74, a phoneme analysis unit 96, and a character determination unit 97. The phoneme analysis unit 96 receives the voice signal 200 from the first AD converter 12. The phoneme analysis unit 96 identifies the characters included in the voice signal 200 by performing phoneme analysis processing (voice recognition processing) on the voice signal 200. The characters included in the voice signal 200 are, for example, "a", "i",... Since known techniques may be used for the phoneme analysis processing (voice recognition processing), the description is omitted here. The phoneme analysis unit 96 outputs the characters included in the voice signal 200 as a character signal 251 to the character determination unit 97.
[0077] The character determination unit 97 receives the detected character information 205 from the myoelectric potential analysis unit 92 and receives the character signal 251 from the phoneme analysis unit 96. The character determination unit 97 compares the characters included in the detected character information 205 with the characters included in the character signal 251. The character determination unit 97 determines a section where both characters match as "match", a section where both characters do not match as "mismatch", and a section where neither character exists as "none". The character determination unit 97 outputs the determination result of the character as a character determination signal 253 to the state determination unit 74.
[0078] FIG. 12 shows the data structure of the table held in the state determination unit 74 according to Embodiment 3. In the speaking section, when the character determination signal 253 is "match", the state determination unit 74 determines that it is the target person's speech, and when the character determination signal 253 is not "match", it determines that there is multiple speech. In the speaking section, when it is determined that the character determination signal 253 is "match", it corresponds to the case where the user's speech is detected based on the voice signal 200. In the non-speaking section, when the character determination signal 253 is "none", the state determination unit 74 determines that there is silence, and when the character determination signal 253 is not "none", it determines that there is other person's speech. In the non-speaking section, when it is determined that the character determination signal 253 is "none", it corresponds to the case where stationary noise is detected based on the voice signal 200. Note that stationary noise is a sound that interferes with a sound that is a target of continuous sound of a certain magnitude.
[0079] Here, in the speaking interval, when the character determination signal 253 is "none", the state determination unit 74 determines that the character determination unit 97 has made some incorrect determination and treats it as "discrepancy". Similarly, in the non-speaking interval, when the character determination signal 253 is "match", the state determination unit 74 determines that the character determination unit 97 has made some incorrect determination and treats it as "discrepancy".
[0080] When the state determination unit 74 determines that it is the target person's speech in the speaking interval, it determines the update of the filter coefficient in the coefficient update unit 60 according to the first frequency. When the state determination unit 74 determines that it is multiple speech in the speaking interval, it determines the update of the filter coefficient in the coefficient update unit 60 according to the second frequency. When the state determination unit 74 determines that it is silent in the non-speaking interval, it determines the update of the filter coefficient in the coefficient update unit 60 according to the third frequency. When the state determination unit 74 determines that it is other person's speech in the non-speaking interval, it determines the update of the filter coefficient in the coefficient update unit 60 according to the second frequency.
[0081] FIG. 13(a), FIG. 13(b), and FIG. 13(c) show the time changes of the sound pressure level, the myoelectric potential analysis result, and the state determination result in the microphone 10 and the vibration sensor 14 according to the third embodiment. FIG. 13(a) shows the time changes of the vibration signal 202 and the voice signal 200 in the same way as FIG. 5(a). The horizontal axis of FIG. 13(a) represents time, and the vertical axis represents the sound pressure level. FIG. 13(b) shows the time change of the speaking interval information 204, which is the myoelectric potential analysis result in the myoelectric potential analysis unit 92, in the same way as FIG. 5(b). The horizontal axis of FIG. 5(b) represents time, and the vertical axis represents the value of the speaking interval information 204. The high-level speaking interval information 204 indicates "speech", and the low-level speaking interval information 204 indicates "silent (non-speaking)". FIG. 13(c) shows the comparison result between the characters included in the detected character information 205 in the character determination unit 97 and the characters included in the character signal 251. The horizontal axis of FIG. 13(c) represents time, and the vertical axis represents the comparison result ("match", "discrepancy", "none").
[0082] Period L11 is the state of the target person's speech. The speech section information 204 indicates speech, and the character determination signal 253, which is the comparison result in the character determination unit 97, indicates "match". As a result, the state determination unit 74 identifies the target person's speech during period L11. Period L12 is a silent state. The speech section information 204 indicates silence, and the character determination signal 253, which is the comparison result in the character determination unit 97, indicates "none". As a result, the state determination unit 74 identifies the silence. Period L13 is the state of the target person's speech, similar to period L11, and period L14 is the silent state, similar to period L12.
[0083] Period L15 is the state of another person's speech. The speech section information 204 indicates silence, and the character determination signal 253, which is the comparison result in the character determination unit 97, indicates "mismatch". As a result, the state determination unit 74 identifies another person's speech. Period L16 is the state of overlapping speech. The speech section information 204 indicates speech, and the character determination signal 253, which is the comparison result in the character determination unit 97, indicates "mismatch". As a result, the state determination unit 74 identifies the overlapping speech. Return to FIG. 10.
[0084] The adaptation control unit 34 receives the determination result (determination signal 252) from the determination unit 94. If the determination result is an update based on the first frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the first frequency. If the determination result is an update based on the second frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the second frequency. If the determination result is an update based on the third frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the third frequency. That is, the adaptation control unit 34 adjusts the update frequency of the filter coefficient in the coefficient update unit 60 based on the myoelectric potential signal 250 and the audio signal 200 that reflect whether it is the user's speech section or non-speech section.
[0085] FIG. 14 is a flowchart showing a processing procedure by the sound collection device 100 according to Embodiment 3. The myoelectric potential sensor 90 acquires a myoelectric potential signal 250 (S150). The myoelectric potential analysis unit 92 determines whether it is a speech section by performing a myoelectric potential analysis process on the myoelectric potential signal 250 (S152). If it is a speech section (Y in S152), the determination unit 94 acquires the voice signal 200 acquired by the microphone 10 and the voice signal 200 digitally-analog converted by the first AD converter 12 (S154). The phoneme analysis unit 96 of the determination unit 94 performs a phoneme analysis process on the voice signal 200 (S156). The character determination unit 97 of the determination unit 94 outputs a character determination signal 253 which is a determination result of characters included in the detected character information 205 which is the character uttered by the subject as a result of the myoelectric potential analysis performed by the myoelectric potential analysis unit 92 and the character signal 241 which is the result of the phoneme analysis process on the voice signal 200 by the phoneme analysis unit 96 (S158). If the character determination signal 253 is not "match" (N in S158), the state determination unit 74 determines that it is a multiple speech (S159) and determines an update at the second frequency (S160). If the character determination signal 253 is "match" (Y in S158), the state determination unit 74 determines that it is a subject speech (S161) and determines an update at the first frequency (S162).
[0086] If it is not a speech section (N in S152), the determination unit 94 acquires the voice signal 200 acquired by the microphone 10 and the voice signal 200 digitally-analog converted by the first AD converter 12 (S164), and the phoneme analysis unit 96 of the determination unit 94 performs a phoneme analysis process on the voice signal 200 (S166). The character determination unit 97 of the determination unit 94 outputs a character determination signal 253 which is a determination result of characters included in the detected character information 205 which is the character uttered by the subject as a result of the myoelectric potential analysis performed by the myoelectric potential analysis unit 92 and the character signal 251 which is the result of the phoneme analysis process on the voice signal 200 by the phoneme analysis unit 96 (S168). If the character determination signal 253 is not "none" (N in S168), the state determination unit 74 determines that it is an other person's speech (S169) and determines an update at the second frequency (S170). If the character determination signal 253 is "none" (Y in S168), the state determination unit 74 determines that it is a silent sound (S171) and determines an update at the third frequency (S172).
[0087] The adaptive filter 18 executes adaptive filter processing according to the control content determined by the state determination unit 74 of the determination unit 94 (S170). If the operation has not ended (N in S172), the process returns to step 150. If the operation has ended (Y in S172), the process ends.
[0088] According to this embodiment, based on the electromyogram analysis result and the voice signal that determine whether it is the user's speaking section or non-speaking section, the frequency of updating the filter coefficient is adjusted, so that the influence of voices or noise components other than the target sound can be suppressed. Further, by imaging the face of the subject and detecting the movement of the face (mouth) of the subject by electromyogram analysis, the speaking section of the subject can be determined with high accuracy. Further, since the speaking section of the subject is determined with high accuracy, the control accuracy of the filter coefficient can be improved.
[0089] Also, when the user's speech is detected based on the voice signal in the speaking section in the electromyogram analysis result, the filter coefficient is updated at the first frequency, so that the filter coefficient can be updated to be closer to the voice signal. Further, when a voice other than the user's speech is detected based on the voice signal in the speaking section in the electromyogram analysis result, the filter coefficient is updated at a second frequency smaller than the first frequency, so that the influence of voices or noise components other than the target sound can be suppressed.
[0090] Also, when stationary noise is detected based on the voice signal in the non-speaking section in the electromyogram analysis result, the filter coefficient is updated at a third frequency smaller than the first frequency, so that the influence of the noise component can be suppressed. Further, when a voice other than the user's speech is detected based on the voice signal in the non-speaking section in the electromyogram analysis result, the filter coefficient is updated at a second frequency smaller than the first frequency, so that the influence of voices or noise components other than the target sound can be suppressed.
[0091] The above has described the present invention based on the embodiments. It is understood by those skilled in the art that these embodiments are illustrative, and various modifications are possible for each of their constituent elements and combinations of each processing process, and such modifications are also within the scope of the present invention.
Explanation of Reference Numerals
[0092] 10 Microphone, 12 First AD converter, 14 Vibration sensor, 16 Second AD converter, 18 Adaptive filter, 20 Subtractor, 22 DA converter, 32 Correlation determination unit, 34 Adaptive control unit, 50 Delay unit, 52 Multiplier, 54 Adder, 60 Coefficient update unit, 70 Threshold setting unit, 72 Correlation value calculation unit, 74 State determination unit, 90 Electromyogram sensor, 92 Electromyogram analysis unit, 94 Determination unit, 96 Phoneme analysis unit, 97 Character determination unit, 100 Sound collection device, 200 Voice signal, 202 Vibration signal, 204 Speech section information, 205 Detected character information, 206 Correlation information, 208 Adaptive control signal, 210 Residual signal, 212 Converted voice signal, 214 Output signal, 216 First threshold, 218 Second threshold, 220 Correlation value, 224 Correlation values for a predetermined number of times, 250 Electromyogram signal, 251 Character signal, 252 Determination signal, 253 Character determination signal.
Claims
1. An adaptive filter that performs an operation based on a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal; A coefficient update unit that updates the filter coefficient based on a residual signal that is a difference between a voice signal acquired by a microphone and the converted voice signal; An adaptive control unit that adjusts the update frequency of the filter coefficient in the coefficient update unit based on an electromyogram signal reflecting whether it is a user's speaking section or a non-speaking section and the voice signal; A sound collection device comprising the above.
2. When the adaptive control unit detects the user's speech based on the voice signal in the speaking section of the electromyogram signal, the adaptive control unit updates the filter coefficient at a first frequency; When the adaptive control unit detects voice other than the user's speech based on the voice signal in the speaking section of the electromyogram signal, the adaptive control unit updates the filter coefficient at a second frequency; The sound collection device according to claim 1, wherein the second frequency is smaller than the first frequency.
3. When the adaptive control unit detects stationary noise based on the voice signal in the non-speaking section of the electromyogram signal, the adaptive control unit updates the filter coefficient at a third frequency; When the adaptive control unit detects voice other than the user's speech based on the voice signal in the non-speaking section of the electromyogram signal, the adaptive control unit updates the filter coefficient at a second frequency; The sound collection device according to claim 2, wherein the third frequency is smaller than the first frequency and larger than the second frequency.
4. Performing an operation based on a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal; Updating the filter coefficient based on a residual signal that is a difference between a voice signal acquired by a microphone and the converted voice signal; Adjusting the update frequency of the filter coefficient based on an electromyogram signal reflecting whether it is a user's speaking section or a non-speaking section and the voice signal; A sound collection method comprising the above.
5. On a computer, Performing an operation based on a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal; Updating the filter coefficient based on a residual signal that is a difference between a voice signal acquired by a microphone and the converted voice signal; A sound collection program that executes a step of adjusting the frequency of updating the filter coefficients based on the electromyogram signal and the voice signal that reflect whether it is a user's speaking section or a non-speaking section.
Citation Information
Patent Citations
Microphone and sound generation method
JP2007251354A