Sound pick-up device, sound pick-up method, and sound pick-up program

The sound collection device uses an adaptive filter and image recognition to update filter coefficients, addressing the issue of non-target audio interference by enhancing audio quality during speaking intervals.

JP2025099814APending Publication Date: 2025-07-03JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023216758
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing sound collection devices struggle to accurately suppress the influence of audio or noise components other than the target sound, particularly when conversation audio from other speakers overlaps, leading to erroneous inclusion of these components in the converted audio.

Method used

A sound collection device utilizing an adaptive filter that updates its filter coefficients based on a residual signal difference between a microphone and vibration sensor signals, adjusted by an adaptive control unit using image recognition to determine speaking or non-speaking sections.

Benefits of technology

Effectively suppresses the influence of non-target audio and noise components, improving audio quality by accurately controlling the approximation of vibration signals to voice signals during speaking intervals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099814000001_ABST
    Figure 2025099814000001_ABST
Patent Text Reader

Abstract

To provide a technique to prevent the influence of a voice other than a target sound or a noise component.SOLUTION: A sound pick-up device 100 includes an adaptive filter 18, a coefficient update unit, and an adaptive control unit 34. The adaptive filter 18 executes operation with a filter coefficient for a vibration signal 202 acquired in a vibration sensor 14, and outputs a converted voice signal 212. The coefficient update unit updates the filter coefficient on the basis of a residual signal that is the difference between a voice signal 200 acquired in a microphone 10 and the converted voice signal 212. The adaptive control unit 34 adjusts the frequency of update of the filter coefficient in the coefficient update unit on the basis of the voice signal 200 and an image recognition result of a determination as to whether an utterance section or a non-utterance section of a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound collection technology, and particularly to a sound collection device, a sound collection method, and a sound collection program using a microphone and a vibration sensor.

Background Art

[0002] A sound collection device includes a microphone that generates an audio signal based on air vibration and a vibration sensor that generates a vibration signal corresponding to the audio signal based on bone vibration, thereby acquiring clear audio in a noisy environment (for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The sound collection device converts the vibration signal generated by the vibration sensor into an audio signal in a filtering unit, approximates the vibration signal to the audio signal originally transmitted through air, and provides an easy-to-listen audio. However, when the target speaker speaks, if the conversation audio of a person other than the speaker overlaps, it will be erroneously adapted to components other than the target sound. Therefore, it is necessary to accurately grasp the case where the audio of a person other than the speaker is mixed into the microphone. If it does not have a function of detecting the audio components and noise components of a person other than the speaker, when increasing the approximation degree to the microphone audio in the speaking section, there is a risk that the influence of the audio or noise components other than the target sound is erroneously added to the converted audio.

[0005] The present invention has been made in view of such a situation, and its purpose is to provide a technology for suppressing the influence of audio or noise components other than the target sound.

Means for Solving the Problems

[0006] In order to solve the above problems, a sound collection device according to an aspect of the present invention includes an adaptive filter that performs an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal, and a coefficient update unit that updates the filter coefficient based on a residual signal that is a difference between the voice signal acquired by a microphone and the converted voice signal, and an adaptive control unit that adjusts the frequency of updating the filter coefficient in the coefficient update unit based on an image recognition result that determines whether it is a user's speaking section or a non-speaking section and the voice signal.

[0007] Another aspect of the present invention is a sound collection method. This sound collection method includes a step of performing an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal, a step of updating the filter coefficient based on a residual signal that is a difference between the voice signal acquired by a microphone and the converted voice signal, and a step of adjusting the frequency of updating the filter coefficient based on an image recognition result that determines whether it is a user's speaking section or a non-speaking section and the voice signal.

[0008] Another aspect of the present invention is a sound collection program. This sound collection program causes a computer to execute a step in which an adaptive filter performs an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal, a step of updating the filter coefficient based on a residual signal that is a difference between the voice signal acquired by a microphone and the converted voice signal, and a step of adjusting the frequency of updating the filter coefficient based on an image recognition result that determines whether it is a user's speaking section or a non-speaking section and the voice signal.

[0009] Note that any combination of the above components, and those obtained by converting the expression of the present invention among a method, an apparatus, a system, a recording medium, a computer program, etc. are also effective as aspects of the present invention.

Advantages of the Invention

[0010] According to the present invention, it is possible to suppress the influence of voices or noise components other than the target sound.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Best Mode for Carrying Out the Invention

[0012] (Example 1) Before specifically describing the present invention, an overview will first be described. This example relates to a sound collection device. The sound collection device includes a microphone that generates an audio signal based on air vibrations, a vibration sensor that generates a vibration signal based on vibrations transmitted to the human body, and an adaptive filter that multiplies a coefficient by the vibration signal to generate a converted audio signal in order to correct the vibration signal to approximate the audio signal. Further, when the sound collection device detects the speaking situation of the target sound by image analysis and detects the mixing of audio components other than the target audio, the update (learning) of the filter coefficient in the adaptive filter is stopped, and it is suppressed that the converted audio signal is approximated to the audio components other than the target sound.

[0013] The mixing of audio components other than the target sound is determined using a threshold value that serves as a reference for the correlation value between the vibration signal and the audio signal. The reference threshold value is variable in conjunction with the correlation between the vibration signal and the audio signal in a non-audio section, that is, under ambient noise. Thereby, the audio quality of the converted audio signal is improved by detecting a higher-precision correlation. Note that the converted audio signal is a digital signal related to audio obtained by performing analog-to-digital conversion on the vibration signal acquired by the vibration sensor and further performing an operation using a predetermined filter coefficient, that is, an audio signal converted based on the vibration signal.

[0014] FIG. 1 shows the configuration of the sound collection device 100 according to Embodiment 1. The sound collection device 100 includes a microphone 10, a first AD converter (Analog to Digital converter) 12, a vibration sensor 14, a second AD converter 16, an adaptive filter 18, a subtractor 20, a DA converter (Digital to Analog converter) 22, a correlation determination unit 32, an adaptive control unit 34, an imaging device 80, and an image recognition unit 82. The imaging device 80 is, for example, a photographing device such as a color camera composed of a lens, an image sensor such as a CMOS (Complementary Metal Oxide Semiconductor), a driver for digitizing a video signal output by the image sensor, a DRAM (Dynamic Random Access Memory) for temporarily storing image data for image processing, and a DSP (Digital Signal Processor).

[0015] The microphone 10 converts air vibrations into an audio signal 200. Since the microphone 10 can accurately acquire the audio signal 200, the audio signal 200 acquired by the microphone 10 is relatively close to the audio signal 200 that a person perceives through the ear. Therefore, by using the audio signal 200 acquired by the microphone 10 as the target value of the vibration signal 202 described below, it is possible to maintain the audio quality of the converted audio signal 212 output by the adaptive filter 18 at a high level.

[0016] The first AD converter 12 performs analog-digital conversion on the audio signal 200 acquired by the microphone 10, and outputs the audio signal 200 converted into a digital signal (hereinafter, this is also referred to as "audio signal 200") to the subtractor 20 and the correlation determination unit 32.

[0017] The vibration sensor 14 acquires the speech signal emitted by a person mainly as the vibration signal 202 transmitted to the human body. The vibration sensor 14 may be a microphone that detects sound or vibration, or may be a camera (visual microphone) that detects vibration from the variation of pixels or the color fluctuation between frames of video to acquire a sound signal. Instead of the vibration sensor 14, a receiving device embedded in the body or a microphone in direct contact with the human body may be used, or a device that measures the vibration on the surface of the human body in a non-contact manner with the human body like a laser displacement meter may be used.

[0018] The second AD converter 16 outputs the vibration signal 202 (hereinafter also referred to as "vibration signal 202") converted into a digital signal, which is obtained by performing analog-digital conversion on the vibration signal 202 acquired by the vibration sensor 14, to the adaptive filter 18 and the correlation determination unit 32.

[0019] Compared with the audio signal 200, the vibration signal 202 has the characteristics of a narrow frequency band and a large bias in the frequency band. Due to this characteristic, the vibration signal 202 becomes a very muffled sound and sounds significantly different from the original audio signal 200. To improve this, in this embodiment, the audio signal 200 acquired at the same time is used as a reference signal, and the adaptive filter 18 is used to make the vibration signal 202 closer to the audio signal 200.

[0020] The adaptive filter 18 receives the vibration signal 202 from the second AD converter 16 and the residual signal 210 from the subtractor 20. The adaptive filter 18 performs an operation on the vibration signal 202 with a filter coefficient based on the residual signal 210, and outputs the operation result as a converted audio signal 212.

[0021] FIG. 2 shows the configuration of the adaptive filter 18 according to Embodiment 1. The adaptive filter 18 includes a first delay element 50a to an (N - 1)-th delay element 50n - 1 collectively referred to as a delay element 50, a first multiplier 52a to an N-th multiplier 52n collectively referred to as a multiplier 52, a first adder 54a to an (N - 1)-th adder 54n - 1 collectively referred to as an adder 54, and a coefficient update unit 60.

[0022] The delay unit 50 delays the vibration signal 202 and sequentially outputs it. The multiplier 52 multiplies the vibration signal 202 and the filter coefficient. The adder 54 sequentially adds the multiplication results in the multiplier 52. The addition result of the (N - 1)th adder 54n - 1 corresponds to the aforementioned operation result. The subtractor 20 calculates a residual signal 210, which is a difference value between the audio signal 200, which is a reference signal, and the converted audio signal 212, which is the operation result, and outputs it to the adaptive filter 18.

[0023] The coefficient update unit 60 updates the filter coefficient based on the audio signal 200 and the operation result by executing, for example, the LMS (Least Mean Square) algorithm as follows.

Equation

[0024] The DA converter 22 outputs a converted audio signal 212 (hereinafter referred to as "output signal 214") converted into an analog signal by performing digital - to - analog conversion on the converted audio signal 212 output from the adaptive filter 18.

[0025] Here, when voice components other than the target voice are mixed in the voice signal 200 acquired by the microphone 10, in the adaptive filter 18, an approximation effect may mistakenly work on the voices of others. Voice components other than the target voice are, for example, the voices of people other than the target person, the voices reproduced by a handset speaker (not shown) at the call destination, and the like. Therefore, in order to improve the accuracy of the process for bringing the vibration signal 202 in the adaptive filter 18 closer to the voice signal 200, it is required to update (learn) the filter coefficients when no voice components other than the target voice are mixed in the voice signal 200 during the speaking interval. As a preferable condition for updating (learning) the filter coefficients, it is desirable that it is the speaking interval of the user. Note that the speaking interval is the time zone in which the speech of the target person is detected.

[0026] In this embodiment, in order to detect whether it is the speaking interval or the non-speaking interval of the target person (user), the imaging device 80 images the face (mouth) of the target person. The imaging device 80 outputs the captured video (image) as an image signal 240 to the image recognition unit 82. The image recognition unit 82 detects the movement of the mouth of the target person from the image signal 240 and specifies the speaking interval by performing image recognition processing on the image signal 240. More specifically, the image recognition unit 82 performs face detection and mouth detection by image analysis, and detects the speaking interval when the amount of the movement vector corresponding to speech exceeds the movement vector amount per unit time of the mouth.

[0027] In order to detect the speaking interval from the image signal 240, for example, technologies such as the literature: Vol.2011-CVIM-177No.13 "Speech Detection by Extraction and Recognition of Lip Region" may be used. The image recognition unit 82 outputs the image recognition result of determining whether it is the speaking interval or the non-speaking interval of the target person to the correlation determination unit 32 as the speaking interval information 204. In order to improve the determination accuracy of the speaking interval, the image recognition unit 82 may use the voice interval determination result by the voice signal 200 by regarding the vibration signal 202 acquired by the vibration sensor 14 as the voice signal 200 as an auxiliary. The effect of preventing false detection due to the movement of the mouth without speech such as yawning can be obtained.

[0028] The human vocal cords generate the types of voice signals 200, and various voice signals 200 are generated by passing through the throat and the mouth opening, which are vocal tract tubes. However, between the vibration signal 202 associated with speech acquired at a predetermined human body part and the voice signal 200 obtained by the microphone 10 acquiring air vibrations at a predetermined position, there is a correlation accompanied by a spatial delay during propagation in space and transmission characteristics due to human body components (bones, muscles, skin, etc.).

Number

[0029] When compensating for a predetermined spatial delay (z1) and the gain correction amount with respect to the reference vibration signal V(t) to align the time axis and sound pressure level of the vibration signal 202 and the voice signal 200, it is expected that they will have similar signal values depending on the type of speech. If it is only the speech signal of the subject, by performing a cross-correlation calculation between the vibration signal 202 and the voice signal 200, a point with high correlation is observed in the vicinity of the time difference z1 + z2. The general formula for cross-correlation is shown as follows.

Number

[0030] In Equation (3), it is observed that the point with the highest correlation has the maximum value Cmax of the correlation value C. Here, τ corresponds to the time difference (z1 + z2) between the speech signal series S(t) and the vibration signal series V(t) with the same signal source. By calculating Equation (3), the maximum value Cmax, which is the correlation value C of the point with the highest correlation between the vibration signal 202 and the speech signal 200, determines the number of samples T according to the sampling frequency from a predetermined analysis time width, gives a shift amount of about several samples to more than a dozen samples before and after around τ, and performs a convolution operation, and is observed as the numerical value of the maximum value of the sum of products.

[0031] The cross-correlation calculation may be performed as follows.

Number

[0032] On the other hand, when the voices of people other than the target speaker or noise components are mixed in, the correlation between the vibration signal 202 and the speech signal 200 decreases, so conversely, a numerical value with low correlation is observed for the correlation value. Therefore, the correlation determination unit 32 determines whether there is a multiple speech state in the time period when the target speaker speaks based on whether the correlation value is equal to or greater than a threshold value.

[0033] The correlation determination unit 32 detects a multiple talk state or a situation where the noise environment has deteriorated based on at least one of the maximum value Cmax of the correlation value C calculated by the above formula (3) or the minimum value Cmin of the correlation value C calculated by the above formula (4). FIG. 3 shows the configuration of the correlation determination unit 32 according to the first embodiment. The correlation determination unit 32 includes a threshold setting unit 70, a correlation value calculation unit 72, and a state determination unit 74. The threshold setting unit 70 receives the speech section information 204 indicating whether it is a speech section or a non-speech section from the image recognition unit 82. The threshold setting unit 70 sets the larger of a predetermined value and the reference correlation value as the first threshold 216 and the smaller of a predetermined value and the reference correlation value as the second threshold 218, based on the correlation value in the case of ambient noise without an external sound source and in a non-speech section.

[0034] The correlation value calculation unit 72 calculates the correlation value 220 between the audio signal 200 and the vibration signal 202 by executing the above formula (3) or formula (4). At this time, in order to improve the accuracy of the correlation value 220, the sound pressure levels of the vibration signal 202 and the audio signal 200 may be aligned in advance. For example, a gain adjustment unit may be provided after the first AD converter 12 in FIG. 2 so that the sound pressure levels of the audio signal 200 during speech are about the same, or the adjustment may be made in the correlation value calculation unit 72.

[0035] The state determination unit 74 receives the first threshold 216 or the second threshold 218 from the threshold setting unit 70 and receives the correlation value 220 from the correlation value calculation unit 72. The state determination unit 74 determines the states of the target person's speech, multiple talk, other person's speech, and silence by comparing the first threshold 216 or the second threshold 218 with the correlation value 220. Further, the state determination unit 74 determines the control content of the coefficient update unit 60 according to the determined state. The target person's speech is a state where only the target person is speaking, the multiple talk is a state of double talk or a state where external noise is mixed in, the other person's speech is a state where only a person other than the target person is speaking, and the silence is a state where there is no speech.

[0036] Figs. 4(a) and 4(b) show the data structure of the table held in the state determination unit 74 according to the first embodiment. When the correlation value 220 is equal to or greater than the first threshold value 216 in the speech section, the state determination unit 74 determines that it is the target person's speech and determines the adaptation activation of the coefficient update unit 60. When the correlation value 220 is less than the first threshold value 216 in the speech section, the state determination unit 74 determines that it is a multiple speech and determines the adaptation stop of the coefficient update unit 60. When the correlation value 220 is equal to or greater than the second threshold value 218 in the non-speech section, the state determination unit 74 determines that it is silence and determines the adaptation save of the coefficient update unit 60. When the correlation value 220 is less than the second threshold value 218 in the non-speech section, the state determination unit 74 determines that it is the speech of others and determines the adaptation stop of the coefficient update unit 60.

[0037] Here, in the adaptation activation, "α" in Expression (1) is set to the first value. The first value is determined in advance. In the adaptation save, "α" in Expression (1) is set to the second value. The second value is smaller than the first value. In the adaptation stop, the update of the filter coefficient in the coefficient update unit 60 is stopped. Return to Fig. 3.

[0038] The vibration signal 202 indicates the vibration of the human body during speech, but compared with the audio signal 200 picked up by the microphone 10, the noise level tends to be high during ambient noise. When determining the correlation between the vibration signal 202 and the audio signal 200, the first threshold value 216 or the second threshold value 218 may be set based on the correlation value 220 during ambient noise without an external sound source and in the non-speech section. The transfer characteristics shown in Expression (2) vary depending on the type of speech, and even when only the speaker is speaking, the correlation value 220 may fluctuate accordingly. In particular, due to the positional relationship between the part of the human body used as the measurement point for picking up the vibration signal 202 during speech and the vocal cords, throat, and mouth where the sound wave is generated, the correlation value 220 during speech may decrease.

[0039] However, since the correlation value 220 obtained during ambient noise and in a non-speaking interval is calculated by a noise component having no correlation between the vibration signal 202 and the voice signal 200, it will not fall below this correlation value 220 during speech. Therefore, the threshold setting unit 70 detects that it is a non-speaking interval based on the signal indicating whether it is a speaking interval or a non-speaking interval, and further determines, for the ambient noise state, that the correlation value 220 changes within a certain width over a predetermined time as a determination material.

[0040] The threshold setting unit 70 acquires, from the correlation value calculation unit 72, the correlation values 224 for a predetermined number of times calculated during ambient noise and in a non-speaking interval. When the variation width of the acquired correlation values 224 for a predetermined number of times is within an allowable range (within several times a predetermined value), it calculates a reference correlation value by averaging the correlation values 224 for a predetermined number of times acquired from the correlation value calculation unit 72, and determines the larger of the reference correlation value plus a predetermined value as the first threshold value 216, and the smaller of the reference correlation value minus a predetermined value as the second threshold value 218.

[0041] FIG. 5(a), FIG. 5(b), and FIG. 5(c) show the time changes of the sound pressure level, the image recognition result, and the correlation value 220 in the microphone 10 and the vibration sensor 14 according to the first embodiment. FIG. 5(a) shows the time changes of the vibration signal 202 and the voice signal 200. The horizontal axis in FIG. 5(a) represents time, and the vertical axis represents the sound pressure level. FIG. 5(b) shows the time change of the speech interval information 204 which is the image recognition result in the image recognition unit 82. The horizontal axis in FIG. 5(b) represents time, and the vertical axis represents the value of the speech interval information 204. High-level speech interval information 204 indicates "speech", and low-level speech interval information 204 indicates "silence (non-speaking)". FIG. 5(c) shows the time change of the correlation value 220 calculated in the correlation value calculation unit 72. The horizontal axis in FIG. 5(c) represents time, and the vertical axis represents the correlation value 220.

[0042] Since period L1 is the state of the subject's speech, in period L1, both the vibration signal 202 and the audio signal 200 have drastically fluctuating sound pressure levels. Since the vibration signal 202 and the audio signal 200 have a similar tendency in the state of the subject's speech, the correlation value 220 also increases. In period L2, since it is a silent state, in period L2, the average ambient noise level of the vibration signal 202 and the average ambient noise level of the audio signal 200 are obtained. Although these are within a certain correlation, since they are uncorrelated noise components with each other, the correlation value 220 becomes small.

[0043] Period L3 is the state of the subject's speech, similar to period L1. Therefore, the vibration signal 202, the audio signal 200, and the correlation value 220 in period L3 have the same tendency as those in period L1. Period L4 is a silent state, similar to period L2. Therefore, the vibration signal 202, the audio signal 200, and the correlation value 220 in period L4 have the same tendency as those in period L2. Since period L5 is the state of others' speech, in period L5, the variation in the sound pressure level of the vibration signal 202 is small, but the sound pressure level of the audio signal 200 fluctuates drastically. Therefore, the correlation value 220 in period L5 significantly decreases. The state of others' speech can also be said to be a state where external noise is mixed in.

[0044] Period L6 is the state of multiple conversations. Since the vibration signal 202 in period L6 only contains the speech information of the subject himself / herself, the part with intense movement on the vertical axis is the speech signal of the target speaker. On the other hand, the audio signal 200 in period L6 contains the speech information of the subject himself / herself and the speech information of others. Therefore, the correlation value 220 in period L6 decreases. Since the speech information of the subject is included during multiple conversations, the correlation value 220 is higher than during others' speech. Period L7 is the state of others' speech, similar to period L5. Therefore, the vibration signal 202, the audio signal 200, and the correlation value 220 in period L7 have the same tendency as those in period L5.

[0045] The speaking intervals are period L1, period L3, and period L6. In the speaking intervals, in order to distinguish the state of the target person's speech and the state of multiple speeches, the first threshold value 216 shown in FIG. 5(b) is set. On the other hand, the non-speaking intervals are period L2, period L4, period L5, and period L7. In the non-speaking intervals, in order to distinguish the silent state and the state of the other person's speech, the second threshold value 218 shown in FIG. 5(b) is set. Returning to FIG. 1.

[0046] The adaptation control unit 34 receives the determination result (correlation information 206) in the correlation determination unit 32. If the determination result is adaptation active, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to set "α" to the first value. If the determination result is adaptation save, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to set "α" to the second value. If the determination result is stop, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to stop updating the filter coefficient. That is, the adaptation control unit 34 adjusts the degree of update of the filter coefficient in the coefficient update unit 60 based on the vibration signal 202 and the audio signal 200.

[0047] The above configuration can be realized hardware-wise by the CPU, memory, and other LSIs of any computer, and software-wise by a program loaded in the memory or the like. However, here, the functional blocks realized by their cooperation are depicted. Therefore, it is understood by those skilled in the art that these functional blocks can be realized in various forms by hardware only, software only, or a combination thereof.

[0048] The operation of the sound collection device 100 with the above configuration will be described. FIG. 6 is a flowchart showing the update procedure of the first threshold value 216 and the second threshold value 218 by the sound collection device 100 according to the first embodiment. This shows the operation of setting the first threshold value 216 and the second threshold value 218 for the correlation value 220 between the vibration signal 202 and the voice signal 200, which is a criterion for determining whether it is a state of multiple conversations. This operation is performed in a state where there is no speech and no environmental noise, and the first threshold value 216 and the second threshold value 218 may be set by executing a series of operations in advance before speech, or the above state may be determined during speech to set the first threshold value 216 and the second threshold value 218.

[0049] The imaging device 80 acquires an image signal 240 (S10). Specifically, when the processing unit in the subsequent adaptive filter 18 is 20 msec, the imaging device 80 acquires at least 20 msec, that is, 50 frames of video per second as the image signal 240. This is because the period of the image recognition process in the image recognition unit 82 is appropriately several tens of msec in accordance with the processing unit of the voice, and in order to determine the voice section by image analysis, the same time resolution as the voice processing is required to observe the temporal transition of the motion vector.

[0050] The image recognition unit 82 determines whether it is a non-speaking section by performing an image recognition process on the image signal 240 acquired by the imaging device 80 (S12). If it is a non-speaking section (Y in S12), the correlation determination unit 32 acquires the voice signal 200 acquired by the microphone 10 and the voice signal 200 digitally-analog converted by the first AD converter 12, and the vibration signal 202 acquired by the vibration sensor 14 and the vibration signal 202 analog-digital converted by the second AD converter 16 (S14), and the correlation value calculation unit 72 of the correlation determination unit 32 calculates the correlation value 220 based on the voice signal 200 and the vibration signal 202 a predetermined number of times according to a predetermined analysis time width and sampling frequency (S16).

[0051] The threshold setting unit 70 of the correlation determination unit 32 acquires the correlation values 224 for a predetermined number of times calculated in step 16 from the correlation value calculation unit 72, and determines whether the variation range of the correlation values is within the allowable range (S18). When the variation range of the correlation values is within the allowable range (Y in S18), that is, when the ambient noise continues, the threshold setting unit 70 updates the first threshold 216 and the second threshold 218 based on the correlation values 224 for a predetermined number of times acquired from the correlation value calculation unit 72 (S20). When the variation range of the correlation values is not within the allowable range (N in S18), step 20 is skipped. When it is not a non-speaking interval (N in S12), steps 14 to 20 are skipped. When the operation has not ended (N in S22), the process returns to step 10. When the operation has ended (Y in S22), the process is terminated.

[0052] FIG. 7 is a flowchart showing a processing procedure by the sound collection device 100 according to the first embodiment. Step 50 is the same as the operation shown in FIG. 6, and here shows the case where an interval in which the ambient noise continues during use is detected and the first threshold 216 and the second threshold 218 are always updated. When setting the first threshold 216 and the second threshold 218 in advance in a situation where the ambient noise continues before an actual conversation starts, step 50 may be omitted.

[0053] The imaging device 80 acquires an image signal 240 (S52). The image recognition unit 82 determines whether it is a speaking interval by performing image recognition processing on the image signal 240 acquired by the imaging device 80 (S54). When it is a speaking interval (Y in S54), the correlation determination unit 32 acquires the voice signal 200 acquired by the microphone 10 and the voice signal 200 digitally-analog converted by the first AD converter 12, and the vibration signal 202 acquired by the vibration sensor 14 and the vibration signal 202 analog-digital converted by the second AD converter 16 (S56), and the correlation value calculation unit 72 of the correlation determination unit 32 calculates a correlation value 220 based on the voice signal 200 and the vibration signal 202 (S58).

[0054] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the first threshold value 216 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the first threshold value 216 (S60). When the correlation value 220 is not greater than or equal to the first threshold value 216 (N in S60), the state determination unit 74 of the correlation determination unit 32 estimates multiple speech (S62) and determines an adaptive stop (S64). When the correlation value 220 is greater than or equal to the first threshold value 216 (Y in S60), the state determination unit 74 of the correlation determination unit 32 estimates the target person's speech (S66) and determines an adaptive active (S68).

[0055] When it is not in the speech section (N in S54), the correlation determination unit 32 causes the microphone 10 to acquire the audio signal 200, and acquires the audio signal 200 that has been digitally - analog - converted by the first AD converter 12, and the vibration sensor 14 acquires the vibration signal 202, and the second AD converter 16 acquires the vibration signal 202 that has been analog - digital - converted (S69). The correlation value calculation unit 72 of the correlation determination unit 32 calculates a correlation value 220 based on the audio signal 200 and the vibration signal 202 (S70).

[0056] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the second threshold value 218 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the second threshold value 218 (S72). When the correlation value 220 is not greater than or equal to the second threshold value 218 (N in S72), the state determination unit 74 of the correlation determination unit 32 estimates other - person speech (S74) and determines an adaptive stop (S76). When the correlation value 220 is greater than or equal to the second threshold value 218 (Y in S72), the state determination unit 74 of the correlation determination unit 32 estimates silence (S78) and determines an adaptive save (S80).

[0057] The adaptive filter 18 executes adaptive filter processing according to the control content determined by the state determination unit 74 of the correlation determination unit 32 (S82). When the operation has not ended (N in S84), it returns to step 50. When the operation has ended (Y in S84), the process is terminated.

[0058] According to this embodiment, by approximating the vibration signal to the characteristics of the voice signal, the influence of ambient noise components can be suppressed. Also, the presence of voices other than the speaker's and sudden noises is sequentially detected from the correlation between the vibration signal and the voice signal, and the degree of approximation to the voice signal is controlled based on the detection results, so that the influence of voices or noise components other than the target voice can be suppressed. Further, since the influence of voices or noise components other than the target voice is suppressed, the mixing of voice components of people other than the speaker is avoided and the voice quality of the speaker can be improved. Also, by imaging the face of the target person and detecting the movement of the face (mouth) of the target person by image recognition, the speaking interval of the target person can be determined with high accuracy. Further, since the speaking interval of the target person is determined with high accuracy, the control accuracy of the filter coefficient can be improved.

[0059] (Embodiment 2) Next, Embodiment 2 will be described. Embodiment 2 relates to the sound collection device 100 in the same manner as Embodiment 1. The sound collection device 100 according to Embodiment 1 adjusts any one of adaptive active, adaptive stop, and adaptive save, that is, the degree of update of the filter coefficient, based on the image recognition result and the correlation value 220 between the vibration signal 202 and the voice signal 200. On the other hand, the sound collection device 100 according to Embodiment 2 adjusts the frequency of update of the filter coefficient based on the image recognition result and the correlation value 220 between the vibration signal 202 and the voice signal 200. The sound collection device 100, the adaptive filter 18, and the correlation determination unit 32 according to Embodiment 2 are of the same type as those in FIGS. 1 to 3. Here, the description will focus on the differences from Embodiment 1.

[0060] The state determination unit 74 receives the first threshold value 216 or the second threshold value 218 from the threshold setting unit 70 and receives the correlation value 220 from the correlation value calculation unit 72. The state determination unit 74 determines the state of the target person speaking, multiple speaking, other person speaking, or silence by comparing the first threshold value 216 or the second threshold value 218 with the correlation value 220. Also, the state determination unit 74 determines the control content of the coefficient update unit 60 according to the determined state.

[0061] Figures 8(a) and 8(b) show the data structure of the table held in the state determination unit 74 according to the second embodiment. When the correlation value 220 is equal to or greater than the first threshold value 216 in the speech section, the state determination unit 74 determines that it is the target person's speech and determines the update by the first frequency of the filter coefficient in the coefficient update unit 60. When the correlation value 220 is less than the first threshold value 216 in the speech section, the state determination unit 74 determines that it is a multiple speech and determines the update by the second frequency of the filter coefficient in the coefficient update unit 60. When the correlation value 220 is equal to or greater than the second threshold value 218 in the non-speech section, the state determination unit 74 determines that it is silent and determines the update by the third frequency of the filter coefficient in the coefficient update unit 60. When the correlation value 220 is less than the second threshold value 218 in the non-speech section, the state determination unit 74 determines that it is the speech of others and determines the update by the second frequency of the filter coefficient in the coefficient update unit 60.

[0062] Here, the second frequency is made smaller than the first frequency, and the third frequency is made smaller than the first frequency and larger than the second frequency. For example, the first frequency is every time, the second frequency is once every 128 times, and the third frequency is between once every 2 times and once every 64 times. The values of each frequency are not limited to these. Return to FIG. 3.

[0063] The adaptation control unit 34 receives the determination result (correlation information 206) in the correlation determination unit 32. If the determination result is an update by the first frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient by the first frequency. If the determination result is an update by the second frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient by the second frequency. If the determination result is an update by the third frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient by the third frequency. That is, the adaptation control unit 34 adjusts the update frequency of the filter coefficient in the coefficient update unit 60 based on the vibration signal 202 and the audio signal 200.

[0064] FIG. 9 is a flowchart showing a processing procedure by the sound collection device 100 according to the second embodiment. Step 100 is the same as the operation shown in FIG. 6, and here it shows the case where a section with continuous ambient noise is detected during use, and the first threshold value 216 and the second threshold value 218 are always updated. If the first threshold value 216 and the second threshold value 218 are set in advance in a situation where ambient noise continues before an actual conversation starts, step 100 may be omitted.

[0065] The imaging device 80 acquires an image signal 240 (S102). The image recognition unit 82 determines whether it is a speech section by performing image recognition processing on the image signal 240 acquired by the imaging device 80 (S104). If it is a speech section (Y in S104), the correlation determination unit 32 acquires the audio signal 200 obtained by the microphone 10 and the audio signal 200 digitally - to - analog - converted by the first AD converter 12, and the vibration signal 202 obtained by the vibration sensor 14 and the vibration signal 202 analog - to - digital - converted by the second AD converter 16 (S106), and the correlation value calculation unit 72 of the correlation determination unit 32 calculates a correlation value 220 based on the audio signal 200 and the vibration signal 202 (S108).

[0066] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the first threshold value 216 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the first threshold value 216 (S110). If the correlation value 220 is not greater than or equal to the first threshold value 216 (N in S110), the state determination unit 74 of the correlation determination unit 32 estimates multiple speech (S112) and determines an update at the second frequency (S114). If the correlation value 220 is greater than or equal to the first threshold value 216 (Y in S110), the state determination unit 74 of the correlation determination unit 32 estimates the target - person speech (S116) and determines an update at the first frequency (S118).

[0067] When it is not the speech period (N in S104), the correlation determination unit 32 acquires the audio signal 200 by the microphone 10, and acquires the audio signal 200 that has been digitally - analog - converted by the first AD converter 12 and the vibration signal 202 that has been analog - digital - converted by the second AD converter 16 (S119). The correlation value calculation unit 72 of the correlation determination unit 32 calculates the correlation value 220 (S120).

[0068] The state determination unit 74 of the correlation determination unit 32 acquires the correlation value 220 calculated by the correlation value calculation unit 72 of the correlation determination unit 32 and the second threshold value 218 set by the threshold value setting unit 70 of the correlation determination unit 32, and compares the correlation value 220 with the second threshold value 218 (S122). When the correlation value 220 is not greater than or equal to the second threshold value 218 (N in S122), the state determination unit 74 of the correlation determination unit 32 estimates that it is other - person speech (S124) and determines an update at the second frequency (S126). When the correlation value 220 is greater than or equal to the second threshold value 218 (Y in S122), the state determination unit 74 of the correlation determination unit 32 estimates that it is silent (S128) and determines an update at the third frequency (S130).

[0069] The adaptive filter 18 executes adaptive filter processing according to the control content determined by the state determination unit 74 of the correlation determination unit 32 (S132). When the operation has not ended (N in S134), it returns to step 100. When the operation has ended (Y in S134), the process is terminated.

[0070] According to this embodiment, since the update frequency of the filter coefficients is adjusted based on the correlation value between the vibration signal and the audio signal, the influence of voices or noise components other than the target sound can be suppressed. Further, when the correlation value is equal to or greater than the first threshold value during the speaking interval, the filter coefficients are updated at a first frequency equal to or greater than the second frequency, so that the filter coefficients can be updated to be closer to the audio signal. Further, when the correlation value is less than the first threshold value during the speaking interval, the filter coefficients are updated at a second frequency less than the first frequency, so that the influence of voices or noise components other than the target sound can be suppressed. Further, when the correlation value is equal to or greater than the second threshold value during the non-speaking interval, the filter coefficients are updated at a third frequency less than the first frequency, so that the influence of the noise components can be suppressed. Further, when the correlation value is less than the second threshold value during the non-speaking interval, the filter coefficients are updated at a second frequency less than the first frequency, so that the influence of voices or noise components other than the target sound can be suppressed. Further, by imaging the face of the subject and detecting the movement of the face (mouth) of the subject by image recognition, the speaking interval of the subject can be determined with high accuracy. Further, since the speaking interval of the subject is determined with high accuracy, the control accuracy of the filter coefficients can be improved.

[0071] (Example 3) Next, Example 3 will be described. Example 3 relates to the sound collection device 100 as before. The sound collection device 100 according to Example 2 adjusts the update frequency of the filter coefficients based on the image recognition result and the correlation value 220 between the vibration signal 202 and the audio signal 200. On the other hand, the sound collection device 100 according to Example 3 adjusts the update frequency of the filter coefficients based on the image recognition result and the phoneme analysis result (speech recognition result) of the audio signal 200. Here, the description will focus on the differences from before.

[0072] FIG. 10 shows the configuration of the sound collection device 100 according to Example 3. The sound collection device 100 includes a microphone 10, a first AD converter 12, a vibration sensor 14, a second AD converter 16, an adaptive filter 18, a subtractor 20, a DA converter 22, an adaptive control unit 34, an imaging device 80, an image recognition unit 82, and a determination unit 84.

[0073] The microphone 10, the first AD converter 12, the vibration sensor 14, the second AD converter 16, the adaptive filter 18, the subtractor 20, the DA converter 22, and the adaptive control unit 34 are the same as before. Further, a determination unit 84 is included instead of the correlation determination unit 32 used heretofore. The first AD converter 12 outputs the audio signal 200 to the subtractor 20 and the determination unit 84, and the second AD converter 16 outputs the vibration signal 202 to the adaptive filter 18.

[0074] The imaging device 80 outputs the captured video (image) as an image signal 240 to the image recognition unit 82 in the same manner as before. The image recognition unit 82 executes image recognition processing on the image signal 240 to detect the movement of the mouth of the subject from the image signal 240 and specify the utterance section. Further, the image recognition unit 82 specifies the characters uttered by the subject from the detected movement of the mouth of the subject. The characters uttered by the subject are, for example, "a", "i",... The image recognition unit 82 outputs the utterance section information 204 indicating whether it is an utterance section or a non-utterance section of the subject, and the detected character information 205 obtained by detecting the characters uttered by the subject from the movement of the mouth of the subject to the determination unit 84.

[0075] The determination unit 84 receives the audio signal 200 from the first AD converter 12 and also receives the utterance section information 204 and the detected character information 205 from the image recognition unit 82. Based on the audio signal 200, the utterance section information 204, and the detected character information 205, the determination unit 84 classifies the current state into any one of the states of the subject's speech, the silent state, the speech of others, and the multi-speech state, and determines the update frequency according to the classified state. The determination unit 84 outputs the determined update frequency as a determination signal 242 to the adaptive control unit 34. The adaptive control unit 34 receives the determination result (determination signal 242) in the determination unit 84. The adaptive control unit 34 adjusts the update frequency of the filter coefficients in the coefficient update unit 60 of the adaptive filter 18 based on the determination result.

[0076] FIG. 11 shows the configuration of the determination unit 84 according to Embodiment 3. The determination unit 84 includes a state determination unit 74, a phoneme analysis unit 86, and a character determination unit 87. The phoneme analysis unit 86 receives the audio signal 200 from the first AD converter 12. The phoneme analysis unit 86 identifies the characters included in the audio signal 200 by performing phoneme analysis processing (speech recognition processing) on the audio signal 200. The characters included in the audio signal 200 are, for example, "a", "i", ···. Since known techniques may be used for the phoneme analysis processing (speech recognition processing), the description is omitted here. The phoneme analysis unit 86 outputs the characters included in the audio signal 200 as a character signal 241 to the character determination unit 87.

[0077] The character determination unit 87 receives the detected character information 205 from the image recognition unit 82 and the character signal 241 from the phoneme analysis unit 86. The character determination unit 87 compares the characters included in the detected character information 205 with the characters included in the character signal 241. The character determination unit 87 determines a section where both characters match as "match", a section where both characters do not match as "mismatch", and a section where neither character exists as "none". The character determination unit 87 outputs the character determination result as a character determination signal 243 to the state determination unit 74.

[0078] FIG. 12 shows the data structure of the table held in the state determination unit 74 according to Embodiment 3. In the speaking interval, when the character determination signal 243 is "match", the state determination unit 74 determines that it is the target person's speech, and when the character determination signal 243 is not "match", it determines that there is multiple speech. Note that in the speaking interval, when it is determined that the character determination signal 243 is "match", it corresponds to the case where the user's speech is detected based on the audio signal 200. In the non-speaking interval, when the character determination signal 243 is "none", the state determination unit 74 determines that there is no sound, and when the character determination signal 243 is not "none", it determines that there is other person's speech. Note that in the non-speaking interval, when it is determined that the character determination signal 243 is "none", it corresponds to the case where stationary noise is detected based on the audio signal 200. Note that stationary noise is noise that interferes with a target sound where a sound of a certain magnitude is continuous.

[0079] Here, when the character determination signal 243 is "none" during the speech section, the state determination unit 74 determines that the character determination unit 87 has made some incorrect determination and treats it as "discrepancy". Similarly, when the character determination signal 243 is "match" during the non-speech section, the state determination unit 74 determines that the character determination unit 87 has made some incorrect determination and treats it as "discrepancy".

[0080] When the state determination unit 74 determines that it is the target person's speech during the speech section, it determines the update of the filter coefficient in the coefficient update unit 60 according to the first frequency. When the state determination unit 74 determines that it is multiple speech during the speech section, it determines the update of the filter coefficient in the coefficient update unit 60 according to the second frequency. When the state determination unit 74 determines that it is silent during the non-speech section, it determines the update of the filter coefficient in the coefficient update unit 60 according to the third frequency. When the state determination unit 74 determines that it is another person's speech during the non-speech section, it determines the update of the filter coefficient in the coefficient update unit 60 according to the second frequency.

[0081] Figs. 13(a), 13(b), and 13(c) show the time changes of the sound pressure level, the image recognition result, and the state determination result in the microphone 10 and the vibration sensor 14 according to the third embodiment. Fig. 13(a) shows the time changes of the vibration signal 202 and the voice signal 200 in the same manner as Fig. 5(a). The horizontal axis in Fig. 13(a) represents time, and the vertical axis represents the sound pressure level. Fig. 13(b) shows the time change of the speech section information 204, which is the image recognition result in the image recognition unit 82, in the same manner as Fig. 5(b). The horizontal axis in Fig. 5(b) represents time, and the vertical axis represents the value of the speech section information 204. The high-level speech section information 204 indicates "speech", and the low-level speech section information 204 indicates "silence (non-speech)". Fig. 13(c) shows the character determination signal 243, which is the comparison result between the characters included in the detected character information 205 in the character determination unit 87 and the characters included in the character signal 241. The horizontal axis in Fig. 13(c) represents time, and the vertical axis represents the comparison result ("match", "discrepancy", "none").

[0082] Period L11 is the state of the target speaker's speech. The speech section information 204 indicates speech, and the character determination signal 243, which is the comparison result in the character determination unit 87, indicates "match". As a result, the state determination unit 74 identifies the target speaker's speech during period L11. Period L12 is a silent state. The speech section information 204 indicates silence, and the character determination signal 243, which is the comparison result in the character determination unit 87, indicates "none". As a result, the state determination unit 74 identifies silence. Period L13 is the state of the target speaker's speech, similar to period L11, and period L14 is the silent state, similar to period L12.

[0083] Period L15 is the state of another person's speech. The speech section information 204 indicates silence, and the character determination signal 243, which is the comparison result in the character determination unit 87, indicates "mismatch". As a result, the state determination unit 74 identifies another person's speech. Period L16 is the state of overlapping speech. The speech section information 204 indicates speech, and the character determination signal 243, which is the comparison result in the character determination unit 87, indicates "mismatch". As a result, the state determination unit 74 identifies overlapping speech. Return to FIG. 10.

[0084] The adaptation control unit 34 receives the determination result (determination signal 242) from the determination unit 84. If the determination result is an update based on the first frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the first frequency. If the determination result is an update based on the second frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the second frequency. If the determination result is an update based on the third frequency, the adaptation control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the third frequency. That is, the adaptation control unit 34 adjusts the update frequency of the filter coefficient in the coefficient update unit 60 based on the image recognition result that determines whether it is the user's speech section or non-speech section and the audio signal 200.

[0085] FIG. 14 is a flowchart showing a processing procedure by the sound collection device 100 according to Embodiment 3. The imaging device 80 acquires an image signal 240 (S150). The image recognition unit 82 executes image recognition processing on the image signal 240 (S152). If it is a speaking interval (Y in S152), the determination unit 84 causes the microphone 10 to acquire the audio signal 200, and acquires the audio signal 200 that has been digitally-analog converted by the first AD converter 12 (S154). The phoneme analysis unit 86 of the determination unit 84 executes phoneme analysis processing on the audio signal 200 (S156). The character determination unit 87 of the determination unit 84 outputs a character determination signal 243, which is the determination result of the characters included in the detected character information 205, which is the character uttered by the target person as a result of the image recognition processing executed by the image recognition unit 82, and the character signal 241 obtained by the phoneme analysis unit 86 performing phoneme analysis processing on the audio signal 200 (S158). If the character determination signal 243 is not "match" (N in S158), the state determination unit 74 determines that it is a multiple talk (S159), and determines an update at the second frequency (S160). If the character determination signal 243 is "match" (Y in S158), the state determination unit 74 determines that it is a target person's speech (S161), and determines an update at the first frequency (S162).

[0086] If it is not a speaking interval (N in S152), the determination unit 84 causes the microphone 10 to acquire the audio signal 200, and acquires the audio signal 200 that has been digitally-analog converted by the first AD converter 12 (S164). The phoneme analysis unit 86 of the determination unit 84 executes phoneme analysis processing on the audio signal 200 (S166). The character determination unit 87 of the determination unit 84 outputs a character determination signal 243, which is the determination result of the characters included in the detected character information 205, which is the character uttered by the target person as a result of the image recognition processing executed by the image recognition unit 82, and the character signal 241 obtained by the phoneme analysis unit 86 performing phoneme analysis processing on the audio signal 200 (S168). If the character determination signal 243 is not "none" (N in S168), the state determination unit 74 determines that it is another person's speech (S169), and determines an update at the second frequency (S170). If the character determination signal 243 is "none" (Y in S168), the state determination unit 74 determines that it is silent (S171), and determines an update at the third frequency (S172).

[0087] The adaptive filter 18 executes an adaptive filter process according to the control content determined by the state determination unit 74 of the determination unit 84 (S170). If the operation has not ended (N in S172), the process returns to step 150. If the operation has ended (Y in S172), the process is terminated.

[0088] According to this embodiment, based on the image recognition result that determines whether it is the user's speaking interval or non-speaking interval and the audio signal, the update frequency of the filter coefficient is adjusted, so that the influence of audio other than the target sound or noise components can be suppressed. Also, by imaging the face of the target person and detecting the movement of the face (mouth) of the target person by image recognition, the speaking interval of the target person can be determined with high accuracy. Further, since the speaking interval of the target person is determined with high accuracy, the control accuracy of the filter coefficient can be improved.

[0089] Also, when the user's speech is detected based on the audio signal in the speaking interval in the image recognition result, the filter coefficient is updated at the first frequency, so that the filter coefficient can be updated to be closer to the audio signal. Also, when audio other than the user's speech is detected based on the audio signal in the speaking interval in the image recognition result, the filter coefficient is updated at a second frequency smaller than the first frequency, so that the influence of audio other than the target sound or noise components can be suppressed.

[0090] Also, when stationary noise is detected based on the audio signal in the non-speaking interval in the image recognition result, the filter coefficient is updated at a third frequency smaller than the first frequency, so that the influence of noise components can be suppressed. Also, when audio other than the user's speech is detected based on the audio signal in the non-speaking interval in the image recognition result, the filter coefficient is updated at a second frequency smaller than the first frequency, so that the influence of audio other than the target sound or noise components can be suppressed.

[0091] As described above, the present invention has been described based on embodiments. It is understood by those skilled in the art that these embodiments are illustrative, and various modifications are possible for each of these constituent elements and combinations of each processing process, and such modifications are also within the scope of the present invention.

Description of Reference Numerals

[0092] 10 microphones, 12 first AD converter, 14 vibration sensor, 16 second AD converter, 18 adaptive filter, 20 subtractor, 22 DA converter, 32 correlation determination unit, 34 adaptive control unit, 50 delay unit, 52 multiplier, 54 adder, 60 coefficient update unit, 70 threshold setting unit, 72 correlation value calculation unit, 74 state determination unit, 80 imaging device, 82 image recognition unit, 84 determination unit, 86 phoneme analysis unit, 87 character determination unit, 100 sound collection device, 200 voice signal, 202 vibration signal, 204 speech section information, 205 detected character information, 206 correlation information, 208 adaptive control signal, 210 residual signal, 212 converted voice signal, 214 output signal, 216 first threshold, 218 second threshold, 220 correlation value, 224 correlation values for a predetermined number of times, 240 image signal, 241 character signal, 242 determination signal, 243 character determination signal.

Claims

1. An adaptive filter that performs an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal, A coefficient update unit that updates the filter coefficient based on a residual signal that is a difference between the voice signal acquired by the microphone and the converted voice signal, An adaptive control unit that adjusts the update frequency of the filter coefficient in the coefficient update unit based on an image recognition result that determines whether it is a user's speaking section or a non-speaking section and the voice signal, A sound collection device comprising the above.

2. When the adaptive control unit detects the user's speech based on the voice signal in the speaking section in the image recognition result, the adaptive control unit updates the filter coefficient at a first frequency, When the adaptive control unit detects voice other than the user's speech based on the voice signal in the speaking section in the image recognition result, the adaptive control unit updates the filter coefficient at a second frequency, The sound collection device according to claim 1, wherein the second frequency is smaller than the first frequency.

3. When the adaptive control unit detects stationary noise based on the voice signal in the non-speaking section in the image recognition result, the adaptive control unit updates the filter coefficient at a third frequency, When the adaptive control unit detects voice other than the user's speech based on the voice signal in the non-speaking section in the image recognition result, the adaptive control unit updates the filter coefficient at a second frequency, The sound collection device according to claim 2, wherein the third frequency is smaller than the first frequency and larger than the second frequency.

4. Performing an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal; Updating the filter coefficient based on a residual signal that is a difference between the voice signal acquired by the microphone and the converted voice signal; Adjusting the update frequency of the filter coefficient based on an image recognition result that determines whether it is a user's speaking section or a non-speaking section and the voice signal; A sound collection method comprising the above.

5. On a computer, The step in which an adaptive filter performs an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal; The step in which the filter coefficient is updated based on a residual signal that is a difference between the voice signal acquired by the microphone and the converted voice signal; A sound collection program that executes a step of adjusting the frequency of updating the filter coefficients based on an image recognition result that determines whether it is a user's speaking section or a non-speaking section and the audio signal.

Citation Information

Patent Citations

  • Microphone and sound generation method

    JP2007251354A