Sound pick-up device, sound pick-up method, and sound pick-up program

The sound collection device uses an adaptive filter with image recognition to distinguish and suppress non-target audio and noise, enhancing audio quality by accurately filtering out non-target speech components.

JP2025099815APending Publication Date: 2025-07-03JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023216759
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing sound collection devices inaccurately incorporate non-target audio and noise components during speech, particularly when multiple speakers are present, degrading audio quality.

Method used

A sound collection device utilizing an adaptive filter that updates filter coefficients based on image recognition of speaking intervals of the target and non-target speakers, adjusting the frequency of updates to suppress non-target audio and noise components.

Benefits of technology

Enhances audio quality by accurately distinguishing target speech from non-target speech and noise, thereby improving the suppression of non-target audio and noise components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025099815000001_ABST
    Figure 2025099815000001_ABST
Patent Text Reader

Abstract

To provide a technique to prevent the influence of a voice other than a target sound or a noise component.SOLUTION: A sound pick-up device 100 includes an adaptive filter 18, a coefficient update unit, and an adaptive control unit 34. The adaptive filter 18 executes operation with a filter coefficient for a vibration signal 202 acquired in a vibration sensor 14, and outputs a converted voice signal 212. The coefficient update unit updates the filter coefficient on the basis of a residual signal that is the difference between a voice signal 200 acquired in a microphone 10 and the converted voice signal 212. The adaptive control unit 34 adjusts the frequency of update of the filter coefficient in the coefficient update unit on the basis of a first image recognition result of a determination as to whether an utterance section or a non-utterance section of a user, and a second image recognition result of a determination as to whether an utterance section or a non-utterance section of a person other than the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound collection technology, and particularly to a sound collection device, a sound collection method, and a sound collection program using a microphone and a vibration sensor.

Background Art

[0002] A sound collection device includes a microphone that generates an audio signal based on air vibration and a vibration sensor that generates a vibration signal corresponding to the audio signal based on bone vibration, thereby obtaining clear audio in a noisy environment (for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The sound collection device converts the vibration signal generated by the vibration sensor into an audio signal in a filtering unit, approximates the vibration signal to the audio signal that is originally transmitted through air, and provides an easy-to-hear audio. However, when the target speaker speaks, if there is overlapping conversation audio of people other than the speaker, it will also be erroneously adapted to components other than the target sound. Therefore, it is necessary to accurately grasp the case where audio other than the speaker is mixed into the microphone. If it does not have a function of detecting audio components and noise components other than the speaker, when increasing the approximation degree to the microphone audio during the speaking section, there is a risk that the influence of audio or noise components other than the target sound will be erroneously added to the converted audio.

[0005] The present invention has been made in view of such a situation, and its object is to provide a technology for suppressing the influence of audio or noise components other than the target sound.

Means for Solving the Problems

[0006] In order to solve the above problems, a sound collection device according to an aspect of the present invention includes an adaptive filter that performs an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal, a coefficient update unit that updates the filter coefficient based on a residual signal that is a difference between the voice signal acquired by a microphone and the converted voice signal, and an adaptive control unit that adjusts the frequency of updating the filter coefficient in the coefficient update unit based on a first image recognition result for determining whether it is a user's speaking section or a non-speaking section and a second image recognition result for determining whether it is a speaking section or a non-speaking section of a person other than the user.

[0007] Another aspect of the present invention is a sound collection method. This method includes a step of performing an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal, a step of updating the filter coefficient based on a residual signal that is a difference between the voice signal acquired by a microphone and the converted voice signal, and a step of adjusting the frequency of updating the filter coefficient based on a first image recognition result for determining whether it is a user's speaking section or a non-speaking section and a second image recognition result for determining whether it is a speaking section or a non-speaking section of a person other than the user.

[0008] Another aspect of the present invention is a sound collection program. This sound collection program causes a computer to execute a step of performing an operation using a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal, a step of updating the filter coefficient based on a residual signal that is a difference between the voice signal acquired by a microphone and the converted voice signal, and a step of adjusting the frequency of updating the filter coefficient based on a first image recognition result for determining whether it is a user's speaking section or a non-speaking section and a second image recognition result for determining whether it is a speaking section or a non-speaking section of a person other than the user.

[0009] Note that any combination of the above components, as well as those obtained by converting the expression of the present invention among a method, an apparatus, a system, a recording medium, a computer program, etc., are also effective as aspects of the present invention.

Effects of the Invention

[0010] According to the present invention, it is possible to suppress the influence of voices or noise components other than the target sound.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0012] (Example 1) Before specifically describing the present invention, an overview will be first given. This embodiment relates to a sound collection device. The sound collection device includes a microphone that generates an audio signal based on air vibrations, a vibration sensor that generates a vibration signal based on vibrations transmitted to the human body, and an adaptive filter that multiplies the vibration signal by a coefficient to generate a converted audio signal in order to correct the vibration signal to approximate the audio signal. Further, the sound collection device detects the presence or absence of a target sound that is the speech of a target person by image analysis, detects the presence or absence of a non-target sound that is the speech of another person by image analysis, and when detecting the mixing of audio components other than the target sound, stops the update (learning) of the filter coefficient in the adaptive filter. Thereby, it is suppressed that the converted audio signal is approximated to an audio component other than the target sound. Note that the converted audio signal is a digital signal related to audio obtained by performing analog-to-digital conversion on the vibration signal acquired by the vibration sensor and further performing an operation using a predetermined filter coefficient, that is, an audio signal converted based on the vibration signal.

[0013] FIG. 1 shows the configuration of a sound collection device 100 according to Embodiment 1. The sound collection device 100 includes a microphone 10, a first AD converter (Analog to Digital converter) 12, a vibration sensor 14, a second AD converter 16, an adaptive filter 18, a subtractor 20, a DA converter (Digital to Analog converter) 22, an adaptive control unit 34, a state determination unit 74, a target person imaging device 80, a target person image recognition unit 82, an other person imaging device 84, a first other person imaging device 84a collectively referred to as the other person imaging device 84, a second other person imaging device 84b, a first other person image recognition unit 86a collectively referred to as the other person image recognition unit 86, and a second other person image recognition unit 86b. The number of the other person imaging device 84 and the other person image recognition unit 86 is not limited to "2". Note that the target person imaging device 80 and the other person imaging device 84 are imaging devices such as color cameras each including, for example, a lens, an image sensor such as a CMOS (Complementary Metal Oxide Semiconductor), a driver that digitizes a video signal output from the image sensor, a DRAM (Dynamic Random Access Memory) that temporarily stores image data for image processing, and a DSP (Digital Signal Processor).

[0014] The microphone 10 converts air vibrations into an audio signal 200. Since the microphone 10 can accurately acquire the audio signal 200, the audio signal 200 acquired by the microphone 10 is relatively close to the audio signal 200 that a person perceives through the ear. Therefore, by using the audio signal 200 acquired by the microphone 10 as the target value of the vibration signal 202 described later, it is possible to keep the audio quality of the converted audio signal 212 output by the adaptive filter 18 at a high level.

[0015] The first AD converter 12 outputs the audio signal 200 converted into a digital signal (hereinafter, this is also referred to as the "audio signal 200") to the subtracter 20 by performing analog-to-digital conversion on the audio signal 200 acquired by the microphone 10.

[0016] The vibration sensor 14 mainly acquires the speech signal emitted by a person as a vibration signal 202 transmitted to the human body. The vibration sensor 14 may be a microphone that detects sound or vibration, or may be a camera (visual microphone) that detects vibration from pixel variations or color fluctuations between frames of video and acquires an audio signal. Instead of the vibration sensor 14, a receiving device embedded in the body or a microphone in direct contact with the human body may be used, or a device that measures the vibration of the human body surface without contact with the human body, such as a laser displacement meter, may be used.

[0017] The second AD converter 16 outputs the vibration signal 202 converted into a digital signal (hereinafter, this is also referred to as the "vibration signal 202") to the adaptive filter 18 by performing analog-to-digital conversion on the vibration signal 202 acquired by the vibration sensor 14.

[0018] The vibration signal 202 has the characteristics that its frequency band is narrower and has a large bias in the frequency band compared with the audio signal 200. Due to this characteristic, the vibration signal 202 becomes a very muffled voice and sounds significantly different from the original audio signal 200. To improve this, in this embodiment, the audio signal 200 acquired at the same time is used as a reference signal, and the adaptive filter 18 is used to make the vibration signal 202 closer to the audio signal 200.

[0019] The adaptive filter 18 receives the vibration signal 202 from the second AD converter 16 and the residual signal 210 from the subtractor 20. Based on the residual signal 210, the adaptive filter 18 performs an operation on the vibration signal 202 using filter coefficients and outputs the operation result as a converted audio signal 212.

[0020] FIG. 2 shows the configuration of the adaptive filter 18 according to Embodiment 1. The adaptive filter 18 includes a first delay element 50a to an (N - 1)th delay element 50n - 1 collectively referred to as a delay element 50, a first multiplier 52a to an Nth multiplier 52n collectively referred to as a multiplier 52, a first adder 54a to an (N - 1)th adder 54n - 1 collectively referred to as an adder 54, and a coefficient update unit 60.

[0021] The delay element 50 delays the vibration signal 202 and sequentially outputs it. The multiplier 52 multiplies the vibration signal 202 by the filter coefficient. The adder 54 sequentially adds the multiplication results at the multiplier 52. The addition result of the (N - 1)th adder 54n - 1 corresponds to the aforementioned operation result. The subtractor 20 calculates a residual signal 210, which is the difference value between the audio signal 200, which is the reference signal, and the converted audio signal 212, which is the operation result, and outputs it to the adaptive filter 18.

[0022] The coefficient update unit 60 updates the filter coefficient based on the audio signal 200 and the operation result by, for example, executing the LMS (Least Mean Square) algorithm as follows.

Equation

[0023] The DA converter 22 outputs a converted voice signal 212 (hereinafter referred to as the "output signal 214") converted into an analog signal by performing digital-to-analog conversion on the converted voice signal 212 output from the adaptive filter 18.

[0024] Here, when a voice component other than the target voice is mixed in the voice signal 200 acquired by the microphone 10, in the adaptive filter 18, an action of approximating to the voice of others may erroneously work. The voice component other than the target voice is, for example, the voice of a person other than the target person, the voice reproduced by a handset speaker (not shown) at the call destination, and the like. Therefore, in order to improve the accuracy of the process for bringing the vibration signal 202 in the adaptive filter 18 closer to the voice signal 200, it is required to update (learn) the filter coefficient when no voice component other than the target voice is mixed in the voice signal 200 in the speaking section. As a preferable condition for updating (learning) the filter coefficient, it is desirable that it is the speaking section of the user. Note that the speaking section is the time zone in which the speech of the target person is detected.

[0025] In this embodiment, in order to detect whether it is the speaking section or the non-speaking section of the target person (user), the target person imaging device 80 images the face (mouth) of the target person. The target person imaging device 80 outputs the captured video (image) as a target person image signal 240 to the target person image recognition unit 82.

[0026] The target person image recognition unit 82 detects the movement of the mouth of the target person from the target person image signal 240 and specifies the speech section by performing image recognition processing on the target person image signal 240. More specifically, the target person image recognition unit 82 performs face detection and mouth detection by image analysis, and detects the speech section when the amount of movement vectors corresponding to speech exceeds the amount of movement vectors per unit time of the mouth. To detect the speech section from the target person image signal 240, for example, technologies such as the literature: Vol.2011-CVIM-177No.13 "Speech Detection by Extraction and Recognition of Lip Region" may be used. The target person image recognition unit 82 outputs the image recognition result that determines whether it is a speech section or a non-speech section of the target person to the state determination unit 74 as the target person speech section information 242. In order to improve the determination accuracy of the speech section, the target person image recognition unit 82 may regard the vibration signal 202 acquired by the vibration sensor 14 as the audio signal 200 and supplementarily use the audio section determination result by the audio signal 200. The effect of preventing false detection due to the movement of the mouth without speech such as yawning can be obtained.

[0027] Also, in this embodiment, in order to detect the presence or absence of the voice of others other than the target person, the first other person imaging device 84a images the face (mouth) of the first other person. The first other person imaging device 84a outputs the captured video (image) to the first other person image recognition unit 86a as the first other person image signal 244a. Also, the second other person imaging device 84b images the face (mouth) of the second other person. The second other person imaging device 84b outputs the captured video (image) to the second other person image recognition unit 86b as the second other person image signal 244b. That is, the other person imaging device 84 performs the same processing as the target person imaging device 80.

[0028] The first other person image recognition unit 86a detects the movement of the mouth of the first other person from the first other person image signal 244a and specifies the speech section of the first other person by performing image recognition processing on the first other person image signal 244a. The first other person image recognition unit 86a outputs the image recognition result that determines whether it is a speech section or a non-speech section of the first other person to the state determination unit 74 as the first other person speech section information 246a.

[0029] The second other-person image recognition unit 86b executes image recognition processing on the second other-person image signal 244b to detect the mouth movement of the second other person from the second other-person image signal 244b and identify the speech interval of the second other person. The second other-person image recognition unit 86b outputs, as the second other-person speech interval information 246b to the state determination unit 74, the image recognition result that determines whether it is a speech interval or a non-speech interval of the second other person. That is, the other-person image recognition unit 86 executes the same processing as the subject image recognition unit 82. The speech interval also includes the time period when the speech of the first other person or the second other person is detected.

[0030] The state determination unit 74 receives the subject speech interval information 242 from the subject image recognition unit 82, receives the first other-person speech interval information 246a from the first other-person image recognition unit 86a, and receives the second other-person speech interval information 246b from the second other-person image recognition unit 86b. The state determination unit 74 acquires information on whether it is a speech interval or a non-speech interval of the subject from the subject speech interval information 242.

[0031] Also, the state determination unit 74 acquires information on whether it is a speech interval or a non-speech interval of the first other person from the first other-person speech interval information 246a, and acquires information on whether it is a speech interval or a non-speech interval of the second other person from the second other-person speech interval information 246b. The state determination unit 74 generates information on whether it is a speech interval or a non-speech interval of the other person by combining the information on whether it is a speech interval or a non-speech interval of the first other person and the information on whether it is a speech interval or a non-speech interval of the second other person. For example, the state determination unit 74 determines that it is a speech interval of the other person when at least one of the first other person and the second other person is in a speech interval, and determines that it is a non-speech interval of the other person when both the first other person and the second other person are in a non-speech interval.

[0032] The state determination unit 74 determines the states of the target person's speech, multiple speech, other person's speech, and silence based on information on whether it is the target person's speech section or non-speech section, and information on whether it is another person's speech section or non-speech section. Further, the state determination unit 74 determines the control content of the coefficient update unit 60 according to the determined state. The target person's speech is a state where only the target person speaks, the multiple speech is a state of double talk or a state where external noise is mixed in, the other person's speech is a state where only a person other than the target person speaks, and the silence is a state without speech. A table is used for the determination of the state and the determination of the control content.

[0033] FIG. 3 shows the data structure of the table held in the state determination unit 74 according to the first embodiment. When it is the target person's speech section and the other person's non-speech section, the state determination unit 74 determines that it is the state of the target person's speech and determines the activation of the adaptation of the coefficient update unit 60. When it is the target person's speech section and the other person's speech section, the state determination unit 74 determines that it is the state of multiple speech and determines the stop of the adaptation of the coefficient update unit 60. When it is the target person's non-speech section and the other person's non-speech section, the state determination unit 74 determines that it is the state of silence and determines the save of the adaptation of the coefficient update unit 60. When it is the target person's non-speech section and the other person's speech section, the state determination unit 74 determines that it is the state of the other person's speech and determines the stop of the adaptation of the coefficient update unit 60.

[0034] Here, in the case of adaptation activation, "α" in Equation (1) is set to the first value. The first value is determined in advance. In the case of adaptation save, "α" in Equation (1) is set to the second value. The second value is a value smaller than the first value. In the case of adaptation stop, the update of the filter coefficient in the coefficient update unit 60 is stopped.

[0035] Figures 4(a), 4(b), and 4(c) show the time variations of the sound pressure level, the target speaker interval information 242, and the other speaker interval information 246 in the microphone 10 and the vibration sensor 14 according to Example 1. Figure 4(a) shows the time variations of the vibration signal 202 and the audio signal 200. The horizontal axis in Figure 4(a) represents time, and the vertical axis represents the sound pressure level. Figure 4(b) shows the time variation of the target speaker interval information 242. The horizontal axis in Figure 4(b) represents time, and the vertical axis represents the value of the target speaker interval information 242. Figure 4(c) shows the time variation of the other speaker interval information 246. The horizontal axis in Figure 4(c) represents time, and the vertical axis represents the value of the other speaker interval information 246. A high level indicates "speech", and a low level indicates "silence (non-speech)". Here, for clarity of explanation, the other speaker interval information 246 is targeted at one other person.

[0036] Since the period L1 is in the state of the target speaker speaking, in the period L1, both the vibration signal 202 and the audio signal 200 have a sound pressure level that fluctuates violently. Also, in the state of the target speaker speaking, the target speaker interval information 242 indicates a high level, and the other speaker interval information 246 indicates a low level. In the period L2, since it is in a silent state, in the period L2, the average ambient noise level of the vibration signal 202 and the average ambient noise level of the audio signal 200 are obtained. Also, in the silent state, both the target speaker interval information 242 and the other speaker interval information 246 indicate a low level.

[0037] The period L3 is in the state of the target speaker speaking, similar to the period L1. The audio signal 200, the vibration signal 202, the target speaker interval information 242, and the other speaker interval information 246 in the period L3 are the same as those in the case of the period L1. The period L4 is in a silent state, similar to the period L2. The audio signal 200, the vibration signal 202, the target speaker interval information 242, and the other speaker interval information 246 in the period L4 are the same as those in the case of the period L2.

[0038] Since period L5 is a state of others' speech, in period L5, the variation in the sound pressure level of the vibration signal 202 is small, but the sound pressure level of the voice signal 200 varies drastically. Also, in the state of others' speech, the target speaker speech section information 242 indicates a low level, and the others' speech section information 246 indicates a high level. The state of others' speech can also be said to be a state of external noise mixing.

[0039] Period L6 is a state of overlapping speech. Since the vibration signal 202 in period L6 contains only the speech information of the target person himself / herself, the part with intense movement on the vertical axis becomes the speech signal of the target speaker. On the other hand, the voice signal 200 in period L6 contains the speech information of the target person himself / herself and the speech information of others. Also, in the state of overlapping speech, both the target speaker speech section information 242 and the others' speech section information 246 indicate a high level. Period L7 is a state of others' speech, similar to period L5. The voice signal 200, vibration signal 202, target speaker speech section information 242, and others' speech section information 246 in period L7 are the same as those in the case of period L5. Return to FIG. 1.

[0040] The adaptation control unit 34 receives the determination result (determination signal 248) from the state determination unit 74. If the determination result is adaptation active, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to set "α" to the first value. If the determination result is adaptation save, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to set "α" to the second value. If the determination result is stop, the adaptation control unit 34 outputs an adaptation control signal 208 instructing the coefficient update unit 60 to stop the update of the filter coefficient. That is, the adaptation control unit 34 adjusts the degree of update of the filter coefficient in the coefficient update unit 60 based on the vibration signal 202 and the voice signal 200.

[0041] The above configuration can be implemented in terms of hardware by the CPU, memory, and other LSIs of any computer, and in terms of software by a program loaded in the memory or the like. Here, however, functional blocks realized by their cooperation are depicted. Therefore, it is understood by those skilled in the art that these functional blocks can be realized in various forms by hardware only, software only, or a combination thereof.

[0042] The operation of the sound collection device 100 with the above configuration will be described. FIG. 5 is a flowchart showing the processing procedure by the sound collection device 100 according to the first embodiment. The target person image recognition unit 82 executes image recognition processing on the target person image signal 240 received from the target person imaging device 80 (S10). When the result of the image recognition processing indicates a speaking interval (Y in S12), the other person image recognition unit 86 executes image recognition processing on the other person image signal 244 received from the other person imaging device 84 (S14). When the image recognition processing indicates other person's speech (Y in S16), the state determination unit 74 estimates multiple speech (S17) and determines an adaptive stop (S18). When the image recognition processing does not indicate other person's speech (N in S16), the state determination unit 74 estimates target person's speech (S19) and determines an adaptive active (S20).

[0043] When the result of the image recognition processing indicates a non-speaking interval (N in S12), the other person image recognition unit 86 executes image recognition processing on the other person image signal 244 received from the other person imaging device 84 (S22). When the image recognition processing indicates other person's speech (Y in S24), the state determination unit 74 estimates other person's speech (S25) and determines an adaptive stop (S26). When the image recognition processing does not indicate other person's speech (N in S24), the state determination unit 74 estimates silence (S27) and determines an adaptive save (S28).

[0044] The adaptive filter 18 executes adaptive filter processing according to the control content determined by the state determination unit 74 (S30). If the operation has not ended (N in S32), the process returns to step 10. If the operation has ended (Y in S32), the process is terminated.

[0045] According to this embodiment, by approximating the vibration signal to the characteristics of the voice signal, the influence of ambient noise components can be suppressed. Also, the presence of voices other than the speaker's and sudden noises is sequentially detected from the movements of the subject's face (mouth) and the faces (mouths) of others, and the degree of approximation to the voice signal is controlled based on the detection results, so that the influence of voices or noise components other than the target voice can be suppressed. Further, since the influence of voices or noise components other than the target voice is suppressed, the mixing of voice components of others is avoided and the voice quality of the speaker can be improved. Also, by imaging the subject's face and the faces of others and detecting the movements of the subject's face (mouth) and the faces (mouths) of others by image recognition, the speaking intervals of the subject and those of others can be determined with high accuracy. Further, since the speaking intervals of the subject and those of others are determined with high accuracy, the control accuracy of the filter coefficients can be improved.

[0046] (Example 2) Next, Example 2 will be described. Example 2 relates to the sound collection device 100 in the same manner as Example 1. The sound collection device 100 according to Example 1 adjusts any one of adaptive active, adaptive stop, and adaptive save, that is, the degree of update of the filter coefficients, based on the image recognition results of the subject image signal 240 and the image recognition results of the other person image signal 244. On the other hand, the sound collection device 100 according to Example 2 adjusts the frequency of update of the filter coefficients based on the image recognition results of the subject image signal 240 and the image recognition results of the other person image signal 244. The sound collection device 100 and the adaptive filter 18 according to Example 2 are of the same type as those in FIGS. 1 and 2. Here, the description will focus on the differences from Example 1.

[0047] The state determination unit 74 acquires information on whether the target person's section is a speaking section or a non-speaking section from the target person speech section information 242. Also, the state determination unit 74 generates information on whether the other person's section is a speaking section or a non-speaking section from the first other person speech section information 246a and the second other person speech section information 246b. The state determination unit 74 determines the states of target person speech, multiple speech, other person speech, and silence based on the information on whether the target person's section is a speaking section or a non-speaking section and the information on whether the other person's section is a speaking section or a non-speaking section. Also, the state determination unit 74 determines the control content of the coefficient update unit 60 according to the determined state. Similar to the first embodiment, a table is used for state determination and determination of control content.

[0048] FIG. 6 shows the data structure of the table held in the state determination unit 74 according to the second embodiment. When the section is the target person's speaking section and the other person's non-speaking section, the state determination unit 74 determines that it is in the state of target person speech and determines the update of the filter coefficient in the coefficient update unit 60 according to the first frequency. When the section is the target person's speaking section and the other person's speaking section, the state determination unit 74 determines that it is in the state of multiple speech and determines the update of the filter coefficient in the coefficient update unit 60 according to the second frequency. When the section is the target person's non-speaking section and the other person's non-speaking section, the state determination unit 74 determines that it is in the state of silence and determines the update of the filter coefficient in the coefficient update unit 60 according to the third frequency. When the section is the target person's non-speaking section and the other person's speaking section, the state determination unit 74 determines that it is in the state of other person speech and determines the update of the filter coefficient in the coefficient update unit 60 according to the second frequency.

[0049] Here, the second frequency is made smaller than the first frequency, and the third frequency is made smaller than the first frequency and larger than the second frequency. For example, the first frequency is every time, the second frequency is once every 128 times, and the third frequency is between once every 2 times and once every 64 times. The values of each frequency are not limited to these. Return to FIG. 1.

[0050] The adaptive control unit 34 receives the determination result (determination signal 248) from the state determination unit 74. If the determination result is an update based on the first frequency, the adaptive control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the first frequency. If the determination result is an update based on the second frequency, the adaptive control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the second frequency. If the determination result is an update based on the third frequency, the adaptive control unit 34 instructs the coefficient update unit 60 to update the filter coefficient according to the third frequency. That is, based on the first image recognition result that determines whether the target person is in a speaking interval or a non-speaking interval, and the second image recognition result that determines whether a person other than the target person is in a speaking interval or a non-speaking interval, the adaptive control unit 34 adjusts the update frequency of the filter coefficient in the coefficient update unit 60.

[0051] Figure 7 is a flowchart showing the processing procedure by the sound collection device 100 according to the second embodiment. The target person image recognition unit 82 performs image recognition processing on the target person image signal 240 received from the target person imaging device 80 (S50). If the result of the image recognition processing indicates a speaking interval (Y in S52), the other person image recognition unit 86 performs image recognition processing on the other person image signal 244 received from the other person imaging device 84 (S54). If the image recognition processing indicates that the other person is speaking (Y in S56), the state determination unit 74 estimates multiple speaking (S57) and determines to update the filter coefficient at the second frequency (S58). If the image recognition processing does not indicate that the other person is speaking (N in S56), the state determination unit 74 estimates that the target person is speaking (S59) and determines to update the filter coefficient at the first frequency (S60).

[0052] When the result of the image recognition process indicates a non-speaking interval (N in S52), the other-person image recognition unit 86 executes an image recognition process on the other-person image signal 244 received from the other-person imaging device 84 (S62). When the image recognition process indicates other-person speech (Y in S64), the state determination unit 74 estimates other-person speech (S65) and determines to update the filter coefficient at the second frequency (S66). When the image recognition process does not indicate other-person speech (N in S64), the state determination unit 74 estimates silence (S67) and determines to update the filter coefficient at the second frequency (S68).

[0053] The adaptive filter 18 executes an adaptive filter process according to the control content determined by the state determination unit 74 (S70). If the operation has not ended (N in S72), the process returns to step 50. If the operation has ended (Y in S72), the process is terminated.

[0054] According to this embodiment, since the update frequency of the filter coefficient is adjusted based on the movement of the face (mouth) of the target person and the movement of the face (mouth) of others, the influence of voices or noise components other than the target sound can be suppressed. Also, when the image recognition result for the target person is the speaking section and the image recognition result for others is the non-speaking section, the filter coefficient is updated at the first frequency, so the filter coefficient can be updated to be closer to the voice signal. Further, when the image recognition result for the target person is the speaking section and the image recognition result for others is also the speaking section, the filter coefficient is updated at a second frequency smaller than the first frequency, so the influence of voices or noise components other than the target sound can be suppressed. Also, when the image recognition result for the target person is the non-speaking section and the image recognition result for others is the non-speaking section, the filter coefficient is updated at a third frequency smaller than the first frequency, so the influence of the noise component can be suppressed. Moreover, when the image recognition result for the target person is the non-speaking section and the image recognition result for others is the speaking section, the filter coefficient is updated at a second frequency smaller than the first frequency, so the influence of voices or noise components other than the target sound can be suppressed. Additionally, by imaging the face of the target person and the face of others and detecting the movement of the face (mouth) of the target person and the movement of the face (mouth) of others through image recognition, the speaking section of the target person and the speaking section of others can be determined with high accuracy. Also, since the speaking section of the target person and the speaking section of others are determined with high accuracy, the control accuracy of the filter coefficient can be improved.

[0055] As described above, the present invention has been described based on the embodiments. It is understood by those skilled in the art that these embodiments are illustrative, and various modifications are possible for the combination of each of these constituent elements and each processing process, and such modifications are also within the scope of the present invention.

[0056] In the sound collection devices 100 in Embodiments 1 and 2, the subject imaging device 80 and one or more other person imaging devices 84 are included, and accordingly, the subject image recognition unit 82 and one or more other person image recognition units 86 are included. However, the present invention is not limited thereto. For example, instead of the subject imaging device 80 and one or more other person imaging devices 84, one imaging device may be included, and instead of the subject image recognition unit 82 and one or more other person image recognition units 86, one image recognition unit may be included. The imaging device acquires an image signal including the subject and one or more other persons and outputs it to the image recognition unit. The image recognition unit extracts the subject and one or more other persons from the image signal by a known technique, and executes the same processing as the subject image recognition unit 82 and one or more other person image recognition units 86 on the extracted subject and one or more other persons. According to this modification example, the degree of freedom in configuration can be improved.

Explanation of Reference Numerals

[0057] 10 Microphone, 12 First AD Converter, 14 Vibration Sensor, 16 Second AD Converter, 18 Adaptive Filter, 20 Subtractor, 22 DA Converter, 34 Adaptive Control Unit, 50 Delay Unit, 52 Multiplier, 54 Adder, 60 Coefficient Update Unit, 74 State Determination Unit, 80 Subject Imaging Device, 82 Subject Image Recognition Unit, 84 Other Person Imaging Device, 86 Other Person Image Recognition Unit, 100 Sound Collection Device, 200 Audio Signal, 202 Vibration Signal, 208 Adaptive Control Signal, 210 Residual Signal, 212 Converted Audio Signal, 214 Output Signal, 240 Subject Image Signal, 242 Subject Speech Interval Information, 244 Other Person Image Signal, 246 Other Person Speech Interval Information, 248 Determination Signal.

Claims

1. An adaptive filter that performs an operation based on a filter coefficient on a vibration signal acquired by a vibration sensor and outputs a converted voice signal, A coefficient update unit that updates the filter coefficient based on a residual signal that is a difference between a voice signal acquired by a microphone and the converted voice signal, An adaptive control unit that adjusts the frequency of updating the filter coefficient in the coefficient update unit based on a first image recognition result that determines whether it is a speaking section or a non-speaking section of the user and a second image recognition result that determines whether it is a speaking section or a non-speaking section of a person other than the user, A sound collection device comprising the same.

2. When the first image recognition result is the speaking section and the second image recognition result is the non-speaking section, the adaptive control unit updates the filter coefficient at a first frequency, When the first image recognition result is the speaking section and the second image recognition result is the speaking section, the adaptive control unit updates the filter coefficient at a second frequency, The sound collection device according to Claim 1, wherein the second frequency is smaller than the first frequency.

3. When the first image recognition result is the non-speaking section and the second image recognition result is the non-speaking section, the adaptive control unit updates the filter coefficient at a third frequency, When the first image recognition result is the non-speaking section and the second image recognition result is the speaking section, the adaptive control unit updates the filter coefficient at the second frequency, The sound collection device according to Claim 2, wherein the third frequency is smaller than the first frequency and larger than the second frequency.

4. A step of performing an operation based on a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal, A step of updating the filter coefficient based on a residual signal that is a difference between a voice signal acquired by a microphone and the converted voice signal, A step of adjusting the frequency of updating the filter coefficient based on a first image recognition result that determines whether it is a speaking section or a non-speaking section of the user and a second image recognition result that determines whether it is a speaking section or a non-speaking section of a person other than the user, A sound collection method comprising the same.

5. On a computer, A step of performing an operation based on a filter coefficient on a vibration signal acquired by a vibration sensor and outputting a converted voice signal, Updating the filter coefficients based on a residual signal that is the difference between the audio signal acquired by the microphone and the converted audio signal; A sound collection program that causes execution of a step of adjusting the frequency of updating the filter coefficients based on a first image recognition result that determines whether it is a speaking section or a non-speaking section of the user, and a second image recognition result that determines whether it is a speaking section or a non-speaking section of another person other than the user.

Citation Information

Patent Citations

  • Microphone and sound generation method

    JP2007251354A