Speech processing system, speech processing apparatus, and speech processing method

Through the multi-microphone system and adaptive filter determination and control, the problem of peripheral voice removal when the number of microphones is insufficient is solved, and efficient target voice extraction and processing volume reduction is achieved.

CN115299074BActive Publication Date: 2025-07-22PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180021337.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-18
Filing Date
2021-02-10
Publication Date
2025-07-22
Estimated Expiration
2041-02-10

AI Technical Summary

Technical Problem

When the number of radio equipment is not enough to match the number of sound sources, it is difficult for the prior art to effectively remove peripheral voices and obtain target voices, and the processing volume is relatively large.

Method used

Multiple microphones are used to obtain the voice signal, and the adaptive filter and determination components are used to determine which voice component is mainly included in the voice signal, and the filter coefficients of the adaptive filter are controlled to remove peripheral voices and reduce the processing amount.

Benefits of technology

Even when the number of microphones is insufficient, peripheral voice can be removed efficiently and target voice can be obtained, reducing processing volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115299074B_ABST
    Figure CN115299074B_ABST
Patent Text Reader

Abstract

The voice processing system includes: at least one first microphone that acquires a first voice signal and outputs a first signal based on the first voice signal, where the first voice signal includes at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; at least one adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; a determination unit that determines which of the first voice component and the second voice component is more included in the first voice signal; and a control unit that controls the filter coefficients of the adaptive filter based on the determination result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a voice processing system, a voice processing device, and a voice processing method. Background Art

[0002] An echo canceller is known that removes surrounding voices and only recognizes the voice of a speaker in a vehicle-mounted voice recognition device or a hands-free call. Patent Document 1 discloses an echo canceller that switches the number of adaptive filters and the number of taps that operate according to the number of sound sources.

[0003] Prior Art Documents

[0004] Patent Documents

[0005] Patent Document 1: Japanese Patent No. 4889810 Gazette Summary of the Invention

[0006] In the case of performing echo cancellation using an adaptive filter, the surrounding voice picked up by a sound collection device is input to the adaptive filter as a reference signal. For example, there is a sound collection device corresponding to each sound source capable of emitting a voice. When one reference signal is output from one sound collection device, the voice included in the reference signal can be determined as the voice generated at the position of the sound source corresponding to the sound collection device that output the reference signal. The target voice can be obtained by subtracting the reference signal from the signal including the target voice while taking into account the generation position of the surrounding voice included in the reference signal.

[0007] On the other hand, when the number of sound collection devices is less than the number of sound sources capable of emitting voices, the voice of multiple sound sources may be included in one reference signal. In this case, it is not possible to determine the position where the voice included in the reference signal is generated only based on the reference signal. Therefore, it is sometimes difficult to remove the surrounding voice to obtain the target voice. It would be beneficial if the surrounding voice could be removed to obtain the target voice even when the number of sound collection devices is less than the number of sound sources. In addition, it is beneficial if the processing amount can be reduced in the process of removing the surrounding voice to obtain the target voice.

[0008] The present disclosure relates to a voice processing system, a voice processing device, and a voice processing method that can solve at least one of the above problems in echo cancellation using an adaptive filter.

[0009] One aspect of the present disclosure relates to a voice processing system including: at least one first microphone that acquires a first voice signal and outputs a first signal based on the first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; at least one adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; a determination unit that determines which of the first voice component and the second voice component is included more in the first voice signal; and a control unit that controls filter coefficients of the adaptive filter based on the determination result.

[0010] One aspect of the present disclosure relates to a voice processing apparatus including: at least one receiving unit that receives a first signal based on a first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; at least one adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; a determination unit that determines which of the first voice component and the second voice component is included more in the first voice signal; and a control unit that controls filter coefficients of the adaptive filter based on the determination result.

[0011] One aspect of the present disclosure relates to a voice processing method including the following steps: receiving a first signal based on a first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; inputting the first signal into at least one adaptive filter, the at least one adaptive filter outputting a passed signal based on the first signal; determining which of the first voice component and the second voice component is included more in the first voice signal; and controlling filter coefficients of the adaptive filter based on the determination result.

[0012] In addition, these general or specific aspects can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium, or can be implemented by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0013] According to the present disclosure, even when the number of sound collection devices is smaller than the number of sound sources that can emit voices, it is possible to remove surrounding voices to obtain a target voice. Alternatively, according to the present disclosure, in a process for removing surrounding voices to obtain a target voice, it is possible to reduce the processing amount. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1This is a diagram showing an example of the schematic structure of the speech processing system in the first embodiment.

[0015] Figure 2 This is a block diagram showing the structure of the speech processing device in the first embodiment.

[0016] Figure 3A This is a diagram showing the time waveform of the speech signal (speech signal C) used in the speech processing device.

[0017] Figure 3B This is a diagram showing the time waveform of the speech signal (first directivity signal) used in the speech processing device.

[0018] Figure 3C This is a diagram showing the time waveform of the speech signal (second directivity signal) used in the speech processing device.

[0019] Figure 4 This is a diagram showing the spectrum of the speech signal used in the speech processing device averaged.

[0020] Figure 5 This is a flowchart showing the operation process of the speech processing device in the first embodiment.

[0021] Figure 6 This is a diagram showing an example of the schematic structure of the speech processing system in the second embodiment.

[0022] Figure 7 This is a block diagram showing the structure of the speech processing device in the second embodiment.

[0023] Figure 8 This is a flowchart showing the operation process of the speech processing device in the second embodiment.

[0024] Figure 9 This is a diagram showing an example of the schematic structure of the speech processing system in the third embodiment.

[0025] Figure 10 This is a block diagram showing the structure of the speech processing device in the third embodiment.

[0026] Figure 11 This is a flowchart showing the operation process of the speech processing device in the third embodiment.

[0027] Figure 12 This is a diagram showing an example of the schematic structure of the speech processing system in the fourth embodiment.

[0028] Figure 13 This is a block diagram showing the structure of the speech processing device in the fourth embodiment.

[0029] Figure 14It is a flowchart showing the operation process of the voice processing device in the fourth embodiment.

[0030] Figure 15A It is a diagram showing an example of the spectrum of a voice signal (first directivity signal) used in the voice processing device.

[0031] Figure 15B It is a diagram showing an example of the spectrum of a voice signal (second directivity signal) used in the voice processing device.

[0032] Figure 15C It is a diagram showing an example of the spectrum of voice signal C used in the voice processing device.

[0033] Figure 15D It is a diagram showing an example of the spectrum of the output signal of the voice processing device.

[0034] Figure 16 It is a diagram showing an example of the schematic structure of the voice processing system in the fifth embodiment.

[0035] Figure 17 It is a block diagram showing the structure of the voice processing device in the fifth embodiment.

[0036] Figure 18 It is a flowchart showing the operation process of the voice processing device in the fifth embodiment.

[0037] Figure 19 It is a diagram showing an example of the schematic structure of the voice processing system in the sixth embodiment.

[0038] Figure 20 It is a block diagram showing the structure of the voice processing device in the sixth embodiment.

[0039] Figure 21 It is a flowchart showing the operation process of the voice processing device in the sixth embodiment. Detailed Embodiments

[0040] Hereinafter, embodiments of the present disclosure will be described in detail with appropriate reference to the drawings. However, sometimes the description of more details than necessary is omitted. In addition, the drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter described in the claims by them.

[0041] (First Embodiment)

[0042] Figure 1FIG. 0 is a diagram showing an example of the schematic configuration of the voice processing system 5 in the first embodiment. The voice processing system 5 is mounted on, for example, a vehicle 10. Hereinafter, an example in which the voice processing system 5 is mounted on the vehicle 10 will be described. A plurality of seats are provided in the passenger compartment of the vehicle 10. The plurality of seats are, for example, four seats including a driver's seat, a front passenger seat, and left and right rear seats. The right seat in the rear seats is an example of the first position. The left seat in the rear seats is an example of the second position. The number of seats is not limited to this. The voice processing system 5 includes a microphone MC1, a microphone MC2, a microphone MC3, and a voice processing device 20. The output of the voice processing device 20 is input to a voice recognition engine (not shown). The voice recognition result of the voice recognition engine is input to the electronic device 50.

[0043] The microphone MC1 picks up the voice spoken by the driver hm1. In other words, the microphone MC1 acquires a voice signal including the voice component spoken by the driver hm1. The microphone MC1 is, for example, disposed on the right side of the overhead console. The microphone MC2 picks up the voice spoken by the occupant hm2. In other words, the microphone MC2 acquires a voice signal including the voice component spoken by the occupant hm2. The microphone MC2 is, for example, disposed on the left side of the overhead console. The microphone MC3 picks up the voices spoken by the occupants hm3 and hm4. In other words, the microphone MC3 acquires a voice signal including the voice components spoken by the occupants hm3 and hm4. The microphone MC3 is, for example, disposed near the center of the rear seats on the ceiling. The microphone MC1 is located farther from the right seat in the rear seats than the microphone MC3. The microphone MC2 is located farther from the left seat in the rear seats than the microphone MC3.

[0044] The arrangement positions of the microphones MC1, MC2, and MC3 are not limited to the examples described. For example, the microphone MC1 may be disposed on the front surface of the right side of the instrument panel. The microphone MC2 may be disposed on the front surface of the left side of the instrument panel.

[0045] Each microphone may be a directional microphone or an omnidirectional microphone. Each microphone may be a small MEMS (Micro Electro Mechanical Systems) microphone or an ECM (Electret Condenser Microphone). Each microphone may also be a microphone capable of beamforming. For example, each microphone may be a microphone array that is directional in the direction of each seat and capable of picking up voices in the direction of the pointing method.

[0046] In the present embodiment, the voice processing system 5 includes a plurality of voice processing devices 20 corresponding to respective microphones. Specifically, the voice processing system 5 includes a voice processing device 21, a voice processing device 22, and a voice processing device 23. The voice processing device 21 corresponds to the microphone MC1. The voice processing device 22 corresponds to the microphone MC2. The voice processing device 23 corresponds to the microphone MC3. Hereinafter, the voice processing device 21, the voice processing device 22, and the voice processing device 23 may be collectively referred to as the voice processing device 20.

[0047] In Figure 1 the structure shown, it is exemplified that the voice processing device 21, the voice processing device 22, and the voice processing device 23 are constituted by respective different hardware, but the functions of the voice processing device 21, the voice processing device 22, and the voice processing device 23 may also be realized by one voice processing device 20. Alternatively, a part of the voice processing device 21, the voice processing device 22, and the voice processing device 23 may be constituted by common hardware, and the remaining parts may be constituted by respective different hardware.

[0048] In the present embodiment, each voice processing device 20 is arranged in each seat near the corresponding microphone. For example, the voice processing device 21 is arranged in the driver's seat, the voice processing device 22 is arranged in the front passenger seat, and the voice processing device 23 is arranged in the rear seat. Each voice processing device 20 may also be arranged in the instrument panel.

[0049] Figure 2 is a block diagram showing the structure of the voice processing system 5 and the structure of the voice processing device 21. As Figure 2 shown, in addition to including the voice processing device 21, the voice processing device 22, and the voice processing device 23, the voice processing system 5 further includes a voice recognition engine 40 and an electronic device 50. The output of the voice processing device 20 is input to the voice recognition engine 40. The voice recognition engine 40 recognizes the voice included in the output signal from at least one voice processing device 20 and outputs a voice recognition result. The voice recognition engine 40 generates a voice recognition result and a signal based on the voice recognition result. The signal based on the voice recognition result is, for example, an operation signal of the electronic device 50. The voice recognition result of the voice recognition engine 40 is input to the electronic device 50. The voice recognition engine 40 may also be a device separate from the voice processing device 20. The voice recognition engine 40 is, for example, arranged inside the instrument panel. The voice recognition engine 40 may also be arranged by being housed inside the seat. Alternatively, the voice recognition engine 40 may be an integrated device embedded in the voice processing device 20.

[0050] A signal output from the speech recognition engine 40 is input to the electronic device 50. The electronic device 50 performs an action corresponding to the operation signal, for example. The electronic device 50 is disposed on the instrument panel of the vehicle 10, for example. The electronic device 50 is, for example, a car navigation device. The electronic device 50 may also be a panel meter, a television, or a portable terminal.

[0051] In Figure 1 , a situation where 4 people are riding in the vehicle is shown, but the number of passengers is not limited to this. The number of passengers only needs to be less than or equal to the maximum load capacity of the vehicle. For example, when the maximum load capacity of the vehicle is 6 people, the number of passengers can be 6 people or less than 5 people.

[0052] The speech processing device 21, the speech processing device 22, and the speech processing device 23 have the same structure and function except for a part of the structure of the filter unit described later. Here, the speech processing device 21 will be described. The speech processing device 21 targets the speech spoken by the driver hm1. Here, the target component is synonymous with the acquisition of the target speech signal. The speech processing device 21 outputs, as an output signal, a speech signal obtained by suppressing the crosstalk component from the speech signal picked up by the microphone MC1. Here, the crosstalk component refers to a noise component including the speech of passengers other than the passenger who speaks the speech that is the target component.

[0053] As Figure 2 shown, the speech processing device 21 includes a speech input unit 29, a directivity control unit 30, a filter unit F1 including a plurality of adaptive filters, a control unit 28 that controls the filter coefficients of the plurality of adaptive filters, and an adder 27.

[0054] The microphones MC1, MC2, and MC3 pick up speech respectively, and output signals of speech signals based on the picked-up speech to the speech input unit 29. The speech signals of the speech picked up by the microphones MC1, MC2, and MC3 are input to the speech input unit 29.

[0055] The microphone MC1 outputs a speech signal A to the speech input unit 29. The speech signal A is a signal including the speech of the driver hm1 and noise, and the noise includes the speech of passengers other than the driver hm1. Here, in the speech processing device 21, the speech of the driver hm1 is the target component, and the noise including the speech of passengers other than the driver hm1 is the crosstalk component. The microphone MC1 corresponds to the second microphone. The speech picked up by the microphone MC1 corresponds to the second speech signal. The speech of passengers other than the driver hm1 includes at least one of the speech emitted by the passenger hm3 and the speech emitted by the passenger hm4. The speech signal A corresponds to the second signal.

[0056] The microphone MC2 outputs a voice signal B to the voice input unit 29. The voice signal B is a signal including the voice of the occupant hm2 and noise, and the noise includes the voices of occupants other than the occupant hm2. The microphone MC2 corresponds to the third microphone. The voice picked up by the microphone MC2 corresponds to the third voice signal. The voices of occupants other than the occupant hm2 include at least one of the voice emitted by the occupant hm3 and the voice emitted by the occupant hm4. The voice signal B corresponds to the third signal.

[0057] The microphone MC3 outputs a voice signal C to the voice input unit 29. The voice signal C is a signal including the voice of the occupant hm3, the voice of the occupant hm4, and noise, and the noise includes the voices of occupants other than the occupant hm3 and the occupant hm4. The microphone MC3 corresponds to the first microphone. The voice picked up by the microphone MC3 corresponds to the first voice signal. The voice emitted by the occupant hm3 corresponds to the first voice component, and the voice emitted by the occupant hm4 corresponds to the second voice component. The voice signal C corresponds to the first signal.

[0058] The voice input unit 29 outputs the voice signal A, the voice signal B, and the voice signal C. The voice input unit 29 corresponds to the receiving unit.

[0059] In the present embodiment, the voice processing device 21 includes one voice input unit 29 that receives the voice signals from all the microphones, but may also include a voice input unit 29 that receives the corresponding voice signal for each microphone. For example, the structure may be as follows: the voice signal of the voice picked up by the microphone MC1 is input to the voice input unit corresponding to the microphone MC1, the voice signal of the voice picked up by the microphone MC2 is input to another voice input unit corresponding to the microphone MC2, and the voice signal of the voice picked up by the microphone MC3 is input to another voice input unit corresponding to the microphone MC3.

[0060] The voice signals A, B, and C output from the voice input unit 29 are input to the directivity control unit 30. The directivity control unit 30 performs directivity control processing using the voice signals A and B. The directivity control processing is, for example, processing for generating a voice signal that more includes the sound in the target direction based on the voice signal. The directivity control processing is, for example, beamforming. Further, the directivity control unit 30 outputs a first directivity signal obtained by performing directivity control processing on the voice signal A. For example, the directivity control unit 30 performs directivity control processing on the voice signal A so that it more includes the sound in the direction from the microphone MC1 toward the driver's seat, thereby obtaining the first directivity signal. In addition, the directivity control unit 30 outputs a second directivity signal obtained by performing directivity control processing on the voice signal B. For example, the directivity control unit 30 performs directivity control processing on the voice signal B so that it more includes the sound in the direction from the microphone MC2 toward the passenger seat, thereby obtaining the second directivity signal.

[0061] In addition, the directivity control unit 30 includes a determination unit 35. The determination unit 35 determines whether a voice component is input to the microphone MC3. For example, when the intensity of the voice signal C is greater than at least one of the intensities of the first directivity signal and the second directivity signal, the determination unit 35 determines that a voice signal is input to the microphone MC3; otherwise, it determines that no voice signal is input to the microphone MC3.

[0062] In addition, the determination unit 35 determines which of the voices of the occupant hm3 and the occupant hm4 is more included in the voice signal C. In the present embodiment, the determination unit 35 determines which of the voices of the occupant hm3 and the occupant hm4 is more included in the voice signal C based on the first directivity signal and the second directivity signal. In other words, the determination unit 35 determines which of the voices of the occupant hm3 and the occupant hm4 is more included in the voice signal C based on the voice signal A and the voice signal B. For example, when the occupant hm3 is speaking and the occupant hm4 is not speaking, the voice of the occupant hm3 is included in the voice signal C and the voice of the occupant hm4 is not included. However, it is difficult to determine which of the voices of the occupant hm3 and the occupant hm4 is included only from the voice signal C. Therefore, the determination unit 35 determines which of the voices of the occupant hm3 and the occupant hm4 is more included in the voice signal C by the following method. Here, "the voice signal C more includes the voice of the occupant hm3" also includes the following case: the voice signal C includes the voice of the occupant hm3 and does not include the voice of the occupant hm4. For example, the determination unit 35 compares the intensity of the first directivity signal with the intensity of the second directivity signal. Then, if the intensity of the first directivity signal is greater than the intensity of the second directivity signal, the determination unit 35 determines that the voice signal C more includes the voice of the occupant hm3. Or, if the intensity of the second directivity signal is greater than the intensity of the first directivity signal, the determination unit 35 determines that the voice signal C more includes the voice of the occupant hm4. The determination unit 35 may also determine which voice is more included in the voice signal C based on the intensity of the first directivity signal and the intensity of the second directivity signal at the time when the voice signal C is the largest. The intensity of a signal is sometimes also referred to as the magnitude or level of the signal.

[0063] In the present embodiment, the determination unit 35 included in the directivity control unit 30 determines whether a voice component is input to the microphone MC3, and which of the voices emitted by the occupant hm3 and the voice emitted by the occupant hm4 is more included in the voice signal C. However, the voice processing device 21 may also include a determination unit 35 separate from the directivity control unit 30. In this case, the determination unit 35 is connected, for example, between the voice input unit 29 and the directivity control unit 30. For example, the function of the determination unit 35 is implemented by a processor executing a program stored in a memory. The function of the determination unit 35 may also be implemented by hardware. Alternatively, the voice processing device 21 may only include the determination unit 35 and not include the directivity control unit 30. For example, the determination unit 35 may determine that a voice signal is input to the microphone MC3 when the intensity of the voice signal C is greater than at least one of the intensity of the voice signal A and the intensity of the voice signal B, and otherwise determine that no voice signal is input to the microphone MC3. Further, for example, the determination unit 35 may also determine which of the voices emitted by the occupant hm3 and the voice emitted by the occupant hm4 is more included in the voice signal C based on the voice signal A and the voice signal B.

[0064] Here, the reason for determining which occupant's voice is more included in the voice signal C by comparing the intensity of the first directivity signal and the intensity of the second directivity signal will be described. The voice emitted by the occupant hm3 at the seat on the right side of the rear seat travels forward, and thus is also picked up by the microphones MC1 and MC2. Among the distance between the seat on the right side of the rear seat and the microphone MC1 and the distance between the seat on the right side of the rear seat and the microphone MC2, the latter is larger. Therefore, the voice emitted by the occupant hm3 attenuates more before being picked up by the microphone MC2. Further, when the directivity control unit 30 performs directivity control processing on the voice signal A, for example, processing is performed to more include the sound in the direction from the microphone MC1 toward the driver's seat. The arrival direction of the voice emitted by the occupant hm3 with respect to the microphone MC1 is closer to the direction from the microphone MC1 toward the driver's seat than the arrival direction of the voice emitted by the occupant hm4 with respect to the microphone MC1. Therefore, when the occupant hm3 speaks, the intensity of the first directivity signal is greater than the intensity of the second directivity signal.

[0065] It can be said that the same applies to the voice emitted by the occupant hm4. That is, the distance between the seat on the left side of the rear seat and the microphone MC1 is greater than the distance between the seat on the left side of the rear seat and the microphone MC2. Therefore, the voice emitted by the occupant hm4 attenuates more before being picked up by the microphone MC1. The arrival direction of the voice emitted by the occupant hm4 with respect to the microphone MC2 is closer to the direction from the microphone MC2 towards the co-pilot seat than the arrival direction of the voice emitted by the occupant hm3 with respect to the microphone MC2. Therefore, when the occupant hm4 speaks, the intensity of the second directivity signal is greater than the intensity of the first directivity signal.

[0066] Use Figure 3A 、 Figure 3B 、 Figure 3C and Figure 4 , to specifically illustrate the determination of which occupant's voice is more contained in the voice signal C. Figure 3A 、 Figure 3B and Figure 3C are the time waveforms of the voice signal C, the first directivity signal, and the second directivity signal output from the directivity control unit 30, respectively. The vertical axis represents time, and the horizontal axis represents amplitude. Figure 3A Two peaks in the time waveform shown are enclosed by a dashed line. In addition, for positions approximately the same as the peaks enclosed by the dashed line in Figure 3A , they are also enclosed by a dashed line in Figure 3B and Figure 3C . By comparing the parts enclosed by the dashed line, it can be seen that: peaks also appear at positions in Figure 3B and Figure 3C that are the same as the peaks appearing in Figure 3A ; and the peaks appearing in Figure 3C are larger than the peaks appearing in Figure 3B . Therefore, it can be seen that: compared with the first directivity signal, the component caused by the voice signal C is more contained in the second directivity signal.

[0067] The result of averaging the spectra of the time waveforms shown in Figure 3B and Figure 3C is Figure 4 . In Figure 4 , the solid line shows the spectrum of the intensity of the first directivity signal, and the dashed line shows the spectrum of the intensity of the second directivity signal. In the example shown in Figure 4 , when calculating the root mean square value of the intensity within a specified time range, the second directivity signal is 3.5 dB larger than the first directivity signal. In this example, it is determined that the voice signal C contains more of the voice emitted by the occupant hm4.

[0068] The determination method of whether the voice signal C contains more of the voice emitted by the occupant hm3 or the voice emitted by the occupant hm4 is not limited to the above method. For example, it may also be that the vehicle 10 has seating information related to whether there is an occupant in each seat, and the determination unit 35 makes a determination based on the seating information received from the vehicle 10. For example, when receiving seating information from the vehicle 10 that there is an occupant in the seat on the right side of the rear seat and there is no occupant in the seat on the left side of the rear seat, the determination unit 35 may determine that the voice signal C contains more of the voice emitted by the occupant hm3.

[0069] Alternatively, it may also be that the vehicle 10 is equipped with a camera for photographing each occupant and an image analysis unit for analyzing the images captured by the camera, and the determination unit 35 makes a determination based on the image analysis results of the image analysis unit. For example, when receiving an image analysis result from the image analysis unit that the occupant hm3 has an open mouth and the occupant hm4 has a closed mouth in the image, the determination unit 35 may determine that the voice signal C contains more of the voice emitted by the occupant hm3.

[0070] Alternatively, the determination unit 35 may also make a determination based on the just-mentioned determination result. For example, when it is determined that the voice signal C contains more of the voice emitted by the occupant hm3, it may continue to be determined that the voice signal C contains more of the voice emitted by the occupant hm3 until the intensity of the voice signal C becomes below a certain level. This is because: when speaking continuously, the possibility that the same occupant continues to speak is high.

[0071] The determination unit 35 outputs the result of determining whether a voice component is input to the microphone MC3 and the result of determining which of the voices of the occupant hm3 and the voice of the occupant hm4 is more included in the voice signal C to the control unit 28. The determination unit 35 outputs the determination result to the control unit 28 as a flag, for example. The flag represents a value of "0" or "1". "0" indicates that no voice component is input to the microphone MC3, and "1" indicates that a voice component is input to the microphone MC3. Alternatively, "0" indicates that the voice signal C more includes the voice of the occupant hm3, and "1" indicates that the voice signal C more includes the voice of the occupant hm4. For example, when the voice signal C more includes the voice of the occupant hm3, the determination unit 35 outputs the flag "1, 0" as the determination result to the control unit 28. The first flag among the two flags in this example represents the result of determining whether a voice component is input to the microphone MC3, and the second flag represents the result of determining which occupant's voice the voice signal more includes. It is also possible that the determination unit 35 can determine the case where the voice signal C more includes the voice of the occupant hm3, the case where the voice signal C more includes the voice of the occupant hm4, and the case where the voice signal C equally includes the voice of the occupant hm3 and the voice of the occupant hm4. The determination unit 35 may also output the result of determining whether a voice component is input to the microphone MC3 and the result of determining which of the voices of the occupant hm3 and the voice of the occupant hm4 is more included in the voice signal C at the same time. Alternatively, the determination unit 35 may output the result of determining whether a voice component is input at the time point when the determination of whether a voice component is input to the microphone MC3 is completed, and then output the result of determining which occupant's voice the voice signal more includes at the time point when the determination of which occupant's voice the voice signal more includes is completed.

[0072] In addition, the directivity control unit 30 outputs the first directivity signal to the adder 27, and outputs the second directivity signal and the voice signal C to the filter unit F1.

[0073] The filter unit F1 includes an adaptive filter F1A, an adaptive filter F1B, and an adaptive filter F1C. An adaptive filter refers to a filter having a function of changing characteristics during signal processing. The filter unit F1 is used for the process of suppressing crosstalk components other than the voice of the driver hm1 included in the voice picked up by the microphone MC1. In the present embodiment, the filter unit F1 includes three adaptive filters, but the number of adaptive filters can be appropriately set based on the number of input voice signals and the processing amount of crosstalk suppression processing. Details of the process of suppressing crosstalk will be described later.

[0074] The second directional signal is input to the adaptive filter F1A as a reference signal. The adaptive filter F1A outputs a passed signal P1A based on the filter coefficient C1A and the second directional signal. When it is determined that the voice signal C contains a relatively large amount of the voice emitted by the occupant hm3, the voice signal C is input to the adaptive filter F1B as a reference signal. The adaptive filter F1B outputs a passed signal P1B based on the filter coefficient C1B and the voice signal C. On the other hand, when it is determined that the voice signal C contains a relatively large amount of the voice emitted by the occupant hm4, the voice signal C is input to the adaptive filter F1C as a reference signal. When the determination unit 35 can determine the case where the voice signal C contains a relatively large amount of the voice emitted by the occupant hm3, the case where the voice signal C contains a relatively large amount of the voice emitted by the occupant hm4, and the case where the voice signal C contains the voice emitted by the occupant hm3 and the voice emitted by the occupant hm4 to the same extent, the filter unit F1 may also include an adaptive filter F1D. When it is determined that the voice signal C contains the voice emitted by the occupant hm3 and the voice emitted by the occupant hm4 to the same extent, the voice signal C is input to the adaptive filter F1D as a reference signal. The adaptive filter F1C outputs a passed signal P1C based on the filter coefficient C1C and the voice signal C. The filter unit F1 adds the passed signal P1A to the passed signal P1B or the passed signal P1C and outputs the result. When the filter unit F1 includes the adaptive filter F1D, the adaptive filter F1D outputs a passed signal P1D based on the filter coefficient C1D and the voice signal C. The filter unit F1 adds any one of the passed signal P1B, the passed signal P1C, and the passed signal P1D to the passed signal P1A and outputs the result. In the present embodiment, the adaptive filter F1A, the adaptive filter F1B, and the adaptive filter F1C are implemented by a processor executing a program. The adaptive filter F1A, the adaptive filter F1B, and the adaptive filter F1C may also be physically separate and different hardware structures.

[0075] Here, an outline of the operation of the adaptive filter will be described. The adaptive filter is a filter for suppressing crosstalk components. For example, in the case where the LMS (Least Mean Square) is used as an update algorithm for the filter coefficient, the adaptive filter is a filter that minimizes a cost function defined by the mean square of the error signal. The error signal referred to here means the difference between the output signal and the target component.

[0076] Here, as the adaptive filter, an FIR (Finite Impulse Response) filter is exemplified. Other types of adaptive filters may also be used. For example, an IIR (Infinite Impulse Response) filter may also be used.

[0077] When the voice processing device 21 uses one FIR filter as the adaptive filter, the error signal, which is the difference between the output signal of the voice processing device 21 and the target component, is represented by the following formula (1).

[0078]

Equation 1

[0079]

[0080] Here, n is the time, e(n) is the error signal, d(n) is the target component, wi is the filter coefficient, x(n) is the reference signal, and l is the tap length. The larger the tap length l is, the more faithfully the adaptive filter can reproduce the acoustic characteristics of the voice signal. In the absence of reverberation, the tap length l can be set to 1. For example, the tap length l is set to a fixed value. For example, when the target component is the voice of the driver hm1, the reference signal x(n) is the second directivity signal and the voice signal C.

[0081] The control unit 28 controls the filter coefficient of the adaptive filter based on the determination result of the determination unit 35. In the present embodiment, the control unit 28 determines which of the adaptive filter F1B and the adaptive filter F1C to input the voice signal C based on the flag as the determination result output from the determination unit 35. The filter coefficient C1B of the adaptive filter F1B is updated to minimize the error signal when the voice signal C contains more voice emitted by the occupant hm3. On the other hand, the filter coefficient C1C of the adaptive filter F1C is updated to minimize the error signal when the voice signal C contains more voice emitted by the occupant hm4. Therefore, it is possible to separately use each adaptive filter according to which voice the voice signal C contains more, and the error signal can be made smaller.

[0082] For example, when the flag "0" is received from the determination unit 35, the control unit 28 determines that the voice signal C contains more voice emitted by the occupant hm3. Then, the control unit 28 controls the filter unit F1 to input the voice signal C to the adaptive filter F1B.

[0083] The adder 27 generates an output signal by subtracting the subtraction signal from the target voice signal output from the voice input unit 29. In the present embodiment, the subtraction signal is the signal obtained by adding the passing signal P1A and the passing signal P1B or the passing signal P1C output from the filter unit F1. The adder 27 outputs the output signal to the control unit 28.

[0084] The control unit 28 outputs the output signal output from the adder unit 27. The output signal of the control unit 28 is input to the speech recognition engine 40. Alternatively, the output signal may be directly input from the control unit 28 to the electronic device 50. When the output signal is directly input from the control unit 28 to the electronic device 50, the control unit 28 and the electronic device 50 may be connected either by wire or wirelessly. For example, the electronic device 50 may be a portable terminal, and the output signal may be directly input from the control unit 28 to the portable terminal via a wireless communication network. The output signal input to the portable terminal may also be output as speech from the speaker of the portable terminal.

[0085] In addition, the control unit 28 updates the filter coefficients of each adaptive filter with reference to the output signal output from the adder unit 27 and the flag as the judgment result output from the determination unit 35.

[0086] First, the control unit 28 determines the adaptive filter to be the update target of the filter coefficients based on the judgment result. Specifically, the control unit 28 sets the adaptive filter to which the voice signal C is input among the adaptive filters F1B and F1C and the adaptive filter F1A as the update target of the filter coefficients. In addition, the control unit 28 does not set the adaptive filters among the adaptive filters F1B and F1C to which the voice signal C is not input as the update target of the filter coefficients. For example, when the flag "0" is received from the determination unit 35, the control unit 28 determines that the voice signal C contains more voice emitted by the occupant hm3. In other words, the control unit 28 determines that the voice signal C is input to the adaptive filter F1B. Then, the control unit 28 sets the adaptive filter F1B as the update target of the filter coefficients and does not set the adaptive filter F1C as the update target of the filter coefficients.

[0087] Then, the control unit 28 updates the filter coefficients of the adaptive filter that is the update target of the filter coefficients so that the value of the error signal in Equation (1) approaches 0.

[0088] The update of the filter coefficients in the case where the LMS is used as the update algorithm will be described. When the filter coefficients w(n) at time n are updated to the filter coefficients w(n + 1) at time n + 1, the relationship between w(n + 1) and w(n) is expressed by the following Equation (2).

[0089]

Equation 2

[0090] w(n + 1) = w(n) - αx(n)e(n) …(2)

[0091] Here, α is the correction coefficient of the filter coefficients. The term αx(n)e(n) corresponds to the update amount.

[0092] In addition, the algorithm for updating the filter coefficients is not limited to LMS, and other algorithms can also be used. For example, algorithms such as ICA (Independent Component Analysis) and NLMS (Normalized Least Mean Square) can also be used.

[0093] When updating the filter coefficients, the control unit 28 sets the intensity of the input reference signal to zero for the adaptive filters that are not the objects of filter coefficient update. For example, when receiving the flag "0" from the determination unit 35, the control unit 28 sets it such that the second directivity signal input to the adaptive filter F1A as the reference signal and the voice signal C input to the adaptive filter F1B as the reference signal are input while maintaining the intensity output from the directivity control unit 30. On the other hand, the control unit 28 sets the intensity of the voice signal C input to the adaptive filter F1C as the reference signal to zero. Here, "setting the intensity of the reference signal input to the adaptive filter to zero" includes suppressing the intensity of the reference signal input to the adaptive filter near zero. In addition, "setting the intensity of the reference signal input to the adaptive filter to zero" also includes setting not to input the reference signal to the adaptive filter. In the adaptive filter where the intensity of the input reference signal is set to zero, adaptive filtering may not be performed. Thus, the processing amount of the crosstalk suppression process using the adaptive filter can be reduced.

[0094] Then, the control unit 28 updates the filter coefficients only for the adaptive filters that are the objects of filter coefficient update, and does not update the filter coefficients for the adaptive filters that are not the objects of filter coefficient update. Thus, the processing amount of the crosstalk suppression process using the adaptive filter can be reduced.

[0095] For example, consider the case where the target seat is set to the driver's seat, and the driver hm1, passengers hm2 and hm4 are silent while passenger hm3 is speaking. At this time, the voices of passengers other than the driver hm1 leak into the voice signal of the voice picked up by the microphone MC1. In other words, the voice signal A contains a crosstalk component. The voice processing device 21 can update the adaptive filter to eliminate the crosstalk component and minimize the error signal. In this case, since no one is speaking at the driver's seat, the error signal is ideally a silent signal. In addition, in the above case where the driver hm1 is speaking, the words spoken by the driver hm1 leak into microphones other than the microphone MC1. In this case, the words spoken by the driver hm1 are not eliminated due to the processing of the voice processing device 21. This is because the words spoken by the driver hm1 contained in the voice signal A are earlier in time than the words spoken by the driver hm1 contained in other voice signals. This is based on the law of causality. Therefore, regardless of whether the voice signal contains the target component or not, the voice processing device 21 updates the adaptive filter to minimize the error signal, thereby reducing the crosstalk component contained in the voice signal A.

[0096] In the present embodiment, the functions of the voice input unit 29, the directivity control unit 30, the filter unit F1, the control unit 28, and the addition unit 27 are implemented by a processor executing a program held in a memory. Alternatively, the voice input unit 29, the directivity control unit 30, the filter unit F1, the control unit 28, and the addition unit 27 may be constituted by different hardware.

[0097] The voice processing device 21 has been described. The voice processing device 22 and the voice processing device 23 have almost the same structure except for the filter unit. The voice processing device 22 uses the voice spoken by the passenger hm2 as the target component. The voice processing device 22 outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC2. Therefore, the difference between the voice processing device 22 and the voice processing device 21 lies in the filter unit that is input with the first directivity signal and the voice signal C. Similarly, the voice processing device 23 uses the voice spoken by the passenger hm3 or hm4 as the target component. The voice processing device 23 outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC3. Therefore, the difference between the voice processing device 23 and the voice processing device 21 lies in the filter unit that is input with the voice signals A, B, and C.

[0098] Figure 5It is a flowchart showing the operation process of the voice processing device 21. First, voice signals A, B, and C are input to the voice input unit 29 (S1). Next, the directivity control unit 30 performs directivity control processing using voice signals A and B to generate a first directivity signal and a second directivity signal (S2). Then, the determination unit 35 determines whether a voice component is input to the microphone MC3 (S3). The determination unit 35 outputs the determination result as a flag to the control unit 28. When the determination unit 35 determines that no voice signal is input to the microphone MC3 (S3: "No"), the control unit 28 sets the intensity of the voice signal C input to the filter unit F1 to zero, and does not change the intensity of the second directivity signal. Then, the filter unit F1 generates a subtraction signal as follows (S4). The adaptive filter F1A passes the second directivity signal and outputs the passed signal P1A. The adaptive filter F1B passes the voice signal C and outputs the passed signal P1B. The adaptive filter F1C passes the voice signal C and outputs the passed signal P1C. The filter unit F1 adds the passed signal P1A, the passed signal P1B, and the passed signal P1C and outputs the result as the subtraction signal. The addition unit 27 subtracts the subtraction signal from the first directivity signal to generate an output signal and outputs the output signal (S5). The output signal is input to the control unit 28 and is output from the control unit 28. Next, the control unit 28 updates the filter coefficients of the adaptive filter F1A based on the output signal so that the target component included in the output signal becomes maximum (S6). Then, the voice processing device 21 performs step S1 again.

[0099] When the determination unit 35 determines that a voice signal has been input to the microphone MC3 (S3: "Yes"), the determination unit 35 determines which of the occupants hm3 and hm4 has emitted the voice component input to the microphone MC3 (S7). In other words, the determination unit 35 determines which of the voices emitted by the occupant hm3 and the voice emitted by the occupant hm4 is more contained in the voice signal C. The determination unit 35 outputs this determination result as a flag to the control unit 28. When the voice signal C more contains the voice emitted by the occupant hm3 (S7: "hm3"), the filter unit F1 generates a subtraction signal as follows (S8). The control unit 28 controls the filter unit F1 so that the voice signal C is input to the adaptive filter F1B. On the other hand, the control unit 28 controls the filter unit F1 so that the voice signal C is input to the adaptive filter F1C in a state where the intensity is zero. In other words, the control unit 28 does not change the intensity of the second directivity signal input to the adaptive filter F1A and the voice signal C input to the adaptive filter F1B, but changes the intensity of the voice signal C input to the adaptive filter F1C to zero. Then, the filter unit F1 generates a subtraction signal by the same operation as in step S4. The addition unit 27 subtracts the subtraction signal from the first directivity signal in the same manner as in step S5, thereby generating an output signal and outputting the output signal (S9). Next, the control unit 28 updates the filter coefficients of the adaptive filter to which the voice signal is input based on the output signal so that the target component contained in the output signal becomes maximum (S10). Specifically, the filter coefficients of the adaptive filter F1A and the adaptive filter F1B are updated. Then, the voice processing device 21 performs step S1 again.

[0100] In the case where it is determined in step S7 that the voice signal C contains a relatively large amount of voice emitted by the occupant hm4 (S7: "hm4"), the filter unit F1 generates a subtraction signal as follows (S11). The control unit 28 controls the filter unit F1 such that the voice signal C is input to the adaptive filter F1C. On the other hand, the control unit 28 controls the filter unit F1 such that the voice signal C is input to the adaptive filter F1B in a state where the intensity is zero. In other words, the control unit 28 does not change the intensity of the second directivity signal input to the adaptive filter F1A and the voice signal C input to the adaptive filter F1C, but changes the intensity of the voice signal C input to the adaptive filter F1B to zero. Then, the filter unit F1 generates a subtraction signal by the same operation as in step S4. The addition unit 27 subtracts the subtraction signal from the first directivity signal in the same manner as in step S5, thereby generating an output signal and outputting the output signal (S9). Next, the control unit 28 updates the filter coefficients of the adaptive filter to which the voice signal is input based on the output signal so that the target component included in the output signal becomes maximum (S10). Specifically, the filter coefficients of the adaptive filter F1A and the adaptive filter F1C are updated. Then, the voice processing device 21 performs step S1 again.

[0101] In the present embodiment, for the adaptive filter to which the voice signal in a state where the input intensity is zero is input, the update of the filter coefficients is not performed. Thereby, compared with the case where the filter coefficients are always updated for all the adaptive filters, the processing amount of the control unit 28 can be reduced. On the other hand, the control unit 28 may always update the filter coefficients for all the adaptive filters. By always updating the filter coefficients for all the adaptive filters, the control unit 28 can always perform the same processing, so the processing becomes simple. In addition, by always updating the filter coefficients for all the adaptive filters, for example, for a certain adaptive filter, even immediately after changing from the state where the input voice signal has a zero intensity to the state where the input voice signal has a non-zero intensity, the filter coefficients can be updated with high accuracy.

[0102] Thus, in the voice processing system 5 of the first embodiment, multiple voice signals are acquired by multiple microphones. Using another voice signal as a reference signal, a subtraction signal generated by an adaptive filter is subtracted from a certain voice signal, thereby accurately obtaining the voice of a specific speaker. In the first embodiment, it is configured to be able to pick up multiple voices with different generation positions using one microphone. Specifically, the microphone MC3 picks up the voices of the occupant hm3 and the occupant hm4 in the rear seat. On this basis, it is determined which of the multiple voices the voice signal based on the picked-up voice contains, and the adaptive filter to which the voice signal is input is changed according to which voice is included. Thus, even in the case of picking up multiple voices using one microphone, the voice signal of the target component can be accurately obtained. Therefore, for example, microphones do not have to be provided for each seat one by one, so the cost can be reduced. In addition, when using an adaptive filter to obtain the target component, compared with the case of using the signals output from the microphones provided in all seats as reference signals, the number of reference signals used in the processing can be reduced. Thereby, the amount of processing for eliminating crosstalk components can be reduced. In addition, for the adaptive filter of the voice signal in a state where the input intensity is zero, the update of the filter coefficients may not be performed. Thus, compared with the case where the filter coefficients of all adaptive filters are always updated, the processing amount can be further reduced.

[0103] (Second Embodiment)

[0104] The voice processing system 5A according to the second embodiment is different from the voice processing system 5 according to the first embodiment in that it includes a voice processing device 20A instead of the voice processing device 20, and includes a microphone MC4. The voice processing device 20A according to the second embodiment is different from the voice processing device 20 according to the first embodiment in that it has an abnormality detection unit and uses the voice signal D.

[0105] The voice processing device 20A according to the second embodiment detects whether there is an abnormality in each microphone, and uses the voice signal output from the microphone in which no abnormality is detected to perform directivity control processing and crosstalk component elimination processing. Next, Figure 6 、 Figure 7 and Figure 8 are used to describe the voice processing device 20A. For the structures and operations that are the same as those described in the first embodiment, the same reference numerals are used, and thus their descriptions are omitted or simplified.

[0106] Using Figure 6 to describe the details of the voice processing system 5A in the second embodiment. Figure 6This is a diagram showing an example of the schematic structure of the voice processing system 5A in the second embodiment. The voice processing system 5 includes a microphone MC1, a microphone MC2, a microphone MC3, a microphone MC4, and a voice processing device 20A. In the present embodiment, the microphone MC3 picks up the voice spoken by the occupant hm3. In other words, the microphone MC3 acquires a voice signal including the voice component spoken by the occupant hm3. The microphone MC3 is arranged, for example, near the right side of the center of the rear seat on the ceiling. In the present embodiment, the microphone MC4 picks up the voice spoken by the occupant hm4. In other words, the microphone MC4 acquires a voice signal including the voice component spoken by the occupant hm4. The microphone MC4 is arranged, for example, near the left side of the center of the rear seat on the ceiling. The microphone MC1 is located at a position farther from the right seat in the rear seat than the microphone MC3. The microphone MC2 is located at a position farther from the left seat in the rear seat than the microphone MC4. The microphone MC4 is located at a position closer to the left seat in the rear seat than the microphone MC3.

[0107] In the present embodiment, the voice processing system 5A includes a plurality of voice processing devices 20A corresponding to the respective microphones. Specifically, the voice processing system 5A includes a voice processing device 21A, a voice processing device 22A, a voice processing device 23A, and a voice processing device 24A. The voice processing device 21A corresponds to the microphone MC1. The voice processing device 22A corresponds to the microphone MC2. The voice processing device 23A corresponds to the microphone MC3. The voice processing device 24A corresponds to the microphone MC4. Hereinafter, the voice processing device 21A, the voice processing device 22A, the voice processing device 23A, and the voice processing device 24A may be collectively referred to as the voice processing device 20A.

[0108] In Figure 6 the structure shown, it is exemplified that the voice processing device 21A, the voice processing device 22A, the voice processing device 23A, and the voice processing device 24A are composed of different hardware, but the functions of the voice processing device 21A, the voice processing device 22A, the voice processing device 23A, and the voice processing device 24A may also be implemented by one voice processing device 20A. Or, a part of the voice processing device 21A, the voice processing device 22A, the voice processing device 23A, and the voice processing device 24A may be composed of common hardware, and the remaining parts may be composed of different hardware.

[0109] In the present embodiment, each voice processing device 20A is arranged in each seat near the corresponding microphone. For example, the voice processing device 21A is arranged in the driver's seat, the voice processing device 22A is arranged in the front passenger seat, the voice processing device 23A is arranged in the seat on the right side of the rear seat, and the voice processing device 24A is arranged in the seat on the left side of the rear seat. Each voice processing device 20A may also be arranged in the instrument panel.

[0110] Figure 7 It is a block diagram showing the structure of the voice processing device 21A. The voice processing device 21A, the voice processing device 22A, the voice processing device 23A, and the voice processing device 24A have the same structure and function except for a part of the structure of the filter unit described later. Here, the voice processing device 21A will be described. The voice processing device 21A targets the voice spoken by the driver hm1. The voice processing device 21A outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC1.

[0111] As Figure 7 shown, the voice processing device 21A includes a voice input unit 29A, an abnormality detection unit 31, a directivity control unit 30A, a filter unit F2 including a plurality of adaptive filters, a control unit 28A that controls the filter coefficients of the adaptive filters of the filter unit F2, and an addition unit 27A.

[0112] Voice signals of voices picked up by the microphones MC1, MC2, MC3, and MC4 are input to the voice input unit 29A. In other words, the microphones MC1, MC2, MC3, and MC4 respectively output signals of voice signals based on the picked-up voices to the voice input unit 29A. Regarding the microphones MC1 and MC2, since they are the same as those in the first embodiment, detailed description thereof will be omitted.

[0113] The microphone MC3 outputs a voice signal C to the voice input unit 29A. The voice signal C is a signal including the voice of the occupant hm3 and noise, and the noise includes the voices of occupants other than the occupant hm3. The microphone MC3 corresponds to the first microphone. In addition, the microphone MC3 corresponds to the fourth microphone. The voice picked up by the microphone MC3 corresponds to the first voice signal. In addition, the voice picked up by the microphone MC3 corresponds to the fourth voice signal. The voice emitted by the occupant hm3 corresponds to the first voice component. The voice signal C corresponds to the first signal. In addition, the voice signal C corresponds to the fourth signal.

[0114] The microphone MC4 outputs a voice signal D to the voice input unit 29A. The voice signal D is a signal including the voice of the occupant hm4 and noise, and the noise includes the voices of occupants other than the occupant hm4. The microphone MC4 corresponds to the first microphone. In addition, the microphone MC4 corresponds to the fifth microphone. The voice picked up by the microphone MC4 corresponds to the first voice signal. In addition, the voice picked up by the microphone MC4 corresponds to the fifth voice signal. The voice emitted by the occupant hm4 corresponds to the second voice component. The voice signal D corresponds to the first signal. In addition, the voice signal D corresponds to the fifth signal.

[0115] The voice input unit 29A outputs a voice signal A, a voice signal B, a voice signal C, and a voice signal D. The voice input unit 29A corresponds to a receiving unit.

[0116] In the present embodiment, the voice processing device 21A includes one voice input unit 29A to which voice signals from all the microphones are input, but it may also include a voice input unit 29A for each microphone to which the corresponding voice signal is input. For example, it may be configured as follows: the voice signal of the voice picked up by the microphone MC1 is input to the voice input unit corresponding to the microphone MC1, the voice signal of the voice picked up by the microphone MC2 is input to another voice input unit corresponding to the microphone MC2, the voice signal of the voice picked up by the microphone MC3 is input to another voice input unit corresponding to the microphone MC3, and the voice signal of the voice picked up by the microphone MC4 is input to another voice input unit corresponding to the microphone MC4.

[0117] The voice signals A, B, C, and D output from the voice input unit 29A are input to the abnormality detection unit 31. The abnormality detection unit 31 detects whether there is an abnormality in the microphones MC3 and MC4, and sends abnormality information related to the abnormalities of the microphones MC3 and MC4 to the control unit 28A. Here, the abnormalities of the microphones include microphone failures, poor connections between the microphones and other devices, and dead batteries in the microphones. The poor connection between the microphone and other devices includes a broken wire of the cable that electrically connects the microphone and other devices. Alternatively, the abnormality detection unit 31 can detect whether there is an abnormality in the microphones MC1 and MC2, and can also send abnormality information related to the abnormalities of the microphones MC1 and MC2 to the control unit 28A. The abnormality detection unit 31 detects, for example, whether there is an abnormality in the microphone corresponding to each voice signal based on each voice signal. The abnormality detection unit 31 determines that there is an abnormality in the microphone corresponding to the voice signal, for example, when the intensity of the voice signal is smaller than a threshold value. The abnormality detection unit 31 can also determine that there is an abnormality in the microphone corresponding to the voice signal when the period during which the intensity of the voice signal is smaller than the threshold value is a fixed length or more, or when the frequency at which the intensity of the voice signal becomes smaller than the threshold value during a fixed period is a fixed number or more. The abnormality detection unit 31 outputs, for example, the determination result of whether there is an abnormality in each microphone as a flag to the control unit 28A. The flag is an example of the abnormality information. For each voice signal, the flag represents a value of "1" or "0". "1" means that it is determined that there is an abnormality in the corresponding microphone, and "0" means that it is not determined that there is an abnormality in the corresponding microphone. For example, when it is determined that there are no abnormalities in the microphones MC1, MC2, and MC4 and it is determined that there is an abnormality in the microphone MC3, the abnormality detection unit 31 outputs the flag "0, 0, 1, 0" as the determination result to the control unit 28. After detecting the abnormalities of each microphone, the abnormality detection unit 31 outputs the voice signals A, B, C, and D to the directivity control unit 30A.

[0118] In the present embodiment, the voice processing device 21A includes one abnormality detection unit 31 to which all voice signals are input, but it is also possible to include an abnormality detection unit 31 to which the corresponding voice signal is input for each voice signal. For example, the voice processing device 21A may be configured to include an abnormality detection unit for the voice signal A, an abnormality detection unit for the voice signal B, an abnormality detection unit for the voice signal C, and an abnormality detection unit for the voice signal D, respectively.

[0119] The voice signals A, B, C, and D output from the abnormality detection unit 31 are input to the directivity control unit 30A. The directivity control unit 30 performs directivity control processing using the voice signals output from the microphones other than the microphone in which an abnormality is detected by the abnormality detection unit 31 and the microphones on the same side as that microphone. The directivity control processing is, for example, beamforming. Here, "on the same side" means that whether it is on the front seat side or on the rear seat side is the same. In the present embodiment, the microphone MC1 and the microphone MC2 are on the same side, and the microphone MC3 and the microphone MC4 are on the same side. For example, when an abnormality of the microphone MC3 is detected, the directivity control unit 30A performs directivity control processing using the voice signals A and B. Then, the directivity control unit 30A outputs two directivity signals obtained by performing directivity control processing using two voice signals. For example, the directivity control unit 30A outputs a first directivity signal obtained by performing directivity control processing on the voice signal A. In addition, the directivity control unit 30A outputs a second directivity signal obtained by performing directivity control processing on the voice signal B. For example, when no abnormality is detected in any microphone, the directivity control unit 30A performs directivity control processing using all the voice signals and outputs the obtained directivity signals. For example, in addition to outputting the first directivity signal and the second directivity signal, the directivity control unit 30A also outputs a third directivity signal obtained by performing directivity control processing on the voice signal C and a fourth directivity signal obtained by performing directivity control processing on the voice signal D. For example, the abnormality detection unit 31 can detect an abnormality of the microphone MC2. When an abnormality of the microphone MC2 is detected, the directivity control unit 30A outputs a third directivity signal obtained by performing directivity control processing on the voice signal C and a fourth directivity signal obtained by performing directivity control processing on the voice signal D.

[0120] In addition, the directivity control unit 30A determines whether a voice component is input to the microphone on the same side as the microphone in which an abnormality is detected. For example, when it is determined that the microphone MC3 is abnormal, the directivity control unit 30A determines that a voice signal is input to the microphone MC4 when the intensity of the voice signal D output from the microphone MC4, which is the microphone on the same side as the microphone MC3, is greater than at least one of the intensities of the first directivity signal and the second directivity signal; otherwise, it determines that no voice signal is input to the microphone MC4.

[0121] In addition, the directivity control unit 30A includes a determination unit 35A. The determination unit 35A determines which occupant's voice is more contained in the voice signal output from the microphone on the same side as the microphone in which an abnormality is detected, based on the voice signal output from the microphone in which no abnormality is detected. Explain the reason for making such a determination. For example, the crosstalk component containing the voice of the occupant hm3 is removed from the target component using the voice signal C output from the microphone MC3. However, when it is determined that there is an abnormality in the microphone MC3, the voice signal C also has an abnormality, so it is difficult to use the voice signal C to remove the crosstalk component containing the voice of the occupant hm3. In this case, since the voice of the occupant hm3 also leaks into the microphone MC4, it is possible to consider using the voice signal D output from the microphone MC4 to remove the crosstalk component containing the voice of the occupant hm3. It is possible that both the voice of the occupant hm3 and the voice of the occupant hm4 leak into the microphone MC4. Therefore, it can be determined which of the voices of the occupant hm3 and the occupant hm4 is more contained in the voice signal D, and if the determination result is that the voice of the occupant hm3 is more contained, the voice signal D is used to remove the crosstalk component containing the voice of the occupant hm3.

[0122] For example, when the determination unit 35A determines that there is an abnormality in the microphone MC3, it determines which of the voices of the occupant hm3 and the occupant hm4 is more contained in the voice signal D based on the first directivity signal and the second directivity signal. In other words, the determination unit 35A determines which of the voices of the occupant hm3 and the occupant hm4 is more contained in the voice signal C based on the voice signal A and the voice signal B. The specific determination method is the same as the method described in the first embodiment.

[0123] The determination unit 35A outputs the result of the determination of which of the voice signal C or the voice signal D more contains the voices of the occupant hm3 and the occupant hm4 to the control unit 28A. The determination unit 35A outputs the determination result to the control unit 28A as a flag, for example. The flag represents a value of "0" or "1". "0" indicates that the voice signal more contains the voice of the occupant hm3, and "1" indicates that the voice signal more contains the voice of the occupant hm4. For example, when it is determined that there are no abnormalities in the microphones MC1, MC2, and MC4 and it is determined that there is an abnormality in the microphone MC3, the directivity control unit 30A sends a flag as the determination result regarding the voice signal D. For example, when it is determined that the voice signal D more contains the voice of the occupant hm3, the directivity control unit 30A outputs the flag "0" as the determination result to the control unit 28A.

[0124] For example, when an abnormality of the microphone MC3 is detected, the directivity control unit 30A outputs a first directivity signal to the addition unit 27A, and outputs a second directivity signal, the voice signal C, and the voice signal D to the filter unit F2.

[0125] In the present embodiment, the determination unit 35A included in the directivity control unit 30A determines whether a voice component is input to the microphone on the same side as the microphone in which the abnormality is detected, and determines which occupant's voice is included more in the voice signal output from the microphone on the same side as the microphone in which the abnormality is detected. However, the voice processing device 21A may also include a determination unit 35A separate from the directivity control unit 30A. In this case, the determination unit 35A is connected, for example, between the abnormality detection unit 31 and the directivity control unit 30A. Alternatively, it may be that the voice processing device 21A only includes the determination unit 35A and does not include the directivity control unit 30A. The structure and function of the determination unit 35A are the same as those of the determination unit described in the first embodiment, and thus the detailed description is omitted.

[0126] The filter unit F2 includes an adaptive filter F2A, an adaptive filter F2B, an adaptive filter F2C, an adaptive filter F2D, and an adaptive filter F2E. The filter unit F2 is used for processing to suppress crosstalk components other than the voice of the driver hm1 included in the voice picked up by the microphone MC1. In the present embodiment, the filter unit F2 includes five adaptive filters, but the number of adaptive filters can be appropriately set based on the number of input voice signals and the processing amount of the crosstalk suppression processing. The details of the crosstalk suppression processing will be described later.

[0127] The second directivity signal is input to the adaptive filter F2A as a reference signal. The adaptive filter F2A outputs a passed signal P2A based on the filter coefficients C2A and the second directivity signal. When it is determined that the microphone MC4 is abnormal and the voice signal C is determined to contain a large amount of voice emitted by the occupant hm3, the voice signal C is input to the adaptive filter F2B as a reference signal. The adaptive filter F2B outputs a passed signal P2B based on the filter coefficients C2B and the voice signal C. It is also possible that even when it is not determined that the microphone MC4 is abnormal, the voice signal C is input to the adaptive filter F2B as a reference signal. On the other hand, when it is determined that the microphone MC4 is abnormal and the voice signal C is determined to contain a large amount of voice emitted by the occupant hm4, the voice signal C is input to the adaptive filter F2C as a reference signal. The adaptive filter F2C outputs a passed signal P2C based on the filter coefficients C2C and the voice signal C. Similarly, when it is determined that the microphone MC3 is abnormal and the voice signal D is determined to contain a large amount of voice emitted by the occupant hm3, the voice signal D is input to the adaptive filter F2D as a reference signal. The adaptive filter F2D outputs a passed signal P2D based on the filter coefficients C2D and the voice signal D. It is also possible that even when it is not determined that the microphone MC3 is abnormal, the voice signal D is input to the adaptive filter F2D as a reference signal. On the other hand, when it is determined that the microphone MC3 is abnormal and the voice signal D is determined to contain a large amount of voice emitted by the occupant hm4, the voice signal D is input to the adaptive filter F2E as a reference signal. The adaptive filter F2E outputs a passed signal P2E based on the filter coefficients C2E and the voice signal D. The filter unit F2 adds and outputs the passed signal P2A, the passed signal P2B or the passed signal P2C, and the passed signal P2D or the passed signal P2E. In the present embodiment, the adaptive filter F2A, the adaptive filter F2B, the adaptive filter F2C, the adaptive filter F2D, and the adaptive filter F2E are implemented by a processor executing a program. The adaptive filter F2A, the adaptive filter F2B, the adaptive filter F2C, the adaptive filter F2D, and the adaptive filter F2E may also be physically separate and different hardware structures.

[0128] In the present embodiment, the filter unit F2 is configured to include two adaptive filters that can receive the voice signal C and two adaptive filters that can receive the voice signal D, and this has been described. The filter unit F2 may also be configured to include two adaptive filters that can receive the second directional signal. For example, the abnormality detection unit 31 may be able to detect an abnormality of the microphone MC2, and the filter unit F2 may include an adaptive filter F2A1 that receives the second directional signal when an abnormality of the microphone MC2 is detected, and an adaptive filter F2A2 that receives the second directional signal when an abnormality of the microphone MC2 is not detected, respectively.

[0129] The control unit 28A controls the filter coefficients of the adaptive filters based on the determination result of the abnormality detection unit 31 and the determination result of the determination unit 35A. In the present embodiment, the control unit 28A determines which of the voice signal C is to be input to the adaptive filter F2B and the adaptive filter F2C based on the flag as the determination result output from the abnormality detection unit 31 and the flag as the determination result output from the determination unit 35A. In addition, in the present embodiment, the control unit 28A determines which of the voice signal D is to be input to the adaptive filter F2D and the adaptive filter F2E based on the flag as the determination result output from the abnormality detection unit 31 and the flag as the determination result output from the determination unit 35A. The filter coefficient C2B of the adaptive filter F2B is updated to minimize the error signal when the voice signal C contains more voice emitted by the occupant hm3. In addition, the filter coefficient C2C of the adaptive filter F2C is updated to minimize the error signal when the voice signal C contains more voice emitted by the occupant hm4. The filter coefficient C2D of the adaptive filter F2D is updated to minimize the error signal when the voice signal D contains more voice emitted by the occupant hm3. In addition, the filter coefficient C2E of the adaptive filter F2E is updated to minimize the error signal when the voice signal D contains more voice emitted by the occupant hm4. Therefore, it is possible to separately use each adaptive filter according to which voice the voice signal C contains more or which voice the voice signal D contains more, and the error signal can be made smaller. When the filter unit F2 includes two adaptive filters that can receive the second directional signal, the control unit 28A may also determine to which adaptive filter the second directional signal is to be input.

[0130] For example, when receiving the flags "0, 0, 1, 0" from the abnormality detection unit 31 and the flag "0" from the determination unit 35A, the control unit 28A determines that there is an abnormality in the microphone MC3 and that the voice signal D contains a relatively large amount of the voice emitted by the occupant hm3. Then, the control unit 28A controls the filter unit F2 so that the voice signal D is input to the adaptive filter F2D.

[0131] The addition unit 27A generates an output signal by subtracting the subtraction signal from the target voice signal output from the self-voice input unit 29. In the present embodiment, the subtraction signal is a signal obtained by adding the passed signal P2A, the passed signal P2B, or the passed signal P2C output from the filter unit F2, and the passed signal P2D or the passed signal P2E. The addition unit 27A outputs the output signal to the control unit 28A.

[0132] The control unit 28A outputs the output signal output from the addition unit 27A. The utilization of the output signal is the same as that in the first embodiment.

[0133] In addition, the control unit 28A updates the filter coefficients of each adaptive filter with reference to the output signal output from the addition unit 27A, the flag as the determination result output from the abnormality detection unit 31, and the flag as the determination result output from the determination unit 35A.

[0134] First, the control unit 28A determines the adaptive filter to be updated as the filter coefficient based on the determination result. Specifically, the control unit 28A sets the adaptive filters F2B, F2C, F2D, and F2E that receive the voice signal and the adaptive filter F2A as the objects for updating the filter coefficients. In addition, the control unit 28A does not set the adaptive filters F2B, F2C, F2D, and F2E that do not receive the voice signal as the objects for updating the filter coefficients. For example, when receiving the flags "0, 0, 1, 0" from the abnormality detection unit 31 and the flag "0" from the determination unit 35A, the control unit 28A determines that there is an abnormality in the microphone MC3 and that the voice signal D contains a relatively large amount of the voice emitted by the occupant hm3. In other words, the control unit 28A determines that the voice signal C is not input to any of the adaptive filters F2B and F2C, the voice signal D is input to the adaptive filter F2D, and the voice signal D is not input to the adaptive filter F2E. Then, the control unit 28A sets the adaptive filter F2D as the object for updating the filter coefficients, and does not set the adaptive filters F2B, F2C, and F2E as the objects for updating the filter coefficients.

[0135] Then, the control unit 28A updates the filter coefficients of the adaptive filter that is the object of update of the filter coefficients so that the value of the error signal in Equation (1) approaches 0. Regarding the specific method of updating the filter coefficients, it is the same as the method described in the first embodiment.

[0136] The control unit 28A updates only the filter coefficients of the adaptive filter that is the object of update of the filter coefficients, and does not update the filter coefficients of the adaptive filter that is not the object of update of the filter coefficients. Thereby, the processing amount of the crosstalk suppression process using the adaptive filter can be reduced.

[0137] In the present embodiment, the functions of the voice input unit 29A, the abnormality detection unit 31, the directivity control unit 30A, the filter unit F2, the control unit 28A, and the addition unit 27A are implemented by a processor executing a program held in a memory. Alternatively, the voice input unit 29A, the abnormality detection unit 31, the directivity control unit 30A, the filter unit F2, the control unit 28A, and the addition unit 27A may be configured by different hardware.

[0138] The voice processing device 21A has been described, but the voice processing devices 22A, 23A, and 24A have almost the same structure except for the filter unit. The voice processing device 22A targets the voice spoken by the occupant hm2. The voice processing device 22A outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC2. Therefore, the difference between the voice processing device 22A and the voice processing device 21A is that it has a filter unit to which the first directivity signal, the voice signal C, and the voice signal D are input. The same applies to the voice processing devices 23A and 24A.

[0139] Figure 8It is a flowchart showing the operation process of the voice processing device 21A. First, voice signals A, B, C, and D are input to the voice input unit 29A (S101). Next, the abnormality detection unit 31 determines whether there is an abnormality in each microphone based on each voice signal (S102). The abnormality detection unit 31 outputs the determined result as a flag to the control unit 28A. When no abnormality is detected from any microphone (S102: "No"), the directivity control unit 30A performs directivity control processing using all voice signals (S103). The directivity control unit 30A outputs a directivity signal to the filter unit F2. The filter unit F2 generates a subtraction signal as follows (S104). The adaptive filter F2A allows the second directivity signal to pass through and outputs the passed signal P2A. The adaptive filter F2B allows the third directivity signal to pass through and outputs the passed signal P2B. The adaptive filter F2D allows the fourth directivity signal to pass through and outputs the passed signal P2D. The filter unit F2 adds the passed signal P2A, the passed signal P2B, and the passed signal P2D and outputs the result as the subtraction signal. The addition unit 27A subtracts the subtraction signal from the first directivity signal to generate an output signal and outputs the output signal (S105). The output signal is input to the control unit 28A and is output from the control unit 28A. Next, the control unit 28A refers to the flag as the determination result output from the abnormality detection unit 31 and the flag as the determination result output from the directivity control unit 30A, and updates the filter coefficients of the adaptive filter F2A, the adaptive filter F2B, and the adaptive filter F2D based on the output signal so that the target component included in the output signal becomes the maximum (S106). Then, the voice processing device 21A performs step S1 again.

[0140] When an abnormality is detected in one of the microphones in step S102 (S102: "Yes"), the abnormality detection unit 31 determines whether the microphone in which the abnormality is detected is the microphone of the target seat. Here, the target seat refers to the seat from which the voice to be the target component is to be acquired. In the voice processing device 21A, the target seat is the driver's seat, and the microphone of the target seat is the microphone MC1. The abnormality detection unit 31 outputs the determined result as a flag to the control unit 28A. When the microphone in which the abnormality is detected is the microphone of the target seat, the control unit 28A sets the intensity of the voice signal A received from the voice input unit 29A to zero and outputs it as the output signal (S108). At this time, the control unit 28A does not update the filter coefficients of the adaptive filter F2A, the adaptive filter F2B, the adaptive filter F2C, the adaptive filter F2D, and the adaptive filter F2E. Then, the voice processing device 21A performs step S101 again.

[0141] When the microphone in which an abnormality is detected in process S107 is not the microphone of the target seat (S107: "No"), the abnormality detection unit 31 determines whether the microphone in which the abnormality is detected is the microphone on the same side as the target seat (S109). When the microphone in which the abnormality is detected is not the microphone on the same side as the target seat (S109: "No"), the abnormality detection unit 31 outputs the determination result as a flag to the control unit 28A. The directivity control unit 30A performs directivity control processing using the voice signal A and the voice signal B to generate a first directivity signal and a second directivity signal (S110). Then, the determination unit 35A determines which voice component is input to the microphone on the same side as the microphone in which the abnormality is detected and in which no abnormality is detected (S111). For example, when an abnormality is detected in the microphone MC3, the determination unit 35A determines which of the voice of the occupant hm3 and the voice of the occupant hm4 is input to the microphone MC4. In other words, the determination unit 35A determines which of the voice of the occupant hm3 and the voice of the occupant hm4 is more included in the voice signal D. The determination unit 35A outputs the determination result as a flag to the control unit 28A. Hereinafter, it is assumed that an abnormality is detected in the microphone MC3 for explanation. When the voice signal D more includes the voice of the occupant hm3 (S111: "hm3"), the filter unit F2 generates a subtraction signal as follows (S112). The adaptive filter F2A passes the second directivity signal and outputs the passed signal P2A. The control unit 28A controls the filter unit F2 so that the voice signal C is input to the adaptive filter F2B in a state where the intensity is zero. In addition, the control unit 28 controls the filter unit F2 so that the voice signal C is input to the adaptive filter F2C in a state where the intensity is zero. On the other hand, the control unit 28A controls the filter unit F2 so that the voice signal D is input to the adaptive filter F2D. In addition, the control unit 28A controls the filter unit F2 so that the voice signal D is input to the adaptive filter F2E in a state where the intensity is zero. In other words, the control unit 28A does not change the intensity of the second directivity signal input to the adaptive filter F2A and the voice signal D input to the adaptive filter F2D, but changes the intensity of the voice signal C input to the adaptive filter F2B, the voice signal C input to the adaptive filter F2C, and the voice signal D input to the adaptive filter F2E to zero. Then, the filter unit F2 generates a subtraction signal by the same operation as in process S104. The addition unit 27A subtracts the subtraction signal from the first directivity signal in the same manner as in process S5, thereby generating an output signal and outputting the output signal (S113). Next, the control unit 28A updates the filter coefficients of the adaptive filter to which the voice signal is input based on the output signal so that the target component included in the output signal becomes the maximum (S114).Specifically, the filter coefficients of the adaptive filter F2A and the adaptive filter F2D are updated. Then, the voice processing device 21 performs the process S101 again.

[0142] In the case where it is determined in the process S111 that the voice signal D contains a large amount of the voice emitted by the occupant hm4 (S111: "hm4"), the filter unit F2 generates a subtraction signal as follows (S115). The adaptive filter F2A passes the second directivity signal and outputs a passed signal P2A. The control unit 28A controls the filter unit F2 so that the voice signal C is input to the adaptive filter F2B in a state where the intensity is zero. In addition, the control unit 28A controls the filter unit F2 so that the voice signal C is input to the adaptive filter F2C in a state where the intensity is zero. On the other hand, the control unit 28A controls the filter unit F2 so that the voice signal D is input to the adaptive filter F2D in a state where the intensity is zero. In addition, the control unit 28A controls the filter unit F2 so that the voice signal D is input to the adaptive filter F2E. In other words, the control unit 28 does not change the intensity of the second directivity signal input to the adaptive filter F2A and the voice signal D input to the adaptive filter F2E, but changes the intensity of the voice signal C input to the adaptive filter F2B, the voice signal C input to the adaptive filter F2C, and the voice signal D input to the adaptive filter F2D to zero. Then, the filter unit F2 generates a subtraction signal by the same operation as in the process S4. The addition unit 27A subtracts the subtraction signal from the first directivity signal in the same manner as in the process S5, thereby generating an output signal and outputting the output signal (S116). Next, the control unit 28A updates the filter coefficients of the adaptive filter to which the voice signal is input based on the output signal so that the target component included in the output signal becomes the maximum (S117). Specifically, the filter coefficients of the adaptive filter F2A and the adaptive filter F2E are updated. Then, the voice processing device 21 performs the process S101 again.

[0143] In addition, when the filter unit F2 includes two adaptive filters that can receive the second directional signal, a part of the previous process is changed as follows. For example, when the abnormality detection unit 31 can detect an abnormality in the microphone MC2, and the filter unit F2 includes an adaptive filter F2A1 that receives the second directional signal when an abnormality in the microphone MC2 is detected, and an adaptive filter F2A2 that receives the second directional signal when no abnormality in the microphone MC2 is detected, it is only necessary to rename the adaptive filter F2A that received the second directional signal in the previous process to the adaptive filter F2A2. The following-described process is performed in the following case: the abnormality detection unit 31 can detect an abnormality in the microphone MC2, and the filter unit F2 includes an adaptive filter F2A1 that receives the second directional signal when an abnormality in the microphone MC2 is detected, and an adaptive filter F2A2 that receives the second directional signal when no abnormality in the microphone MC2 is detected.

[0144] When the microphone in which an abnormality is detected in the process S109 is the microphone on the same side as the target seat, the abnormality detection unit 31 outputs the determination result as a flag to the control unit 28A. In this example, an abnormality in the microphone MC2 is detected. The directivity control unit 30A performs directivity control processing using the voice signal C and the voice signal D to generate a third directional signal and a fourth directional signal (S118). Then, the determination unit 35A determines which voice component is input to the microphone on the same side as the microphone in which an abnormality is detected and in which no abnormality is detected (S119). For example, when an abnormality is detected in the microphone MC2, the determination unit 35A determines which of the voice emitted by the driver hm1 and the voice emitted by the passenger hm2 is input to the microphone MC1. In other words, the determination unit 35A determines which of the voice emitted by the driver hm1 and the voice emitted by the passenger hm2 is included more in the voice signal A. The determination unit 35A outputs the determination result as a flag to the control unit 28A.

[0145] When the voice signal A includes more of the voice emitted by the passenger hm2, the control unit 28A sets the intensity of the voice signal A to zero and outputs it as an output signal (S108). At this time, the control unit 28A does not update the filter coefficients of the adaptive filter F2A1, the adaptive filter F2A2, the adaptive filter F2B, the adaptive filter F2C, the adaptive filter F2D, and the adaptive filter F2E. Then, the voice processing device 21A performs the process S101 again.

[0146] In the case where the voice signal A contains more voice emitted by the driver hm1, the filter unit F2 generates a subtraction signal as follows (S120). The control unit 28A controls the filter unit F2 such that the voice signal B is input to the adaptive filter F2A1 in a state where the intensity is zero. On the other hand, the control unit 28A controls the filter unit F2 such that the third directivity signal is input to the adaptive filter F2B. In addition, the control unit 28A controls the filter unit F2 such that the fourth directivity signal is input to the adaptive filter F2D. In other words, the control unit 28A does not change the intensity of the third directivity signal input to the adaptive filter F2B and the fourth directivity signal input to the adaptive filter F2D, but changes the intensity of the voice signal B input to the adaptive filter F2A1 to zero. The adaptive filter F2B passes the third directivity signal and outputs a passed signal P2B. The adaptive filter F2D passes the fourth directivity signal and outputs a passed signal P2D. The filter unit F2 adds the passed signal P2B and the passed signal P2D and outputs the result as a subtraction signal. The addition unit 27A subtracts the subtraction signal from the voice signal A to generate an output signal and outputs the output signal (S121). The output signal is input to the control unit 28A and output from the control unit 28A. Next, the control unit 28A refers to the flag as the determination result output from the abnormality detection unit 31 and the flag as the determination result output from the determination unit 35A, and updates the filter coefficients of the adaptive filter F2B and the adaptive filter F2D based on the output signal so that the target component included in the output signal becomes the maximum (S122). Then, the voice processing device 21A performs the process S101 again.

[0147] In addition, an example in the case where the abnormality detection unit 31 can detect abnormalities of the microphones MC1 and MC2 has been described, but it may be that the abnormality detection unit 31 can detect abnormalities of only the microphones MC3 and MC4. In this case, in the Figure 8 flowchart shown, the processes S107, S108, S109, and the processes S118 to S122 are omitted.

[0148] In the present embodiment, for the adaptive filter of the voice signal in the state where the input strength is zero, the update of the filter coefficients is not performed. Thus, compared with the case where the filter coefficients of all adaptive filters are always updated, the processing amount of the control unit 28A can be reduced. On the other hand, the control unit 28A may also always update the filter coefficients of all adaptive filters. By always updating the filter coefficients of all adaptive filters, the control unit 28A can perform the same processing all the time, so the processing becomes simple. In addition, by always updating the filter coefficients of all adaptive filters, for example, for a certain adaptive filter, even immediately after changing from the state of the voice signal with zero input strength to the state of the voice signal with non-zero input strength, the filter coefficients can be updated with high precision.

[0149] Thus, in the voice processing system 5A in the second embodiment, multiple voice signals are also acquired by multiple microphones. Using other voice signals as reference signals, the subtraction signal generated by using the adaptive filter is subtracted from a certain voice signal, thereby accurately obtaining the voice of a specific speaker. In addition, in the second embodiment, even when an abnormality is detected in a part of the microphones, the crosstalk component can be eliminated based on the voice leaking into other microphones. Thus, even when an abnormality occurs in the microphone, the voice of a specific speaker can be accurately obtained. In addition, in the second embodiment, when using the adaptive filter to obtain the target component, the voice signal output from the microphone in which the abnormality is detected is not used as the reference signal. Thus, the amount of processing for eliminating the crosstalk component can be reduced. In addition, for the adaptive filter of the voice signal in the state where the input strength is zero, the update of the filter coefficients may not be performed. Thus, compared with the case where the filter coefficients of all adaptive filters are always updated, the processing amount can be further reduced.

[0150] (Third Embodiment)

[0151] The difference between the voice processing system 5B according to the third embodiment and the voice processing system 5A according to the second embodiment is that it includes a voice processing device 20B instead of the voice processing device 20A and does not include the directivity control unit 30A.

[0152] The voice processing device 20B according to the third embodiment detects whether there is an abnormality in each microphone, and uses the voice signal output from the microphone in which no abnormality is detected to perform the processing of eliminating the crosstalk component. Next, use Figure 9 、 Figure 10 and Figure 11This is used to explain the voice processing device 20B. For the structures and operations that are the same as those described in the first and second embodiments, the same reference numerals are used, and thus their descriptions are omitted or simplified.

[0153] Use Figure 9 to explain the details of the voice processing system 5B in the third embodiment. Figure 9 FIG. is an example showing a schematic structure of the voice processing system 5B in the third embodiment. The voice processing system 5B includes a microphone MC1, a microphone MC2, a microphone MC3, a microphone MC4, and a voice processing device 20B. In the present embodiment, the microphone MC1 is disposed, for example, at the right-side handle of the driver's seat. In the present embodiment, the microphone MC2 is disposed, for example, at the left-side handle of the front passenger seat. In the present embodiment, the microphone MC3 is disposed, for example, at the right-side handle of the rear seat. In the present embodiment, the microphone MC4 is disposed, for example, at the left-side handle of the rear seat. The microphone MC1 is located at a position farther from the right seat in the rear seat than the microphone MC3. The microphone MC2 is located at a position farther from the left seat in the rear seat than the microphone MC4. The microphone MC4 is located at a position closer to the left seat in the rear seat than the microphone MC3.

[0154] In the present embodiment, the voice processing system 5B includes a plurality of voice processing devices 20B corresponding to the respective microphones. Specifically, the voice processing system 5B includes a voice processing device 21B, a voice processing device 22B, a voice processing device 23B, and a voice processing device 24B. The voice processing device 21B corresponds to the microphone MC1. The voice processing device 22B corresponds to the microphone MC2. The voice processing device 23B corresponds to the microphone MC3. The voice processing device 24B corresponds to the microphone MC4. Hereinafter, the voice processing device 21B, the voice processing device 22B, the voice processing device 23B, and the voice processing device 24B may be collectively referred to as the voice processing device 20B.

[0155] In Figure 9 the structure shown, it is exemplified that the voice processing device 21B, the voice processing device 22B, the voice processing device 23B, and the voice processing device 24B are constituted by respective different hardware, but the functions of the voice processing device 21B, the voice processing device 22B, the voice processing device 23B, and the voice processing device 24B may also be realized by one voice processing device 20B. Alternatively, a part of the voice processing device 21B, the voice processing device 22B, the voice processing device 23B, and the voice processing device 24B may be constituted by common hardware, and the remaining parts may be constituted by respective different hardware.

[0156] Also in the present embodiment, each voice processing device 20B is arranged in each seat near the corresponding microphone.

[0157] Figure 10 FIG. is a block diagram showing the structure of the voice processing device 21B. The voice processing device 21B, the voice processing device 22B, the voice processing device 23B, and the voice processing device 24B have the same structure and function except for a part of the structure of the filter unit described later. Here, the voice processing device 21B will be described. The voice processing device 21B targets the voice spoken by the driver hm1. The voice processing device 21B outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC1.

[0158] As Figure 10 shown, the voice processing device 21B includes a voice input unit 29B, an abnormality detection unit 31B, a filter unit F3 including a plurality of adaptive filters, a control unit 28B that controls the filter coefficients of the adaptive filters of the filter unit F3, and an addition unit 27B.

[0159] The microphones MC1, MC2, MC3, MC4, and the voice input unit 29B are the same as those in the second embodiment, and thus the description thereof is omitted.

[0160] In the present embodiment, the abnormality detection unit 31B includes a determination unit 35B. The determination unit 35B has the following function: based on the voice signal output from the microphone where no abnormality is detected, it determines which occupant's voice is more contained in the voice signal output from the microphone on the same side as the microphone where the abnormality is detected.

[0161] For example, when it is determined that the microphone MC3 is abnormal, the determination unit 35B determines which of the voices of the occupant hm3 and the occupant hm4 is more contained in the voice signal D based on the voice signal A and the voice signal B. The specific determination method is the same as the methods described in the first embodiment and the second embodiment. The structure and function of the determination unit 35B are the same as those of the determination unit described in the first embodiment, and thus the detailed description thereof is omitted.

[0162] The abnormality detection unit 31B outputs the result of the determination of whether there is an abnormality in each microphone to the control unit 28B. The determination unit 35B outputs the result of the determination of whether the voice signal C or the voice signal D contains more of the voice emitted by the occupant hm3 or the voice emitted by the occupant hm4 to the control unit 28B. The determination unit 35B outputs the determination result to the control unit 28B as a flag, for example. The flag represents a value of "0" or "1". "1" means that it is determined that there is an abnormality in the corresponding microphone, and "0" indicates that it is not determined that there is an abnormality in the corresponding microphone. Alternatively, "0" indicates that the voice signal contains more of the voice emitted by the occupant hm3, and "1" indicates that the voice signal contains more of the voice emitted by the occupant hm4. For example, when it is determined that there is no abnormality in the microphones MC1, MC2, and MC4, and it is determined that there is an abnormality in the microphone MC3, and when it is determined that the voice signal D contains more of the voice emitted by the occupant hm3, the determination unit 35B outputs the flag "0, 0, 1, 0, 0" as the determination result to the control unit 28B. The first four of the five flags in this example represent the result of the determination of whether there is an abnormality in the microphone, and the last one represents the result of the determination of which occupant's voice the voice signal contains more. The output of the result of the determination of whether there is an abnormality in the microphone by the abnormality detection unit 31B and the output of the result of the determination of which occupant's voice the voice signal contains more by the determination unit 35B can be simultaneous. Alternatively, it can also be that the abnormality detection unit 31B outputs the result of the determination of whether there is an abnormality in the microphone as a flag at the time point when the determination of whether there is an abnormality in the microphone is completed, and then the determination unit 35B outputs the result of the determination of which occupant's voice the voice signal contains more as a flag at the time point when the determination of which occupant's voice the voice signal contains more is completed.

[0163] After detecting the abnormality of each microphone, the abnormality detection unit 31B outputs the voice signal A, the voice signal B, the voice signal C, and the voice signal D to the filter unit F3.

[0164] The filter unit F3 includes an adaptive filter F3A, an adaptive filter F3B, an adaptive filter F3C, an adaptive filter F3D, and an adaptive filter F3E. The filter unit F3 is used for processing to suppress crosstalk components other than the voice of the driver hm1 included in the voice picked up by the microphone MC1. The filter unit F3 in the present embodiment is the same as the filter unit F2 in the second embodiment except that the voice signal B is input to the adaptive filter F3A instead of the second directional signal, so the detailed description is omitted. The adaptive filter F3A outputs a passed signal P3A based on the filter coefficient C3A and the voice signal B. The adaptive filter F3B outputs a passed signal P3B based on the filter coefficient C3B and the voice signal C. The adaptive filter F3C outputs a passed signal P3C based on the filter coefficient C3C and the voice signal C. The adaptive filter F3D outputs a passed signal P3D based on the filter coefficient C3D and the voice signal D. The adaptive filter F3E outputs a passed signal P3E based on the filter coefficient C3E and the voice signal D. Also in the present embodiment, the filter unit F3 may also be configured with two adaptive filters capable of inputting the voice signal B. For example, it may be that the abnormality detection unit 31B can detect an abnormality of the microphone MC2, and the filter unit F2 includes an adaptive filter F3A1 that inputs the voice signal B when an abnormality of the microphone MC2 is detected, and an adaptive filter F3A2 that inputs the voice signal B when an abnormality of the microphone MC2 is not detected, respectively.

[0165] The control unit 28B controls the filter coefficients of the adaptive filter based on the determination result of the abnormality detection unit 31B. In the present embodiment, the control unit 28B determines which of the adaptive filters F3B and F3C to input the voice signal C based on the flags output as the determination results from the abnormality detection unit 31B and the determination unit 35B. In addition, in the present embodiment, the control unit 28B determines which of the adaptive filters F3D and F3E to input the voice signal D based on the flags output as the determination results from the abnormality detection unit 31B and the determination unit 35B. The control of the filter coefficients is the same as that of the control unit 28A in the second embodiment, so the detailed description is omitted.

[0166] The addition unit 27B generates an output signal by subtracting the subtraction signal from the target voice signal output from the voice input unit 29B. In the present embodiment, the subtraction signal is a signal obtained by adding the passed signal P3A, the passed signal P3B or the passed signal P3C, and the passed signal P3D or the passed signal P3E output from the filter unit F3. The addition unit 27B outputs the output signal to the control unit 28B.

[0167] The control unit 28B outputs the output signal output from the addition unit 27B. The utilization of the output signal is the same as that in the first embodiment.

[0168] In addition, the control unit 28B updates the filter coefficients of each adaptive filter with reference to the output signal output from the addition unit 27B, the flag as the determination result output from the abnormality detection unit 31B, and the flag as the determination result output from the determination unit 35B. The update of the filter coefficients is the same as that of the control unit 28A in the second embodiment, so the detailed description is omitted.

[0169] In the present embodiment, the functions of the voice input unit 29B, the abnormality detection unit 31B, the filter unit F3, the control unit 28B, and the addition unit 27B are implemented by a processor executing a program held in a memory. Alternatively, the voice input unit 29B, the abnormality detection unit 31B, the filter unit F3, the control unit 28B, and the addition unit 27B may be constituted by different hardware.

[0170] The voice processing device 21B has been described, but the voice processing devices 22B, 23B, and 24B have almost the same structure except for the filter unit. The voice processing device 22B targets the voice spoken by the occupant hm2. The voice processing device 22B outputs, as an output signal, the voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC2. Therefore, the difference between the voice processing device 22B and the voice processing device 21B is that it has a filter unit to which the voice signals A, C, and D are input. The same applies to the voice processing devices 23B and 24B.

[0171] Figure 11It is a flowchart showing the operation process of the voice processing device 21B. First, voice signals A, B, C, and D are input to the voice input unit 29B (S201). Next, the abnormality detection unit 31B determines whether there is an abnormality in each microphone based on each voice signal (S202). The abnormality detection unit 31B may also output the determined result as a flag to the control unit 28B at this time point. When no abnormality is detected from any microphone, the abnormality detection unit 31B outputs all the voice signals to the filter unit F3. The filter unit F3 generates a subtraction signal as follows (S203). The adaptive filter F3A allows the voice signal B to pass through and outputs the passed signal P3A. The adaptive filter F3B allows the voice signal C to pass through and outputs the passed signal P3B. The adaptive filter F3D allows the voice signal D to pass through and outputs the passed signal P3D. The filter unit F3 adds the passed signals P3A, P3B, and P3D and outputs the result as the subtraction signal. The addition unit 27B subtracts the subtraction signal from the voice signal A, thereby generating an output signal and outputting the output signal (S204). The output signal is input to the control unit 28B and is output from the control unit 28B. Next, the control unit 28B refers to the flag as the determination result output from the abnormality detection unit 31B, and updates the filter coefficients of the adaptive filters F3A, F3B, and F3D based on the output signal so that the target component included in the output signal becomes the maximum (S205). Then, the voice processing device 21B performs the process S201 again.

[0172] When an abnormality is detected in one of the microphones in the process S202 (S202: "Yes"), the abnormality detection unit 31B determines whether the microphone in which the abnormality is detected is the microphone of the target seat (S206). At this time point, the abnormality detection unit 31B may also output the determined result as a flag to the control unit 28B. When the microphone in which the abnormality is detected is the microphone of the target seat (S206: "Yes"), the control unit 28B sets the intensity of the voice signal A received from the voice input unit 29B to zero and outputs it as the output signal (S207). At this time, the control unit 28B does not update the filter coefficients of the adaptive filters F3A, F3B, F3C, F3D, and F3E. Then, the voice processing device 21B performs the process S201 again.

[0173] When the microphone in which an abnormality is detected in step S206 is not the microphone of the target seat (S206: "No"), the abnormality detection unit 31B determines whether the microphone in which the abnormality is detected is the microphone on the same side as the target seat (S208). When the microphone in which the abnormality is detected is not the microphone on the same side as the target seat (S208: "No"), the abnormality detection unit 31B may also output the determination result as a flag to the control unit 28B at this point in time. The determination unit 35B determines which voice component is input to the microphone on the same side as the microphone in which the abnormality is detected and in which no abnormality is detected (S209). Hereinafter, it is assumed that an abnormality is detected in the microphone MC3 for explanation. Since it is the same as the second embodiment hereinafter, detailed description is omitted. When it is determined that the voice signal D contains a large amount of the voice emitted by the occupant hm3, the filter unit F3 uses the adaptive filter F3A and the adaptive filter F3D to generate a subtraction signal (S210). The addition unit 27B subtracts the subtraction signal from the voice signal A in the same manner as in step S4, thereby generating an output signal and outputting the output signal (S211). Next, the control unit 28B updates the filter coefficients of the adaptive filter to which the voice signal is input based on the output signal so that the target component included in the output signal becomes the maximum (S212). Then, the voice processing device 21B performs step S201 again.

[0174] When it is determined in step S209 that the voice signal D contains a large amount of the voice emitted by the occupant hm4 (S209: "hm4"), the filter unit F3 uses the adaptive filter F3A and the adaptive filter F3E to generate a subtraction signal (S213). The addition unit 27B subtracts the subtraction signal from the voice signal A in the same manner as in step S4, thereby generating an output signal and outputting the output signal (S214). Next, the control unit 28A updates the filter coefficients of the adaptive filter to which the voice signal is input based on the output signal so that the target component included in the output signal becomes the maximum (S215). Then, the voice processing device 21B performs step S201 again.

[0175] In addition, when the filter unit F3 includes two adaptive filters that can receive the voice signal B, a part of the previous process is changed as follows. For example, when the abnormality detection unit 31B can detect an abnormality in the microphone MC2, and the filter unit F3 includes an adaptive filter F3A1 that receives the voice signal B when an abnormality in the microphone MC2 is detected, and an adaptive filter F3A2 that receives the voice signal B when no abnormality in the microphone MC2 is detected, the adaptive filter F3A that received the second directional signal in the previous process is simply renamed as the adaptive filter F3A2. The following process is performed when the abnormality detection unit 31B can detect an abnormality in the microphone MC2, and the filter unit F3 includes an adaptive filter F3A1 that receives the voice signal B when an abnormality in the microphone MC2 is detected, and an adaptive filter F3A2 that receives the voice signal B when no abnormality in the microphone MC2 is detected.

[0176] When the microphone in which an abnormality is detected in the process S208 is the microphone on the same side as the target seat, the abnormality detection unit 31B outputs the determination result as a flag to the control unit 28B. In this example, an abnormality in the microphone MC2 is detected. Then, the determination unit 35B determines which voice component is input to the microphone on the same side as the microphone in which an abnormality is detected and in which no abnormality is detected (S216). For example, when an abnormality is detected in the microphone MC2, the determination unit 35B determines which of the voice of the driver hm1 and the voice of the passenger hm2 is input to the microphone MC1. In other words, the determination unit 35B determines which of the voice of the driver hm1 and the voice of the passenger hm2 is included more in the voice signal A. The determination unit 35B outputs the determination result as a flag to the control unit 28B.

[0177] When the voice signal A includes more of the voice of the passenger hm2, the control unit 28B sets the intensity of the voice signal A to zero and outputs it as an output signal (S207). At this time, the control unit 28B does not update the filter coefficients of the adaptive filter F3A1, the adaptive filter F3A2, the adaptive filter F3B, the adaptive filter F3C, the adaptive filter F3D, and the adaptive filter F3E. Then, the voice processing device 21B performs the process S201 again.

[0178] When the voice signal A contains more voice emitted by the driver hm1, the filter unit F3 generates a subtraction signal (S217) as follows. The control unit 28B controls the filter unit F3 so that the voice signal B is input to the adaptive filter F3A1 in a state where the intensity is zero. On the other hand, the control unit 28B controls the filter unit F3 so that the voice signal C is input to the adaptive filter F3B. In addition, the control unit 28B controls the filter unit F3 so that the voice signal D is input to the adaptive filter F3D. In other words, the control unit 28B does not change the intensity of the voice signal C input to the adaptive filter F3B and the voice signal D input to the adaptive filter F3D, but changes the intensity of the voice signal B input to the adaptive filter F3A1 to zero. The adaptive filter F3B passes the voice signal C and outputs a passed signal P3B. The adaptive filter F3D passes the voice signal D and outputs a passed signal P3D. The filter unit F3 adds the passed signal P3B and the passed signal P3D and outputs the result as a subtraction signal. The addition unit 27B subtracts the subtraction signal from the voice signal A to generate an output signal and outputs the output signal (S218). The output signal is input to the control unit 28B and is output from the control unit 28B. Next, the control unit 28B refers to the flag as the determination result output from the abnormality detection unit 31B, and updates the filter coefficients of the adaptive filter F3B and the adaptive filter F3D based on the output signal so that the target component included in the output signal becomes the maximum (S219). Then, the voice processing device 21B performs the process S201 again.

[0179] In addition, an example in the case where the abnormality detection unit 31B can detect abnormalities of the microphones MC1 and MC2 has been described, but it may also be the case where the abnormality detection unit 31B can detect abnormalities of only the microphones MC3 and MC4. In this case, in the Figure 11 flowchart shown, the processes S206, S207, S208, and processes S216 to S219 can be omitted.

[0180] In the present embodiment, for an adaptive filter for a voice signal in a state where the input intensity is zero, the update of the filter coefficients is not performed. Thus, compared with the case where the filter coefficients of all adaptive filters are always updated, the processing amount of the control unit 28B can be reduced. On the other hand, the control unit 28B may always update the filter coefficients of all adaptive filters. By always updating the filter coefficients of all adaptive filters, the control unit 28B can always perform the same processing, so the processing becomes simple. In addition, by always updating the filter coefficients of all adaptive filters, for example, for a certain adaptive filter, even immediately after changing from a state where a voice signal with an input intensity of zero is input to a state where a voice signal with a non-zero input intensity is input, the filter coefficients can be updated with high accuracy.

[0181] Thus, in the voice processing system 5B in the third embodiment, the same effect as that of the voice processing system 5A in the second embodiment can also be obtained.

[0182] (Fourth Embodiment)

[0183] The difference between the voice processing system 5C according to the fourth embodiment and the voice processing system 5 according to the first embodiment is that a voice processing device 20C is provided instead of the voice processing device 20. The voice processing device 20C according to the fourth embodiment performs crosstalk component cancellation processing using the voice signal output from the microphone without determining which occupant's voice is input to the microphone capable of inputting voices emitted by a plurality of occupants. Hereinafter, Figure 12 、 Figure 13 and Figure 14 are used to describe the voice processing device 20C. For the structures and operations that are the same as those described in the first embodiment, the same reference numerals are used, and thus their descriptions are omitted or simplified.

[0184] Using Figure 12 to describe the details of the voice processing system 5C in the fourth embodiment. Figure 12 is a diagram showing an example of the schematic structure of the voice processing system 5C in the fourth embodiment. The voice processing system 5C includes a microphone MC1, a microphone MC2, a microphone MC3, and a voice processing device 20C. Regarding the microphone MC1, the microphone MC2, and the microphone MC3, they are the same as those in the first embodiment, and thus the description thereof is omitted.

[0185] In the present embodiment, the voice processing system 5C includes a plurality of voice processing devices 20C corresponding to the respective microphones. Specifically, the voice processing system 5C includes a voice processing device 21C, a voice processing device 22C, and a voice processing device 23C. The voice processing device 21C corresponds to the microphone MC1. The voice processing device 22C corresponds to the microphone MC2. The voice processing device 23C corresponds to the microphone MC3. Hereinafter, the voice processing device 21C, the voice processing device 22C, and the voice processing device 23C may be collectively referred to as the voice processing device 20C.

[0186] In Figure 12 the structure shown, it is exemplified that the voice processing device 21C, the voice processing device 22C, and the voice processing device 23C are constituted by different hardware respectively, but the functions of the voice processing device 21C, the voice processing device 22C, and the voice processing device 23C may also be realized by one voice processing device 20C. Alternatively, a part of the voice processing device 21C, the voice processing device 22C, and the voice processing device 23C may be constituted by common hardware, and the remaining parts may be constituted by different hardware respectively.

[0187] Also in the present embodiment, each voice processing device 20C is arranged in each seat near the corresponding microphone. Regarding the position of the voice processing device 20C, for example, it is the same as that in the first embodiment.

[0188] Figure 13 is a block diagram showing the structure of the voice processing device 21C. The voice processing device 21C, the voice processing device 22C, and the voice processing device 23C have the same structure and function except for a part of the structure of the filter unit described later. Here, the voice processing device 21C will be described. The voice processing device 21C takes the voice spoken by the driver hm1 as the target component. The voice processing device 21C outputs the voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC1 as the output signal.

[0189] As Figure 13 shown, the voice processing device 21C includes a voice input unit 29C, a directivity control unit 30C, a filter unit F4 including a plurality of adaptive filters, a control unit 28C for controlling the filter coefficients of the plurality of adaptive filters, and an adder 27C.

[0190] The voice input unit 29C is the same as the voice input unit 29 in the first embodiment, and thus the description thereof is omitted.

[0191] The voice signals A, B, and C output from the voice input unit 29C are input to the directivity control unit 30C. The directivity control unit 30C performs directivity control processing using the voice signals A and B. Then, the directivity control unit 30C outputs a first directivity signal obtained by performing directivity control processing on the voice signal A. In addition, the directivity control unit 30C outputs a second directivity signal obtained by performing directivity control processing on the voice signal B. The directivity control unit 30C outputs the first directivity signal to the addition unit 27C, and outputs the second directivity signal and the voice signal C to the filter unit F4.

[0192] In addition, the directivity control unit 30C determines whether a voice component is input to the microphone MC3. For example, the directivity control unit 30C determines that a voice signal is input to the microphone MC3 when the intensity of the voice signal C is greater than at least one of the intensities of the first directivity signal and the second directivity signal, and otherwise determines that no voice signal is input to the microphone MC3.

[0193] The directivity control unit 30C outputs the result of the determination as to whether a voice component is input to the microphone MC3 to the control unit 28C. The directivity control unit 30C outputs the result of the determination to the control unit 28C as a flag, for example. The flag represents a value of "0" or "1". "0" indicates that no voice component is input to the microphone MC3, and "1" indicates that a voice component is input to the microphone MC3.

[0194] In the present embodiment, the directivity control unit 30C determines whether a voice component is input to the microphone MC3. However, it may also be that the voice processing device 21C includes a speech determination unit as a determination unit separate from the directivity control unit 30C, and the speech determination unit makes the determination. In this case, the speech determination unit is connected between the voice input unit 29C and the directivity control unit 30C, for example. Alternatively, it may also be that the voice processing device 21C only includes the speech determination unit and does not include the directivity control unit 30C. The structure and function of the speech determination unit are the same as those of the determination unit 35 described in the first embodiment, and thus the detailed description is omitted.

[0195] The filter unit F4 includes an adaptive filter F4A and an adaptive filter F4B. The filter unit F4 is used for processing to suppress crosstalk components other than the voice of the driver hm1 included in the voice picked up by the microphone MC1. In the present embodiment, the filter unit F4 includes two adaptive filters, but the number of adaptive filters can be appropriately set based on the number of input voice signals and the processing amount of crosstalk suppression processing. The details of the crosstalk suppression processing will be described later.

[0196] A second directional signal is input to the adaptive filter F4A as a reference signal. The adaptive filter F4A outputs a passed signal P4A based on the filter coefficients C4A and the second directional signal. A voice signal C is input to the adaptive filter F4B as a reference signal. In the present embodiment, whether the voice signal C contains a large amount of voice emitted by the occupant hm3 or a large amount of voice emitted by the occupant hm4, the voice signal C is input to the adaptive filter F4B. The adaptive filter F4B outputs a passed signal P4B based on the filter coefficients C4B and the voice signal C. The filter unit F4 adds the passed signal P4A and the passed signal P4B and outputs the result. In the present embodiment, the adaptive filter F4A and the adaptive filter F4B are implemented by a processor executing a program. The adaptive filter F4A and the adaptive filter F4B may also be physically separate and different hardware structures.

[0197] The adder 27C subtracts the subtraction signal from the target voice signal output from the voice input unit 29C, thereby generating an output signal. In the present embodiment, the subtraction signal is a signal obtained by adding the passed signal P4A and the passed signal P4B output from the filter unit F4. The adder 27C outputs the output signal to the control unit 28C.

[0198] The control unit 28C outputs the output signal output from the adder 27C. The utilization of the output signal is the same as that in the first embodiment.

[0199] In addition, the control unit 28C updates the filter coefficients of each adaptive filter with reference to the output signal output from the adder 27C. Specifically, the control unit 28C updates the filter coefficients of the adaptive filter F4A and the adaptive filter F4B so that the value of the error signal in Equation (1) approaches 0. The specific method for updating the filter coefficients is the same as the method described in the first embodiment.

[0200] In the present embodiment, the functions of the voice input unit 29C, the directivity control unit 30C, the filter unit F4, the control unit 28C, and the adder 27C are implemented by a processor executing a program held in a memory. Alternatively, the voice input unit 29C, the directivity control unit 30C, the filter unit F4, the control unit 28C, and the adder 27C may also be constituted by different hardware.

[0201] The voice processing device 21C has been described. However, the voice processing devices 22C and 23C have almost the same structure except for the filter section. The voice processing device 22C targets the voice spoken by the occupant hm2. The voice processing device 22C outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC2. Therefore, the difference between the voice processing device 22C and the voice processing device 21C is that it has a filter section to which the first directional signal and the voice signal C are input. The same applies to the voice processing device 23C.

[0202] Figure 14 is a flowchart showing the operation process of the voice processing device 21C. First, the voice signals A, B, and C are input to the voice input section 29C (S301). Next, the directivity control section 30C performs directivity control processing using the voice signals A and B to generate the first directional signal and the second directional signal (S302). Then, the directivity control section 30C determines whether a voice component has been input to the microphone MC3 (S303). The directivity control section 30C outputs the determination result as a flag to the control section 28C. When the directivity control section 30C determines that no voice signal has been input to the microphone MC3 (S303: "No"), the control section 28C sets the intensity of the voice signal C input to the filter section F4 to zero and does not change the intensity of the second directional signal. Then, the filter section F4 generates a subtraction signal as follows (S304). The adaptive filter F4A passes the second directional signal and outputs the passed signal P4A. The adaptive filter F4B passes the voice signal C and outputs the passed signal P4B. The filter section F4 adds the passed signal P4A and the passed signal P4B and outputs the result as the subtraction signal. The addition section 27C subtracts the subtraction signal from the first directional signal to generate and output an output signal (S305). The output signal is input to the control section 28C and is output from the control section 28C. Next, the control section 28C updates the filter coefficients of the adaptive filter F4A based on the output signal so that the target component included in the output signal becomes maximum (S306). Then, the voice processing device 21 repeats the process of step S301.

[0203] When the directivity control unit 30C determines that a voice signal has been input to the microphone MC3 (S303: "Yes"), the filter unit F4 generates a subtraction signal as follows (S307). The control unit 28C controls the filter unit F4 so that the voice signal C is input to the adaptive filter F4B. Then, the filter unit F4 generates a subtraction signal through the same operation as in step S304. The addition unit 27C subtracts the subtraction signal from the first directivity signal in the same manner as in step S305, thereby generating an output signal and outputting the output signal (S308). Next, the control unit 28C updates the filter coefficients of the adaptive filter to which the voice signal is input based on the output signal so that the target component included in the output signal becomes maximum (S310). Specifically, the filter coefficients of the adaptive filter F4A and the adaptive filter F4B are updated. Then, the voice processing device 21C performs step S301 again.

[0204] In the present embodiment, for an adaptive filter in a state where the input intensity of the voice signal is zero, the update of the filter coefficients is not performed. Thereby, compared with the case where the filter coefficients are always updated for all adaptive filters, the processing amount of the control unit 28C can be reduced. On the other hand, the control unit 28C may always update the filter coefficients for all adaptive filters. By always updating the filter coefficients for all adaptive filters, the control unit 28C can always perform the same processing, so the processing becomes simple. In addition, by always updating the filter coefficients for all adaptive filters, for example, for a certain adaptive filter, even immediately after changing from a state where the input intensity of the voice signal is zero to a state where the input intensity of the voice signal is not zero, the filter coefficients can be updated with high accuracy.

[0205] An example of each voice signal and output signal in the voice processing device 21C is shown in FIG. 15. Figure 15A The spectrum of the first directivity signal is shown, Figure 15B The spectrum of the second directivity signal is shown, Figure 15C The spectrum of the voice signal C is shown, Figure 15D The spectrum of the output signal is shown. In FIG. 15, an example is shown in the case where the driver hm1, the passengers hm2, hm3, and hm4 are speaking simultaneously, and the driver hm1 intermittently speaks a specific word while the other passengers are chatting continuously. In addition, in the first directivity signal and the second directivity signal, directivity control processing is being performed, so the S / N ratio is higher than that of the voice signal C. By comparing Figure 15A with Figure 15D it can be seen that by performing the process of suppressing the crosstalk component, the S / N ratio becomes higher in the output signal than in the first directivity signal.

[0206] Thus, also in the voice processing system 5C of the fourth embodiment, multiple voice signals are acquired by multiple microphones. Using another voice signal as a reference signal, a subtraction signal generated by an adaptive filter is subtracted from a certain voice signal, thereby accurately obtaining the voice of a specific speaker. In the fourth embodiment, it is configured to be able to pick up multiple voices with different generation positions using one microphone. Specifically, the microphone MC3 picks up the voices of the occupant hm3 and the occupant hm4 in the rear seat. On this basis, regardless of whether the voice signal C output from the microphone MC3 contains the voice of the occupant hm3 or the voice of the occupant hm4, the voice signal C is input to the adaptive filter F4B. Thus, even in the case of picking up multiple voices using one microphone, the voice signal of the target component can be accurately obtained. Therefore, for example, microphones do not need to be provided for each seat one by one, so the cost can be reduced. In addition, when using an adaptive filter to obtain the target component, compared with the case of using the signals output from the microphones provided in all seats as reference signals, the number of reference signals used in the process can be reduced. Thereby, the amount of processing for eliminating crosstalk components can be reduced. In addition, in the fourth embodiment, the process of determining which occupant's voice is included in the voice signal is not performed, nor is a structure adopted in which adaptive filters are separately used according to the occupant whose voice is included in the voice signal. Therefore, the amount of processing for eliminating crosstalk components can be reduced, and the structure of the voice processing system 5C can also be simplified. In addition, for the adaptive filter of the voice signal in a state where the input intensity is zero, the update of the filter coefficient may not be performed. Thereby, compared with the case where the filter coefficients of all adaptive filters are always updated, the processing amount can be further reduced.

[0207] (Fifth Embodiment)

[0208] The difference between the voice processing system 5D according to the fifth embodiment and the voice processing system 5C according to the fourth embodiment is that a voice processing device 20D is provided instead of the voice processing device 20C. The voice processing device 20D according to the fifth embodiment inputs the voice signal output from the microphone capable of inputting the voices of multiple occupants to multiple adaptive filters. The multiple adaptive filters include an adaptive filter that supports the case where the voice of one occupant input to the microphone is input, and an adaptive filter that supports the case where the voice of other occupants input to the microphone is input. The voice processing device 20D determines which adaptive filter can make the crosstalk component smaller, and uses the adaptive filter that can make the crosstalk component smaller to perform the process of eliminating the crosstalk component. Next, use Figure 16 , Figure 17 and Figure 18To describe the voice processing device 20D. For the structures and operations that are the same as those described in the first and fourth embodiments, the same reference numerals are used, and thus their descriptions are omitted or simplified.

[0209] Use Figure 16 To describe the details of the voice processing system 5D in the fifth embodiment. Figure 16 It is a diagram showing an example of the schematic structure of the voice processing system 5D in the fifth embodiment. The voice processing system 5D includes a microphone MC1, a microphone MC2, a microphone MC3, and a voice processing device 20D. Regarding the microphone MC1, the microphone MC2, and the microphone MC3, they are the same as those in the first embodiment, and thus their descriptions are omitted.

[0210] In the present embodiment, the voice processing system 5D includes a plurality of voice processing devices 20D corresponding to the respective microphones. Specifically, the voice processing system 5D includes a voice processing device 21D, a voice processing device 22D, and a voice processing device 23D. The voice processing device 21D corresponds to the microphone MC1. The voice processing device 22D corresponds to the microphone MC2. The voice processing device 23D corresponds to the microphone MC3. Hereinafter, the voice processing device 21D, the voice processing device 22D, and the voice processing device 23D may be collectively referred to as the voice processing device 20D.

[0211] In Figure 16 In the structure shown, it is exemplified that the voice processing device 21D, the voice processing device 22D, and the voice processing device 23D are constituted by different hardware, respectively, but the functions of the voice processing device 21D, the voice processing device 22D, and the voice processing device 23D may also be realized by one voice processing device 20D. Alternatively, a part of the voice processing device 21D, the voice processing device 22D, and the voice processing device 23D may be constituted by common hardware, and the remaining parts may be constituted by different hardware, respectively.

[0212] Also in the present embodiment, each voice processing device 20D is arranged in each seat near the corresponding microphone. Regarding the position of the voice processing device 20D, for example, it is the same as that in the first embodiment.

[0213] Figure 17 It is a block diagram showing the structure of the voice processing device 21D. The voice processing device 21D, the voice processing device 22D, and the voice processing device 23D have the same structures and functions except for a part of the structure of the filter unit described later. Here, the voice processing device 21D is described. The voice processing device 21D targets the voice spoken by the driver hm1. The voice processing device 21D outputs, as an output signal, the voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC1.

[0214] As Figure 17 shown, the voice processing device 21D includes a voice input unit 29D, a directivity control unit 30D, a filter unit F5 including a plurality of adaptive filters, a control unit 28D that controls the filter coefficients of the plurality of adaptive filters, and an addition unit 27D.

[0215] The voice input unit 29D is the same as the voice input unit 29 of the first embodiment, and thus the description thereof is omitted.

[0216] The directivity control unit 30D is the same as the directivity control unit 30C of the fourth embodiment, and thus the description thereof is omitted. The voice processing system 5D may also include a speech determination unit as a determination unit. In the case of including the speech determination unit, the voice processing system 5D may not include the directivity control unit 30D.

[0217] The filter unit F5 includes an adaptive filter F5A, an adaptive filter F5B, an adaptive filter F5C, and an adaptive filter F5D. The filter unit F5 is used for processing to suppress crosstalk components other than the voice of the driver hm1 included in the voice received by the microphone MC1. In the present embodiment, the filter unit F5 includes four adaptive filters, but the number of adaptive filters can be appropriately set based on the number of input voice signals and the processing amount of the crosstalk suppression processing. Details of the processing for suppressing crosstalk are described later.

[0218] A second directional signal is input to the adaptive filter F5A as a reference signal. The adaptive filter F5A outputs a passed signal P5A based on the filter coefficient C5A and the second directional signal. A voice signal C is input to the adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D as a reference signal. The adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D correspond to "two or more adaptive filters". The adaptive filter F5B corresponds to the first adaptive filter. The adaptive filter F5C corresponds to the second adaptive filter. The adaptive filter F5D corresponds to the third adaptive filter. The adaptive filter F5B outputs a passed signal P5B based on the filter coefficient C5B and the voice signal C. The passed signal P5B corresponds to the first passed signal. The adaptive filter F5C outputs a passed signal P5C based on the filter coefficient C5C and the voice signal C. The passed signal P5C corresponds to the second passed signal. The adaptive filter F5D outputs a passed signal P5D based on the filter coefficient C5D and the voice signal C. The filter unit F5 outputs a subtracted signal SSA obtained by adding the passed signal P5A and the passed signal P5B, a subtracted signal SSB obtained by adding the passed signal P5A and the passed signal P5C, and a subtracted signal SSC obtained by adding the passed signal P5A and the passed signal P5D. The subtracted signal SSA corresponds to the first subtracted signal. The subtracted signal SSB corresponds to the second subtracted signal. In the present embodiment, the adaptive filter F5A, the adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D are implemented by a processor executing a program. The adaptive filter F5A, the adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D may also be physically separate different hardware structures.

[0219] The filter coefficient C5B of the adaptive filter F5B is updated to minimize the error signal when the voice signal C contains more voice emitted by the occupant hm3. In addition, the filter coefficient C5C of the adaptive filter F5C is updated to minimize the error signal when the voice signal C contains more voice emitted by the occupant hm4. On the other hand, the filter coefficient C5D of the adaptive filter F5D is updated to minimize the error signal when the voice signal C contains both the voice emitted by the occupant hm3 and the voice emitted by the occupant hm4.

[0220] In the present embodiment, the filter unit F5 includes the adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D as the adaptive filters to which the voice signal C is input, but may also include only the adaptive filter F5B and the adaptive filter F5C as the adaptive filters to which the voice signal C is input. In this case, the processing amount of crosstalk cancellation described later can be reduced.

[0221] The adder unit 27D generates an output signal by subtracting the subtraction signal from the first directivity signal of the voice signal as a target output from the voice input unit 29D. In the present embodiment, an output signal OSA when using the subtraction signal SSA, an output signal OSB when using the subtraction signal SSB, and an output signal OSC when using the subtraction signal SSC are generated respectively. The output signal OSA corresponds to the first output signal. The output signal OSB corresponds to the second output signal. The adder unit 27D outputs the output signal OSA, the output signal OSB, and the output signal OSC to the control unit 28D.

[0222] The control unit 28D determines the output signal with the smallest error signal by referring to the output signal OSA, the output signal OSB, and the output signal OSC output from the adder unit 27D. For example, when the voice signal C contains a lot of voice emitted by the occupant hm3, the error signal becomes the smallest in the output signal OSA. For example, when the voice signal C contains a lot of voice emitted by the occupant hm4, the error signal becomes the smallest in the output signal OSB. For example, when the voice signal C contains both the voice emitted by the occupant hm3 and the voice emitted by the occupant hm4, the error signal becomes the smallest in the output signal OSC. Then, the control unit 28D updates the filter coefficients of the adaptive filter used when generating the output signal with the smallest error signal. The specific method for updating the filter coefficients is the same as the method described in the first embodiment.

[0223] In addition, the control unit 28D outputs the output signal with the smallest error signal among the output signal OSA, the output signal OSB, and the output signal OSC. The utilization of the output signal is the same as that in the first embodiment.

[0224] In the present embodiment, the functions of the voice input unit 29D, the directivity control unit 30D, the filter unit F5, the control unit 28D, and the adder unit 27D are implemented by a processor executing a program held in a memory. Alternatively, the voice input unit 29D, the directivity control unit 30D, the filter unit F5, the control unit 28D, and the adder unit 27D may be constituted by different hardware.

[0225] The voice processing device 21D has been described. However, the voice processing devices 22D and 23D have almost the same structure except for the filter section. The voice processing device 22D targets the voice spoken by the occupant hm2. The voice processing device 22D outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC2. Therefore, the difference between the voice processing device 22D and the voice processing device 21D is that it has a filter section that is input with the first directivity signal and the voice signal C. The same applies to the voice processing device 23D.

[0226] Figure 18 is a flowchart showing the operation process of the voice processing device 21D. First, the voice signals A, B, and C are input to the voice input section 29D (S401). Next, the directivity control section 30D performs directivity control processing using the voice signals A and B to generate the first directivity signal and the second directivity signal (S402). Then, the directivity control section 30D determines whether a voice component has been input to the microphone MC3 by the same method as in the first embodiment (S403). The directivity control section 30D outputs the determination result as a flag to the control section 28D. When the directivity control section 30D determines that no voice signal has been input to the microphone MC3 (S403: "No"), the control section 28D sets the intensity of the voice signal C input to the filter section F5 to zero and does not change the intensity of the second directivity signal. Then, the filter section F5 generates a subtraction signal as follows (S404). The adaptive filter F5A passes the second directivity signal and outputs the passed signal P5A. The adaptive filter F5B passes the voice signal C and outputs the passed signal P5B. The adaptive filter F5C passes the voice signal C and outputs the passed signal P5C. The adaptive filter F5D passes the voice signal C and outputs the passed signal P5D. The filter section F5 adds the passed signals P5A, P5B, P5C, and P5D and outputs the result as the subtraction signal. The addition section 27D subtracts the subtraction signal from the first directivity signal to generate and output an output signal (S405). The output signal is input to the control section 28D and output from the control section 28D. Next, the control section 28D updates the filter coefficients of the adaptive filter F5A based on the output signal so that the target component included in the output signal becomes maximum (S406). Then, the voice processing device 21 repeats step S1 again.

[0227] When the directivity control unit 30D determines that a voice signal has been input to the microphone MC3 (S403: "Yes"), the control unit 28D controls the filter unit F5 so that the voice signal C is input to the adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D, respectively. In other words, the control unit 28D does not change the intensity of the second directivity signal input to the adaptive filter F5A and the voice signal C input to the adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D. Then, the filter unit F5 generates a subtraction signal as follows (S407). The filter unit F5 generates a subtraction signal SSA obtained by adding the passing signal P5A and the passing signal P5B, a subtraction signal SSB obtained by adding the passing signal P5A and the passing signal P5C, and a subtraction signal SSC obtained by adding the passing signal P5A and the passing signal P5D, and outputs them to the addition unit 27D. The addition unit 27D generates an output signal as follows and outputs it to the control unit 28D (S408). The addition unit 27D subtracts the subtraction signal SSA from the first directivity signal to generate an output signal OSA and outputs it to the control unit 28D. The addition unit 27D subtracts the subtraction signal SSB from the first directivity signal to generate an output signal OSB and outputs it to the control unit 28D. In addition, the addition unit 27D subtracts the subtraction signal SSC from the first directivity signal to generate an output signal OSC and outputs it to the control unit 28D. Next, the control unit 28D determines in which case the error signal becomes the minimum based on the output signal OSA, the output signal OSB, and the output signal OSC (S409). When it is determined that the error signal becomes the minimum when the adaptive filter F5B is used, the control unit 28D updates the filter coefficients of the adaptive filter to which the voice signal is input so that the target component included in the output signal OSA becomes the maximum (S410). Specifically, the filter coefficients of the adaptive filter F5A and the adaptive filter F5B are updated. Then, the voice processing device 21D performs the process S401 again.

[0228] When it is determined in the process S409 that the error signal becomes the minimum when the adaptive filter F5C is used, the control unit 28D updates the filter coefficients of the adaptive filter to which the voice signal is input so that the target component included in the output signal OSB becomes the maximum (S411). Specifically, the filter coefficients of the adaptive filter F5A and the adaptive filter F5C are updated. Then, the voice processing device 21D performs the process S401 again.

[0229] When it is determined in step S409 that the error signal becomes minimum when the adaptive filter F5D is used, the control unit 28D updates the filter coefficients of the adaptive filter for the input speech signal so that the target component included in the output signal OSC becomes maximum (S412). Specifically, the filter coefficients of the adaptive filter F5A and the adaptive filter F5D are updated. Then, the speech processing device 21D performs step S401 again.

[0230] In the present embodiment, for the adaptive filter for the speech signal in the state where the input intensity is zero, the update of the filter coefficients is not performed. Thereby, compared with the case where the filter coefficients are always updated for all the adaptive filters, the processing amount of the control unit 28D can be reduced. On the other hand, the control unit 28D may always update the filter coefficients for all the adaptive filters. By always updating the filter coefficients for all the adaptive filters, since the control unit 28D can always perform the same processing, the processing becomes simple. In addition, by always updating the filter coefficients for all the adaptive filters, for example, for a certain adaptive filter, even immediately after changing from the state of the speech signal with the input intensity of zero to the state of the speech signal with the input intensity not zero, the filter coefficients can be updated with high accuracy.

[0231] Thus, also in the voice processing system 5D of the fifth embodiment, multiple voice signals are acquired by multiple microphones. Using another voice signal as a reference signal, a subtraction signal generated using an adaptive filter is subtracted from a certain voice signal, thereby accurately obtaining the voice of a specific speaker. In the fifth embodiment, it is configured to be able to pick up multiple voices with different generation positions using one microphone. Specifically, the voice processing system 5D uses the microphone MC3 to pick up the voices of the occupant hm3 and the occupant hm4 in the rear seat. On this basis, output signals are respectively generated when the voice signal C is input to the adaptive filter F5B, the adaptive filter F5C, and the adaptive filter F5D. The voice processing system 5D determines the output signal when the error signal becomes the minimum. Thus, even in the case of picking up multiple voices using one microphone, the voice signal of the target component can be accurately obtained. Therefore, for example, microphones do not need to be provided for each seat one by one, so the cost can be reduced. In addition, when using an adaptive filter to obtain the target component, compared with the case where the signals output from the microphones provided in all seats are used as reference signals, the number of reference signals used in the processing can be reduced. Thereby, the amount of processing for canceling the crosstalk component can be reduced. In addition, in the fifth embodiment, the process of determining which occupant's voice is included in the voice signal is not performed. Therefore, the amount of processing for canceling the crosstalk component can be reduced. In addition, for the adaptive filter in a state where the input intensity is zero for the voice signal, the update of the filter coefficients may not be performed. Thereby, compared with the case where the filter coefficients are always updated for all adaptive filters, the processing amount can be further reduced.

[0232] (Sixth Embodiment)

[0233] The difference between the voice processing system 5E according to the sixth embodiment and the voice processing system 5A according to the second embodiment is that the voice processing device 20E is provided instead of the voice processing device 20A. The voice processing device 20E according to the sixth embodiment uses the signal obtained by summing the voice signals output from multiple microphones as a reference signal to perform the process of canceling the crosstalk component. Next, Figure 19 , Figure 20 and Figure 21 are used to describe the voice processing device 20E. For the structures and operations that are the same as those described in the first and second embodiments, the same reference numerals are used, thereby omitting or simplifying the description thereof.

[0234] Using Figure 19 to describe the details of the voice processing system 5E in the sixth embodiment. Figure 19This is a diagram showing an example of the schematic structure of the voice processing system 5E in the sixth embodiment. The voice processing system 5E includes a microphone MC1, a microphone MC2, a microphone MC3, a microphone MC4, and a voice processing device 20E. Regarding the microphones MC1, MC2, MC3, and MC4, they are the same as those in the second embodiment, so the description thereof is omitted.

[0235] In this embodiment, the voice processing system 5E includes a plurality of voice processing devices 20E corresponding to the respective microphones. Specifically, the voice processing system 5E includes a voice processing device 21E, a voice processing device 22E, a voice processing device 23E, and a voice processing device 24E. The voice processing device 21E corresponds to the microphone MC1. The voice processing device 22E corresponds to the microphone MC2. The voice processing device 23E corresponds to the microphone MC3. The voice processing device 24E corresponds to the microphone MC4. Hereinafter, the voice processing device 21E, the voice processing device 22E, the voice processing device 23E, and the voice processing device 24E may be collectively referred to as the voice processing device 20E.

[0236] In Figure 19 In the structure shown, it is exemplified that the voice processing device 21E, the voice processing device 22E, the voice processing device 23E, and the voice processing device 24E are constituted by different hardware respectively, but the functions of the voice processing device 21E, the voice processing device 22E, the voice processing device 23E, and the voice processing device 24E may also be realized by one voice processing device 20E. Alternatively, a part of the voice processing device 21E, the voice processing device 22E, the voice processing device 23E, and the voice processing device 24E may be constituted by common hardware, and the remaining parts may be constituted by different hardware respectively.

[0237] In this embodiment, each voice processing device 20E is arranged in each seat near the corresponding microphone. Regarding the position of the voice processing device 20E, for example, it is the same as that in the second embodiment.

[0238] Figure 20 This is a block diagram showing the structure of the voice processing device 21E. The voice processing device 21E, the voice processing device 22E, the voice processing device 23E, and the voice processing device 24E have the same structure and function except for a part of the structure of the filter unit described later. Here, the voice processing device 21E is described. The voice processing device 21E targets the voice spoken by the driver hm1. The voice processing device 21E outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC1.

[0239] As Figure 20As shown, the voice processing device 21E includes a voice input unit 29E, a directivity control unit 30E, a filter unit F6 including a plurality of adaptive filters, a control unit 28E that controls the filter coefficients of the adaptive filters of the filter unit F6, and an addition unit 27E.

[0240] The voice input unit 29E is the same as the voice input unit 29A of the second embodiment, and thus the description thereof is omitted.

[0241] The voice signals A, B, C, and D output from the voice input unit 29E are input to the directivity control unit 30E. The directivity control unit 30E performs directivity control processing using the voice signals output from the microphone near the seat of the target occupant and the microphone on the same side as the microphone. In the voice processing device 21E, since the voice spoken by the driver hm1 is the target, the directivity control unit 30E performs directivity control processing using the voice signals A and B. Then, the directivity control unit 30E outputs two directivity signals obtained by performing directivity control processing using the two voice signals. For example, the directivity control unit 30E outputs a first directivity signal obtained by performing directivity control processing on the voice signal A. In addition, the directivity control unit 30E outputs a second directivity signal obtained by performing directivity control processing on the voice signal B. The directivity control unit 30E may also perform directivity control processing using all the voice signals and output the obtained directivity signals. For example, in addition to outputting the first directivity signal and the second directivity signal, the directivity control unit 30E also outputs a third directivity signal obtained by performing directivity control processing on the voice signal C and a fourth directivity signal obtained by performing directivity control processing on the voice signal D.

[0242] In addition, the directivity control unit 30E determines whether a voice component is input to the microphone on the side different from the microphone near the seat of the target occupant. Specifically, the directivity control unit 30E determines whether a voice component is input to the microphones MC3 and MC4. For example, when the intensity of the voice signal C is greater than at least one of the intensities of the first directivity signal and the second directivity signal, the directivity control unit 30 determines that a voice signal is input to the microphone MC3, otherwise, it determines that no voice signal is input to the microphone MC3. The same applies to the microphone MC4.

[0243] In the present embodiment, the directivity control unit 30E determines whether a voice component is input to a microphone on a side different from the microphone near the seat of the target occupant. However, the voice processing device 21E may include a speech determination unit as a determination unit separate from the directivity control unit 30E, and the speech determination unit makes the determination. In this case, the speech determination unit is connected, for example, between the voice input unit 29E and the directivity control unit 30E. The structure and function of the speech determination unit are the same as those of the determination unit described in the first embodiment, and thus detailed description thereof is omitted. When the speech determination unit is provided, the voice processing system 5E may not include the directivity control unit 30E.

[0244] The filter unit F6 includes an adaptive filter F6A and an adaptive filter F6B. The filter unit F6 is for processing to suppress crosstalk components other than the voice of the driver hm1 included in the voice picked up by the microphone MC1. In the present embodiment, the filter unit F6 includes two adaptive filters, but the number of adaptive filters can be appropriately set based on the number of input voice signals and the processing amount of the crosstalk suppression processing. Details of the crosstalk suppression processing will be described later.

[0245] A second directivity signal is input to the adaptive filter F6A as a reference signal. The adaptive filter F6A outputs a passed signal P6A based on the filter coefficient C6A and the second directivity signal. A voice signal C and a voice signal D are input to the adaptive filter F6B as reference signals. The adaptive filter F6B outputs a passed signal P6B based on the filter coefficient C6B, the voice signal C, and the voice signal D. The adaptive filter F6B corresponds to "the adaptive filter to which the first signal and the second signal are input". The filter unit F6 adds the passed signal P6A and the passed signal P6B and outputs the result. In the present embodiment, the adaptive filter F6A and the adaptive filter F6B are implemented by a processor executing a program. The adaptive filter F6A and the adaptive filter F6B may also be physically separate different hardware structures.

[0246] The addition unit 27E generates an output signal by subtracting the subtraction signal from the first directivity signal of the target voice signal output from the voice input unit 29E. In the present embodiment, the subtraction signal is a signal obtained by adding the passed signal P6A and the passed signal P6B output from the filter unit F6. The addition unit 27E outputs the output signal to the control unit 28E.

[0247] The control unit 28E outputs the output signal output from the addition unit 27E. The output signal of the control unit 28E is input to the speech recognition engine 40. Alternatively, the output signal may be directly input from the control unit 28E to the electronic device 50. When the output signal is directly input from the control unit 28E to the electronic device 50, the control unit 28E and the electronic device 50 may be connected by a wired method or a wireless method. For example, the electronic device 50 may be a portable terminal, and the output signal may be directly input from the control unit 28E to the portable terminal via a wireless communication network. The output signal input to the portable terminal may also be output as speech from the speaker of the portable terminal.

[0248] In addition, the control unit 28E updates the filter coefficients of each adaptive filter based on the output signal output from the addition unit 27E. The control unit 28E updates the filter coefficients of each adaptive filter so that the value of the error signal in Equation (1) approaches 0. The specific method for updating the filter coefficients is the same as the method described in the first embodiment.

[0249] In the present embodiment, the functions of the voice input unit 29E, the directivity control unit 30E, the filter unit F6, the control unit 28E, and the addition unit 27E are implemented by a processor executing a program held in a memory. Alternatively, the voice input unit 29E, the directivity control unit 30E, the filter unit F6, the control unit 28E, and the addition unit 27E may be constituted by different hardware.

[0250] The voice processing device 21E has been described. However, the voice processing devices 22E, 23E, and 24E have almost the same structure except for the filter unit. The voice processing device 22E targets the voice spoken by the occupant hm2. The voice processing device 22E outputs, as an output signal, a voice signal obtained by suppressing the crosstalk component from the voice signal picked up by the microphone MC2. Therefore, the difference between the voice processing device 22E and the voice processing device 21E is that it has a filter unit to which the first directivity signal, the voice signal C, and the voice signal D are input. The same applies to the voice processing devices 23E and 24E.

[0251] Figure 21It is a flowchart showing the operation process of the voice processing device 21E. First, voice signals A, B, C, and D are input to the voice input unit 29E (S501). Next, the directivity control unit 30E performs directivity control processing using voice signals A and B to generate a first directivity signal and a second directivity signal (S502). Then, the directivity control unit 30E determines whether a voice component is input to the microphone MC3 or the microphone MC4 by the same method as in the first embodiment (S503). The directivity control unit 30E outputs the determination result as a flag to the control unit 28E. When the directivity control unit 30E determines that no voice signal is input to the microphone MC3 or the microphone MC4 (S503: "No"), the control unit 28E sets the intensities of the voice signals C and D input to the filter unit F6 to zero, and does not change the intensity of the second directivity signal. Then, the filter unit F6 generates a subtraction signal as follows (S504). The adaptive filter F6A passes the second directivity signal and outputs the passed signal P6A. The adaptive filter F6B passes the voice signals C and D and outputs the passed signal P6B. The filter unit F6 adds the passed signal P6A and the passed signal P6B and outputs the result as the subtraction signal. The addition unit 27E subtracts the subtraction signal from the first directivity signal, thereby generating and outputting an output signal (S505). The output signal is input to the control unit 28E and output from the control unit 28E. Next, the control unit 28E updates the filter coefficients of the adaptive filter F6A based on the output signal so that the target component included in the output signal becomes maximum (S506). Then, the voice processing device 21E performs step S501 again.

[0252] When the directivity control unit 30E determines in step S503 that a voice signal has been input to the microphone MC3 or the microphone MC4 (S503: "Yes"), the control unit 28E controls the filter unit F6 so that the voice signal C and the voice signal D are input to the adaptive filter F6B while maintaining their intensities unchanged. In other words, the control unit 28E does not change the intensity of the second directivity signal input to the adaptive filter F6A, nor the intensities of the voice signal C and the voice signal D input to the adaptive filter F6B. The filter unit F6 generates a subtraction signal obtained by adding the passed signal P6A and the passed signal P6B, and outputs it to the addition unit 27E (S507). The addition unit 27E subtracts the subtraction signal from the first directivity signal to generate an output signal and outputs it to the control unit 28E (S508). The control unit 28E updates the filter coefficients of the adaptive filter to which the voice signal is input so that the target component included in the output signal becomes maximum (S509). Specifically, the filter coefficients of the adaptive filter F6A and the adaptive filter F6B are updated. Then, the voice processing device 21E performs step S501 again.

[0253] In the present embodiment, the filter coefficients of the adaptive filter to which a voice signal in a state with zero intensity is input are not updated. Thus, compared with the case where the filter coefficients of all adaptive filters are always updated, the processing amount of the control unit 28E can be reduced. On the other hand, the control unit 28E may always update the filter coefficients of all adaptive filters. By always updating the filter coefficients of all adaptive filters, the control unit 28E can always perform the same processing, so the processing becomes simple. In addition, by always updating the filter coefficients of all adaptive filters, for example, for a certain adaptive filter, even immediately after changing from a state where a voice signal with zero intensity is input to a state where a voice signal with non-zero intensity is input, the filter coefficients can be updated with high precision.

[0254] Thus, also in the voice processing system 5E of the sixth embodiment, multiple voice signals are acquired by multiple microphones. Using another voice signal as a reference signal, a subtraction signal generated using an adaptive filter is subtracted from a certain voice signal, thereby accurately obtaining the voice of a specific speaker. In the sixth embodiment, a signal obtained by adding multiple voice signals is used as the reference signal. Thereby, voice signals can be independently picked up at each seat. At the same time, compared with the case where all signals obtained at each seat are used as the reference signal, the amount of processing for canceling crosstalk components can be reduced. Specifically, the voice processing system 5E uses the microphone MC3 and the microphone MC4 to independently pick up the voices of the occupant hm3 and the occupant hm4 in the rear seats. On this basis, the voice processing system 5E inputs both the voice signal C and the voice signal D into the adaptive filter F6B to be used as the reference signal. In addition, in the sixth embodiment, the process of determining which occupant's voice is included in the voice signal is not performed. Therefore, the amount of processing for canceling crosstalk components can be reduced. In addition, for an adaptive filter of a voice signal in a state where the input intensity is zero, the update of the filter coefficients may not be performed. Thereby, compared with the case where the filter coefficients are always updated for all adaptive filters, the amount of processing can be further reduced.

[0255] Item 1 (Fourth Embodiment)

[0256] A voice processing system, comprising:

[0257] A first microphone that acquires a first voice signal and outputs a first signal based on the first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position;

[0258] An adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; and

[0259] A control unit that controls filter coefficients of the adaptive filter,

[0260] Whether the first voice signal includes the first voice component or the first voice signal includes the second voice component, the first signal is input to the adaptive filter.

[0261] Item 2 (Fifth Embodiment)

[0262] A voice processing system, comprising:

[0263] A first microphone that acquires a first voice signal and outputs a first signal based on the first voice signal, where the first voice signal includes at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position;

[0264] A second microphone located at a position farther from the first position than the first microphone, the second microphone acquiring a second voice signal and outputting a second signal based on the second voice signal, where the second voice signal includes at least one of the first voice component and the second voice component;

[0265] A third microphone located at a position farther from the second position than the first microphone, the third microphone acquiring a third voice signal and outputting a third signal based on the third voice signal, where the third voice signal includes at least one of the first voice component and the second voice component;

[0266] Two or more adaptive filters that are input with the first signal and output a passed signal based on the first signal;

[0267] A control unit that controls filter coefficients of the two or more adaptive filters; and

[0268] An addition unit that subtracts a subtracted signal based on the passed signal from the second signal or the third signal,

[0269] The two or more adaptive filters include a first adaptive filter and a second adaptive filter,

[0270] The first adaptive filter is input with the first signal and outputs a first passed signal based on the first signal,

[0271] The second adaptive filter is input with the first signal and outputs a second passed signal based on the first signal,

[0272] The addition unit outputs a first output signal obtained by subtracting a first subtracted signal based on the first passed signal from the second signal or the third signal, and a second output signal obtained by subtracting a second subtracted signal based on the second passed signal from the second signal or the third signal,

[0273] The control unit determines which of the first adaptive filter and the second adaptive filter to use when generating the subtracted signal based on the first output signal and the second output signal.

[0274] Item 3

[0275] The voice processing system according to Item 2, wherein

[0276] When the first voice signal includes the first voice component, the first signal is input to the first adaptive filter.

[0277] When the first voice signal includes the second voice component, the first signal is input to the second adaptive filter.

[0278] Item 4

[0279] The voice processing system according to Item 3, wherein

[0280] The two or more adaptive filters include a third adaptive filter.

[0281] When the first voice signal includes the first voice component and the second voice component, the first signal is input to the third adaptive filter.

[0282] Item 5 (Sixth Embodiment)

[0283] A voice processing system, comprising:

[0284] A first microphone that acquires a first voice signal and outputs a first signal based on the first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position;

[0285] A second microphone that is located at a position farther from the second position than the first microphone, the second microphone acquires a second voice signal and outputs a second signal based on the second voice signal, the second voice signal including at least one of the first voice component and the second voice component;

[0286] A third microphone that is located at a position farther from the first position than the first microphone, or at a position farther from the second position than the second microphone, the third microphone acquires a third voice signal and outputs a third signal based on the third voice signal, the third voice signal including at least one of the first voice component and the second voice component;

[0287] An adaptive filter that is input with the first signal and the second signal and outputs a passed signal based on the first signal and the second signal; and

[0288] An adder that subtracts a subtracted signal based on the passed signal from the third signal.

[0289] Item 6

[0290] The voice processing system according to Item 5, comprising:

[0291] A fourth microphone located at a position farther from the first microphone and the second microphone than the second position, the fourth microphone acquiring a fourth voice signal and outputting a fourth signal based on the fourth voice signal, the fourth voice signal including at least one of the first voice component and the second voice component; and

[0292] A directivity control unit that performs a directivity control process on the third signal to output a first directivity signal and performs a directivity control process on the fourth signal to output a second directivity signal,

[0293] The third microphone is located at a position farther from the first position than the first microphone.

[0294] Explanation of reference numerals

[0295] 5: Voice processing system; 10: Vehicle; 20, 21, 22, 23: Voice processing device; 27: Addition unit; 28: Control unit; 29: Voice input unit; 30: Directivity control unit; 31: Abnormality detection unit; F1: Filter unit; F1A, F1B, F1C: Adaptive filter; 40: Voice recognition engine; 50: Electronic device.

Claims

1. A voice processing system, comprising: At least one first microphone that acquires a first voice signal and outputs a first signal based on the first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; At least one adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; A determination unit that determines which of the first voice component and the second voice component is included more in the first voice signal; A control unit that controls the filter coefficients of the adaptive filter based on the result of the determination; A second microphone that is located at a position farther from the first position than at least one of the first microphones, the second microphone acquiring a second voice signal and outputting a second signal based on the second voice signal, the second voice signal including at least one of the first voice component and the second voice component; And A third microphone that is located at a position farther from the second position than at least one of the first microphones, the third microphone acquiring a third voice signal and outputting a third signal based on the third voice signal, the third voice signal including at least one of the first voice component and the second voice component, The determination unit determines which of the first voice component and the second voice component is included more in the first voice signal based on the second signal and the third signal.

2. The voice processing system according to claim 1, wherein A directivity control unit is provided that outputs a first directivity signal obtained by performing a directivity control process on the second signal and outputs a second directivity signal obtained by performing a directivity control process on the third signal.

3. The voice processing system according to claim 2, wherein The determination unit determines which of the first voice component and the second voice component is included more in the first voice signal based on the first directivity signal and the second directivity signal.

4. The voice processing system according to claim 2 or 3, wherein The directivity control unit includes the determination unit.

5. The voice processing system according to any one of claims 1 to 3, wherein The at least one first microphone includes: A fourth microphone that acquires a fourth voice signal and outputs a fourth signal based on the fourth voice signal, the fourth voice signal including at least one of the first voice component and the second voice component; And A fifth microphone that is located at a position closer to the second position than the fourth microphone, the fifth microphone acquiring a fifth voice signal and outputting a fifth signal based on the fifth voice signal, the fifth voice signal including at least one of the first voice component and the second voice component, The voice processing system includes an abnormality detection unit that detects whether there is an abnormality in the at least one first microphone and sends abnormality information related to the abnormality of the at least one first microphone to the control unit. The control unit controls the filter coefficients of the adaptive filter based on the abnormality information and the determination result.

6. The voice processing system according to claim 5, wherein when the determination unit detects an abnormality in the fourth microphone, the control unit sets the intensity of the fourth signal input to the adaptive filter to zero. when the determination unit detects an abnormality in the fifth microphone, the control unit sets the intensity of the fifth signal input to the adaptive filter to zero.

7. The voice processing system according to claim 5, wherein the abnormality detection unit includes the determination unit.

8. A voice processing system, comprising: at least one first microphone that acquires a first voice signal and outputs a first signal based on the first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; at least one adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; a determination unit that determines which of the first voice component and the second voice component is included more in the first voice signal; and a control unit that controls the filter coefficients of the adaptive filter based on the determination result, wherein the at least one first microphone includes: a fourth microphone that acquires a fourth voice signal and outputs a fourth signal based on the fourth voice signal, the fourth voice signal including at least one of the first voice component and the second voice component; and a fifth microphone that is located closer to the second position than the fourth microphone, the fifth microphone acquires a fifth voice signal and outputs a fifth signal based on the fifth voice signal, the fifth voice signal including at least one of the first voice component and the second voice component. The voice processing system includes an abnormality detection unit that detects whether there is an abnormality in the at least one first microphone and sends abnormality information related to the abnormality of the at least one first microphone to the control unit. The control unit controls the filter coefficients of the adaptive filter based on the abnormality information and the determination result.

9. The voice processing system according to claim 8, wherein when the determination unit detects an abnormality in the fourth microphone, the control unit sets the intensity of the fourth signal input to the adaptive filter to zero. when the determination unit detects an abnormality in the fifth microphone, the control unit sets the intensity of the fifth signal input to the adaptive filter to zero.

10. The voice processing system according to claim 8, wherein The abnormality detection unit includes the determination unit.

11. A voice processing apparatus, comprising: At least one first receiving unit that receives a first signal based on a first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; At least one adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; A determination unit that determines which of the first voice component and the second voice component the first voice signal contains more; A control unit that controls filter coefficients of the adaptive filter based on a result of the determination; A second receiving unit that is located at a position farther from the first position than at least one of the first receiving units, the second receiving unit acquiring a second voice signal and outputting a second signal based on the second voice signal, the second voice signal including at least one of the first voice component and the second voice component; And A third receiving unit that is located at a position farther from the second position than at least one of the first receiving units, the third receiving unit acquiring a third voice signal and outputting a third signal based on the third voice signal, the third voice signal including at least one of the first voice component and the second voice component, The determination unit determines which of the first voice component and the second voice component the first voice signal contains more based on the second signal and the third signal.

12. A voice processing apparatus, comprising: At least one first receiving unit that receives a first signal based on a first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; At least one adaptive filter that is input with the first signal and outputs a passed signal based on the first signal; A determination unit that determines which of the first voice component and the second voice component the first voice signal contains more; and A control unit that controls filter coefficients of the adaptive filter based on a result of the determination, wherein the at least one first receiving unit includes: A fourth receiving unit that acquires a fourth voice signal and outputs a fourth signal based on the fourth voice signal, the fourth voice signal including at least one of the first voice component and the second voice component; And A fifth receiving unit that is located at a position closer to the second position than the fourth receiving unit, the fifth receiving unit acquiring a fifth voice signal and outputting a fifth signal based on the fifth voice signal, the fifth voice signal including at least one of the first voice component and the second voice component, The voice processing device includes an abnormality detection unit that detects whether there is an abnormality in the at least one first receiving unit and sends abnormality information related to the abnormality of the at least one first receiving unit to the control unit. The control unit controls the filter coefficients of the adaptive filter based on the abnormality information and the result of the determination.

13. A voice processing method is executed by a voice processing device. The voice processing method includes the following steps: Receiving a first signal based on a first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; Inputting the first signal into at least one adaptive filter, and the at least one adaptive filter outputs a passed signal based on the first signal; Receiving a second voice signal at a position farther from the first position than the position where the first signal is received, and outputting a second signal based on the second voice signal, the second voice signal including at least one of the first voice component and the second voice component; Receiving a third voice signal at a position farther from the second position than the position where the first signal is received, and outputting a third signal based on the third voice signal, the third voice signal including at least one of the first voice component and the second voice component; Based on the second signal and the third signal, determining which of the first voice component and the second voice component the first voice signal contains more; and Based on the result of the determination, controlling the filter coefficients of the adaptive filter.

14. A voice processing method is executed by a voice processing device. The voice processing method includes the following steps: Receiving, by at least one receiving unit, a first signal based on a first voice signal, the first voice signal including at least one of a first voice component generated at a first position and a second voice component generated at a second position different from the first position; Inputting the first signal into at least one adaptive filter, and the at least one adaptive filter outputs a passed signal based on the first signal; Receiving, by the at least one receiving unit, a fourth voice signal and outputting a fourth signal based on the fourth voice signal, the fourth voice signal including at least one of the first voice component and the second voice component; Receiving, by the at least one receiving unit, a fifth voice signal at a position closer to the second position than the position where the fourth voice signal is received, and outputting a fifth signal based on the fifth voice signal, the fifth voice signal including at least one of the first voice component and the second voice component; Determining which of the first voice component and the second voice component the first voice signal contains more; Detecting whether there is an abnormality in the at least one receiving unit and outputting abnormality information related to the abnormality of the at least one receiving unit; and And Based on the abnormal information and the determined result, control the filter coefficients of the adaptive filter.

Citation Information

Patent Citations

  • JP1973089810A

  • Voice processing device

    US20170352349A1