Speaker position detection system and in-vehicle communication support system

The speaker position detection system uses a microphone array and filter settings to calculate the speaker's position within the vehicle, addressing the challenge of cost-effective detection and improving speech output naturalness.

JP7696677B2Active Publication Date: 2025-06-23ALPS ALPINE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021116455
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-14
Publication Date
2025-06-23
Estimated Expiration
2041-07-14

Smart Images

  • Figure 0007696677000001
    Figure 0007696677000001
  • Figure 0007696677000002
    Figure 0007696677000002
  • Figure 0007696677000003
    Figure 0007696677000003
Patent Text Reader

Abstract

To provide a "speaker location detection system and an in-vehicle communication support system" which detect a speaker's location with a low-cost configuration.SOLUTION: An input processing section provided for each microphone in each seat of a car generates a signal S() representing a component in a microphone output V() that is correlated with an output V() of the other two microphones aligned in the front-front, left, and right using a filter W(), and an uncorrelated / correlated level calculation section 15 calculates, from each output V() and each signal S(), an uncorrelated level U() of each microphone output V() of each microphone and the outputs V() of the other two microphones aligned in the front-front, left, and right of the output V() of each microphone. A sound source coordinate calculation section 16 estimates a speaker position X in the left-right direction from the uncorrelated level U() with the outputs V() of the microphones aligned to the left and right, and estimates the speaker position Y in the front-front direction from the uncorrelated level U() with the outputs V() of the microphones aligned to the front-front, for a microphone in which both of uncorrelated levels U() with the outputs V() of the other two microphones exceed thresholds.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to a technology for assisting communication by in-vehicle speech.

Background Art

[0002] As a technology for assisting communication by in-vehicle speech, there is known a technology in which the speech voice of a speaker is picked up by a microphone, and the speech voice with the gain adjusted so that the listener can clearly hear it is output from a speaker (for example, Patent Document 1).

[0003] Also known is a technology in which a target sound picked up by a microphone arranged at a first position is converted into a target sound that would be picked up when a microphone is arranged at a second position using a filter (for example, Patent Document 2).

[0004] Here, in this technology, in advance, a microphone is actually arranged at the second position, and a transfer function that minimizes the difference between the output of the microphone arranged at the first position and the output of the microphone arranged at the second position is obtained using an adaptive filter and set as the transfer function of the above filter.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] According to the technology for assisting communication by in-vehicle speech described above, even when the distance between users is close, such as when the speaker leans forward towards the listener and the listener can clearly hear the speech without the speaker outputting the speech voice, the speaker still outputs the speech voice. Also, in such a case, when the speaker outputs the speech voice, the listener will hear an unnatural speech voice in which the speech voice directly coming from the speaker overlaps with the speech voice coming through the speaker and delayed compared to the directly coming speech voice.

[0007] Therefore, it is conceivable to detect the position of the speaker and control the presence or absence of output of the speech voice from the speaker and the gain of the speech voice output from the speaker according to the position of the speaker. As a method for detecting the position of the speaker, a method of detecting the sound source position of the speech voice using a microphone array, a method of detecting the position of the speaker's mouth using a stereo camera, etc. can be considered, but any method requires the addition of a special configuration and causes a significant increase in cost.

[0008] Therefore, the present invention aims to detect the position of the speaker with a configuration that does not cause a relatively large increase in cost in an in-vehicle communication support system that picks up the user's speech voice with a microphone and outputs it from a speaker towards other users.

Means for Solving the Problem

[0009] To achieve the above object, the present invention provides a first seat seat, and the aforesaid a first seat seat and a second seat seat arranged side by side in the left-right direction, and the aforesaid a first seat seat and a the third seatIn a speaker position detection system mounted on an automobile having [specific components not provided in the original], a first microphone which is a microphone arranged near the first seat, a second microphone which is a microphone arranged near the second seat, a third microphone which is a microphone arranged near the third seat, a filter for the second microphone in which a transfer function for approximately extracting a component correlated with the output of the first microphone from the output of the second microphone is set, a filter for the third microphone in which a transfer function for approximately extracting a component correlated with the output of the first microphone from the output of the third microphone is set, an uncorrelated level calculation means for calculating a left-right uncorrelated level by subtracting the output of the filter for the second microphone from the output of the first microphone and calculating a front-back uncorrelated level by subtracting the output of the filter for the third microphone from the output of the first microphone, and a speaker position calculation means for calculating the left-right position in the vehicle interior of the automobile that matches the left-right position of the sound source from which the calculated left-right uncorrelated level value is obtained as the left-right position of the speaker and calculating the front-back position in the vehicle interior of the automobile that matches the front-back position of the sound source from which the calculated front-back uncorrelated level value is obtained as the front-back position of the speaker.

[0010] Here, the speaker position detection system may be provided with a left-right uncorrelated level map showing the correspondence between each value of the left-right uncorrelated level and the left-right position in the vehicle interior of the automobile that matches the left-right position of the sound source from which the value of the left-right uncorrelated level is obtained, and a front-back uncorrelated level map showing the correspondence between each value of the front-back uncorrelated level and the front-back position in the vehicle interior of the automobile that matches the front-back position of the sound source from which the value of the front-back uncorrelated level is obtained. The speaker position calculation means may calculate the left-right position in the vehicle interior of the automobile that matches the left-right position of the sound source from which the calculated left-right uncorrelated level value is obtained according to the correspondence shown by the left-right uncorrelated level map, and calculate the front-back position in the vehicle interior of the automobile that matches the front-back position of the sound source from which the calculated front-back uncorrelated level value is obtained according to the correspondence shown by the front-back uncorrelated level map.

[0011] In addition, in this speaker position detection system, the transfer function of the second microphone filter is set in advance by using, as an error, the difference between the output of a first adaptive filter that takes the output of the second microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, and causing the first adaptive filter to perform an adaptive operation. The transfer function of the first adaptive filter that has converged is set as the transfer function of the second microphone filter. The transfer function of the third microphone filter may be set in advance by using, as an error, the difference between the output of a second adaptive filter that takes the output of the third microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, and causing the second adaptive filter to perform an adaptive operation. The transfer function of the second adaptive filter that has converged is set as the transfer function of the third microphone filter.

[0012] Alternatively, this speaker position detection system may be provided with a first filter for the first microphone having a transfer function set to approximately extract a component correlated with the output of the second microphone from the output of the first microphone, a second filter for the first microphone having a transfer function set to approximately extract a component correlated with the output of the third microphone from the output of the first microphone, and a correlation level calculation means for calculating a left-right correlation level by adding the output of the first filter for the first microphone and the output of the second microphone filter, and calculating a front-back correlation level by adding the output of the second filter for the first microphone and the output of the third microphone filter. In the speaker position calculation means, the left-right position inside the vehicle of the automobile that matches the left-right direction position of the sound source from which the calculated left-right uncorrelated level value and the calculated left-right correlation level value are obtained is calculated as the left-right direction position of the speaker, and the front-back position inside the vehicle of the automobile that matches the front-back direction position of the sound source from which the calculated front-back uncorrelated level value and the calculated front-back correlation level value are obtained is calculated as the front-back direction position of the speaker.

[0013] In this case, the speaker position detection system includes a left-right uncorrelated level map showing the correspondence between each value of the left-right uncorrelated level and the left-right position in the vehicle interior of the automobile that matches the left-right position of the sound source from which the value of the left-right uncorrelated level is obtained, a front-rear uncorrelated level map showing the correspondence between each value of the front-rear uncorrelated level and the front-rear position in the vehicle interior of the automobile that matches the front-rear position of the sound source from which the value of the front-rear uncorrelated level is obtained, a left-right correlated level map showing the correspondence between each value of the left-right correlated level and the left-right position in the vehicle interior of the automobile that matches the left-right position of the sound source from which the value of the left-right correlated level is obtained, and a front-rear correlated level map showing the correspondence between each value of the front-rear correlated level and the front-rear position in the vehicle interior of the automobile that matches the front-rear position of the sound source from which the value of the front-rear correlated level is obtained. In the speaker position calculation means, according to the correspondence shown by the left-right uncorrelated level map and the correspondence shown by the left-right correlated level map, the left-right position in the vehicle interior of the automobile that matches the left-right position of the sound source from which the calculated value of the left-right uncorrelated level and the calculated value of the left-right correlated level are obtained is calculated, and according to the correspondence shown by the front-rear uncorrelated level map and the correspondence shown by the front-rear correlated level map, the front-rear position in the vehicle interior of the automobile that matches the front-rear position of the sound source from which the calculated value of the front-rear uncorrelated level and the calculated value of the front-rear correlated level are obtained may be calculated.

[0014] Also, in this case, for the speaker position detection system, the transfer function of the filter for the second microphone is set in advance by using, as an error, the difference between the output of the first adaptive filter that takes the output of the second microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, and causing the first adaptive filter to perform an adaptation operation. The transfer function of the converged first adaptive filter is set as the transfer function of the filter for the second microphone. The transfer function of the filter for the third microphone is set in advance by using, as an error, the difference between the output of the second adaptive filter that takes the output of the third microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, and causing the second adaptive filter to perform an adaptation operation. The transfer function of the converged second adaptive filter is set as the transfer function of the filter for the third microphone. The transfer function of the first filter for the first microphone is set in advance by using, as an error, the difference between the output of the third adaptive filter that takes the output of the first microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the second microphone, and causing the third adaptive filter to perform an adaptation operation. The transfer function of the converged third adaptive filter is set as the transfer function of the first filter for the first microphone. The transfer function of the second filter for the first microphone may be set in advance by using, as an error, the difference between the output of the fourth adaptive filter that takes the output of the first microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the third microphone, and causing the fourth adaptive filter to perform an adaptation operation. The transfer function of the converged fourth adaptive filter is set as the transfer function of the second filter for the first microphone.

[0015] Here, in the above speaker position detection system, in the speaker position calculation means, when both the magnitude of the left-right decorrelation level and the magnitude of the front-back decorrelation level are greater than a predetermined threshold value, it may be configured to determine that there is a speaker in the first size and seat and calculate the position of the speaker in the left-right direction and the position in the front-back direction. seat

[0016] The present invention also provides an in-vehicle communication support system including the above-mentioned speaker position detection system. This in-vehicle communication support system includes a speaker and an output processing means for processing to output the sound picked up by the first microphone to the speaker, and the output processing means switches at least one of whether or not to output the sound picked up by the first microphone to the speaker and the gain of the sound picked up by the first microphone to be output to the speaker, depending on the speaker position represented by the left-right position and the front-rear position of the speaker calculated by the speaker position calculation means.

[0017] According to the above-described speaker position detection system, the in-vehicle communication support system can detect the speaker position without adding any special configuration for detecting the speaker position. seat The speaker's position can be detected using the output of a microphone provided for picking up the speech of the person at the seat. Effect of the Invention

[0018] As described above, according to the present invention, in an in-car communication support system that picks up a user's speech with a microphone and outputs it from a speaker to other users, the position of the speaker can be detected with a configuration that does not incur a relatively high cost. [Brief description of the drawings]

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0020] Hereinafter, embodiments of the present invention will be described. FIG. 1 shows the configuration of an in-vehicle communication support system according to the present embodiment. The in-vehicle communication support system is a system mounted on an automobile, and as shown in the figure, it includes four microphones, namely, microphone MFR, microphone MFL, microphone MBR, and microphone MBL, and four speakers, namely, speaker SPFR, speaker SPFL, speaker SPBR, and speaker SPBL. Further, the in-vehicle communication support system includes a speaker position calculation unit 1, a control unit 2, and an output processing unit 3.

[0021] Here, as shown in FIG. 2, the microphone MFR is arranged to pick up the speech voice of a user sitting in the right front seat on the right side of the right front seat of the automobile, the microphone MFL is arranged to pick up the speech voice of a user sitting in the left front seat on the left side of the left front seat of the automobile, the microphone MBR is arranged to pick up the speech voice of a user sitting in the right rear seat on the right side of the right rear seat of the automobile, and the microphone MBL is arranged to pick up the speech voice of a user sitting in the left rear seat on the left side of the left rear seat of the automobile.

[0022] Also, the speaker SPFR is arranged to output sound toward the user sitting in the right front seat on the right side of the front right seat of the vehicle, the speaker SPFL is arranged to output sound toward the user sitting in the left front seat on the left side of the front left seat of the vehicle, the speaker SPBR is arranged to output sound toward the user sitting in the right rear seat on the right side of the rear right seat of the vehicle, and the speaker SPBL is arranged to output sound toward the user sitting in the left rear seat on the left side of the rear left seat of the vehicle.

[0023] Returning to FIG. 1, the speaker position calculation unit 1 calculates the speaker position (X, Y) from the outputs of the respective microphones and outputs it to the control unit 2. The output processing unit 3, according to the control unit 2, adjusts the gain for each speaker of the output of the audio signal input from each microphone to each speaker, and controls the presence or absence of the output of the audio signal input from each microphone to each speaker. The control unit 2 controls the input / output processing unit based on the speaker position (X, Y) calculated by the speaker position calculation unit 1. In this control, the audio signal input from the microphone that picks up the speech of the user in the seat where there is no speaker is controlled so that the output processing unit 3 does not output it to any speaker. Also, the audio signal input from the microphone that picks up the speech of the user in the seat where the speaker is located is not output to the speaker that outputs sound toward the user in the seat where the speaker is located and the speaker that outputs sound toward the user in the seat within a predetermined distance from the speaker position (X, Y), and the output processing unit 3 is controlled so that only the remaining speakers output it to the speakers that output sound toward the users in the seats far from the speaker position (X, Y) with a larger gain.

[0024] Next, the configuration of the speaker position calculation unit 1 is shown in FIG. 3. As shown in the figure, the speaker position calculation unit 1 includes four input processing units corresponding to the four microphones, namely, an MFR input processing unit 11 that processes the output V(MFR) of the microphone MFR, an MFL input processing unit 12 that processes the signal V(MFL) output from the microphone MFL, an MBR input processing unit 13 that processes the signal V(MBR) output from the microphone MBR, and an MBL input processing unit 14 that processes the signal V(MBL) output from the microphone MBL.

[0025] In addition, the speaker position calculation unit 1 includes an uncorrelated / correlation level calculation unit 15, a sound source coordinate calculation unit 16, and an uncorrelated / correlation level map storage unit 17. In addition, the MFR input processing unit 11 includes a filter W(MFR_MFL) 111 and a filter W(MFR_MBR) 112 that take the output V(MFR) of the microphone MFR as input, and outputs V(MFR), S(MFR_MFL) which is the output of the filter W(MFR_MFL), and S(MFR_MBR) which is the output of the filter W(MFR_MBR) to the uncorrelated / correlation level calculation unit 15. Also, the MFL input processing unit 12 includes a filter W(MFL_MFR) 121 and a filter W(MFL_MBL) 122 that take the output V(MFL) of the microphone MFL as input, and outputs V(MFL), S(MFL_MFR) which is the output of the filter W(MFL_MFR), and S(MFL_MBL) which is the output of the filter W(MFL_MBL) to the uncorrelated / correlation level calculation unit 15.

[0026] In addition, the MBR input processing unit 13 includes a filter W(MBR_MBL) 131 and a filter W(MBR_MFR) 132 that take the output V(MBR) of the microphone MBR as input, and outputs V(MBR), S(MBR_MBL) which is the output of the filter W(MBR_MBL), and S(MBR_MFR) which is the output of the filter W(MBR_MFR) to the uncorrelated / correlation level calculation unit 15. Also, the MBL input processing unit 14 includes a filter W(MBL_MBR) 141 and a filter W(MBL_MFL) 142 that take the output V(MBL) of the microphone MBL as input, and outputs V(MBL), S(MBL_MBR) which is the output of the filter W(MBL_MBR), and S(MBL_MFL) which is the output of the filter W(MBL_MFL) to the uncorrelated / correlation level calculation unit 15.

[0027] Here, the filter W(MA_MB) approximately extracts a component in the output V(MA) of the microphone MA that correlates with the output V(MB) of the microphone MB, and outputs it as S(MA_MB). That is, for example, the filter W(MFR_MFL) 111 of the MFR input processing unit 11 approximately extracts a component in the output V(MFR) of the microphone MFR that correlates with the output V(MFL) of the microphone MFL, and outputs it as S(MFR_MFL), and the filter W(MFR_MBR) 112 of the MFR input processing unit 11 approximately extracts a component in the output V(MFR) of the microphone MFR that correlates with the output V(MBR) of the microphone MBR, and outputs it as the signal S(MFR_MBR).

[0028] Here, the transfer function of each filter W() of each input processing unit is calculated and set in advance through tuning. The tuning for setting the transfer functions of the filter W(MA_MB) and the filter W(MB_MA) is performed according to the configuration shown in FIG. 4a. As shown in the figure, this configuration includes a tuning speaker TSPR, a microphone MA, a microphone MB, a first adaptive filter 41, a second adaptive filter 42, a first adder 43, and a second adder 44.

[0029] Also, the first adaptive filter 41 includes a first variable filter 411 and a first adaptive algorithm execution unit 412 that updates the transfer function (filter coefficient) of the second variable filter by an adaptive algorithm such as LMS, and the second adaptive filter 42 includes a second variable filter 421 and a third adaptive algorithm execution unit 422 that updates the transfer function (filter coefficient) of the second variable filter 421 by an adaptive algorithm such as LMS.

[0030] Here, the speaker TSPR is centrally arranged in the passenger compartment of the automobile, and the tuning is performed while outputting a predetermined tuning voice from the speaker TSPR. The sound collected by microphone MA is output to the first adder 43 through the first variable filter 411 of the first adaptive filter 41, and the sound collected by microphone MB is output to the second adder 44 through the second variable filter 421 of the second adaptive filter 42.

[0031] The first adder 43 subtracts the output of the first variable filter 411 from the output of microphone MB, outputs it as error e1 to the first adaptive algorithm execution unit 412 of the first adaptive filter 41.

[0032] The first adaptive algorithm execution unit 412 executes an adaptive algorithm such as LMS, and updates the transfer function of the first variable filter 411 so that the error e1 is minimized. The second adder 44 subtracts the output of the second variable filter 421 from the output of microphone MA, outputs it as error e2 to the second adaptive algorithm execution unit of the second adaptive filter 42. The second adaptive algorithm execution unit executes an adaptive algorithm such as LMS, and updates the transfer function of the second variable filter 421 so that the error e2 is minimized. And, if the transfer functions of the first variable filter 411 and the first variable filter 411 converge by the above operations, the converged transfer function of the first variable filter 411 is set as the transfer function of filter W(MA_MB), and the converged transfer function of the second variable filter 421 is set as the transfer function of filter W(MB_MA).

[0033] Here, in a state where the transfer functions of the first variable filter 411 and the second variable filter 421 have converged, the transfer function of the first variable filter 411 extracts the component most correlated with the sound collected by microphone MB from the sound collected by microphone MA, and the transfer function of the first variable filter 411 extracts the component most correlated with the sound collected by microphone MA from the sound collected by microphone MB.

[0034] Note that, for actual tuning, each combination of (microphone MFR, microphone MFL), (microphone MFR, microphone MBR), (microphone MFL, microphone MBL), and (microphone MBR, microphone MBL) is tuned as the combination of (microphone MA, microphone MB) in the configuration of FIG. 4a.

[0035] Thereby, for example, as shown in FIG. 4b, by tuning the combination of (microphone MFR, microphone MFL) as the combination of (microphone MA, microphone MB) in the configuration of FIG. 4a, the transfer function of the filter W(MFR_MFL) 111 of the MFR input processing unit 11 is obtained as the transfer function of the converged first variable filter 411, and the transfer function of the filter W(MFL_MFR) 121 of the MFL input processing unit 12 is obtained as the transfer function of the converged second variable filter 421.

[0036] Next, the uncorrelated / correlated level calculation unit 15 calculates the left-right uncorrelated level U(MFR_MFL) and the front-back uncorrelated level U(MFR_MBR) of the microphone MFR, the left-right uncorrelated level U(MFL_MFR) and the front-back uncorrelated level U(MFL_MBL) of the microphone MFL, the left-right uncorrelated level U(MBR_MBL) and the front-back uncorrelated level U(MBR_MFR) of the microphone MBR, the left-right uncorrelated level U(MBL_MBR) and the front-back uncorrelated level U(MBL_MFL) of the microphone MBL, the left-right correlation level C(MFR / MFL), the left-right correlation level C(MBR / MBL), the front-back correlation level C(MFR / MBR), and the front-back correlation level C(MFL / MBL) from each signal input from the four input processing units, and outputs them to the sound source coordinate calculation unit 16.

[0037] The left-right uncorrelated level U(MA_MB) of microphone MA is calculated from V(MA)-S(MB_MA), and the front-back uncorrelated level U(MA_MC) of microphone MA is calculated from V(MA)-S(MC_MA). The left-right uncorrelated level U(MA_MB) of microphone MA represents the degree of uncorrelation of the output V(MA) of microphone MA with respect to the output V(MB) of microphone MB arranged in the left-right direction with respect to the microphone MA, and the front-back uncorrelated level U(MA_MC) of microphone MA represents the degree of uncorrelation of the output V(MA) of microphone MA with respect to the output V(MC) of microphone MC arranged in the front-back direction with respect to the microphone MA.

[0038] Therefore, for example, the left-right uncorrelated level U(MFR_MFL) of microphone MFR is calculated by V(MFR)-S(MFL_MFR), and represents the degree of uncorrelation of the output V(MFR) of microphone MFR with respect to the output V(MFL) of microphone MFL arranged in the left-right direction with respect to the microphone MFR. The front-back uncorrelated level U(MFR_MBR) of microphone MFR is calculated by V(MFR)-S(MBR_MFR), and represents the degree of uncorrelation of the output V(MFR) of microphone MFR with respect to the output V(MBR) of microphone MBR arranged in the front-back direction with respect to the microphone MFR.

[0039] Next, the left-right correlation level C(MA / MB) is calculated by {S(MA_MB)+S(MB_MA)} / 2, and the front-back correlation level C(MA / MC) is calculated by {S(MA_MC)+S(MC_MA)} / 2. The left-right correlation level C(MA / MB) represents the degree of correlation between the output V(MA) of microphone MA arranged side by side and the output V(MB) of microphone MB, and the front-back correlation level C(MA / MC) represents the degree of correlation between the output V(MA) of microphone MA arranged front and back and the output V(MC) of microphone MC.

[0040] Therefore, for example, the degree of correlation C(MFR / MFR) representing the degree of correlation between V(MFR) and V(MFL), which are the outputs of the left and right microphones MFR and MFL, is calculated by {S(MFR_MFL) + S(MFL_MFR)} / 2, and the degree of correlation C(MFR / MBR) representing the degree of correlation between V(MFR) and V(MBR), which are the outputs of the front and rear microphones MFR and MBR, is calculated by {S(MFR_MBR) + S(MBR_MFR)} / 2.

[0041] Next, the sound source position calculation unit detects the presence or absence of a speaker from the signal output by the uncorrelated / correlated level calculation unit 15 and calculates the speaker position (X, Y), and outputs the calculated speaker position (X, Y) to the control unit 2.

[0042] The presence or absence of a speaker is detected by determining that there is no speaker if there is no microphone MA for which both the left-right uncorrelated level U() and the front-rear uncorrelated level U() are equal to or greater than a predetermined threshold, and determining that there is a speaker at the seat where the microphone MA picks up the speech if such a microphone MA exists.

[0043] Therefore, for example, when both the left-right uncorrelated level U(MFR_MFL) and the front-rear uncorrelated level U(MFR_MBR) of the microphone MFR are equal to or greater than the threshold value, a speaker in the right front seat where the microphone MFR picks up the speech is detected.

[0044] Next, the calculation of the speaker position (X, Y) in the sound source position calculation unit is performed using the uncorrelated level map and the correlation level map stored in advance in the uncorrelated / correlated level map storage unit 17. As the uncorrelated level maps, eight uncorrelated level maps corresponding to the left-right uncorrelated level U(MFR_MFL) of the microphone MFR, the front-back uncorrelated level U(MFR_MBR) of the microphone MFR, the left-right uncorrelated level U(MFL_MFR) of the microphone MFL, the front-back uncorrelated level U(MFL_MBL) of the microphone MFL, the left-right uncorrelated level U(MBR_MBL) of the microphone MBR, the front-back uncorrelated level U(MBR_MFR) of the microphone MBR, the left-right uncorrelated level U(MBL_MBR) of the microphone MBL, and the front-back uncorrelated level U(MBL_MFL) of the microphone MBL are stored.

[0045] Each uncorrelated level map shows the relationship between each position in the vehicle interior and the value of the corresponding left-right / front-back uncorrelated level U() for the sound with that position as the sound source position. Each uncorrelated level map is created in advance by simulation, actual measurement, etc., and is stored in the uncorrelated / correlated level map storage unit 17.

[0046] Also, as the correlated level maps, four correlated level maps corresponding to each of the left-right correlated levels C(MFR / MFL), left-right correlated levels C(MBR / MBL), front-back correlated levels C(MFR / MBR), and front-back correlated levels C(MFL / MBL) are stored.

[0047] correlation level map shows the relationship between each position in the vehicle interior and the value of the corresponding left-right / front-back correlated level C() for the sound with that position as the sound source position, and each correlated level map is created in advance by simulation, actual measurement, etc., and is stored in the uncorrelated / correlated level map storage unit 17. is stored.

[0048] Here, an example of such uncorrelated level maps and correlated level maps is shown in FIG. 6. For example, FIG. 6a1 is an example of an uncorrelated level map corresponding to the left-right uncorrelated level U(MFR_MFL) of the microphone MFR, and FIG. 6a2 is an example of the front-back uncorrelated level U(MFR_MBR) of the microphone MFR. Also, FIG. 6b1 is an example of a correlation level map corresponding to the left-right correlation level C(MFR / MFL), and FIG. 6b2 is an example of the front-back correlation level C(MFR / MBR).

[0049] The direction from left to right in each figure represents the direction from left to right of the automobile, and the direction from top to bottom represents the direction from front to back of the automobile. Also, in each figure, PMFR represents the position of the microphone MFR, PMFL represents the position of the microphone MFL, PMBR represents the position of the microphone MBR, and PMBL represents the position of the microphone MBL.

[0050] Also, the coordinates in each figure represent the distance between the left and right microphones and the distance between the front and back microphones as 1, indicating that the darker the color, the smaller the level value, and the closer to white, the larger the level value.

[0051] The calculation of the speaker position (X, Y) using such an uncorrelated level map and a correlation level map is performed as follows. That is, a microphone for which both the left-right uncorrelated level U() and the front-back uncorrelated level U() are equal to or greater than a threshold value is defined as microphone MA, and a microphone arranged side by side with microphone MA on the left and right is defined as microphone MB. Among the ranges of the coordinates in the left-right direction on the line connecting the position PMA of microphone MA and the position PMB of microphone MB in the uncorrelated level map corresponding to the left-right uncorrelated level U(MA_MB), the range that is considered to be the most appropriate as the range of the speaker position where the value is the same as the left-right uncorrelated level U(MA_MB) is calculated as the range BX of the coordinates in the left-right direction of the speaker position. The validity as the range of the speaker position is determined such that the closer the range is to the standard speaking position of a user sitting on the seat where the speech sound is picked up by microphone MA, the higher the validity.

[0052] Also, with the microphone MA and the microphones arranged in front of and behind it designated as the microphone MC, among the uncorrelated level maps corresponding to the front-back uncorrelated level U(MA_MC), the range of the speaker position within the front-back direction coordinate range where the value is the same as the front-back uncorrelated level U(MA_MC) on the line connecting the position PMA of the microphone MA and the position PMC of the microphone MC is calculated as the left-right direction coordinate range BY of the speaker position, which is considered to be the most appropriate range.

[0053] Next, the coordinates in the left-right direction within the coordinate range BX on the line connecting the position PMA of the microphone MA and the position PMB of the microphone MB on the left-right correlation level C(MA / MB), where the value is the same as the left-right correlation level C(MA / MB), are calculated as the left-right direction coordinates X of the speaker position, and the coordinates in the front-back direction within the coordinate range BY on the line connecting the position PMA of the microphone MA and the position PMC of the microphone MC on the front-back correlation level C(MA / MC), where the value is the same as the front-back correlation level C(MA / MC), are calculated as the front-back direction coordinates Y of the speaker position. When there is no left-right correlation level C(MA / MB), the left-right correlation level C(MB / MA) is used instead. When there is no front-back correlation level C(MA / MC), the front-back correlation level C(MC / MA) is used instead.

[0054] Then, the set of calculated coordinates (X, Y) is taken as the speaker position. Therefore, for example, when the microphone for which both the left-right uncorrelated level U() and the front-back uncorrelated level U() are equal to or greater than the threshold value is the microphone MFR, since the microphones arranged on the left and right of the microphone MFR are the microphones MFL, among the uncorrelated level maps corresponding to the left-right uncorrelated level U(MFR_MFL) in FIG. 7a1, the range of the speaker position within the left-right direction (x direction in the figure) coordinate range of the position where the value is the same as the left-right uncorrelated level U(MFR_MFL) on the line RL connecting the position PMFR of the microphone MFR and the position PMFL of the microphone MFL is calculated as the left-right direction coordinate BX of the speaker position, which is considered to be the most appropriate range.

[0055] Note that Th in the figure indicates the position of the same value as the threshold value, and in FIG. 7a1, the area above Th is the area of values equal to or greater than the threshold value. Also, since the microphones arranged before and after the microphone MFR are the microphones MBR, the range of the speaker position in the coordinate range in the front-back direction (y direction in the figure) of the position where the value is the same as the front-back uncorrelated level U(MFR_MBR) on the uncorrelated level map corresponding to the front-back uncorrelated level U(MFR_MBR) between the position PMFR of the microphone MFR and the position PMBR of the microphone MBR on the line FB connecting them is calculated as the front-back coordinate range BY of the speaker position as the most appropriate range.

[0056] Note that in Fig. 7a2, the area on the right side of Th is the area with values above the threshold. Then, the horizontal coordinates within the coordinate range BX on the line RL connecting the position PMFR of the microphone MFR and the position PMFL of the microphone MFL on the correlation level map corresponding to the left-right correlation level C(MFR / MFL) shown in Fig. 7b1, where the value is the same as the left-right correlation level C(MFR / MFL), are calculated as the horizontal coordinates X of the speaker position, and the vertical coordinates within the coordinate range BY on the line FB connecting the position PMFR of the microphone MFR and the position PMBR of the microphone MBR on the correlation level map corresponding to the front-back correlation level C(MFR / MBR) shown in Fig. 7b2, where the value is the same as the front-back correlation level C(MFR / MBR), are calculated as the vertical coordinates Y of the speaker position.

[0057] Then, the set of calculated coordinates (X, Y) is taken as the speaker position. As described above, the calculation of the speaker position (X, Y) using the uncorrelated level map and the correlation level map has been explained. However, in cases where the speaker position (X, Y) can be calculated with sufficient resolution using only the uncorrelated level map, instead of the above process of calculating the speaker position (X, Y) using the uncorrelated level map and the correlation level map, a process of calculating the speaker position (X, Y) using only the uncorrelated level map may be performed.

[0058] That is, in this case, a microphone for which both the left - right uncorrelated level U() and the front - back uncorrelated level U() are equal to or greater than the threshold value is defined as microphone MA, and a microphone arranged side - by - side with microphone MA on the left and right is defined as microphone MB. Among the coordinates in the left - right direction of the position on the line connecting the position PMA of microphone MA and the position PMB of microphone MB in the uncorrelated level map corresponding to the left - right uncorrelated level U(MA_MB), the coordinate that is considered to be the most appropriate as the coordinate of the speaker position is calculated as the left - right direction coordinate X of the speaker position. The validity as the coordinate of the speaker position is determined such that the closer the coordinate is to the standard speaking position of the user sitting on the seat where the voice of speech is picked up by microphone MA, the higher the validity.

[0059] Also, a microphone arranged in front of and behind microphone MA is defined as microphone MC. Among the coordinates in the front - back direction of the position on the line connecting the position PMA of microphone MA and the position PMC of microphone MC in the uncorrelated level map corresponding to the front - back uncorrelated level U(MA_MC), the coordinate that is considered to be the most appropriate as the coordinate of the speaker position is calculated as the front - back direction coordinate Y of the speaker position.

[0060] Then, the set of calculated coordinates (X, Y) is taken as the speaker position. Therefore, for example, when the microphone for which both the left - right uncorrelated level U() and the front - back uncorrelated level U() are equal to or greater than the threshold value is microphone MFR, since the microphones arranged side - by - side with microphone MFR on the left and right are microphone MFL, among the coordinates in the left - right direction (x - direction in the figure) of the position on the line RL connecting the position PMFR of microphone MFR and the position PMFL of microphone MFL in the uncorrelated level map corresponding to the left - right uncorrelated level U(MFR_MFL), the coordinate that is considered to be the most appropriate as the coordinate of the speaker position is calculated as the left - right direction coordinate X of the speaker position.

[0061] Also, since the microphones arranged before and after the microphone MFR are the microphones MBR, among the uncorrelated level maps corresponding to the front-back uncorrelated level U(MFR_MBR) in FIG. 8a2, the coordinates in the front-back direction (the y direction in the figure) of the position where the value is the same as the front-back uncorrelated level U(MFR_MBR) on the line FB connecting the position PMFR of the microphone MFR and the position PMBR of the microphone MBR are calculated as the coordinates of the speaker position, which are considered to be the most appropriate, as the coordinates Y in the front-back direction of the speaker position.

[0062] Then, the calculated coordinate pair (X, Y) is taken as the speaker position. The embodiments of the present invention have been described above. As described above, according to the present embodiment, the in-vehicle communication support system can detect the speaker position using the outputs of the microphones MFR, MFL, MBR, and MBL provided for picking up the voice of the speaker at each seat without adding a special configuration for detecting the speaker position. Also, since the transfer functions of the filters of the speaker position calculation unit 1 are fixed, an increase in the processing load is also suppressed.

[0063] Here, the above embodiment may be configured to use the technology of converting the target sound picked up by the microphone arranged at the first position into the target sound that would be picked up if the microphone were arranged at the second position using the above-described filter, so as to further improve the accuracy of speaker position calculation.

[0064] In this case, for example, as shown in FIG. 9, virtual microphones VM1-VM6 are set at positions on the left side of the right front seat, on the right side of the left front seat, on the left side of the right rear seat, on the right side of the left rear seat, on the right side of the right seat and on the left side of the left seat between the front seat and the rear seat in the front-back direction. Sounds that would be picked up if the virtual microphones VM1-VM6 were real microphones are generated from the outputs of the microphones MFR, MFL, MBR, and MBL, and the speaker position is calculated using the uncorrelated levels and correlation levels of the outputs of the microphones MFR, MFL, MBR, MBL, and the virtual microphones VM1-VM6.

[0065] In addition, when applying the above-described embodiments to an automobile equipped with a seat surface sensor that detects the presence or absence of a user's seating at each seat, or to an automobile equipped with a camera that captures the state of the user at each seat, the outputs of these seat surface sensors and cameras may be used in combination to calculate the presence or absence of a speaker and the speaker position.

Explanation of Reference Numerals

[0066] 1... Speaker position calculation unit, 2... Control unit, 3... Output processing unit, 11... MFR input processing unit, 12... MFL input processing unit, 13... MBR input processing unit, 14... MBL input processing unit, 15... Uncorrelated / correlation level calculation unit, 16... Sound source coordinate calculation unit, 17... Uncorrelated / correlation level map storage unit, 41... First adaptive filter, 42... Second adaptive filter, 43... First adder, 44... Second adder, 111... Filter W (MFR_MFL), 112... Filter W (MFR_MBR), 121... Filter W (MFL_MFR), 122... Filter W (MFL_MBL), 131... Filter W (MBR_MBL), 132... Filter W (MBR_MFR), 141... Filter W (MBL_MBR), 142... Filter W (MBL_MFL), 411... First variable filter, 412... First adaptive algorithm execution unit, 421... Second variable filter, 422... Third adaptive algorithm execution unit, MA... Microphone, MB... Microphone, MFR... Microphone, MFL... Microphone, MBR... Microphone, MBL... Microphone, SPFR... Speaker, SPFL... Speaker, SPBR... Speaker, SPBL... Speaker, TSPR... Tuning speaker.

Claims

1. A speaker position detection system mounted on a vehicle having a first seat, a second seat arranged side by side with the first seat in the left-right direction, and a third seat arranged in the front-rear direction with the first seat, a first microphone which is a microphone arranged near the first seat, a second microphone which is a microphone arranged near the second seat, a third microphone which is a microphone arranged near the third seat, a second microphone filter in which a transfer function for approximately extracting a component correlated with the output of the first microphone from the output of the second microphone is set, a third microphone filter in which a transfer function for approximately extracting a component correlated with the output of the first microphone from the output of the third microphone is set, an uncorrelated level calculation means for calculating, as a left-right uncorrelated level, a magnitude obtained by subtracting the magnitude of the output of the second microphone filter from the magnitude of the output of the first microphone, and calculating, as a front-rear uncorrelated level, a magnitude obtained by subtracting the magnitude of the output of the third microphone filter from the magnitude of the output of the first microphone, a left-right uncorrelated level map showing the correspondence between each value of the left-right uncorrelated level and the left-right position in the vehicle interior of the sound source from which the value of the left-right uncorrelated level is obtained, a front-rear uncorrelated level map showing the correspondence between each value of the front-rear uncorrelated level and the front-rear position in the vehicle interior of the sound source from which the value of the front-rear uncorrelated level is obtained, and a speaker position calculation means, wherein the speaker position calculation means, among the left-right positions of the sound source from which the calculated value of the left-right uncorrelated level indicated by the left-right uncorrelated level map on the line connecting the first microphone and the second microphone is obtained, calculates, as the left-right position of the speaker who is the user sitting on the first seat, a position closer to the standard utterance position of the user sitting on the first seat, A speaker position detection system, characterized in that a position closer to a standard speaking position of a user sitting in the first seat is calculated as the front-back direction position of the speaker among the front-back direction positions of sound sources at which values of the calculated front-back uncorrelated level indicated by the front-back uncorrelated level map on a line connecting the first microphone and the third microphone are obtained. **Claim 2** A speaker position detection system mounted on an automobile having a first seat, a second seat arranged side by side with the first seat in the left-right direction, and a third seat arranged in the front-back direction with the first seat, comprising: A first microphone, which is a microphone arranged near the first seat; A second microphone, which is a microphone arranged near the second seat; A third microphone, which is a microphone arranged near the third seat; A second microphone filter in which a transfer function for approximately extracting a component correlated with the output of the first microphone from the output of the second microphone is set; A third microphone filter in which a transfer function for approximately extracting a component correlated with the output of the first microphone from the output of the third microphone is set; Uncorrelated level calculation means for calculating, as a left-right uncorrelated level, a magnitude obtained by subtracting the magnitude of the output of the second microphone filter from the magnitude of the output of the first microphone, and calculating, as a front-back uncorrelated level, a magnitude obtained by subtracting the magnitude of the output of the third microphone filter from the magnitude of the output of the first microphone; A left-right uncorrelated level map showing a correspondence between each value of the left-right uncorrelated level and a left-right direction position inside the automobile of a sound source at which the value of the left-right uncorrelated level is obtained; A front-back uncorrelated level map showing a correspondence between each value of the front-back uncorrelated level and a front-back direction position inside the automobile of a sound source at which the value of the front-back uncorrelated level is obtained; A first filter for the first microphone in which a transfer function for approximately extracting a component correlated with the output of the second microphone from the output of the first microphone is set; A second filter for the first microphone in which a transfer function for approximately extracting a component correlated with the output of the third microphone from the output of the first microphone is set; Calculate the magnitude obtained by adding the magnitude of the output of the first filter for the first microphone and the magnitude of the output of the filter for the second microphone as the left-right correlation level, and calculate the magnitude obtained by adding the magnitude of the output of the second filter for the first microphone and the magnitude of the output of the filter for the third microphone as the front-back correlation level; a correlation level calculation means; A left-right correlation level map showing the correspondence between each value of the left-right correlation level and the position in the left-right direction inside the vehicle of the sound source from which the value of the left-right correlation level is obtained; A front-back correlation level map showing the correspondence between each value of the front-back correlation level and the position in the front-back direction inside the vehicle of the sound source from which the value of the front-back correlation level is obtained; It has a speaker position calculation means, The speaker position calculation means is, Among the ranges of the left-right positions of the sound source from which the calculated value of the left-right uncorrelation level shown in the left-right uncorrelation level map is obtained on the line connecting the first microphone and the second microphone, the range of the position closer to the standard speaking position of the user sitting in the first seat is calculated as the range of the left-right position of the speaker who is the user sitting in the first seat, Among the ranges of the front-back positions of the sound source from which the calculated value of the front-back uncorrelation level shown in the front-back uncorrelation level map is obtained on the line connecting the first microphone and the third microphone, the range of the position closer to the standard speaking position of the user sitting in the first seat is calculated as the range of the front-back position of the speaker, On the line connecting the first microphone and the second microphone, within the calculated range of the left-right position of the speaker, the position in the left-right direction inside the vehicle of the sound source from which the calculated value of the left-right correlation level shown in the left-right correlation level map is obtained is calculated as the left-right position of the speaker, On the line connecting the first microphone and the third microphone, within the calculated range of the front-back position of the speaker, the position in the front-back direction inside the vehicle of the sound source from which the calculated value of the front-back correlation level shown in the front-back correlation level map is obtained is calculated as the front-back position of the speaker. A speaker position detection system characterized by this.

3. The speaker position detection system according to claim 1, wherein the transfer function of the second microphone filter is set in advance by using, as an error, the difference between the output of a first adaptive filter that takes the output of the second microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, causing the first adaptive filter to perform an adaptation operation, and setting the transfer function of the converged first adaptive filter as the transfer function of the second microphone filter; the transfer function of the third microphone filter is set in advance by using, as an error, the difference between the output of a second adaptive filter that takes the output of the third microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, causing the second adaptive filter to perform an adaptation operation, and setting the transfer function of the converged second adaptive filter as the transfer function of the third microphone filter. A speaker position detection system characterized by the above.

4. The speaker position detection system according to claim 2, wherein the transfer function of the second microphone filter is set in advance by using, as an error, the difference between the output of a first adaptive filter that takes the output of the second microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, causing the first adaptive filter to perform an adaptation operation, and setting the transfer function of the converged first adaptive filter as the transfer function of the second microphone filter; the transfer function of the third microphone filter is set in advance by using, as an error, the difference between the output of a second adaptive filter that takes the output of the third microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the first microphone, causing the second adaptive filter to perform an adaptation operation, and setting the transfer function of the converged second adaptive filter as the transfer function of the third microphone filter; The transfer function of the first filter for the first microphone is set in advance by using, as an error, the difference between the output of a third adaptive filter that takes the output of the first microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the second microphone, causing the third adaptive filter to perform an adaptation operation, and setting the transfer function of the converged third adaptive filter as the transfer function of the first filter for the first microphone. The transfer function of the second filter for the first microphone is set in advance by using, as an error, the difference between the output of a fourth adaptive filter that takes the output of the first microphone as an input while outputting a predetermined tuning sound from a predetermined position inside the vehicle, and the output of the third microphone, causing the fourth adaptive filter to perform an adaptation operation, and setting the transfer function of the converged fourth adaptive filter as the transfer function of the second filter for the first microphone. A speaker position detection system characterized by this.

5. The speaker position detection system according to claim 1, 2, 3, or 4, wherein the speaker position calculation means determines that there is a speaker in the first seat when both the magnitude of the left-right uncorrelated level and the magnitude of the front-rear uncorrelated level are greater than a predetermined threshold value, and calculates the position of the speaker in the left-right direction and the position in the front-rear direction. A speaker position detection system characterized by this.

6. An in-vehicle communication support system including the speaker position detection system according to claim 1, 2, 3, 4, or 5, a speaker, and output processing means for performing a process of outputting the sound collected by the first microphone to the speaker, wherein the output processing means switches at least one of the presence or absence of output of the sound collected by the first microphone to the speaker and the gain of the sound collected by the first microphone output to the speaker according to the speaker position represented by the position of the speaker in the left-right direction and the position in the front-rear direction calculated by the speaker position calculation means. An in-vehicle communication support system characterized by this.

Citation Information

Patent Citations

  • Directional controller for microphone

    JP1993207117A

  • Active muffler

    JP2001142469A

  • In-vehicle conversation assisting device

    JP2002051392A

  • Loudspeaker direction detecting circuit

    JP2002315089A

  • Sound collection device, control method thereof, and control program thereof

    JP2016005181A