System and method for preparing a reference signal for an acoustic echo canceller
By downselecting and prefiltering speaker channels and separating them into subbands, the echo canceller's performance is enhanced, addressing computational inefficiencies and convergence issues in multi-channel systems.
Patent Information
- Application Number
- JP2023573441
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-04
- Filing Date
- 2022-06-02
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Echo cancellers with a large number of channels are computationally expensive and have slow convergence speeds, and software and hardware limitations impose constraints on the maximum number of reference channels, necessitating a reduction in channels without affecting echo cancellation performance.
Optimize reference channel inputs by downselecting a smaller set of speaker channels, prefiltering them to approximate acoustic signals at the microphone, and separating them into subbands to generate a reference signal for the echo canceller, using methods that maximize multicoherence and echo return loss enhancement metrics.
Reduces computational cost and improves convergence speed while maintaining effective echo cancellation performance by optimizing the reference channels based on conditions within the vehicle cabin.
Smart Images

Figure 0007729921000015 
Figure 0007729921000016 
Figure 0007729921000017
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 17 / 339,332, filed June 4, 2021, and entitled "Systems and Methods for Preparing Reference Signals For an Acoustic Echo Canceller," the entire disclosure of which is incorporated herein by reference. [Background technology]
[0002] The present disclosure generally relates to systems and methods for preparing a reference signal for an acoustic echo canceller. Summary of the Invention
[0003] All embodiments and features mentioned below can be combined in any technically possible manner.
[0004] According to one aspect, a method for preparing a reference signal for an echo cancellation system disposed within a vehicle includes receiving a plurality of drive signals, each drive signal being provided to an associated transducer of a plurality of acoustic transducers, where the associated acoustic transducer converts the drive signal into an acoustic signal, each of the plurality of acoustic transducers being disposed within the vehicle such that each acoustic signal is audible within a passenger compartment of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone disposed within the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; and outputting the summed reference signal to the echo cancellation system.
[0005] In one embodiment, the method further includes summing a second subset of the plurality of filtered signals to generate a second summed reference signal, and outputting the second summed reference signal to the echo cancellation system.
[0006] In one embodiment, the multiple filters are selected from a set of filters according to conditions within the vehicle.
[0007] In one embodiment, the state is determined according to at least one of seat position, window position, number of occupants, occupant position, and door position.
[0008] According to one aspect, a method for preparing a reference signal for an echo cancellation system located in a vehicle includes receiving a plurality of drive signals, each drive signal being supplied to an associated transducer of a plurality of acoustic transducers located in the vehicle, whereby the associated acoustic transducer converts the drive signal into an acoustic signal; separating each of the plurality of drive signals into a plurality of frequency subbands; and providing a first selection of the plurality of drive signals to the echo cancellation system as a reference signal for a first frequency subband of the plurality of frequency subbands, the first selection including at least a subset of the plurality of drive signals; and providing a second selection of the plurality of drive signals to the echo cancellation system as a reference signal for a second frequency subband of the plurality of frequency subbands, the second selection including a subset of the plurality of drive signals, the second selection including a subset of the plurality of drive signals, the first selection and the second selection differing by at least one drive signal.
[0009] In one embodiment, at least one of the first selection of the plurality of drive signals and at least one of the second selection of the plurality of drive signals are summed before being provided to the echo cancellation system as a reference signal.
[0010] In one embodiment, a first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
[0011] In one embodiment, the first selection of the plurality of drive signals is filtered with a respective one of the plurality of filters before being provided to the echo cancellation system as a reference signal, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located in the vehicle such that each of the first selection of the plurality of drive signals estimates a respective acoustic signal in a first subband at the microphone.
[0012] In one embodiment, a first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
[0013] In one embodiment, the multiple filters are selected from a set of filters according to conditions within the vehicle.
[0014] According to another aspect, a non-transitory storage medium storing program code that, when executed by a processor, prepares a reference signal for an echo cancellation system disposed in a vehicle, the program code, when executed, includes receiving a plurality of drive signals, each drive signal provided to an associated transducer of a plurality of acoustic transducers, where the associated acoustic transducer converts the drive signal into an acoustic signal, each of the plurality of acoustic transducers disposed in the vehicle such that each acoustic signal is audible within a cabin of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone disposed in the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; and outputting the summed reference signal to the echo cancellation system.
[0015] In one embodiment, the program code further includes summing a second subset of the plurality of filtered signals to generate a second summed reference signal, and outputting the second summed reference signal to the echo cancellation system.
[0016] In one embodiment, the multiple filters are selected from a set of filters according to conditions within the vehicle.
[0017] In one embodiment, the state is determined according to at least one of seat position, window position, number of occupants, occupant position, and door position.
[0018] According to another aspect, a non-transitory storage medium storing program code that, when executed by a processor, prepares a reference signal for an echo cancellation system disposed in a vehicle, the program code, when executed, includes the steps of: receiving a plurality of drive signals, each drive signal being supplied to an associated transducer of a plurality of acoustic transducers disposed in the vehicle, such that the associated acoustic transducer converts the drive signal into an acoustic signal; separating each of the plurality of drive signals into a plurality of frequency subbands; and, for a first frequency subband of the plurality of frequency subbands, providing a first selection of the plurality of drive signals to the echo cancellation system as a reference signal, the first selection including at least a subset of the plurality of drive signals; and, for a second frequency subband of the plurality of frequency subbands, providing a second selection of the plurality of drive signals to the echo cancellation system as a reference signal, the second selection including a subset of the plurality of drive signals, the second selection including a subset of the plurality of drive signals, the first selection and the second selection differing by at least one drive signal.
[0019] In one embodiment, at least one of the first selection of the plurality of drive signals and at least one of the second selection of the plurality of drive signals are summed before being provided to the echo cancellation system as a reference signal.
[0020] In one embodiment, a first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
[0021] In one embodiment, the first selection of the plurality of drive signals is filtered with a respective one of the plurality of filters before being provided to the echo cancellation system as a reference signal, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located in the vehicle such that each of the first selection of the plurality of drive signals estimates a respective acoustic signal in a first subband at the microphone.
[0022] In one embodiment, a first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
[0023] In one embodiment, the multiple filters are selected from a set of filters according to conditions within the vehicle. [Brief explanation of the drawings]
[0024] In the drawings, like reference numbers generally refer to the same parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of various aspects. [Figure 1] 1 depicts a schematic diagram of an echo cancellation system, according to one embodiment; [Figure 2] 1 is a schematic diagram of a speaker system in a vehicle, according to one embodiment. [Figure 3] FIG. 2 is a schematic diagram of a reference signal generator according to one embodiment. [Figure 4] 4 is a flowchart of a method for generating a reference signal for an echo canceller, according to an embodiment. [Figure 5] FIG. 2 is a schematic diagram of a reference signal generator according to one embodiment. [Figure 6] 4 is a flowchart of a method for generating a reference signal for an echo canceller, according to an embodiment. [Figure 7] 10 illustrates a plot of selected speaker signals across multiple subbands according to one embodiment. [Figure 8A]4 is a flowchart of a method for generating a reference signal for an echo canceller, according to one embodiment. [Figure 8B] 4 is a flowchart of a method for generating a reference signal for an echo canceller, according to one embodiment. [Figure 8C] 4 is a flowchart of a method for generating a reference signal for an echo canceller, according to one embodiment. [Figure 8D] 4 is a flowchart of a method for generating a reference signal for an echo canceller, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0025] Echo cancellers that implement a large number of channels are computationally expensive and have slow convergence speeds. Furthermore, when a third-party product performs echo cancellation, software and / or hardware limitations may impose limits on the maximum number of reference channels. As a result, it is desirable to reduce the number of channels in an echo canceller without affecting its echo cancellation performance.
[0026] Various examples described in this disclosure relate to systems and methods for optimizing reference channel inputs to an acoustic echo canceller in a multi-channel system. Some examples downselect the number of speaker channels to a smaller set of reference channels to improve performance, reduce computational cost, or comply with the number of reference channels required by the echo canceller. In some examples, the speaker channels are prefiltered to approximate the acoustic signal received at the microphone before being summed. Some examples separate the speaker channels into a set of subbands and downselect within the subbands to generate a reference signal for the echo canceller.
[0027] 1 shows an exemplary multi-channel acoustic echo cancellation system 100. Speaker channels 102a-102M each receive a drive signal u1(n),...,uM (n) to speakers (104a to 104M (alternatively called acoustic transducers)), where M is the total number of speaker channels, drive signals, and speakers. M (n) is composed of one or more program content signals such as music, navigation commands, voice assistance, etc. Each speaker 102a to 102M receives a drive signal u1(n)-u M (n) into respective acoustic signals that are audible within the vehicle cabin 106. (As used in this disclosure, a speaker may be any transducer suitable for receiving an electrical signal and converting it into an acoustic signal that is audible within the vehicle cabin.)
[0028] The echo cancellation system 100 further includes at least one microphone 108 positioned within the vehicle cabin 106 to receive audio signals from at least one user seated within the vehicle. However, due to its location, the microphone 108 receives signals other than audio signals, including acoustic signals generated by the speakers 104a-104M and noise within the vehicle cabin (e.g., road noise). The microphone signal y(n) can therefore be expressed as: y(n)=s(n)+d(n)+v(n) (1) where s(n) is the desired signal (typically a speech signal) and v(n) is the noise signal.
[0029]
number
[0030] In operation, echo canceller 110 is an N-channel echo canceller and, as such, generates drive signals u1(n)-u1(n) in accordance with reference signal generator 112, as described in more detail below. M (n) or otherwise selected from the drive signals u1(n)-uM (n). As mentioned above, a large number of reference channels increases computational complexity and slows convergence. Vehicles with high-quality audio typically use a large number of speakers. Figure 2 shows an example of such a vehicle in which the number of speaker channels equals 14 (three dash speakers, three speakers in each of the front doors, one channel in each of the rear doors, two back speakers, and a woofer).
[0031] Thus, in the example of FIG. 1, the reference signal generator 112 selects from the M speaker signals and / or converts them into N reference signals x1(n),...,x N (n), where N≦M, and fed to an N-channel acoustic echo canceller. From the N signals, the echo canceller 102 generates an echo signal
[0032]
number
[0033]
number
[0034] Combined echo signal
[0035]
number
[0036]
number
[0037]
number
[0038] Drive signal u1(n)-u M It should be understood that drive signal u1(n)-u(n) may be the result of various upstream processing stages, including equalization, upmixing, routing, and / or sound stage rendering. Alternatively, or in addition, drive signal u1(n)-u M It should be understood that the estimated speech signal s(n) may undergo additional processing before being converted into an acoustic signal by the speakers 104a-104M. If the drive signal undergoes additional processing, such as equalization, before being provided to the speakers 104a-104M, the adaptive filters 114a-114N attempt to distinguish between the modified echo path and the additional processing that the reference signal undergoes before being converted by the speakers 104a-104M. Furthermore, it should be understood that the estimated speech signal s(n) may undergo further processing, such as with a post-filter, to improve the estimated speech signal s(n).
[0039] One metric that can be used to evaluate the performance of the generated reference channel is the combined reference x1(n),.,x NThe multicoherence C_xy between the input and the microphone signal y(n) is the multicoherence C_xy between the input and the microphone signal y(n). Multicoherence yields a value between 0 and 1 that represents the level of linearity between the input and the output. In the absence of a noise signal v(n) and a desired signal s(n), a multicoherence of 1 means that a linear relationship between the input and the output can be found, achieving perfect cancellation of the echo signal d(n). As a result, the solution that maximizes multicoherence is considered optimal.
[0040] Another metric that can be used to evaluate the performance of the generated reference is the echo return loss enhancement (ERLE) when the generated signal is used as a reference channel in a multi-channel acoustic echo canceller and the microphone signal is used as the input signal. The ERLE can be used to evaluate the actual level of echo cancellation and the convergence speed of the adaptive filter when the echo path is changed. While this disclosure describes optimizing the reference channel with respect to multi-coherence, it will be understood that any suitable metric for determining the performance of an echo canceller can be used instead to optimize the reference channel generation (e.g., ERLE).
[0041] Drive signal u1(n)-u M (n) to generate a reference signal x1(n)-x NSeveral methods for rendering (n) are described below. These methods can be executed by one or more processors, such as a digital signal processor, in communication with non-transitory stored program code for executing the methods. In some examples, the reference signal generator can be implemented in a processor separate from the echo canceller, receiving drive signals to be converted into acoustic signals and preparing, from these drive signals, reference signals that are input to the echo cancellation processor / device. For example, a processor can be added to a vehicle with an existing echo cancellation unit to generate a reference signal for the existing unit. In an alternative example, the same processor that implements the functionality of the echo canceller can be used to generate the reference signal from the speaker drive signal.
[0042] In a first method, a subset of drive signals consisting of N drive signals is selected from the M drive signals and used to form the reference signal. In one embodiment, the selection of the N drive signals from the drive signals is selected to maximize the total multicoherence between the selected input and microphone signals.
[0043]
number
[0044]
number
[0045] In other words, the selection of the drive signal that maximizes the multicoherence metric is deemed optimal and is selected to generate the reference channel.
[0046] 2 shows one example of these combinations, where M=14 and N=5. Thus, in this example, speakers 104a-104m are implemented as speakers 104a-104n. Since N=5, five drive signals are selected from the 14 available drive signals. In one example, the drive signals for speakers 104d, 104f, 104h, 104j, and 104l may be selected as the set of drive signals that provides the greatest total multicoherence between the selection and the microphone signals.
[0047] Typically, the selection of the speaker drive signals that provide the greatest multicoherence is performed during the system design phase. That is, while the vehicle's audio system is being designed, various tests can be performed to determine which set of speaker drive signals provides the greatest multicoherence with the microphone signals, and the selection is not changed after the vehicle is sold. However, in an alternative example, the reference signal generator 112 may be configured to generate the drive signals u1(n)-u M The multicoherence of various selections of (n) can be compared to select, during runtime, the set of N drive signals as the reference signal that provides the best multicoherence with the microphone signals. This set of N drive signals can be updated (e.g., periodically) as conditions in the vehicle cabin change. Additionally or alternatively, the set of N drive signals can be updated to account for changes in program content (e.g., music, announcements, etc.) being played through speakers 104a-104N.
[0048] In another example, the drive signals of a group of speakers can be summed into the same reference channel. The drive signals of speakers in the same group are combined using summation as follows:
[0049]
number
[0050] For example, the speakers in Figure 2 can be divided into four groups of speakers: front left, front right, rear left, and rear right, as well as a center channel. More specifically, in this example, the front left group includes speakers 104k-104n, the front right group includes speakers 104b and 104e, the rear left group includes speakers 104i and 104j, the rear right group includes speakers 104f and 104g, and the center channel (group) includes speakers 104a and 104h. The drive signals for each of these groups can be summed together to form a single reference channel from each group.
[0051] Another example of this method is shown in the block diagram of Figure 3. Figure 3 shows a simplified example in which M=4 and N=2, and thus speakers 104a-104M include speakers 104a-104d. As shown, in this example, reference channel generator 112 includes summation blocks 302a, 302b, which sum the drive signals for speakers 104a-104b and the drive signals for speakers 104c-104d (i.e., drive signals u1(n)-u4(n)) to provide reference signals x1(n) and x2(n).
[0052] Typically, the grouped speakers are co-located within the vehicle cabin, i.e., they form a set of adjacent speakers located within a particular structure or adjacent structures within the vehicle cabin (i.e., doors, pillars, etc.) As explained above, the groups may be selected to maximize multi-coherence between the generated reference signals and the microphone signals.
[0053] Note that this example generally assumes that speakers in the same group with high mutual coherence do not have overlapping frequency ranges, since such overlapping frequency ranges could result in constructive or destructive summation between the drive signals. Thus, for example, a center channel group consisting of tweeter speaker 104a and woofer speaker 104h can be added without frequency overlap. However, in most cases, this condition depends on the tuning of the audio system and cannot always be guaranteed.
[0054] Thus, if two speakers share similar content and have high mutual coherence, only one of the speakers needs to be included in the combined criteria. Note, however, that speakers with low mutual coherence can be grouped together, even if they share the same frequency band. Thus, in a variation of the example of FIG. 2, some of speakers 104a-104m may be omitted from the combined speaker group. For example, returning to FIG. 2, to avoid grouping speakers with high mutual coherence within similar frequency bands, the front left group includes speakers 104k-104m, the front right group includes speakers 104c-104e, the rear left group includes speaker 104i, the rear right group includes speaker 104g, and the center channel (group) includes speakers 104a and 104h. Compared to the previous grouping, loudspeakers 104c, 104f, 104j, and 105n are excluded to avoid summing loudspeakers with high mutual coherence and overlapping frequencies.
[0055] When selecting speakers to sum within a group, the speaker group combination with the highest total multicoherence relative to the microphone signal is generally considered optimal. The speaker drive signals to be summed and the grouping of the selected speaker drive signals (including the grouping of all speaker drive signals) are generally selected during the design phase to maximize multicoherence between the generated reference signal and the microphone signal, but in alternative examples, the selected speaker drive signals and the grouping of the speaker signals can be adjusted during runtime to maximize multicoherence in consideration of changing conditions in the vehicle cabin and / or changing program content (e.g., music, announcements, etc.). Thus, the selection of speaker drive signals and the grouping of speaker drive signals can be updated during runtime (e.g., periodically) to maximize multicoherence between the reference signal and the microphone signal.
[0056] Figure 4 shows a flowchart of a method 400 for generating a reference signal in accordance with the example of summing drive signals described in connection with Figure 3. It should be appreciated that method 400 may be implemented to provide a reference signal for any suitable echo cancellation system, such as echo cancellation system 100.
[0057] In step 402, a plurality of drive signals are received, each of which is provided to an associated transducer of a plurality of acoustic transducers, which converts the drive signal into an acoustic signal. The plurality of drive signals may be the result of upstream processing such as equalization, upmixing, routing, and / or sound stage rendering, and such processes are known in the art. Additionally, the drive signals may undergo further processing before being converted into acoustic signals by the acoustic transducers.
[0058] In step 404, at least a subset of the multiple drive signals is summed to generate a summed reference signal. Typically, multiple subsets of drive signals are summed to generate multiple summed reference signals that are provided to the echo cancellation system in step 406. Which drive signals are summed with which drive signals, i.e., the drive signals that make up any of the summed subsets of drive signals, may be selected to maximize a metric such as multicoherence between the resulting summed reference signal and the microphones. Generally speaking, as discussed above, if a speaker has high mutual coherence with another speaker and operates within the same frequency range (or at least overlapping frequency ranges), at least one of these speakers may be excluded from any subsets that are summed to generate the reference signal.
[0059] In step 406, the summed reference signals are output to an echo canceller, such as echo canceller 110, although any suitable echo canceller can be used. These reference signals are typically filtered by the echo canceller to provide an estimated echo signal that is subtracted from the microphone signal to render an estimated speech signal.
[0060] Alternatively, the selected drive signals are first filtered before being combined into a reduced number of signals, as shown in the following equation:
[0061]
number
[0062] In this example, the difference between equation (4) and equation (5) is k (n) is the corresponding impulse response h 0,k (n) is the filtering. 0,k(n) is chosen to maximize the multicoherence between the generated drive signal and the microphone signal. Assuming there is no desired signal s(n) and no noise signal v(n) and no changes in the echo path, h(n) for k=1,...,M 0,k (n)=h k (n) is set, h1(n)...h M (n) results in multicoherence equal to 1, making this solution optimal. This applies to any number of reference channels, as long as the drive signals for all loudspeakers are included when generating the reference signal. However, it is still generally desirable to group co-located loudspeakers together to improve the tracking ability of the adaptive filter.
[0063] In other words, in this example, each drive signal is filtered using a fixed filter that estimates the transfer function between the respective loudspeaker (i.e., the loudspeaker to which the drive signal is applied) and the microphone. FIG. 5 shows an example implementation of drive signals filtered according to Equation 5. The set of fixed filters 602a-602N, here fixed filters 602a-602d, performs an estimate of the transfer function from the associated loudspeaker to the microphone (i.e., the echo path). Thus, fixed filter 602a receiving drive signal u1(n) performs an estimate of the transfer function from loudspeaker 104a to microphone 108, for example, such that the output of filter 602a represents an estimate of the acoustic signal generated by loudspeaker 104a at microphone 108. Similarly, fixed filter 602b receiving drive signal u2(n) performs an estimate of the transfer function from loudspeaker 104b to microphone 108, for example, so that the output of filter 602b represents an estimate of the acoustic signal produced by loudspeaker 104b at microphone 108, and so on.
[0064] This example effectively pre-filters the drive signals to generate reference signals, thus achieving the function of filters 114a-114N (for all M drive signals, where M>N) before they are input to echo canceller 110. (In certain examples, not all drive signals are filtered, but filtering all drive signals generally provides best results.) Thus, in this example, filters 114a-114N are not adapted to estimate the echo path from speakers 104a-104M, but rather to estimate the difference between the transfer functions of fixed filters 602a-602N, since there is typically some difference between the predetermined transfer function of the fixed filters and the actual echo path between each speaker 104a-104N and the microphone.
[0065] In reality, the echo path often changes. To account for this, the transfer functions from all loudspeakers to the microphones can be measured in advance for different conditions (e.g., seat position, window position, number of occupants, occupant position, and door position). These measured transfer functions are then combined to form a filter h that is applied in all conditions. 0,k In the example of FIG. 5, for example, the estimated transfer function
[0066]
number
[0067]
number
[0068]
number
[0069] Alternatively, a particular transfer function h 0,k A set of filters 602a-602N implementing (n) may be stored for each interior state of the vehicle. These filters may be stored and implemented in response to changing conditions within the vehicle. Thus, in the example of FIG. 5, one set of filters 602a-602d may be stored and implemented for one interior state (e.g., one seat position). Meanwhile, different sets of pre-stored fixed filters 602a-602d may be loaded and implemented for different interior states using different sets of transfer functions appropriate for the different states. In this way, the fixed filters 602a-602d appropriately represent the transfer functions for a particular interior state. While this method generally produces better results than finding a combined transfer function representing multiple interior states, it requires additional storage and greater processing power to implement because it must receive inputs representing the interior state and then load and implement the correct set of filters from a repository of filters.
[0070] Additionally, if the drive signals undergo additional processing, such as equalization, before being provided to the speakers 104a-104M, the fixed filters 602a-602d can further account for the additional processing. In this example, different fixed filters can be stored and loaded for different types of program content that use different equalization settings.
[0071] Figure 6 shows a flowchart of a method 600 for generating a reference signal according to the example described in connection with Figure 5. It should be appreciated that method 600 may be implemented to provide a reference signal for any suitable echo cancellation system, such as echo cancellation system 100.
[0072] In step 602, a plurality of drive signals are received, each of which is provided to an associated transducer of the plurality of acoustic transducers such that the associated acoustic transducer converts the drive signal into an acoustic signal. The plurality of drive signals may be the result of upstream processing such as equalization, upmixing, routing, and / or sound stage rendering, and such processes are known in the art. Additionally, the drive signals may undergo further processing before being converted into acoustic signals by the acoustic transducers.
[0073] In step 604, each drive signal is filtered using a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located within the vehicle, such that each of the plurality of filtered signals estimates a respective acoustic signal at the microphone. The transfer function implemented by each filter may be a combined transfer function between the associated acoustic transducer (i.e., the acoustic transducer receiving the filtered drive signal) and the microphone under various conditions within the vehicle. Alternatively, a set of filters implementing transfer functions for various conditions within the vehicle (e.g., seat position, window position, number of occupants, occupant position, and door position) may be stored and loaded according to the current conditions within the vehicle.
[0074] At least a subset of the multiple filtered signals are summed in step 606 to generate a summed reference signal. This step is typically to reduce the number of reference signals fed to the echo canceller in order to reduce processing time or to adhere to input constraints imposed by the echo canceller. In this step, it is typically desirable to sum groups of drive signals fed to co-located speakers.
[0075] In step 608, the summed reference signals are output to an echo canceller, such as echo canceller 110, although any suitable echo canceller can be used. These reference signals are typically filtered by the echo canceller to provide an estimated echo signal that is subtracted from the microphone signal to render an estimated speech signal.
[0076] While the above-described systems and methods for generating reference signals have been described for all frequencies, the optimal solutions generated using these methods are not necessarily optimal in each frequency bin. For example, a subset of speaker drive signals selected to optimize multi-coherence between the reference signal and microphones within the 100 Hz to 600 Hz frequency range may perform well for that range but poorly for another frequency range, such as 5000 Hz to 5500 Hz. To account for this, the entire frequency spectrum can be divided into subbands, and the optimal solution can be found for each subband using the methods described above. The objective is still to find a solution that maximizes the total coherence in each of the subbands.
[0077] For example, the optimal selection of the loudspeaker driving signal in the jth subband is obtained as follows:
[0078]
number
[0079] Similarly, as described in connection with FIGS. 3 and 4, an optimal solution can be obtained for each subband by summing groups of drive signals within each subband. For example, drive signals can be grouped and summed within each subband to optimize multicoherence between the resulting reference signal and microphone signal. For example, in the simple example of FIG. 3, in one subband, the drive signals for speakers 104a and 104b can be summed, and the drive signals for speakers 104c and 104d can be summed. Meanwhile, in a different subband, drive signals 104a and 104c can be summed, and 104b and 104d can be summed to optimize multicoherence for each subband. Similarly, as described above, when high mutual coherence exists, it may be beneficial to omit certain speaker drive signals that have frequency overlap. Therefore, grouped speaker drive signals and groups of speaker drive signals can be selected for each subband to optimize multicoherence within that subband.
[0080] Furthermore, as described in connection with FIGS. 5 and 6, filters can be selected within each subband to optimize multicoherence within the subband. Different filters can be selected to maximize multicoherence with the subband. Thus, at least one set of filters can be implemented for each subband. As described above, the filters for each subband can combine transfer functions for various conditions to arrive at a transfer function that performs well for most conditions. Alternatively, a set of filters can be stored and loaded for each subband to account for changing conditions within the vehicle cabin. Thus, one set of filters within a subband can be implemented for one condition, and another set of filters can be loaded from storage and implemented to account for different conditions within the vehicle cabin. Thus, the implemented filters can be optimized within each subband to maximize multicoherence between the reference signal and the microphone signal.
[0081] If the echo canceller is designed to receive input reference signals that are not subbanded, the subbanded signals can be combined to render the desired number of reference signals. Thus, in the example of Figure 7, the optimized reference signals in the subbands can be summed across frequencies to generate five reference signals. Alternatively, if the echo canceller is configured to receive subbanded signals, the optimal selected signals can be provided directly as subbanded reference signals.
[0082] Furthermore, the method for optimizing the multicoherence of the reference signal generation may vary across subbands. For example, the filter and sum method (i.e., as described in connection with FIGS. 5 and 6) may work well at low frequencies because the echo path undergoes smaller changes at those frequencies. However, this method suffers from performance degradation at higher frequencies due to larger changes in the echo path at those frequencies. On the other hand, the speaker drive signal downselection method may work well at high frequencies because the number of speakers that can reproduce at those frequencies is limited. Therefore, the optimal solution may be a hybrid between different methods, in which solutions at different subbands are generated using different methods. Thus, in one embodiment, an optimal filter may be implemented in the low-frequency subbands, but at high frequencies, a limited number of speaker drive signals may be selected that optimize the multicoherence within those subbands. In fact, it should be understood that the above-described methods may be implemented within each subband to provide a reference signal within each subband in any manner technically possible.
[0083] It should further be appreciated that multiple microphones may be used to capture the user's voice. In these cases, an acoustic echo canceller may be implemented for each microphone, and the resulting estimated speech signals output from the echo cancellers may be combined into a single output estimated speech signal. For example, if P microphones are used to capture the user's voice, a set of P echo cancellers, each associated with a respective microphone, may be used to cancel the echo in the associated microphone. Similar to echo canceller 110, the P echo cancellers operate by estimating the echo path from each speaker to its associated microphone. A respective reference signal generator (i.e., from the P reference signal generators) may be used to select or combine speaker signals to obtain a reference signal for each echo canceller. Each reference signal generator may use a method for selecting or combining speaker signals, as described above in connection with FIGS. 1-7 and below in connection with FIGS. 8A-8D.
[0084] Alternatively, the microphone signal y(n) may be the result of combining signals output from an array of microphones using a beamformer. The resulting combined microphone signal may be treated as a single microphone signal, and echo signals within the combined microphone signal may be canceled using methods for selecting or combining speaker signals, such as those described above in connection with Figures 1-7, and methods described below in connection with Figures 8A-8D.
[0085] Figure 8 shows a flowchart of a method 800 for generating a reference signal according to the example described in connection with Figure 7. It should be understood that method 800 may be implemented to provide a reference signal for any suitable echo cancellation system, such as echo cancellation system 100.
[0086] In step 802, a plurality of drive signals are received, each of which is provided to an associated transducer of the plurality of acoustic transducers such that the associated acoustic transducer converts the drive signal into an acoustic signal. The plurality of drive signals may be the result of upstream processing such as equalization, upmixing, routing, and / or sound stage rendering, and such processes are known in the art. Additionally, the drive signals may undergo further processing before being converted into acoustic signals by the acoustic transducers.
[0087] In step 804, each of the multiple drive signals is separated into multiple frequency sub-bands. Any suitable method for sub-banding the drive signals may be used. For example, each drive signal may be filtered with a parallel bank of bandpass filters, with each filter in the bank of bandpass filters implementing a cutoff frequency to generate a different sub-band of the drive signal.
[0088] At least one reference signal is generated in each subband in step 806. The reference signals generated in each subband in this step can be generated using any combination of the methods described above. Various methods for generating reference signals in each subband are described below in connection with steps 806b-806d. In some examples, different methods may be used for different subbands to optimize the resulting reference signals and / or reduce computation time.
[0089] In a first method, shown in step 806b of FIG. 8B, within each subband, a subset of drive signals is selected as the reference signal. In this example, a different subset of drive signals is selected for each subband such that the subsets between at least two bands differ by at least one drive signal. Here, the subsets may be selected to maximize multicoherence between the selected reference signal and the microphones within each subband.
[0090] In a second method, in step 806c, at least a subset of the filtered signals is summed within at least one subband to generate at least a summed reference signal. Typically, within a subset, multiple subsets of drive signals are summed to generate multiple summed reference signals that are provided to the echo cancellation system in step 808. Which drive signals are summed with which drive signals, i.e., the drive signals that make up any of the summed subsets of drive signals, can be selected to maximize a metric such as multicoherence between the resulting summed reference signal and the microphones. Generally speaking, as described above, if a speaker has high mutual coherence with another speaker and operates within the same frequency range (or at least overlapping frequency ranges), at least one of these speakers can be excluded from any subsets that are summed to generate the reference signal within the subband.
[0091] In a third method, in step 806d, each drive signal is filtered within at least one subband using a respective filter from a plurality of filters to generate a plurality of filtered signals, each of which approximates a transfer function from an associated acoustic transducer to a microphone disposed within the vehicle, such that each of the plurality of filtered signals estimates a respective acoustic signal within the subband at the microphone. The transfer function implemented by each filter may be a combined transfer function between the associated acoustic transducer (i.e., the acoustic transducer receiving the filtered drive signal) and the microphone under various conditions within the vehicle. Alternatively, a set of filters implementing transfer functions for various conditions within the vehicle (e.g., seat position, window position, number of occupants, occupant position, and door position) may be stored and loaded according to the current conditions within the vehicle. Furthermore, if the drive signal undergoes additional processing, such as equalization, before being provided to the speakers, the plurality of filters may further take into account the additional processing. In this example, different filters may be stored and loaded for different types of program content that use different equalization settings.
[0092] In step 808, the reference signals in each subband are output to an echo canceller, such as echo canceller 110, although any suitable echo canceller can be used. These reference signals are typically filtered by the echo canceller to provide an estimated echo signal that is subtracted from the microphone signal to render an estimated speech signal.
[0093] The above methods are typically selected during the design phase to maximize a metric such as multicoherence between the reference signal and the microphone for each subband, although alternatively, these methods may be periodically updated or modified to account for changes in cabin conditions / program content.
[0094] Additionally, while method 400 has been described for a single microphone, the steps of method 400 can be repeated for multiple microphones, each with its own echo canceller, i.e., a reference signal can be generated for each echo canceller associated with each separate microphone. Alternatively, the microphone signal can be the result of combining signals output from an array of microphones using a beamformer. The resulting composite microphone signal can be treated as a single microphone signal, and method 400 can be employed to cancel echo signals within the composite microphone signal.
[0095] The mathematical formulas provided in this disclosure are simplified solely for the purpose of illustrating the principles of aspects of the present invention and should not be construed as exclusive or limiting in any way. Furthermore, variations of the mathematical formulas are contemplated and are within the spirit and scope of the present disclosure.
[0096] Regarding the use of symbols herein, an uppercase letter, e.g., H, generally represents a term, signal, or quantity in the frequency or spectral domain, and a lowercase letter, e.g., h, generally represents a term, signal, or quantity in the time domain. The relationship between the time domain and the frequency domain is generally well known and explained at least in the field of Fourier mathematics or Fourier analysis, and therefore will not be presented herein. Additionally, signals, transfer functions, or other terms or quantities represented by symbols herein may be operated on, considered, or analyzed in analog or discrete form. In the case of time-domain terms or quantities, the analog time index, e.g., t, and / or the discrete sample index, e.g., n, may be interchanged or omitted in various cases. Similarly, in the frequency domain, the analog frequency index, e.g., f, and the discrete frequency index, e.g., k, are most often omitted. Furthermore, the relationships and calculations disclosed herein generally may exist or be performed in either the time domain or the frequency domain, and in either the analog domain or the discrete domain, as will be understood by those skilled in the art. Therefore, various examples to illustrate all possible variations in the time or frequency domain, and in the analog or discrete domain, are not presented herein.
[0097] The functionality or portions thereof, and various modifications thereof (hereinafter "functionality") described herein may be implemented, at least in part, via a computer program product (e.g., a computer program tangibly embodied in an information carrier, such as one or more non-transitory machine-readable media or storage devices, for execution by or to control the operation of one or more data processing devices, e.g., a programmable processor, a computer, multiple computers, and / or programmable logic components).
[0098] The computer program may be written in any form of programming language, including compiled or interpreted languages, and may be arranged in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may be arranged to be executed on one computer, on multiple computers at one site, or distributed across multiple sites and interconnected by a network.
[0099] The operations associated with implementing all or a portion of the functionality may be performed by one or more programmable processors executing one or more computer programs to perform the functions of reference signal selection or combination. All or a portion of the functionality may be implemented as special purpose logic circuitry, such as an FPGA and / or an ASIC (application-specific integrated circuit).
[0100] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory, a random access memory, or both. Elements of a computer include a processor for executing instructions and one or more memory devices for storing instructions and data.
[0101] While several embodiments of the present invention have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions and / or achieving the results, and / or advantages of one or more of the present inventions described herein, and each of such variations and / or modifications is deemed to be within the scope of the embodiments of the present invention described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications for which the teachings of the present invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the present invention described herein. Accordingly, it should be understood that the foregoing embodiments are presented by way of example only, and that, within the scope of the appended claims and their equivalents, embodiments of the present invention may be practiced otherwise than as specifically described and claimed. The inventive embodiments of the present disclosure relate to each individual feature, system, article, material, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials and / or methods, if such features, systems, articles, materials and / or methods are not mutually inconsistent, is included within the inventive scope of this disclosure.
Claims
1. 1. A method for preparing a reference signal for an echo cancellation system located in a vehicle, comprising: receiving a plurality of drive signals, each drive signal provided to an associated transducer of a plurality of acoustic transducers, the associated acoustic transducer converting the drive signal into an acoustic signal, each of the plurality of acoustic transducers positioned within the vehicle such that each acoustic signal is audible within a passenger compartment of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located within the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; outputting the summed reference signal to an echo cancellation system; summing a second subset of the plurality of filtered signals to generate a second summed reference signal; outputting the second summed reference signal to the echo cancellation system; A method comprising:
2. A method for preparing a reference signal for an echo cancellation system located in a vehicle, comprising: receiving a plurality of drive signals, each drive signal provided to an associated transducer of a plurality of acoustic transducers, the associated acoustic transducer converting the drive signal into an acoustic signal, each of the plurality of acoustic transducers positioned within the vehicle such that each acoustic signal is audible within a passenger compartment of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located within the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; outputting the summed reference signal to an echo cancellation system; Including, The method, wherein the plurality of filters are selected from a set of filters according to conditions within a vehicle.
3. The method of claim 2 , wherein the state is determined according to at least one of seat position, window position, number of occupants, occupant position, and door position.
4. 1. A method for preparing a reference signal for an echo cancellation system located in a vehicle, comprising: receiving a plurality of drive signals, each drive signal being provided to an associated transducer of a plurality of acoustic transducers disposed within the vehicle, the associated acoustic transducer converting the drive signal into an acoustic signal; separating each of the plurality of drive signals into a plurality of frequency sub-bands; providing a first selection of the plurality of drive signals as a reference signal to the echo cancellation system for a first frequency subband of the plurality of frequency subbands, the first selection being obtained from a subset of the plurality of drive signals; providing a second selection of the plurality of drive signals to the echo cancellation system as a reference signal for a second frequency subband of the plurality of frequency subbands, the second selection being obtained from a subset of the plurality of drive signals, the first selected subset and the second selected subset differing by at least one drive signal; A method comprising:
5. 5. The method of claim 4, wherein at least one of the first selection of the plurality of drive signals and at least one of the second selection of the plurality of drive signals are summed before being provided as a reference signal to the echo cancellation system.
6. The method of claim 4 , wherein the first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
7. 5. The method of claim 4, wherein the first selection of the plurality of drive signals is filtered with a respective filter of a plurality of filters before being provided to the echo cancellation system as a reference signal, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located in the vehicle such that each of the first selection of the plurality of drive signals estimates a respective acoustic signal in the first frequency subband at the microphone.
8. The method of claim 7 , wherein the first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
9. The method of claim 7 , wherein the plurality of filters are selected from a set of filters according to conditions within a vehicle.
10. 1. A non-transitory storage medium storing program code that, when executed by a processor, prepares a reference signal for an echo cancellation system located in a vehicle, the program code, when executed, comprising: receiving a plurality of drive signals, each drive signal provided to an associated transducer of a plurality of acoustic transducers, the associated acoustic transducer converting the drive signal into an acoustic signal, each of the plurality of acoustic transducers positioned within the vehicle such that each acoustic signal is audible within a passenger compartment of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located within the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; outputting the summed reference signal to an echo cancellation system; summing a second subset of the plurality of filtered signals to generate a second summed reference signal; outputting the second summed reference signal to the echo cancellation system; Non-transitory storage media, including 11. A non-transitory storage medium storing program code that, when executed by a processor, prepares a reference signal for an echo cancellation system located in a vehicle, the program code, when executed, performing: receiving a plurality of drive signals, each drive signal provided to an associated transducer of a plurality of acoustic transducers, the associated acoustic transducer converting the drive signal into an acoustic signal, each of the plurality of acoustic transducers positioned within the vehicle such that each acoustic signal is audible within a passenger compartment of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located within the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; outputting the summed reference signal to an echo cancellation system; Including, The plurality of filters are selected from a set of filters according to conditions within the vehicle.
12. The non-transitory storage medium of claim 11 , wherein the state is determined according to at least one of seat position, window position, number of occupants, occupant position, and door position.
13. 1. A non-transitory storage medium storing program code that, when executed by a processor, prepares a reference signal for an echo cancellation system located in a vehicle, the program code, when executed, comprising: receiving a plurality of drive signals, each drive signal being provided to an associated transducer of a plurality of acoustic transducers disposed within the vehicle, the associated acoustic transducer converting the drive signal into an acoustic signal; separating each of the plurality of drive signals into a plurality of frequency sub-bands; providing a first selection of the plurality of drive signals as a reference signal to the echo cancellation system for a first frequency subband of the plurality of frequency subbands, the first selection being obtained from a subset of the plurality of drive signals; providing a second selection of the plurality of drive signals to the echo cancellation system as a reference signal for a second frequency subband of the plurality of frequency subbands, the second selection being obtained from a subset of the plurality of drive signals, the first selected subset and the second selected subset differing by at least one drive signal; Non-transitory storage media, including
14. 14. The non-transitory storage medium of claim 13, wherein at least one of the first selection of the plurality of drive signals and at least one of the second selection of the plurality of drive signals are summed before being provided as a reference signal to the echo cancellation system.
15. 14. The non-transitory storage medium of claim 13, wherein the first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
16. 14. The non-transitory storage medium of claim 13, wherein the first selection of the plurality of drive signals is filtered with a respective filter of a plurality of filters before being provided to the echo cancellation system as a reference signal, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone disposed in the vehicle such that each of the first selection of the plurality of drive signals estimates a respective acoustic signal in the first frequency subband at the microphone.
17. 17. The non-transitory storage medium of claim 16, wherein the first selected subset of the plurality of drive signals are summed before being provided to the echo cancellation system.
18. The non-transitory storage medium of claim 16 , wherein the plurality of filters are selected from a set of filters according to conditions within a vehicle.
19. A method for preparing a reference signal for an echo cancellation system located in a vehicle, comprising: receiving a plurality of drive signals, each drive signal provided to an associated transducer of a plurality of acoustic transducers, the associated acoustic transducer converting the drive signal into an acoustic signal, each of the plurality of acoustic transducers positioned within the vehicle such that each acoustic signal is audible within a passenger compartment of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located within the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; outputting the summed reference signal to an echo cancellation system; Including, the subset is formed by selecting the plurality of filtered signals corresponding to drive signals selected to maximize total multicoherence between the drive signals and microphone signals, or by selecting the plurality of filtered signals corresponding to drive signals of speakers grouped based on speaker positions within the vehicle. method.
20. A non-transitory storage medium storing program code that, when executed by a processor, prepares a reference signal for an echo cancellation system located in a vehicle, the program code, when executed, performing: receiving a plurality of drive signals, each drive signal provided to an associated transducer of a plurality of acoustic transducers, the associated acoustic transducer converting the drive signal into an acoustic signal, each of the plurality of acoustic transducers positioned within the vehicle such that each acoustic signal is audible within a passenger compartment of the vehicle; filtering each drive signal with a respective filter of a plurality of filters to generate a plurality of filtered signals, each of the plurality of filters approximating a transfer function from an associated acoustic transducer to a microphone located within the vehicle, whereby each of the plurality of filtered signals estimates a respective acoustic signal at the microphone; summing at least a subset of the plurality of filtered signals to generate a summed reference signal; outputting the summed reference signal to an echo cancellation system; Including, the subset is formed by selecting the plurality of filtered signals corresponding to drive signals selected to maximize total multicoherence between the drive signals and microphone signals, or by selecting the plurality of filtered signals corresponding to drive signals of speakers grouped based on speaker positions within the vehicle. Non-transitory storage medium.
Citation Information
Patent Citations
Echo erasing method for subband multichannel voice communication conference
JP1998093680A
Echo suppressor, echo suppressing method, echo suppressor program, and its record medium
JP2006246397A
Echo canceler
JP2012070385A
Systems and methods for canceling road noise in a microphone signal
US10839786B1
Systems and methods for canceling echo in a microphone signal
US20200413191A1