Directionality induced robust acoustic echo canceller adaptation

By using a combination of parallel and orthogonal domain filter weights in the vehicle audio system, the erroneous convergence problem of the acoustic echo canceller in the double-speech scenario is solved, improving speech quality and simplifying double-speech detection, achieving faster convergence and better audio signal processing.

CN120977327APending Publication Date: 2025-11-18GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410934503.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-15
Filing Date
2024-07-12
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In vehicle audio systems, acoustic echo cancellers may converge to incorrect solutions in dual-speech scenarios, resulting in artifacts on the output such as cancellation of the desired signal, musical pitch, and reverberation effects.

Method used

By employing a combination of parallel and orthogonal domain filter weights, the input signal is divided into parallel domain signals and orthogonal domain signals through the vehicle control module. Appropriate step sizes are selected to adapt the filter weights. Combined with beamformer parameters and tuning state parameters, an acoustic echo canceller operation is performed to reduce residual echo.

Benefits of technology

It effectively suppresses erroneous convergence of acoustic echo cancellers, improves speech quality, simplifies dual speech detection, and achieves faster convergence and better audio signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977327A_ABST
    Figure CN120977327A_ABST
Patent Text Reader

Abstract

The invention provides a robust acoustic echo canceller adaptation caused by directivity. An example vehicle audio system includes a vehicle speaker configured to generate audio, a plurality of microphones, and a vehicle control module configured to receive an input signal via the plurality of microphones, divide the input signal into a parallel domain signal and an orthogonal domain signal, selecting a constant step value for the orthogonal domain filter weight and a variable step value for the parallel domain filter weight, adapting the orthogonal domain filter weight according to the orthogonal domain signal and the constant step value, and adapting the parallel domain filter weight according to the parallel domain signal and the variable step value, the adapted orthogonal domain filter weights and the adapted parallel domain filter weights are grouped to define a total filter weight, and the total filter weight is applied to the received input signal to perform a signal processing operation on the received input signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The information provided in this section is for the purpose of generally presenting the context of the disclosure. The work of the presently named inventors, to the extent the work is described in this section, as well as aspects of the description that can not otherwise qualify as prior art at the time of

[0002] The present disclosure relates generally to vehicle audio systems, including audio signal processing using parallel and orthogonal domain filter weights. BACKGROUND

[0003] During a telephone call, a double talk scenario can occur when both a far-end talker and a near-end talker are speaking at the same time. During a double talk event, an acoustic echo canceller in the speech processing chain can converge to a wrong solution, which can cause artifacts such as cancellation of the desired signal, musical tones, and reverberation effects on the output of the acoustic echo canceller, where the original sound is heard together with a delayed version of the same sound. SUMMARY

[0004] An example vehicle audio system includes at least one vehicle speaker configured to generate audio within a vehicle interior, a plurality of microphones each configured to obtain audio within the vehicle interior, and a vehicle control module configured to receive an input signal via the plurality of microphones, separate the input signal received via the plurality of microphones into a parallel domain signal and an orthogonal domain signal, select a constant step size value for the orthogonal domain filter weights and a variable step size value for the parallel domain filter weights, adapt the orthogonal domain filter weights according to the orthogonal domain signal and the constant step size value, adapt the parallel domain filter weights according to the parallel domain signal and the variable step size value, combine the adapted orthogonal domain filter weights and the adapted parallel domain filter weights to define total filter weights, and apply the total filter weights to the received input signal to perform a signal processing operation on the received input signal prior to an audio output of the received input signal.

[0005] In some examples, the vehicle control module is configured to control the at least one vehicle speaker to output an audio signal based on the input signal modified by the total filter weights.

[0006] In some examples, the vehicle control module is configured to separate the input signal into the parallel domain signal and the orthogonal domain signal by applying a parallel projection to the input signal received via the plurality of microphones and applying an orthogonal projection to the input signal received via the plurality of microphones.

[0007] In some examples, the vehicle control module is configured to obtain a source signal steering vector from at least one of the beamformer parameters and the specified tuning state parameters, and to calculate a parallel projection and an orthogonal projection based on the source signal steering vector.

[0008] In some examples, the parallel projection is defined as parallel to a target near-end audio source. In some examples, the orthogonal projection is defined as orthogonal to the target near-end audio source.

[0009] In some examples, the vehicle control module is configured to apply a greater variable step size value during a first time period in which a double talk condition exists in the input signal than a second time period in which the double talk condition exists in the input signal.

[0010] In some examples, the vehicle control module is configured to perform an acoustic echo canceller (AEC) operation to determine a residual echo value by subtracting a product of the total filter weights and a reference signal from the input signal received via the plurality of microphones.

[0011] In some examples, the vehicle control module is configured to adapt the orthogonal domain filter weights and the parallel domain filter weights using at least one of normalized least mean squares (NLMS), recursive least squares (RLS), or affine projection. In some examples, the plurality of microphones are arranged in a linear array within the vehicle.

[0012] An example method of processing audio signals inside a vehicle includes receiving, by a vehicle control module, an input signal from a plurality of microphones, each of the plurality of microphones configured to obtain audio inside the vehicle; separating the input signal received via the plurality of microphones into a parallel domain signal and an orthogonal domain signal; selecting a constant step size value for orthogonal domain filter weights and a variable step size value for parallel domain filter weights; adapting the orthogonal domain filter weights from the orthogonal domain signal and the constant step size value; adapting the parallel domain filter weights from the parallel domain signal and the variable step size value; combining the adapted orthogonal domain filter weights and the adapted parallel domain filter weights to define total filter weights; and applying the total filter weights to the received input signal to perform a signal processing operation on the received input signal prior to an audio output of the received input signal.

[0013] In some examples, the method includes controlling at least one vehicle loudspeaker to output an audio signal based on the input signal modified by the total filter weights.

[0014] In some examples, separating the input signal into the parallel domain signal and the orthogonal domain signal includes applying a parallel projection to the input signal received via the plurality of microphones and applying an orthogonal projection to the input signal received via the plurality of microphones.

[0015] In some examples, the method includes obtaining a source signal steering vector from at least one of the beamformer parameters and the specified tuning state parameters, and computing a parallel projection and an orthogonal projection based on the source signal steering vector.

[0016] In some examples, the parallel projection is defined as parallel to the target near-end audio source. In some examples, the orthogonal projection is defined as orthogonal to the target near-end audio source.

[0017] In some examples, adapting the parallel-domain filter weights includes applying a larger variable step size value during a first time period in which the double talk condition exists in the input signal than a second time period in which the double talk condition exists in the input signal.

[0018] In some examples, the method includes performing an acoustic echo canceller (AEC) operation to determine a residual echo value by subtracting a product of the total filter weights and a reference signal from an input signal received via the plurality of microphones.

[0019] In some examples, adapting the orthogonal-domain filter weights and adapting the parallel-domain filter weights includes using at least one of normalized least mean squares (NLMS), recursive least squares (RLS), or affine projection. In some examples, the plurality of microphones are arranged in a linear array within the vehicle.

[0020] The following solutions are provided:

[0021] 1. A vehicle audio system comprising:

[0022] at least one vehicle loudspeaker configured to generate audio inside the vehicle;

[0023] a plurality of microphones each configured to obtain audio inside the vehicle; and

[0024] a vehicle control module configured to:

[0025] receive an input signal via the plurality of microphones;

[0026] divide the input signal received via the plurality of microphones into a parallel-domain signal and an orthogonal-domain signal;

[0027] select a constant step size value for the orthogonal-domain filter weights and a variable step size value for the parallel-domain filter weights;

[0028] adapt the orthogonal-domain filter weights from the orthogonal-domain signal and the constant step size value;

[0029] adapt the parallel-domain filter weights from the parallel-domain signal and the variable step size value;

[0030] combining the adapted orthogonal domain filter weights and the adapted parallel domain filter weights to define total filter weights; and

[0031] applying the total filter weights to the received input signal to perform a signal processing operation on the received input signal prior to audio output of the received input signal.

[0032] 2. The vehicle audio system of Scheme 1, wherein the vehicle control module is configured to control the at least one vehicle loudspeaker to output an audio signal based on the input signal modified by the total filter weights.

[0033] 3. The vehicle audio system of Scheme 1, wherein the vehicle control module is configured to separate the input signal into a parallel domain signal and an orthogonal domain signal by:

[0034] applying a parallel projection to the input signal received via the plurality of microphones; and

[0035] applying an orthogonal projection to the input signal received via the plurality of microphones.

[0036] 4. The vehicle audio system of Scheme 3, wherein the vehicle control module is configured to:

[0037] obtaining a source signal steering vector as a function of at least one of the beamformer parameters and the specified tuning state parameter; and

[0038] calculating the parallel projection and the orthogonal projection based on the source signal steering vector.

[0039] 5. The vehicle audio system of Scheme 3, wherein the parallel projection is defined as parallel to a target near-end audio source.

[0040] 6. The vehicle audio system of Scheme 3, wherein the orthogonal projection is defined as orthogonal to the target near-end audio source.

[0041] 7. The vehicle audio system of Scheme 1, wherein the vehicle control module is configured to apply a greater variable step size value during a first time period in which a double talk condition exists in the input signal compared to a second time period in which the double talk condition exists in the input signal.

[0042] 8. The vehicle audio system of Scheme 1, wherein the vehicle control module is configured to perform an acoustic echo canceller (AEC) operation to determine a residual echo value by subtracting a product of the total filter weights and a reference signal from the input signal received via the plurality of microphones.

[0043] 9. The vehicle audio system of paragraph 1, wherein the vehicle control module is configured to adapt the orthogonal domain filter weights and adapt the parallel domain filter weights using at least one of normalized least mean squares (NLMS), recursive least squares (RLS), or affine projection.

[0044] 10. The vehicle audio system of paragraph 1, wherein the plurality of microphones are arranged in a linear array within the vehicle.

[0045] 11. A method of processing interior vehicle audio signals, the method comprising:

[0046] receiving, by a vehicle control module, input signals from a plurality of microphones, each of the plurality of microphones configured to obtain audio of an interior of a vehicle;

[0047] dividing the input signals received via the plurality of microphones into parallel domain signals and orthogonal domain signals;

[0048] selecting a constant step size value for the orthogonal domain filter weights and a variable step size value for the parallel domain filter weights;

[0049] adapting the orthogonal domain filter weights according to the orthogonal domain signals and the constant step size value;

[0050] adapting the parallel domain filter weights according to the parallel domain signals and the variable step size value;

[0051] combining the adapted orthogonal domain filter weights and the adapted parallel domain filter weights to define total filter weights; and

[0052] applying the total filter weights to the received input signals to perform a signal processing operation on the received input signals prior to audio output of the received input signals.

[0053] 12. The method of paragraph 11, further comprising controlling at least one vehicle speaker to output an audio signal based on the input signals modified by the total filter weights.

[0054] 13. The method of paragraph 11, wherein dividing the input signals into parallel domain signals and orthogonal domain signals comprises:

[0055] applying a parallel projection to the input signals received via the plurality of microphones; and

[0056] applying an orthogonal projection to the input signals received via the plurality of microphones.

[0057] 14. The method of paragraph 13, further comprising:

[0058] obtaining a source signal steering vector according to at least one of a beamformer parameter and a specified tuning state parameter; and

[0059] Parallel and orthogonal projections are calculated based on the source signal steering vector.

[0060] 15. The method according to Scheme 13, wherein parallel projection is defined as parallel to the target near-end audio source.

[0061] 16. The method according to Scheme 13, wherein orthogonal projection is defined as orthogonal to the target near-end audio source.

[0062] 17. The method according to Scheme 11, wherein the adaptive parallel domain filter weights include: applying a larger variable step size value during the first time period in which the dual speech condition exists in the input signal compared to the second time period in which the dual speech condition exists in the input signal.

[0063] 18. The method of claim 11 further includes performing an acoustic echo canceller (AEC) operation to determine the residual echo value by subtracting the product of the total filter weight and the reference signal from the input signal received via multiple microphones.

[0064] 19. The method according to Scheme 11, wherein adapting the orthogonal domain filter weights and adapting the parallel domain filter weights includes using at least one of Normalized Least Mean Square (NLMS), Recursive Least Squares (RLS), or Affine Projection.

[0065] 20. The method according to claim 11, wherein a plurality of microphones are arranged in a linear array within the vehicle.

[0066] Further applications of this disclosure will become apparent from the detailed description, claims, and accompanying drawings. The detailed description and specific examples are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description

[0067] This disclosure will become more fully understood in light of the detailed description and accompanying drawings.

[0068] Figure 1 This is a diagram of an example vehicle that includes a vehicle audio system.

[0069] Figure 2 This is a block diagram depicting an example signal processing system that includes an audio echo canceller.

[0070] Figure 3 It is a description Figure 2 A block diagram of an example signal in a signal processing system.

[0071] Figure 4 It is a flowchart depicting an example process of performing audio signal processing using parallel domain signals and orthogonal domain signals.

[0072] Figure 5 is a flowchart depicting an example process for determining a variable step size value for a process of Figure 4 is a flowchart of an example process for determining a variable step size value for a process of

[0073] In the drawings, reference numerals can be repeated among the figures for like and / or identical elements. DETAILED DESCRIPTION

[0074] In some examples, a signal processing chain can include an acoustic echo canceller (AEC), e.g., for processing audio signals associated with a vehicle interior. During a telephone call, a double talk (DT) scenario occurs when both a far-end (FE) talker and a near-end (NE) talker speak at the same time. During a double talk scenario, adaptation provided by an AEC adaptive filter (AF) can be reduced or stopped to inhibit or prevent the adaptive filter from converging to a false solution. This can potentially result in sub-optimal convergence, leading to a false solution that can cause artifacts such as cancellation of desired signals, musical tones, and reverberation effects on the output of the AEC.

[0075] In some example embodiments, knowledge of the desired near-end source location can be used to split the AEC filter into two domains, one domain parallel to the desired source and the other domain orthogonal to the desired near-end source. Since the orthogonal domain signal can not include any desired speech (e.g., due to the domain being orthogonal to the near-end source), the orthogonal domain can be adapted with little or no restriction even during DT conditions. In the parallel domain, double talk detection can be simpler relative to the original domain due to higher signal echo ratio (SER) levels in the parallel domain (e.g., because the orthogonal domain signal has been separated from the parallel domain signal).

[0076] Reference is now made to Figure 1 The vehicle 10 comprises front wheels 12 and rear wheels 13. In Figure 1 The drive unit 14 selectively outputs torque to the front wheels 12 and / or the rear wheels 13 via drive lines 16, 18, respectively. The vehicle 10 can comprise different types of drive units. For example, the vehicle can be an electric vehicle, such as a battery electric vehicle (BEV), a hybrid vehicle, or a fuel cell vehicle, a vehicle comprising an internal combustion engine (ICE), or other types of vehicles.

[0077] Some examples of the drive unit 14 can comprise any suitable electric motor, power inverter, and motor controller configured to control power switches within the power inverter to regulate motor speed and torque during propulsion and / or regeneration. During propulsion or regeneration, a battery system provides power to or receives power from the electric motor of the drive unit 14 via the power inverter.

[0078] While in Figure 1The middle vehicle 10 includes one drive unit 14, although the vehicle 10 can have other configurations. For example, two separate drive units can drive the front wheels 12 and the rear wheels 13, one or more separate drive units can drive separate wheels, etc. It will be appreciated that other vehicle configurations and / or drive units can be used.

[0079] The vehicle control module 20 can be configured to control operation of one or more vehicle components, such as the drive unit 14 (e.g., by commanding a torque setting of the electric motor of the drive unit 14). The vehicle control module 20 can receive inputs for controlling the vehicle components, such as signals received from a steering wheel, an acceleration pedal, a brake pedal, etc. The vehicle control module 20 can monitor telematics of the vehicle for safety purposes, such as vehicle speed, vehicle location, vehicle braking and acceleration, etc.

[0080] The vehicle control module 20 can receive signals from any suitable components for monitoring one or more aspects of the vehicle, including one or more vehicle sensors (e.g., a camera, a microphone, a pressure sensor, a steering wheel position sensor, a brake sensor, a location sensor such as a global positioning system (GPS) antenna, a wheel height and / or position sensor, an accelerometer, etc.). Some sensors can be configured to monitor a current motion of the vehicle, an acceleration of the vehicle, a braking of the vehicle, a current steering direction of the vehicle, a current height and / or position of one or more wheels, etc.

[0081] In some examples, the vehicle microphones 22 are configured to capture audio signals from inside the vehicle 10. For example, multiple microphones (e.g., at least two microphones, at least four microphones, at least eight microphones, etc.) can be arranged inside the vehicle 10, one or more devices inside the vehicle 10 can include a microphone, etc. In some examples, the vehicle microphones 22 can be arranged in a linear array, and can include any suitable microphone structure or component suitable for capturing and transmitting audio signals (e.g., converting acoustic mechanical audio signals into electrical signals).

[0082] The vehicle 10 includes multiple vehicle speakers 24, which can be configured to generate audio signals inside the vehicle 10. For example, a passenger or driver of the vehicle can use the vehicle microphones 22 and the vehicle speakers 24 to make a call (e.g., a hands-free telephone call), where the vehicle microphones 22 capture speech of the passenger or driver, and the vehicle speakers 24 generate audio signals (which can be referred to as far-end (FE) signals) based on speech of another person on the other end of the telephone call.

[0083] The vehicle control module 20 can communicate with another device through a wireless communication interface, which can include one or more wireless antennas for transmitting and / or receiving wireless communication signals. For example, the wireless communication interface can communicate via any suitable wireless communication protocol, including but not limited to vehicle-to-everything (V2X) communication, Wi-Fi communication, wireless area network (WAN) communication, cellular communication, personal area network (PAN) communication, short-range wireless communication (e.g., Bluetooth), etc. The wireless communication interface can communicate with a remote computing device through one or more wireless and / or wired networks. With respect to vehicle-to-vehicle (V2X) communication, the vehicle 10 can include one or more V2X transceivers (e.g., V2X signal transmitting and / or receiving antennas).

[0084] Figure 2 is a block diagram depicting an example signal processing system 200 including an acoustic echo canceller 202. A call can occur between a far-end audio source 204 (e.g., a loudspeaker at the other end of a telephone call) and a near-end audio source 206 (e.g., a vehicle occupant such as a driver or passenger). Although described with reference to a vehicle, Figure 2 other examples can be used in other non-vehicle settings.

[0085] As Figure 2 shown, the system 200 can include a loudspeaker 208 (e.g., a vehicle interior loudspeaker) configured to generate an audio signal based on an input signal from the far-end audio source 204. For example, speech from a far-end talker can be converted by the loudspeaker 208 from an electrical signal to an acoustic mechanical sound signal that is audible to a vehicle occupant.

[0086] The system 200 includes a plurality of microphones 210. The microphones 210 can be configured to obtain audio signals from within the vehicle. For example, the microphones 210 can pick up audio from the near-end audio source 206 and convert acoustic mechanical sounds (e.g., vehicle occupant speech) to electrical audio signals.

[0087] The microphones 210 can include any suitable microphone components and arrangement, such as a linear array of microphones. Although Figure 2 an array of four microphones is shown, other embodiments can include more or fewer microphones (e.g., at least two microphones, at least eight microphones, etc.), and the microphones can be in different arrangements relative to one another and within the vehicle.

[0088] As Figure 2 shown, the microphones 210 can capture audio from the loudspeaker 208. For example, the microphones 210 can be primarily designed to pick up speech from the near-end audio source 206 (e.g., a vehicle occupant loudspeaker), but other sounds can also be introduced into the vehicle interior that are not intended to be captured by the microphones 210.

[0089] As Figure 2 shown, a voice processing chain can be used to at least partially remove the audio signal from the loudspeaker 208 in the signal received at the microphone 210. The voice processing chain can include any suitable voice processing elements, such as an acoustic echo canceller 202, a beamformer 212, etc. In some example embodiments, these components can be part of a vehicle control module.

[0090] Figure 3 is a block diagram of an example signal in a signal processing system depicting Figure 2 In the example diagram of Figure 3 , Z represents the output residual echo, D is the input signal (e.g., received from the far-end audio source 204), X is the reference signal, and W represents the adaptive filter weights. In this example, the AEC operation can be:

[0091] Z = D - W H X

[0092] When there is no active speech from the near-end audio source 206, the AEC can adapt the adaptive filter weights W using the following equation:

[0093] W opt = arg min W E{|Z| 2} = R -1 P

[0094] where R is the reference autocorrelation matrix R = E{XX H} and P is the cross-correlation matrix P = E{XD H}.

[0095] In some examples, the optimal adaptive filter weights W opt can be split into two domains W || and W ⊥ , W || is the domain parallel to the near-end audio source 206, and W ⊥ is the domain orthogonal to the near-end audio source 206. For example:

[0096] W opt = W || + W ⊥

[0097] To adapt separately, the parallel and orthogonal projections T || and T ⊥ can be applied to the input signal D:

[0098] D || = T || D; D ⊥ = T ⊥ D

[0099] where T || and T ⊥ are defined as the parallel and orthogonal projection functions. The steering vector S of the desired source can be used to construct the parallel and orthogonal projection T || and T ⊥ . The near-end desired steering signal S can be obtained in any suitable manner, e.g., using available parameters of the beamformer 212, using calculations from a pre-tuning phase, etc. An example of the parallel and orthogonal projection matrices for a single-rank steering vector S is:

[0100] T || = SS H / S H S; T ⊥ = I - T ||

[0101] Any suitable adaptive algorithm can be used to solve for W || and W ⊥ , e.g., normalized least mean square (NLMS), recursive least square (RLS), affine projection algorithm (APA), etc. Using NLMS as an example, the adaptive iteration step can be:

[0102] W(n + l) = W(n) + μX(n)Z * (n)‖X(n)‖ 2

[0103] where μ is the adaptive step size, also known as the learning rate. The learning rate determines the duration of an epoch. Because the near-end desired signal can not be present in D ⊥ , the following example equation can be used to solve for W ⊥ with a constant adaptive step size μ ⊥ , enabling fast and deep convergence:

[0104]

[0105] Because the near-end signal can be present in D || , a variable step size (VSS) μ || (n) can be used to solve for W || to prevent W || from diverging:

[0106]

[0107] In some examples, for situations where double talk is not present, a higher μ || (n) value can be selected or used, enabling fast convergence. When double talk is present, a lower μ || (n) value can be selected or used.(n) values and adjusting μ as appropriate || (n) values to avoid filter divergence. μ || (n) values can be estimated in any suitable manner, such as based on X(n) energy levels, X(n) correlation with D || (n), directionality of Z || (n), etc.

[0108] Due to the projection, the echo level in D || may be lower and the SER level can be higher when compared to D || , making it easier to estimate μ || (n). This can reduce the number of false detections, enable faster convergence, and produce better speech quality. These concepts can also be used in other adaptive algorithm processes.

[0109] The adaptive algorithm, optimization criteria, rule set, etc. can be selected based on the suitability of the orthogonal and parallel models. For example, these options can be selected based on the system identification method. For example, if these features fit an autoregressive moving average (ARMA) model, a more suitable rule set can be selected based on the system parameterization.

[0110] Referring again to Figure 3 , element G can be an impulse response. For example, in an ideal case, the adaptive filter weights W can be G. The reference signal X can represent the sound played in the loudspeaker 208, and the system 200 can be used to attempt to cancel the reference signal X (e.g., intended to cancel the reference signal). For example, the reference signal X can be multiplied by the adaptive filter weights W, resulting in a signal Y that can then cancel the echo signal E present in D

[0111] Although element G can be unknown, the system 200 can be configured to solve for element G. S can represent an acoustic function between the near-end audio source 206 and the microphone 210, and can not include the speech itself. For example, the acoustic function S can depend on the position of the speaker’s mouth relative to the microphone 210.

[0112] Figure 4 is a flowchart depicting an example process for performing audio signal processing using parallel and orthogonal domain signals. The process can be performed by, for example, the vehicle control module 20 of Figure 1 . At 304, the process begins by receiving an input signal via a plurality of microphones, such as the vehicle microphones 22 of the vehicle 10 in Figure 1 , or the microphone 210 in Figure 2 and 3 .

[0113] At 408, control separates the input signal into a parallel domain signal and a quadrature domain signal. Control then selects a constant step size value for the quadrature domain filter weights and a variable step size value for the parallel domain filter weights at 412.

[0114] At 416, the vehicle control module is configured to adapt the quadrature domain filter weights as a function of the quadrature domain signal and the constant step size value. The above references Figure 2 and Figure 3 describe example details of adapting the quadrature domain filter weights.

[0115] At 420, the vehicle control module is configured to adapt the parallel domain filter weights as a function of the parallel domain signal and the variable step size value. The above references Figure 2 and Figure 3 describe further details regarding adapting the parallel domain filter weights, and the below references Figure 5 further describe these details.

[0116] At 424, control is configured to combine the adapted quadrature and parallel domain filter weights to define overall filter weights. Then, at 428, the overall filter weights are applied to a received input signal. For example, the vehicle control module can be configured to generate an output audio signal based on the received input signal after modifying the input signal based on the adapted filter weights.

[0117] Figure 5 is a flowchart depicting an example process for determining a variable step size value for a process of Figure 4 The process can be performed by, for example, the vehicle control module 20 of Figure 1 At 504, the process begins by obtaining a parallel domain filter weight value for a current time step (e.g., a value that is currently being used for adaptive filtering on a parallel domain signal).

[0118] Control then determines whether a double talk audio signal condition exists at 508 (e.g., by determining or detecting double talk using any suitable sensor and / or signal processing techniques). If a double talk condition exists at 508, control proceeds to 512 to select a first value for the variable step size.

[0119] If a double talk condition does not exist at 508, control proceeds to 516 to select a second value for the variable step size that is greater than the first value. In this way, the vehicle control module can use a smaller step size for adaptation when a double talk condition exists and a larger step size for adaptation when a double talk condition does not exist. At 520, control calculates an updated parallel domain filter weight value using the selected step size value.

[0120] As described above, in some examples, the input signal is split into two domains (orthogonal and parallel to the desired source), and different adaptations are used in each domain. For example, in the orthogonal domain, the desired source (e.g., a near-end talker) can not be present, or can only be present with residual. When adaptation is not needed to stop, a constant learning rate step size can be used, as the desired talker is not present in the orthogonal signal.

[0121] In the parallel domain, the desired source can be present, but at different levels of echo (e.g., when the orthogonal domain signal is removed). A variable step size can be used, where a higher value is selected when the desired talker is not present, and a smaller (or zero) value is selected when the desired talker is present. The desired talker can refer to the person who is talking, where the system intends to transmit their speech, but cancel any echo. For example, the adaptive filter should ignore the desired source (e.g., a person in a vehicle talking into a microphone array), but continue to filter the echo even when the desired source is talking.

[0122] Splitting the input signal into two dimensions allows for little consideration of the desired source in the orthogonal domain (e.g., as the desired source can not be present in the orthogonal domain signal), while there should be lower echo in the parallel domain, making it easier to detect the presence of a double talk scenario.

[0123] While some example embodiments described herein relate to a vehicle interior and vehicle microphones and speakers, the signal processing techniques can be used in other suitable settings, such as a room or playing music on a speaker while trying to talk with a smartphone (e.g., where it is desired to cancel the music when the smartphone hears the speech of the talker). In some examples, the techniques described herein can modify the adaptation learning rate while continuously or as desired applying filtering to the input signal. The adaptive filter can use polynomial fitting, weights that fit a recursive equation, etc.

[0124] The foregoing description is merely illustrative in nature and is not intended to limit the disclosure, its application or uses. The broad teachings of the disclosure can be implemented in a variety of forms. Therefore, while this disclosure includes particular examples, the true scope of the disclosure should not be so limited since other modifications will become apparent upon a study of the drawings, specification, and claims. It should be understood that one or more steps within a method can be executed in different order (or concurrently) without altering the principles of the

[0125] Spatial and functional relationships between elements (for example, between modules, circuit elements, semiconductor layers, etc.) are described using various terms, including "connected," "engaged," "coupled," "adjacent," "next to," "on top of," "over," "under," and "disposed." Unless explicitly described as being "direct," a relationship between a first and a second element described in the above disclosure can also be an indirect relationship where one or more other intervening elements are present (spatially or functionally) between the first and second elements. As used herein, the phrase at least one of A, B, and C should be construed to mean a logical A OR B OR C using the non-exclusive logical OR set operator. It should be interpreted also to mean at least one of A and at least one of B and at least one of C.

[0126] In the drawings, the direction indicated by an arrow of an arrowhead generally indicates the flow of information (such as data or instructions) of interest in the illustration. For example, when elements A and B exchange a variety of information, but the information transmitted from element A to element B is relevant to the illustration, an arrow can be directed from element A to element B. This one-way arrow does not mean that no other information is transmitted from element B to element A. Also, for the information transmitted from element A to element B, element B can send a request for the information or a reception acknowledgement to element A.

[0127] In this application, including the following claims, the term "module" or the term "controller" can be replaced with the term "circuit." The term "module" can refer to or include portions of an application-specific integrated circuit (ASIC), a digital, analog, or mixed analog / digital discrete circuit; a digital, analog, or mixed analog / digital integrated circuit; a combination of combinations of them; an execution code of a processor circuit (shared, dedicated, or group) that executes the code; a memory circuit (shared, dedicated, or group) that stores execution code of the processor circuit; other suitable hardware components that provide the described function; or a combination of some or all of the above, such as in a system on a chip.

[0128] A module can include one or more interface circuits. In some examples, the interface circuits can include wired or wireless interfaces that are connected to a local area network (LAN), the Internet, a wide area network (WAN), or combinations thereof. The functionality of any given module of the present disclosure can be distributed among multiple modules that are connected via interface circuits. For example, a plurality of modules can allow load balancing. In another example, a server (also known as remote, or cloud) module can accomplish some functionality on behalf of a client module.

[0129] The term code, as used above, can include software, firmware, and / or microcode, and can refer to programs, routines, functions, classes, data structures, and / or objects. The term shared processor circuitry encompasses a single processor circuitry executing some or all code from multiple modules. The term group processor circuitry encompasses a processor circuitry combined with an additional processor circuitry executing some or all code from one or more modules. A reference to multiple processor circuitries encompasses multiple processor circuitries on discrete dies, multiple processor circuitries on a single die, multiple cores of a single processor circuitry, multiple threads of a single processor circuitry, or a combination of one or more of the above. The term shared memory circuitry encompasses a single memory circuitry storing some or all code from multiple modules. The term group memory circuitry encompasses a memory circuitry combined with an additional memory storing some or all code from one or more modules.

[0130] The term memory circuitry is a subset of the term computer readable medium. The term computer readable medium, as used herein, does not encompass transitory propagating signals per se (e.g., electromagnetic waves propagating through a medium or over a medium, such as on a carrier or the like) The term computer readable medium can therefore be considered tangible and non-transitory. Non-limiting examples of non-transitory, tangible computer readable media are nonvolatile memory circuits (e.g., flash memory, erasable programmable read only memory (EPROM), or electrically erasable programmable read only memory (EEPROM)), volatile memory circuits (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM)), magnetic storage media (e.g., magnetic tapes or magnetic hard drives), and optical storage media (e.g., optical discs, optical fiber media, and the like).

[0131] The apparatus and methods described in this application can be partly or entirely implemented by special purpose computers created by configuring general purpose computers to execute one or more specific functions embodied in computer programs. The above-described functional blocks, flowchart components, and other elements are to be understood as software specifications that can be translated into computer programs by a skilled artisan or programmer.

[0132] A computer program includes processor-executable instructions stored on at least one non-transitory computer-readable medium. A computer program can also include or rely on stored data. A computer program can encompass a Basic Input-Output system (BIOS) that interacts with hardware of the special purpose computer, device drivers that interact with particular devices of the special purpose computer, one or more operating systems, user applications, background services, background applications, etc.

[0133] A computer program can include: (i) a descriptive text to be interpreted, such as HTML (HyperText Markup Language), XML (Extensible Markup Language), or JSON (JavaScript Object Notation); (ii) assembly code; (iii) object code that is generated by a compiler from source code; (iv) source code that is executed by an interpreter; (v) source code that is compiled and executed by a just-in-time compiler; etc. By way of example only, source code can be written using a syntax from a language of the group including C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, Java®, Fortran, Perl, Pascal, Curl, OCaml, JavaScript®, HTML5 (HyperText Markup Language Version 5), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Flash®, Visual Basic®, Lua, MATLAB, SIMULINK, and SPHINX®. Fortran, Perl, Pascal, Curl, OCaml, JavaScript, HTML, CSS, DHTML, and Visual Basic, Java,.NET, C#, C++, and / or Python. Lua, MATLAB, SIMULINK, and

Claims

1. A vehicle audio system, comprising: At least one vehicle speaker is configured to generate audio inside the vehicle; Multiple microphones, each configured to acquire audio from inside the vehicle; as well as The vehicle control module is configured as follows: Input signals are received via multiple microphones; The input signal received from multiple microphones is divided into parallel domain signal and orthogonal domain signal; Choose a constant step size for the weights of the orthogonal domain filter and a variable step size for the weights of the parallel domain filter; The orthogonal domain filter weights are adapted based on the orthogonal domain signal and a constant step size value. The parallel domain filter weights are adapted based on the parallel domain signal and the variable step size value. Combine the adapted orthogonal domain filter weights and the adapted parallel domain filter weights to define the total filter weights; as well as Before the audio output of the received input signal, the total filter weights are applied to the received input signal to perform signal processing operations on the received input signal.

2. The vehicle audio system of claim 1, wherein the vehicle control module is configured to control at least one vehicle speaker to output an audio signal based on an input signal modified by the total filter weights.

3. The vehicle audio system of claim 1, wherein the vehicle control module is configured to divide the input signal into a parallel domain signal and an orthogonal domain signal by: Parallel projection is applied to input signals received via multiple microphones; and Orthogonal projection is applied to the input signal received via multiple microphones.

4. The vehicle audio system according to claim 3, wherein the vehicle control module is configured as follows: The source signal steering vector is obtained based on at least one of the beamformer parameters and specified tuning state parameters; and Parallel and orthogonal projections are calculated based on the source signal steering vector.

5. The vehicle audio system of claim 3, wherein parallel projection is defined as parallel to the target near-end audio source.

6. The vehicle audio system of claim 3, wherein orthogonal projection is defined as orthogonal to the target near-end audio source.

7. The vehicle audio system of claim 1, wherein the vehicle control module is configured to apply a larger variable step size value during the first time period in which the dual speech condition exists in the input signal, compared to the second time period in which the dual speech condition exists in the input signal.

8. The vehicle audio system of claim 1, wherein the vehicle control module is configured to perform acoustic echo canceller (AEC) operation to determine the residual echo value by subtracting the product of the total filter weight and the reference signal from the input signal received via a plurality of microphones.

9. The vehicle audio system of claim 1, wherein the vehicle control module is configured to use at least one of Normalized Least Mean Square (NLMS), Recursive Least Squares (RLS), or Affine Projection to adapt the orthogonal domain filter weights and the parallel domain filter weights.

10. A method for processing in-vehicle audio signals, the method comprising: The vehicle control module receives input signals from multiple microphones, each of which is configured to acquire audio from inside the vehicle. The input signal received from multiple microphones is divided into parallel domain signal and orthogonal domain signal; Choose a constant step size for the weights of the orthogonal domain filter and a variable step size for the weights of the parallel domain filter. The orthogonal domain filter weights are adapted based on the orthogonal domain signal and a constant step size value. The parallel domain filter weights are adapted based on the parallel domain signal and the variable step size value. Combine the adapted orthogonal domain filter weights and the adapted parallel domain filter weights to define the total filter weights; as well as Before the audio output of the received input signal, the total filter weights are applied to the received input signal to perform signal processing operations on the received input signal.