In-vehicle voice positioning system and method based on dynamic noise reduction of T-shaped microphone array
By using a T-shaped microphone array and dynamic noise reduction technology, the problems of low accuracy and poor robustness of in-vehicle voice positioning have been solved, achieving high-precision, real-time voice positioning and adapting to complex in-vehicle environments.
Patent Information
- Application Number
- CN202511924648.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-02-27
AI Technical Summary
Existing in-vehicle voice positioning technology suffers from low positioning accuracy, poor robustness, insufficient real-time performance, and high cost in complex environments. In particular, it is difficult to achieve high-precision voice positioning under in-vehicle multipath effects and noise interference.
A T-shaped microphone array is used, combined with a dynamic noise reduction module and a TDOA estimation module. The time delay difference is calculated by a generalized cross-correlation and phase transformation weighting method, and the sound source coordinates are solved by the least squares method. The filter parameters are dynamically adjusted to adapt to the noise environment at different vehicle speeds.
Achieve high-precision, robust, and real-time voice positioning in complex in-vehicle environments, with positioning errors stable within the range of 10 to 20 centimeters, meeting the real-time requirements of in-vehicle voice interaction and driver status monitoring.
Smart Images

Figure CN121583256A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to an in-vehicle speech positioning system and method based on dynamic noise reduction of a T-shaped microphone array, and belongs to the technical field of sound source positioning and in-vehicle speech signal processing. BACKGROUND
[0002] With the rapid development of intelligent networked vehicles, the in-vehicle speech interaction system has become a core component of modern vehicle human-computer interaction. Precise speech positioning technology not only enables personalized interaction of "who speaks, who wakes up", but also effectively distinguishes the instructions of the driver and the passenger, which is of great significance to improving driving safety and interaction experience.
[0003] However, the in-vehicle environment is a typical complex acoustic scene, which faces multiple challenges:
[0004] First, the space is small and there are many reflecting surfaces, resulting in severe reverberation and multipath effect;
[0005] Second, the background noise is complex and diverse, including engine noise, road noise, wind noise and other steady-state noise, as well as air conditioning, music and other non-steady-state noise;
[0006] Third, due to the installation location restriction, the in-vehicle microphone array usually needs to be designed in a small size, which puts higher requirements on the precision and real-time performance of traditional beamforming and other positioning methods.
[0007] At present, linear or circular microphone arrays are mostly used for in-vehicle speech positioning. Although the linear array is accurate for horizontal azimuth estimation, it has limited two-dimensional plane positioning capability; although the circular array can improve the two-dimensional positioning performance, it requires a larger physical size and has high algorithm complexity and poor real-time performance. In addition, the existing methods have poor positioning stability under the interference of engine noise, windowed wind noise and other interferences, and are easily affected by multipath effect and overlapping of multiple speech. Some vehicle models use a multi-microphone distribution scheme, which can improve the positioning effect, but has high hardware cost and complex installation and maintenance.
[0008] Therefore, there is an urgent need for a speech positioning solution that can achieve high precision, high robustness and dynamic adaptability in a complex in-vehicle environment. SUMMARY
[0009] The application aims to solve the problems of low positioning accuracy, poor robustness, insufficient real-time performance and high system cost of existing in-vehicle speech positioning technology due to complex environmental noise, severe reverberation, limited array size and static rigidity of traditional noise reduction strategies, and provides an in-vehicle speech positioning system and method based on dynamic noise reduction of a T-shaped microphone array.
[0010] The in-vehicle speech positioning system based on dynamic noise reduction of a T-shaped microphone array according to the application comprises:
[0011] A T-shaped microphone array is composed of four omnidirectional MEMS microphones, two of which are symmetrically arranged on both sides of the center axis of the front row of the vehicle, and the third microphone is located at the midpoint of the center axis, which together form a horizontal branch; the fourth microphone is arranged vertically with the midpoint microphone, forming a vertical branch.
[0012] A signal acquisition and preprocessing module is used for synchronous acquisition, framing, and windowing processing of four audio signals.
[0013] A dynamic noise reduction module dynamically adjusts filter parameters according to real-time vehicle speed information to perform frequency domain noise reduction on audio signals.
[0014] A TDOA estimation module calculates the time delay difference between microphone pairs based on generalized cross-correlation and phase transform weighting method.
[0015] A positioning solution module constructs a hyperbolic equation set based on the time delay difference and solves the sound source coordinates using the least squares method.
[0016] Preferably, the layout of the T-shaped microphone array satisfies the following spacing relationships:
[0017] The distance between the fourth microphone and the two end microphones of the horizontal branch is 6.3 cm.
[0018] The distance between the third microphone and the two end microphones of the horizontal branch is 4 cm.
[0019] The distance between the fourth microphone and the third microphone is 4 cm.
[0020] Preferably, the dynamic noise reduction module includes:
[0021] An IIR high-pass filter is used to attenuate low-frequency components of wind noise and engine low-frequency noise.
[0022] An FIR notch filter with a dynamically adjusted notch frequency band according to vehicle speed:
[0023] When the vehicle speed is between 40 km / h and 80 km / h, the notch frequency band is 210 Hz~290 Hz, which is used to suppress wind noise at medium vehicle speed.
[0024] When the vehicle speed is higher than 80 km / h, the notch frequency band is 280 Hz~520 Hz, which is used to suppress strong wind noise and tire noise.
[0025] Preferably, the signal acquisition and preprocessing module uses a sampling rate of 48 kHz, 16-bit quantization, a frame length of 20 ms, and a Hamming window for windowing processing to smooth the signal edges.
[0026] The time domain expression of the window function of the Hamming window is:
[0027] ;
[0028] wherein, denotes the length of the window function, i.e. the frame length,
[0029] denotes the index of the current sample point, ,
[0030] and all denote coefficients of the window function, which for a standard Hamming window take the values:
[0031] , .
[0032] Preferably, all microphones in the T-shaped microphone array are fixed by shockproof supports, and the dynamic range covers 30 dB to 120 dB.
[0033] The in-vehicle voice positioning method based on the T-shaped microphone array dynamic noise reduction provided by the application specifically comprises the following steps:
[0034] Four audio signals are collected by the T-shaped microphone array and are synchronously pretreated;
[0035] Filter parameters are dynamically selected according to real-time vehicle speed, and the signals are subjected to noise reduction processing;
[0036] The signals subjected to noise reduction are subjected to amplitude normalization, and the generalized cross-correlation functions between all microphone pairs are calculated;
[0037] The time delay difference between each microphone pair is estimated based on a phase transformation weighting method;
[0038] A plane coordinate system is established with one microphone as the coordinate origin, a hyperbolic equation set is constructed according to the time delay difference, and the sound source position is solved by using the least square method.
[0039] Preferably, the specific method for amplitude normalization of the signals subjected to noise reduction comprises:
[0040] ;
[0041] wherein, denotes the new amplitude obtained after amplitude normalization processing on the original sample point ,
[0042] denotes the original amplitude of a single sample point in the time domain waveform of a single audio signal after noise reduction,
[0043] denotes the maximum value of the amplitude of the signal in the current processing frame,
[0044] denotes the minimum value of the amplitude of the signal in the current processing frame;
[0045] denotes the data is shifted to start from the minimum value,
[0046] denotes the original range of the signal is calculated, and linear scaling is implemented to interval;
[0047] multiplied by 2 and reduced by 1, which denotes linear transformation of the value in the interval to interval, and normalization is completed.
[0048] Preferably, the specific method of calculating the generalized cross-correlation function between all pairs of microphones comprises:
[0049] the cross-correlation function of the two signals and is:
[0050] ;
[0051] wherein, denotes a time extension parameter, and denote the time-domain audio signals collected by two microphones numbered and respectively.
[0052] Preferably, the specific method of estimating the time delay difference between each pair of microphones based on the phase transform weighting method comprises:
[0053] The phase transform PHAT weighting highlights the phase information to sharpen the correlation peak:
[0054] ;
[0055] wherein, denotes the cross-correlation function after the phase transform weighting,
[0056] is the Fourier transform result of , is the Fourier transform result of , denotes the complex conjugate of ,
[0057] This represents a complex exponential function and signifies the time shift factor in the frequency domain.
[0058] Then calculate the time delay difference between microphone pairs. :
[0059] .
[0060] Preferably, the specific method for establishing a planar coordinate system with one of the microphones as the origin, constructing a hyperbolic equation system based on the time delay difference, and solving for the sound source location using the least squares method includes:
[0061] Establish a planar coordinate system with the position of the third microphone as the origin;
[0062] Let the coordinates of the target sound source be... ;
[0063] The positions of the four microphones are as follows , , , ;
[0064] Let the wind speed be ;
[0065] The system of equations for the intersection points of hyperbolas:
[0066]
[0067] Solving the system of equations yields the coordinates of the sound source location. ;
[0068] in, This indicates the time difference of arrival between the first and second microphones.
[0069] This indicates the time difference of arrival between the first and third microphones.
[0070] This indicates the time difference of arrival between the third and fourth microphones.
[0071] Advantages of the present invention: The in-vehicle voice localization system and method based on dynamic noise reduction of a T-shaped microphone array proposed in this invention have the following advantages:
[0072] 1. High positioning accuracy and stable and controllable error: Through the optimized T-shaped microphone array layout, the spatial resolution is effectively improved within a limited physical size. Combined with the high-precision PHAT weighted time delay estimation algorithm and least squares optimized positioning solution, the positioning error of the system in complex in-vehicle environments can be stably controlled within the range of 10 to 20 centimeters.
[0073] 2. Strong environmental adaptability and excellent robustness: A new dynamic noise reduction strategy is proposed, which can adaptively adjust the filter parameters according to the real-time vehicle speed and specifically suppress the dominant noise under different working conditions. This mechanism greatly enhances the system's resistance to time-varying noise. Compared with traditional noise reduction algorithms, it has higher accuracy and more stable error fluctuation trend, and does not appear the phenomenon of sudden positioning accuracy drop.
[0074] 3. Good real-time performance and meeting interactive requirements: The optimization is carried out for 20ms audio frame, which ensures that the end-to-end delay from signal collection to output positioning result is extremely low, fully meeting the stringent requirements of real-time performance for vehicle-mounted voice interaction, driver state monitoring and other applications. BRIEF DESCRIPTION OF DRAWINGS
[0075] Figure 1 is the flow chart of the in-vehicle voice positioning based on the dynamic noise reduction of the T-shaped microphone array according to the present application;
[0076] Figure 2 is the layout schematic diagram of the T-shaped microphone array according to the present application;
[0077] Figure 3 is the positioning performance analysis comparison diagram of the ordinary noise reduction algorithm and the dynamic noise reduction algorithm for specific test points in the vehicle;
[0078] Figure 4 is the positioning error distribution comparison diagram;
[0079] Figure 5 is the positioning error cumulative distribution function diagram;
[0080] Figure 6 is the positioning error distribution histogram;
[0081] Figure 7 is the positioning error change trend diagram. DETAILED DESCRIPTION
[0082] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0083] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0084] The present application will be further described below with reference to the drawings and specific embodiments, but not as a limitation of the present application.
[0085] Embodiment 1:
[0086] The present embodiment will be described below Figure 1 And Figure 2 The present embodiment is a voice positioning system in a vehicle based on dynamic noise reduction of a T-shaped microphone array, which comprises:
[0087] The T-shaped microphone array is composed of four omnidirectional MEMS microphones, two of which are symmetrically arranged on both sides of the center axis of the front row of the vehicle, and the third microphone is located at the midpoint of the center axis, and the three together form a horizontal arm; the fourth microphone is arranged in a vertical direction with the midpoint microphone, forming a vertical arm;
[0088] The signal acquisition and preprocessing module is used for synchronous acquisition, framing and windowing processing of four audio signals;
[0089] The dynamic noise reduction module dynamically adjusts the filter parameters according to the real-time vehicle speed information to perform frequency domain noise reduction on the audio signal;
[0090] The TDOA estimation module calculates the time delay difference between the microphone pairs based on the generalized cross-correlation and phase transformation weighting method;
[0091] The positioning solution module constructs a hyperbolic equation set according to the time delay difference and solves the sound source coordinates by using the least squares method.
[0092] Further, the layout of the T-shaped microphone array satisfies the following distance relationship:
[0093] The distance between the fourth microphone and the two end microphones of the horizontal arm is 6.3 cm;
[0094] The distance between the third microphone and the two end microphones of the horizontal arm is 4 cm;
[0095] The distance between the fourth microphone and the third microphone is 4 cm.
[0096] Further, the dynamic noise reduction module comprises:
[0097] An IIR high-pass filter is used to attenuate the low-frequency components of wind noise and engine low-frequency noise;
[0098] An FIR notch filter, whose notch frequency band is dynamically adjusted according to the vehicle speed:
[0099] When the vehicle speed is 40km / h to 80km / h, the notch frequency band is 210Hz~290Hz, which is used to suppress wind noise at medium vehicle speed;
[0100] When the vehicle speed is higher than 80km / h, the notch frequency band is 280Hz~520Hz, which is used to suppress strong wind noise and tire noise.
[0101] Further, the signal acquisition and preprocessing module adopts a 48kHz sampling rate, 16-bit quantization, a frame length of 20ms, and Hamming windowing processing for smoothing the signal edges.
[0102] The time domain expression of the window function of the Hamming window is:
[0103] ;
[0104] wherein, denotes the length of the window function, i.e., the frame length,
[0105] denotes the index of the current sampling point, ,
[0106] and both denote the coefficients of the window function, and for a standard Hamming window, the value is:
[0107] , .
[0108] Further, all the microphones in the T-shaped microphone array are fixed by shockproof supports, and the dynamic range covers 30dB to 120dB.
[0109] Embodiment 2:
[0110] The present embodiment will be described below Figures 5-7 The in-vehicle voice positioning method based on the T-shaped microphone array dynamic noise reduction, which specifically includes:
[0111] Four audio signals are collected by the T-shaped microphone array and are synchronously preprocessed;
[0112] Filter parameters are dynamically selected according to the real-time vehicle speed, and the signals are processed for noise reduction;
[0113] The amplitude of the noise-reduced signal is normalized, and the generalized cross-correlation function between all microphone pairs is calculated;
[0114] The time delay difference between each microphone pair is estimated based on the phase transformation weighting method;
[0115] A plane coordinate system is established with one microphone as the coordinate origin, a hyperbolic equation set is constructed according to the time delay difference, and the least square method is used to solve the sound source position.
[0116] Further, the specific method for amplitude normalization of the noise-reduced signal includes:
[0117] ;
[0118] wherein, denotes the new amplitude obtained after amplitude normalization processing of the original sampling point ,
[0119] denotes the original amplitude of a single sampling point in the time-domain waveform of a single-channel audio signal after noise reduction,
[0120] denotes the maximum value of the amplitude of the signal in all sampling points of the current processing frame,
[0121] denotes the minimum value of the amplitude of the signal in all sampling points of the current processing frame;
[0122] denotes the data is shifted to start from the minimum value,
[0123] denotes the original range of the signal is calculated, and linear scaling to interval is implemented;
[0124] multiplied by 2 and reduced by 1, denotes linear transformation of the value in the interval to the interval, and normalization is completed.
[0125] Further, the specific method of calculating the generalized cross-correlation function between all pairs of microphones includes:
[0126] the cross-correlation function of two signals and is:
[0127] ;
[0128] wherein, denotes a time extension parameter, and denote the time-domain audio signals collected by two microphones numbered and .
[0129] Further, the specific method of estimating the time delay difference between each pair of microphones based on the phase transform weighting method includes:
[0130] The phase transform PHAT weighting highlights the phase information to sharpen the correlation peak:
[0131] ;
[0132] wherein, denotes the cross-correlation function after phase transform weighting,
[0133] For the Fourier transform result, For the Fourier transform result, Indicates The complex conjugate of,
[0134] Indicates the complex exponential function, represents the time shift factor in the frequency domain;
[0135] Further calculate the time delay difference between the microphone pair :
[0136] .
[0137] Further, the method for establishing a plane coordinate system with one of the microphones as the coordinate origin, constructing a hyperbolic equation set according to the time delay difference, and solving the specific position of the sound source by using the least square method comprises:
[0138] A plane coordinate system is established with the position of the third microphone as the coordinate origin;
[0139] Let the coordinates of the target sound source be ;
[0140] The positions of the four microphones are , , , ;
[0141] Let the wind speed be ;
[0142] Hyperbolic intersection equation set:
[0143]
[0144] Solve the equation set to obtain the coordinates of the sound source position ;
[0145] Wherein, Indicates the time difference between the first microphone and the second microphone of the microphone pair,
[0146] Indicates the time difference between the first microphone and the third microphone of the microphone pair,
[0147] Indicates the time difference between the third microphone and the fourth microphone of the microphone pair.
[0148] In the present application, the hardware core of the system is composed of a microphone array, a data acquisition module and a signal processing unit.
[0149] Figure 1 is the flow chart of the in-vehicle sound source positioning system. The chart shows the complete algorithm process and data flow from the beginning of sound signal collection, through preprocessing, dynamic noise reduction, generalized cross-correlation calculation, TDOA estimation, to the final calculation of the sound source position coordinates in the form of block diagram and arrow connection.
[0150] Microphone array: As shown in Figure 2 , the chart specifically shows the two-dimensional plane arrangement of the four microphones near the center axis of the front row of the vehicle in the form of a top view diagram, including the composition of horizontal and vertical arms. A four-microphone array configuration is adopted, in which microphones No. 2, No. 3, and No. 4 are symmetrically arranged along the center axis of the front row of the vehicle (e.g., below the front windshield, in front of the roof console), forming a horizontal arm structure for estimating the horizontal azimuth angle of the sound source. Microphone No. 1 is vertically arranged at the central position of the front row, connected to the midpoint of the horizontal arm, constituting a vertical arm for assisting in estimating the vertical information of the sound source or enhancing spatial resolution.
[0151] The distance between each arm is optimized to balance spatial resolution and phase ambiguity. The specific parameters are:
[0152] Distance between microphone No. 1 and microphones No. 2 and No. 4: 6.3 cm
[0153] Distance between microphone No. 1 and microphone No. 3: 4 cm
[0154] Distance between microphone No. 3 and microphones No. 2 and No. 4: 4 cm
[0155] All microphones are fixed through shockproof supports, with a dynamic range covering 30-120 dB to ensure accurate capture of the human voice frequency band.
[0156] Data acquisition and preprocessing system: TITMS320C6748 floating-point DSP chip (1 GHz main frequency, 4-core parallel) combined with ADS8364 chip (6-channel synchronization, 16-bit resolution, 48 kHz sampling rate) is used for data acquisition. The system generates a synchronous sampling pulse signal through FPGA to ensure that the time synchronization error of the 4-channel ADC is less than 50 ns. DMA ping-pong operation is used to transfer audio data to the DSP ring buffer, with a single positioning period of 20 ms (960 sampling points).
[0157] Window function processing: To avoid the problem of frequency spectrum leakage when performing Fourier transform on a finite-length signal, a Hamming window is used to smooth the signal edges. The window function time-domain expression is as follows:
[0158]
[0159] where, denotes the length of the window function, i.e. the frame length,
[0160] denotes the index of the current sample point, ,
[0161] and denote the coefficients of the window function, which for a standard Hamming window take the values:
[0162] , .
[0163] Dynamic noise reduction scheme: Based on vehicle speed (CAN bus access to the automotive diagnostic interface to read real-time vehicle speed information) to dynamically adjust the design of IIR + FIR cascade filter.
[0164] Where the IIR filter (4 order Chebyshev high-pass) set the cutoff frequency of 150Hz, to attenuate the low frequency components of wind noise and engine low frequency noise.
[0165] FIR notch filter (128 order linear phase) using dynamic adjustment scheme:
[0166] When the vehicle speed is between 40km / h and 80km / h, the notch window is designed to be 210~290Hz, mainly used to suppress the significant wind noise at medium speed.
[0167] When the vehicle speed is greater than 80km / h, the notch window is designed to be 280Hz~520Hz, mainly used to suppress strong wind noise and tire noise.
[0168] Sound source localization algorithm:
[0169] Amplitude normalization: To avoid the influence of amplitude difference on cross-correlation peak detection, amplitude normalization is used. The calculation formula is:
[0170]
[0171] where, denotes the new amplitude obtained after amplitude normalization processing of the original sample point ,
[0172] denotes the original amplitude of a single sample point in the time domain waveform of a single audio signal after noise reduction,
[0173] denotes the maximum value of the amplitude of the signal in all sample points of the current processing frame,
[0174] denotes the minimum value of the amplitude of the signal in all sample points of the current processing frame;
[0175] represents the data is shifted to start with the minimum value,
[0176] represents the signal original range is calculated, linear scaling to interval is achieved.
[0177] multiplied by 2 and reduced by 1, represents the value of the interval is linearly transformed to interval, completing the normalization.
[0178] Cross-correlation calculation:
[0179] The cross-correlation function of two signals and is:
[0180] ;
[0181] where, represents the time extension parameter, and respectively represent the time-domain audio signals collected by two microphones numbered and .
[0182] TDOA estimation: introduce PHAT (Phase Transformation) weighting to highlight phase information to sharpen the correlation peak:
[0183]
[0184] where, represents the cross-correlation function after phase transformation weighting,
[0185] is the Fourier transform result of , is the Fourier transform result of , represents the complex conjugate of ,
[0186] represents the complex exponential function, representing the time shift factor in the frequency domain;
[0187] and then calculate the time delay difference between the microphone pair :
[0188] .
[0189] Sound source positioning solution:
[0190] Establish a mathematical model: take the third microphone position as the coordinate origin to establish a plane coordinate system;
[0191] Let the coordinates of the target sound source be ;
[0192] The positions of the four microphones are , , , ;
[0193] Let the wind speed be .
[0194] Construct an equation set: according to the geometric principle, the distance difference between the sound source and two microphones is equal to the sound speed multiplied by the time delay difference between them. Thus, for each pair of microphones, a hyperbolic equation can be established, and using the multiple groups of independent equations calculated, a nonlinear equation set can be constructed:
[0195]
[0196] Solve the position coordinates: the above equation set is usually overdetermined or nonlinear, and it is difficult to solve directly. The least squares method is used for optimization and solving, that is, to find a set of values, so that the sum of the squares of the difference between the left and right sides of all hyperbolic equations is minimized. This process can be efficiently implemented in DSP through mature numerical algorithms such as the Newton-Gauss iteration method, and finally the two-dimensional plane coordinates of the sound source are output.
[0197] As shown in Figures 3-7 , it is a positioning performance analysis comparison diagram of the ordinary noise reduction algorithm and the dynamic noise reduction algorithm. Figure 3 is a positioning result comparison diagram of a specific test point in the vehicle. This diagram shows the actual positioning effect of the system at a predetermined test position (coordinate point) in the vehicle in the form of a two-dimensional plane scatter diagram. The diagram clearly marks the positions of the four microphones (black squares), the predetermined coordinates of the real sound source (red circles), the positioning results obtained by using the ordinary noise reduction algorithm (blue cross points), and the positioning results obtained by using the dynamic noise reduction algorithm of the present application (green triangles). By directly comparing the offset distances between the two types of positioning results and the real sound source, the significant advantages of the dynamic noise reduction method of the present application in improving the accuracy of single-point positioning are clearly verified.
[0198] Figure 4 is a positioning error distribution comparison diagram, Figure 5 is a positioning error cumulative distribution function diagram, Figure 6 is a positioning error distribution histogram, Figure 7is a positioning error change trend chart. In the real vehicle test, the dynamic noise reduction positioning method of the application is compared with the traditional method using fixed parameter noise reduction. The test sound source is placed in different positions in the vehicle, and the vehicle drives at different speeds. Figures 3-7 The experimental results show that the positioning error of the method of the application can be stabilized in the range of 10-20 cm under different vehicle speeds and noise environments, and the error curve is smooth. While the positioning error of the traditional method significantly increases and fluctuates violently when the vehicle speed changes (especially at high speed). This fully verifies that the application effectively improves the precision, stability and environmental adaptability of in-vehicle voice positioning by array optimization and dynamic noise reduction cooperation.
[0199] Although the application is described herein with reference to particular embodiments, it should be understood that these examples are merely illustrative of the principles and applications of the present application. It should therefore be understood that numerous modifications can be made to the illustrative embodiments and that other arrangements can be devised without departing from the spirit and scope of the present application as defined by the appended claims. It should be understood that the features described in connection with one embodiment can be used in conjunction with other embodiments described herein. It should also be understood that features described in connection with separate embodiments can be used in combination with each other.
Claims
1. An in-vehicle voice positioning system based on dynamic noise reduction using a T-shaped microphone array, characterized in that, It includes: The T-shaped microphone array consists of four omnidirectional MEMS microphones. Two microphones are symmetrically arranged on both sides of the center axis of the front row of the vehicle, and the third microphone is located at the midpoint of the center axis. Together, the three form a horizontal support arm. The fourth microphone is aligned with the microphone at the midpoint along the vertical direction to form a vertical support arm. The signal acquisition and preprocessing module is used to synchronously acquire, frame, and window the four audio signals. The dynamic noise reduction module dynamically adjusts the filter parameters based on real-time vehicle speed information to perform frequency domain noise reduction on the audio signal; The TDOA estimation module calculates the time delay difference between microphone pairs based on a generalized cross-correlation and phase transform weighting method. The positioning and calculation module constructs a hyperbolic equation system based on the time delay difference and uses the least squares method to solve for the sound source coordinates.
2. The in-vehicle voice positioning system based on dynamic noise reduction using a T-shaped microphone array according to claim 1, characterized in that, The layout of the T-shaped microphone array satisfies the following spacing relationship: The distance between the fourth microphone and the microphones at both ends of the horizontal support arm is 6.3cm; The distance between the third microphone and the microphones at both ends of the horizontal support arm is 4cm. The distance between the fourth microphone and the third microphone is 4cm.
3. The in-vehicle voice positioning system based on dynamic noise reduction using a T-shaped microphone array according to claim 1, characterized in that, The dynamic noise reduction module includes: An IIR high-pass filter is used to attenuate low-frequency components of wind noise and low-frequency engine noise; An FIR notch filter whose notch bandwidth is dynamically adjusted according to vehicle speed: When the vehicle speed is between 40km / h and 80km / h, the notch filter frequency band is 210Hz~290Hz, which is used to suppress wind noise at medium vehicle speeds. When the vehicle speed is above 80km / h, the notch filter frequency band is 280Hz~520Hz, which is used to suppress strong wind noise and tire noise.
4. The in-vehicle voice positioning system based on dynamic noise reduction using a T-shaped microphone array according to claim 1, characterized in that, The signal acquisition and preprocessing module uses a 48kHz sampling rate, 16-bit quantization, and a frame length of 20ms. It also uses a Hamming window for windowing to smooth signal edges. The time-domain expression of the Hamming window function is as follows: ; in, This indicates the length of the window function, i.e., the frame length. Indicates the index of the current sampling point. , and All represent the coefficients of the window function. For the standard Hamming window, the values are: , 。 5. The in-vehicle voice positioning system based on dynamic noise reduction using a T-shaped microphone array according to claim 1, characterized in that, All microphones in the T-shaped microphone array are fixed with shockproof brackets, and the dynamic range covers 30dB to 120dB.
6. A positioning method based on the in-vehicle voice positioning system with dynamic noise reduction based on a T-shaped microphone array as described in any one of claims 1-5, characterized in that, This positioning method specifically includes: Four audio signals were acquired using a T-shaped microphone array and then synchronously preprocessed. The filter parameters are dynamically selected based on the real-time vehicle speed to perform noise reduction processing on the signal; The amplitude of the noise-reduced signal is normalized, and the generalized cross-correlation function between all microphone pairs is calculated. The time delay difference between each microphone pair is estimated based on the phase transform weighting method; A planar coordinate system is established with one of the microphones as the origin. A hyperbola equation system is constructed based on the time delay difference, and the location of the sound source is solved using the least squares method.
7. The in-vehicle voice localization method based on dynamic noise reduction of a T-shaped microphone array according to claim 6, characterized in that, The specific method for amplitude normalization of the denoised signal includes: ; in, Indicates the original sampling points The new amplitude obtained after amplitude normalization This represents the original amplitude of a single sampling point in the time-domain waveform of a single audio signal after noise reduction. This indicates the maximum amplitude of the signal across all sampling points in the current processing frame. This indicates the minimum amplitude of the signal across all sampling points in the current processing frame; This means shifting the data to a position starting from the minimum value. This represents the original range of the calculated signal, achieving linear scaling to... interval; Multiply by 2 and subtract 1, which means to The values of the interval are linearly transformed to The interval is normalized.
8. The in-vehicle voice localization method based on dynamic noise reduction of a T-shaped microphone array according to claim 6, characterized in that, The specific method for calculating the generalized cross-correlation function among all microphone pairs includes: Two signals and cross-correlation function for: ; in, Indicates the time extension parameter, and They represent the numbers respectively. and The time-domain audio signal collected by the two microphones.
9. The in-vehicle voice localization method based on dynamic noise reduction of a T-shaped microphone array according to claim 6, characterized in that, The specific method for estimating the time delay difference between microphone pairs based on the phase transform weighting method includes: Phase information highlighted by PHAT phase transformation is introduced to sharpen relevant peaks: ; in, This represents the cross-correlation function after phase transformation and weighting. for The Fourier transform result, for The Fourier transform result, express The complex conjugate, This represents a complex exponential function and a time shift factor in the frequency domain; Then calculate the time delay difference between microphone pairs. : 。 10. The in-vehicle voice localization method based on dynamic noise reduction of a T-shaped microphone array according to claim 6, characterized in that, The specific method for establishing a planar coordinate system with one of the microphones as the origin, constructing a hyperbolic equation system based on the time delay difference, and solving for the sound source location using the least squares method includes: Establish a planar coordinate system with the position of the third microphone as the origin; Let the coordinates of the target sound source be... ; The positions of the four microphones are as follows , , , ; Let the wind speed be ; The system of equations for the intersection points of hyperbolas: Solving the system of equations yields the coordinates of the sound source location. ; in, This indicates the time difference of arrival between the first and second microphones. This indicates the time difference of arrival between the first and third microphones. This indicates the time difference of arrival between the third and fourth microphones.