Acoustic wave focusing sensing method and system based on distributed vehicle-mounted microphone array
Through the acoustic wave focus perception method of distributed vehicle microphone arrays, dual-band signal and three-dimensional beamforming technology are used to solve the problems of small aperture and wide signal beam of the vehicle microphone array, and multi-objective non-contact sensing and intelligent cockpit functions are achieved.
Patent Information
- Application Number
- CN202510484881.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-26
AI Technical Summary
The existing vehicle-mounted microphone array has small aperture, wide signal beam and low direction, making it difficult to achieve multi-objective and anti-interference intelligent cockpit perception.
Using a distributed vehicle microphone array, channel impulse response is generated through dual-band ZC sequence signals, and combined with differential frequency processing and three-dimensional beamforming technology, the focus of the acoustic signal and the estimation of the target position are achieved.
Multi-objective contactless perception, including multi-objective breath detection, gesture recognition and driver driving behavior monitoring, provides location-based multi-objective perception information to support a wider range of smart cockpit applications.
Smart Images

Figure CN120544590A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of contactless intelligent perception, and in particular relates to a sound wave focusing perception method and system based on a distributed vehicle-mounted microphone array. Background Art
[0002] With the rapid development of intelligent vehicles and connected car technology, in-vehicle microphones are becoming increasingly important as a bridge between the driver and the vehicle's intelligent systems. These microphones are typically embedded in the vehicle's interior and connected to the navigation system, voice assistant, and in-car phone functions. They are designed to provide convenient features such as hands-free calling and voice command operation, thereby enhancing driving safety and the entertainment experience. With the continuous advancement of technology and continuous product upgrades, the performance of in-vehicle microphones has significantly improved. For example, AAC Technologies launched an intelligent cockpit speech acquisition system (2023. AAC Technologies Launches a Complete Set of Automotive MEMS Microphone Modules to Accelerate Its Automotive Business.), which uses six microphone array modules arranged around the vehicle, each containing at least two microphones. By optimizing microphone layout and processing algorithms, this system achieves high-precision speech recognition and interaction in scenarios such as zoned calls, voice control, and active noise cancellation.
[0003] In recent years, the application of in-car microphones has expanded beyond traditional communication and entertainment functions to include intelligent sensing and human-computer interaction. Su Y et al. proposed a contactless child presence detection method based on an in-car distributed speaker and microphone system (Su Y, Zhang F, et al. 2024. Embracing Distributed Acoustic Sensing in Car Cabin for Children Presence Detection. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 8, 1, Article 16 (March 2024), 28 pages.), which is a fundamental feature of future autonomous vehicles.
[0004] When investigating other existing in-vehicle acoustic wave sensing systems, it was found that most studies can monitor the driver's vital signs by placing a mobile phone on the dashboard and facing the driver, such as the respiratory symptom detection based on acoustic sensing in driving environments proposed by Wu Y et al. (Wu Y, Li F, Xie Y, et al. SymListener: Detecting Respiratory Symptoms via Acoustic Sensing in Driving Environments[J]. ACM Transactions on Sensor Networks, 2023, 19(1): 1-21.). However, due to the limited number of microphones on devices such as mobile phones or smart speakers (usually only 2 to 6, limited by the size of the device), the aperture of the microphone array is only a few centimeters, and the signal beam is wide and has low directivity. In contrast, on-board distributed microphones naturally form a large array around the entire vehicle, with a larger aperture (the distance between microphones can be up to 1.5 meters) and a larger number of microphones (usually more than 10). This makes it possible to achieve narrower signal beams and higher target angle resolution. By focusing the acoustic wave signals on different targets, a multi-target, anti-interference, and fine-grained intelligent cockpit perception system can be realized. Summary of the Invention
[0005] Based on a distributed vehicle-mounted microphone array, the present invention provides a method and system for acoustic wave focusing perception based on a distributed vehicle-mounted microphone array. The system uses vehicle-mounted speakers to transmit acoustic wave signals and vehicle-mounted microphones to receive acoustic wave signals reflected from targets, achieving acoustic signal energy focusing at the target location. Using the distributed microphones in the vehicle cabin, the occupancy of targets within the vehicle is modeled, enabling estimation of multiple target locations. Furthermore, based on the target location, the acoustic wave signal is further focused on the target, enabling non-contact sensing of the target, such as multi-target breathing detection, multi-target gesture recognition, and driver behavior monitoring.
[0006] To achieve the above objectives, the technical solution of the present invention includes the following contents.
[0007] A sound wave focusing perception method based on a distributed vehicle-mounted microphone array, the method comprising:
[0008] Based on the dual-band ZC sequence signal, a channel impulse response cir1 and a channel impulse response cir2 of two frequency bands are generated; wherein the dual-band ZC sequence signal is transmitted based on the vehicle-mounted distributed multi-speaker and received by the distributed microphone array;
[0009] By performing difference frequency processing on the channel impulse responses of the two frequency bands, a difference frequency channel impulse response is obtained, and the number and area of the targets in the cockpit are extracted based on the difference frequency channel impulse response;
[0010] A microphone in the microphone array is used as a reference microphone, and a polar coordinate system is constructed with the reference microphone as the origin. The azimuth angle θ and pitch angle of the center of the target area in the polar coordinate system are obtained through three-dimensional beamforming technology.
[0011] Based on the azimuth angle θ and the pitch angle The steering vector is compensated on the channel impulse response cir1 or the channel impulse response cir2, and the compensated channel impulse response is used to extract the perceptual information of the target.
[0012] Furthermore, based on the dual-band ZC sequence signal, channel impulse responses cir1 and cir2 of two frequency bands are generated, including:
[0013] Two band-pass filters are used to band-pass filter the dual-band ZC sequence signal;
[0014] A cross-correlation operation is performed on the filtered signal of each frequency band and the dual-band ZC sequence signal to obtain a channel impulse response cir1 and a channel impulse response cir2.
[0015] Furthermore, the number and area of the targets in the cockpit are extracted based on the beat frequency channel impulse response, including:
[0016] Divide the three-dimensional space of the vehicle cabin into three-dimensional grids;
[0017] Perform multipath signal decomposition on the beat frequency channel impulse response of each microphone, and map each multipath signal to the corresponding distance grid in the three-dimensional grid according to the distance between the target and the microphone;
[0018] The multipath signals from all microphones in each grid are superimposed to obtain the energy intensity of the grid;
[0019] According to the energy intensity of each grid, the number and area of the targets in the cockpit are obtained.
[0020] Furthermore, a microphone in the microphone array is used as a reference microphone, and a polar coordinate system is constructed with the reference microphone as the origin. The azimuth angle θ and pitch angle of the target area center in the polar coordinate system are obtained by three-dimensional beamforming technology. include:
[0021] Establish a three-dimensional coordinate system based on the origin;
[0022] The distributed microphone array is modeled on the YOZ plane in the three-dimensional coordinate system, and the i-th microphone M is set i The rectangular coordinates are (x i ,y i ,z i ), the polar coordinates are The rectangular coordinates of a grid in the target K area are (x K ,y K ,z K ), the polar coordinates are And the conversion relationship between rectangular coordinates and polar coordinates is
[0023] In the three-dimensional coordinate system and polar coordinate system, the received signal phase difference ΔΦ is modeled i and the path length difference Δd i The mathematical expression between: wherein the received signal phase difference ΔΦ i is the i-th microphone M i Phase difference of the received signal from the reference microphone, path length difference Δd i For this grid, the number of microphones M is i The path length difference to the reference microphone;
[0024] Based on the received signal phase difference ΔΦ i and the path length difference Δd i The mathematical expression between the two models models the i-th microphone M i Phase compensation coefficient ω i The mathematical expression of
[0025] Mathematical expression for the steering vector W that models beamforming; where the steering vector W = [ω1,ω2,,...,ω N ];
[0026] By compensating the steering vector W on the difference frequency channel impulse response of the grid, an expression for the beamformed signal is generated;
[0027] The azimuth angle θ and the elevation angle in the expression of the traversal beamforming signal are By identifying the peak of the two-dimensional angle spectrum, the azimuth angle θ and pitch angle corresponding to the central grid in the target K area are obtained.
[0028] Furthermore, it is characterized in that the path length difference
[0029] Furthermore, in the three-dimensional coordinate system and the polar coordinate system, the received signal phase difference ΔΦ i and the path length difference Δd iThe mathematical expression between is:
[0030] Furthermore, the i-th microphone M i Phase compensation coefficient ω i The mathematical expression is: Wherein, Δf represents the frequency of the difference frequency channel impulse response, and c represents the speed of sound.
[0031] A sound wave focusing perception system based on a distributed vehicle-mounted microphone array, the system comprising:
[0032] A signal preprocessing module is configured to generate a channel impulse response cir1 and a channel impulse response cir2 of two frequency bands based on a dual-band ZC sequence signal, wherein the dual-band ZC sequence signal is transmitted based on a vehicle-mounted distributed multi-speaker and received using a distributed microphone array;
[0033] A target occupancy assessment module is used to perform difference frequency processing on the channel impulse responses of the two frequency bands to obtain a difference frequency channel impulse response, and extract the number and area of the targets in the cockpit based on the difference frequency channel impulse response;
[0034] The spatial acoustic wave focusing module is used to use one microphone in the microphone array as a reference microphone, construct a polar coordinate system with the reference microphone as the origin, and obtain the azimuth angle θ and pitch angle of the target area center in the polar coordinate system through three-dimensional beamforming technology. Based on the azimuth angle θ and the pitch angle The steering vector is compensated on the channel impulse response cir1 or the channel impulse response cir2, and the compensated channel impulse response is used to extract the perceptual information of the target.
[0035] An electronic device comprises: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements any of the above-mentioned sound wave focusing perception methods based on a distributed vehicle-mounted microphone array.
[0036] A computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement any of the above-mentioned sound wave focusing perception methods based on a distributed vehicle-mounted microphone array.
[0037] Compared with the prior art, the present invention has the following beneficial effects.
[0038] The present invention utilizes distributed vehicle-mounted speakers to transmit and distributed vehicle-mounted microphone arrays to receive acoustic signals, analyzes and processes the echoes, and focuses the acoustic signals on multiple targets in the vehicle to achieve non-contact perception. The present invention makes full use of the distributed acoustic devices already deployed in the vehicle, does not require any additional hardware costs, and can be used for non-contact perception in a convenient manner. At the same time, the dual-band signal transmission design and the difference frequency-based signal processing method proposed in the present invention suppress the phase ambiguity problem in the angle spectrum and eliminate the influence of the interference side lobe on the target main lobe; the proposed vehicle target occupancy grid construction technology makes full use of the diversity of multi-microphone received signals, constructs the dynamic energy distribution caused by multiple targets in the vehicle space, and further extracts the number and position information of the targets in the vehicle; the three-dimensional beamforming technology adjusts and combines the phases of the received signals of microphones at different positions by combining the pitch angle and horizontal angle information of the three-dimensional space, so as to achieve the focusing of the acoustic signal at each position target in the three-dimensional space, and further can extract the perception signals of multiple targets (such as breathing frequency, gestures, etc.) at the same time. The proposed acoustic wave focusing perception system based on a distributed vehicle-mounted microphone array enables non-contact monitoring and identification of multiple targets within the vehicle, including multi-target breathing detection, multi-target gesture recognition, and driver behavior monitoring. This system provides a fundamental spatial perception capability, providing position-based multi-target perception information for the vehicle, thereby supporting a wider range of in-vehicle intelligent applications. This invention is easily applicable to fields such as intelligent sensing and smart cockpits. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Flowchart corresponding to the system of the present invention.
[0040] Figure 2 This is a modeling diagram of a distributed microphone array in three-dimensional space.
[0041] Figure 3 Figure 2 shows the deployment of distributed loudspeaker and microphone elements in a vehicle.
[0042] Figure 4 These are three embodiments of the present invention. DETAILED DESCRIPTION
[0043] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.
[0044] The present invention discloses an acoustic wave focusing perception system based on a distributed vehicle-mounted microphone array. First, a vehicle-mounted distributed loudspeaker is designed to transmit a dual-band ZC sequence signal and receive it using a distributed microphone array. Then, the received signal is bandpass filtered to remove noise outside the required range and extract signals from the two frequency bands respectively. Then, a cross-correlation operation is performed on the transmitted signal and the two frequency band signals respectively to obtain channel impulse responses (CIR1 and CIR2). The two CIR signals are then difference-frequency-measured to construct the occupancy status of the targets in the cabin and extract the number and position of the targets. Finally, based on the proposed three-dimensional beamforming technology, the acoustic wave signal is focused on the target at each position to extract and enhance the target's perception signals (such as breathing signals, gestures, etc.).
[0045] like Figure 1 As shown in the figure, this system consists of three main steps: signal preprocessing, target occupancy assessment, and spatial acoustic wave focusing. The first step, signal preprocessing, mainly involves bandpass filtering the received signal and performing an autocorrelation operation to obtain CIR signals in two frequency bands. The second step, target occupancy assessment, mainly involves difference frequency signal processing and occupancy grid construction to obtain target count and location information. The third step, spatial acoustic wave focusing, mainly uses three-dimensional beamforming technology to focus the signal at the target to obtain an enhanced perception signal.
[0046] Step one: signal preprocessing.
[0047] The ZC (Zadoff-Chu) sequence is used as the baseband signal of the transmitted signal. The generation expression of the ZC sequence is as follows:
[0048]
[0049] Among them, u is the root sequence, u∈{1,2,…,(N zc -1)},N zc is the length of the ZC sequence, u and N zc The relationship of mutual prime must be satisfied, that is, gcd(u,N zc )=1 means that the sequence position n=0,1,…,N zc -1, q is a constant, and q∈Z, j is the unit of the imaginary part of the complex signal.
[0050] Because the spacing between on-board microphones far exceeds half the wavelength of the acoustic signal, significant phase ambiguity is generated during the subsequent beamforming process, interfering with target recognition. Therefore, the present invention designs a unique dual-band transmission signal, modulating the baseband signal to high frequencies of 15kHz-17kHz and 18kHz-20kHz before transmitting. The center frequencies of the two frequency bands are f1 = 16kHz and f2 = 19kHz, respectively.
[0051] To obtain channel information, namely the channel impulse response (CIR) for perception, the present invention first bandpass filters the received signal of the distributed microphone array and uses two corresponding bandpass filters to extract the signals of the two frequency bands respectively. Then, the two filtered received signals of the two frequency bands are cross-correlated with the transmitted signals of each frame to obtain the corresponding CIRs (i.e., cir1 and cir2). The CIR signals of the two frequency bands are mathematically expressed as follows:
[0052]
[0053] Where A1 and A2 represent the amplitudes of the two CIR signals, d0 represents the initial path length from the target to the microphone, Δd represents the change in the target path over time, and c represents the speed of sound.
[0054] Step 2: Target occupancy assessment.
[0055] The present invention notes that the movement of targets within the cabin can lead to dynamic changes in the energy of the acoustic signal. Therefore, the present invention proposes an occupancy grid construction technique to represent the dynamic energy distribution in the entire three-dimensional space of the cabin, and further extract the number and location information of targets. The specific steps are as follows:
[0056] Step 1: Signal processing based on difference frequency. Divide the two frequency band CIR signals obtained above and construct a CIR signal based on difference frequency to obtain a "virtual" low-frequency signal with a longer wavelength. The mathematical expression of the CIR signal based on difference frequency is as follows:
[0057]
[0058] Here, Δf = f2 - f1, and A′ = A2 / A1. This difference frequency operation is performed on each microphone in the distributed microphone array, and the target space occupancy grid is then constructed.
[0059] Step 2: Multipath signal mapping for a single microphone. First, the present invention divides the entire three-dimensional space of the cabin into a 2cm×2cm×2cm grid. The difference frequency CIR signal from each microphone is then subjected to multipath signal decomposition. Based on the distance between the target and the microphone, each multipath signal is mapped one-to-one to the grid cell of the corresponding distance in the three-dimensional space.
[0060] Step 3: Multi-microphone signal energy superposition. This method superimposes the multipath signals from all microphones within each grid cell, using this sum to represent the energy intensity of that grid cell. Therefore, the energy distribution across the entire three-dimensional space reflects the target's occupancy within the cabin. By setting an energy intensity threshold, the target's corresponding area can be effectively identified, further enabling the extraction of target quantity and location information.
[0061] Step 3: Spatial sound wave focusing.
[0062] Based on the aforementioned target quantity and location information, the present invention further uses three-dimensional beamforming technology to focus the acoustic signal on the target at each location. The specific steps of the three-dimensional beamforming technology are as follows:
[0063] Step 1: Phase difference modeling. Figure 2 As shown in FIG, the present invention first models the distributed microphone array (16 microphones) in the YOZ plane of the three-dimensional coordinate system. The microphone at the lower left corner of the array is located at the origin of the coordinate system and serves as the reference microphone. Let the rectangular coordinates of the i-th microphone be (x i ,y i ,z i ), the polar coordinate system is The rectangular coordinates of target K are (x K ,y K ,z K ), the polar coordinate system is The conversion relationship between rectangular coordinates and polar coordinates is (taking target K as an example):
[0064]
[0065] Then the distance between target K and the i-th microphone can be expressed as:
[0066]
[0067] Since the microphone is in the YOZ plane, x i =0,θ i =90°.
[0068] Therefore, the phase difference between the received signal of the i-th microphone and the reference microphone is linearly related to the path length difference between the target and the reference microphone. The mathematical expression model is as follows:
[0069]
[0070] Step 2: Steering vector construction. To accurately focus on the target, the present invention uses the previously proposed CIR signal based on the difference frequency to perform three-dimensional beamforming. In this case, the wavelength of the CIR signal is λ = c / Δf. Based on the phase difference caused by the path length difference between the target and different microphones, that is, formula (7), the phase compensation coefficient of the i-th microphone can be expressed as:
[0071]
[0072] The steering vector of beamforming can be expressed as:
[0073] W=[ω1,ω2,,...,ω N ] (9)
[0074] Where N represents the total number of distributed microphones.
[0075] Step 3: Acoustic signal focusing. The present invention compensates the steering vector on the signal in the target grid to obtain a beamforming signal, which is mathematically expressed as:
[0076]
[0077] Among them, cir Δf,i (n, t) represents the CIR signal based on the difference frequency calculated by the i-th microphone receiving signal, CIR Δf =[cir Δf,1 (n,t),cir Δf,2 (n,t),...,cir Δf,N (n,t)] T , (·) T Represents the transpose operation of the matrix. By traversing the azimuth angle θ and the pitch angle When the phase of the steering vector is complementary to the phase of the target signal, a peak will be generated in the two-dimensional angle spectrum, which is the center of the target. At this time, the phase alignment between the acoustic wave signals at the target center realizes signal focusing. The steering vector compensation is based on the single-frequency CIR signal, that is, the steering vector based on formula (9) Compensation is performed on the signal cir1(b,t) or cir2(n,t) to enhance the target perception signal and further extract the required perception information (such as breathing rate, gesture movement, etc.).
[0078] Example Introduction
[0079] The embodiment of the present invention uses an acoustic device embedded in a real car to implement the system of the present invention. The acoustic device includes 4 Bose speakers and 16 distributed microphones. Figure 3 As shown in the figure, in a typical deployment of in-vehicle acoustic components, four speakers are located on the A-pillar (two) and the B-pillar (two). Sixteen microphones are evenly spaced around the entire roof. The speakers and microphones are all connected to a sound card (PXUA216MB-DL2-M / XMOS) to play and collect sound data (transmitted signal parameters are: f1 = 16 kHz, f2 = 19 kHz, B = 2 kHz, T = 0.0427 s). The sound card is then connected to a laptop computer, which receives the sound signals in real time at a 48 kHz sampling rate and further processes them using MATLAB.
[0080] like Figure 4As shown, it includes: (1) multi-target breathing detection; (2) multi-target gesture recognition; (3) driver driving behavior monitoring. The present invention demonstrates the system performance through three applications:
[0081] (1) Multi-target breathing detection: The purpose of multi-target breathing detection is to detect the breathing of multiple passengers at the same time. This application aims to monitor the breathing rate of all passengers in the car in real time, so as to detect abnormal conditions such as apnea, shortness of breath, etc. in time, thereby providing protection for the health of passengers, which is especially important for long-distance driving. Specifically, the present invention detects the breathing signal of each target, compares it with the true value, and then calculates the average breathing rate estimation error of multiple targets as a system performance evaluation indicator. The average breathing rate detection error of the system in scenarios involving 1-4 targets is 0.58 times / minute. Even when there are 4 targets in the car, the average absolute error of the breathing rate estimation does not exceed 0.85 times / minute.
[0082] (2) Multi-target gesture recognition: Multi-target gesture recognition technology can make vehicles more intelligent and provide personalized services and functions based on the gesture habits and preferences of different passengers. The present invention evaluated five gesture types, namely pulling, pushing, push-pull, drawing a triangle, and tapping twice. The present invention collected 1,000 samples for the five gestures to form a data set, and then used a method based on dynamic time matching for gesture recognition. The acoustic focusing technology proposed in the present invention successfully solved the problems of multi-target interference and system instability, and finally achieved an average gesture recognition accuracy of 94.5%.
[0083] (3) Driver driving behavior monitoring: By real-time monitoring of the driver's dangerous driving behaviors, such as fatigue driving and distracted driving, early warnings can be issued in time to remind the driver to pay attention to safety, thereby effectively preventing the occurrence of traffic accidents. The evaluation experiment of this system focuses on four common dangerous driving behaviors: blinking, nodding, yawning (representing three types of fatigue driving behaviors) and turning the head for a long time (representing one type of distracted driving behavior). The present invention focuses the signal on the driver's head and can extract the movement characteristics of the head from the perception signal, especially the distance change and movement duration. Taking into account the signal patterns generated by different driving behaviors, the present invention applies a method based on dynamic time matching to measure the similarity between the current signal and the signal in the template library, and effectively identify driving behavior. This system collected 500 sets of data and achieved an average behavior detection accuracy of 95.9%.
[0084] By using the three application examples described above, this paper demonstrates the system's ability to accurately detect spatial occupancy and multi-target perception, along with strong robustness. This paper argues that focused acoustic signal perception based on distributed acoustic devices is a fundamental capability that can provide vehicles with location-based, multi-target perception information, thereby supporting a wider range of smart cockpit applications.
[0085] The above embodiments are provided for the purpose of describing the present invention only and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the present invention are intended to be within the scope of the present invention.
Claims
1. A sound wave focusing perception method based on a distributed vehicle-mounted microphone array, characterized in that: The method comprises: Based on the dual-band ZC sequence signal, a channel impulse response cir1 and a channel impulse response cir2 of two frequency bands are generated; wherein the dual-band ZC sequence signal is transmitted based on the vehicle-mounted distributed multi-speaker and received by the distributed microphone array; By performing difference frequency processing on the channel impulse responses of the two frequency bands, a difference frequency channel impulse response is obtained, and the number and area of the targets in the cockpit are extracted based on the difference frequency channel impulse response; A microphone in the microphone array is used as a reference microphone, and a polar coordinate system is constructed with the reference microphone as the origin. The azimuth angle θ and pitch angle of the center of the target area in the polar coordinate system are obtained through three-dimensional beamforming technology. Based on the azimuth angle θ and the pitch angle The steering vector is compensated on the channel impulse response cir1 or the channel impulse response cir2, and the compensated channel impulse response is used to extract the perceptual information of the target.
2. The method according to claim 1, characterized in that Based on the dual-band ZC sequence signal, the channel impulse response cir1 and the channel impulse response cir2 of the two frequency bands are generated, including: Two band-pass filters are used to band-pass filter the dual-band ZC sequence signal; A cross-correlation operation is performed on the filtered signal of each frequency band and the dual-band ZC sequence signal to obtain a channel impulse response cir1 and a channel impulse response cir2.
3. The method according to claim 1, characterized in that The number and area of the targets in the cockpit are extracted based on the beat frequency channel impulse response, including: Divide the three-dimensional space of the vehicle cabin into three-dimensional grids; Perform multipath signal decomposition on the beat frequency channel impulse response of each microphone, and map each multipath signal to the corresponding distance grid in the three-dimensional grid according to the distance between the target and the microphone; The multipath signals from all microphones in each grid are superimposed to obtain the energy intensity of the grid; According to the energy intensity of each grid, the number and area of the targets in the cockpit are obtained.
4. The method according to claim 3, characterized in that A microphone in the microphone array is used as a reference microphone, and a polar coordinate system is constructed with the reference microphone as the origin. The azimuth angle θ and pitch angle of the center of the target area in the polar coordinate system are obtained through three-dimensional beamforming technology. include: Establish a three-dimensional coordinate system based on the origin; The distributed microphone array is modeled on the YOZ plane in the three-dimensional coordinate system, and the i-th microphone M is set i The rectangular coordinates are (x i ,y i , z i ), the polar coordinates are The rectangular coordinates of a grid in the target K area are (x K ,y K , z K ), polar coordinates are (d K ,θ K , ), and the conversion relationship between rectangular coordinates and polar coordinates is In the three-dimensional coordinate system and polar coordinate system, the received signal phase difference ΔΦ is modeled i and the path length difference Δd i The mathematical expression between: wherein the received signal phase difference ΔΦ i is the i-th microphone M i Phase difference of the received signal from the reference microphone, path length difference Δd i For this grid, the number of microphones M is i The path length difference to the reference microphone; Based on the received signal phase difference ΔΦ i and the path length difference Δd i The mathematical expression between the two models models the i-th microphone M i Phase compensation coefficient ω i The mathematical expression of Mathematical expression for the steering vector W that models beamforming; where steering vector W = [ω1, ω2, ..., ω N ]; By compensating the steering vector W on the difference frequency channel impulse response of the grid, an expression for the beamformed signal is generated; The azimuth angle θ and the elevation angle in the expression of the traversal beamforming signal are By identifying the peak of the two-dimensional angle spectrum, the azimuth angle θ and pitch angle corresponding to the central grid in the target K area are obtained.
5. The method according to claim 4, characterized in that The path length difference 6. The method according to claim 4, characterized in that In the three-dimensional coordinate system and the polar coordinate system, the received signal phase difference ΔΦ i and the path length difference Δd i The mathematical expression between is:
7. The method according to claim 4, characterized in that The i-th microphone M i Phase compensation coefficient ω i The mathematical expression is: Wherein, Δf represents the frequency of the difference frequency channel impulse response, and c represents the speed of sound.
8. An acoustic wave focusing perception system based on a distributed vehicle-mounted microphone array, characterized in that: The system comprises: A signal preprocessing module is configured to generate a channel impulse response cir1 and a channel impulse response cir2 of two frequency bands based on a dual-band ZC sequence signal, wherein the dual-band ZC sequence signal is transmitted based on a vehicle-mounted distributed multi-speaker and received using a distributed microphone array; A target occupancy assessment module is used to perform difference frequency processing on the channel impulse responses of the two frequency bands to obtain a difference frequency channel impulse response, and extract the number and area of the targets in the cockpit based on the difference frequency channel impulse response; The spatial acoustic wave focusing module is used to use one microphone in the microphone array as a reference microphone, construct a polar coordinate system with the reference microphone as the origin, and obtain the azimuth angle θ and pitch angle of the target area center in the polar coordinate system through three-dimensional beamforming technology. Based on the azimuth angle θ and the pitch angle The steering vector is compensated on the channel impulse response cir1 or the channel impulse response cir2, and the compensated channel impulse response is used to extract the perceptual information of the target.
9. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the sound wave focusing perception method based on a distributed vehicle-mounted microphone array as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the sound wave focusing perception method based on a distributed vehicle-mounted microphone array as described in any one of claims 1 to 7.