Vehicle external sound processing method and system and storage device

By using a spatially distributed microphone array and advanced signal processing technology, it achieves accurate acquisition and high-fidelity reproduction of natural sounds outside the vehicle, solving the problem of not being able to selectively listen to sounds outside the vehicle in a closed environment and providing an immersive interactive experience.

CN121838799APending Publication Date: 2026-04-10JINGDIAN AUTOMOTIVE ELECTRONICS (HUIZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing vehicles cannot effectively filter and reproduce natural sounds from outside the vehicle in a closed state, resulting in energy waste and physical discomfort caused by opening windows.

Method used

By deploying a spatially distributed microphone array to collect external sounds, and using beamforming, adaptive filtering, and deep learning algorithms for noise suppression and target sound separation, high-fidelity reconstruction and output are achieved.

Benefits of technology

Specific natural sounds are precisely extracted and reproduced within a sealed cabin, avoiding temperature fluctuations and physical discomfort caused by opening windows, thus providing an immersive auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838799A_ABST
    Figure CN121838799A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle external sound processing method and system and a storage device. The method comprises the steps that original environment sound signals in multiple directions outside a vehicle are collected; based on a target sound pointing instruction, performing beam forming processing on the original environment sound signal, and extracting a first sound signal in a specific direction; performing noise suppression and target sound separation processing on the first sound signal to obtain a second sound signal; and performing high-fidelity audio reconstruction and amplification on the second sound signal to obtain an audio signal of the target sound, and outputting the audio signal to an in-vehicle loudspeaker. According to the method provided by the invention, directional screening and noise reduction of the environment sound outside the vehicle and high-fidelity playback in the cabin can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent cockpit technology, and in particular to a method, system and storage device for processing external sounds of a vehicle. Background Technology

[0002] In the modern automotive industry, NVH (Noise, Vibration, and Harshness) performance optimization has become a core indicator for enhancing vehicle luxury and ride comfort. Through the widespread application of technologies such as multi-layered soundproof glass, high-density sealing strips, and acoustic packaging materials, vehicles can effectively block external environmental noise (such as traffic noise and wind noise) in a sealed state, creating a highly quiet cabin space for occupants. However, this fully isolated design has significant limitations in certain user experience scenarios. For example, in scenarios such as outdoor camping, forest driving, lakeside relaxation, or stargazing at night, occupants expect to perceive natural white noise such as birdsong, insect chirping, flowing streams, and rustling leaves to achieve immersive interaction with the environment. However, occupants must open the windows to access natural sound sources. This leads to multiple problems, such as a sharp drop in air conditioning / heating efficiency due to open windows, resulting in significant energy waste; the intrusion of harmful substances such as mosquitoes, dust, and pollen into the cabin; and the weakening of warning information such as approaching wild animals and sudden external noises by the window barrier. Existing vehicle audio technology cannot achieve directional filtering and noise reduction of external ambient sounds, or high-fidelity playback within the cabin. Summary of the Invention

[0003] This application provides a method, system, and storage device for processing external sounds of a vehicle to solve the above-mentioned technical problems.

[0004] A method for processing external vehicle sound includes: acquiring raw ambient sound signals from multiple directions outside the vehicle; performing beamforming processing on the raw ambient sound signals based on a target sound pointing command to extract a first sound signal from a specific direction; performing noise suppression and target sound separation processing on the first sound signal to obtain a second sound signal; performing high-fidelity audio reconstruction and amplification on the second sound signal to obtain an audio signal of the target sound, and outputting it to an in-vehicle speaker.

[0005] This solution allows users to selectively listen to specific natural sounds within the enclosed cabin, such as birdsong to the left or the sound of a stream in front, enjoying an immersive environmental interaction without opening windows. It avoids temperature fluctuations (sudden drops in air conditioning / heating efficiency) and physical discomfort (such as wind noise) caused by opening windows, ensuring passengers can comfortably experience natural soundscapes even in extreme weather or sensitive environments (such as pollen season), addressing a core pain point in outdoor leisure scenarios while maintaining cabin tranquility and luxury. Specifically, it collects raw sound through a multi-directional microphone array and performs beamforming processing based on the target sound direction command to accurately extract the first sound signal from a specific direction. This achieves spatial sound filtering, allows for customization of the sound source direction, and enhances the personalization and flexibility of the interaction. Simultaneously, advanced signal processing algorithms suppress noise in the first sound signal, effectively removing interference such as wind noise and engine sounds, while separating the pure target sound, overcoming the inability to distinguish between environmental noise and the target sound source. High-fidelity reconstruction and amplification of the second sound signal ensures the realistic reproduction of natural sound details, providing an immersive auditory experience through the in-vehicle speaker system.

[0006] Furthermore, the acquisition of raw ambient sound signals from multiple directions outside the vehicle includes: acquiring raw ambient sound signals from multiple directions outside the vehicle through a spatially distributed microphone array deployed outside the vehicle.

[0007] In this solution, the spatially distributed microphone array includes a front microphone integrated at the front of the vehicle for collecting sound from the front area; a rear microphone integrated at the rear of the vehicle for collecting sound from the rear area; and side microphones integrated on both sides of the vehicle for collecting sound from the sides. This spatially distributed microphone array enables 360° omnidirectional sound source acquisition.

[0008] Furthermore, the beamforming process for the original ambient sound signal includes: determining the target sound direction; determining the relative delay of the received signals of each microphone in the microphone array based on the target sound direction; performing delay compensation processing on the received signals of each microphone based on the relative delay; and performing fusion processing on the compensated sound signals of each microphone based on the microphone array weights.

[0009] In this scheme, the target sound direction can be determined by the user or automatically by a sound source localization algorithm. After determining the target sound direction, the sound source position is locked. By calculating the distance difference between the sound source and each microphone, the relative delay of the received signal of each microphone is determined, and the signal delay is dynamically adjusted so that the same sound signal from the target direction on all microphone channels is perfectly aligned in time, achieving phase synchronization. By assigning different importance to the microphone signals at different positions, a weighting coefficient is multiplied before adding the signals of each microphone, and the amplitude of the signal of that channel is adjusted on the basis of delay compensation. All microphone signals that have undergone delay and weighting are summed to output an audio signal that is enhanced from a specific direction while the sound from other directions is suppressed. That is, the target sound is coherently superimposed due to the same phase, which significantly enhances the signal amplitude, while noise from non-target directions is canceled and suppressed in the superposition because the phase cannot be aligned.

[0010] Furthermore, the beamforming processing of the original ambient sound signal further includes: calculating the spatial characteristics of the current ambient noise field based on the original ambient sound signal acquired in real time by the microphone array, and then determining the microphone array weights.

[0011] In this scheme, an adaptive beamforming algorithm is used to beamform the original ambient sound signal. Instead of using preset fixed weights, the optimal weights are dynamically calculated based on the real-time acquired sound field environment. This actively and accurately suppresses the strongest interference sources while ensuring lossless sound in the target direction. Specifically, based on the acquired audio signal, the spatial correlation characteristics of the current noise field are calculated and estimated. With signal distortion-free operation in the target direction as a constraint and minimizing the total output power of the array as the objective, a set of optimal microphone array weights can be obtained.

[0012] Further, the step of performing noise suppression and target sound separation processing on the first sound signal to obtain the second sound signal includes: extracting time-frequency features from the first sound signal; determining a mask value corresponding to the target sound in the first sound signal based on the time-frequency features; and separating and extracting the second sound signal from the first sound signal based on the mask value.

[0013] In this scheme, time-frequency features are extracted from the first sound signal, converting the one-dimensional time-domain waveform signal into a two-dimensional time-frequency representation (time-frequency feature map), which serves as the input to a pre-trained deep learning model. The deep learning model estimates the probability of the target sound's presence or energy proportion for each time-frequency point, using methods such as an ideal soft mask or a ratio mask. The ideal soft mask typically has a value range between [0, 1], representing the posterior probability that the time-frequency point belongs to the target sound. The ratio mask typically has a value range between [0, ∞), accurately representing the energy ratio between the target sound and the mixed sound. The mask value predicted by the deep learning model is applied to the first sound signal, filtering it in the time-frequency domain and reconstructing a clean time-domain signal. Specifically, the predicted mask value is multiplied by the time-frequency representation of the first sound signal to obtain the filtered complex spectrum. An inverse short-time Fourier transform is then performed on the filtered complex spectrum to obtain the second sound signal in the time domain, i.e., the enhanced target sound.

[0014] Furthermore, before performing noise suppression and target sound separation processing on the first sound signal, the method further includes: performing adaptive filtering processing on the first sound signal to suppress non-stationary environmental noise signals in the first sound signal.

[0015] In this scheme, non-stationary noise (such as sudden traffic horns and construction machinery noise) has time-varying characteristics, which traditional fixed filtering parameters cannot effectively track. Adaptive filtering, by updating the filter coefficients in real time, can initially suppress non-stationary environmental noise (such as sudden traffic horns and construction machinery noise), reducing the noise modeling pressure of subsequent separation models. Using the LMS / NLMS adaptive algorithm, the temporal phase characteristics of the target sound can be completely preserved while suppressing noise, avoiding signal distortion caused by pure frequency domain filtering. Combined with the frequency domain adaptive filter, high-frequency details of natural sounds such as birdsong and streams can be specifically preserved. Specifically, ambient pure noise is collected through a microphone array as a reference noise signal; the coefficients of the adaptive filter are dynamically updated using the least mean square or normalized least mean square algorithm, so that the output of the filter can best predict the noise components in the first sound signal; the output of the adaptive filter is subtracted from the first sound signal to generate an enhanced signal that has been suppressed by non-stationary noise, and this enhanced signal is used as the input for subsequent deep learning model processing.

[0016] Furthermore, before performing noise suppression and target sound separation processing on the first sound signal, the method further includes: performing echo cancellation processing on the first sound signal to eliminate echo noise signals generated by the in-vehicle audio playback in the first sound signal.

[0017] In this scheme, echo cancellation processing is added before noise suppression and target sound separation, which can eliminate acoustic / circuit echo interference at the source and reduce the noise modeling pressure of subsequent separation models. Specifically, the audio signal currently being played or about to be played by the in-vehicle audio system is obtained as a reference signal; based on the reference signal, an acoustic echo path of sound from the in-vehicle speakers to the microphone array is simulated through an adaptive filter to generate an echo estimation signal; the echo estimation signal is subtracted from the first sound signal to eliminate the acoustic echo generated by the in-vehicle sound being amplified and fed back to the microphone, and the echo-free signal is used as the input for subsequent deep learning model processing.

[0018] Furthermore, before performing noise suppression and target sound separation processing on the first sound signal, the method further includes: extracting the steady-state background noise spectrum from the first sound signal, and subtracting the steady-state background noise spectrum from the spectrum of the first sound signal to suppress the background noise signal in the first sound signal.

[0019] In this scheme, the steady-state background noise spectrum is extracted and pre-filtered through spectral subtraction or Wiener filtering to remove long-term steady-state interference such as wind noise and low-frequency engine noise, reducing the noise modeling pressure of subsequent separation models. Specifically, taking spectral subtraction as an example, a short-time Fourier transform is performed on the first sound signal to obtain its short-time amplitude spectrum or power spectrum; the estimated value of the noise spectrum is subtracted from the spectrum of the first sound signal, and the result of the subtraction is subject to appropriate over-subtraction factor and spectral lower limit control; the processed spectrum is combined with the original phase and an inverse short-time Fourier transform is performed to reconstruct the time-domain signal that has suppressed the steady-state background noise, and this signal is used as the input for subsequent deep learning model processing.

[0020] Based on the same concept, a vehicle external sound processing system is also proposed, comprising: a microphone array including multiple microphones arranged in different spatial directions outside the vehicle for acquiring original ambient sound signals from multiple directions outside the vehicle; a processing module for performing beamforming processing on the original ambient sound signals based on a target sound pointing command to extract a first sound signal in a specific direction; performing noise suppression and target sound separation processing on the first sound signal to obtain a second sound signal; and an audio output module for performing high-fidelity audio reconstruction and amplification on the second sound signal to obtain an audio signal of the target sound, and outputting it to an in-vehicle speaker.

[0021] Based on the same concept, a computer storage device is also proposed, which stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to execute the aforementioned vehicle external sound processing method.

[0022] Compared with the prior art, the beneficial effects of this application are as follows: This solution allows users to selectively listen to specific natural sounds within the enclosed cabin, such as birdsong to the left or the sound of a stream in front, enjoying an immersive environmental interaction without opening windows. It avoids temperature fluctuations (sudden drops in air conditioning / heating efficiency) and physical discomfort (such as wind noise) caused by opening windows, ensuring passengers can comfortably experience natural soundscapes even in extreme weather or sensitive environments (such as pollen season), addressing a core pain point in outdoor leisure scenarios while maintaining cabin tranquility and luxury. Specifically, it collects raw sound through a multi-directional microphone array and performs beamforming processing based on the target sound direction command to accurately extract the first sound signal from a specific direction. This achieves spatial sound filtering, allows for customization of the sound source direction, and enhances the personalization and flexibility of the interaction. Simultaneously, advanced signal processing algorithms suppress noise in the first sound signal, effectively removing interference such as wind noise and engine sounds, while separating the pure target sound, overcoming the inability to distinguish between environmental noise and the target sound source. High-fidelity reconstruction and amplification of the second sound signal ensures the realistic reproduction of natural sound details, providing an immersive auditory experience through the in-vehicle speaker system. Attached Figure Description

[0023] Figure 1 This is a flowchart of the vehicle external sound processing method described in this application.

[0024] Figure 2 This is a flowchart illustrating the beamforming process of the original environmental sound signal as described in this application.

[0025] Figure 3 This is a flowchart illustrating the noise suppression and target sound separation processes performed on the first sound signal as described in this application.

[0026] Figure 4 This is a schematic diagram of the vehicle external sound processing system described in this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application. Example 1

[0028] like Figure 1 As shown, this embodiment provides a method for processing external vehicle sounds, including the following steps S10 to S40.

[0029] S10 collects raw ambient sound signals from multiple directions outside the vehicle.

[0030] S20, based on the target sound pointing instruction, performs beamforming processing on the original environmental sound signal to extract the first sound signal in a specific direction.

[0031] S30, perform noise suppression and target sound separation processing on the first sound signal to obtain the second sound signal.

[0032] S40 performs high-fidelity audio reconstruction and amplification on the second sound signal to obtain the audio signal of the target sound, and outputs it to the in-vehicle speakers.

[0033] It should be noted that in step S10, the raw ambient sound signals from multiple directions outside the vehicle are collected by a spatially distributed microphone array deployed outside the vehicle. Specifically, the spatially distributed microphone array includes a front microphone integrated at the front of the vehicle for collecting sound from the front area, which can be integrated around the windshield or inside the rearview mirror housing, containing 2-3 high-performance MEMS microphone units; a rear microphone integrated at the rear of the vehicle for collecting sound from the rear area, which can be integrated above the license plate frame or near the rear window, containing 2-3 high-performance MEMS microphone units; and side microphones integrated on both sides of the vehicle for collecting sound from the sides, which can be integrated on the B-pillar or outside the windows, with at least one high-performance MEMS microphone unit on each side. This spatially distributed microphone array enables 360° omnidirectional sound source acquisition, providing a hardware foundation for sound source localization and multi-array data fusion.

[0034] In step S20, the target sound pointing instruction includes the direction instruction of the target sound, which is selected by the user or automatically determined by the sound source localization algorithm. When the target direction is selected by the user, an azimuth and pitch angle instruction can be generated through the interactive interface of the central control screen (such as clicking on a direction in the surround vehicle view). When the direction is automatically determined by the sound source localization algorithm, a low-computational-load global sound source monitoring algorithm, such as the time delay estimation method based on the generalized cross-correlation function, continuously runs to scan the data of each subarray in real time, quickly detect and mark several potential sound sources of interest with sudden energy increases or typical spectral characteristics and their approximate directions. When a specific type of sound source is detected or the user triggers the function to enhance the sound in a certain direction, the system starts a high-precision sound source localization algorithm. This algorithm uses the time delay information of the entire large-aperture distributed array, employs a subspace method, or a controllable beamforming scanning method based on maximum output power, to perform a fine search in the approximate azimuth area of ​​the target, calculate the precise three-dimensional coordinates of the sound source (relative to the vehicle coordinate system), and automatically generate the corresponding target sound direction instruction.

[0035] It should be noted that the target sound pointing command includes not only direction commands but also sound type commands.

[0036] like Figure 2 As shown, in step S20, beamforming processing is performed on the original ambient sound signal, specifically including the following steps S201 to S204. S201, Determine the direction of the target sound.

[0037] S202, determine the relative delay of the received signal of each microphone in the microphone array according to the direction of the target sound.

[0038] S203 performs delay compensation processing on the signals received by each microphone based on relative delay.

[0039] S204 performs fusion processing on the audio signals after compensation processing of each microphone based on the microphone array weights.

[0040] The target sound direction can be determined by the user or automatically by a sound source localization algorithm. After determining the target sound direction, the sound source position is locked and converted into a unit direction vector relative to the array's geometric center. Based on the sound speed model and the target direction vector, the distance difference between the sound source and each microphone is calculated to determine the relative delay of the received signal at each microphone. The signal delay is dynamically adjusted to ensure that the same sound signal from the target direction is perfectly aligned in time on all microphone channels, achieving phase synchronization. By assigning different importance to the microphone signals at different locations, a weighting coefficient is multiplied before adding the signals from each microphone to adjust the amplitude of the signal in that channel based on delay compensation. All delayed and weighted microphone signals are summed to output an audio signal that is enhanced from a specific direction while suppressing sound from other directions. That is, the target sound is coherently superimposed due to the same phase, resulting in a significant increase in signal amplitude, while noise from non-target directions is canceled out and suppressed during superposition because the phases cannot be aligned.

[0041] In step S204, the microphone array weights are dynamically changed based on the real-time acquired sound field environment. Specifically, the microphone array weights are determined by calculating the spatial characteristics of the current environmental noise field based on the original environmental sound signals acquired in real time by the microphone array.

[0042] An adaptive beamforming algorithm is used to process the raw ambient sound signal. Instead of using preset fixed weights, it dynamically calculates the optimal weights based on the real-time acquired sound field environment. This actively and accurately suppresses the strongest interference sources while ensuring lossless sound in the target direction. Specifically, based on the acquired audio signal, the spatial correlation characteristics of the current noise field are calculated and estimated. With lossless signal in the target direction as a constraint and minimizing the total output power of the array as the objective, an optimal set of microphone array weights can be obtained.

[0043] like Figure 3 As shown, in step S30, the first sound signal is subjected to noise suppression and target sound separation processing to obtain the second sound signal, including the following steps S301 to S303.

[0044] S301, extract time-frequency features from the first sound signal.

[0045] S302, based on time-frequency characteristics, determine the mask value corresponding to the target sound in the first sound signal.

[0046] S303, based on the mask value, separate and extract the second sound signal from the first sound signal.

[0047] Specifically, time-frequency features are extracted from the first sound signal, converting the one-dimensional time-domain waveform signal into a two-dimensional time-frequency representation (time-frequency feature map), which serves as input to a pre-trained deep learning model. Using the deep learning model, the probability of the target sound's presence or energy proportion is estimated for each time-frequency point, such as using an ideal soft mask or a ratio mask. The value range of an ideal soft mask is typically between [0, 1], representing the posterior probability that the time-frequency point belongs to the target sound. The value range of a ratio mask is typically between [0, ∞), accurately representing the energy ratio between the target sound and the mixed sound. The mask values ​​predicted by the deep learning model are applied to the first sound signal, filtering it in the time-frequency domain, and reconstructing a clean time-domain signal. Specifically, the predicted mask values ​​are multiplied by the time-frequency representation of the first sound signal to obtain the filtered complex spectrum. An inverse short-time Fourier transform is then performed on the filtered complex spectrum to obtain the second sound signal in the time domain, i.e., the enhanced target sound.

[0048] It should be noted that the pre-trained deep learning model analyzes the time-frequency features of the input to predict a time-frequency mask that can accurately separate the target sound. The deep learning model can be a convolutional recurrent neural network (CRNN). Specifically, the front-end CNN (2D convolutional layers) is responsible for extracting local time-frequency pattern features (such as harmonic stripes and impact transients), possessing translation invariance and effectively learning the local structure of the sound; the back-end RNN (bidirectional LSTM / GRU) is responsible for modeling long-term time-series dependencies and understanding the dynamic changes in the sound. Of course, the deep learning model can also be a recurrent neural network (RNN) or a temporal convolutional network (TCN).

[0049] Before performing noise suppression and target sound separation processing on the first sound signal in step S30, the method further includes: performing adaptive filtering processing on the first sound signal to suppress non-stationary environmental noise signals in the first sound signal.

[0050] Non-stationary noise (such as sudden traffic horns and construction machinery noise) has time-varying characteristics, which traditional fixed filtering parameters cannot effectively track. Adaptive filtering, by updating the filter coefficients in real time, can initially suppress non-stationary environmental noise (such as sudden traffic horns and construction machinery noise), reducing the noise modeling pressure of subsequent separation models. Using LMS / NLMS adaptive algorithms can completely preserve the temporal phase characteristics of the target sound while suppressing noise, avoiding signal distortion caused by pure frequency domain filtering. Combined with frequency domain adaptive filters, high-frequency details of natural sounds such as birdsong and streams can be specifically preserved. Specifically, ambient pure noise is collected through a microphone array as a reference noise signal; the coefficients of the adaptive filter are dynamically updated using the least mean square or normalized least mean square algorithm, so that the output of the filter can best predict the noise components in the first sound signal; the output of the adaptive filter is subtracted from the first sound signal to generate an enhanced signal that has been suppressed by non-stationary noise, and this enhanced signal is used as the input for subsequent deep learning model processing.

[0051] Before performing noise suppression and target sound separation processing on the first sound signal in step S30, the process may further include: performing echo cancellation processing on the first sound signal to eliminate echo noise signals generated by the in-vehicle audio playback in the first sound signal.

[0052] Adding echo cancellation processing before noise suppression and target sound separation can eliminate acoustic / circuit echo interference at its source, reducing the noise modeling burden on subsequent separation models. Specifically, the audio signal currently being played or about to be played by the in-vehicle audio system is acquired as a reference signal; based on the reference signal, an acoustic echo path from the in-vehicle speakers to the microphone array is simulated using an adaptive filter to generate an echo estimation signal; the echo estimation signal is subtracted from the first sound signal to eliminate the acoustic echo generated by the in-vehicle sound being amplified and fed back to the microphone, and the echo-free signal is used as the input for subsequent deep learning model processing.

[0053] In step S30, before performing noise suppression and target sound separation processing on the first sound signal, the method may further include: extracting the steady-state background noise spectrum from the first sound signal and subtracting the steady-state background noise spectrum from the spectrum of the first sound signal to suppress the background noise signal in the first sound signal.

[0054] Steady-state background noise spectra are extracted and pre-filtered using spectral subtraction or Wiener filtering to remove long-term steady-state interference such as wind noise and low-frequency engine rumble, reducing the noise modeling burden on subsequent separation models. Specifically, taking spectral subtraction as an example, a short-time Fourier transform is performed on the first sound signal to obtain its short-time amplitude spectrum or power spectrum; the estimated value of the noise spectrum is subtracted from the spectrum of the first sound signal, and the result of the subtraction is subject to appropriate over-subtraction factors and spectral lower limits; the processed spectrum is combined with the original phase and an inverse short-time Fourier transform is performed to reconstruct the time-domain signal that has suppressed steady-state background noise, and this signal is used as the input for subsequent deep learning model processing.

[0055] In step S40, the second sound signal is reconstructed and amplified using high fidelity audio to obtain the audio signal of the target sound. Specifically, a high-definition audio decoder, such as the ESP9038PRO, can be used to perform high-definition audio decoding and output high-quality audio. It intelligently compresses the dynamic range to avoid sudden loudness changes and performs frequency response correction based on the vehicle's internal acoustic characteristics to compensate for inherent defects in the vehicle's speaker-cabin system. Its audio gain control capability allows for volume adjustment as needed, amplifying the target sound. Furthermore, the ESP9038PRO's built-in digital signal processor provides rich audio gain control options, allowing for adjustments to parameters such as balance, volume, and gain. A power amplification system, such as the SABRE9601K+TPA3118 amplifier, further amplifies the audio signal to drive the speakers. The SABRE9601K, acting as a preamplifier and line driver, provides a high-quality, low-distortion line-level signal, while the TPA3118 amplifier, as a high-efficiency power amplifier, drives the vehicle speakers and can be flexibly configured to drive all vehicle speakers or speakers in specific areas.

[0056] This embodiment allows users to selectively listen to specific natural sounds within the enclosed cabin, such as birdsong to the left or the sound of a stream in front, enjoying an immersive environmental interaction without opening windows. It avoids temperature fluctuations (sudden drops in air conditioning / heating efficiency) and physical discomfort (such as wind noise) caused by opening windows, ensuring occupants can comfortably experience natural soundscapes even in extreme weather or sensitive environments (such as pollen season), addressing a core pain point in outdoor leisure scenarios while maintaining the cabin's tranquility and luxury. Specifically, it uses a multi-directional microphone array to collect raw sound and performs beamforming processing based on the target sound direction command to accurately extract the first sound signal from a specific direction. This achieves spatial sound filtering, allows for customization of the sound source direction, and enhances the personalization and flexibility of the interaction. Simultaneously, advanced signal processing algorithms are used to suppress noise in the first sound signal, effectively removing interference such as wind noise and engine noise, while separating the pure target sound, overcoming the inability to distinguish between environmental noise and the target sound source. High-fidelity reconstruction and amplification of the second sound signal ensures the realistic reproduction of natural sound details, providing an immersive auditory experience through the in-vehicle speaker system. Example 2

[0057] like Figure 4 As shown, this embodiment proposes a vehicle external sound processing system, including: a microphone array 100, a processing module 200, an audio output module 300, and a vehicle speaker 400.

[0058] The microphone array 100 includes multiple microphones arranged in different spatial directions outside the vehicle to collect raw ambient sound signals from multiple directions outside the vehicle. Specifically, the spatially distributed microphone array includes a front microphone integrated at the front of the vehicle for collecting sound from the front area, which can be integrated around the windshield or inside the rearview mirror housing, containing 2-3 high-performance MEMS microphone units; a rear microphone integrated at the rear of the vehicle for collecting sound from the rear area, which can be integrated above the license plate frame or near the rear window, containing 2-3 high-performance MEMS microphone units; and side microphones integrated on both sides of the vehicle for collecting sound from the sides, which can be integrated on the B-pillar or the outside of the windows, with at least one high-performance MEMS microphone unit on each side. This spatially distributed microphone array enables 360° omnidirectional sound source acquisition, providing a hardware foundation for sound source localization and multi-array data fusion.

[0059] The processing module 200, electrically connected to the microphone array 100, includes a system control and scheduling unit 201 and a central signal processing unit 202. The system control and scheduling unit employs a low-power microcontroller based on a Cortex-M4 core and integrating a DSP instruction set, used for running system control, power management, and diagnostic programs. The central signal processing unit employs a high-performance digital signal processor ADSP-21569, used for beamforming processing of the original ambient sound signal based on the target sound pointing command, extracting a first sound signal from a specific direction; and performing noise suppression and target sound separation processing on the first sound signal to obtain a second sound signal. An audio output module 300, electrically connected to the signal processing module 200, is used to perform high-fidelity audio reconstruction and amplification on the second sound signal to obtain the audio signal of the target sound, and output it to the in-vehicle speaker 400. The audio output module 300 includes: a high-fidelity audio decoding unit 301, employing an audio decoding chip ESP9038PRO, used to decode and optimize the sound quality of the processed digital audio signal; and a power amplification unit 302, employing a combination of an audio driver chip SABRE9601 and a Class D power amplifier TPA3118, used to amplify the decoded audio signal and drive the in-vehicle speaker 400. The 400 car speaker is located inside the cabin and is used to play decoded audio. It can be headphones or an external speaker.

[0060] In addition, the vehicle's external sound processing system also includes a central control screen 500, which integrates a user control interface to receive the user's direction or type of the target sound and transmit the command to the processing module 200.

[0061] The vehicle's external sound processing system also includes a power module 600, which includes a main power supply, a backup power supply, and a power management unit. The main power supply is the vehicle's battery, which provides a stable voltage through a DC-DC converter. The backup power supply uses a 2000mAh battery and automatically switches when the main power supply is insufficient. The power management unit uses intelligent power control to adjust the system power consumption according to the vehicle's status. Example 3

[0062] This application also provides a computer storage device. The methods described in the embodiments of this application can be implemented in hardware or firmware, or implemented as computer instruction code that can be recorded on a computer storage device, or implemented as computer instruction code downloaded via a network and originally stored in a remote computer storage device or a non-transitory machine-readable computer storage device and then stored in a local computer storage device. Thus, the methods described herein can be stored in software processing on a computer storage device using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The computer storage device can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the computer storage device may also include combinations of the above types of memory. It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer instruction code. When the computer instruction code is accessed and executed by the computer, processor, or hardware, the above-described vehicle external sound processing method is implemented.

[0063] Although exemplary embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above exemplary embodiments are merely illustrative and are not intended to limit the scope of this application. Various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of this application. All such changes and modifications are intended to be included within the scope of this application as claimed in the appended claims.

[0064] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0065] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed.

[0066] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules according to the embodiments of this application. This application can also be implemented as an apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0067] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0068] Although the description of this application has been made in conjunction with the specific embodiments described above, it will be apparent to those skilled in the art that many substitutions, modifications, and variations can be made based on the foregoing. Therefore, all such substitutions, modifications, and variations are included within the spirit and scope of the appended claims.

Claims

1. A method for processing external vehicle sounds, characterized in that, The method includes: Collect raw ambient sound signals from multiple directions outside the vehicle; Based on the target sound pointing command, beamforming processing is performed on the original environmental sound signal to extract the first sound signal in a specific direction; The first sound signal is subjected to noise suppression and target sound separation processing to obtain the second sound signal; The second sound signal is reconstructed and amplified using high fidelity audio to obtain the audio signal of the target sound, which is then output to the in-vehicle speaker.

2. The vehicle external sound processing method according to claim 1, characterized in that, The acquisition of raw ambient sound signals from multiple directions outside the vehicle includes: A spatially distributed microphone array deployed outside the vehicle is used to collect raw ambient sound signals from multiple directions outside the vehicle.

3. The vehicle external sound processing method according to claim 2, characterized in that, The beamforming process for the original environmental sound signal includes: Determine the direction of the target sound; Based on the direction of the target sound, determine the relative delay of the signal received by each microphone in the microphone array; Based on the relative delay, delay compensation processing is performed on the signals received by each microphone. Based on the microphone array weights, the audio signals after compensation processing of each microphone are fused.

4. The vehicle external sound processing method according to claim 3, characterized in that, The beamforming process for the original environmental sound signal further includes: Based on the raw ambient sound signals collected in real time by the microphone array, the spatial characteristics of the current ambient noise field are calculated, and then the weights of the microphone array are determined.

5. The vehicle external sound processing method according to claim 1, characterized in that, The step of performing noise suppression and target sound separation processing on the first sound signal to obtain the second sound signal includes: Extract time-frequency features from the first sound signal; Based on the time-frequency characteristics, determine the mask value corresponding to the target sound in the first sound signal; Based on the mask value, the second sound signal is extracted from the first sound signal.

6. The vehicle external sound processing method according to claim 1, characterized in that, Before performing noise suppression and target sound separation processing on the first sound signal, the method further includes: performing adaptive filtering processing on the first sound signal to suppress non-stationary environmental noise signals in the first sound signal.

7. The vehicle external sound processing method according to claim 1, characterized in that, Before performing noise suppression and target sound separation processing on the first sound signal, the method further includes: performing echo cancellation processing on the first sound signal to eliminate echo noise signals generated by the in-vehicle audio playback in the first sound signal.

8. The vehicle external sound processing method according to claim 1, characterized in that, Before performing noise suppression and target sound separation processing on the first sound signal, the method further includes: extracting the steady-state background noise spectrum from the first sound signal, and subtracting the steady-state background noise spectrum from the spectrum of the first sound signal to suppress the background noise signal in the first sound signal.

9. A vehicle external sound processing system, characterized in that, include: A microphone array, comprising multiple microphones arranged in different spatial directions outside the vehicle, for acquiring raw ambient sound signals from multiple directions outside the vehicle; The processing module is used to perform beamforming processing on the original ambient sound signal based on the target sound pointing command, extract a first sound signal in a specific direction, and perform noise suppression and target sound separation processing on the first sound signal to obtain a second sound signal. An audio output module is used to perform high-fidelity audio reconstruction and amplification on the second sound signal to obtain the audio signal of the target sound, and output it to the in-vehicle speaker.

10. A computer storage device, characterized in that, The computer storage device stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 8.