Microphone performance test method, apparatus, device, medium, and program product

CN122846014APending Publication Date: 2026-09-29SHENZHEN BLUEWHALE INTEL-CONNECTIVITY S&T CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611157860.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]本发明提供一种麦克性能测试方法、装置、设备、介质及程序产品,以解决在噪声环境下出现播音不清晰的问题

Benefits of technology

[0009]上述麦克性能测试方法、装置、设备、介质及程序产品所提供的一个方案中,通过车内外部预置的声音感知系统同步采集原始声学信号,基于该信号对车载麦克系统进行噪声干扰模拟,生成受干扰预测信息,依据预测信息对系统进行发声适应性分析,得出理论最优发声模式,按该模式驱动麦克系统发声测试,并用同一声音感知系统采集结果,分析得到当前噪声环境下的真实发声性能,本发明构建了真实噪声采集—干扰预测—参数寻优—实车验证的闭环测试链路,基于实际非稳态噪声场主动探索系统可达的最佳拾音参数,而非被动评估固定设置,充分挖掘降噪与波束成形潜力,测试结果高度贴近真实驾驶体验,可精准指导麦克风系统标定与座舱声学优化,解决了在噪声环境下出现播音不清晰的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122846014A_ABST
    Figure CN122846014A_ABST
Patent Text Reader

Abstract

The application is suitable for the field of audio control technology, and provides a microphone performance test method, device, equipment, medium and program product. The sound perception system preset in the vehicle body is used to collect the sound signals inside and outside the vehicle body to obtain original perception information. Noise interference simulation is performed on the vehicle-mounted microphone system according to the original perception information to generate interference prediction information of the vehicle-mounted microphone system. Sound emission adaptability analysis is performed on the vehicle-mounted microphone system based on the interference prediction information to obtain a theoretical sound emission mode of the vehicle-mounted microphone system. Sound emission test is performed on the vehicle-mounted microphone system according to the theoretical sound emission mode, and the sound emission test result is obtained through the sound perception system to analyze the sound emission performance information of the vehicle-mounted microphone system in the current noise environment, thereby solving the problem of unclear broadcast in the noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio control technology, and in particular to a microphone performance testing method, apparatus, device, medium, and program product. Background Technology

[0002] As the core interactive interface of the smart cockpit, the in-vehicle microphone system directly determines the quality of in-vehicle calls and the success rate of voice control. When a vehicle is in motion, it faces a complex sound field formed by wind noise, road noise, powertrain noise, and cabin reverberation, which seriously interferes with the microphone's extraction of near-field speech. In existing technologies, fixed voice parameters are usually used to control the in-vehicle microphone system for voice playback, which can easily lead to unclear broadcasting in noisy environments. Summary of the Invention

[0003] This invention provides a microphone performance testing method, apparatus, equipment, medium, and program product to solve the problem of unclear broadcasting in noisy environments.

[0004] In a first aspect, embodiments of this application provide a microphone performance testing method, including: The sound sensing system pre-installed in the vehicle body collects sound signals from inside and outside the vehicle to obtain raw sensing information; Based on the original sensing information, noise interference simulation is performed on the vehicle microphone system to generate interference prediction information for the vehicle microphone system. Based on the interference prediction information, a sound production adaptability analysis is performed on the vehicle microphone system to obtain the theoretical sound production mode of the vehicle microphone system. The vehicle microphone system is tested for sound output based on the theoretical sound output mode, and the sound output test results are obtained through the sound perception system to analyze the sound output performance information of the vehicle microphone system in the current noise environment.

[0005] Secondly, embodiments of this application provide a microphone performance testing device, comprising: The sound perception module is used to collect sound signals from inside and outside the vehicle through a sound perception system pre-installed in the vehicle body to obtain raw perception information; The interference prediction module is used to simulate noise interference on the vehicle microphone system based on the original sensing information to generate interference prediction information for the vehicle microphone system. The adaptive analysis module is used to perform sound adaptation analysis on the vehicle microphone system based on the interference prediction information to obtain the theoretical sound production mode of the vehicle microphone system. The sound production test module is used to perform a sound production test on the vehicle microphone system according to the theoretical sound production mode, and to obtain the sound production test results through the sound perception system in order to analyze and obtain the sound production performance information of the vehicle microphone system in the current noise environment.

[0006] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described microphone performance testing method.

[0007] Fourthly, embodiments of this application provide a readable storage medium storing a computer program that, when executed by a processor, implements the steps of the microphone performance testing method described above.

[0008] Fifthly, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, enables the implementation of the steps of the microphone performance testing method described above.

[0009] In one of the solutions provided by the aforementioned microphone performance testing methods, devices, equipment, media, and programs, raw acoustic signals are synchronously collected through a pre-installed sound perception system inside and outside the vehicle. Based on these signals, noise interference is simulated on the in-vehicle microphone system to generate interference prediction information. Based on the prediction information, the system's sound adaptation is analyzed to derive the theoretically optimal sound mode. The microphone system is then driven to perform sound testing according to this mode. The results are collected using the same sound perception system, and the actual sound performance under the current noise environment is analyzed. This invention constructs a closed-loop test link of real noise acquisition—interference prediction—parameter optimization—real vehicle verification. Based on the actual unsteady noise field, it actively explores the optimal pickup parameters that the system can achieve, rather than passively evaluating fixed settings. It fully explores the potential of noise reduction and beamforming. The test results are highly close to the real driving experience and can accurately guide microphone system calibration and cabin acoustic optimization, solving the problem of unclear broadcasting in noisy environments. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a microphone performance testing method according to an embodiment of the present invention; Figure 2 yes Figure 1 A schematic diagram of the implementation process of step S10; Figure 3 yes Figure 1 A schematic diagram of the implementation process of step S20; Figure 4 yes Figure 1 A schematic diagram of the implementation process of step S30; Figure 5 yes Figure 1 A schematic diagram of the implementation process of step S40; Figure 6 This is a schematic diagram of a microphone performance testing device according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that, as used in this specification and the appended claims, the term "and / or" refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0014] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0015] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0016] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0017] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0018] In one embodiment, such as Figure 1 As shown, a microphone performance testing method is provided, including the following steps: S10: Collect sound signals from inside and outside the vehicle body through a sound perception system pre-installed on the vehicle body to obtain raw perception information; S20: Based on the original sensing information, perform noise interference simulation on the vehicle microphone system to generate interference prediction information for the vehicle microphone system; S30: Based on the interference prediction information, perform a sound production adaptation analysis on the vehicle microphone system to obtain the theoretical sound production mode of the vehicle microphone system; S40: Perform a sound emission test on the vehicle microphone system according to the theoretical sound emission mode, and obtain the sound emission test results through the sound perception system to analyze and obtain the sound emission performance information of the vehicle microphone system in the current noise environment.

[0019] In this embodiment, the sound perception system is a multi-channel acoustic sensor network pre-installed and calibrated on the vehicle body. The network includes microphone arrays distributed outside the vehicle (such as below the exterior rearview mirror, inside the front bumper, near the shark fin on the roof, etc.) and microphone arrays distributed inside the vehicle (such as the roof, A-pillar, seat headrest area, above the center console, etc.). All microphones record in a multi-channel manner in sync with the vehicle's audio bus, collecting sound signals of the vehicle under target operating conditions (idling, constant speed, rapid acceleration, cruising at different speeds, etc.) to form raw perception information containing timestamps and channel location labels.

[0020] The raw sensing information includes at least the following multi-dimensional acoustic components: wind noise outside the vehicle, road noise, powertrain noise, surrounding traffic noise, as well as occupant voices, air conditioning noise, and structural vibration radiation. To ensure the accuracy of subsequent simulations, real-time CAN bus data (such as vehicle speed and engine speed) can be recorded during the acquisition process to label the operating condition dependence of noise.

[0021] By employing multi-channel synchronous acquisition covering both inside and outside the vehicle, the time-frequency characteristics and spatial distribution information of the real sound field can be fully preserved. This raw perception information serves as the noise base for subsequent noise interference simulation and as a reference benchmark for the evaluation system during the final sound emission test. This avoids scene distortion caused by using a single microphone or standard noise library, making the entire testing process rooted in the acoustic environment of the actual vehicle.

[0022] The vehicle microphone system refers to the call or voice interaction pickup system under test, which includes at least one pickup microphone and an embedded signal processing link (such as beamforming, echo cancellation, noise suppression, etc.). The local sound field representation centered on the installation location of the vehicle microphone system is extracted from the original perceived information. By measuring the transfer function of the microphone array inside and outside the vehicle, the sound propagation model from the equivalent position of each noise source to each pickup element of the vehicle microphone system is obtained.

[0023] Using a sound field synthesis algorithm, the environmental noise component in the original perceived information (the residual background noise after removing near-end speech through speech activity detection) is used as the excitation source to reconstruct the noisy signal at the microphone diaphragm of the vehicle microphone system in the software environment. The phase difference, coherence and reflection reverberation characteristics between each array element are fully taken into account during the simulation process to obtain a multi-channel simulated noisy signal.

[0024] These simulated noisy signals are input into a signal processing algorithm model that is completely equivalent to the vehicle microphone system (or directly input into the actual hardware-in-the-loop interface), and noise reduction, beamforming and other processing are performed. The processed time-domain signal or frequency-domain characteristics are output. A set of interference prediction information is calculated based on the processed signal. This information includes segmented signal-to-noise ratio, speech intelligibility index, speech recognition confidence reduction rate, echo residual energy, frequency response distortion, etc. It can also predict whether the signal output by the microphone system will experience clipping or nonlinear distortion at a specific near-end speech sound pressure level.

[0025] By using acoustic simulation based on real multi-channel noise fields, the signal conditions of the pickup front end under various noise scenarios can be quickly reproduced without actual road testing or interference with the vehicle microphone system hardware. Quantitative predictions covering both voice quality and recognition accuracy can be obtained. This approach greatly improves the coverage and efficiency of test scenarios. At the same time, the prediction results directly reflect the shortcomings of existing hardware and algorithms in dealing with this noise environment, providing a clear direction for optimization in the subsequent search for the optimal sound generation mode.

[0026] The purpose of acoustic adaptability analysis is to find the system operating parameters and acoustic response mode that enable the vehicle microphone system to achieve optimal sound pickup performance under the current noise interference. The theoretical sound production mode includes the pickup beam direction and width of the microphone array, the gain coefficient of each array element, the noise reduction algorithm mode and intensity level, the automatic gain control target level, and the echo cancellation half-duplex switching threshold.

[0027] An adjustable parameter set is pre-established for the vehicle microphone system, and a comprehensive performance cost function is defined. This cost function can be set according to the application scenario. For example, for hands-free calling, the focus is on voice quality perception evaluation and echo suppression, and for speech recognition, the focus is on word recognition accuracy. The generated interference prediction information is used as the current environmental constraint. The parameter optimization algorithm (such as Bayesian optimization, genetic algorithm or simple grid search) is used to iterate in the parameter space. Each iteration repeats the signal processing simulation and calculates the interference prediction information until the cost function converges or reaches the set performance threshold. The parameter combination that meets the threshold and has the best comprehensive performance is selected as the theoretical sound production mode.

[0028] If, after sufficient searching, all parameter combinations fail to bring the key performance indicators to the minimum requirements, then this operating condition is marked as performance-limited in the theoretical sound production mode, and the configuration closest to the target is provided for further hardware optimization reference. Manually adjusting microphone system parameters repeatedly based on measured noise is not only time-consuming but also makes it difficult to find a globally optimal solution. Adaptive analysis based on interference prediction information can quickly compare hundreds or thousands of pickup patterns within the simulation domain, automatically identifying the optimal sound production mode for the specific noise environment. The resulting theoretical sound production mode has clear physical meaning and fully leverages the adjustable potential of the hardware and software, providing a targeted and executable calibration benchmark for the final real-world sound production test.

[0029] Configure the theoretical sound production mode parameters into the actual vehicle microphone system to put the microphone system into this working mode. At the same time, maintain the same noise conditions when the vehicle collects the original perception information, or use the vehicle audio system to play back the noise environment to keep the background sound field unchanged.

[0030] The vehicle's microphone system emits predefined test data, such as male and female artificial speech sequences conforming to the ITU-T P.501 standard, frequency sweep signals, or specific keyword commands, through a reference sound source in the sound perception system (such as an artificial mouth fixed to the driver's headrest, or by playing standard test signals through the vehicle's speakers). In this state, the vehicle's microphone system picks up the sound and outputs the processed signal.

[0031] The in-vehicle microphone array of the sound perception system synchronously records the direct sound of the test sound source and the reverberation sound in the cabin during the sound generation test. At the same time, it extracts the processed signal output from the digital audio interface of the in-vehicle microphone system and summarizes these records into the sound generation test results.

[0032] By comparing the microphone system output signal with known test source signals, objective indicators such as frequency response, total harmonic distortion, delay introduced by signal processing, weighted signal-to-noise ratio, voice quality perception evaluation, and correct response rate of voice recognition commands are calculated. Combined with the sound field data recorded in the vehicle, the spatial selectivity after beamforming, the consistency of the directional pattern with the theoretical design, and the presence of abnormal howling or residual echoes can also be evaluated. By combining these indicators, a report on the actual sound performance of the vehicle microphone system under the optimal theoretical sound production mode in the current noise environment is formed.

[0033] Since theoretical sound production modes are derived from simulation predictions, their effectiveness and robustness must be verified in the real physical world. Using the same sound perception system to complete the test signal emission and result acquisition avoids introducing additional measurement uncertainties. By comparing the deviation between actual performance and predicted performance, the acoustic model and signal processing model can also be calibrated and iterated, enabling the entire testing method to have continuous self-optimization capabilities. The final sound production performance information can serve as a direct basis for vehicle factory microphone performance calibration, noise reduction algorithm version evaluation, and cabin acoustic design.

[0034] In one embodiment, such as Figure 2 As shown, in step S10, sound signals from inside and outside the vehicle are collected through a pre-installed sound perception system to obtain raw perception information. This specifically includes the following steps: S11: The sound sensing system, composed of several sound sensors, collects sound signals from inside and outside the vehicle to obtain the acoustic signals corresponding to each sound sensor. S12: Generate a corresponding spatial marker for the acoustic signal according to the setting position of the sound sensor, and generate a corresponding time marker according to the acquisition time of the acoustic signal; S13: Arrange the acoustic signals according to the time and spatial markers to obtain the original sensory information.

[0035] In this embodiment, a distributed sensor network consisting of multiple sound sensors is pre-deployed on the vehicle body. These sound sensors include, for example, MEMS microphones or electret microphones. External sensors are installed below the exterior rearview mirrors, inside the front bumper, near the roof shark fin, and in the area of ​​the rear license plate light, etc., to pick up wind noise, road noise, powertrain radiated noise, and surrounding traffic noise. Internal sensors are installed in the roof, A-pillars, B-pillars, seat headrests, and above the center console, etc., to pick up occupant voices, air conditioning vent noise, structural vibration radiated sound, and interior reverberation components. All sensors are connected to a multi-channel synchronous data acquisition unit via shielded cables. This acquisition unit operates based on a high-precision clock source synchronized with the vehicle's CAN bus or the vehicle's overall clock.

[0036] During the test, signal acquisition from all sound sensors is initiated simultaneously. Analog-to-digital conversion is performed at a uniform sampling rate (e.g., 48kHz or 96kHz) to obtain the acoustic signal corresponding to each sound sensor. These acoustic signals completely record the sound pressure waveform that changes over time at the sensor's location.

[0037] By employing multiple sound sensors distributed inside and outside the vehicle for synchronous acquisition, the entirety of the vehicle's noise environment can be captured in a spatial dimension. A single microphone can only record the sound pressure at its location and cannot distinguish the incident direction and spatial distribution of sound. However, multi-channel synchronous acquisition preserves the phase difference, amplitude difference, and coherence between signals at different locations, providing irreplaceable multi-dimensional basic data for subsequent accurate reconstruction of the actual sound field at each array element of the onboard microphone system. If sound information from a certain direction is missing, the noise interference simulation will miss important directional noise components, leading to prediction distortion.

[0038] During the installation of the sound sensors, the precise installation position (x, y, z coordinates) of each sensor relative to the vehicle coordinate system and the pointing angle of its acoustic axis are calibrated using a three-dimensional coordinate measuring device. This position information is stored in the test system configuration file. Simultaneously with or after the acquisition of acoustic signals in step S11, the signal processing module reads the spatial information corresponding to each sensor from the configuration file and generates a spatial marker for each acoustic signal. This spatial marker can be a data structure containing fields such as coordinates, angles, and the region it belongs to (e.g., "exterior left front wheel arch" or "interior driver's headrest").

[0039] The acquisition unit uses a global synchronization clock to stamp each sampling frame or data block with a high-precision timestamp, generating a time stamp. This time stamp can be accurate to the microsecond level and is aligned with the time axis of operating condition data such as vehicle speed and engine speed recorded by the vehicle's CAN bus. The spatial and time stamps are stored as metadata associated with the corresponding acoustic signal stream.

[0040] Spatial marking gives the original acoustic signal a clear spatial attribute, enabling any subsequent analysis targeting a specific noise source or pickup location to accurately select the corresponding sensor signal. It also facilitates the construction of the acoustic transfer function from the external noise source to the internal pickup point. Temporal marking ensures the synchronization accuracy of all channel signals at the sub-millisecond level, eliminating the sampling delay inconsistencies that may exist between different sensors. This precise spatiotemporal alignment is the foundation for subsequent multi-channel sound field reconstruction, noise interference simulation, and phase-sensitive beamforming prediction. If the synchronization is off, it will lead to errors in the spatial cues in the simulated noisy signal, resulting in a complete distortion of the noise reduction performance prediction.

[0041] All acoustic signals carrying time and spatial markers are integrated. Using the time markers as a reference, the sampling points of acoustic signals from all channels are strictly aligned to ensure that the sampling values ​​at the same moment logically correspond to the same physical moment. The signals are arranged according to preset grouping rules and spatial markers. For example, the external sensor signals are arranged into an external noise matrix according to the spatial order from the front to the rear of the vehicle and from left to right; the internal sensor signals are arranged into an internal sound field matrix according to the seat position and height. During the arrangement process, the independent waveform data of each channel is retained, and a structured data file containing multi-channel synchronous waveforms, sampling rate, channel spatial distribution map, spatial markers of each channel, time start and end markers, and test condition information is generated. This file is the original perception information.

[0042] By arranging the spatiotemporally aligned acoustic signals in an orderly manner, the original perceptual information obtained is no longer a collection of scattered, independent recording segments, but a comprehensive dataset that fully describes the spatiotemporal structure of the sound field inside and outside the vehicle during the test period. This structured data organization greatly facilitates the direct extraction of noise excitation sources by spatial region, the calculation of the contribution of specific transmission paths, and the provision of accurate multi-channel source signals for sound field synthesis algorithms.

[0043] In one embodiment, such as Figure 3 As shown, step S20, which involves simulating noise interference on the vehicle microphone system based on the original sensing information to generate interference prediction information for the vehicle microphone system, specifically includes the following steps: S21: Analyze the distribution characteristics of acoustic signals at each moment in the original sensing information to obtain continuous acoustic distribution characteristics; S22: Analyze the distribution change trend of continuous acoustic distribution characteristics, and predict the acoustic distribution characteristics in future time periods based on the analysis results to obtain acoustic distribution prediction information. S23: Obtain the setting position and setting form of each microphone unit in the vehicle microphone system, and simulate the sound interference effect of the setting position and setting form of each microphone unit based on the acoustic distribution prediction information to obtain the interference prediction information of the vehicle microphone system.

[0044] In this embodiment, the obtained raw sensing information is received. This information is a multi-channel acoustic signal matrix that has been spatiotemporally aligned and arranged. Short-time Fourier transform is performed on the signals of each sound sensor with a preset time window length and step (e.g., frame length 20 milliseconds, step 10 milliseconds) to obtain the time spectrum of each channel. Using the spatial markers of each channel, at each frame time, the acoustic distribution characteristics of the preset area inside the vehicle (e.g., the driver's head area, the passenger area, and the rear passenger area) and the equivalent surface of the noise source outside the vehicle are calculated. The acoustic distribution characteristics may include the spatial distribution spectrum of sound pressure level of each frequency band, sound intensity vector field, the location estimation and intensity of the main noise source obtained based on beamforming or acoustic holography methods, the sound field coherence matrix between each region, and acoustic parameters such as reverberation time reflecting the reverberation characteristics of the sound field. These features calculated frame by frame are concatenated in time order to form a continuous acoustic distribution feature sequence. The feature vector at each moment is attached with a time marker and a corresponding operating condition label.

[0045] Acoustic signals are waveform data, which are not convenient to use directly for prediction and simulation. By analyzing frame by frame to extract high-level features such as energy distribution, directionality and reverberation of the sound field, the complex and ever-changing original acoustic signal can be transformed into a sequence of acoustic distribution features with clear physical meaning and reduced dimensionality. This continuous acoustic distribution feature not only describes the real-time dynamic changes of the noise environment, but also retains spatial information, providing a structured input for the next step of trend prediction. This allows the prediction model to focus on the changing laws of the essential properties of the sound field, rather than the fluctuations of the original sampling points, thereby improving the accuracy and robustness of the prediction.

[0046] The obtained continuous acoustic distribution feature sequence is input into a pre-trained trend prediction model. This prediction model can employ a recurrent neural network based on a long short-term memory network or a gated recurrent unit, or an adaptive predictor based on a Kalman filter. The model's input includes not only the acoustic distribution features of the current moment and a past period (such as the spatial distribution of sound pressure level and the location of the main noise sources), but also operating condition prediction signals obtained from the vehicle's CAN bus, such as accelerator pedal opening, engine speed change rate, and vehicle speed change trend. The model outputs acoustic distribution prediction information for a future period (e.g., 0.5 to 2 seconds). This information is consistent in form with the acoustic distribution features, including the predicted spatial distribution of sound pressure levels in each frequency band, the migration trajectory of the main noise sources, changes in reverberation characteristics, etc., and is accompanied by prediction confidence. If the vehicle is in a constant-speed steady-state condition, the prediction model will mainly provide a steady noise prediction based on historical features. If a transient operating condition trend such as rapid acceleration is detected, the model combines the operating condition input to quickly predict the sharp increase in wind noise and powertrain noise.

[0047] The vehicle noise environment is not completely random; its changes have a definite dynamic relationship with the vehicle's driving conditions and the external environment, and it has a certain inertia. By analyzing the continuous changing trend of acoustic distribution characteristics, we can capture the law of noise field evolution from the current state to the future state. This predictive ability allows subsequent interference simulations not only to target the noise that has already occurred, but also to obtain the upcoming noise field distribution in advance. This provides a time window for parameter pre-adjustment of the vehicle microphone system, which can avoid the instantaneous degradation of communication quality caused by the lag of noise reduction algorithms or sound pickup mode switching behind noise mutations. It is a key foundation for achieving adaptive and forward-looking sound mode optimization.

[0048] The precise installation position of each microphone unit in the vehicle coordinate system, the sensitivity and directivity pattern of each unit, and the array configuration (such as linear, circular, or distributed) are read from the design parameters of the vehicle microphone system under test. The generated acoustic distribution prediction information for a future target time is extracted, which provides the predicted sound pressure spectrum and noise source direction estimation at each point in the vehicle space.

[0049] Based on this, the sound interference effect of each microphone unit is simulated. It is assumed that at a future target time, a target speech signal with known characteristics (test speech) will be emitted by the speaker of the in-vehicle microphone system or by a passenger in the vehicle, while a predicted noise environment exists. The transfer function of the direct sound and early reflected sound of the target speech signal from the emission point to the diaphragm of each microphone unit is calculated using a sound propagation model. Using acoustic distribution prediction information, a predicted noise signal at the diaphragm of each microphone unit is synthesized. These two signals are superimposed to obtain a simulated noisy pickup signal. This simulated noisy signal is input into a software model (including beamforming, noise suppression, echo cancellation, etc.) equivalent to the actual signal processing link of the in-vehicle microphone system for full-link processing simulation. After processing, a set of interference prediction information is calculated, including key performance indicators such as predicted output signal-to-noise ratio, predicted speech intelligibility index, predicted speech recognition confidence, predicted echo residual, and predicted frequency response fluctuation. This interference prediction information characterizes the expected actual pickup performance of the in-vehicle microphone system under the current parameters in the upcoming noise environment.

[0050] The location and form of the microphone unit determine its spatial sampling physical characteristics. Different installation locations have fundamental differences in their sensitivity to noise and their ability to pick up target speech. By combining the predicted acoustic distribution field with the acoustic characteristics of the microphone unit for simulation, it is possible to realistically reproduce the signal that each microphone unit will actually hear in the future spatial sound field. Using this predictive simulation of the future, the degree of degradation of various performance indicators of the microphone system can be predicted before the noise actually arrives.

[0051] In one embodiment, such as Figure 4As shown, step S30, which involves performing a sound adaptation analysis on the vehicle microphone system based on the interference prediction information to obtain the theoretical sound production mode of the vehicle microphone system, specifically includes the following steps: S31: Based on the sound performance data of the vehicle microphone system, simulate the sound effect of the vehicle microphone system with standard sound parameters to obtain standard sound effect information; S32: Based on the interference prediction information of the vehicle microphone system, the interference effect simulation is performed on the standard sound effect information to obtain the expected performance information of the standard sound parameters under noise interference. S33: Evaluate the expected performance information to optimize the standard sound parameters until the final optimized sound parameters are obtained, so as to establish the theoretical sound mode of the vehicle microphone system.

[0052] In this embodiment, the sound performance data of the vehicle microphone system is acquired. This data includes the frequency response curve, sensitivity, directivity diagram, dynamic range, self-noise level of each element in the microphone array, as well as the mathematical model and adjustable parameter range of the embedded signal processing algorithms (such as beamforming, echo cancellation, noise suppression, automatic gain control, etc.). Based on this, a set of preset standard sound parameters is selected as the initial state of the simulation. These standard sound parameters can be the system's factory default values ​​or the parameter set saved in the current operating state, such as the conventional angle of the main lobe of the beam pointing to the driver's mouth area, moderate noise reduction level, and echo cancellation in full-duplex mode.

[0053] Using acoustic simulation software or a self-developed simulation engine, assuming an ideal acoustic environment with no background noise or only extremely low background noise, a test signal sequence conforming to international standards is emitted from a reference sound source inside the vehicle (such as an artificial mouth at the driver's headrest or a specific speaker). These signals include, for example, artificial speech recommended by ITU-T P.501, logarithmic sweep signals, etc. Combining sound performance data and standard sound parameters, the processed signal output after passing through the entire signal link is simulated and calculated. Standard sound effect information is extracted from this signal. This standard sound effect information includes, but is not limited to, frequency response flatness, total harmonic distortion plus noise, speech intelligibility index, speech recognition accuracy, echo loss enhancement value, output signal delay, etc., which characterize the pure sound performance benchmark of the system itself when it is not affected by external noise.

[0054] Standard sound quality information calibrates the baseline performance of an in-vehicle microphone system under ideal acoustic conditions. It isolates the variable of noise interference and reflects only the sound pickup quality determined by the system's hardware and default algorithm parameters. By establishing this performance baseline first, any performance degradation caused by the subsequent introduction of noise interference can be clearly attributed to the noise environment, rather than inherent limitations of the system itself.

[0055] The generated interference prediction information is used as a noise environment constraint to evaluate the expected performance under the current standard sound parameters in the simulation domain. The interference prediction information includes the predicted spatial distribution of noise at the future target time, noise energy in each frequency band, and directional interference characteristics. The test signal sequence is mixed with the predicted noise signal corresponding to the diaphragm of each microphone unit using a sound field superposition method to form a simulated noisy input signal. This noisy input signal is then input to a signal processing simulation link with the exact same configuration as the standard sound parameters to obtain the processed output signal.

[0056] By comparing the output signal with the known original test signal, a set of performance indicators under noise interference is calculated, which is the expected performance information of the standard vocal parameters under noise interference. This information includes the predicted output signal-to-noise ratio, the predicted decrease in speech intelligibility index, the predicted increase in speech recognition word error rate, the predicted frequency response distortion, and the predicted increase in echo residual energy. If necessary, the masking effect and speech distortion caused by noise can also be visualized and analyzed in the time and frequency domain.

[0057] By combining predicted noise interference with the system's standard sound output mode, the system can simulate how it will perform in an upcoming noisy environment if it continues to operate with the current standard sound output parameters without any parameter adjustments. This proactive interference effect simulation can expose the weaknesses of standard parameters under specific noise conditions without consuming real-vehicle road test resources and time. For example, it can reveal serious deficiencies in the signal-to-noise ratio of certain frequency bands or speech recognition failures caused by beam pointing failing to avoid strong interference sources. This provides a quantitative basis for deciding whether parameter optimization is necessary and from which dimensions, avoiding blind adjustments.

[0058] A multi-objective cost function is defined, which determines the weight of each performance indicator based on the application scenario of the vehicle microphone system. For example, for hands-free calling scenarios, the voice quality perception evaluation and echo suppression level have higher weights in the cost function; for voice recognition scenarios, the recognition accuracy and response latency have higher weights. Substitute each indicator in the expected performance information into the cost function to calculate the comprehensive performance cost of the current standard voice parameters.

[0059] Initiate the parameter optimization iterative process. The optimizable vocal parameters constitute a multi-dimensional parameter space, including the beam pointing angle and beamwidth of the microphone array, the gain weighting coefficient of each array element, the suppression depth and frequency band division of the noise reduction algorithm, the target level of automatic gain control, the nonlinear processing intensity of echo cancellation, and even the trigger threshold for voice activity detection. Using methods such as Bayesian optimization, genetic algorithm, particle swarm optimization, or simple gradient descent, new parameter combinations are searched within this parameter space. For each set of candidate parameters, the interference effect is repeatedly simulated to obtain its corresponding expected performance information and calculate the cost function value. The iteration continues until the cost function converges to an acceptable range or reaches the preset maximum number of iterations.

[0060] When a set of parameters not only meets the minimum performance threshold (e.g., predicted output signal-to-noise ratio greater than 15dB, predicted speech recognition accuracy higher than 95%), but also minimizes the cost function, this set of parameters is selected as the optimized sound generation parameters. These optimized sound generation parameters, along with their expected performance, constitute the theoretical sound generation mode of the vehicle microphone system under the current predicted noise environment. If, after sufficient searching, no parameter combination can make the key indicators meet the minimum requirements, it is marked as performance-limited in the theoretical sound generation mode, and the compromise parameter configuration closest to the target is given.

[0061] Manual tuning of numbers is almost impossible to find the global optimal solution in a short time in a complex multi-channel noise reduction system, especially when the noise environment changes dynamically. By adopting a cost function-based and automatic optimization method, the optimal sound pickup strategy for a specific sound field pattern can be efficiently explored in a virtual noise environment constructed from predictive information. The theoretical sound generation mode obtained in this way fully taps the potential adjustment capability of the vehicle microphone system's hardware and software, and is the optimal solution to adapt to changes in environmental noise at the lowest cost.

[0062] In one embodiment, such as Figure 5 As shown, in step S40, the vehicle microphone system is tested for sound output according to the theoretical sound output mode, and the sound output test results are obtained through the sound perception system to analyze and obtain the sound output performance information of the vehicle microphone system in the current noise environment. Specifically, this includes the following steps: S41: Analyze the testing methods of the vehicle microphone system based on the theoretical sound production mode to obtain the testing methods of the vehicle microphone system for the theoretical sound production mode; S42: Drive the vehicle microphone system to play test voice according to the test method, and collect the played test voice through the sound perception system to obtain test perception information; S43: Analyze the voice playback quality of the test perception information to obtain the sound performance information of the vehicle microphone system in the current noise environment.

[0063] In this embodiment, a theoretical sound emission mode for the current predicted noise environment is received. This theoretical sound emission mode includes a set of determined optimized sound emission parameters, such as: the precise pointing angle and width of the array beam main lobe, the gain allocation coefficient of each array element, the specific working mode and suppression intensity level of the noise reduction algorithm, the target response curve of automatic gain control, etc. These parameters are analyzed and mapped, and converted into a specific set of test instructions that can be driven and executed, which is the test method.

[0064] The algorithm-level parameters are mapped to hardware-level configuration instructions. For example, beam pointing angle and width are converted into complex weighting factors for each microphone channel and loaded into the beamforming module; noise reduction mode and suppression intensity are converted into calling parameters and threshold settings for the corresponding algorithm library; and the automatic gain control target level is set to the reference level value of the hardware codec. Simultaneously, based on the guidance instructions for speaker playback that may be included in the theoretical sound production mode, the type, sound pressure level, playback sequence, and playback channel of the test signal are determined. For example, if the theoretical sound production mode specifies that occupants need to be guided to increase the volume, the test method includes instructions to play the test data at a preset higher sound pressure level; if the theoretical sound production mode sets a specific beam scanning range, the test method includes the sequential playback of test signals at the corresponding spatial locations. This test method is ultimately output as a script or instruction sequence that can be directly parsed and executed by the test execution system.

[0065] The theoretical sound production mode is an abstract optimal solution defined at the parameter space and algorithm level. It cannot directly drive physical hardware to perform actions. By parsing the theoretical sound production mode into specific, hardware-oriented testing methods, a precise mapping bridge can be established between the virtual domain of parameter optimization and the physical test domain of the actual vehicle. This conversion step ensures that the optimal parameters obtained through optimization can be accurately deployed to the real vehicle microphone system, avoiding the inability to reproduce the theoretical optimal mode in practice due to parameter conversion errors or hardware interface mismatches. This ensures seamless connection and high fidelity from simulation optimization to real vehicle verification.

[0066] To ensure the vehicle is kept under the same target noise conditions as the original sensing information collected, such as by reproducing driving noise using a chassis dynamometer and wind tunnel equipment, or by using the vehicle audio system to play back the recorded environmental noise with high fidelity, the background sound field is consistent with the predicted noise environment targeted by the theoretical sound generation mode.

[0067] The test execution module parses the generated test script and performs the following operations: On the one hand, it sends configuration instructions to the in-vehicle microphone system, writing optimized sound parameters into the registers or parameter files of each module in its signal processing link, so that it enters the theoretical sound mode; on the other hand, it controls the preset reference sound source in the vehicle (e.g., a standard artificial mouth installed in the driver's headrest, or calling a specific channel of the in-vehicle high-fidelity speaker system) to play a standard test speech sequence according to the signal type, sound pressure level and timing specified in the test method. The test speech sequence may include artificial speech recommended by ITU-T P.501, a keyword list that conforms to the speech recognition test specification, a logarithmic sweep signal, and continuous sentences with speech rate and intonation changes required for specific noise suppression evaluation.

[0068] During this process, the in-vehicle microphone array of the sound perception system maintains a synchronized working state, responsible for collecting test voice signals played from multiple locations. The specific collection content includes the direct sound and reverberation signals that are played by the driver's artificial mouth, propagated through the air to reach each in-vehicle microphone, the processed signal output by the digital audio interface of the in-vehicle microphone system itself, and the original electrical signal as a reference. The sound perception system records the signals of all channels synchronously and adds spatial and temporal markers to form test perception information. This test perception information completely records the acoustic data of the entire process of test voice from the sound source to the output of the microphone system in a real noise environment and under the theoretical sound production mode.

[0069] By integrating configuration and data acquisition into a single process and using the same sound perception system as during initial noise acquisition, the consistency of acoustic benchmarks throughout the test loop is ensured, avoiding the introduction of additional system errors. Furthermore, playback and data acquisition within a realistically reproduced noise field verify whether the theoretically superior sound generation mode, which performs well in simulation, remains effective when encountering various non-ideal factors in a real sound field (such as sensor tolerances, installation structure vibrations, and actual reflections from in-vehicle scattering objects). The acquired multi-channel test perception information includes not only system output but also spatial sound field information, providing a complete data foundation for comprehensive and objective performance analysis.

[0070] From the acquired test perception information, the reference source signal and the output signal of the vehicle microphone system are extracted and preprocessed, such as time alignment and level normalization. A series of objective sound quality evaluation algorithms are used to calculate performance indicators. The analysis dimensions include at least the following three aspects: First, speech quality analysis, comparing the output signal with the reference source signal, calculating the Perceptual Speech Quality Evaluation (PESQ) or Objective Speech Quality Assessment (POLQA) score, frequency response deviation, total harmonic distortion plus noise (THD+N), segmented signal-to-noise ratio, etc., to evaluate the effect of theoretical speech mode on speech fidelity improvement. Second, speech intelligibility and recognition performance analysis, using the Speech Transmission Index (STI) to evaluate intelligibility, and using the speech recognition engine to conduct actual tests on the output signal, statistically analyzing indicators such as keyword recognition accuracy, false wake-up rate, and response latency, directly reflecting the actual performance of the voice interaction function. Third, spatial sound pickup characteristic analysis: using the sound field information recorded by the multi-channel sound perception system, by comparing the spatial response of the theoretical beam pointing with that of the actual beamforming output, the directivity pattern deviation and sidelobe suppression level are calculated to assess whether there is abnormal howling or residual echo.

[0071] By comparing the various indicators obtained from the above analysis with the standard sound effect information and the expected performance information, a comprehensive conclusion is formed, which is the sound performance information of the vehicle microphone system in the current noise environment. This information specifically indicates the degree to which the theoretical sound mode achieves its performance in the real environment, such as whether the voice quality score meets expectations, whether the recognition accuracy meets product specifications, whether the beamforming is aligned with the target and effectively suppresses interference, etc. If there are performance deviations, the source of the deviations can also be analyzed to provide a basis for the iterative calibration of the model.

[0072] Voice playback quality is the ultimate reflection of the performance of an in-vehicle microphone system, directly determining the user's subjective call experience and the success rate of voice interaction. Through multi-dimensional and standardized objective evaluation, the assessment of system performance can be upgraded from qualitative judgment to quantitative analysis, which is repeatable and comparable. In particular, comparing and analyzing the actual test results with the simulation prediction values ​​not only verifies the effectiveness of this optimization, but also provides feedback correction signals for the entire prediction-optimization-test closed loop. The resulting sound performance information can serve as a direct and reliable basis for product factory calibration, horizontal comparison of different noise reduction algorithm versions, and selection of the best cabin acoustic solution during the design phase.

[0073] In one embodiment, the step of analyzing the voice playback quality of the test perception information to obtain the sound performance information of the vehicle microphone system in the current noise environment includes: S411: The test perception information is decomposed into speech frames to obtain a speech frame sequence, and the speech frame sequence is subjected to lightweight mean processing and attention pooling processing to obtain a global speech vector. S412: Perform cross-attention fusion on the global speech vector to obtain a fusion vector, and evaluate the sound quality clarity and noise suppression level of the test speech based on the fusion vector to obtain speech performance information.

[0074] From the test perception information, the processed speech signal output by the vehicle microphone system in the theoretical sound production mode is extracted. The signal is pre-emphasized, framed and windowed. The frame length is set to 25 milliseconds and the frame shift is 10 milliseconds to obtain a speech frame sequence. Multidimensional acoustic features are extracted from each frame, such as 40-dimensional Mel frequency cepstral coefficients and their first and second order differences, or 128-dimensional Mel spectrogram features, to form a frame-level feature vector sequence.

[0075] Frame-level feature vectors are input into a lightweight convolutional neural network or a low-parameter Transformer encoder for local context modeling, resulting in enhanced frame-level representations. Two processing pathways are then executed in parallel: a lightweight averaging pathway, which calculates the average across all frame-level representations along the time dimension to obtain a fixed-length statistical vector capturing the average spectral energy distribution and global timbre features of the entire speech segment; and an attention pooling pathway, which introduces a set of learnable attention parameters to perform a weighted summation of the frame-level representation sequence. Specifically, a scalar attention score is calculated for each frame representation, normalized, and then weighted across all frames to obtain a pooling vector focusing on key articulation segments, high-energy phonemes, and clear transient components in the speech. This attention mechanism automatically highlights the segments in the speech signal most important for human auditory quality perception.

[0076] The statistical vector obtained by lightweight averaging is concatenated with the pooling vector obtained by attention pooling, and then dimensionality is compressed and fused through a fully connected layer to generate a fixed-length global speech vector. This global speech vector integrates the overall energy distribution information and temporal saliency information of the speech signal, serving as a compact representation of the test speech. Speech frame decomposition transforms continuous waveforms into acoustic feature sequences that are easy to analyze, which is a necessary prerequisite for refined speech quality assessment. Lightweight averaging can preserve the global spectral features of speech with extremely low computational overhead, which is particularly effective in distinguishing the overall frequency masking effect caused by background noise. Attention pooling, by simulating the human ear's emphasis on important speech segments, can significantly improve the correlation between the assessment results and subjective listening experience. The combination of the two to generate a global speech vector ensures processing efficiency without losing key details that affect clarity judgment, providing a complete and dimensionally controllable input for subsequent deep fusion assessment.

[0077] In addition to the obtained global speech vector, a reference global speech vector is also pre-obtained. This reference vector can be obtained by using the direct sound reference signal recorded by the sound perception system or the known original pure test speech, after the same frame decomposition, lightweight mean processing and attention pooling processing.

[0078] The global speech vector (to be evaluated) and the reference global speech vector are input into a cross-attention fusion module. In this module, the reference vector is used as the query and the global speech vector to be evaluated is used as the key and value. The cross-attention weight between the two is calculated. This operation can adaptively locate the differences between the speech to be evaluated and the clean reference speech in various feature dimensions, and generate a set of differential representations that emphasize distortion and residual noise components. The differential representation is then residually connected and layer normalized with the original global speech vector to obtain the fused vector.

[0079] The fused vector is fed into a multi-task evaluation head composed of fully connected layers. The evaluation head outputs a speech clarity score and a noise suppression level score in parallel. The speech clarity score, for example, maps to PESQ or POLQA prediction scores, reflecting the intelligibility, brightness, and distortion of the speech itself. The noise suppression level score quantifies the proportion of residual noise energy in the processed signal and the degree of noise suppression, for example, reflecting the mean and variance of segmented noise suppression. Combining these two scores and their respective confidence intervals forms the sound performance information. This information can be directly used to determine whether the pickup quality of the in-vehicle microphone system, optimized by the theoretical sound production mode, meets the standards in the current noise environment.

[0080] Features extracted solely from noisy processed signals are insufficient to distinguish whether speech distortion is caused by residual noise or excessive noise reduction. Through cross-attention fusion, the system can precisely align and compare the speech to be evaluated with known clean references, clearly separating the two aberration components: residual noise and speech distortion. This allows for the objective quantification of both speech clarity and noise suppression levels.

[0081] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0082] In one embodiment, a microphone performance testing device is provided, which corresponds one-to-one with the microphone performance testing methods described in the above embodiments. For example... Figure 6 As shown, the functional modules of this microphone performance testing device are described in detail below: The sound perception module is used to collect sound signals from inside and outside the vehicle through a sound perception system pre-installed in the vehicle body to obtain raw perception information; The interference prediction module is used to simulate noise interference on the vehicle microphone system based on the original sensing information to generate interference prediction information for the vehicle microphone system. The adaptive analysis module is used to perform sound adaptation analysis on the vehicle microphone system based on the interference prediction information to obtain the theoretical sound production mode of the vehicle microphone system. The sound production test module is used to perform a sound production test on the vehicle microphone system according to the theoretical sound production mode, and to obtain the sound production test results through the sound perception system in order to analyze and obtain the sound production performance information of the vehicle microphone system in the current noise environment.

[0083] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0084] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0085] This application also provides a computer device, such as... Figure 7 As shown, the computer device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in any of the above method embodiments, or when the processor executes the computer program, it implements the functions of each module / unit in the above device embodiments.

[0086] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device.

[0087] Those skilled in the art will understand that Figure 7 The computer device described is merely an example and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.

[0088] The aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), or Field Programmable Gate Arrays (FPGAs). Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0089] The memory can be an internal storage unit of the computer device, such as a hard drive or RAM. The memory can also be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal and external storage units of the computer device.

[0090] This application also provides a readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0091] This application provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.

[0092] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0093] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0094] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0095] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0097] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for testing microphone performance, characterized in that, include: The sound sensing system pre-installed in the vehicle body collects sound signals from inside and outside the vehicle to obtain raw sensing information; Based on the original sensing information, noise interference simulation is performed on the vehicle microphone system to generate interference prediction information for the vehicle microphone system. Based on the interference prediction information, a sound production adaptability analysis is performed on the vehicle microphone system to obtain the theoretical sound production mode of the vehicle microphone system. The vehicle microphone system is tested for sound output based on the theoretical sound output mode, and the sound output test results are obtained through the sound perception system to analyze the sound output performance information of the vehicle microphone system in the current noise environment.

2. The microphone performance testing method as described in claim 1, characterized in that, The steps for obtaining raw perception information by collecting sound signals from inside and outside the vehicle through a pre-installed sound perception system include: A sound perception system composed of several sound sensors is used to collect sound signals from inside and outside the vehicle to obtain the acoustic signals corresponding to each sound sensor. Based on the location of the sound sensor, a corresponding spatial marker is generated for the acoustic signal, and a corresponding time marker is generated based on the acquisition time of the acoustic signal. The acoustic signals are arranged according to time and spatial markers to obtain the original sensory information.

3. The microphone performance testing method as described in claim 1, characterized in that, The steps of simulating noise interference on the vehicle microphone system based on the original perceived information to generate interference prediction information for the vehicle microphone system include: The acoustic signal distribution characteristics at each moment are analyzed on the original sensing information to obtain continuous acoustic distribution characteristics; The distribution trend of continuous acoustic distribution characteristics is analyzed, and the acoustic distribution characteristics of future time periods are predicted based on the analysis results, thus obtaining acoustic distribution prediction information. The location and configuration of each microphone unit in the vehicle microphone system are obtained, and the effect of sound interference on the location and configuration of each microphone unit is simulated based on the acoustic distribution prediction information to obtain the interference prediction information of the vehicle microphone system.

4. The microphone performance testing method as described in claim 1, characterized in that, The steps for performing a sound adaptation analysis on the vehicle microphone system based on the interference prediction information to obtain the theoretical sound production mode of the vehicle microphone system include: Based on the sound performance data of the vehicle microphone system, the sound effect of the vehicle microphone system with standard sound parameters is simulated to obtain standard sound effect information. Based on the interference prediction information of the vehicle microphone system, the interference effect simulation is performed on the standard sound effect information to obtain the expected performance information of the standard sound parameters under noise interference. The expected performance information is evaluated to optimize the standard sound parameters until the final optimized sound parameters are obtained, so as to establish the theoretical sound mode of the vehicle microphone system.

5. The microphone performance testing method as described in claim 1, characterized in that, The steps of conducting sound production tests on the vehicle microphone system based on the theoretical sound production mode, and obtaining the sound production test results through the sound perception system to analyze and obtain sound production performance information of the vehicle microphone system in the current noise environment include: Based on the analysis of the test methods for the vehicle microphone system according to the theoretical sound production mode, the test methods for the vehicle microphone system for the theoretical sound production mode are obtained. The vehicle microphone system is driven to play test voice according to the test method described above, and the test voice played is collected by the sound perception system to obtain test perception information. The test perception information is analyzed to determine the voice playback quality, thereby obtaining the sound performance information of the vehicle microphone system in the current noise environment.

6. The microphone performance testing method as described in claim 5, characterized in that, The steps of analyzing the voice playback quality of the test perception information to obtain the sound performance information of the vehicle microphone system in the current noise environment include: The test perception information is decomposed into speech frames to obtain a speech frame sequence, and the speech frame sequence is subjected to lightweight mean processing and attention pooling processing to obtain a global speech vector. Cross-attention fusion is performed on the global speech vector to obtain a fusion vector. Based on the fusion vector, the sound quality clarity and noise suppression level of the test speech are evaluated to obtain speech performance information.

7. A microphone performance testing device, characterized in that, include: The sound perception module is used to collect sound signals from inside and outside the vehicle through a sound perception system pre-installed in the vehicle body to obtain raw perception information; The interference prediction module is used to simulate noise interference on the vehicle microphone system based on the original sensing information to generate interference prediction information for the vehicle microphone system. The adaptive analysis module is used to perform sound adaptation analysis on the vehicle microphone system based on the interference prediction information to obtain the theoretical sound production mode of the vehicle microphone system. The sound production test module is used to perform a sound production test on the vehicle microphone system according to the theoretical sound production mode, and to obtain the sound production test results through the sound perception system in order to analyze and obtain the sound production performance information of the vehicle microphone system in the current noise environment.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the microphone performance testing method as described in any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the microphone performance testing method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, enables the implementation of the steps of the microphone performance testing method as described in any one of claims 1 to 6.