Vehicle cabin sound field reconstruction method based on space energy diagram constraint and deep optimization

By introducing spatial energy map constraints and depth optimization into the sound field reconstruction method in the automotive cabin, the problems of weak spatial positioning ability and artifact noise interference in complex environments of traditional methods are solved, and high-fidelity and robust sound field reconstruction results are achieved.

CN121397451APending Publication Date: 2026-01-23PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511274888.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing automotive cabin sound field reconstruction technologies suffer from weak spatial positioning capabilities, insufficient robustness, and severe artifact noise interference in complex environments. Traditional methods cannot meet the requirements for high-fidelity spatial audio reproduction.

Method used

A cabin sound field reconstruction method based on spatial energy map constraints and depth optimization is adopted. By measuring the impulse response of the speaker array, spatial energy maps of the actual sound field and the ideal sound field are generated, an objective function is constructed, and the control filter coefficients are iteratively optimized using a neural network to achieve sound field reconstruction.

Benefits of technology

It significantly improves the sound quality and spatial positioning accuracy of sound field reconstruction, reduces artifact noise interference, and provides a more stable immersive in-vehicle spatial audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397451A_ABST
    Figure CN121397451A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle cabin sound field reconstruction method based on space energy diagram constraint and deep optimization. The method comprises the following steps: 1) measuring the impulse response of a vehicle-mounted loudspeaker array in a target vehicle cabin in a listening area; (2) space scanning is conducted on the impulse response and the impulse response simulated in the ideal anechoic chamber scene through a delay and beam forming method, sound energy distribution in all preset directions is calculated, and a space energy diagram of an actual sound field and a space energy diagram of a target ideal sound field are generated; 3) calculating the space energy diagram similarity between the two space energy diagrams, and constructing a target function; 4) constructing a neural network, taking random noise as input, outputting a control filter coefficient, and iteratively optimizing the neural network by using a target function; and 5) according to the neural network after iterative optimization, outputting a control filter coefficient of each loudspeaker channel, and loading the control filter coefficient to a sound system in the target vehicle cabin to realize sound field reconstruction in the target vehicle cabin. The invention has better subjective tone quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of 3D audio and spatial sound field reproduction technology, specifically involving a method for reconstructing the cabin sound field based on spatial energy map constraints and depth optimization. Background Technology

[0002] With the continuous advancement of autonomous driving technology and in-vehicle entertainment systems, the car cabin has gradually evolved from a traditional driving space into an intelligent living space integrating communication, entertainment, and information interaction. Users' expectations for in-vehicle audio systems have also upgraded from basic voice announcements and stereo playback to spatial audio reproduction systems with a sense of spatial presence and immersive experience. Under this trend, automakers have continuously improved the hardware configuration of audio systems, increasing the number of speakers in the cabin from single digits to over twenty, thus providing a physical basis for more precise sound field control and spatial positioning. Simultaneously, with the popularization of new energy vehicles and the development of vehicle quietness engineering, engine noise and environmental noise have been significantly reduced, creating a better listening environment for achieving high-fidelity spatial audio. However, compared to typical ideal acoustic environments such as home theaters and headphones, the car cabin, as a closed, small, and complex space, still faces many acoustic challenges. First, the interior walls of the car cabin are composed of various materials such as metal, glass, and fabric, which easily lead to complex multiple reflections, scattering, and absorption of sound waves during propagation, resulting in a mixture of early reflections and late reverberation, severely interfering with the reproduced sound quality. Secondly, due to the limitations of vehicle structure and spatial arrangement, the installation positions of loudspeakers are usually asymmetrical and non-uniform, with significant differences in distance, directivity, and frequency response between different units. These factors not only make sound field modeling difficult but also make it difficult to directly apply traditional sound field reconstruction algorithms.

[0003] Early sound field reconstruction methods, such as Wave Field Synthesis (WFS), Higher Order Ambisonics (HOA), and Vector-based Amplitude Panning (VBAP), typically relied on ideal assumptions such as uniform frequency response of the speaker array, regular spatial distribution (e.g., spherical uniformity), and a sound field propagation environment that is reflection-free or approximately anechoic. However, the actual automotive cabin environment deviates significantly from these assumptions, resulting in poor performance of such algorithms in in-vehicle applications, with both sound reproduction quality and spatial localization being significantly affected. To compensate for the impact of the cabin environment on acoustic propagation, playback methods based on Acoustic Channel Equalization (ACE) have been proposed. These methods measure the impulse responses (IRs) between each speaker and multiple control points, and calculate compensation filters to restore the desired sound effect. Traditional equalization methods typically rely on spatial sampling and direct inversion using least-squares algorithms. However, these methods suffer from poor spatial robustness, exhibiting degradation in performance within regions between control points, leading to distorted sound perception or location drift. To improve spatial consistency in playback, some studies propose aligning, clustering, or extracting prototypes from the IR values ​​of multiple control points to construct a unified spatial averaging filter. While this strategy improves the stability of the equalization results to some extent, it reduces the accuracy of spatial location due to the fuzzy processing of phase information. Furthermore, if appropriate temporal constraints are not applied in the filter design, the equalized response may produce significant pre-ringing artifacts, severely damaging the sound quality of the reconstructed sound field. Further, equalization techniques based on WFS and HOA apply the Fourier-Bessel sound field expansion principle to sound field reconstruction, achieving both a flat frequency response and accurate spatial location. However, due to the non-minimum phase characteristics of the acoustic channel, the reconstructed sound field is also prone to uncontrollable pre-ringing artifact noise. Although minimum-phase control filters can be used to mitigate these noises to some extent, their performance falls far short of expectations. To suppress this noise, some methods attempt to extract common zeros from the impulse responses of multiple control points and design mixed-phase control filters using spectral decomposition of polynomial matrices to achieve an equalized response that does not introduce front-resonance. Other methods start from the energy envelope and constrain the shape of the equalized IR, retaining only noise components that are insensitive to auditory perception. While these methods can improve sound quality to some extent, they still cannot fully solve the problem of high-fidelity spatial audio reproduction in the cockpit because they do not explicitly model and constrain spatial localization.

[0004] In summary, there is a significant contradiction between the complexity of the acoustic environment in current automotive cabins and traditional sound field reconstruction techniques. Algorithms relying on ideal conditions or regular structures cannot meet the reconstruction needs under non-ideal conditions; hardware optimization solutions are costly and have poor adaptability. Therefore, there is an urgent need for a method that can accurately reconstruct sound in complex spaces while improving high sound quality and spatial positioning capabilities to achieve a more flexible, robust, and high-fidelity immersive in-vehicle spatial audio experience. Summary of the Invention

[0005] Existing automotive cabin sound field reconstruction technologies mainly rely on physical models or frequency / time domain-based equalization optimization methods. These methods suffer from weak spatial localization capabilities, insufficient robustness, and severe artifact noise interference in complex cabin environments. To address these issues, this invention proposes a cabin sound field reconstruction method based on a combination of spatial energy map constraints and depth optimization.

[0006] The technical solution adopted in this invention is:

[0007] A method for reconstructing the cabin sound field based on spatial energy map constraints and depth optimization includes the following steps:

[0008] 1) Measure the impulse response of the onboard speaker array inside the target vehicle cabin within the listening area;

[0009] 2) The impulse response is spatially scanned using the delay and sum beamforming method to calculate the sound energy distribution in each predetermined direction and generate a spatial energy map of the actual sound field; the impulse response of the simulated ideal anechoic chamber scenario in the target vehicle cabin is spatially scanned using the delay and sum beamforming method to calculate the sound energy distribution in each predetermined direction and generate a spatial energy map of the target ideal sound field.

[0010] 3) Calculate the spatial energy map similarity between the actual sound field and the target ideal sound field, and construct an objective function; the objective function includes the similarity of impulse in the time domain response of the actual sound field and the target ideal sound field, the overall amplitude of artifact noise, the flatness of the frequency domain response, the loudspeaker operating frequency band, and the spatial energy map similarity;

[0011] 4) Construct a neural network with random noise as input and output control filter coefficients, and iteratively optimize the neural network using the objective function;

[0012] 5) The control filter coefficients of each speaker channel are output from the neural network after iterative optimization and loaded into the audio system in the target vehicle cabin to realize the reconstruction of the sound field in the target vehicle cabin.

[0013] Preferably, the method for measuring the impulse response of the vehicle-mounted speaker array within the listening area of ​​the target vehicle cabin is as follows:

[0014] 11) Set up a spherical microphone array at the selected listening position in the target vehicle cabin, and select a speaker from the vehicle speaker array as a reference speaker;

[0015] 12) Excitation signals are sequentially sent to the L speaker channels in the vehicle speaker array, and the response signals are collected using the spherical microphone array;

[0016] 13) Repeat step 12) L times to obtain a total of Q×L response signals from the speakers to the microphones; Q is the total number of microphones in the spherical microphone array;

[0017] 14) The acquired response signals are decoupled from the excitation signals by deconvolution to obtain the impulse response c between each pair of loudspeakers l and microphones q. ql (t), the impulse response c ql The length of (t) is L c .

[0018] Preferably, the excitation signal includes a maximum length sequence for achieving timestamp alignment between each speaker channel and an exponential sweep signal for measuring each frequency response; in step 12), the reference speaker is first driven to play the maximum length sequence signal, and then the speaker under test is driven to play the exponential sweep signal after a set delay time.

[0019] Preferably, the method for constructing the objective function is as follows:

[0020] 1) The sound field reconstruction problem is modeled as a multi-input multi-output linear time-invariant system; the linear time-invariant system includes S virtual sound sources, L speaker channels corresponding to the vehicle speaker array, and Q microphones corresponding to the spherical microphone array;

[0021] 2) Configure a length L for each speaker channel h Control filter h ls (t), then the total impulse response from the sound source s to the microphone q in the linear time-invariant system. Where * represents convolution operation, g qs The length of (t) is L g =L c +L h -1;

[0022] 3) Simulate and obtain the ideal transmission response d from the virtual sound source s to the microphone q. qs (t), ideal transmission response d qs (t) has a length of L g ;

[0023] 4) Based on the acoustic channel C from speaker l to microphone q qlConstruct L g ×L h Convolution matrix C ql Then the total convolution matrix of the linear time-invariant system C QL The convolution matrix constructed based on the acoustic channel from the Lth loudspeaker to the Qth microphone; the total response g of the linear time-invariant system. s =Ch s ; h Ls The control filter coefficients for the loudspeaker L corresponding to the sound source S;

[0024] 5) Construct four loss functions directly related to the reconstructed sound quality, and then weight and merge them to obtain the basic loss function.

[0025]

[0026] The first item measures the relationship between the total response and the target response in the time domain main window. The error within; the second term is used to penalize the pre- and post-ringing artifact noise components in the total response, combined with the window function. The first term is the energy of the total response; the second term is based on the standard deviation of the amplitude spectrum, while controlling the peak and valley values ​​in the spectrum to improve the spectral flatness; the third term limits the frequency range of the control filter through the frequency domain mask ξ. The discrete-time Fourier transform matrix; |·| p Indicate l p Norm, ⊙ denotes element-wise multiplication, λ1 to λ4 are weighting coefficients;

[0027] 6) Construct the objective function in,· s This is a spatial energy diagram of the actual sound field. The spatial energy diagram of the target ideal sound field; λ5 is the weighting coefficient.

[0028] Preferably, the neural network has the same length L as the control filter. h A 5-layer fully connected neural network with equal width.

[0029] Preferably, the method for training the neural network is as follows: a fixed set of random numbers is used as the input of the neural network, and the output is used as the control filter coefficients h for all speaker channels. s The backpropagation algorithm is employed, and the Adam optimizer automatically adjusts the neural network parameters to gradually converge the total loss calculated from the objective function. Then, the trained neural network is used to output the optimal control filter coefficients for each speaker channel. The filter set was then applied to the actual in-vehicle audio system to reconstruct the target sound field.

[0030] The main steps of this invention include:

[0031] 1) Acoustic channel measurement: Multi-channel microphones are arranged around the listening position (usually the driver's head position in the target vehicle cabin) to collect the impulse response of the vehicle speaker array in the actual cabin environment;

[0032] 2) Optimization Function Construction: Using the Delay-and-Sum Beamforming method, the multi-channel responses measured in the target vehicle cabin and simulated in an ideal anechoic chamber scenario are spatially scanned to calculate the sound energy distribution in each predetermined direction, generating spatial energy maps of the actual sound field and the target ideal sound field, thereby calculating their similarity. In general, the optimization function comprises a weighted combination of five loss parameters: the similarity of the impulse in the time-domain response, the overall amplitude of artifact noise (time-domain envelope), the flatness of the frequency-domain response, the loudspeaker operating frequency band, and the similarity of the spatial energy map. This combination serves as the objective function for training the deep optimization network.

[0033] 3) Deep optimization network design: A neural network is constructed, using random noise as input, and the control filter coefficients controlled by the network output are used to further calculate the specific values ​​of the optimization function. Simultaneously, the non-convex pseudo-infinite norm constraint is approximated by a differentiable large p-norm, and the spectral extremum constraint is replaced by a bidirectional spectral deviation metric. A backpropagation algorithm is used to iteratively optimize the network parameters to obtain a control filter that can ensure both flat sound quality and accurate spatial positioning reconstruction in complex vehicle cabin environments. Experiments have demonstrated that the method of this invention has the following advantages:

[0034] Superior subjective sound quality. In subjective tests based on auditory perception, listeners generally reported that the sound field generated by the method of this invention had higher sound quality, with subjective scores significantly higher than traditional equalization methods.

[0035] Enhanced spatial orientation. By introducing spatial energy map constraints, this method effectively reconstructs the energy distribution of the target sound field, significantly improving the accuracy and stability of sound image localization. Listeners experience clearer perception of sound source direction, smaller errors, and stronger auditory consistency in localization tests, resulting in a substantial improvement in subjective localization accuracy.

[0036] The sound consistency is superior to traditional methods. Compared with traditional equalization strategies based on time-domain or frequency-domain optimization, this method achieves a more stable and continuous sound field performance across multiple control points, effectively mitigating the impact of spatial location on sound quality and positioning perception.

[0037] The optimization function design is more general. By introducing a deep optimization solution method, the traditional method's restriction that the optimization function must be a convex function is relaxed, thus broadening the design scope of optimization functions. Attached Figure Description

[0038] Figure 1 This is a block diagram of the sound field reconstruction system.

[0039] Figure 2 This is a flowchart of the SPMnet method and a diagram of the calculation process of each optimization function.

[0040] Figure 3 This is a schematic diagram of the speaker arrangement inside the cabin of the experimental car.

[0041] Figure 4 It is a semi-simulated SSPM chart for objective evaluation.

[0042] Figure 5 It is an objectively measured SSPM chart.

[0043] Figure 6 It is a picture of a violin based on subjective evaluation;

[0044] (a) Sound quality, (b) Sense of direction. Detailed Implementation

[0045] The following description, in conjunction with the accompanying drawings and embodiments, introduces a method for reconstructing the cabin sound field based on spatial energy map constraints and depth optimization provided by the present invention:

[0046] Step 1: Measurement of the impulse response of each speaker in the cabin to the spherical microphone array

[0047] To obtain the true characteristics of the vehicle cabin acoustic system, this invention measures the acoustic channels of the vehicle speaker array within the listening area and models them using a Finite Impulse Response (FIR) filter. Specifically, the process includes the following steps: First, a spherical microphone array is positioned near the driver's head, with the array center located at the intended listening position. The array contains Q microphone channels, all of which are approximately uniformly arranged on the sphere. Next, excitation signals are sequentially emitted to the L speaker channels in the vehicle system. These excitation signals include a Maximum Length Sequence (MLS) for aligning the timestamps between speaker channels and an Exponential Sinusoidal Sweeps (ESS) signal for measuring the frequency response. The testing process for each speaker channel specifically involves first driving a reference speaker (usually the center speaker) to play the maximum length sequence signal, and then, after a fixed time delay, driving the speaker under test to play the exponential sweep signal. This process is repeated L times to obtain a total of Q×L recording files from the speakers to the microphones. If measurements of other nearby locations are required, the center of the spherical microphone array can be moved and the above operation repeated. Then, the impulse response is calculated and processed. The acquired response signal is deconvolved from the excitation signal to obtain the impulse response c between each pair of loudspeakers l and microphones q. ql (t), the length of the impulse response is denoted as L. c It is important to note that L c The choice of size depends on the environment of the car cabin, but since the reverberation time T in most car cabins is... 60 Less than 100 milliseconds, so L is selected. c =4096 is usually sufficient.

[0048] Step 2: Construct an optimization function based on the collected impulse response.

[0049] The impulse response between the speakers and the spherical microphone array in the cabin is completed. ql After measuring (t), this invention models the sound field reconstruction problem as a multi-input multi-output linear time-invariant (LTI) system. The system includes S virtual sound sources, L speaker channels, and Q microphones distributed around the listening position. Each speaker channel is configured with a length of L... h FIR control filter h ls (t), then the total impulse response g from the sound source s to the microphone q qs (t) can be expressed as:

[0050]

[0051] The as symbol represents the convolution operation, and the output response is g. qs The length of (t) is L g =L c +L h -1. Since the filter design for each sound source in this invention is performed independently, the following derivation is conducted under a single sound source scenario (s=1). Ideal target response d qs (t) represents the ideal transmission response from the sound source s to the microphone q in a reflection-free environment (e.g., an anechoic chamber), which is also of length L. g The signal is usually obtained through simulation. To construct the overall optimization problem, the variables are vectorized. The filter coefficients representing the sound source s corresponding to the loudspeaker l are expressed as:

[0052] h ls =[h ls (0),…,h ls (L h -1)] T

[0053] The measured acoustic channel C characterizing the path from speaker l to microphone q... ql Represented as

[0054] C ql =[c ql (0),…,c ql (L c -1)] T

[0055] Define the acoustic channel C from speaker l to microphone q. ql The constructed L g ×L h The convolution matrix is ​​C ql The system components can then be represented in the following block matrix form:

[0056]

[0057] Where d Qs The total response of the sound source s to the microphone Q, obtained from simulation under an anechoic environment, is characterized, and the total response of the system can be simplified as:

[0058] g s =Ch s

[0059] This invention first constructs four loss functions directly related to the reconstructed sound quality, and then combines them in the following weighted form as follows:

[0060]

[0061] Here, std() is the standard deviation function, and the first term measures the difference between the total response and the target response in the time domain main window. The error within; the second term is used to penalize the pre- and post-ringing artifact noise components in the total response, combined with the window function. The control imposes stricter constraints on the impulse response to concentrate the energy of the total response; the third term, based on the amplitude spectrum standard deviation, controls the peaks and valleys in the spectrum to improve spectral flatness; the fourth term limits the frequency range of the control filter through a frequency domain mask ξ. The specific design of the mask ξ depends on the operating characteristics of the loudspeaker used in the car cabin. Generally speaking, the mask ξ should not be wider than the normal operating frequency band of the loudspeaker to avoid possible nonlinear distortion. A discrete-time Fourier transform matrix of appropriate dimensions; |·| p Indicate l p The norm, in practice, is usually set to p = 20 to effectively achieve the goal. ⊙ represents element-wise multiplication. λ1 to λ4 are the weights of the guiding constraint functions.

[0062] To further enhance the spatial localization capability of the reconstructed sound field, this invention proposes an explicit constraint based on the Spatial Power Map (SPM). This constraint utilizes delay and beamforming methods to continuously estimate the spatial power map, which reflects the acoustic energy distribution of the sound field in different directions, during the control filter design process. By minimizing the difference between this spatial power map and the spatial power map of the target sound field, the spatial localization capability of the reconstructed sound field is improved. The calculation process for the spatial power map value in direction b is as follows:

[0063]

[0064] in G represents qs Fourier transform of (t); This represents the beamforming weights used in direction b, which can be easily calculated from the array manifold employing the microphone array. Further matrixing the above calculation process, Γ... bs By piecing them together sequentially, we can directly obtain the calculation method for the entire space energy map. The above process can be rewritten as follows:

[0065]

[0066] in The microphone spectrum is spliced ​​together, and the total time-domain response g is obtained. qs (t) is obtained through discrete Fourier transform; Ω is the beamforming matrix, predefined in all directions and combined with the microphone; Γ s Spatial energy map for reconstructing the sound field; ideal spatial energy map From the target response d sIt is obtained through beamforming with the same weight. Therefore, the spatial energy map constraint can be defined as:

[0067]

[0068] This explicit constraint reconstructs the directional distribution of the sound field, avoiding implicit modeling of orientation solely based on the impulse response waveform, thus more effectively reconstructing the orientation and spatial consistency of the sound image. The complete filter design objective function proposed in this invention is as follows:

[0069]

[0070] λ5 controls the influence of the spatial energy map error term on the total loss. This optimization objective, by introducing spatial energy constraints and combining them with traditional sound quality control mechanisms, constructs a sound field reconstruction filter design method that can ensure a flat frequency response, controllable artifact noise, and improved spatial positioning capabilities.

[0071] Step 3: Construct a neural network and train the control filter

[0072] This step addresses the optimization objective established in the previous stage by employing a neural network for automatic filter coefficient design. The specific implementation method is as follows: First, initialize a set of neural networks for generating filter coefficients. The network is the same as the control filter length L. h A 5-layer fully connected neural network with equal width. The network input is a fixed set of random numbers, and the output size h is... s To maintain consistency, the FIR control filter coefficient h is used for all speaker channels. s The loss function for network training is the optimization objective function defined in the previous step. During training, new optimization objective function values ​​are continuously calculated, and the backpropagation algorithm is used. The Adam optimizer automatically adjusts the neural network parameters, gradually bringing the total loss to convergence. For data from a single microphone array, the training process typically converges within 10,000 epochs; while for data from multiple microphone arrays, they can be treated as different batches for training, typically converging within 20,000 epochs. Finally, after network training is complete, the optimal FIR filter coefficients for each speaker channel are directly output from the trained neural network. By loading this set of filters into the actual car audio system, the target sound field can be reconstructed.

[0073] Method evaluation experiment

[0074] The method of this invention evaluated the sound quality and spatial positioning performance of the reconstructed sound field in a real-world automotive cabin environment using multiple objective indicators. The experiment employed 17 speakers (11 independent channels in total) arranged within the automotive cabin as a playback array (see...). Figure 3The microphone array was mounted near the driver's head. Five measurement locations were set up in the experiment to comprehensively examine spatial robustness.

[0075] In the objective experimental evaluation section, this invention uses normalized perceivable reverberation quantization (nPRQ) and spectral deviation (SD) of the reconstructed sound field impulse response to evaluate the sound quality of the reconstructed sound field from the perspectives of temporal artifact noise suppression and frequency domain response flatness. The nPRQ index is divided into front-resonance and rear-resonance components, and SD covers six octave bands to comprehensively measure the flatness and artifact suppression capability of the reconstructed sound field. This invention uses a stacked spatial power map (SSPM) to evaluate the spatial localization of the reconstructed sound field. Comparison methods include: the proposed method SPMnet and its multi-location version SPMnet-3, the traditional convex optimization method (CVX), the FD method based on frequency domain deconvolution, and the original vehicle audio system (ORI). All methods use the target impulse response of the virtual sound image as a reference and are measured and evaluated at five spatial locations. The results of the sound quality evaluation are shown in Table 1, where smaller nPRQ and SD indices represent lower perceptible ringing noise and higher spectral flatness, respectively. The results show that, except for the FD method, all algorithms significantly reduced pre-ringing and post-ringing artifact noise and improved spectral flatness. The CVX method performed best in terms of after-ringing suppression and spectral consistency. The method proposed in this invention has advantages in suppressing pre-ringing artifact noise. After introducing multi-location microphone data, it achieved optimal results in all temporal artifact suppression indices and further improved spatial consistency. Compared with the single-location method of this invention, the multi-location version of this invention is more robust in sound quality evaluation, indicating that multi-point spatial constraints have a positive effect on overall sound quality improvement.

[0076] Table 1. nPRQ and SD indices for the original system and various methods.

[0077]

[0078] The method of this invention was further evaluated in terms of spatial domain playback accuracy. To this end, Stacked Spatial Power Maps (SSPM) were constructed to analyze the directional rendering effect of the virtual sound source. SSPM calculates the spatial energy distribution at multiple angles using a delay-summation beamformer, with the horizontal axis representing the beamforming angle, the vertical axis representing the direction of the virtual sound source, and the brightness representing the energy intensity in that direction. Figure 6The SSPM of various methods at reference position O under semi-simulation conditions is presented. The target response exhibits a clear main diagonal, representing the desired spatial distribution. The original system and frequency domain methods (ORI and FD) do not show a clear diagonal structure, indicating weak directional focusing ability; the NN and CVX methods based on impulse response peak control have weak main diagonals in some regions; the SPMnet and SPMnet-3 methods proposed in this invention show a brighter and structurally continuous main diagonal, indicating that a concentrated diagonal energy distribution can be formed under these conditions. Overall brightness and contrast decrease, and the main diagonal becomes thinner or locally interrupted, possibly due to variations in sound propagation and measurement uncertainties in the cabin environment. At position O, some methods can still maintain a relatively clear main diagonal structure in the 240°–330° range. At positions L and LL, the main diagonal is blurred, and vertical fringes increase; at position R, some methods show intermittent diagonal recovery, possibly related to the geometric symmetry of this position. At position RR, some methods show diagonal shift or blurring, and reduced spatial consistency. Different methods show differences in the stability of spatial directionality. Some methods exhibit significant energy aliasing at certain angles (e.g., the main diagonal shifts to one side or a strong off-diagonal structure appears), indicating problems with unidentifiable angles or inaccurate spatial mapping during sound source reproduction. The method of this invention preserves a clear main diagonal structure at multiple locations, exhibits high continuity in directional rendering, and demonstrates stability in spatial domain response.

[0079] In the subjective evaluation, two sets of listening experiments were conducted to assess the perception effect of the method of this invention in a real vehicle cabin environment, targeting sound quality and spatial positioning performance respectively. The tests were conducted using the vehicle's existing speaker system, with reference signals played through headphones to provide a stable, high-fidelity control. A total of 13 subjects with normal hearing participated, with an average age of 24 years, 11 of whom had relevant listening experiment experience. The experiment employed a modified multi-excitation method (MUltiStimulus test with Hidden Reference and Anchor, MUSHRA), evaluating attributes including sound quality and spatial positioning. Two types of anchor signals corresponded to sound quality and orientation tasks, respectively. The listening materials included speech, symphonic music, allegro, string music, and impulse noise. Figure 6A violin graph showing the subjective test scores is presented. The sound quality test involved 4 materials, 6 angles, and 3 methods (plus anchor points); the spatial awareness test involved 3 materials, 12 angles, and 3 methods (plus anchor points). Testing methods included SPMnet-3, SPMnet, and CVX. Data from 12 participants were selected for the sound quality assessment. After standardization of scores, repeated measures ANOVA (RM-ANOVA) was performed, with Huynh-Feldt correction. The main effect of the methods was significant [F(2.566,677.425)=1244.462, p<0.001]. There was no significant difference between SPMnet and SPMnet-3 [p=0.732], but both were significantly better than CVX. A significant interaction existed between methods and materials [p=0.017]. Paired-samples t-tests for each material showed that SPMnet-3 was significantly better than CVX. The experimental results were not significant for the other two and three interactions.

[0080] In the spatial positioning assessment, data from another 12 subjects were selected. The method had a significant impact on positioning scores [F(3,1188)=1428.046,p<0.001]. Post-hoc paired-samples t-tests showed that SPMnet-3>SPMnet>CVX, with statistically significant differences among the three (p<0.001). SPMnet-3 scores were concentrated in the high-score range, showing strong consistency. There was a significant interaction between the method and the angle [F(33,1188)=6.292,p<0.001]. The results showed that SPMnet-3 achieved significantly higher scores in multiple directions; CVX scored the lowest in most angles, indicating higher subjective positioning uncertainty; SPMnet scores were in between, showing no significant difference from SPMnet-3 in some angles. The experimental results were not significant in the other two and three interactions. The spatial positioning experiment results show that the methods exhibit stable differences in perceived clarity of orientation reconstruction. SPMnet-3 maintains a high score under different angles and material conditions; SPMnet shows similar performance at some angles; CVX scores fluctuate significantly, with some subjective perceptions of spatial ambiguity or inconsistency in certain directions. These results are consistent with the trends in objective SSPM evaluations, indicating that multi-location control and spatial power constraints have the potential to improve orientation rendering perception.

[0081] In summary, this invention proposes an optimization method, SPMnet, for in-vehicle sound field reconstruction. By introducing a spatial power map constraint from delay-summation beamforming, it jointly achieves acoustic equalization and direction reconstruction, thereby improving spatial positioning performance. For the non-convex optimization problem introduced by this constraint, a deep optimization strategy is employed to solve it, overcoming the dependence of traditional methods on convexity. Objective and subjective experiments conducted in a real vehicle cabin environment demonstrate that the proposed method maintains clear and stable direction perception at multiple locations while also achieving good sound quality. Both objective and subjective evaluations validate the effectiveness of this invention.

[0082] Although specific embodiments and accompanying drawings of the invention have been disclosed for illustrative purposes to aid in understanding and implementing the invention, those skilled in the art will understand that various substitutions, variations, and modifications are possible without departing from the spirit and scope of the invention and the appended claims. Therefore, the invention should not be limited to the content disclosed in the preferred embodiments and accompanying drawings.

Claims

1. A method for reconstructing the cabin sound field based on spatial energy map constraints and depth optimization, comprising the following steps: 1) Measure the impulse response of the onboard speaker array inside the target vehicle cabin within the listening area; 2) The impulse response is spatially scanned using the delay and sum beamforming method to calculate the sound energy distribution in each predetermined direction and generate a spatial energy map of the actual sound field; the impulse response of the simulated ideal anechoic chamber scenario in the target vehicle cabin is spatially scanned using the delay and sum beamforming method to calculate the sound energy distribution in each predetermined direction and generate a spatial energy map of the target ideal sound field. 3) Calculate the spatial energy map similarity between the actual sound field and the target ideal sound field, and construct an objective function; the objective function includes the similarity of impulse in the time domain response of the actual sound field and the target ideal sound field, the overall amplitude of artifact noise, the flatness of the frequency domain response, the loudspeaker operating frequency band, and the spatial energy map similarity; 4) Construct a neural network with random noise as input and output control filter coefficients, and iteratively optimize the neural network using the objective function; 5) The control filter coefficients of each speaker channel are output from the neural network after iterative optimization and loaded into the audio system in the target vehicle cabin to realize the reconstruction of the sound field in the target vehicle cabin.

2. The method according to claim 1, characterized in that, The method for measuring the impulse response of the vehicle-mounted speaker array within the listening area of ​​the target vehicle cabin is as follows: 11) Set up a spherical microphone array at the selected listening position in the target vehicle cabin, and select a speaker from the vehicle speaker array as a reference speaker; 12) Excitation signals are sequentially sent to the L speaker channels in the vehicle speaker array, and the response signals are collected using the spherical microphone array; 13) Repeat step 12) L times to obtain a total of Q×L response signals from the speakers to the microphones; Q is the total number of microphones in the spherical microphone array; 14) The collected response signals are decoupled from the excitation signals by means of deconvolution, resulting in an impulse response c ql (t) between each pair of loudspeaker i and microphone q ql (t) has a length L c .

3. The method according to claim 2, characterized in that, The excitation signal includes a maximum length sequence for achieving timestamp alignment between each speaker channel and an exponential sweep signal for measuring each frequency response; in step 12), the reference speaker is first driven to play the maximum length sequence signal, and then the speaker under test is driven to play the exponential sweep signal after a set delay time.

4. The method according to claim 2 or 3, characterized in that, The method for constructing the objective function is as follows: 1) The sound field reconstruction problem is modeled as a multi-input multi-output linear time-invariant system; the linear time-invariant system includes S virtual sound sources, L speaker channels corresponding to the vehicle speaker array, and Q microphones corresponding to the spherical microphone array; 2) Configure a length L for each speaker channel h Control filter h ls (t), then the total impulse response from the sound source s to the microphone q in the linear time-invariant system. s = 1,…,S; q = 1,…,Q; where * denotes convolution operation, g qs The length of (t) is L g =L c +L h -1; 3) simulate an ideal transfer response d from virtual sound source s to microphone q qs (t), ideal transfer response d qs (t) length L g ; 4) Based on the acoustic channel C from speaker l to microphone q ql Construct L g ×L h Convolution matrix C ql Then the total convolution matrix of the linear time-invariant system C QL The convolution matrix constructed based on the acoustic channel from the Lth loudspeaker to the Qth microphone; the total response g of the linear time-invariant system. s =Ch s ; h Ls The control filter coefficients for the loudspeaker L corresponding to the sound source S; 5) Construct four loss functions directly related to the reconstructed sound quality, and then weight and merge them to obtain the basic loss function. The first item measures the relationship between the total response and the target response in the time domain main window. The error within; the second term is used to penalize the pre- and post-ringing artifact noise components in the total response, combined with the window function. The first term is the energy of the total response; the second term is based on the standard deviation of the amplitude spectrum, while controlling the peak and valley values ​​in the spectrum to improve the spectral flatness; the third term limits the frequency range of the control filter through the frequency domain mask ξ. The discrete-time Fourier transform matrix; |·| p express Norm, ⊙ denotes element-wise multiplication, λ1 to λ4 are weighting coefficients; 6) Construct the objective function Among them, Γ s This is a spatial energy diagram of the actual sound field. The spatial energy diagram of the target ideal sound field; λ5 is the weighting coefficient.

5. The method according to claim 1, characterized in that, The neural network is of the same length L as the control filter h 5-layer fully connected neural network with equal-width.

6. The method according to claim 1 or 5, characterized in that, The method for training the neural network is as follows: a fixed set of random numbers is used as the input to the neural network, and the output is used as the control filter coefficients h for all speaker channels. s The backpropagation algorithm is employed, and the Adam optimizer automatically adjusts the neural network parameters to gradually converge the total loss calculated from the objective function. Then, the trained neural network is used to output the optimal control filter coefficients for each speaker channel. The filter set was then applied to the actual car audio system to reconstruct the target sound field.