Partition evaluation and prediction method and system based on intelligent cabin sound field

By combining objective acoustic parameters and subjective evaluation scores with a CNN neural network, a sound field zoning effect prediction model is constructed, which solves the problem of the lack of subjective evaluation in the sound field zoning of intelligent cockpits, realizes efficient and accurate sound field zoning effect prediction, and improves design efficiency.

CN121936048APending Publication Date: 2026-04-28SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI UNIV OF ENG SCI
Filing Date
2026-01-17
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies lack subjective listening experience evaluation methods in smart cockpit sound field zoning, resulting in a lack of subjective evaluation dimensions, making it difficult to bridge the differences between subjective and objective factors, and affecting the predictability and usability of sound field zoning effects.

Method used

A sound field zoning effect prediction model is constructed by combining CNN neural network with objective acoustic parameters and subjective evaluation scores. By collecting and integrating noise signals, calculating objective parameters and subjective evaluation values ​​of sound field zoning, training samples are generated, and an intelligent prediction method for sound field zoning effect is established.

Benefits of technology

It achieves good generalization ability while ensuring the accuracy of evaluation, and can stably predict the subjective effect of sound field zoning, reduce the number of subjective tests, reduce costs and improve design efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121936048A_ABST
    Figure CN121936048A_ABST
Patent Text Reader

Abstract

The invention discloses a sound field partition evaluation and prediction method and system based on an intelligent cabin. The method comprises the steps that noise signals in a vehicle in front of and behind sound field partitions under different working conditions are collected in a semi-anechoic room; calculating objective parameters of sound field partitions according to the noise signals; evaluating and scoring the noise by an evaluator to obtain a corresponding sound field partition evaluation value; a training sample is generated according to the sound field partition objective parameters and the sound field partition evaluation values, a CNN neural network is constructed and trained based on the training sample, and a partition effect prediction model of the intelligent cabin sound field is obtained; and performing sound field partition evaluation prediction by using the partition effect prediction model. According to the method, good generalization ability is achieved while evaluation accuracy is guaranteed, continuous subjective evaluation values can be stably predicted, and an effective basis can be provided for tuning and subjective hearing sense design of a sound field partition system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a method and system for evaluating and predicting the sound field zoning of an intelligent cockpit. Background Technology

[0002] With the development of in-vehicle entertainment systems, traditional single-sound-field playback methods can no longer meet the personalized listening needs of different passengers. Intelligent cockpit sound field zoning technology has emerged to address this, establishing acoustic bright and dark zones to achieve spatial isolation of audio content within the same vehicle. Although various sound field optimization control algorithms have made some progress, and there are relatively complete zoning effect prediction systems, the lack of subjective listening experience and the absence of subjective evaluation dimensions makes it difficult to bridge the differences between subjective and objective factors, which makes it difficult to guarantee the usability of prediction results in practical applications. Research on prediction methods that combine subjective and objective dimensions has been extensive in the field of sound quality evaluation; however, because the research objects in this area differ from those in sound field zoning effect prediction, and because the data differences between bright and dark zones in sound field zoning are significant, it is difficult to directly refer to relevant research on sound quality evaluation. Therefore, a complete prediction method for sound field zoning is lacking. Summary of the Invention

[0003] This invention aims to address the shortcomings of existing technologies and provides the following solution: A method for evaluating and predicting the sound field zoning of a smart cockpit includes the following steps: Noise signals were collected from the front and rear of the vehicle under different operating conditions in a semi-anechoic chamber; Calculate the objective parameters of the sound field partitioning based on the noise signal; The noise is evaluated and scored by evaluators to obtain the corresponding sound field zone evaluation value; Training samples are generated based on the objective parameters of the sound field partition and the evaluation value of the sound field partition. A CNN neural network is constructed and trained based on the training samples to obtain a partition effect prediction model of the sound field of the intelligent cockpit. The sound field zoning evaluation prediction is performed using the zoning effect prediction model.

[0004] Preferably, the method for acquiring the noise signal includes: In the semi-anechoic chamber, noise audio before and after the sound field partitioning are played respectively. The first noise signal inside the vehicle was collected when the evaluators were in different sitting positions, and the second noise signal inside the vehicle was collected when the evaluators were distributed in different positions. The noise signal is obtained by integrating the first noise signal and the second noise signal.

[0005] Preferably, the objective parameters of the sound field partitioning include: the ratio of the average sound potential energy of the bright area to the dark area, the percentage of the normalized sound field mean square error in the bright area, the sound pressure level, the loudness, and the sharpness. The sound pressure level, the loudness, and the sharpness are extracted directly from the noise signal; The ratio of the average acoustic potential energy of the bright area to that of the dark area is: in, AC This represents the ratio of the average acoustic potential energy of the bright area to that of the dark area. q This indicates the speaker array drive signal. zB This represents the electroacoustic transfer function in the bright region. zD Represents the electroacoustic transfer function within the dark region. H Indicates taking the conjugate; The percentage of the normalized sound field mean square error in the bright area is: in, MSE This represents the percentage of the normalized mean square error of the sound field in the bright area. pB Indicates the reconstructed sound field of the bright area. pBT Indicates the target sound field in the bright area. f Indicates frequency, This represents solving for the square of the L2 norm.

[0006] Preferably, the method for obtaining the sound field partition evaluation value includes: Multiple evaluators used a rating scale to quantify and score the noise signals before and after the sound field partitioning, resulting in multiple subjective evaluation scores. Calculate the correlation coefficients of the multiple subjective evaluation scores to perform data validation: in, r Represents the correlation coefficient. Xi and Yi This represents two scores given by the same evaluator to a noise sample. and This represents the average of two scores given by the same evaluator for all noise samples. The data of evaluators with low correlation coefficients are removed to obtain the sound field zoning evaluation values.

[0007] Preferably, the CNN neural network includes: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a flattening layer, a fully connected layer, and an output layer; The first convolutional layer uses a 3×3 convolutional kernel and has 12 output channels to receive the input feature map; The first pooling layer uses a 2×2 pooling window with a stride of 2, which halves the feature map space size. The second convolutional layer uses a 3×3 convolutional kernel and has 24 output channels; The second pooling layer uses a 2×2 pooling window with a stride of 2, which again halves the feature map space size; The flattening layer flattens the pooled multidimensional feature map into a one-dimensional feature vector for processing by the fully connected layer; The fully connected layer contains 64 neurons, which map the flattened feature vectors to a 64-dimensional feature space. The output layer ultimately outputs a subjective evaluation value.

[0008] The present invention also provides a smart cockpit sound field zoning evaluation and prediction system, which applies the above-mentioned method and includes: a noise signal acquisition module, an objective data calculation module, a subjective evaluation value acquisition module, a model building module, and a prediction module; The noise signal acquisition module is used to acquire noise signals inside the vehicle before and after the sound field partition under different working conditions in a semi-anechoic chamber. The objective data calculation module is used to calculate the objective parameters of the sound field partitioning based on the noise signal; The subjective evaluation value acquisition module obtains the corresponding sound field zone evaluation value by evaluating personnel to score the noise. The model building module is used to generate training samples based on the objective parameters of the sound field partition and the evaluation value of the sound field partition, and to build and train a CNN neural network based on the training samples to obtain a partition effect prediction model of the sound field of the intelligent cockpit. The prediction module uses the zoning effect prediction model to perform sound field zoning evaluation and prediction.

[0009] Preferably, the workflow of the noise signal acquisition module includes: In the semi-anechoic chamber, noise audio before and after the sound field partitioning are played respectively. The first noise signal inside the vehicle was collected when the evaluators were in different sitting positions, and the second noise signal inside the vehicle was collected when the evaluators were distributed in different positions. The noise signal is obtained by integrating the first noise signal and the second noise signal.

[0010] Preferably, the objective parameters of the sound field partitioning include: the ratio of the average sound potential energy of the bright area to the dark area, the percentage of the normalized sound field mean square error in the bright area, the sound pressure level, the loudness, and the sharpness. The sound pressure level, the loudness, and the sharpness are extracted directly from the noise signal; The ratio of the average acoustic potential energy of the bright area to that of the dark area is: in, AC This represents the ratio of the average acoustic potential energy of the bright area to that of the dark area. q This indicates the speaker array drive signal. zB This represents the electroacoustic transfer function in the bright region. zD Represents the electroacoustic transfer function within the dark region. H Indicates taking the conjugate; The percentage of the normalized sound field mean square error in the bright area is: in, MSE This represents the percentage of the normalized mean square error of the sound field in the bright area. pB Indicates the reconstructed sound field of the bright area. pBT Indicates the target sound field in the bright area. f Indicates frequency, This represents solving for the square of the L2 norm.

[0011] Preferably, the workflow of the subjective evaluation value acquisition module includes: Multiple evaluators used a rating scale to quantify and score the noise signals before and after the sound field partitioning, resulting in multiple subjective evaluation scores. Calculate the correlation coefficients of the multiple subjective evaluation scores to perform data validation: in, r Represents the correlation coefficient. Xi and Yi This represents two scores given by the same evaluator to a noise sample. and This represents the average of two scores given by the same evaluator for all noise samples. The data of evaluators with low correlation coefficients are removed to obtain the sound field zoning evaluation values.

[0012] Preferably, the CNN neural network includes: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a flattening layer, a fully connected layer, and an output layer; The first convolutional layer uses a 3×3 convolutional kernel and has 12 output channels to receive the input feature map; The first pooling layer uses a 2×2 pooling window with a stride of 2, which halves the feature map space size. The second convolutional layer uses a 3×3 convolutional kernel and has 24 output channels; The second pooling layer uses a 2×2 pooling window with a stride of 2, which again halves the feature map space size; The flattening layer flattens the pooled multidimensional feature map into a one-dimensional feature vector for processing by the fully connected layer; The fully connected layer contains 64 neurons, which map the flattened feature vectors to a 64-dimensional feature space. The output layer ultimately outputs a subjective evaluation value.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention introduces a Convolutional Neural Network (CNN) to establish a mapping relationship between objective acoustic parameters and subjective ratings through automatic image feature extraction and regression learning, thereby achieving intelligent prediction of sound field zoning effects. Taking the zoned sound field as the research object, the invention collects the zoned sound field signals and calculates acoustic indices as model inputs. It then combines subjective listening evaluation scores as supervision signals to construct a CNN-based sound field zoning effect prediction model, providing efficient and scalable algorithmic support for intelligent cockpit acoustic optimization. Results show that this invention, while ensuring evaluation accuracy, possesses good generalization ability and can stably predict continuous subjective evaluation values, providing a valid basis for the optimization of sound field zoning systems and subjective listening design.

[0014] This invention is not only for evaluating the effects of generated sound fields, but also for using CNN models to predict the effects of sound field partitioning under different strategies or without prior adjustment, thereby reducing the number of subjective tests, lowering costs and improving design efficiency. Attached Figure Description

[0015] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a layout diagram of the in-vehicle speakers and microphones according to an embodiment of the present invention; Figure 3 This is a color block matrix diagram used as input to the CNN neural network in an embodiment of the present invention; Figure 4 This is a diagram of the CNN neural network structure according to an embodiment of the present invention; Figure 5 This is a prediction result diagram of the partitioning effect prediction model in an embodiment of the present invention; Figure 6 This is a loss curve of the partitioning effect prediction model in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] Example 1 In this embodiment, as Figure 1 As shown, a method for evaluating and predicting the sound field zoning of an intelligent cockpit includes the following steps: S1. Collect noise signals from the front and rear of the vehicle under different operating conditions in a semi-anechoic chamber.

[0020] The method for collecting noise signals includes: playing noise audio before and after sound field partitioning in a semi-anechoic chamber; collecting the first noise signal inside the vehicle when the evaluator is in different sitting positions (upright or leaning), and the second noise signal inside the vehicle when the evaluator is in different positions (driver's seat, passenger seat, left rear seat, right rear seat); and integrating the first and second noise signals to obtain the noise signal.

[0021] Specifically, such as Figure 2As shown, an LMS data acquisition unit, power amplifier, loudspeakers (Infinity, frequency range: 50~20000Hz), and PCB microphones were used to collect noise signals in the smart cockpit under different operating conditions. Eight loudspeakers were positioned above the dashboard, at the doors, and on the roof. PCB microphones were positioned at the ear positions of the driver and passengers. Noise signals were collected from the microphone locations in the cockpit under three different conditions: when there were passengers in the driver's and front passenger seats, when there were passengers in the driver's and rear left seats, and when there were passengers in the driver's and rear right seats. For each condition, two sets of data were collected: one for upright seating and one for reclining seating, resulting in a total of six sets of noise signals for six operating conditions. Additionally, a control group was set up for each condition, using audio from the loudspeakers without sound field zoning, collecting three sets of noise signals for the three operating conditions. The eight PCB microphones were divided into two groups, one in the acoustic bright area and one in the dark area. The acoustic bright zone refers to an area where sound wave energy is concentrated, sound pressure level is high, and sound propagation is clear with minimal distortion. The acoustic dark zone, on the other hand, is the opposite. In practical applications, the acoustic bright zone should be located in the area where the user needs to hear the in-car audio, while the dark zone should be located in the area where the user needs a quiet experience. Therefore, the areas on either side of the driver's side headrest are defined as the acoustic bright zone, and the areas on either side of the rear right side headrest are defined as the acoustic dark zone. The four microphones were positioned in the acoustic bright zone and the four microphones in the acoustic dark zone, respectively, in this experiment. In summary, a total of 36 sets of in-car noise signals were collected under 9 different operating conditions across 4 recording channels.

[0022] S2. Calculate the objective parameters of the sound field partition based on the noise signal.

[0023] Objective parameters for sound field zoning include: the ratio of the average sound potential energy of the bright area to that of the dark area ( AC ), percentage of normalized mean square error of the sound field in the bright area ( MSE ), sound pressure level, loudness, and sharpness. AC It is the sound energy contrast between the bright and dark areas, defined as the ratio of the average sound energy density of the bright area to that of the dark area. The greater the difference in sound energy between the bright and dark areas, the smaller the crosstalk between the areas. MSE It is the percentage of the normalized mean square error of the bright area sound field, defined as the ratio of the mean square sum of the bright area sound field reconstruction errors to the mean square sum of the target sound field, and is used to measure the quality of the bright area sound field reconstruction performance.

[0024] Sound pressure level, loudness, and sharpness are extracted directly from the noise signal. The ratio of the average sound potential energy of the bright area to that of the dark area is: in, AC This represents the ratio of the average acoustic potential energy of the bright area to that of the dark area. q This indicates the speaker array drive signal. zB This represents the electroacoustic transfer function in the bright region. zD Represents the electroacoustic transfer function within the dark region. H This indicates taking the conjugate; the percentage of normalized mean square error of the sound field in the bright area is: in, MSE This represents the percentage of the normalized mean square error of the sound field in the bright area. pB Indicates the reconstructed sound field of the bright area. pBT Indicates the target sound field in the bright area. f Indicates frequency, This represents solving for the square of the L2 norm.

[0025] Specifically, noise signal data collected from the semi-anechoic chamber vehicle test is extracted to a host PC. After obtaining a one-dimensional time-series graph of the noise signal sound pressure level in 1ms, the objective parameters of sound field zoning are first calculated. The audio signal data is 11 seconds long, and the loudness, sharpness, and sound pressure level of 9 groups of signals are calculated with a window of 0.1 seconds. Then, the sound pressure level spectrum is obtained by performing a Fourier transform on the one-dimensional time-series graph of the noise signal sound pressure level. The remaining objective parameters of sound field zoning are calculated using the time-domain and frequency-domain signals, namely the ratio of average acoustic potential energy (AC) between the bright and dark areas, the percentage of normalized mean square error (MSE) of the bright area, and the sound pressure level. Then, the set of results suitable for use as the input dataset for the subsequent CNN neural network prediction model is selected from the time and frequency domains and integrated with the calculated objective parameters of sound field zoning to prepare for the subsequent creation of the image dataset. In this embodiment, since the objective parameters of sound field zoning generally adopt the time dimension, the same time-domain signal is used as the result.

[0026] S3. The evaluation personnel will evaluate and score the noise to obtain the corresponding sound field zone evaluation value.

[0027] The method for obtaining sound field zoning evaluation values ​​includes: having multiple evaluators use a rating scale to quantify and score the noise signals before and after sound field zoning, thus obtaining multiple subjective evaluation scores; and calculating the correlation coefficient of the multiple subjective evaluation scores for data verification. in, r Represents the correlation coefficient. Xi and Yi This represents two scores given by the same evaluator to a noise sample. and This represents the average of two scores given by the same evaluator to all noise samples. Data from evaluators with lower correlation coefficients are removed to obtain the sound field zoning evaluation values. Evaluators score the noise level in the bright zone using pleasantness as an indicator and in the dark zone using annoyance as an indicator.

[0028] Specifically, a 20-member evaluation team was randomly selected from ordinary individuals aged 23 to 40, with a male-to-female ratio of 3:1, and all without hearing impairments. The evaluation team underwent initial auditory training to familiarize themselves with the experimental procedure. Then, 36 audio samples—9 sets of bright-area audio and 9 sets of dark-area audio from 4 recording channels—were played in different orders. Each audio set was divided into 11 one-second segments, with corresponding bright-area and dark-area audio segments played sequentially within the same time period. The 20 evaluators then took turns providing subjective evaluations and scores, resulting in 396 evaluation values ​​(198 for bright-area and 198 for dark-area). The subjective evaluation scoring method employed a rating scale, selecting pleasure as the evaluation indicator for the bright-area segment and annoyance as the indicator for the dark-area segment. Both evaluation indicators require participants to score based on their subjective feelings. Pleasure primarily reflects the extent to which participants feel pleasure, i.e., whether they find the sound pleasant, natural, comfortable, and relaxing. Annoyance, on the other hand, reflects the extent to which participants feel annoyed or uncomfortable, mainly reflecting the negative feelings the sound evokes, such as harshness, noise, oppression, or distraction. After obtaining the scores, the newly proposed subjective evaluation indicator for sound field zoning, "Comprehensive Zoning Comfort," is calculated based on the two sets of subjective evaluation scores. The difference between pleasure and annoyance is taken as the final score for Comprehensive Zoning Comfort. To ensure the reliability of the evaluators' scores, it is necessary to ensure that each evaluator uses the same standard when scoring the same set of sample audio in two playbacks. Therefore, a correlation coefficient needs to be calculated for data verification to ensure that the scores given by each evaluator for the same audio are not significantly different. In this experiment, the Pearson correlation coefficient is used, and the scores from evaluators with a correlation coefficient less than 0.6 are removed to obtain the final subjective evaluation score.

[0029] S4. Generate training samples based on the objective parameters and evaluation values ​​of the sound field partitions. Construct and train a CNN neural network based on the training samples to obtain a prediction model for the partitioning effect of the sound field in the intelligent cockpit.

[0030] In this embodiment, 10 sets of data are read in 0.1-second time windows. Each set of data contains 12 columns of features, such as... Figure 3As shown, the parameters include the average acoustic potential energy ratio (AC), the percentage of mean square error (MSE) of the normalized sound field in the bright area, the sound pressure level in the bright and dark areas, and the loudness and sharpness of each channel. Since there are 8 recording channels for the noise signal in the bright and dark areas under one working condition, two sets of data containing 12 columns of features can be obtained. Then, a unified global normalization strategy is adopted for all similar features. Based on the 5% and 95% quantiles of all valid data as the lower and upper limits of normalization, respectively, the features such as AC, MSE, sound pressure level, loudness, and sharpness are mapped to the [0,1] interval according to the normalization grouping to ensure that the feature scales of different data are consistent. After that, the normalized continuous time series data is divided into 10×12 two-dimensional color block matrix feature maps with a window of 1 second. In order to improve the generalization ability of the model, Gaussian noise, random brightness perturbation, contrast adjustment, and other data augmentation operations are applied to each original feature map, and its corresponding subjective evaluation label is inherited. Finally, all the original and enhanced feature maps are stored together, and label files containing image file names and subjective evaluation values, as well as label files that can be directly called by MATLAB, are generated, thus completing the construction of feature map samples suitable for training convolutional neural networks from raw acoustic data.

[0031] The CNN neural network consists of: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a flattening layer, a fully connected layer, and an output layer. The first convolutional layer uses a 3×3 kernel with 12 output channels to receive the input feature map. The first pooling layer uses a 2×2 pooling window with a stride of 2, halving the feature map space size. The second convolutional layer uses a 3×3 kernel with 24 output channels. The second pooling layer uses a 2×2 pooling window with a stride of 2, again halving the feature map space size. The flattening layer flattens the multi-dimensional feature map after pooling into a one-dimensional feature vector for processing by the fully connected layer. The fully connected layer contains 64 neurons, mapping the flattened feature vector to a 64-dimensional feature space. Finally, the output layer outputs a subjective evaluation value. The activation layers in the model's network architecture use the ReLU (Rectified Linear Unit) activation function, defined as: During network training, Half Mean Squared Error (HMSE) is used as the regression loss function, defined as: in, yp Indicates the predicted value. yt This represents the actual value.

[0032] Feature maps are input into the network via the input module and normalized within a predefined feature space to ensure consistency in numerical scale across samples. In the first stage of the network, shallow convolutional structures extract local acoustic feature patterns from the feature maps, including basic spatial relationships between sound pressure, loudness, sharpness, and subjective evaluation parameters. Convolutional operations slide within local neighborhoods, learning the variations in local parameters in different directions and combinations through multiple convolutional kernels. Subsequent ReLU nonlinear activation further enhances feature representation capabilities, enabling the network to recognize nonlinear acoustic patterns.

[0033] After initial feature extraction, the feature maps undergo spatial compression via a pooling module. This process suppresses local fluctuations, reduces redundant information, and improves the model's robustness to minor data drift and noise. The compressed features then proceed to deeper convolutional structures. In this stage, the network no longer focuses solely on simple local changes but abstractly represents the combined relationships between multiple acoustic parameters and the dynamic change patterns within a time window. This is achieved by increasing the channel dimension to obtain higher-level sound field partitioning features. Following deep convolution, pooling is performed again to preserve key structures at a more compact scale, enabling the network to comprehensively judge acoustic patterns across a wider receptive field.

[0034] After multi-level convolution and pooling, the network flattens the resulting multi-dimensional feature map into a one-dimensional feature vector for subsequent global mapping. The flattened feature vector contains global acoustic information from the input feature map, which is further mapped and integrated through a fully connected structure, enabling the network to learn the non-linear correspondence between different combinations of acoustic parameters and subjective evaluation values. A Dropout random deactivation mechanism is introduced into the fully connected layer, randomly discarding neurons with a probability of 0.2 to reduce the risk of overfitting and thus improve the model's generalization ability. After fully connected mapping, the network finally generates a prediction of the subjective evaluation value through a single output node, achieving end-to-end regression from multi-parameter acoustic features to continuous scores.

[0035] During training, the training, validation, and test sets are divided into three groups at a ratio of 70%, 15%, and 15%, respectively. During training, feature maps undergo forward propagation operations such as convolution, activation, pooling, flattening, and fully connected layers to obtain predicted values. Half-mean squared error (HMSE) is used as the regression loss function to calculate the HMSE loss between the predicted values ​​and the ground truth labels, and gradient backpropagation is performed through automatic differentiation. After training, the final partitioning effect prediction model and training history are saved for automatic prediction of subjective evaluation values ​​for sound field partitioning.

[0036] The partitioning effect prediction model constructed in this invention can automatically output subjective evaluation prediction values ​​after inputting acoustic features. The prediction results are as follows: Figure 5As shown, the model achieved RMSEs of 0.53, 0.60, and 0.76 on the training, validation, and test sets, respectively, with correlation coefficients... r The values ​​are 0.86, 0.83, and 0.76, respectively. The specific evaluation results are shown in Table 1.

[0037] Table 1 This indicates that the model has good prediction accuracy and generalization ability, and can stably predict the subjective effects of sound field zoning. Here, the root mean square error (RMSE) is the root mean square of the prediction error, and the correlation coefficient is... r The correlation coefficient indicates the degree of linear correlation between predicted and true values. Mean Absolute Percentage Error (MAPE) is the average percentage of relative error. Smaller RMSE and MAPE values ​​indicate better model training results. r The closer the value is to 1, the more accurate the prediction. The loss curve is as follows: Figure 6 As shown. Therefore, the training results indicate that the model has good evaluation accuracy, and the differences between the training set, validation set, and test set are small, indicating that the model also has good generalization ability. This embodiment can not only evaluate the performance of sound field partitioning, but also predict the subjective listening results under different partitioning strategies without subjective evaluation.

[0038] S5. Use the zoning effect prediction model to evaluate and predict the sound field zoning.

[0039] Example 2 In this embodiment, a smart cockpit sound field zoning evaluation and prediction system includes: a noise signal acquisition module, an objective data calculation module, a subjective evaluation value acquisition module, a model construction module, and a prediction module.

[0040] The noise signal acquisition module is used to acquire noise signals from the front and rear of the vehicle under different operating conditions in a semi-anechoic chamber.

[0041] The workflow of the noise signal acquisition module includes: playing noise audio before and after sound field partitioning in a semi-anechoic chamber; acquiring the first noise signal inside the vehicle when the evaluator is in different sitting positions, and the second noise signal inside the vehicle when the evaluator is distributed in different positions; and integrating the first and second noise signals to obtain the noise signal.

[0042] The objective data calculation module is used to calculate the objective parameters of the sound field partition based on the noise signal.

[0043] Objective parameters for sound field zoning include: the ratio of average sound potential energy between the bright and dark areas, the percentage of normalized mean square error of the sound field in the bright area, sound pressure level, loudness, and sharpness. Sound pressure level, loudness, and sharpness are extracted directly from the noise signal; the ratio of average sound potential energy between the bright and dark areas is: in, AC This represents the ratio of the average acoustic potential energy of the bright area to that of the dark area. q This indicates the speaker array drive signal. zB This represents the electroacoustic transfer function in the bright region. zD Represents the electroacoustic transfer function within the dark region. H This indicates taking the conjugate; the percentage of normalized mean square error of the sound field in the bright area is: in, MSE This represents the percentage of the normalized mean square error of the sound field in the bright area. pB Indicates the reconstructed sound field of the bright area. pBT Indicates the target sound field in the bright area. f Indicates frequency, This represents solving for the square of the L2 norm.

[0044] The subjective evaluation value acquisition module allows evaluators to assess and score the noise, thereby obtaining the corresponding sound field zone evaluation value.

[0045] The workflow of the subjective evaluation value acquisition module includes: quantifying and scoring the noise signals before and after the sound field partitioning by multiple evaluators using a rating scale to obtain multiple subjective evaluation scores; calculating the correlation coefficient of the multiple subjective evaluation scores for data verification. in, r Represents the correlation coefficient. Xi and Yi This represents two scores given by the same evaluator to a noise sample. and This represents the average score of two ratings given by the same evaluator for all noise samples; data from evaluators with lower results in the correlation coefficient calculation are removed to obtain the sound field zoning evaluation value.

[0046] The model building module is used to generate training samples based on the objective parameters and evaluation values ​​of the sound field partitions. Based on the training samples, a CNN neural network is built and trained to obtain a partition effect prediction model for the sound field of the intelligent cockpit.

[0047] The CNN neural network consists of: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a flattening layer, a fully connected layer, and an output layer. The first convolutional layer uses a 3×3 convolutional kernel with 12 output channels to receive the input feature map. The first pooling layer uses a 2×2 pooling window with a stride of 2, halving the feature map space size. The second convolutional layer uses a 3×3 convolutional kernel with 24 output channels. The second pooling layer uses a 2×2 pooling window with a stride of 2, again halving the feature map space size. The flattening layer flattens the multi-dimensional feature map after pooling into a one-dimensional feature vector for processing by the fully connected layer. The fully connected layer contains 64 neurons, mapping the flattened feature vector to a 64-dimensional feature space. Finally, the output layer outputs a subjective evaluation value.

[0048] The prediction module uses a zone effect prediction model to evaluate and predict the sound field zones.

[0049] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for evaluating and predicting the sound field zoning of an intelligent cockpit, characterized in that, Includes the following steps: Noise signals were collected from the front and rear of the vehicle under different operating conditions in a semi-anechoic chamber; Calculate the objective parameters of the sound field partitioning based on the noise signal; The noise is evaluated and scored by evaluators to obtain the corresponding sound field zone evaluation value; Training samples are generated based on the objective parameters of the sound field partition and the evaluation value of the sound field partition. A CNN neural network is constructed and trained based on the training samples to obtain a partition effect prediction model of the sound field of the intelligent cockpit. The sound field zoning evaluation prediction is performed using the zoning effect prediction model.

2. The method for evaluating and predicting the sound field zoning of an intelligent cockpit according to claim 1, characterized in that, The method for acquiring the noise signal includes: In the semi-anechoic chamber, noise audio before and after the sound field partitioning are played respectively. The first noise signal inside the vehicle was collected when the evaluators were in different sitting positions, and the second noise signal inside the vehicle was collected when the evaluators were distributed in different positions. The noise signal is obtained by integrating the first noise signal and the second noise signal.

3. The method for evaluating and predicting the sound field zoning of an intelligent cockpit according to claim 1, characterized in that, The objective parameters of the sound field partitioning include: the ratio of the average sound potential energy of the bright area to the dark area, the percentage of the normalized sound field mean square error in the bright area, the sound pressure level, loudness, and sharpness. The sound pressure level, the loudness, and the sharpness are extracted directly from the noise signal; The ratio of the average acoustic potential energy of the bright area to that of the dark area is: in, AC This represents the ratio of the average acoustic potential energy of the bright area to that of the dark area. q Indicates the speaker array drive signal. zB This represents the electroacoustic transfer function in the bright region. zD Represents the electroacoustic transfer function within the dark region. H Indicates taking the conjugate; The percentage of the normalized sound field mean square error in the bright area is: in, MSE This represents the percentage of the normalized mean square error of the sound field in the bright area. pB Indicates the reconstructed sound field of the bright area. pBT Indicates the target sound field in the bright area. f Indicates frequency, This represents solving for the square of the L2 norm.

4. The method for evaluating and predicting the sound field zoning of an intelligent cockpit according to claim 1, characterized in that, The methods for obtaining the sound field partition evaluation values ​​include: Multiple evaluators used a rating scale to quantify and score the noise signals before and after the sound field partitioning, resulting in multiple subjective evaluation scores. Calculate the correlation coefficients of the multiple subjective evaluation scores to perform data validation: in, r Represents the correlation coefficient. Xi and Yi This represents two scores given by the same evaluator to a noise sample. and This represents the average of two scores given by the same evaluator for all noise samples. The data of evaluators with low correlation coefficients are removed to obtain the sound field zoning evaluation values.

5. The method for evaluating and predicting the sound field zoning of an intelligent cockpit according to claim 1, characterized in that, The CNN neural network includes: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a flattening layer, a fully connected layer, and an output layer; The first convolutional layer uses a 3×3 convolutional kernel and has 12 output channels to receive the input feature map; The first pooling layer uses a 2×2 pooling window with a stride of 2, which halves the feature map space size. The second convolutional layer uses a 3×3 convolutional kernel and has 24 output channels; The second pooling layer uses a 2×2 pooling window with a stride of 2, which again halves the feature map space size; The flattening layer flattens the pooled multidimensional feature map into a one-dimensional feature vector for processing by the fully connected layer; The fully connected layer contains 64 neurons, which map the flattened feature vectors to a 64-dimensional feature space. The output layer ultimately outputs a subjective evaluation value.

6. A smart cockpit sound field zoning evaluation and prediction system, wherein the system applies the method described in any one of claims 1-5, characterized in that, include: The module includes a noise signal acquisition module, an objective data calculation module, a subjective evaluation value acquisition module, a model building module, and a prediction module. The noise signal acquisition module is used to acquire noise signals inside the vehicle before and after the sound field partition under different working conditions in a semi-anechoic chamber. The objective data calculation module is used to calculate the objective parameters of the sound field partitioning based on the noise signal; The subjective evaluation value acquisition module obtains the corresponding sound field zone evaluation value by evaluating personnel to score the noise. The model building module is used to generate training samples based on the objective parameters of the sound field partition and the evaluation value of the sound field partition, and to build and train a CNN neural network based on the training samples to obtain a partition effect prediction model of the sound field of the intelligent cockpit. The prediction module uses the zoning effect prediction model to perform sound field zoning evaluation and prediction.

7. The intelligent cockpit sound field zoning evaluation and prediction system according to claim 6, characterized in that, The workflow of the noise signal acquisition module includes: In the semi-anechoic chamber, noise audio before and after the sound field partitioning are played respectively. The first noise signal inside the vehicle was collected when the evaluators were in different sitting positions, and the second noise signal inside the vehicle was collected when the evaluators were distributed in different positions. The noise signal is obtained by integrating the first noise signal and the second noise signal.

8. The intelligent cockpit sound field zoning evaluation and prediction system according to claim 6, characterized in that, The objective parameters of the sound field partitioning include: the ratio of the average sound potential energy of the bright area to the dark area, the percentage of the normalized sound field mean square error in the bright area, the sound pressure level, loudness, and sharpness. The sound pressure level, the loudness, and the sharpness are extracted directly from the noise signal; The ratio of the average acoustic potential energy of the bright area to that of the dark area is: in, AC This represents the ratio of the average acoustic potential energy of the bright area to that of the dark area. q Indicates the speaker array drive signal. zB This represents the electroacoustic transfer function in the bright region. zD Represents the electroacoustic transfer function within the dark region. H Indicates taking the conjugate; The percentage of the normalized sound field mean square error in the bright area is: in, MSE This represents the percentage of the normalized mean square error of the sound field in the bright area. pB Indicates the reconstructed sound field of the bright area. pBT Indicates the target sound field in the bright area. f Indicates frequency, This represents solving for the square of the L2 norm.

9. The intelligent cockpit sound field zoning evaluation and prediction system according to claim 6, characterized in that, The workflow of the subjective evaluation value acquisition module includes: Multiple evaluators used a rating scale to quantify and score the noise signals before and after the sound field partitioning, resulting in multiple subjective evaluation scores. Calculate the correlation coefficients of the multiple subjective evaluation scores to perform data validation: in, r Represents the correlation coefficient. Xi and Yi This represents two scores given by the same evaluator to a noise sample. and This represents the average of two scores given by the same evaluator for all noise samples. The data of evaluators with low correlation coefficients are removed to obtain the sound field zoning evaluation values.

10. The intelligent cockpit sound field zoning evaluation and prediction system according to claim 6, characterized in that, The CNN neural network includes: a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a flattening layer, a fully connected layer, and an output layer; The first convolutional layer uses a 3×3 convolutional kernel and has 12 output channels to receive the input feature map; The first pooling layer uses a 2×2 pooling window with a stride of 2, which halves the feature map space size. The second convolutional layer uses a 3×3 convolutional kernel and has 24 output channels; The second pooling layer uses a 2×2 pooling window with a stride of 2, which again halves the feature map space size; The flattening layer flattens the pooled multidimensional feature map into a one-dimensional feature vector for processing by the fully connected layer; The fully connected layer contains 64 neurons, which map the flattened feature vectors to a 64-dimensional feature space. The output layer ultimately outputs a subjective evaluation value.