Sound quality evaluation method and terminal device
Patent Information
- Application Number
- CN202610865287.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]然而,传统技术中直接采用全部24个Bark子带特征建模,存在特征维度偏高、冗余信息繁杂的问题,会影响噪声评价的精度
Smart Images

Figure CN122738540A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sound quality evaluation technology, and in particular to a sound quality evaluation method and terminal equipment. Background Technology
[0002] As people's living standards improve, users are paying more and more attention to the comfort of using equipment, and the noise level generated during equipment operation has become a key concern.
[0003] Currently, sound quality assessment often adopts a subjective-objective integrated assessment method. First, noise signals are collected, and psychoacoustic indicators such as loudness and sharpness are extracted along with Bark domain frequency band characteristics. Combined with subjective listening scores, a mapping model between objective characteristics and subjective feelings is built to quantitatively predict noise sound quality.
[0004] However, traditional techniques that directly use all 24 Bark subband features for modeling suffer from high feature dimensionality and redundant information, which can affect the accuracy of noise evaluation. Summary of the Invention
[0005] This application provides a sound quality evaluation method and terminal device that can retain the target acoustic characteristics of noise signals while reducing the number of frequency bands and psychoacoustic indicators, thereby improving the efficiency and accuracy of sound quality evaluation.
[0006] In a first aspect, some embodiments provide a sound quality evaluation method, including:
[0007] Noise sample signals of the device under test under different operating conditions are acquired, and the psychoacoustic features of the noise sample signals are extracted in different candidate frequency bands; the different candidate frequency bands include the full frequency band and each auditory critical frequency band;
[0008] Based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands, the target acoustic index is selected from the psychoacoustic indexes.
[0009] Based on the psychoacoustic characteristics of each noise sample signal for the same auditory critical band under different target acoustic indicators, auditory critical bands that meet the merging conditions are merged.
[0010] Based on the target acoustic characteristics of each noise sample signal in the full frequency band and the combined auditory critical frequency band, the sound quality of the equipment noise under test is evaluated; the target acoustic characteristics are determined based on the psychoacoustic characteristics under the corresponding target acoustic indicators.
[0011] In the above embodiments, compared to traditional techniques that use the characteristics of all frequency bands under all psychoacoustic indicators to evaluate the sound quality of noise, the above process, on the one hand, selects the main indicators with higher discriminative power and reflecting human auditory perception as target acoustic indicators based on the differences between the psychoacoustic characteristics of the noise sample signal under various psychoacoustic indicators. This can retain the core psychoacoustic indicators while reducing the amount of subsequent data processing, avoiding a decrease in evaluation accuracy. On the other hand, while retaining the overall noise characteristics of the entire frequency band, the auditory critical frequency bands with similar characteristics or similar auditory perception are reasonably merged, which can further reduce the feature dimensionality while retaining the core noise auditory characteristics. In other words, the entire process can retain the target acoustic characteristics of the noise signal while reducing the number of frequency bands and psychoacoustic indicators, improving the efficiency and accuracy of sound quality evaluation.
[0012] Secondly, some embodiments also provide a sound quality evaluation device, including:
[0013] The feature extraction module is used to acquire noise sample signals of the device under test under different operating conditions, and extract the psychoacoustic features of the noise sample signals under different candidate frequency bands; the different candidate frequency bands include the full frequency band and each auditory critical frequency band;
[0014] The index selection module is used to select the target acoustic index from the psychoacoustic indices based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands.
[0015] The frequency band merging module is used to merge the auditory critical frequency bands that meet the merging conditions based on the psychoacoustic characteristics of each noise sample signal for the same auditory critical frequency band under different target acoustic indicators.
[0016] The sound quality evaluation module is used to evaluate the sound quality of the equipment noise of the device under test based on the target acoustic characteristics of each noise sample signal in the full frequency band and the combined auditory critical frequency band; the target acoustic characteristics are determined based on the psychoacoustic characteristics under the corresponding target acoustic indicators.
[0017] Thirdly, some embodiments also provide a terminal device, including:
[0018] At least one processor, and configured as follows:
[0019] Noise sample signals of the device under test under different operating conditions are acquired, and the psychoacoustic features of the noise sample signals under different candidate frequency bands are extracted; the different candidate frequency bands include the full frequency band and each auditory critical frequency band;
[0020] Based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands, the target acoustic index is selected from the psychoacoustic indexes.
[0021] Based on the psychoacoustic characteristics of each noise sample signal for the same auditory critical band under different target acoustic indicators, auditory critical bands that meet the merging conditions are merged.
[0022] Based on the target acoustic characteristics of each noise sample signal in the full frequency band and the combined auditory critical frequency band, the sound quality of the equipment noise under test is evaluated; the target acoustic characteristics are determined based on the psychoacoustic characteristics under the corresponding target acoustic indicators.
[0023] Fourthly, some embodiments also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the methods provided in some embodiments of the first aspect.
[0024] Fifthly, some embodiments also provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods provided in some embodiments of the first aspect. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 A schematic flowchart illustrating the sound quality evaluation method provided in some embodiments of this application;
[0027] Figure 2 This application provides schematic diagrams showing microphone deployment locations for some embodiments.
[0028] Figure 3 A flowchart illustrating the steps for selecting target acoustic parameters provided in some embodiments of this application;
[0029] Figure 4 A flowchart illustrating the auditory critical band merging steps provided in some embodiments of this application;
[0030] Figure 5 A flowchart illustrating the auditory critical band merging steps provided for other embodiments of this application;
[0031] Figure 6 A flowchart illustrating the relevant threshold update steps provided in some embodiments of this application;
[0032] Figure 7A flowchart illustrating the reference threshold update steps provided in some embodiments of this application;
[0033] Figure 8 Internal structural diagrams of the sound quality evaluation device provided in some embodiments of this application;
[0034] Figure 9 This is an internal structural diagram of a computer device provided in some embodiments of this application. Detailed Implementation
[0035] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.
[0036] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0037] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0038] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0039] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0040] Extracting psychoacoustic features by simulating human hearing characteristics in the Bark domain is the mainstream technical approach for sound quality assessment. However, traditional sound quality assessment mainly relies on noise signal acquisition, psychoacoustic index extraction, subjective listening scoring, and modeling mapping prediction. Typically, the Bark domain is divided into 24 critical frequency bands (i.e., the aforementioned auditory critical frequency bands), and then all psychoacoustic features of the noise signal in the full frequency band and the 24 critical frequency bands are extracted as input features for the model. However, using the full frequency band plus the 24 critical frequency bands, as well as psychoacoustic features of all dimensions, for sound quality assessment results in high feature dimensionality and a large proportion of redundant information. This not only increases model training time and reduces overall computational efficiency, but also easily leads to model overfitting, weakening the algorithm's generalization performance in different scenarios. At the same time, sensitive frequency band features closely related to human subjective perception in high-dimensional data are easily covered by irrelevant information, further reducing the prediction accuracy of the subjective-objective mapping model, i.e., the sound quality assessment model. However, there is still a lack of perfect solutions for how to reduce redundant frequency bands and / or reduce feature dimensions while preserving core auditory features, and to balance computational efficiency and model prediction accuracy.
[0041] Based on this, in some embodiments, a sound quality evaluation method is provided. This sound quality evaluation method can be implemented by a terminal device.
[0042] In an exemplary embodiment, taking the application of the sound quality evaluation method to a terminal device as an example, the terminal device includes at least one processor for performing matters related to sound quality evaluation, such as noise signal preprocessing; further, the terminal device may also include a sound acquisition device. The sound acquisition device can be used to receive noise signals emitted by the device under test. The device under test can be a common device that emits noise during operation, such as an air conditioner or refrigerator. In this embodiment, an air conditioner is used as an example for illustration.
[0043] like Figure 1 As shown, the sound quality evaluation method for a processor applied in a terminal includes the following steps:
[0044] S110: Acquire noise sample signals of the device under test under different operating conditions, and extract the psychoacoustic features of the noise sample signals under different candidate frequency bands.
[0045] Taking an air conditioner as an example, the different operating conditions may include, but are not limited to, cooling, heating, air supply, and dehumidification. In some embodiments, a sound acquisition device can be used to collect noise sample signals of the device under different operating conditions, and the noise sample signals collected by the sound acquisition device can be obtained.
[0046] The sound acquisition device may include, but is not limited to, a microphone. Optionally, multiple microphones can be deployed in the noise sample signal acquisition environment, and an air conditioner prototype can be installed. Then, the air conditioner prototype can be controlled to operate under different operating conditions in the same environment. The noise signals generated by the air conditioner prototype under different operating conditions can be collected synchronously by each microphone, thereby completing the acquisition of the noise sample signal.
[0047] This application does not impose any limitations on the microphone placement. For example, since the noise level may vary depending on the location of the device under test (DUT) during operation, the microphone placement stage needs to consider factors such as the noise radiation direction, sound source distance, sound field distribution, and device structure to improve the comprehensiveness of the noise sample signal. For example, the microphone can be deployed around the DUT, or directly in front of, to the side of, or above the main noise radiation surface of the DUT. It is understood that different types of DUTs have different noise generation characteristics; therefore, the microphone placement scheme can differ for different DUTs.
[0048] Taking an air conditioner as an example, the air conditioner can be divided into an indoor unit and an outdoor unit. The microphone deployment diagram for collecting noise sample signals from the indoor and outdoor units can be shown as follows: Figure 2 As shown. For the indoor unit, microphones can be deployed on the air supply side, air return side, and the listening area where people are usually active, to comprehensively collect airflow noise and noise signals generated by the unit's operation; for the outdoor unit, microphones can be deployed around the unit body at the locations corresponding to the main noise sources such as the fan outlet and the compressor, to completely collect noise signals generated by the outdoor unit's operation and airflow disturbances.
[0049] Among them, the different candidate frequency bands include the full frequency band and each auditory critical frequency band (that is, 24 Bark domains).
[0050] Among them, psychoacoustic features can be used to characterize the auditory stimulation and subjective perception characteristics of noise on the human ear. For example, psychoacoustic features may include, but are not limited to, at least one of loudness, sharpness, roughness, fluctuation intensity, pitch, clarity index, tone modulation, prominence, and noise annoyance.
[0051] In one alternative implementation, for any noise sample signal, a Bark domain critical band decomposition can be obtained to simulate the frequency characteristics of the human ear basilar membrane, thereby determining the acoustic signal of the noise sample signal under different candidate frequency bands; for the acoustic signal of the noise sample signal under each candidate frequency band, the acoustic signal is input into a psychological feature extraction model to obtain the psychoacoustic features of the noise sample signal under the corresponding candidate frequency band.
[0052] S120: Select the target acoustic index from the psychoacoustic indices based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands.
[0053] Among them, target acoustic indicators can be understood as psychoacoustic indicators that have a high correlation with subjective auditory perception and significant differences in characteristics under different working conditions.
[0054] Among them, psychoacoustic features can be the conversion of noise sample signals into psychoacoustic index values that conform to the characteristics of human hearing.
[0055] In one alternative implementation, a target acoustic index can be selected from psychoacoustic indices based on the statistical values (e.g., mean and standard deviation) of the psychoacoustic characteristics of each noise sample signal under different candidate frequency bands for the same psychoacoustic index, and the difference between these statistical values and the characteristic benchmark values under different operating conditions. For example, when the statistical value exceeds a preset operating condition difference threshold, the corresponding psychoacoustic index is used as the target acoustic index.
[0056] In another alternative implementation, the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands can be input into the index selection model. The index selection model can then comprehensively evaluate the working condition discrimination and subjective correlation of each psychoacoustic index to select the target acoustic index.
[0057] To facilitate understanding, taking W as an example with candidate frequency bands including the entire domain and 24 Bark domains, and combining the psychoacoustic characteristics of the noise sample signals in the table below, the selection of target acoustic indicators is illustrated.
[0058] Sample number Bark domain Loudness (sone) Sharpness (acum) Roughness (asper) jitter (vacil) ... Sample 1 Global ... ... ... ... ... Sample 1 Bark 1 2.1 0.8 0.3 0.05 ... Sample 1 Bark 2 2.3 0.9 0.4 0.06 ... ... ... ... ... ... ... ... Sample 1 Bark 24 0.5 0.2 0.1 0.01 ...
[0059] The table records the psychoacoustic characteristics of the same noise sample signal (Sample 1) across the entire frequency range and in each candidate frequency band from Bark1 to Bark24, including loudness, sharpness, roughness, and jitter. In the Bark1 band, the sample has a loudness of 2.1 sone, a sharpness of 0.8 acum, a roughness of 0.3 asper, and a jitter of 0.05 vacil; in the Bark2 band, the loudness is 2.3 sone, the sharpness is 0.9 acum, the roughness is 0.4 asper, and the jitter is 0.06 vacil; and in the Bark24 band, the loudness is 0.5 sone, the sharpness is 0.2 acum, the roughness is 0.1 asper, and the jitter is 0.01 vacil.
[0060] For the same psychoacoustic index, the psychoacoustic feature values of all noise sample signals in all candidate frequency bands are traversed, and the statistical values such as the mean and variance of the index's features in different Bark frequency bands are calculated to evaluate the index's ability to characterize frequency band differences and noise characteristics.
[0061] Taking the psychoacoustic characteristics of Sample 1 as an example, calculations show that the mean loudness is 1.633 with a variance of 0.973; the mean sharpness is 0.633 with a variance of 0.143; the mean roughness is 0.267 with a variance of 0.023; and the mean jitter is 0.040 with a variance of 0.0007. Comparing the values of each psychoacoustic indicator with statistical values reveals that loudness and sharpness exhibit large fluctuations across candidate frequency bands, effectively distinguishing noise characteristics from different Bark frequency bands; roughness is next; and jitter, with its overall low value and weak differences between frequency bands, has poor ability to distinguish noise frequency band characteristics. Based on these reasons, loudness, sharpness, and roughness are selected as the target acoustic indicators, while jitter, with its insufficient distinguishing ability, is discarded.
[0062] S130: Based on the psychoacoustic characteristics of each noise sample signal for the same auditory critical band under different target acoustic indicators, merge the auditory critical bands that meet the merging conditions.
[0063] The merging conditions can be determined based on a large number of experiments. The main principle is that the psychoacoustic characteristics of the frequency bands to be merged are similar, and the noise perception characteristics are similar or even basically the same. For example, between two pairs of auditory critical frequency bands, the statistical difference, spatial distance, or correlation of the acoustic characteristics of each target are all less than a preset threshold.
[0064] In one optional implementation, the characteristic mean of the auditory critical frequency band under each target acoustic index can be determined separately, and the difference between the characteristic mean of any two adjacent auditory critical frequency bands under the same target acoustic index can be determined sequentially. If the number of target acoustic indices with a difference less than a preset difference threshold exceeds a preset number, or even if the difference between the mean values corresponding to all target acoustic indices is less than the preset difference threshold, then the two sets of auditory critical frequency bands are determined to meet the merging condition, and the two are merged into a new frequency band unit.
[0065] To facilitate understanding, we will continue with the example above, which includes target acoustic indicators such as loudness, sharpness, and roughness. First, we compare Bark1 and Bark2: we calculate the feature differences item by item. The loudness difference is 0.2, the sharpness difference is 0.1, and the roughness difference is 0.1. The differences of the three target acoustic indicators do not exceed the preset threshold (e.g., the preset threshold is 0.5), proving that the psychoacoustic characteristics of the two frequency bands are relatively similar. Therefore, Bark1 and Bark2 are merged into a joint frequency band.
[0066] Next, Bark1 was compared with Bark24, and Bark2 with Bark24 respectively: the loudness difference between Bark1 and Bark24 was 1.6, the sharpness difference was 0.6, and the roughness difference was 0.2; the loudness difference between Bark2 and Bark24 was 1.8, the sharpness difference was 0.7, and the roughness difference was 0.3. The differences in multiple target acoustic indicators were greater than the preset threshold, which proved that the acoustic characteristics between frequency bands were significantly different and did not meet the merging conditions. Therefore, they could not be merged.
[0067] Experimental data shows that the acoustic characteristics of frequency bands that are far apart are usually more different. Therefore, in order to improve the efficiency of frequency band merging, only adjacent auditory critical frequency bands are usually traversed, compared and merged, skipping combinations of frequency bands that are far apart, thus reducing the amount of invalid calculations.
[0068] S140, based on the target acoustic characteristics of each noise sample signal in the full frequency band and the combined auditory critical frequency band, evaluates the sound quality of the equipment noise of the device under test.
[0069] The target acoustic features are determined based on the psychoacoustic features under the corresponding target acoustic indicators. For example, for a frequency band that has been merged, the psychoacoustic features of all its sub-bands under each target acoustic indicator can be extracted, and for any psychoacoustic indicator, the overall features of the merged frequency band can be calculated by means of averaging, weighted averaging, or normalization, and used as the target acoustic features.
[0070] Using the previous example, taking the target acoustic features as the corresponding mean values, Bark1 and Bark2 have been merged into a joint frequency band. The original features of this joint frequency band in terms of loudness, sharpness, and roughness are 2.1sone / 2.3sone, 0.8acum / 0.9acum, and 0.3asper / 0.4asper, respectively. After calculating the mean values, the corresponding values are obtained and used as the target acoustic features of this merged frequency band, that is, loudness 2.2sone, sharpness 0.85acum, and roughness 0.35asper.
[0071] In one alternative implementation, the target acoustic features of the full-band and combined auditory critical band corresponding to the noise of the device under test can be input into a pre-trained sound quality evaluation model, and the model can output the corresponding sound quality evaluation results to evaluate the sound quality of the device noise of the device under test.
[0072] In the aforementioned sound quality evaluation method, compared to traditional techniques that use the characteristics of all frequency bands under all psychoacoustic indicators to evaluate noise sound quality, this process, on the one hand, selects key indicators with higher discriminative power and reflecting human auditory perception as target acoustic indicators based on the differences in psychoacoustic characteristics of noise sample signals under various psychoacoustic indicators. This reduces the amount of subsequent data processing while retaining core psychoacoustic indicators, avoiding a decrease in evaluation accuracy. On the other hand, while preserving the overall noise characteristics of the entire frequency band, it reasonably merges auditory critical frequency bands with similar characteristics or similar auditory perception, further reducing the feature dimensionality while retaining the core noise auditory characteristics. In other words, the entire process can reduce the number of frequency bands and psychoacoustic indicators while preserving the target acoustic characteristics of the noise signal, improving the efficiency and accuracy of sound quality evaluation.
[0073] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the process of selecting a target acoustic index from psychoacoustic indices based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index in different candidate frequency bands is described.
[0074] See Figure 3 The steps for selecting the target acoustic parameters shown include:
[0075] S310, for any psychoacoustic index, determines the degree of characteristic difference of noise sample signals under different operating conditions under the psychoacoustic index.
[0076] The degree of difference in the aforementioned characteristics can reflect the ability of the psychoacoustic index to distinguish the noise characteristics of different working conditions, and can be used as a basis for judging target acoustic indicators.
[0077] In one alternative implementation, the global mean of the psychoacoustic characteristics of all noise sample signals under the psychoacoustic index can be determined; the local mean of the psychoacoustic characteristics of noise sample signals under different operating conditions under the psychoacoustic index can be determined; and the degree of characteristic difference of noise sample signals under different operating conditions under the psychoacoustic index can be determined based on the number of noise sample signals under each operating condition and the mean deviation between the local mean and the global mean under the corresponding operating conditions.
[0078] The global mean is used to characterize the overall characteristic level of the psychoacoustic index across all operating conditions; the local mean is used to characterize the characteristic level of the psychoacoustic index under a single operating condition.
[0079] In one optional implementation, the characteristic difference between the local mean and the global mean under each operating condition can be determined first, and the product of this characteristic difference and the number of noise sample signals under the corresponding operating condition can be used as the operating condition deviation. The operating condition deviations under all operating conditions are then integrated to obtain a sum of operating condition deviations. This sum is divided by the difference between the total number of operating condition types and 1 to obtain the degree of characteristic difference corresponding to the psychoacoustic index. For example, the degree of characteristic difference of noise sample signals under different operating conditions under this psychoacoustic index can be determined using the following formula:
[0080] ;
[0081] In the formula, D represents the degree of difference in the characteristics of noise sample signals under different operating conditions under this psychoacoustic index; M represents the total number of operating condition types; and C represents the Cth type of operating condition. This represents the number of noise sample signals under Class C operating conditions; This represents the local mean of the psychoacoustic characteristics of the noise sample signal under the above psychoacoustic indices under the Class C operating condition; This represents the global mean of the psychoacoustic characteristics of all noise sample signals under this psychoacoustic index.
[0082] S320 determines the characteristic fluctuation degree of different noise sample signals under the same operating conditions in terms of psychoacoustic indicators.
[0083] Among them, the aforementioned characteristic fluctuation level can characterize the discrete fluctuation level of the noise sample signal under the corresponding psychoacoustic index within a single operating condition, and can reflect the stability of the psychoacoustic index under the same operating condition. It can be used as a basis for screening target acoustic indicators.
[0084] In one optional implementation, for any operating condition, the characteristic deviation between the psychoacoustic characteristics of each noise sample signal under the psychoacoustic index and the corresponding local mean can be determined; based on the characteristic deviation, the total number of all noise sample signals, and the total number of categories of the operating condition, the characteristic fluctuation degree of different noise sample signals under the psychoacoustic index under the operating condition can be determined.
[0085] Among them, the aforementioned characteristic deviation can characterize the degree of deviation between the psychoacoustic characteristics of a single noise sample signal and the overall level of the corresponding operating condition, and can reflect the fluctuation of the sample within the same operating condition.
[0086] In one optional implementation, for any operating condition, the characteristic deviation between the psychoacoustic characteristics of each noise sample signal under the aforementioned psychoacoustic index and the corresponding local mean of that operating condition can be determined. The characteristic deviations of all noise sample signals under a single operating condition are squared and summed to obtain the sum of squares of the within-group deviations for that operating condition. Then, the sums of squares of the within-group deviations for all types of operating conditions are summarized to obtain the total sum of squares of the within-group deviations for all samples. The total number of noise sample signals and the total number of operating condition types for all operating conditions are obtained. The degrees of freedom are obtained by subtracting the total number of operating condition types from the total number of samples. The total sum of squares of the within-group deviations is divided by the degrees of freedom to finally obtain the characteristic fluctuation degree corresponding to the psychoacoustic index. For example, the characteristic fluctuation degree of different noise sample signals under any operating condition under the psychoacoustic index can be determined based on the following formula:
[0087] ;
[0088] In the formula, S represents the characteristic fluctuation degree of different noise sample signals under the psychoacoustic index under operating condition C; N represents the total number of samples; M represents the total number of operating condition types; and C represents the Cth type of operating condition. This represents the number of noise sample signals under Class C operating conditions; This represents the psychoacoustic characteristics of the nth noise sample signal under the above psychoacoustic indices during the C-type operating condition. This represents the local mean of the psychoacoustic characteristics of the noise sample signal under the above psychoacoustic indices under the Class C operating condition.
[0089] S330, the reference coefficient of the psychoacoustic index is determined based on the ratio between the degree of characteristic difference and the degree of characteristic fluctuation.
[0090] Among them, the index reference coefficient can characterize the reference value of the corresponding psychoacoustic index for evaluating the sound quality of equipment noise. The higher the index reference coefficient, the higher the reference value of the corresponding psychoacoustic index; the lower the index reference coefficient, the lower the reference value of the corresponding psychoacoustic index.
[0091] For example, the reference coefficient for this psychoacoustic index can be determined using the following formula:
[0092] ;
[0093] In the formula, D represents the degree of difference in the characteristics of noise sample signals under different operating conditions under this psychoacoustic index; M represents the total number of operating condition types; and C represents the Cth type of operating condition. This represents the number of noise sample signals under Class C operating conditions; This represents the local mean of the psychoacoustic characteristics of the noise sample signal under the above psychoacoustic indices under the Class C operating condition; S represents the global mean of the psychoacoustic characteristics of all noise sample signals under this psychoacoustic index; S represents the degree of characteristic fluctuation of different noise sample signals under the psychoacoustic index under operating condition C; N represents the total number of samples. This represents the psychoacoustic characteristics of the nth noise sample signal under the above psychoacoustic indices during the C-type operating condition.
[0094] S340, select psychoacoustic indicators whose reference coefficients exceed the reference threshold as target acoustic indicators.
[0095] The reference threshold can be determined based on human experience, through extensive experimentation, or according to the required accuracy of sound quality evaluation. For example, different reference thresholds can be set for different levels of sound quality evaluation accuracy. For instance, for high accuracy requirements, the reference threshold could be 15; for general requirements, it could be 10; and for preliminary screening, it could be 5.
[0096] It is understandable that updating the reference threshold directly affects the selection of target acoustic indicators, and thus the results of sound quality evaluation. Therefore, after selecting target acoustic indicators using noise sample signals, the rationality of the selection can be verified, and the reference threshold can be adjusted if it is unreasonable, to ensure that the selected target acoustic indicators are typical and to guarantee the accuracy and reliability of subsequent sound quality evaluation results. The process of updating the reference threshold will be described in subsequent embodiments.
[0097] It should be noted that, although in Figure 3 In the flowchart shown, there is a specific order between S310 and S320. However, in the actual execution process, the execution order of the two is not limited. S310 can be executed first and then S320, or S320 can be executed first and then S310, or S310 and S320 can be executed simultaneously.
[0098] In the above embodiments, by quantifying the degree of characteristic differences between different operating conditions and the degree of characteristic fluctuation within the same operating condition, and using the ratio of the two as the index reference coefficient, the index reference coefficient can objectively measure the ability of each psychoacoustic index to distinguish the noise generated by the equipment under different operating conditions, thereby making the selected target acoustic index typical, so as to ensure the accuracy and reliability of the subsequent sound quality evaluation results.
[0099] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the process of merging auditory critical frequency bands that meet the merging conditions according to the psychoacoustic characteristics of each noise sample signal for the same auditory critical frequency band under different target acoustic indicators is refined.
[0100] See Figure 4 The illustrated auditory critical band combining steps include:
[0101] S410, determine the normalized features of the psychoacoustic features of each noise sample signal under different auditory critical frequency bands and under each target acoustic index.
[0102] In one alternative implementation, the psychoacoustic features corresponding to each noise sample signal under different auditory critical frequency bands and under each target acoustic index can be input into the normalization model to obtain the corresponding normalized features.
[0103] In one alternative implementation, for the same target acoustic index under any auditory critical band, the mean and standard deviation of the corresponding features can be determined based on the psychoacoustic characteristics of each noise sample signal under the auditory critical band and the target acoustic index; based on the mean and standard deviation of the features, the psychoacoustic features under the auditory critical band and the target acoustic index are normalized to obtain the normalized features of the corresponding noise sample signal.
[0104] For example, for any target acoustic index, the mean value of the corresponding feature can be determined based on the following formula:
[0105] ;
[0106] In the formula, represents the feature mean of target acoustic index i in the Bark domain b; N represents the total number of samples; This represents the psychoacoustic characteristics of the nth noise sample signal in the bth Bark domain under the target acoustic index i.
[0107] For example, for any target acoustic index, the corresponding feature standard deviation can be determined based on the following formula:
[0108] ;
[0109] In the formula, The standard deviation of the target acoustic index i in the Bark domain b represents the characteristic of the target acoustic index i; N represents the total number of samples. This represents the psychoacoustic characteristics of the nth noise sample signal in the bth Bark domain under the target acoustic index i. Let represent the characteristic mean of the target acoustic index i in the Bark domain b.
[0110] To facilitate understanding, the following example, which includes loudness as a target acoustic index, will be used to illustrate the feature normalization process. This example contains three sets of noise samples and three auditory critical frequency bands: Bark1, Bark2, and Bark3. The mean-standard deviation normalization method is used for calculation. First, the mean and standard deviation of the loudness features of all samples under each auditory critical frequency band are calculated, and then substituted into the normalization formula to obtain the corresponding normalized features.
[0111] Sample number Bark1 Bark2 Bark3 Sample 1 6.2 5.5 4.1 Sample 2 7.0 5.9 4.5 Sample 3 6.6 5.7 4.3
[0112] Based on the data in the table above, calculations show that the mean loudness of the Bark1 band is 6.6 with a standard deviation of 0.4, and the normalized features of the three sample groups are -1.0, 1.0, and 0, respectively. The mean loudness of the Bark2 band is 5.7 with a standard deviation of 0.2, and the mean loudness of the Bark3 band is 4.3 with a standard deviation of 0.2. All samples were calculated using the same rules, resulting in the normalized features of the loudness indices for all samples at different auditory threshold frequency bands. The corresponding normalized features are shown in the table below:
[0113] Sample number Bark1 Bark2 Bark3 Sample 1 -1.0 -1.0 -1.0 Sample 2 1.0 1.0 1.0 Sample 3 0 0 0
[0114] S420, determine the feature similarity of adjacent auditory critical frequency bands under the corresponding target acoustic index based on the degree of deviation between the mean values of the normalized features of adjacent auditory critical frequency bands under the same target acoustic index.
[0115] In one alternative implementation, the mean value of the normalized features of each sample in the same auditory critical band can be determined under the same target acoustic index.
[0116] Referring again to the example above, under the loudness index, the mean of the normalized features of the three auditory critical bands, Bark1, Bark2, and Bark3, is all 0. Correspondingly, feature similarity is quantified based on the difference between the means. For example, if the difference between the normalized feature means of Bark1 and Bark2 is 0, their feature deviation is 0, and the calculated feature similarity is 1; similarly, if the difference between the normalized feature means of Bark2 and Bark3 is 0, their feature deviation is 0, and the corresponding feature similarity is also 1. If the difference between the normalized feature means of two Bark domains is 1, the corresponding feature similarity is less than 1. In this embodiment, no limitations are placed on the mapping method between the mean difference and feature similarity.
[0117] S430 merges adjacent auditory critical frequency bands based on the relationship between the feature similarity and the similarity threshold under each target acoustic index.
[0118] The similarity threshold can be determined based on human experience or through extensive experimentation; this application does not impose any limitations on it. For example, adjacent auditory critical frequency bands can be merged when the feature similarity under each target acoustic index exceeds the similarity threshold.
[0119] For example, when the similarity threshold is 0.8 and the target acoustic index is only loudness, Bark1 and Bark2, and Bark2 and Bark3 can all be merged.
[0120] In the above embodiments, the features of each frequency band and each acoustic index are first normalized to unify the data scale. Then, the similarity is calculated using the difference in the mean of the normalized features of adjacent frequency bands, and frequency band merging is completed by combining it with a preset threshold. This can effectively merge frequency bands with similar acoustic characteristics, reducing the total amount of data and simplifying the subsequent analysis process while retaining the key features of noise.
[0121] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the process of merging auditory critical frequency bands that meet the merging conditions according to the psychoacoustic characteristics of each noise sample signal under different target acoustic indicators for the same auditory critical frequency band is described in detail.
[0122] See Figure 5 The illustrated auditory critical band combining steps include:
[0123] S510, for any target acoustic index, determines the characteristic deviation between the psychoacoustic characteristics of each noise sample signal under different auditory critical frequency bands and under the target acoustic index and the corresponding characteristic mean.
[0124] S520: Based on the characteristic deviations of each noise sample signal under adjacent auditory critical frequency bands and under the target acoustic index, determine the correlation coefficient of the target acoustic index between adjacent auditory critical frequency bands.
[0125] For example, the correlation coefficient of the target acoustic index between adjacent auditory critical frequency bands can be determined based on the following formula:
[0126] ;
[0127] In the formula, Let represent the correlation coefficient between the b-th and b+1-th auditory critical frequency bands under the i-th acoustic index; N is the total number of noise samples. Let be the psychoacoustic feature value of the nth noise sample under the bth auditory critical frequency band and the ith target acoustic index; Let be the mean psychoacoustic characteristics of all noise samples under the b-th auditory critical frequency band and the i-th target acoustic index; This refers to the characteristic deviation of the nth sample obtained in the aforementioned steps under the corresponding frequency band and index.
[0128] S530 merges adjacent auditory critical frequency bands based on the relationship between the correlation coefficient and the correlation threshold of each target acoustic index.
[0129] For example, adjacent auditory critical frequency bands can be merged when the correlation coefficients of all target acoustic indicators exceed the correlation threshold. The correlation threshold can be determined based on human experience or through extensive experimentation; this application does not impose any limitations on it.
[0130] Although Figure 5 In the flowchart shown, there is a specific order between S510 and S520. However, in the actual execution process, the execution order of the two is not limited. S510 can be executed first and then S520, or S520 can be executed first and then S510, or S510 and S520 can be executed simultaneously.
[0131] The above embodiments provide another method for frequency band merging. By calculating the deviation of sample features relative to the frequency band mean, and solving the correlation coefficient of the index based on the deviation data of adjacent frequency bands, it can objectively reflect the strength of the correlation between different auditory critical frequency bands on the same acoustic index. By merging frequency bands in this way, adjacent frequency bands with similar acoustic characteristics can be effectively merged, reducing redundant data dimensions, reducing the computational cost of subsequent sound quality analysis, and at the same time completely preserving the core auditory characteristics of the noise signal.
[0132] Optionally, to improve merging accuracy, in some embodiments, adjacent auditory critical frequency bands can be merged when the correlation coefficients of the indicators corresponding to each target acoustic indicator exceed the correlation threshold and the feature similarity under each target acoustic indicator exceeds the similarity threshold.
[0133] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the method for updating the above-mentioned related thresholds is described.
[0134] See Figure 6 The relevant threshold update steps shown include:
[0135] S610, based on the feature mean and feature standard deviation, normalizes the verification sample signals under the corresponding auditory critical frequency band and the corresponding target acoustic index, and merges the auditory critical frequency bands corresponding to the verification sample signals using the merging rules determined based on the training sample signals.
[0136] In one optional implementation, the collected noise sample signals can be divided into training sample signals and validation sample signals. The training sample signals are used to statistically obtain the feature mean and feature standard deviation corresponding to each auditory critical frequency band and each target acoustic index, and combined with the sample feature deviation and the correlation coefficient of adjacent frequency bands, to summarize the frequency band merging judgment rules that can be directly reused.
[0137] During the verification phase, the psychoacoustic features of the verification sample signals can be read, and the mean and standard deviation obtained during the training phase can be directly used to normalize the features of the verification samples. Subsequently, the frequency band merging rules obtained during training are used to determine the correlation between adjacent frequency bands for each auditory critical frequency band corresponding to the verification sample, and frequency band merging is completed to obtain the composite frequency band result corresponding to the verification sample, so that the frequency band determination criteria of the training process and the verification process are consistent.
[0138] S620: For any composite frequency band obtained after merging, determine the average correlation coefficient between the composite frequency band and each auditory critical frequency band contained in the composite frequency band under each target acoustic index, and update the correlation threshold according to the relationship between the average correlation coefficient and the correlation threshold.
[0139] It should be noted that the method for determining the correlation coefficient of the indicators has been described in the above embodiments and will not be repeated here. The average correlation coefficient is the average of the correlation coefficients of the indicators under each target acoustic indicator.
[0140] For example, the relationship between the average correlation coefficient and the correlation threshold can be compared, and the correlation threshold can be updated based on the following:
[0141] If the average correlation coefficient is significantly higher than the correlation threshold, it indicates that the current correlation threshold is too lenient and the merging accuracy is insufficient; the correlation threshold can be appropriately increased. If the average correlation coefficient is lower than the correlation threshold, it indicates that the threshold is too strict and may prevent the merging of reasonable frequency bands; the correlation threshold can be appropriately decreased. The comparison and adjustment of all composite frequency bands are completed in this manner, and the merging rules are updated based on the adjusted correlation threshold.
[0142] In the above embodiments, the relevant thresholds are updated using the verification sample signal, which can update the merging rules in a timely manner when the merging rules are unreasonable, thereby improving the accuracy of frequency band merging.
[0143] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the method for updating the above-mentioned reference threshold is described.
[0144] See Figure 7 The reference threshold update steps shown include:
[0145] S710, for the same target acoustic index under any auditory critical frequency band, determines the corresponding feature mean and feature standard deviation based on the psychoacoustic characteristics of each noise sample signal under the auditory critical frequency band and the target acoustic index.
[0146] It should be noted that the process of determining the feature mean and feature standard deviation has been described in the above embodiments and will not be repeated here.
[0147] S720 normalizes the verification sample signals under the corresponding auditory critical frequency band and the corresponding target acoustic index according to the characteristic mean and characteristic standard deviation, and determines the characteristic difference and characteristic fluctuation of the verification sample signals under different operating conditions under each target acoustic index based on the normalization result.
[0148] It should be noted that the process of normalizing the validation sample signal and determining the feature density and feature fluctuation can be found in the process of processing the training sample signal, and will not be repeated here.
[0149] S730 determines the verification reference coefficients corresponding to the relevant psychoacoustic indicators based on the degree of feature difference and the degree of feature fluctuation.
[0150] Similarly, the ratio of the degree of feature difference to the degree of feature fluctuation is used as the verification reference coefficient for the corresponding psychoacoustic index.
[0151] S740, update the reference threshold based on the relationship between the verification reference coefficient and the indicator reference coefficient.
[0152] For example, the verification reference coefficient can be compared with the indicator reference coefficient to determine the reference threshold, and the reference threshold can be updated in combination with a preset ratio standard:
[0153] If the verification reference coefficient is greater than or equal to 80% of the reference coefficient, it means that the current model's discrimination ability remains stable and there is no need to adjust the reference threshold. If the verification reference coefficient is less than 80% of the reference coefficient, it means that the model's discrimination ability fluctuates and the judgment effect is unstable. In this case, the reference threshold should be adjusted and the target acoustic index should be reselected based on the adjusted reference threshold.
[0154] In the above embodiments, updating the reference threshold using the verification sample signal can promptly update the target acoustic index when the target acoustic index is not selected reasonably, thereby improving the selection accuracy of the target acoustic index.
[0155] Based on the same inventive concept, some embodiments also provide a sound quality evaluation device for implementing the sound quality evaluation method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more sound quality evaluation device embodiments provided below can be found in the limitations of the sound quality evaluation method described above, and will not be repeated here.
[0156] In one exemplary embodiment, such as Figure 8 As shown, a sound quality evaluation device is provided, including: a feature extraction module 810, an index selection module 820, a frequency band merging module 830, and a sound quality evaluation module 840. Among them,
[0157] The feature extraction module 810 is used to acquire noise sample signals of the device under test under different operating conditions and extract the psychoacoustic features of the noise sample signals under different candidate frequency bands; the different candidate frequency bands include the full frequency band and each auditory critical frequency band.
[0158] The index selection module 820 is used to select the target acoustic index from the psychoacoustic indices based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands.
[0159] The frequency band merging module 830 is used to merge auditory critical frequency bands that meet the merging conditions based on the psychoacoustic characteristics of each noise sample signal for the same auditory critical frequency band under different target acoustic indicators.
[0160] The sound quality evaluation module 840 is used to evaluate the sound quality of the equipment noise of the device under test based on the target acoustic characteristics of each noise sample signal in the full frequency band and the combined auditory critical frequency band; the target acoustic characteristics are determined based on the psychoacoustic characteristics under the corresponding target acoustic indicators.
[0161] In one embodiment, the index selection module 820 includes a first determining unit for determining the degree of characteristic difference of noise sample signals under different operating conditions under any psychoacoustic index; a second determining unit for determining the degree of characteristic fluctuation of different noise sample signals under the same operating condition under the psychoacoustic index; a third determining unit for determining the index reference coefficient of the psychoacoustic index based on the ratio between the degree of characteristic difference and the degree of characteristic fluctuation; and an index selection unit for selecting psychoacoustic indices whose index reference coefficients exceed a reference threshold as target acoustic indices.
[0162] In one embodiment, the first determining unit is specifically configured to determine the global mean of the psychoacoustic characteristics of all noise sample signals under psychoacoustic indices; determine the local mean of the psychoacoustic characteristics of noise sample signals under different operating conditions under psychoacoustic indices; and determine the degree of characteristic difference of noise sample signals under different operating conditions under psychoacoustic indices based on the number of noise sample signals under each operating condition and the mean deviation between the local mean and the global mean under the corresponding operating conditions.
[0163] In one embodiment, the second determining unit is specifically used to determine, for any operating condition, the characteristic deviation between the psychoacoustic characteristics of each noise sample signal under the psychoacoustic index and the corresponding local mean under the operating condition; and to determine the characteristic fluctuation degree of different noise sample signals under the psychoacoustic index under the operating condition based on the characteristic deviation, the total number of all noise sample signals, and the total number of operating condition categories.
[0164] In one embodiment, the band merging module 830 includes a normalization unit for determining the normalized features of the psychoacoustic features of each noise sample signal under different auditory critical bands and under each target acoustic index; a fourth determination unit for determining the feature similarity of adjacent auditory critical bands under the corresponding target acoustic index based on the degree of deviation between the mean of the normalized features of adjacent auditory critical bands under the same target acoustic index; and a first merging unit for merging adjacent auditory critical bands based on the relationship between the feature similarity and the similarity threshold under each target acoustic index.
[0165] In one embodiment, the normalization unit is specifically used to determine the mean and standard deviation of the corresponding features based on the psychoacoustic characteristics of each noise sample signal under the same target acoustic index in any auditory critical band; and to normalize each psychoacoustic feature under the auditory critical band and target acoustic index based on the mean and standard deviation of the features to obtain the normalized features of the corresponding noise sample signal.
[0166] In one embodiment, the band combining module 830 includes a fifth determining unit, configured to determine the feature deviation between the psychoacoustic characteristics and the corresponding feature mean of each noise sample signal under different auditory critical bands and under the target acoustic index for any target acoustic index; a sixth determining unit, configured to determine the index correlation coefficient between adjacent auditory critical bands based on the feature deviation of each noise sample signal under adjacent auditory critical bands and under the target acoustic index; and a second combining unit, configured to combine adjacent auditory critical bands based on the magnitude relationship between the index correlation coefficient and the correlation threshold corresponding to each target acoustic index.
[0167] In one embodiment, the noise sample signal is the training sample signal; the sound quality evaluation device further includes a similarity threshold update module, which is used to normalize the verification sample signals under the corresponding auditory critical frequency band and the corresponding target acoustic index according to the feature mean and feature standard deviation, and to merge the auditory critical frequency bands corresponding to the verification sample signals according to the merging rules determined according to the training sample signals; for any composite frequency band obtained after merging, the average correlation coefficient between the composite frequency band and each auditory critical frequency band contained in the composite frequency band under each target acoustic index is determined, and the correlation threshold is updated according to the relationship between the average correlation coefficient and the correlation threshold.
[0168] In one embodiment, the noise sample signal is the training sample signal; the sound quality evaluation device further includes a reference threshold update module, used to determine the corresponding feature mean and feature standard deviation for the same target acoustic index under any auditory critical frequency band, based on the psychoacoustic characteristics of each noise sample signal under the auditory critical frequency band and the target acoustic index; normalize the verification sample signal under the corresponding auditory critical frequency band and the corresponding target acoustic index based on the feature mean and feature standard deviation, and determine the feature difference degree and feature fluctuation degree of the verification sample signal under different operating conditions under each target acoustic index based on the feature difference degree and feature fluctuation degree; determine the verification reference coefficient corresponding to the corresponding psychoacoustic index based on the feature difference degree and feature fluctuation degree; and update the reference threshold based on the magnitude relationship between the verification reference coefficient and the index reference coefficient.
[0169] In one exemplary embodiment, a computer device is provided, which may be a terminal device, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores PCB product-related data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a sound quality evaluation method.
[0170] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0171] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0172] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0173] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0174] It should be noted that the user information (including but not limited to user behavior data) and data (including but not limited to data used for analysis, stored data, and displayed data) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0175] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0176] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0177] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for evaluating sound quality, characterized in that, The method includes: Noise sample signals of the device under test under different operating conditions are acquired, and the psychoacoustic features of the noise sample signals are extracted in different candidate frequency bands; the different candidate frequency bands include the full frequency band and each auditory critical frequency band. Based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands, a target acoustic index is selected from the psychoacoustic indexes. Based on the psychoacoustic characteristics of each noise sample signal for the same auditory critical band under different target acoustic indicators, auditory critical bands that meet the merging conditions are merged. The sound quality of the device under test is evaluated based on the target acoustic characteristics of each noise sample signal in the full frequency band and the combined auditory critical frequency band; the target acoustic characteristics are determined based on the psychoacoustic characteristics under the corresponding target acoustic indicators.
2. The method according to claim 1, characterized in that, The step of selecting a target acoustic index from the psychoacoustic indices based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index in different candidate frequency bands includes: For any psychoacoustic index, determine the degree of characteristic difference of noise sample signals under different operating conditions under the psychoacoustic index; determine the degree of characteristic fluctuation of different noise sample signals under the same operating condition under the psychoacoustic index. The reference coefficient of the psychoacoustic index is determined based on the ratio between the degree of difference in the feature and the degree of fluctuation in the feature. Psychoacoustic indices whose reference coefficients exceed the reference threshold are selected as target acoustic indices.
3. The method according to claim 2, characterized in that, The determination of the degree of characteristic difference of noise sample signals under different operating conditions under the psychoacoustic index includes: Determine the global mean of the psychoacoustic characteristics of all noise sample signals under the aforementioned psychoacoustic index; determine the local mean of the psychoacoustic characteristics of noise sample signals under different operating conditions under the aforementioned psychoacoustic index. Based on the number of noise sample signals under each operating condition and the mean deviation between the local mean and the global mean under the corresponding operating condition, the degree of characteristic difference of the noise sample signals under different operating conditions under the psychoacoustic index is determined.
4. The method according to claim 3, characterized in that, The determination of the characteristic fluctuation degree of different noise sample signals under the same operating condition under the psychoacoustic index includes: For any operating condition, determine the characteristic deviation between the psychoacoustic characteristics of each noise sample signal under the psychoacoustic index and the corresponding local mean under the operating condition; Based on the characteristic deviation, the total number of all noise sample signals, and the total number of operating condition categories, the characteristic fluctuation degree of different noise sample signals under the psychoacoustic index is determined.
5. The method according to any one of claims 1-4, characterized in that, The step of merging auditory critical frequency bands that meet the merging conditions based on the psychoacoustic characteristics of each noise sample signal under different target acoustic indices for the same auditory critical frequency band includes: Determine the normalized features of the psychoacoustic features of each noise sample signal under different auditory critical frequency bands and under each target acoustic index; The feature similarity of adjacent auditory critical frequency bands under the corresponding target acoustic index is determined based on the degree of deviation between the mean values of the normalized features of adjacent auditory critical frequency bands under the same target acoustic index. Based on the relationship between the feature similarity and the similarity threshold under each target acoustic index, adjacent auditory critical frequency bands are merged.
6. The method according to claim 5, characterized in that, The determination of the normalized features of the psychoacoustic features corresponding to each of the noise sample signals under different auditory critical frequency bands and under each target acoustic index includes: For the same target acoustic index under any auditory critical frequency band, the mean and standard deviation of the corresponding features are determined based on the psychoacoustic characteristics of each noise sample signal under the auditory critical frequency band and the target acoustic index. Based on the feature mean and the feature standard deviation, the psychoacoustic features of the auditory critical frequency band and the target acoustic index are normalized to obtain the normalized features of the corresponding noise sample signal.
7. The method according to claim 5, characterized in that, The step of merging auditory critical frequency bands that meet the merging conditions based on the psychoacoustic characteristics of each noise sample signal under different target acoustic indices for the same auditory critical frequency band includes: For any target acoustic index, determine the characteristic deviation between the psychoacoustic characteristics and the corresponding characteristic mean of each noise sample signal under different auditory critical frequency bands and under the target acoustic index; based on the characteristic deviation of each noise sample signal under adjacent auditory critical frequency bands and under the target acoustic index, determine the index correlation coefficient of the target acoustic index between adjacent auditory critical frequency bands. Based on the relationship between the correlation coefficient and the correlation threshold of each target acoustic index, adjacent auditory critical frequency bands are merged.
8. The method according to claim 7, characterized in that, The noise sample signal is the training sample signal; the update method of the relevant threshold includes: Based on the feature mean and feature standard deviation, the verification sample signals under the corresponding auditory critical frequency band and the corresponding target acoustic index are normalized, and the auditory critical frequency bands corresponding to the verification sample signals are merged using the merging rules determined based on the training sample signals. For any composite frequency band obtained after merging, determine the average correlation coefficient between the composite frequency band and each auditory critical frequency band contained in the composite frequency band under each target acoustic index, and update the correlation threshold according to the relationship between each average correlation coefficient and the correlation threshold.
9. The method according to claim 2, characterized in that, The noise sample signal is the training sample signal; the update method of the reference threshold includes: For the same target acoustic index under any auditory critical frequency band, the mean and standard deviation of the corresponding features are determined based on the psychoacoustic characteristics of each noise sample signal under the auditory critical frequency band and the target acoustic index. Based on the mean and standard deviation of the features, the verification sample signals under the corresponding auditory critical frequency band and the corresponding target acoustic index are normalized. Based on the normalization results, the degree of feature difference and feature fluctuation of the verification sample signals under different operating conditions under each target acoustic index are determined. Based on the degree of difference and fluctuation of the features, the verification reference coefficients corresponding to the psychoacoustic indicators are determined. The reference threshold is updated based on the relationship between the verification reference coefficient and the indicator reference coefficient.
10. A terminal device, characterized in that, The device includes: At least one processor, and configured as follows: Noise sample signals of the device under test under different operating conditions are acquired, and the psychoacoustic features of the noise sample signals are extracted in different candidate frequency bands; the different candidate frequency bands include the full frequency band and each auditory critical frequency band. Based on the psychoacoustic characteristics of each noise sample signal for the same psychoacoustic index under different candidate frequency bands, a target acoustic index is selected from the psychoacoustic indexes. Based on the psychoacoustic characteristics of each noise sample signal for the same auditory critical band under different target acoustic indicators, auditory critical bands that meet the merging conditions are merged. The sound quality of the device under test is evaluated based on the target acoustic characteristics of each noise sample signal in the full frequency band and the combined auditory critical frequency band; the target acoustic characteristics are determined based on the psychoacoustic characteristics under the corresponding target acoustic indicators.