A method for evaluating environmental noise based on annoyance perception index

Through XGBoost model training and sound source annoyance weighting, the problem that traditional noise evaluation methods fail to reflect the differences in annoyance perception is solved, and an accurate environmental noise evaluation method is provided, which is suitable for the assessment of the sound environment quality in ecological cities.

CN119559968BActive Publication Date: 2025-10-03HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411633353.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-10-03
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

Traditional environmental noise evaluation methods fail to effectively reflect the differences in people's perception of annoyance caused by environmental noise with different semantics, and are unable to meet the needs of ecological cities.

Method used

The XGBoost model is used to train the environmental noise annoyance evaluation value prediction model. By obtaining the annoyance evaluation value of the environmental noise sample, combining it with the age of the target population, converting it into the sound source annoyance, and weighting it with the sound level annoyance, the final environmental noise annoyance is obtained.

Benefits of technology

It achieves a more accurate description of the annoying effects of environmental noise on people and provides a more precise method for evaluating the quality of ecological urban sound environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559968B_ABST
    Figure CN119559968B_ABST
Patent Text Reader

Abstract

The present invention provides an environmental noise evaluation method based on an annoyance perception index, comprising: determining an annoyance evaluation value of an acquired environmental noise sample; training an XGBoost model using the environmental noise sample and the annoyance evaluation value as training samples to obtain a trained model; obtaining an original audio signal to be tested and inputting the original audio signal into the trained model to obtain an annoyance evaluation value of the environmental noise; converting the environmental noise annoyance evaluation value into the annoyance of the sound source of the environmental noise based on the age of the target population; and finally, the present invention weights the annoyance caused by the sound source and the annoyance caused by the sound level in the environmental noise, and calculates the environmental noise annoyance in decibels. Using the human annoyance level as an indicator can more accurately describe the impact of environmental noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of environmental noise evaluation, and in particular to an environmental noise evaluation method based on annoyance perception index. Background Art

[0002] As cities move toward ecological goals, environmental noise is increasingly encompassing natural and human-induced sounds, sources that people readily perceive. Environmental noise has long since evolved beyond the noise of industrialized cities. While traffic, machinery, and the clamor of urban life remain the primary environmental noise sources of ecological cities, natural and human-induced sounds cannot be ignored. Noise assessment techniques that simply treat all environmental noise sources as pollution sources no longer adequately reflect the distressing effects of environmental noise. Currently, the persistently high number of noise complaints and the declining noise monitoring data indicate that traditional environmental noise assessment methods are no longer sufficient to meet the needs of social development.

[0003] The World Health Organization's 21st-century acoustic environment quality guidelines use "annoyance" as a key indicator to measure community noise environments. Soundscape research indicates that the quality of a sound environment depends on how people perceive environmental noise. Annoyance is a key indicator for evaluating sound environment quality and reflects the connotation of noise. Therefore, the fundamental criterion for evaluating environmental noise is the perceived annoyance caused by environmental noise. Unlike traditional noise evaluation, which uses sound level as the sole factor, soundscape considers all factors that cause annoyance. The source of ambient noise has been shown to play a key role, particularly when the ambient noise level is in the 50dB–70dB range. Numerous studies have shown that the perceived annoyance of ambient noise with different meanings varies significantly. When evaluating ambient noise, the traditional method of using equivalent sound level as the primary metric is no longer sufficient to meet the needs of modern ecological cities.

[0004] To distinguish between the natural and social sounds that people enjoy, and the urban sounds of traffic, production, and daily life that they dislike, evaluating noise based on people's perceived annoyance is key to measuring the acoustic quality of an eco-city. To address this need, developing more accurate environmental noise assessment methods tailored to ecological and environmental conditions is crucial. Summary of the Invention

[0005] In order to overcome the deficiencies of the prior art, the present invention aims to provide an environmental noise evaluation method based on annoyance perception index, which can more accurately describe the annoyance level of environmental noise.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] An environmental noise evaluation method based on annoyance perception index, comprising:

[0008] Determine an annoyance evaluation value of the sample according to the acquired environmental noise sample;

[0009] The environmental noise samples and the annoyance evaluation values ​​are used as training samples to train the XGBoost model to obtain a trained environmental noise annoyance evaluation value prediction model;

[0010] Acquire an original audio signal to be tested, and input the original audio signal into the environmental noise annoyance evaluation value prediction model to obtain an annoyance evaluation value of the environmental noise;

[0011] converting the annoyance evaluation value of the environmental noise into the sound source annoyance of the environmental noise according to the age of the target population;

[0012] The sound source annoyance degree and the sound level annoyance degree corresponding to the original audio signal are weighted to obtain a final environmental noise annoyance degree.

[0013] Preferably, determining the annoyance evaluation value of the sample according to the acquired environmental noise sample includes:

[0014] A calibrated binaural measurement system is used to obtain environmental noise samples;

[0015] The 10 seconds with the most obvious sound source characteristics in the environmental noise sample are intercepted as a research segment, and a subjective evaluation experiment on the annoyance of environmental noise is carried out to obtain the annoyance evaluation value of the sample.

[0016] Preferably, the environmental noise samples and the annoyance evaluation values ​​are used as training samples to train the XGBoost model to obtain a trained environmental noise annoyance evaluation value prediction model, including:

[0017] Performing audio preprocessing on the environmental noise sample to obtain a preprocessed audio signal;

[0018] extracting a spectrogram reflecting perceptual characteristics from the preprocessed audio signal;

[0019] The spectrum graph that can reflect the perception characteristics and the annoyance evaluation value are input into the XGBoost model for training to obtain a trained environmental noise annoyance evaluation value prediction model.

[0020] Preferably, extracting a spectrogram reflecting perceptual characteristics from the preprocessed audio signal comprises:

[0021] Extracting the power spectrum of the preprocessed audio signal at different time segments;

[0022] The power spectra of different time segments are processed to obtain a Fourier transform matrix, and the Fourier transform matrix is ​​determined as the frequency spectrum that can reflect the perception characteristics.

[0023] Preferably, converting the annoyance evaluation value of the environmental noise into the sound source annoyance of the environmental noise according to the age of the target population includes:

[0024] If the target population is in the adult range, the sound source annoyance level is determined using the formula y = -1.190x + 4.183, where x is the annoyance rating of the ambient noise and y is the value of the sound source annoyance level.

[0025] If the target population is in the elderly range, the sound source annoyance degree is determined using the formula y=-1.000x+4.253;

[0026] If the age of the target population is in the teenage range, the sound source annoyance degree is determined using the formula y=-1.496x+9.622.

[0027] Preferably, the obtained sound source annoyance is weighted with the sound level annoyance to obtain the environmental noise annoyance, wherein the sound level annoyance is obtained by measuring with an acoustic measuring device.

[0028] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0029] The present invention provides an environmental noise evaluation method based on an annoyance perception index, comprising: determining an annoyance evaluation value of an acquired environmental noise sample; training an XGBoost model using the environmental noise sample and the annoyance evaluation value as training samples to obtain a trained environmental noise annoyance evaluation value prediction model; obtaining an original audio signal to be measured and inputting the original audio signal into the environmental noise annoyance evaluation value prediction model to obtain an environmental noise annoyance evaluation value; converting the environmental noise annoyance evaluation value into an environmental noise source annoyance value based on the age of a target population; and weighting the sound source annoyance value and the sound level annoyance value corresponding to the original audio signal to obtain a final environmental noise level. The present invention uses the weighted "annoyance value" generated by the sound source in the environmental noise and the "annoyance value" generated by the sound level to form the environmental noise annoyance value, which can more accurately describe a person's annoyance perception of environmental noise. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0031] Figure 1A flow chart of a method provided by an embodiment of the present invention;

[0032] Figure 2 Schematic diagram of audio extraction using short-time Fourier transform provided in an embodiment of the present invention;

[0033] Figure 3 A schematic diagram of the linear relationship between the sound source annoyance degree and the annoyance evaluation value provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0035] The purpose of the present invention is to provide an environmental noise evaluation method based on annoyance perception index, which can more accurately describe the noisiness level of environmental noise.

[0036] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] Figure 1 A flow chart of the method provided in the embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides an environmental noise evaluation method based on annoyance perception index, comprising:

[0038] Step 100: determining an annoyance evaluation value of the sample according to the acquired environmental noise sample;

[0039] Step 200: Using the environmental noise sample and the annoyance evaluation value as training samples to train the XGBoost model, thereby obtaining a trained environmental noise annoyance evaluation value prediction model;

[0040] Step 300: obtaining an original audio signal to be tested, and inputting the original audio signal into the environmental noise annoyance evaluation value prediction model to obtain an annoyance evaluation value of the environmental noise;

[0041] Step 400: converting the annoyance evaluation value of the environmental noise into the sound source annoyance of the environmental noise according to the age of the target population;

[0042] Step 500: performing weighting according to the sound source annoyance degree and the ambient noise level corresponding to the original audio signal to obtain a final ambient noise level.

[0043] Specifically, process 1 of this embodiment first establishes a prediction model for the noise annoyance rating (NPS) of ambient noise. Specifically, the XGBoost "extreme gradient boosting" model is used to obtain the noise source annoyance rating (NPS) based on subjective evaluation experiments. The following is the complete solution:

[0044] 1. Use a calibrated binaural measurement system (e.g., an artificial head) to sample ambient noise samples. Select the 10-second segment with the most distinct sound source characteristics from the sample and conduct a subjective evaluation experiment on the ambient noise source noise level (Note: This experimental method is based on the method in "GBT42473-2023: Evaluation and Prediction Method for Acoustic Noise Annoyance") to obtain the subjects' annoyance ratings. These ambient noise samples and the subjects' ratings will serve as training samples to build the XGBoost model.

[0045] 2. Use the original audio signal as input variable;

[0046] 3. The trained XGBoost model will predict the annoyance evaluation value of the ambient noise corresponding to the audio signal based on the input original audio signal (that is, the annoyance perception value of the sound source semantics).

[0047] Specifically, the XGBoost model consists of three parts: audio preprocessing, extracting spectrograms that reflect perceptual features, and machine learning regression:

[0048] (1) Audio preprocessing. Audio preprocessing is divided into three processes: sampling, encoding, and normalization. The original audio signal is a time-continuous voltage signal. The original audio signal needs to be converted into a discrete digital signal through sampling before it can be processed by the computer. This process is calculated using the following formula:

[0049] s d (n) = s c (n / f)

[0050] Where f is the sampling frequency, which is set to 44100 Hz here; n is the serial number of the sampled digital signal; s(n) is the sampled digital signal, which is further encoded into a 16-bit binary number by the computer with a sampling accuracy of 1 / 2 15 Taking into account human auditory perception, the project set the sampling accuracy to 10 seconds. Therefore, the model preprocessed audio for 10 seconds. After sampling, a total of 441,000 frames of signal were generated, and the encoded data size was 882kB. To align with the subjective evaluation experimental conditions and eliminate the influence of sound level, the equivalent sound level of the 10 seconds of preprocessed audio was normalized to 60dB.

[0051] (2) Extracting a spectrum that reflects perceptual features as input for machine learning. The frequency envelope information in the spectrum is an important source of information for people to perceive different sound sources. It basically reflects the amplitude, frequency domain, and time domain characteristics of the sound. Therefore, this technology uses the spectrum as a prerequisite for judging the noise level of environmental noise sources. The model extracts spectral features from the preprocessed audio signal as input for subsequent machine learning regression. The extraction process includes: short-time Fourier transform (STFT), calculation of power spectrum, and calculation of cepstral coefficients.

[0052] Figure 2 This figure shows the power spectrum of audio at different time segments extracted using the short-time Fourier transform. The figure uses the time dimension as a benchmark, cutting out the audio digital signal at different time segments in the form of a sliding window.

[0053] The window width is 2048 frames, about 46.4ms; the sliding step is 512 frames, about 11.6ms. The cutting process is shown in the following formula.

[0054] s(n,m)=s d (n+m(N s -1))

[0055] Among them, m is the serial number of the time window, n is the serial number of the frame in the time window, N s The sliding step is 512. The audio signal on each time segment needs to be processed by the window function to compensate for the spectrum leakage caused by not meeting the signal stationary assumption. w (n,m) performs discrete Fourier transform, as shown below

[0056]

[0057] Where k = 0, 1, 2, ..., N-1 represents the frequency, and N is the number of signal frames within a time window. The modulus of S(k,m) is the energy of the signal at different frequencies within the time window, i.e., the power spectrum. The argument of S(k,m) is the phase of the signal at different frequencies within the time window.

[0058] In order to simulate the difference in sensitivity of the human ear to sounds of different frequencies, the project uses a Mel filter bank to process the audio power spectrum of different time periods to obtain the Fourier transform matrix S at the Mel scale frequency. mel (k,m). The cepstral coefficient summarizes the frequency envelope of the Mel power spectrum curve over a period of time, as shown in the following formula, and is generated by the discrete cosine transform of the Fourier transform result matrix.

[0059]

[0060] Where n=0,1,2,...12,Cn is the cosine coefficient, Otherwise, it is 1. After obtaining the Mel-frequency cepstral coefficient matrix, the matrix is ​​downsampled in the time series direction to obtain its six-dimensional parameters: mean, first-order difference mean, first-order difference variance, second-order difference mean, skewness, and kurtosis, which is a 13×6 Mel-frequency cepstral coefficient matrix.

[0061] (3) Machine learning regression acquisition. Machine learning regression uses the extracted Mel-cepstrum coefficient matrix as input variables and uses the XGBoost model to integrate multiple decision trees after regression learning to establish a logical connection between the Mel-cepstrum coefficient matrix features and the annoyance rating value (NPS) of the environmental noise source. For a single decision tree, its regression is a supervised learning algorithm used to establish a prediction model. The basic principle is to gradually split the data set through a series of binary judgments, divide the data into multiple subsets, and perform regression prediction on each subset. Decision tree regression recursively divides the data set into smaller subsets and fits a linear regression model on each subset, and finally combines the linear regression models of all subsets into an overall model.

[0062] A total of 363 audio samples with NPS labels were used to build the model. The accuracy calculation formula is as follows:

[0063]

[0064] Among them, N correct The sum of the number of samples whose absolute error is less than σ=1 (σ is the variance of the NPS of the subjective evaluation test subjects), that is, the number of samples with accurate prediction. total It represents the total number of samples. After training and testing the samples, the model achieved a mean absolute error (MAE) of 0.57, a root mean square error (RMSE) of 0.73, and an accuracy of 84% at ±σ=1.

[0065] Specifically, XGBoost, short for eXtreme Gradient Boosting, is an ensemble learning algorithm based on Gradient Boosting Decision Trees (GBDT). It improves prediction performance by using multiple decision trees trained through regression. Here's how XGBoost works:

[0066] 1. Initialize the model: First, build an initial decision tree model and fit the model using the training data.

[0067] 2. Calculate the residual: For each sample in the training dataset, calculate the error between its true value and the model's predicted value to obtain the sample's residual. The residual refers to the deviation in the model's prediction for the sample in the current round.

[0068] 3. Build a new decision tree: In each iteration, the residuals are used to build a new decision tree model to fit the residual values. The new decision tree model is built on the basis of the previous model, and the residuals are learned to further improve the model's predictive ability.

[0069] 4. Update model prediction: Integrate the newly constructed decision tree model with the previous model to obtain an updated model. The final prediction value is obtained by adding the prediction results of each model.

[0070] 5. Adjust the learning rate: To improve model accuracy, XGBoost also introduces the concept of learning rate. The learning rate controls the contribution of each new model in the ensemble. By adaptively adjusting the learning rate, you can further improve model performance.

[0071] 6. Iterative training: Repeat the above steps for multiple rounds of iterative training. Each round builds a new decision tree model to fit the residuals from the previous round, and then integrates the new model with the previous one. Through this iterative process, the model gradually improves and predictive performance increases.

[0072] Furthermore, process 2 of this embodiment is to convert the noise perception value of the environmental noise source obtained in process 1 from a psychological description score into an acoustic physical quantity with the same unit as the sound pressure level index commonly used in current environmental noise evaluation, namely the noise source annoyance level of environmental noise (abbreviated as "NS") through a structural model. This conversion process aims to combine the subjective evaluation results with the objective physical measurement standard (sound pressure level expressed in decibels) to provide a more comprehensive, objective and quantifiable environmental noise evaluation index. Since different people have different annoyance perceptions, the specific structural model is as follows: Figure 3 shown.

[0073] Considering that the main users of a large number of places in the city are adults, the formula is as follows:

[0074] y=-1.190x+4.183

[0075] Where x is the noise source annoyance rating (NPS), and y is the noise source annoyance rating (NS). When the user population is elderly, the following formula can be used:

[0076] y=-1.000x+4.253

[0077] When the applicable population of the venue is teenagers, the following formula can be used:

[0078] y=-1.496x+9.622

[0079] Optionally, process 3 of this embodiment weights the "annoyance" generated by the sound source semantics and the "annoyance" generated by the sound level to form the environmental noise annoyance level (NL), providing an accurate method for environmental noise evaluation in ecological environments. SL stands for "Source Level" and is the sound level annoyance level. The formula is (in dB):

[0080] NL=SL+NS

[0081] Generally speaking, ambient noise level measurement is achieved by measuring the perceived annoyance decibel value (i.e., sound source annoyance) and sound level annoyance caused by the semantics of the sound sources in the environment. The sound source annoyance is obtained using the structural model provided in step 2, while the sound level annoyance is the equivalent sound level over a specific time period and can be provided by currently available, qualified acoustic measurement equipment.

[0082] The beneficial effects of the present invention are as follows:

[0083] Compared with the existing evaluation index that only uses the "annoyance level" of sound level to measure environmental noise, the present invention uses the "annoyance level" caused by the sound source in the environmental noise and the "annoyance level" caused by the sound level to weight them as the environmental noise annoyance level, which can more accurately describe the annoyance effect of environmental noise on people.

[0084] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0085] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A method for evaluating environmental noise based on annoyance perception index, characterized in that: include: Determine an annoyance evaluation value of the sample according to the acquired environmental noise sample; The environmental noise samples and the annoyance evaluation values ​​are used as training samples to train the XGBoost model to obtain a trained environmental noise annoyance evaluation value prediction model; Obtaining an original audio signal of the environmental noise to be measured, and inputting the original audio signal into the environmental noise annoyance evaluation value prediction model to obtain an annoyance evaluation value of the environmental noise; converting the annoyance evaluation value of the environmental noise into the sound source annoyance of the environmental noise according to the age of the target population; weighting the sound source annoyance degree and the sound level annoyance degree corresponding to the original audio signal to obtain a final environmental noise annoyance degree; The XGBoost model is trained using the environmental noise sample and the annoyance evaluation value as training samples to obtain a trained environmental noise annoyance evaluation value prediction model, including: Performing audio preprocessing on the environmental noise sample to obtain a preprocessed audio signal; extracting a spectrogram reflecting perceptual characteristics from the preprocessed audio signal; Inputting the spectrum graph that can reflect the perception characteristics and the annoyance evaluation value into the XGBoost model for training to obtain a trained environmental noise annoyance evaluation value prediction model; Converting the annoyance evaluation value of the environmental noise into the sound source annoyance of the environmental noise according to the age of the target population includes: If the target population is in the adult range, the sound source annoyance level is determined using the formula y=-1.190x+4.183; where x is the annoyance rating of the ambient noise, and y is the numerical value of the sound source annoyance level; If the target population is in the elderly range, the sound source annoyance degree is determined using the formula y=-1.000x+4.253; If the target population is in the teenage age range, the sound source annoyance degree is determined using the formula y=-1.496x+9.

622.

2. The environmental noise evaluation method based on annoyance perception index according to claim 1, characterized in that: Determining an annoyance evaluation value of the sample based on the acquired environmental noise sample, including: A calibrated binaural measurement system is used to obtain environmental noise samples; The 10 seconds with the most obvious sound source characteristics in the environmental noise sample are intercepted as a research segment, and a subjective evaluation experiment of environmental noise is carried out to obtain the annoyance evaluation value of the sample.

3. The environmental noise evaluation method based on annoyance perception index according to claim 1, characterized in that: Extracting a spectrogram reflecting perceptual characteristics from the preprocessed audio signal includes: Extracting the power spectrum of the preprocessed audio signal at different time segments using a short-time Fourier transform method; The power spectra of different time segments are processed to obtain a Fourier transform matrix, and the Fourier transform matrix is ​​used to obtain a frequency spectrum that can reflect the perception characteristics.

4. The environmental noise evaluation method based on annoyance perception index according to claim 1, characterized in that: The obtained sound source annoyance degree and the sound level annoyance degree are weighted to obtain the environmental noise annoyance degree, wherein the sound level annoyance degree is obtained by measuring with an acoustic measuring device.

Citation Information

Patent Citations

  • Matching standard sample method for evaluating audio injection noise suppression annoyance effect

    CN115662466A

  • Method of Measuring Annoyance Caused by Noise in an Audio Signal

    US20080267425A1