Method, device and equipment for measuring environmental sound psychoacoustic decibel value and storage medium
Patent Information
- Application Number
- CN202311088115.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-08-28
AI Technical Summary
由此可见,需要引入以描述声压级为评价环境声的参数,在环境噪声监测工作中存在较大的局限性,出现了因自然声太多而超标现象
[0013]This application provides a method, apparatus, computer equipment, and storage medium for measuring the psychological decibel value of ambient sound. It proposes, for the first time, an accurate psychological decibel value reflecting environmental noise pollution. By experimentally obtaining psychological decibel values from audio data under different environments and using them as a dataset for model training, feature extraction is performed, and a solution model is established between Mel-frequency cepstral coefficients and psychological decibel values. The psychological decibel value of environmental noise can be obtained through this solution model. Furthermore, the method for measuring the psychological decibel value of ambient sound used in this application makes the environmental noise measurement results more scientific and accurate.
Smart Images

Figure CN117077087B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of noise detection technology, and in particular to a method, apparatus, equipment and storage medium for measuring the psychodecibel value of ambient sound. Background Technology
[0002] Currently, environmental noise is measured using equivalent sound level as an indicator, describing the "noise intensity" resulting from the magnitude of environmental sound energy. However, research confirms that when the equivalent sound level is between 50-70 dB, the "noise intensity" of environmental sound is affected by both the sound level and the sound source. Under the same equivalent sound level, natural sounds have a higher "noise intensity" than mechanical sounds, even though their decibel values are the same, which clearly does not conform to people's psychological perception of environmental noise intensity. Therefore, introducing a parameter describing sound pressure level to evaluate environmental noise has significant limitations in environmental noise monitoring, leading to instances where natural sounds are too prevalent and exceed standards. In environmental noise monitoring, it is essential to distinguish between ecological and non-ecological sounds and define noise standards based on both sound level and sound source indicators. This approach is more conducive to ecological environment protection and the construction of eco-cities. Summary of the Invention
[0003] The purpose of this application is to provide a method, apparatus, computer equipment, and storage medium for measuring the psychological decibel value of ambient sound, which can make the environmental noise measurement results more scientific and accurate.
[0004] To address the aforementioned technical problems, this application provides a method for measuring the psychological decibel value of ambient sound. The method includes: acquiring audio data from different ambient sound sources; for the audio data from different ambient sound sources, under the same sound pressure level, experimentally labeling the psychological decibel value of the audio data, and extracting the Mel-frequency cepstral coefficients of all the audio data in the audio dataset; establishing a psychological decibel value model between the Mel-frequency cepstral coefficients and the psychological decibel value based on an FC-Boost regression model; and obtaining the psychological decibel value of the environment to be measured based on the psychological decibel value model.
[0005] Specifically, the method of labeling the psychological decibel value of the audio data from different environmental sound sources under the same sound pressure level by means of experiments includes: using pink noise energy as a standard to calibrate the different energy levels of the audio data under human perception, thereby obtaining the psychological decibel value of the audio data.
[0006] The extraction of Mel-frequency cepstral coefficients from all audio data in the audio dataset includes: performing a short-time Fourier transform on the audio data to obtain the power spectrum of the audio data at different time segments; and processing the power spectrum of the audio data at different time segments to obtain the Mel-frequency cepstral coefficients.
[0007] The step of establishing the solution model for the Mel inverse coefficients and the psychological decibel value based on the FC-Boost regression model includes: selecting Mel inverse coefficients that meet preset conditions from all the Mel inverse coefficients to divide the audio dataset into multiple audio data subsets; fitting each audio data subset to obtain the prediction result of the audio data subset; and combining the prediction results of all the audio data subsets to obtain the FC-Boost regression model.
[0008] The step of combining the prediction results of all the audio data subsets to obtain the FC-Boost regression model includes: constructing an initial FC-Boost model and training the initial FC-Boost model with the audio data in the audio dataset; calculating the error between the true value of each sample sound in the audio dataset and the predicted value of the initial FC-Boost model as the residual value of the sample sound; constructing a new FC-Boost sub-model using the residual value of each sample sound to fit the residual value of the next round of model training; and iterating continuously and adding the prediction results of all the FC-Boost sub-models to obtain the final composite FC-Boost model.
[0009] To address the aforementioned technical problems, this application also provides a device for measuring the psychological decibel value of ambient sound. The device includes a microphone, a sound acquisition unit, a central control unit, and a data transmission module. The microphone acquires audio data from different ambient sound sources and transmits the ambient audio data to the sound acquisition unit for audio preprocessing to obtain a digital signal. The digital signal is then transmitted to the central control unit for feature extraction, model training, and model calculation to obtain the psychological decibel value. Finally, the data is transmitted to a backend server via the data transmission module.
[0010] To address the aforementioned technical problems, this application also provides a computer device, including a memory and a processor. The memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the method for measuring the ambient sound psychodecibel value as described in any of the preceding claims.
[0011] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the following technical solution: the computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the aforementioned method for measuring the ambient sound psychodecibel value.
[0012] Compared with the prior art, the embodiments of this application have the following main advantages:
[0013] This application provides a method, apparatus, computer equipment, and storage medium for measuring the psychological decibel value of ambient sound. It proposes, for the first time, an accurate psychological decibel value reflecting environmental noise pollution. By experimentally obtaining psychological decibel values from audio data under different environments and using them as a dataset for model training, feature extraction is performed, and a solution model is established between Mel-frequency cepstral coefficients and psychological decibel values. The psychological decibel value of environmental noise can be obtained through this solution model. Furthermore, the method for measuring the psychological decibel value of ambient sound used in this application makes the environmental noise measurement results more scientific and accurate. Attached Figure Description
[0014] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating one implementation method of the method for measuring the psychological decibel value of environmental sound in this application;
[0016] Figure 2 This is a flowchart illustrating an embodiment of step S300 of this application;
[0017] Figure 3 This is a flowchart illustrating an embodiment of step S400 of this application;
[0018] Figure 4 This is a flowchart illustrating one embodiment of step S430 of this application;
[0019] Figure 5 This is a schematic diagram of one embodiment of the environmental sound psychodecibel value measuring device of this application;
[0020] Figure 6 This is a schematic diagram of one embodiment of the computer device of this application. Detailed Implementation
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The terms “comprising” and “having,” and any variations thereof, in the specification, claims, and foregoing drawings, are intended to cover non-exclusive inclusion. The terms “first,” “second,” etc., in the specification, claims, or foregoing drawings are used to distinguish different objects and not to describe a particular order.
[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0024] It should be noted that this application proposes for the first time a parameter for evaluating the "noise level" of environmental sound sources—the psychological decibel. Defined as a parameter that measures the annoyance caused by environmental sound sources based on people's perception of the "noise level" of these sources. This parameter reflects the noise level value of the environmental sound source, converting the impact of the sound source on people into the equivalent sound pressure level by referencing sound pressure level units. The psychological decibel is a parameter describing the noise level of a sound source. Its initial introduction will fill the gap in the inaccuracy of evaluating environmental noise level solely based on sound pressure level, especially for environmental sounds in the 50-70 decibel range. The sound source plays a crucial role in the perception of environmental noise level, and existing sound pressure level measurements cannot reflect this effect, making this application significant. The initial introduction of this parameter will make environmental noise measurement results more scientific and accurate. The evaluation and measurement of the environmental sound psychological decibel value in this application are described in detail below with reference to the accompanying drawings.
[0025] Please combine Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the method for measuring the psychodecibel value of ambient sound in this application. Figure 1 The method for measuring the psychological decibel value of ambient sound provided in this application includes the following steps:
[0026] S100 acquires audio data from different environmental sound sources.
[0027] Optionally, in this application, audio data of sound sources under different target environments can be collected by recording to obtain the audio dataset required for the experiment.
[0028] S200 uses an experimental method to label the psychological decibel values of audio data from different environmental sound sources under the same sound pressure level, in order to obtain an audio dataset labeled with psychological decibel values.
[0029] Optionally, in this application, pink noise energy is used as a standard to calibrate the different energy levels of the audio data under human perception, thereby obtaining the psychological decibel value of the audio data.
[0030] In a specific application scenario of this application, 11 urban environmental sound sources under different target environments were recorded as experimental sample sounds, and pink noise was used as the contrast sound (pink noise). Among them, pink noise is a sound without semantic meaning, which is a common noise, similar to the sound of a car driving, and is used as a standard sound for noise loudness.
[0031] Furthermore, each audio data point in the audio dataset is labeled with a psychological decibel value, where the psychological decibel value is a decibel value defined as the level of noise to a person from different audio data. Further, sample audio data of contrast sound (pink noise) and audio data under different target environments are obtained. Eleven experimental sample audio samples and one contrast sound (pink noise) sample are combined into 11 groups as experimental materials, and the duration of the segments of the experimental sample audio and the contrast sound (pink noise) can be uniformly controlled at 10 seconds. Of course, in other embodiments, the number of experimental sample audio samples and the duration of the segments of the experimental sample audio and the contrast sound (pink noise) can be other numbers, which are not specifically limited here.
[0032] Furthermore, the contrast sound (pink noise) and the sample sound are adjusted to the same sound pressure level. In this embodiment, to eliminate the influence of sound pressure level, only the influence of the sound source on human psychological perception is considered. It is necessary to uniformly adjust the experimental sample sound and the contrast sound (pink noise) to the same sound pressure level. In the specific embodiment of this application, the two are adjusted to the same sound pressure level of 60 decibels.
[0033] Furthermore, the psychological decibel value of the sample sounds is labeled based on the perceived noise level of the sample sounds. Specifically, the experimenter listens to 11 sets of sounds (sample sounds and contrast sounds (pink noise)). The first sound in each set is the contrast sound (pink noise). After listening to the sample sound, the experimenter remembers the perceived noise level of the contrast sound (pink noise). Then, the experimenter listens to the second sound (sample sound) and compares its perceived noise level with that of the contrast sound (pink noise). Subsequently, the experimenter listens to the contrast sound (pink noise) again. During the listening process, based on the experimenter's psychological perception of the sound, the contrast sound (pink noise) is adjusted to a sound pressure level that corresponds to the same psychological perception. This sound pressure level is the psychological decibel value of the experimental sample, and the psychological decibel value of the sample sound is labeled. Similarly, the psychological decibel values of all sample sounds are labeled using the above method to obtain the audio dataset required for model training as described below. It is understood that the psychological decibel values of the sample sounds in this application can also be obtained through automatic labeling of psychological decibel values via learning methods, which is not specifically limited here. Based on the definition of psychological decibel value in the above steps, audio data of different target environments can be obtained and labeled with psychological decibel values to obtain an audio dataset.
[0034] S300 extracts the Mel-frequency inversion coefficients of all audio data in the audio dataset.
[0035] Optionally, in this application, the labeled experimental data obtained from the above experiments are used as model training data, and the Mel-frequency chopper coefficients (MFCC features) of all audio data in the audio dataset need to be further extracted.
[0036] MFCC is a method for converting outdoor ambient sound to the frequency domain and extracting its features using Mel scale and cepstral coefficients. Physically, it represents the spectral envelope of the audio signal, a crucial characteristic of the audio signal. The purpose of extraction is to eliminate information redundancy in the audio signal, facilitating the establishment of a connection between audio and the perceptual characteristics of the human auditory system. The process of extracting Mel cepstral coefficients includes three steps: Short-Time Fourier Transform (STFT), calculation of the Mel power spectrum, and calculation of the cepstral coefficients. Please refer to... Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of step S300 of this application, as shown below. Figure 2 The step S300 provided in this application further includes the following sub-steps:
[0037] S310 performs a short-time Fourier transform on the audio data to obtain the power spectrum of the audio data at different time segments.
[0038] The short-time Fourier transform extracted the power spectrum of the audio data at different time segments. In the time dimension, the audio digital signals at different time segments were segmented using a sliding window, with a window width of 2048 frames (approximately 46.4 ms) and a sliding step size of 512 frames (approximately 11.6 ms).
[0039] The audio signal at each time segment needs to be windowed to compensate for spectral leakage caused by the failure to meet the signal stationarity assumption. The windowed audio signal s... w Perform a discrete Fourier transform on (n,m)
[0040] The amplitude of S(k,m) is the energy of the signal at different frequencies over a period of time, i.e., the power spectrum. By combining the power spectra of different time segments and performing logarithmic processing, we can obtain the energy in decibels of the audio at different times and frequencies.
[0041] S320 processes the power spectrum of the audio data at different time segments to obtain the Mel-frequency cepstral coefficients of the audio data.
[0042] Furthermore, the audio signal at each time segment needs to be windowed to compensate for spectral leakage caused by the failure to meet the signal stationarity assumption. A discrete Fourier transform is performed on the windowed audio signal to obtain the power spectrum of the audio data at different time segments. By combining the power spectra of different time segments and performing logarithmic processing, the energy in decibels of the audio at different times and frequencies can be obtained. In this application, to simulate the difference in human ear sensitivity to different frequencies, a Mel filter bank is used to process the audio power spectrum at different time periods to obtain the Fourier transform matrix S at the Mel-scale frequency. mel (k,m). The reciprocal coefficient is a summary of the frequency envelope of the Mel power spectrum curve over a period of time, as shown in Equation (3), and is generated by the discrete cosine transform of the power spectrum.
[0043]
[0044] Where m is the sequence number of the time segment, n is the sequence number of the signal frame within the time segment, and N s Let k = 0, 1, 2, ..., N-1 represent the frequency band, and N be the number of signal frames within a time segment. n = 0, 1, 2, ..., 19, C n Cosine coefficient, The value for other cases is 1.
[0045] Thus, the Mel-frequency inverse coefficients of all audio data in the audio dataset can be obtained using the method described above.
[0046] Furthermore, after obtaining the Mel cepstral coefficient matrix, the project team compressed the matrix along the time series direction (m) into five dimensions: mean, 25th percentile, 75th percentile, slope, and kurtosis. The final input is a 20*5 Mel cepstral coefficient matrix.
[0047] S400 establishes a psychological decibel value model based on the FC-Boost regression model, which establishes the relationship between the Mel reciprocal coefficient and the psychological decibel value.
[0048] Please combine Figure 3 , Figure 3 This is a flowchart illustrating an embodiment of step S400 of this application, as shown below. Figure 3 The step S400 provided in this application further includes the following sub-steps:
[0049] S410, select Mel inverted frequency coefficients that meet preset conditions from all Mel inverted frequency coefficients to divide the audio dataset into multiple audio data subsets.
[0050] Understandably, for a single FC-Boost algorithm, regression is a supervised learning algorithm used to build a predictive model. Its basic principle is to progressively split the dataset using a series of binary judgments, dividing the data into multiple subsets, and then performing regression prediction on each subset. FC-Boost regression recursively divides the dataset into smaller subsets, fits a linear regression model to each subset, and finally combines the linear regression models of all subsets into a single overall model.
[0051] The specific steps of FC-Boost regression in this application are as follows: Selecting Mel-frequency inversion coefficients that meet preset conditions from all Mel-frequency inversion coefficients of the audio data to divide the audio dataset into multiple audio data subsets. A detailed description follows:
[0052] 1. Select the optimal feature as the root node. Select the optimal feature from all Mel-frequency cepstral coefficients (MFCC features) as the root node to divide the frequency dataset into two subsets.
[0053] 2. Use this feature to divide the frequency dataset into two subsets. Based on the selected optimal feature, divide the dataset into two subsets, one subset containing data points that satisfy the feature and the other subset containing data points that do not satisfy the feature.
[0054] 3. Recursively construct FC-Boost: For each subset, repeat the above two steps, select an optimal feature, divide the subset into two smaller subsets, until a certain termination condition is met.
[0055] S420, fits each subset of audio data to obtain the prediction result for the subset of audio data.
[0056] Furthermore, a regression method is used to fit each subset to obtain the prediction result for that subset. In this application, a regression method (such as linear regression) can be used to fit each of the above subsets to obtain the prediction results for all subsets.
[0057] S430 combines the prediction results of all audio data subsets to obtain the FC-Boost regression model.
[0058] Furthermore, the prediction results of all subsets are combined into a single overall model, which is the FC-Boost regression model. Optionally, this application uses the gradient boosting tree method to integrate multiple FC-Boost models that have undergone regression learning to establish the relationship between the Mel-frequency inverse coefficient and the psychological decibel value.
[0059] Please combine further Figure 4 , Figure 4 This is a flowchart illustrating an embodiment of step S430 of this application, as shown below. Figure 4Step S430 provided in this application further includes the following sub-steps:
[0060] S431, construct the initial FC-Boost model and train the initial FC-Boost model using audio data from the audio dataset.
[0061] Optionally, an initial FC-Boost model is constructed and trained using audio data from the audio dataset.
[0062] S432, calculate the error between the true value of each sample sound in the audio dataset and the predicted value of the initial FC-Boost model, and use it as the residual value of the sample sound.
[0063] Furthermore, for each sample sound in the audio dataset, the error between its true value and the prediction value of the initial FC-Boost model is calculated as the residual value of that sample sound.
[0064] S433: Construct a new FC-Boost sub-model using the residual values of each sample sound to fit the residual values of the next round of model training.
[0065] Furthermore, in the next training round, a new FC-Boost model is constructed using the residual values to fit the residual values. This process is repeated iteratively, with a new FC-Boost model constructed in each round to fit the residuals from the previous round.
[0066] S434 iterates continuously and adds the prediction results of all FC-Boost sub-models to obtain the final composite FC-Boost model.
[0067] By iterating continuously and summing the prediction results of all FC-Boost sub-models, the final prediction value is obtained, which is the final composite FC-Boost model.
[0068] Optionally, the main idea behind using gradient boosting trees in this application is to gradually improve the model through continuous iteration, thereby enhancing its predictive performance. In each iteration, the newly built FC-Boost model is constructed based on the previous iteration, thus more accurately capturing the details and complex relationships in the sample data. Simultaneously, gradient boosting trees can adaptively adjust the learning rate in each iteration, further improving the model's accuracy.
[0069] S500 obtains the psychological decibel value of the environment under test based on the psychological decibel value model.
[0070] It is understandable that after the model is established, the outdoor sound is subjected to feature extraction and model analysis to obtain the psychological decibel value of the environmental noise. Optionally, in this application, the noise source in the environment under test is transmitted to the sound acquisition unit by the microphone sound analog signal for audio preprocessing. Optionally, the original audio signal is a continuous voltage signal, which needs to be sampled and converted into a digital signal before it can be processed by the computer. As shown in formula (4), for the original audio signal s c (t):
[0071] s d (n)=s c (n / f) (2)
[0072] Where f is the sampling frequency, which is 44100Hz in this application, n is the sampling sequence, and s(n) is the sampled digital signal.
[0073] Furthermore, the digital signal is further encoded into a 16-bit binary number by a computer for processing. In one specific embodiment of this application, each segment of audio data is 10 seconds long, resulting in 441,000 frames of signal after sampling, and the encoded data is 882kB in size. To eliminate the influence of sound volume on the annoyance measurement, the decibel levels of all sounds are normalized to a uniform decibel level. After extracting the Mel-frequency inverse coefficients from the preprocessed audio data, model analysis is performed to obtain the psychological decibel value of the environmental noise.
[0074] Understandably, the machine learning model in this application uses the labeled experimental data obtained from the aforementioned experiments as training data. After extracting the Mel-frequency cepstral coefficients (MFCC features) from the audio data, it uses machine learning methods to learn and establish the relationship between the MFCC features and the psychological decibel value, i.e., to establish a psychological decibel value calculation model. After the model is established, outdoor sound can be analyzed by the model after feature extraction to obtain the psychological decibel of environmental noise. After extracting the Mel-frequency cepstral coefficients, this model uses a gradient boosting tree method to integrate multiple FC-Boost algorithms that have undergone regression learning to establish the relationship between the Mel-frequency cepstral coefficients and the psychological decibel value.
[0075] The above-described implementation method, for the first time, accurately reflects the psychological decibel value of environmental noise pollution. It obtains psychological decibel values from audio data under different environments through experiments, uses these values as a training dataset for the model, extracts features, and establishes a calculation model between Mel-frequency cepstral coefficients and psychological decibel values. The psychological decibel value of environmental noise can be obtained through this model. Furthermore, the method for measuring the psychological decibel value of environmental sound used in this application makes the environmental noise measurement results more scientific and accurate.
[0076] Please combine further Figure 5, Figure 5 This is a schematic diagram of one embodiment of the environmental sound psychodecibel measurement device of this application, as shown below. Figure 5 The ambient sound psychodecibel measuring device 100 in this application includes a microphone 110, a sound acquisition device 120, a central control unit 130, and a data transmission module 140.
[0077] The microphone 110 collects audio data from different environmental sound sources and transmits the environmental audio data to the sound acquisition unit 120 for audio preprocessing to obtain a digital signal. The digital signal is then transmitted to the central control host 130 for feature extraction, model training, and model calculation to obtain the psychological decibel value. Finally, the data is transmitted to the backend server via the data transmission module 140. Furthermore, the environmental sound psychological decibel measurement device of this application also includes a heat dissipation module 150 and a display module 160, used for heat dissipation of the device and display of relevant information, respectively.
[0078] It is understood that the method for measuring psychological decibel values (i.e., subjective evaluation experiments) in this application uses feature extraction and model learning to calculate the subjective evaluation experiment. Then, the feature extraction algorithm and model are deployed on the industrial control computer of the device, and the measuring device does not need to perform subjective evaluation experiments or build models.
[0079] The above-described implementation method, for the first time, accurately reflects the psychological decibel value of environmental noise pollution. It obtains psychological decibel values from audio data under different environments through experiments, uses these values as a training dataset for the model, extracts features, and establishes a calculation model between Mel-frequency cepstral coefficients and psychological decibel values. The psychological decibel value of environmental noise can be obtained through this model. Furthermore, the method for measuring the psychological decibel value of environmental sound used in this application makes the environmental noise measurement results more scientific and accurate.
[0080] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference] for details. Figure 6 , Figure 6 This is a basic structural block diagram of the computer device in this embodiment.
[0081] The computer device 300 includes a memory 301, a processor 302, and a network interface 303 that are interconnected via a system bus. It should be noted that... Figure 6Only a computer device 300 with components 301-303 is shown in this document; however, it should be understood that implementation of all shown components is not required, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0082] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0083] The memory 301 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 301 may be an internal storage unit of the computer device 300, such as the hard disk or memory of the computer device 300. In other embodiments, the memory 301 may also be an external storage device of the computer device 300, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 300. Of course, the memory 301 may also include both the internal storage unit and the external storage device of the computer device 300. In this embodiment, the memory 301 is typically used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions for interface calling methods. Furthermore, the memory 301 can also be used to temporarily store various types of data that have been output or will be output.
[0084] In some embodiments, the processor 302 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 302 is typically used to control the overall operation of the computer device 300. In this embodiment, the processor 302 is used to execute computer-readable instructions stored in the memory 301 or to process data, such as executing computer-readable instructions.
[0085] The network interface 303 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 300 and other electronic devices.
[0086] The above-described implementation method, for the first time, accurately reflects the psychological decibel value of environmental noise pollution. It obtains psychological decibel values from audio data under different environments through experiments, uses these values as a training dataset for the model, extracts features, and establishes a calculation model between Mel-frequency cepstral coefficients and psychological decibel values. The psychological decibel value of environmental noise can be obtained through this model. Furthermore, the method for measuring the psychological decibel value of environmental sound used in this application makes the environmental noise measurement results more scientific and accurate.
[0087] This application also provides another embodiment, namely, a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the method for measuring the ambient sound psychodecibel value as described above.
[0088] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0089] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A method for measuring the psychological decibel value of ambient sound, characterized in that, The measurement method includes: S100 acquires audio data from different environmental sound sources; S200, for the audio data from different environmental sound sources under the same sound pressure level, the psychological decibel value of the audio data is labeled by experimental method to obtain an audio dataset labeled with psychological decibel value; S300, Extract the Mel-frequency inverse coefficients of all audio data in the audio dataset; S400, based on the FC-Boost regression model, establishes a psychological decibel value model between the Mel reciprocal coefficient and the psychological decibel value, including: S410, Select Mel-frequency inversion coefficients that meet preset conditions from all Mel-frequency inversion coefficients to divide the audio dataset into multiple audio data subsets: Specifically as follows: ① Select the optimal feature as the root node: Select the optimal feature from all Mel cepstral coefficients as the root node to divide the audio dataset into two subsets; ② Use this feature to divide the audio dataset into two subsets: Based on the selected optimal feature, divide the dataset into two subsets, one subset containing data points that satisfy the feature and the other subset containing data points that do not satisfy the feature; ③ Recursively construct FC-Boost: For each subset, repeat steps ① and ② above, select an optimal feature, divide the subset into two smaller subsets, until a certain termination condition is met; S420, A regression method is used to fit each subset of audio data to obtain the prediction results for the subset of audio data; S430 combines the prediction results of all subsets into a single overall model, which is the FC-Boost regression model; Specifically, the gradient boosting tree method is used to integrate multiple FC-Boosts that have undergone regression learning to establish the relationship between Mel-frequency inverse coefficients and psychological decibel values; Specifically, it includes the following sub-steps: S431, Construct the initial FC-Boost model and train the initial FC-Boost model using audio data from the audio dataset; S432, calculate the error between the true value of each sample sound in the audio dataset and the predicted value of the initial FC-Boost model, and use it as the residual value of the sample sound; S433, a new FC-Boost model is constructed using the residual value of each sample sound to fit the residual value of the previous round of model training. This process is repeated iteratively, with a new FC-Boost model constructed in each round to fit the residual value of the previous round. S434, continuously iterates and sums the prediction results of all FC-Boost models to obtain the final FC-Boost regression model; S500 obtains the psychological decibel value of the environment under test based on the psychological decibel value model.
2. The measurement method according to claim 1, characterized in that, The method of labeling the audio data from different environmental sound sources with psychological decibel values under the same sound pressure level through experiments to obtain an audio dataset labeled with psychological decibel values includes: using pink noise energy as a standard to calibrate the different energy levels of the audio data under human perception, thereby obtaining the psychological decibel values of the audio data.
3. The measurement method according to claim 1, characterized in that, Extracting the Mel-frequency cepstral coefficients of all audio data in the audio dataset includes: performing a short-time Fourier transform on the audio data to obtain the power spectrum of the audio data at different time segments; and processing the power spectrum of the audio data at different time segments to obtain the Mel-frequency cepstral coefficients.
4. A measuring device for implementing the method for measuring the ambient sound psychodecibel value as described in any one of claims 1 to 3, characterized in that, The measuring device includes a microphone, a sound acquisition unit, a central control unit, and a data transmission module. The microphone acquires audio data from different environmental sound sources and transmits the environmental audio data to the sound acquisition unit for audio preprocessing to obtain a digital signal. The digital signal is then transmitted to the central control unit for feature extraction, model training, and model calculation to obtain the psychological decibel value. Finally, the data is transmitted to the backend server through the data transmission module.
5. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the method for calculating the ambient sound psycho-decibel value as described in any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the method for measuring the ambient sound psychodecibel value as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Ecological noise source measurement method and device, terminal and storage medium
CN114387987A
Environmental noise evaluation method based on annoyance perception index
CN119559968A