Method for locating and classifying environmental noise sources based on bayesian inference

By combining Bayesian inference with dynamic noise thresholding, time-frequency domain detection, and spherical acoustic array localization, along with a noise target classification network, high-precision detection and classification of environmental noise is achieved. This solves the problems of accuracy and correlation in noise source localization and identification in existing technologies, and improves the robustness and adaptability of the system.

CN119694344BActive Publication Date: 2026-04-10QINGDAO MINGDE ENVIRONMENTAL PROTECTION INSTR CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO MINGDE ENVIRONMENTAL PROTECTION INSTR CO LTD
Filing Date
2024-12-11
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing noise source localization and identification classification technologies suffer from problems such as misjudging excessive noise and low classification accuracy in complex environments, and the localization and identification classification lack correlation.

Method used

By employing a Bayesian inference-based approach, combining dynamic noise threshold setting, time-frequency joint detection, spherical acoustic array localization, and noise target classification network, high-precision noise source detection, localization, and classification are achieved through the fusion of localization and classification results via Bayesian inference.

Benefits of technology

It improves the accuracy of noise detection and the reliability of classification, reduces the false detection rate and the missed detection rate, enhances the robustness and adaptability of the system, and is suitable for complex noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119694344B_ABST
    Figure CN119694344B_ABST
Patent Text Reader

Abstract

The application discloses an environmental noise source positioning and classification method based on Bayesian inference, and belongs to the technical field of environmental noise detection. The steps are as follows: a dynamic noise threshold is set, and whether the noise exceeds the threshold is detected through noise time-frequency domain joint detection; a spherical acoustic array is used for noise source positioning; noise source identification and classification are carried out based on a noise target classification model to obtain the probability of each noise category; and the positioning result and the classification result are fused based on Bayesian inference to obtain the final noise classification result. The application improves the accuracy of noise monitoring, realizes accurate noise source positioning, improves the reliability of noise classification, enhances the robustness and adaptability of the system, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of environmental noise monitoring, and particularly relates to an environmental noise source positioning and classification method based on Bayesian inference. BACKGROUND

[0002] Currently, with the increasing complexity of environmental noise year by year, the existing noise source positioning and identification classification technology may have the problem of misjudgment of noise exceeding the standard or can only provide noise intensity and cannot further classify the specific types and positions of noise sources when facing complex environmental noise; in addition, the existing noise positioning and identification classification technology is independent, that is, the positioning and identification classification are not related, which will lead to a relatively low noise source classification accuracy to some extent. SUMMARY

[0003] In view of the above problems in the prior art, the application provides an environmental noise source positioning and classification method based on Bayesian inference, which is reasonable in design, solves the problems of the prior art, and has good effects.

[0004] The technical scheme of the application is as follows:

[0005] The environmental noise source positioning and classification method based on Bayesian inference comprises the following steps:

[0006] S1, setting a dynamic noise threshold;

[0007] S2, detecting whether the noise exceeds the standard through noise time-frequency domain joint detection;

[0008] S3, positioning the noise source by using a spherical acoustic array;

[0009] S4, identifying and classifying the noise source based on a noise target classification network to obtain the probability of each noise category;

[0010] S5, fusing the positioning result of S3 and the classification result of S4 based on Bayesian inference to obtain the final noise classification result.

[0011] Further, the S1 is specifically:

[0012] First, in a time window [t, t+Δt], noise signals x(t) are collected; the noise signals are divided into N frames, and each frame has a length L;

[0013] The energy value of each frame is calculated as:

[0014]

[0015] wherein, E n is the energy value of the nth frame, x n(i) is the value of the i-th sampling point of the n-th frame, n = 1, 2, …, N, i = 1, 2, …, L;

[0016] The average energy value E is calculated as:

[0017]

[0018] The standard deviation σ is calculated as: E

[0019]

[0020] The dynamic noise threshold T is set as:

[0021]

[0022] Wherein, K is an adjustment coefficient.

[0023] Further, the S2 comprises the following sub-steps:

[0024] S2.1, first, a preliminary detection is performed, and a time-domain energy detection is used to screen out possible excessive audio frames, specifically:

[0025] The continuous sound signal b(t) is segmented by frame length L and frame shift R to obtain a series of frames b n (i), wherein n is the frame number, b n (i) is the value of the i-th sampling point of the n-th frame;

[0026] Each frame b n (i) is multiplied by a window function ω(i), and the expression is:

[0027] c n (i) = b n (i) * ω(i); (5)

[0028] Wherein, c n (i) is the windowed signal;

[0029] The short-time energy E' is calculated as: n

[0030]

[0031] The short-time energy is first-order recursive smoothed, and the expression is:

[0032]

[0033] Wherein, is the smoothed short-time energy, and α is a constant, 0 < α < 1;

[0034] If ​​Continue to perform the frequency domain energy detection, if Determine that the noise is normal;

[0035] S2.2, the frequency domain energy detection is specifically:

[0036] Perform fast Fourier transform on each frame meeting , to obtain a spectrum X n (k), and calculate the spectrum amplitude as:

[0037]

[0038] Wherein, X n (k) is the spectrum of the nth frame, and |X n (k)| is the spectrum amplitude of the nth frame;

[0039] Calculate the spectrum energy, and accumulate the energy of the noise target frequency range [k1, k2], the expression is:

[0040]

[0041] Wherein, is the spectrum energy of the nth frame;

[0042] If , determine that the current noise is over standard, if , determine that the noise is normal.

[0043] Further, the S3 includes the following sub-steps:

[0044] S31, uniformly distribute M microphones on a spherical surface with a radius r to form a spherical acoustic array structure, take the array center as the origin to establish a three-dimensional rectangular coordinate system (x, y, z) and a corresponding spherical coordinate system Wherein r is the radial distance; theta is the azimuth angle, the value range is [0, 2pi); phi is the pitch angle, the value range is [0, pi).

[0045] Suppose that the sound source is located at an unknown position P s =(x s ,y s ,z s ), and the position of the mth microphone is P m =(x m ,y m ,z m ), then the distance d m from the sound source to the mth microphone is:

[0046] d m =‖P s -P m ‖;(10) ​

[0047] The signal received by the mth microphone x m (t) is:

[0048] x m (t) = a m s(t-τ m )+n m (t);(11)

[0049] wherein a m is an attenuation coefficient, d m is the distance between the sound source signal and the microphone; s(t-τ m ) is the signal of the source signal after a propagation delay τ m reaching the mth microphone, c is the speed of sound, and n m (t) is the noise of the mth microphone;

[0050] S32, time difference estimation is performed, a microphone is first selected as a reference microphone, denoted as the m0th microphone, and the cross-correlation function of the signals received by the mth microphone and the reference microphone is calculated is:

[0051]

[0052] The maximum value of the cross-correlation function is calculated, and the time delay corresponding to the maximum value is the relative time difference The expression is:

[0053]

[0054] A time difference matrix is constructed, each microphone and the reference microphone form a microphone pair, the time difference of all microphone pairs is collected to form a time difference matrix Δτ;

[0055] S33, the relationship between the time difference and the distance difference is:

[0056]

[0057] For each microphone pair, the following equation is established:

[0058]

[0059] wherein m = 1, 2, …, M and m ≠ m0; a nonlinear equation set is constructed;

[0060] Finally, a linearization method is used for position estimation, for far-field approximation, it is assumed that the distance of the sound source is much larger than the diameter of the spherical acoustic array structure, the nonlinear equation set is linearized, and the least square method is used to solve the sound source position P s ;

[0061] The sound source position vector is:

[0062] d = P s - P0; (17)

[0063] where P0 is the array center position, usually the origin;

[0064] The azimuth angle theta is: theta = arctan2(d y ,d x ), where d x ,d y are the x and y components of d, respectively;

[0065] The elevation angle phi is: where d z is the z component of d.

[0066] Further, a training set containing multiple noise categories is constructed, and the noise target recognition classification network is trained and tested using the training set,

[0067] The noise target recognition classification network includes an input layer, a preprocessing layer, a feature extraction layer, a sequence modeling layer, a full connection unit, and an output layer;

[0068] The input layer input is the noise exceeding signal q(t) determined after noise time-frequency domain joint detection, with a length size of (FN, 1), where FN is the sampling point number of the audio signal;

[0069] The preprocessing layer performs denoising and normalization processing on the audio signal, and converts the time domain signal into a cepstrum feature through a mel-frequency cepstrum coefficient;

[0070] The feature extraction layer sequentially includes a convolution layer 1, a maximum pooling layer 1, a convolution layer 2, a maximum pooling layer 2, a residual block 1, and a residual block 2;

[0071] The convolution layer 1 is a 1D convolution layer, containing 64 filters, with a convolution kernel size of 3 and a ReLU activation function;

[0072] The maximum pooling layer 1 and the maximum pooling layer 2 both have a size of 2;

[0073] The convolution layer 2 is a 1D convolution layer, containing 128 filters, with a convolution kernel size of 3 and a ReLU activation function;

[0074] The residual block 1 includes a convolution layer 3 and a convolution layer 4, both of which are 1D convolution layers, containing 256 filters, with a convolution kernel size of 3 and a ReLU activation function;

[0075] ​The residual block 2 includes a convolution layer 5 and a convolution layer 6, the convolution layer 5 and the convolution layer 6 are 1D convolution layers, containing 512 filters, the convolution kernel size is 3, and the activation function is ReLU;

[0076] The sequence modeling layer includes a bidirectional LSTM layer 1 and a bidirectional LSTM layer 2, the bidirectional LSTM layer 1 includes 128 units, and the bidirectional LSTM layer 2 includes 256 units;

[0077] The full connection unit includes a flattening layer, a full connection layer 1, a Dropout layer 1, a full connection layer 2 and a Dropout layer 2;

[0078] The flattening layer expands a three-dimensional feature map into a one-dimensional vector;

[0079] The full connection layer 1 includes 512 neurons, and the activation function is ReLU;

[0080] The Dropout rate of the Dropout layer 1 is 0.5;

[0081] The full connection layer 2 includes 256 neurons, and the activation function is ReLU;

[0082] The Dropout rate of the Dropout layer 2 is 0.5;

[0083] The output layer is a full connection layer 3, the number of output neurons is the number of categories, and the activation function is Softmax.

[0084] Further, the S5 includes the following sub-steps:

[0085] S51, processing q(t) by using a noise target recognition classification network, and outputting a prior probability of each noise category, denoted as P(Class i |NetworkOut,Loc);

[0086] S52, assuming that, under each noise category, the positioning feature obeys a Gaussian distribution:

[0087]

[0088] Wherein, μ i is the mean of the pitch angle under the i-th noise category; σ i is the standard deviation of the pitch angle under the i-th noise category;

[0089] S53, for the observed pitch angle , calculating the probability of observing the pitch angle under each noise category, denoted as The expression is:

[0090]

[0091] S54, applying Bayes formula, the posterior probability of each noise category is calculated as:

[0092]

[0093] Where j represents the index when traversing all categories, that is, there are C noise categories in the output category of the neural network, then j=1, 2, …C;

[0094] S55, according to the posterior probability obtained by calculation, the noise category with the maximum probability is selected as the final noise classification result.

[0095] The beneficial technical effects brought by the present application are:

[0096] The present application combines dynamic threshold setting, time-frequency domain joint detection, spherical acoustic array positioning, neural network classification and fusion method based on probability model, realizes high-precision detection, positioning and classification of environmental noise. Its beneficial technical effects mainly reflect in the following aspects:

[0097] Improve the accuracy of noise detection: 1, dynamic threshold setting: according to the real-time acquisition of environmental noise data, dynamically adjust the noise threshold, adapt to the change of complex noise environment, reduce the false rejection rate and the false rejection rate. 2 Time-frequency domain joint detection: through the dual detection of time domain and frequency domain, the possible noise exceeding the standard is accurately determined, which further improves the accuracy of detection.

[0098] Realize accurate noise source positioning: use spherical acoustic array and time difference estimation technology to accurately obtain the coordinates, azimuth, pitch angle and distance of noise source relative to the array center. Through effective positioning of noise source in the environment, reliable data support is provided for subsequent noise source management and control.

[0099] Improve the reliability of noise classification: 1, use deep learning model to identify and classify noise, which can accurately distinguish different types of noise sources. 2 Bayesian inference fusion: combine the positioning result with the classification result, use Bayesian inference to correct the classification probability, improve the accuracy and reliability of classification.

[0100] Enhance the robustness and adaptability of the system: 1, multi-information fusion: combine acoustic characteristics and positioning characteristics, the system can maintain high performance in complex and variable noise environment. 2 Reduce false rejection and false rejection: through comprehensive decision and fusion algorithm, reduce the false rejection rate and false rejection rate in the process of noise detection and classification.

[0101] It has broad application prospects: it is mainly used in the field of environmental monitoring, and is suitable for urban noise monitoring, traffic noise management, industrial noise control and other fields.

[0102] This invention achieves high-precision detection, localization, and classification of environmental noise, significantly improving the system's accuracy, reliability, and intelligence. Its beneficial technical effects include reducing false positives and false negatives, improving localization and classification accuracy, and enhancing the system's robustness and adaptability, making it a promising candidate for widespread applications. Attached Figure Description

[0103] Figure 1 This is a flowchart of the environmental noise source localization and classification method based on Bayesian inference in this invention;

[0104] Figure 2 This is a flowchart of the joint time-frequency domain noise detection process in this invention to determine whether noise exceeds the standard;

[0105] Figure 3 This is an installation diagram of the environmental monitoring equipment used in the application of the method of the present invention;

[0106] Figure 4 This is the output image from the camera of the environmental monitoring equipment;

[0107] (a) is the image before the noise source was detected; (b) is the image after the noise was detected.

[0108] Figure 5 A diagram showing the noise classification results output by environmental monitoring equipment; Detailed Implementation

[0109] The specific embodiments of the present invention will be further described below with reference to specific examples:

[0110] Methods for locating and classifying environmental noise sources based on Bayesian inference, such as Figure 1 and 2 As shown, it includes the following sub-steps:

[0111] S1. Set the dynamic noise threshold;

[0112] First, within the time window [t, t+Δt], a noise signal x(t) is acquired; this noise signal is then divided into N frames, each with a length of L.

[0113] The energy value for each frame is calculated as follows:

[0114]

[0115] Among them, E n Let x be the energy value of the nth frame. n (i) represents the value of the i-th sampling point in the n-th frame, where n = 1, 2, ..., N, and i = 1, 2, ..., L;

[0116] The average energy value E is calculated as follows:

[0117]

[0118] Calculate the standard deviation σ E for:

[0119]

[0120] Set the dynamic noise threshold T as:

[0121]

[0122] Where K is the adjustment coefficient, which can be 2 or 3 depending on the actual situation.

[0123] Based on real-time collected environmental noise data, the noise threshold is dynamically adjusted to adapt to complex noise environment changes and improve detection accuracy.

[0124] S2. To improve detection accuracy and reduce false positive and false negative rates, noise is detected using a combined time-frequency domain method to determine if noise exceeds the standard. This includes the following sub-steps:

[0125] S2.1 First, a preliminary detection is performed, using temporal energy detection to quickly filter out possible out-of-range audio frames, specifically:

[0126] The continuous acoustic signal b(t) is divided into a frame length L and a frame shift R to obtain a series of frames b. n (i), where n is the frame number, b n (i) represents the value of the i-th sampling point in the n-th frame;

[0127] For each frame b n (i) Multiplied by the window function ω(i), the expression is:

[0128] c n (i)=b n (i)*ω(i); (5)

[0129] Among them, c n (i) is the windowed signal. This process is used to reduce the "leakage effect" that may occur when the signal is analyzed in the frequency domain and helps to obtain a more accurate spectrum estimate.

[0130] Calculate the short-time energy E′ n for:

[0131]

[0132] To reduce fluctuations, a first-order recursive smoothing is applied to the short-time energy, expressed as:

[0133]

[0134] wherein, is the smoothed short-time energy, and a is a constant, 0 < a < 1, usually taking a value of 0.9;

[0135] If continue to perform the frequency domain energy detection, if determine that the noise is normal;

[0136] S2.2, the frequency domain energy detection is specifically:

[0137] performing fast Fourier transform on each frame meeting to obtain a frequency spectrum X n (k), and calculating a frequency spectrum amplitude as:

[0138]

[0139] wherein, X n (k) is the frequency spectrum of the nth frame, and |X n (k)| is the frequency spectrum amplitude of the nth frame;

[0140] calculating the frequency spectrum energy, and performing energy accumulation on a noise target frequency range [k1, k2], and the expression is:

[0141]

[0142] wherein, is the frequency spectrum energy of the nth frame;

[0143] If determine that the current noise is over-standard, if determine that the noise is normal.

[0144] S3, using a spherical acoustic array to perform noise source positioning, including the following sub-steps:

[0145] S31, uniformly distributing M microphones on a spherical surface with a radius of r, taking the array center as the origin, the array center being the center of the sphere, establishing a three-dimensional rectangular coordinate system (x, y, z) and a corresponding spherical coordinate system wherein r is a radial distance; θ is an azimuth angle, taking a value range of [0, 2π); φ is a pitch angle, taking a value range of [0, π);

[0146] assuming that a sound source is located at an unknown position P s =(x s ,y s ,z s ), and the position of the mth microphone is P m =(x m ,y​m z m ), the distance from the sound source to the mth microphone is:

[0147] d m =‖P s -P m ‖;(10)

[0148] The signal received by the mth microphone is:

[0149] x m (t)=a m s(t-τ m )+n m (t);(11)

[0150] where a m is the attenuation coefficient, d m is the distance between the sound source signal and the microphone; τ m is the propagation delay, c is the speed of sound, about 343 m / s, n m (t) is the noise of the mth microphone;

[0151] S32, time difference estimation is performed, first select a microphone as a reference microphone, usually the first microphone m0 is selected, the cross-correlation function of the signal received by the mth microphone and the reference microphone is calculated as:

[0152]

[0153] The maximum value of the cross-correlation function corresponds to the time delay, that is, the relative time difference The expression is:

[0154]

[0155] The time difference matrix is constructed, and the time differences of all microphone pairs are collected to form the time difference matrix Δτ;

[0156] S33, geometric relationship is established, and the relationship between the time difference and the distance difference is:

[0157]

[0158] For each microphone pair, the following equation is established:

[0159]

[0160] where m=1, 2, …, M, and m≠m0; a nonlinear equation set is constructed;

[0161] Finally, linearization method or iterative optimization method is used for position estimation.

[0162] In the linearization method, for the far-field approximation, the nonlinear equation set is linearized by assuming that the sound source distance is much larger than the array size, and the sound source position P is solved using the least squares method s .

[0163] In the iterative optimization method, an initial position is randomly set first, and the Levenberg-Marquardt algorithm is used for iterative optimization to minimize the residual as follows:

[0164]

[0165] The sound source position vector is:

[0166] d = P s - P0; (17)

[0167] where P0 is the center position of the array, usually the origin;

[0168] The azimuth angle θ is: θ = arctan2(d y ,d x ), where d x ,d y are the x and y components of d, respectively;

[0169] The pitch angle is: where d z is the z component of d.

[0170] S4, identifying and classifying the noise source based on the noise target classification network to obtain the probability of each noise category;

[0171] The noise target identification and classification network includes an input layer, a preprocessing layer, a feature extraction layer, a sequence modeling layer, a full connection unit, and an output layer;

[0172] The input layer inputs the audio signal, and the length dimension is set to (FN, 1), where FN is the number of sampling points of the audio signal;

[0173] The preprocessing layer performs denoising and normalization processing on the audio signal, converts the time domain signal into a cepstrum feature through a mel frequency cepstrum coefficient (MFCC), and for example, the feature dimension obtained by calculating the MFCC feature is FM (for example, a 13-dimensional vector);

[0174] The feature extraction layer sequentially includes a convolution layer 1, a maximum pooling layer 1, a convolution layer 2, a maximum pooling layer 2, a residual block 1, and a residual block 2;

[0175] The convolution layer 1 is a 1D convolution layer, which includes 64 filters, the convolution kernel size is 3, and the activation function is ReLU;

[0176] The size of the max-pooling layer 1 and the max-pooling layer 2 is 2;

[0177] The convolutional layer 2 is a 1D convolutional layer, includes 128 filters, the convolution kernel size is 3, and the activation function is ReLU;

[0178] The residual block 1 includes the convolutional layer 3 and the convolutional layer 4, the convolutional layer 3 and the convolutional layer 4 are 1D convolutional layers, include 256 filters, the convolution kernel size is 3, and the activation function is ReLU;

[0179] The residual block 2 includes the convolutional layer 5 and the convolutional layer 6, the convolutional layer 5 and the convolutional layer 6 are 1D convolutional layers, include 512 filters, the convolution kernel size is 3, and the activation function is ReLU;

[0180] The sequence modeling layer includes the bidirectional LSTM layer 1 and the bidirectional LSTM layer 2, the bidirectional LSTM layer 1 includes 128 units, and the bidirectional LSTM layer 2 includes 256 units;

[0181] The full connection unit includes the flattening layer, the full connection layer 1, the Dropout layer 1, the full connection layer 2 and the Dropout layer 2;

[0182] The flattening layer expands the three-dimensional feature map into a one-dimensional vector;

[0183] The full connection layer 1 includes 512 neurons, and the activation function is ReLU;

[0184] The Dropout rate of the Dropout layer 1 is 0.5;

[0185] The full connection layer 2 includes 256 neurons, and the activation function is ReLU;

[0186] The Dropout rate of the Dropout layer 2 is 0.5;

[0187] The output layer is the full connection layer 3, the number of output neurons is the number of categories, and the activation function is Softmax.

[0188] S5, fusing the positioning result of S3 and the classification result of S4 based on Bayesian inference to obtain a final noise classification result.

[0189] The Bayesian theorem describes the relationship between the posterior probability P(A|B) of an event A, the prior probability P(A) and the likelihood function P(B|A) under the condition of knowing B as follows:

[0190]

[0191] where P(A) is the prior probability of event A; P(B|A) is the probability of B given event A (likelihood function); P(B) is the marginal probability of B, which can be computed by the law of total probability:

[0192] P(B) = ∑ i P(B|A i )P(A i );(19)

[0193] where event A is the class of noise source Class i , where i represents the class number. For example, "car horn", "drone buzzing", etc.

[0194] The observation data B includes the classification result of the neural network and the positioning result (such as the pitch angle and the azimuth angle).

[0195] S5 includes the following sub-steps:

[0196] S51, processing the acoustic signal by using the noise target recognition classification network, and outputting the prior probability of each noise class, denoted as P(Class i |NetworkOut, Loc), wherein NetworkOut represents the classification result of the neural network;

[0197] S52, establishing a likelihood function of the positioning feature, the positioning result is associated with the noise class, and the probability of observing a specific positioning feature under each class, i.e. the likelihood function, needs to be established;

[0198] Suppose that under each noise class, the positioning feature follows a Gaussian distribution:

[0199]

[0200] where μ i is the mean of the pitch angle under the i-th noise class; σ i is the standard deviation of the pitch angle under the i-th noise class;

[0201] For example, in the parameter setting, the μ Car of car horn (Car Horn) is 0 degrees, and the σ Car is 10 degrees; the μ Drone of drone buzzing (Drone Buzzing) is 75 degrees, and the σ Drone is 10 degrees.

[0202] S53, calculating the likelihood value, for the observed pitch angle, calculating the probability of observing the pitch angle under each noise class, denoted as The expression is:

[0203]

[0204] S54, applying Bayes formula, calculating the posterior probability of each noise category as:

[0205]

[0206] Wherein, j represents the index when traversing all categories, that is, the output category of the neural network has a total of C noise categories, then j = 1, 2, … C;

[0207] S55, according to the calculated posterior probability, selecting the noise category with the maximum probability as the final noise classification result.

[0208] The method of the present application is applied in an environmental noise monitoring device, which is integrated with a camera, as shown in Figure 3 When the spherical acoustic array detects noise exceeding the standard, the azimuth, elevation and distance of the noise source are calculated, the camera rotates to the position according to the positioning information to take pictures of the noise exceeding the standard behavior, wherein the noise source target is located in the middle of the video image, represented by a white strip, as shown in Figure 4 (a) and (b), and the corresponding noise source classification result is as shown in Figure 5 .

[0209] Of course, the above description is not a limitation of the present application, and the present application is not limited to the above examples, and the changes, modifications, additions or replacements made by the person skilled in the art within the essential scope of the present application should also belong to the protection scope of the present application.

Claims

1. A method for environmental noise source localization and classification based on Bayesian inference, characterized in that, The method comprises the following steps: S1, setting a dynamic noise threshold; S2, detecting whether the noise exceeds the threshold through a noise time-frequency domain, and if the noise exceeds the threshold, continuing to perform S3; S3, positioning the noise source by using a spherical acoustic array; S4, identifying and classifying the noise source based on a noise target classification network to obtain the probability of each noise category; S5, fusing the positioning result of S3 and the classification result of S4 based on Bayesian inference to obtain the final noise classification result; S5 comprises the following sub-steps: S51, using the noise target recognition classification network on the processed image, output the prior probability of each noise category, denoted as ; S52, assuming that under each noise category, the positioning feature obeys a Gaussian distribution: ; (20) wherein is the average value of the pitch angle under the first noise class; is the standard deviation of the pitch angle under the first noise class; S53、For the observed pitch angle , the probability of observing this pitch angle under each noise class is computed, denoted as , the expression is: ;(21) S54, applying the Bayesian formula to calculate the posterior probability of each noise category as: ;(22) wherein, denotes the index when traversing all classes, i.e. there are C classes of noise in the output classes of the neural network, then j = 1, 2,... C; S55, according to the calculated posterior probability, selecting the noise category with the maximum probability as the final noise classification result.

2. The method of claim 1, wherein, The S1 is specifically: First, in the time window Inside, noise signals are collected. ; divide the noise signal into There are 3 frames, each with a length of 1. ; The energy value of each frame is calculated as: ;(1) wherein is the energy value of the frame, is the value of the sample point of the frame, , ; calculating the average energy value is: ;(2) Computing the standard deviation is: ;(3) Setting dynamic noise threshold is: ;(4) wherein is a modulation coefficient.

3. The method of claim 2, wherein, The S2 comprises the following sub-steps: S2.1, first, preliminary detection is performed, and time-domain energy detection is used to screen out possible exceeding audio frames, specifically: segmenting a continuous acoustic signal into frame lengths and frame shifts to obtain a series of frames wherein is a frame number, is a value of a sample point of the frame . for each frame multiplying by a window function with the expression ;(5) wherein is the windowed signal; Computing short-time energy is: ;(6) First-order recursive smoothing is performed on the short-time energy, and the expression is: ;(7) wherein is the smoothed short-time energy, is a constant, ; If , then the frequency domain energy detection is continued, and if , then it is determined that the noise is normal. S2.2, frequency-domain energy detection is specifically: The fast Fourier transform is performed on each frame that meets the criteria to obtain a frequency spectrum The frequency spectrum amplitude is calculated as: ;(8) wherein, is the spectrum of the frame, is the spectrum amplitude of the frame; Computing the spectral energy, for the noise target frequency range Energy accumulation is performed, expressed as: ;(9) wherein is the first frame of the spectrum energy; If , it is determined that the current noise is over-standard, and if , it is determined that the noise is normal.

4. The method of claim 3, wherein, The S3 comprises the following sub-steps: S31, uniformly distributing on a spherical surface with a radius of r a microphone, constituting a spherical acoustic array structure, taking the center of the array as the origin, establishing a three-dimensional rectangular coordinate system and a corresponding spherical coordinate system , wherein is a radial distance; is an azimuth angle, the value range is ; is an elevation angle, the value range is ; Let the sound source be located at an unknown position , the position of the th microphone is , then the distance from the sound source to the th microphone is: ;(10) The first signal received by the first microphone is: ;(11) wherein is the attenuation coefficient, , is the distance between the sound source signal and the microphone; is the signal of the source signal after a propagation delay to the th microphone, , is the sound velocity, is the noise of the th microphone; S32, time difference estimation, first select a microphone as a reference microphone, the reference microphone is recorded as the first microphone, the cross-correlation function of the signal received by the first microphone and the reference microphone is calculated : ;(12) The time delay corresponding to the maximum of the cross-correlation function is the relative time difference The expression is: ;(13) Constructing the time difference matrix, each microphone and the reference microphone constitute a microphone pair, collect all the time difference of microphone pairs, form the time difference matrix ; S33, relationship between time difference and distance difference is: ;(14) For each microphone pair, the following equation is established: ;(15) wherein and ≠ ; constructing a system of nonlinear equations; Finally, the linearization method is used to estimate the position of the sound source. For the far-field approximation, the distance between the sound source and the spherical acoustic array is assumed to be much larger than the diameter of the spherical acoustic array. The nonlinear equations are linearized, and the least square method is used to solve the position of the sound source ; The sound source position vector is: ;(17) wherein is the array center position, typically the origin; azimuth angle is: wherein are of and components; pitch angle is: wherein is of components.

5. The method of claim 4, wherein, A training set containing multiple noise categories is constructed, and the noise target identification and classification network is trained and tested by using the training set, The noise target identification and classification network comprises an input layer, a preprocessing layer, a feature extraction layer, a sequence modeling layer, a full connection unit and an output layer; The input layer input is a signal determined as exceeding noise after noise time-frequency domain joint detection , set the length size as , wherein is the sampling point number of the audio signal; The preprocessing layer performs denoising and normalization processing on the audio signal, and converts the time domain signal into a cepstrum feature through a mel-frequency cepstrum coefficient; The feature extraction layer comprises a convolution layer 1, a maximum pooling layer 1, a convolution layer 2, a maximum pooling layer 2, a residual block 1 and a residual block 2 in sequence; The convolution layer 1 is a 1D convolution layer, contains 64 filters, the convolution kernel size is 3, and the activation function is ReLU; The maximum pooling layer 1 and the maximum pooling layer 2 both have a size of 2; The convolution layer 2 is a 1D convolution layer, contains 128 filters, the convolution kernel size is 3, and the activation function is ReLU; The residual block 1 comprises a convolution layer 3 and a convolution layer 4, and the convolution layer 3 and the convolution layer 4 are 1D convolution layers, contain 256 filters, the convolution kernel size is 3, and the activation function is ReLU; The residual block 2 comprises a convolution layer 5 and a convolution layer 6, and the convolution layer 5 and the convolution layer 6 are 1D convolution layers, contain 512 filters, the convolution kernel size is 3, and the activation function is ReLU; The sequence modeling layer comprises a bidirectional LSTM layer 1 and a bidirectional LSTM layer 2, the bidirectional LSTM layer 1 comprises 128 units, and the bidirectional LSTM layer 2 comprises 256 units; The full connection unit comprises a flattening layer, a full connection layer 1, a Dropout layer 1, a full connection layer 2 and a Dropout layer 2; The flattening layer expands the three-dimensional feature map into a one-dimensional vector; The full connection layer 1 comprises 512 neurons, and the activation function is ReLU; The Dropout rate of the Dropout layer 1 is 0.5; The full connection layer 2 includes 256 neurons, and the activation function is ReLU; The Dropout rate of the Dropout layer 2 is 0.5; The output layer is a full connection layer 3, the number of output neurons is the number of categories, and the activation function is Softmax.

Citation Information

Patent Citations

  • Noise source event detection method and device, electronic equipment and storage medium

    CN115312075A

  • Audio processing method and device, storage medium and intelligent glasses

    CN115775564A