Animal cry accurate identification method of integrated animal cry sensor
Through the integrated animal call sensor method, the combined noise reduction model and deep learning model are used to process multi-channel audio signals, which solves the problem of incomplete sound recognition characteristics and low accuracy, and achieves accurate recognition of target animal calls in complex environments.
Patent Information
- Application Number
- CN202510338116.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the sound recognition and extraction analysis features are incomplete, resulting in low accuracy of sound recognition, especially in complex natural environments, it is difficult to separate the call of target animals from background noise.
The integrated animal call sensor method is adopted to obtain multi-channel audio signals, use the joint noise reduction model to perform wavelet decomposition and optimization processing, combine the time-frequency feature model and the spatial feature model for fusion processing, and finally use the deep learning model to process the time-frequency spatial feature signal and the optimized wavelet coefficient to obtain the target audio signal.
Effectively separating the target sound source from background noise improves the accuracy of sound recognition, can distinguish species with similar frequency spectrum but different vocal locations, and is suitable for feature analysis in complex environments.
Smart Images

Figure CN119964582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and in particular to a method for accurately identifying animal calls using an integrated animal call sensor. Background Art
[0002] Accurately identifying and analyzing the calls of specific species in natural environments has long been a research hotspot in the fields of ecology, animal behavior, and agricultural monitoring. Traditional sound collection and processing methods mainly rely on a single microphone or a simple array system to capture sound signals in the environment. However, due to the complexity and diversity of background noise in nature, these methods are often difficult to effectively separate the calls of target animals from the noisy background. In this embodiment, the sound of wind, rain, and other creatures may mask the target call, thereby affecting the accuracy of recognition.
[0003] In terms of animal call feature recognition, existing research has mostly focused on extracting basic parameters in the time domain or frequency domain, such as amplitude, frequency, and duration. However, these parameters are too simple for the high diversity and complexity of calls of different types of animals, and the extracted features are not comprehensive enough, which limits the accuracy of sound recognition. Summary of the invention
[0004] The present invention provides an animal call accurate recognition method and system of an integrated animal call sensor, which are used to solve the defects of incomplete sound recognition extraction and analysis features and low sound recognition accuracy in the prior art.
[0005] The present invention provides an animal call accurate recognition method of an integrated animal call sensor, comprising:
[0006] Acquire a multi-channel audio signal of a sound source, and use a joint noise reduction model to perform L-level wavelet decomposition joint optimization processing on the multi-channel audio signal to obtain optimized wavelet coefficients;
[0007] The multi-channel audio signal is fused using a time-frequency feature model and a space feature model to obtain a time-frequency space feature signal;
[0008] The time-frequency-space feature signal and the optimized wavelet coefficients are processed using a deep learning model to obtain a target audio signal.
[0009] According to a method for accurately identifying animal sounds of an integrated animal sound sensor provided by the present invention, the method uses a joint noise reduction model to perform L-level wavelet decomposition and joint optimization processing on the multi-channel audio signal to obtain optimized wavelet coefficients, including:
[0010] Set up a noisy signal model:
[0011] y=x+n
[0012] Among them, x is a clean signal, n~N(0,σ 2 ) is Gaussian noise, y is the observed signal;
[0013] Perform L-level wavelet decomposition on the observed signal, and obtain the approximate coefficient a at each level l and detail factor d l :
[0014]
[0015] Initial value a 0 =y
[0016] is the lth-order low-pass filter with a cutoff frequency of f low ;
[0017] is the first-order high-pass filter with a cutoff frequency of f high ;
[0018] Output the decomposition coefficient set of all channels
[0019] Enter the decomposition coefficient
[0020] Design a joint objective function with multiple levels of constraints:
[0021]
[0022] It is the data fidelity term, which ensures the consistency between the estimated signal and the observed signal;
[0023] is a sparse regularization term, through l 1 The norm promotes the sparsity of wavelet coefficients;
[0024] is the high-frequency coefficient approximation term, which constrains the high-frequency components of the restored signal to be equal to the high-frequency components of the observed signal. l Close to, λ, β are weight parameters, and the high-frequency coefficient d of the observed signal l Contains some real signal components;
[0025] Iteratively optimize the joint objective function and decomposition coefficients to update x, z l , μ l ;
[0026] Reconstruct the signal by inverse wavelet transform:
[0027]
[0028] in is the optimized wavelet coefficient, k means there are k coefficients in each level.
[0029] According to a method for accurately identifying animal calls using an integrated animal call sensor provided by the present invention, the multi-channel audio signal is fused using a time-frequency feature model and a spatial feature model, and the time-frequency feature model processing includes:
[0030] Pre-emphasize the original time domain signal x(t) to enhance the high-frequency component, divide the time domain signal with enhanced high-frequency component into frames, and add a Hamming window to each frame;
[0031] Calculate the amplitude spectrum and Mel filter of each frame of the time domain signal, and calculate the MFCC and Mel spectrum features.
[0032] According to a method for accurately identifying animal calls using an integrated animal call sensor provided by the present invention, the multi-channel audio signal is fused using a time-frequency feature model and a spatial feature model, and the spatial feature model processing includes:
[0033] Set the multi-channel audio signal of the sound source to:
[0034] X=A(θ)S+N
[0035] Where A(θ)=[a(θ 1 ),…,a(θ K )] is the guidance matrix,
[0036] a(θ)=[1,e -j2πdsinθ / λ ,…,e -j2π(M-1)dsin θ / λ ] T ;
[0037] Perform covariance matrix decomposition on the signal:
[0038]
[0039] Among them U n is the noise subspace;
[0040] Using spatial spectrum calculation:
[0041]
[0042] For the i,j microphone pair, calculate the generalized cross-correlation (GCC-PHAT):
[0043]
[0044] Take the peak position T ij As a delay estimate, construct a spatial contrast matrix:
[0045] C=[Tij] M×M .
[0046] According to a method for accurately identifying animal calls using an integrated animal call sensor provided by the present invention, the multi-channel audio signal is fused using a time-frequency feature model and a spatial feature model, and the fusion process includes:
[0047] Time-frequency feature mapping:
[0048] The 39-dimensional MFCC and the 64-dimensional Mel spectrum are concatenated into a 103-dimensional vector and arranged by time frame into a matrix FTF∈R T×103 , expanded to 128×128 by bilinear interpolation:
[0049]
[0050] Spatial feature polar coordinate mapping:
[0051] Step 1: Convert the spatial spectrum to polar coordinates and convert P MUSIC (θ) is discretized into 128 directions according to the angle θ∈[0°,360°), mapped to the polar coordinate grid (r,θ), with a fixed radius r=1;
[0052] Step 2: Spatial contrast matrix encoding, transform the delay τ of C ij Converted to angle difference Δθ ij Fill in the corresponding position of the polar coordinate grid;
[0053] Step 3: Interpolate to generate an image, fill the sparse polar coordinate data into a 128×128 matrix F through nearest neighbor interpolation Space ;
[0054] Feature map concatenation:
[0055] Concatenate time-frequency and spatial features along the channel dimension:
[0056]
[0057] Among them, 128 is the time frame, the frequency point is 128, and the feature dimension is 2.
[0058] According to a method for accurately identifying animal sounds of an integrated animal sound sensor provided by the present invention, the method uses a deep learning model to process the time-frequency-space feature signal and the optimized wavelet coefficients to obtain a target audio signal, including:
[0059] CNN branch: Input time-frequency-space features into the deep learning model to extract local time-frequency-space characteristics;
[0060] FCN branch: Input optimized wavelet coefficients into the deep learning model to capture the characteristics of the target sound source;
[0061] The time-frequency-spatial features of the CNN branch and the features of the target sound source of the FCN branch are fused to output the target audio signal.
[0062] According to the method for accurately identifying animal calls of an integrated animal call sensor provided by the present invention, before fusing the multi-channel audio signal using the time-frequency feature model and the space feature model, the method further includes:
[0063] Input a multi-channel audio signal, filter it through the A-weighted transfer function, and obtain the amplitude of the filtered signal;
[0064] The amplitude of the filtered signal is adjusted by dynamic range compression to obtain a signal with normalized amplitude;
[0065] The frequency domain signal energy after A-weighted filtering is processed by A-weighted sound pressure level calculation to obtain sound pressure level data for dynamic compression parameter adjustment;
[0066] Finally, the sound pressure level data is processed by Kalman filter noise tracking to obtain the noise-reduced sound pressure level data;
[0067] The noise-reduced sound pressure level data is optimized through the loss function until the bit error rate meets the preset conditions.
[0068] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for accurately identifying animal calls of the integrated animal call sensor described in any one of the above technical solutions are implemented.
[0069] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method for accurately identifying animal calls using the integrated animal call sensor described in any one of the above technical solutions are implemented.
[0070] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method for accurately identifying animal calls using the integrated animal call sensor described in any one of the above technical solutions.
[0071] The animal call accurate identification method of the integrated animal call sensor provided by the present invention obtains optimized wavelet coefficients and time-frequency space feature signals by processing multi-channel audio signals, effectively separates the target sound source from the background noise, and can distinguish species with similar frequency spectra but different sounding positions. The optimized wavelet coefficients and time-frequency space feature signals are processed by a deep learning model, and the influence of frequency domain and space on audio signal processing is taken into account. The feature analysis in complex environments is more comprehensive, and the accuracy of the target audio signal is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0073] Figure 1 It is a flow chart of a method for accurately identifying animal sounds using an integrated animal sound sensor provided by the present invention. DETAILED DESCRIPTION
[0074] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0075] Combine the following Figure 1 The animal sound accurate recognition method of the integrated animal sound sensor of the present invention comprises:
[0076] Acquire a multi-channel audio signal of a sound source, and use a joint noise reduction model to perform L-level wavelet decomposition joint optimization processing on the multi-channel audio signal to obtain optimized wavelet coefficients;
[0077] The multi-channel audio signal is fused using a time-frequency feature model and a space feature model to obtain a time-frequency space feature signal;
[0078] The time-frequency-space feature signal and the optimized wavelet coefficients are processed using a deep learning model to obtain a target audio signal.
[0079] By processing multi-channel audio signals, we obtain optimized wavelet coefficients and time-frequency-space feature signals, effectively separate the target sound source from the background noise, and distinguish species with similar spectra but different vocalization positions. We use a deep learning model to process the optimized wavelet coefficients and time-frequency-space feature signals, taking into account the impact of frequency domain and space on audio signal processing, making the feature analysis in complex environments more comprehensive and improving the accuracy of the target audio signal.
[0080] Specifically, the sound collection module is composed of a linearly arranged six-channel MEMS microphone array, with a spacing of d = 8 ± 0.5 cm between adjacent microphones, a frequency response covering 10 Hz to 60 kHz, and the array meeting the spatial sampling constraints:
[0081]
[0082] When d = 8cm, the highest frequency without aliasing is:
[0083] f alias =c / (2d)=343 / (2×0.08)=2143Hz<<60kHz
[0084] The anti-aliasing filter eliminates high-frequency interference to obtain pure animal sound signals.
[0085] Yes, set the cutoff frequency to 2143Hz and use a Butterworth low-pass filter to eliminate high-frequency aliasing interference.
[0086] Preferably, the joint noise reduction model is used to perform L-level wavelet decomposition and joint optimization processing on the multi-channel audio signal to obtain optimized wavelet coefficients, the frequency band characteristics of the chicken's call are concentrated in 200Hz~2kHz, the wavelet decomposition level L=3 is set, and the cutoff frequency of each low-pass filter is adjusted to 5kHz, 2.5kHz, 1.25kHz to better match the low-frequency characteristics, including:
[0087] Set up a noisy signal model:
[0088] y=x+n
[0089] Among them, x is a clean signal, n~N(0,σ 2 ) is Gaussian noise, y is the observed signal, and the target is y≈x;
[0090] Perform L-level wavelet decomposition on the observed signal, and obtain the approximate coefficient a at each level l and detail factor d l :
[0091]
[0092] Initial value a 0 =y
[0093] is the lth-order low-pass filter with a cutoff frequency of f low ;
[0094] is the first-order high-pass filter with a cutoff frequency of f high ;
[0095] Output the decomposition coefficient set of all channels
[0096] Enter the decomposition coefficient
[0097] Design a joint objective function with multiple levels of constraints:
[0098]
[0099] It is the data fidelity term, which ensures the consistency between the estimated signal and the observed signal; is a sparse regularization term, through l 1 The norm promotes the sparsity of wavelet coefficients;
[0100] is the high-frequency coefficient approximation term, which constrains the high-frequency components of the restored signal to be equal to the high-frequency components of the observed signal. l close to, λ, β are weight parameters, and the high-frequency coefficient d of the observed signal l Contains some real signal components, and retains weak but important details such as texture and weak edges during denoising through approximation constraints;
[0101] Iteratively optimize the joint objective function and decomposition coefficients to update x, z l , μ l ;
[0102] Adopt ADMM (Alternating Direction Method of Multipliers) decomposition optimization and introduce auxiliary variables get:
[0103]
[0104] Where, l = 1, ... L;
[0105] Then, construct the augmented Lagrangian function:
[0106]
[0107] Among them, μl is the Lagrange multiplier, ρ>0 is the penalty parameter;
[0108] Update the variables alternately:
[0109] 1) Update x
[0110] Fixed z l
[0111]
[0112] The closed-form solution is obtained by differentiation:
[0113]
[0114] Update x to ensure that the recovered signal x is close to the observed signal y, forcing the high-frequency components of x Auxiliary variables close to current estimates And through the multiplier Correct deviations, coordinate global signal fidelity with local high-frequency component consistency, and avoid over-reliance on constraints at a certain level;
[0115] 2) Update
[0116] Fix x and solve:
[0117]
[0118] Combine quadratic terms:
[0119]
[0120] in
[0121] Solved by soft threshold shrinkage:
[0122]
[0123] Update z, l 1 The term forces most of the wavelet coefficients to be zero, suppressing noise, and the quadratic term constrains z l Close to the observed high frequency coefficient d l and current estimates
[0124] 3) Update the multiplier μ l
[0125]
[0126] Update μ l , the multiplier μ l Accumulation Constraints The violation of , promotes the subsequent iterations to gradually satisfy the constraints, and increases ρ to accelerate convergence, but it may make the sub-problem difficult to solve;
[0127] Reconstruct the signal by inverse wavelet transform:
[0128]
[0129] in is the optimized wavelet coefficient, k means there are k coefficients in each level.
[0130] Specifically, the spatial frequency domain noise reduction algorithm performs the following steps:
[0131] Generalized Sidelobe Canceller Beamforming:
[0132] w=R -1 a(θ)]a H (θ)R -1 a(θ)] -1
[0133] The noise covariance matrix update period Δt≤100ms;
[0134] Improved spectral subtraction:
[0135] |X enhanced (k)| 2 =max(|Y(k)| 2 -α|N(k)| 2 ,γ|N(k)| 2 ),(α=2.0,γ=0.05)
[0136] Parameter optimization:
[0137] Oversubscription factor α = 2.0: Balance noise suppression and signal distortion,
[0138] Residual factor γ = 0.05: retain weak call components,
[0139] Improve the signal-to-noise ratio by 9.2dB.
[0140] Furthermore, the fusing process of the multi-channel audio signal using the time-frequency feature model and the space feature model includes:
[0141] Time-frequency feature model processing:
[0142] Pre-emphasis and frame windowing, retaining α = 0.95, and adding an additional low-frequency compensation filter to avoid low-frequency distortion caused by high-frequency enhancement in view of the low-frequency characteristics of chicken calls. MFCC is used to pre-emphasize the original time domain signal x(t), and the coverage range of the Mel filter group is adjusted to 100Hz~3kHz, which is more suitable for the energy distribution of chicken calls and enhances the high-frequency components.
[0143] x pre (t)=x(t)-αx(t-1),α∈[0.95,0.97]
[0144] Divide the time domain signal with enhanced high-frequency components into frames, add a Hamming window to each frame,
[0145]
[0146] Where N is the frame length;
[0147] Calculate the amplitude spectrum and Mel filter of each frame of time domain signal, use Fourier transform (FFT) to calculate the amplitude spectrum of each frame to calculate MFCC and Mel spectrum features,
[0148]
[0149] k=0,1,…,N-1
[0150] Through 40 Mel filters H m (k) Smoothing the spectrum:
[0151]
[0152] Calculate MFCC and Mel spectrum features, take logarithm and perform discrete cosine transform (DCT) to get F MFCC :
[0153]
[0154] i=1,2,....,13
[0155] Total dimensions: 13 static coefficients + 13 first-order differences + 13 second-order differences = 39 dimensions,
[0156] Take the output energy of the Mel filter bank to form a 64-dimensional feature:
[0157] F Mel =E(m)
[0158] m=1,2,…,64
[0159] The MFCC and Mel spectrum of the two target sound sources are combined to obtain a 103-dimensional feature vector by concatenating 39-dimensional and 64-dimensional features:
[0160] F fused =[F MFCC ,E Mel ]∈R 103 .
[0161] Spatial feature model processing:
[0162] By extracting the features of poultry calls, we get the feature information, parameters and simplified original waveform sampling signals, and obtain the time-frequency feature mapping and spatial feature polar space mapping. For the chicken house environment, the sound source is close, usually <5m, so the covariance matrix update period of the MUSIC algorithm is adjusted to 50ms to improve the real-time positioning performance, and the distance error is controlled within 5m to 2%;
[0163] MUSIC spatial spectrum calculation:
[0164]
[0165] Resolution test:
[0166] 10kHz signal, angular resolution up to 0.8° (S / N ratio 15dB)
[0167] Sound source distance error: When the distance is 10m, the sound source distance error is less than 5%;
[0168] Generalized cross-correlation and time delay estimation: Assume that the multi-channel audio signal of the sound source is:
[0169] X=A(θ)S+N
[0170] Where A(θ)=[a(θ 1 ),…,a(θ K )] is the guidance matrix,
[0171] a(θ)=[1,e -j2πdsinθ / λ ,…,e -j2π(M-1)dsinθ / λ ] T ;
[0172] Perform covariance matrix decomposition on the signal:
[0173]
[0174] Among them U n is the noise subspace;
[0175] Using spatial spectrum calculation:
[0176]
[0177] For the i,j microphone pair, calculate the generalized cross-correlation (GCC-PHAT):
[0178]
[0179] Take the peak position T ij As a delay estimate, construct a spatial contrast matrix:
[0180] C=[T ij ] M×M :
[0181] Fusion processing:
[0182] Time-frequency feature map (channel 1):
[0183] The 39-dimensional MFCC and the 64-dimensional Mel spectrum are concatenated into a 103-dimensional vector and arranged by time frame into a matrix FTF∈R T×103 , expanded to 128×128 by bilinear interpolation:
[0184]
[0185] Spatial feature polar coordinate mapping (channel 2):
[0186] Step 1: Convert the spatial spectrum to polar coordinates and convert P MUSIC (θ) is discretized into 128 directions according to the angle θ∈[0°,360°), mapped to the polar coordinate grid (r,θ), with a fixed radius r=1;
[0187] Step 2: Spatial contrast matrix encoding, transform the delay τ of C ij Converted to angle difference Δθ ij Fill in the corresponding position of the polar coordinate grid;
[0188] Step 3: Interpolate to generate an image, fill the sparse polar coordinate data into a 128×128 matrix F through nearest neighbor interpolation Space ;
[0189] The time-frequency feature matrix F of the target sound source TF With the spatial feature matrix F Space Consistent, through the bilinear interpolation adjustment method, F TF The size of is expanded to 128×128 matrix, and we get
[0190] The bilinear interpolation adjustment method is expressed as:
[0191]
[0192] The purpose is to adjust the size of the time-frequency feature matrix to be consistent with the spatial feature matrix.
[0193] After feature fusion, the time-frequency feature matrix And the spatial feature matrix F Space Concatenated along the channel dimension to form a 128×128×2 feature map F Fusion
[0194] Feature map concatenation:
[0195] Concatenate time-frequency and spatial features along the channel dimension:
[0196]
[0197] Among them, 128 is the time frame, the frequency point is 128, and the feature dimension is 2.
[0198] Preferably, the deep learning model is used to process the time-frequency-space feature signal and the optimized wavelet coefficients to obtain the target audio signal, using a labeled data set containing the calls of poultry such as chickens, ducks, and geese, and adding common noises in chicken houses, such as fan sounds and feeder vibration sounds, including:
[0199] CNN branch: Input time-frequency-space features into the deep learning model to extract local time-frequency-space characteristics;
[0200] Since the duration of chicken calls is short, about 0.5 to 1 second, the frame length is shortened to 20 ms to capture more detailed time-varying features;
[0201] FCN branch: Input optimized wavelet coefficients into the deep learning model to capture the characteristics of the target sound source;
[0202] Feature fusion layer: The time-frequency-spatial features of the CNN branch and the features of the target sound source of the FCN branch are merged to output the target audio signal.
[0203] CNN branches:
[0204] From the time-frequency feature F Fusion Extract high-level features from
[0205] Because the 2-layer convolution kernel is R 128×128×2 , 2 is the number of output channels
[0206] Output the feature map h 1 =F Fusion , b l For bias top;
[0207] Use ReLU activation function and then perform maximum pooling layer
[0208] Finally, after multiple layers of convolution, activation and pooling, high-order time-frequency features are obtained. Where D 1 is the dimension after flattening.
[0209] FCN Branch:
[0210] Design the weight of layer 2 as Bias is
[0211] The output is:
[0212] Initial input
[0213] Post-activation function
[0214] Use Sigmoid function to enhance nonlinearity:
[0215]
[0216] Finally, after multiple layers of full connection, the global minimum wave feature is obtained
[0217] Finally, the two features are combined:
[0218] The two branches are weighted:
[0219] h=σ(W C h 1 +W ω h 2 +b)
[0220] in is a learnable parameter
[0221] Where σ(z) is usually the Tanh function:
[0222]
[0223] Output layer:
[0224] The target signal dimension is obtained by mapping through the fully connected layer:
[0225]
[0226] in b o ∈R N , N is the target signal length.
[0227] Then optimize it
[0228] Objective function:
[0229] Minimize the prediction signal The mean square error with the true signal y
[0230]
[0231] Backward Propagation:
[0232] Calculate the gradient by chain rule and update W C ,W O ,W ω etc:
[0233]
[0234] Similarly, the gradients are passed back layer by layer to the CNN and FCN branches.
[0235] Then the optimization algorithm is performed;
[0236] Using Adam optimizer, the update rule is:
[0237]
[0238] in, are bias-corrected first- and second-order moment estimates.
[0239] Preferably, before the fusing process of the multi-channel audio signal by using the time-frequency feature model and the space feature model, the method further comprises:
[0240] Input a multi-channel audio signal, filter it through the A-weighted transfer function, and get the amplitude of the filtered signal.
[0241]
[0242] Through A-weighted filtering, the sound signal in the human ear sensitive frequency band (1-4kHz) can be highlighted, which is in line with the human auditory perception characteristics;
[0243] The amplitude of the filtered signal is adjusted through dynamic range compression to obtain a signal with amplitude standardization, so that the dynamic range of the signal can meet the needs of subsequent processing.
[0244]
[0245] Performance Verification:
[0246] Transient response: attack time 20ms, release time 500ms
[0247] Total harmonic distortion: <0.5%@1kHz sine input
[0248] When the sine input is 1kHz, the total harmonic distortion is less than 0.5%;
[0249] The frequency domain signal energy after A-weighted filtering is processed through A-weighted sound pressure level calculation to obtain sound pressure level data for dynamic compression parameter adjustment. The frequency response of the A-weighted filter complies with the IEC61672 standard, and its amplitude-frequency characteristic is defined as:
[0250]
[0251] Simulate the attenuation characteristics of the human ear for low and high frequencies, highlight the sensitive frequency band of the human ear 1-4kHz, signal filtering and energy integration
[0252] Apply A-weighted filtering to the input sound pressure signal S(f) (frequency domain representation) and calculate the weighted energy:
[0253]
[0254] Where S(f)=F{s(t)} is the Fourier transform of the time domain signal s(t);
[0255] Sound Pressure Level (SPL) Calculation
[0256] With reference sound pressure p ref =20μPa as the benchmark, calculate the A-weighted sound pressure level:
[0257]
[0258] Derivation and verification:
[0259] RMS value of sound pressure therefore:
[0260]
[0261] Dynamic range compression function
[0262] Gain Curve Definition
[0263] The dynamic range compression function G(L) adaptively adjusts the gain according to the current sound pressure level L:
[0264]
[0265] Among them, L th is the compression threshold, which is 60dB in this embodiment. When it is lower than this threshold, the maximum gain G is used. max Amplify weak signals,
[0266] G min : minimum gain,
[0267] k: Compression ratio slope, which controls the sensitivity of gain to changes in sound pressure level;
[0268] Kalman filter estimation of slope parameter k
[0269] State equation: Assuming the ambient noise level L noise Slowly varying, modeled as a first-order Markov process:
[0270]
[0271] ω (t) ~N(0,Q)
[0272] Observation equation: through the current sound pressure level Observation noise level:
[0273]
[0274] v (t)~N(0,R)
[0275] Finally, the sound pressure level data is processed through Kalman filter noise tracking to obtain the noise-reduced sound pressure level data. The Kalman filter realizes the real-time estimation of the noise level of bird calls through the prediction-update mechanism. Combined with the dynamically adjusted compression slope k, the system can adapt to environmental changes and balance denoising and signal fidelity.
[0276] predict:
[0277]
[0278] P (t∣t-1) =P (t-1∣t-1) +Q
[0279] State equation assumption: Ambient noise level L noise The changes are slow and can be modeled as a first-order Markov process:
[0280]
[0281] w (t) ~N(0,Q)
[0282] Among them, w (t) is the process noise, reflecting the random fluctuations in the noise level,
[0283] Prediction state: Since the noise level is assumed to change slowly, the estimated value of the previous moment is directly inherited Ignore the mean of the process noise (mean is zero).
[0284] Prediction covariance: Covariance P (t|t-1) Increasing the process noise variance Q, indicating that the uncertainty of the prediction accumulates over time;
[0285] The noise-reduced sound pressure level data is optimized through the loss function until the bit error rate meets the preset conditions.
[0286] renew:
[0287]
[0288] P (t∣t) =(1-K (t) ) (t∣t-1)
[0289] Observation model: Assuming sound pressure level measurements is the sum of the true noise level and the observed noise:
[0290]
[0291] v (t) ~N(0,R)
[0292] Kalman gain K (t) : Balance the weights of predictions and observations:
[0293] If the observation noise R is small (the measurement is accurate), K (t) →1, trust new observations more
[0294] If the prediction covariance P (t|t-1) Large (high prediction uncertainty), K (t) →1, more dependent on observation correction.
[0295] Status update: via residual Corrected predicted value, residual weight is L (t) Decide.
[0296] Covariance update: updated covariance P (t|t) decreases, reflecting the reduction in uncertainty due to new observations.
[0297] Adaptive adjustment of slope k:
[0298] The estimated noise level Dynamically adjust the compression ratio:
[0299]
[0300] ΔL=L max -L th
[0301] Where L max The maximum sound pressure level allowed by the system, such as 90dB;
[0302] The slope k determines the rate at which the gain decreases with increasing sound pressure level. The smaller k is, the more gradual the gain decreases, and the more gentle the dynamic range compression is. The larger k is, the steeper the gain decreases, and the more aggressive the compression is.
[0303] Jointly optimize acoustic recognition (such as classification loss L class ) and physical parameter regression (such as sound pressure level error Azimuth error L θ ):
[0304]
[0305] Weight allocation: Determine α, β, γ, λ through grid search or automatic parameter adjustment;
[0306] Bit error rate analysis:
[0307] Dynamic loudness normalization reduces the wireless transmission bit error rate through the following mechanism: Since the dynamic range of chicken calls is narrow, 40-80dB, the compression threshold L is adjusted. th=60dB, compression slope k = 0.5, to avoid feature loss caused by over-compression:
[0308] Amplitude stability: Compresses the dynamic range to avoid signal clipping or a drop in signal-to-noise ratio (SNR).
[0309] Noise suppression: Kalman filtering tracks environmental noise in real time and increases the proportion of effective signals.
[0310] Quantization consistency: The normalized signal amplitude adapts to the linear operating range of the ADC / DAC.
[0311] Bit Error Rate (BER) Theoretical Model:
[0312]
[0313] Among them, E b / N 0 is the signal-to-noise ratio, E after dynamic compression b / N 0 Improve by at least 3dB and reduce BER to 10 -5 the following;
[0314] The amplitude of the chicken call signal is normalized to [-1,1], adapted to the 12-bit ADC quantization range, and the bit error rate (BER) is reduced to 5×10 -6 , meeting the low-power IoT transmission requirements.
[0315] An electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other through the communications bus. The processor may call logic instructions in the memory to execute an animal call precision recognition method of an integrated animal call sensor, and the packaging structure adopts an aluminum alloy linear cavity and a built-in temperature and humidity compensation sensor to protect each module and provide a stable working environment.
[0316] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0317] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the animal call accurate identification method of the integrated animal call sensor provided by the above-mentioned methods.
[0318] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the animal call accurate recognition method of the integrated animal call sensor provided by the above-mentioned methods.
[0319] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0320] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0321] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for accurately identifying animal sounds using an integrated animal sound sensor, characterized in that: include: Acquire a multi-channel audio signal of a sound source, and use a joint noise reduction model to perform L-level wavelet decomposition joint optimization processing on the multi-channel audio signal to obtain optimized wavelet coefficients; The multi-channel audio signal is fused using a time-frequency feature model and a space feature model to obtain a time-frequency space feature signal; The time-frequency-space feature signal and the optimized wavelet coefficients are processed using a deep learning model to obtain a target audio signal.
2. The method for accurately identifying animal sounds using an integrated animal sound sensor according to claim 1, characterized in that: The method of performing L-level wavelet decomposition and joint optimization processing on the multi-channel audio signal using the joint noise reduction model to obtain optimized wavelet coefficients includes: Set up a noisy signal model: y=x+n Among them, x is a clean signal, n~N(0,σ 2 ) is Gaussian noise, y is the observed signal; Perform L-level wavelet decomposition on the observed signal, and obtain the approximate coefficient a at each level l and detail factor d l : Initial value a0 = y is the lth-order low-pass filter with a cutoff frequency of f low ; is the first-order high-pass filter with a cutoff frequency of f high ; Output the decomposition coefficient set of all channels Enter the decomposition coefficient Design a joint objective function with multiple levels of constraints: It is the data fidelity term, which ensures the consistency between the estimated signal and the observed signal; is a sparse regularization term, which promotes the sparsity of wavelet coefficients through the l1 norm; is the high-frequency coefficient approximation term, which constrains the high-frequency components of the restored signal to be equal to the high-frequency components of the observed signal. l Close to, λ, β are weight parameters, and the high-frequency coefficient d of the observed signal l Contains some real signal components; Iteratively optimize the joint objective function and decomposition coefficients to update x, z l , μ l ; Reconstruct the signal by inverse wavelet transform: in is the optimized wavelet coefficient, k means there are k coefficients in each level.
3. The method for accurately identifying animal sounds using an integrated animal sound sensor according to claim 1, characterized in that: The multi-channel audio signal is subjected to fusion processing by using a time-frequency feature model and a space feature model, and the time-frequency feature model processing includes: Pre-emphasize the original time domain signal x(t) to enhance the high-frequency component, divide the time domain signal with enhanced high-frequency component into frames, and add a Hamming window to each frame; Calculate the amplitude spectrum and Mel filter of each frame of the time domain signal, and calculate the MFCC and Mel spectrum features.
4. The method for accurately identifying animal sounds using an integrated animal sound sensor according to claim 3, characterized in that: The multi-channel audio signal is fused using a time-frequency feature model and a spatial feature model, wherein the spatial feature model processing includes: Set the multi-channel audio signal of the sound source to: X=A(θ)S+N Among them, A(θ)=[a(θ1),…,a(θ K )] is the guidance matrix, Perform covariance matrix decomposition on the signal: Among them U n is the noise subspace; Using spatial spectrum calculation: For the i,j microphone pair, calculate the generalized cross-correlation (GCC-PHAT): Take the peak position T ij As a delay estimate, construct a spatial contrast matrix: C=[T ij ] M×M 。 5. The method for accurately identifying animal sounds using an integrated animal sound sensor according to claim 4, characterized in that: The multi-channel audio signal is subjected to fusion processing by using a time-frequency feature model and a space feature model, and the fusion processing includes: Time-frequency feature mapping: The 39-dimensional MFCC and the 64-dimensional Mel spectrum are concatenated into a 103-dimensional vector and arranged by time frame into a matrix FTF∈R T×103 , expanded to 128×128 by bilinear interpolation: Polar coordinate mapping of spatial features: Step 1: Convert the spatial spectrum to polar coordinates and convert P MUSIC (θ) is discretized into 128 directions according to the angle θ∈[0°,360°), mapped to the polar coordinate grid (r,θ), with a fixed radius r=1; Step 2: Spatial contrast matrix encoding, transform the delay τ of C ij Converted to angle difference Δθ ij Fill in the corresponding position of the polar coordinate grid; Step 3: Interpolate to generate an image, fill the sparse polar coordinate data into a 128×128 matrix F through nearest neighbor interpolation Space ; Feature map concatenation: Concatenate time-frequency and spatial features along the channel dimension: Among them, 128 is the time frame, the frequency point is 128, and the feature dimension is 2.
6. The method for accurately identifying animal sounds using an integrated animal sound sensor according to claim 1, characterized in that: The method of using a deep learning model to process the time-frequency-space feature signal and the optimized wavelet coefficients to obtain a target audio signal includes: CNN branch: Input time-frequency-space features into the deep learning model to extract local time-frequency-space characteristics; FCN branch: Input optimized wavelet coefficients into the deep learning model to capture the characteristics of the target sound source; The time-frequency-spatial features of the CNN branch and the features of the target sound source of the FCN branch are fused to output the target audio signal.
7. The method for accurately identifying animal sounds using an integrated animal sound sensor according to claim 1, characterized in that: Before the fusion processing of the multi-channel audio signal using the time-frequency feature model and the space feature model, the method further includes: Input a multi-channel audio signal, filter it through the A-weighted transfer function, and obtain the amplitude of the filtered signal; The amplitude of the filtered signal is adjusted by dynamic range compression to obtain a signal with normalized amplitude; The frequency domain signal energy after A-weighted filtering is processed by A-weighted sound pressure level calculation to obtain sound pressure level data for dynamic compression parameter adjustment; Finally, the sound pressure level data is processed by Kalman filter noise tracking to obtain the noise-reduced sound pressure level data; The noise-reduced sound pressure level data is optimized through the loss function until the bit error rate meets the preset conditions.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for accurately identifying animal calls using the integrated animal call sensor according to any one of claims 1 to 7 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for accurately identifying animal calls using the integrated animal call sensor according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for accurately identifying animal calls using the integrated animal call sensor according to any one of claims 1 to 7 are implemented.