A method for sound source localization of a bionic owl asymmetric ear structure
By incorporating a biomimetic owl-like asymmetrical ear structure and multimodal signal processing, combined with deep learning algorithms, the problem of insufficient sound source localization accuracy in small arrays was solved, achieving high-precision sound source localization and dynamic tracking in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2026-06-18
- Publication Date
- 2026-07-21
AI Technical Summary
Existing sound source localization technologies suffer from low resolution in small arrays, strong environmental noise interference, and slow dynamic sound source tracking, especially in the low-frequency band where localization accuracy is insufficient. Furthermore, traditional bionic models have failed to achieve end-to-end collaborative design.
By adopting a biomimetic owl asymmetrical ear structure, an asymmetrical microphone array is constructed. Combining multimodal signal processing and deep learning algorithms, the joint estimation and dynamic tracking of sound source azimuth and pitch angle are achieved through auricle-guided multi-scale wavelet denoising, spatial adaptive filtering, and extended Kalman filtering.
It significantly improves spatial resolution and positioning accuracy under small array conditions, reduces noise interference, and achieves end-to-end three-dimensional sound source localization, making it suitable for complex environments and dynamic sound source scenarios.
Smart Images

Figure CN122430792A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source localization, specifically to a sound source localization method based on a biomimetic owl asymmetrical ear structure. Background Technology
[0002] With the rapid development of drone technology, the application of drones has experienced explosive growth, but this has also brought many significant regulatory and governance challenges. The safety management of low-altitude airspace faces unprecedented pressure and challenges. However, traditional drone detection methods, such as radar detection, radio spectrum detection, and optical detection, suffer from problems such as large blind spots, strong susceptibility to environmental interference, and limited identification capabilities when dealing with low-altitude, low-speed, and small drone targets. Therefore, drone detection and identification technology based on acoustic characteristics, with its advantages of passive detection, immunity to electromagnetic interference, all-weather operation, and target identification potential, has gradually become a research hotspot and a key area for technological breakthroughs in the field of low-altitude security.
[0003] Currently, research on bionic acoustics has obvious limitations: on the one hand, existing bionic models only simulate a single asymmetry and do not consider the multidimensional asymmetry of the biological auditory system; on the other hand, research on bionic acoustic localization usually optimizes hardware design and signal processing algorithms separately, failing to achieve end-to-end collaborative design. At the same time, existing localization technologies perform well in the high-frequency band, but their localization accuracy is insufficient in the low-frequency band (<500Hz), while many important sound sources (such as mechanical fault sounds and specific biological sounds) contain rich low-frequency components.
[0004] Therefore, researching a biomimetic sound source localization technology is of great significance for the practicality and deployability of identifying and locating drones. Summary of the Invention
[0005] This invention aims to solve the problems of low resolution of small arrays, strong environmental noise interference, and slow dynamic sound source tracking in existing sound source localization technologies, and provides a multimodal sound source localization method based on the asymmetrical auricle structure of a biomimetic owl.
[0006] This invention solves the above-mentioned technical problems through the following technical solution: a sound source localization method based on a biomimetic owl asymmetrical ear structure, the method comprising the following steps:
[0007] S1. Construct a left ear assembly and a right ear assembly with asymmetric acoustic transmission characteristics, arrange microphone arrays in the left ear assembly and the right ear assembly respectively, and perform amplitude calibration, delay compensation and time synchronization on the sampling data of each microphone channel;
[0008] S2. Perform noise suppression, spatial filtering, and frequency domain feature extraction on the multi-channel acoustic signals acquired by the microphone array to obtain preprocessed acoustic signals;
[0009] S3. Calculate the arrival time difference between the left ear component and the right ear component and between the microphone channels based on the preprocessed acoustic signal, and obtain the initial estimation result of the sound source location by combining the asymmetric acoustic transmission characteristics of the left ear component and the right ear component.
[0010] S4. Based on the initial estimation result of the sound source orientation, control the bionic head to rotate, so that the microphone array forms a changing positioning baseline relative to the sound source, and record the head posture information and multi-frame arrival time difference data during the rotation process.
[0011] S5. Input the multi-frame arrival time difference data, frequency domain features, head posture information and asymmetric acoustic transmission features into the sound source localization neural network model, and output the azimuth angle estimation result, pitch angle estimation result and location confidence of the sound source.
[0012] S6. The extended Kalman filter algorithm is used to fuse the azimuth estimation results, pitch estimation results, position confidence and time series observation data to obtain the three-dimensional position and motion trajectory of the dynamic sound source.
[0013] The positive and progressive effects of this invention are as follows:
[0014] By incorporating the asymmetrical auricle structure of a biomimetic owl and an active head rotation mechanism, the spatial resolution capability is significantly enhanced under small array conditions, and the accuracy of joint estimation of azimuth and pitch angles is greatly improved, breaking through the physical constraints of traditional uniform arrays when size is limited. Through auricle-guided multi-scale wavelet denoising and spatial adaptive filtering, robust positioning performance is maintained in complex reverberation and noise environments, without the need for additional auxiliary sound sources or knowledge of the sound field environment throughout the process.
[0015] By fusing frequency domain-time difference of arrival hybrid features with a CNN-BiLSTM deep learning model, an end-to-end mapping from acoustic signals to three-dimensional spatial coordinates is achieved, significantly reducing the localization error compared to the traditional GCC-PHAT method. Combined with temporal trajectory tracking using extended Kalman filtering, smooth and continuous motion estimation is achieved in dynamic sound source scenarios. This method is compatible with mobile platforms such as drones and robots, and does not require the deployment of fixed microphone arrays or dedicated anechoic environments. It can be directly integrated into intelligent security, low-altitude detection, and robot hearing systems, meeting the real-time deployment requirements under complex field conditions. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method of the present invention.
[0017] Figure 2 This is a box plot showing the time difference of arrival estimation error under different noise environments according to the present invention.
[0018] Figure 3This is a schematic curve illustrating the effect of the bionic head rotation strategy of this invention on the optimization of the positioning baseline.
[0019] Figure 4 This is a comparison chart of the three-dimensional trajectory error (RMSE) of the CNN-BiLSTM+EKF fusion localization results of this invention.
[0020] Figure 5 This is a bar chart comparing the positioning accuracy of the present invention with that of traditional microphone array positioning methods in multiple scenarios. Detailed Implementation
[0021] The present invention will be further illustrated by way of embodiments below, but the present invention is not limited to the scope of the embodiments.
[0022] See Figures 1 to 5 A sound source localization method based on a biomimetic owl asymmetrical ear structure, the method comprising the following steps:
[0023] S1. Construct a left ear assembly and a right ear assembly with asymmetric acoustic transmission characteristics. Arrange microphone arrays in the left ear assembly and the right ear assembly respectively, and perform amplitude calibration, delay compensation, and time synchronization on the sampling data of each microphone channel. Specifically, this includes:
[0024] S11. Bionic auricle structure design and modeling
[0025] Based on the natural asymmetrical ear canal height difference, cavity shape difference, and external auricle refraction angle difference of owls, this invention constructs a biomimetic structure with three-dimensional geometric differences between the left and right auricles. The height difference between the left and right ear canals is set as follows: ;
[0026] The height of the left ear canal. The height of the right ear canal is determined by setting an auricular refractive surface, which makes the propagation path of the incident sound wave different in the left and right ear canals, thereby generating a phase difference and amplitude difference with enhanced directionality and improving the vertical positioning accuracy.
[0027] S12. The microphone array is embedded inside the auricle in a circular arrangement, and the array spacing is selected according to the target frequency range.
[0028] Circular array geometric model:
[0029] Assume the microphone array radius is r, and the 8 microphones are evenly distributed:
[0030] The coordinates of microphone m: ;
[0031] in, For the first The three-dimensional spatial coordinates of the microphone. The radius of the circular microphone array arrangement. For the first The circumferential azimuth angle corresponding to each microphone. Number the microphones;
[0032] Microphone array arrangement and synchronous acquisition: N microphones are arranged inside the auricle, and their position vectors are denoted as: ;
[0033] The left and right lateral axes, The front and rear depth axes, The microphones are synchronized via a hardware clock along the vertical axis, with a sampling frequency of fs. After gain calibration and phase compensation, a multi-channel acoustic acquisition matrix is formed. .
[0034] S2. Calculate the arrival time difference between the left and right ear components and between the microphone channels based on the preprocessed acoustic signal, and obtain an initial estimation result of the sound source location by combining the asymmetric acoustic transmission characteristics of the left and right ear components, specifically including:
[0035] S21. Set the spatial filtering window to suppress environmental reverberation through array spatial correlation;
[0036] Array covariance matrix:
[0037] ;
[0038] in In the formula, This is the frequency domain autocovariance matrix of the received signal. Angular frequency, This is the frequency domain vector of the multi-channel microphone signal. This is the conjugate transpose. For mathematical expectation;
[0039] Spatial filter settings:
[0040] In the formula, Find the inverse of the noise covariance matrix. is the array steering vector, and s is the spatial orientation parameter of the target sound source;
[0041] in As the guiding vector, For noise covariance matrix estimation, Let be the angular loading factor, where For matrix trace, For the number of microphones, For coefficients, The imaginary unit, For the source of students To the Mach geometric propagation delay, This is the matrix transpose.
[0042] The final filtered signal is:
[0043] In the formula These are the adaptive spatial domain filtering weight coefficients;
[0044] Adaptive noise suppression and spatial filtering
[0045] Using the array spatial correlation matrix: ;
[0046] Constructing spatial filter weights: , where d is a pointing vector;
[0047] Output spatial filtering signal: This achieves suppression of background noise and reflected sound;
[0048] S22. Multi-scale wavelet packet decomposition is used, and soft thresholding is applied to the coefficients of each level to reconstruct and obtain the denoised signal.
[0049] The signal is decomposed into L layers using Daubechies wavelets: ;
[0050] Perform soft thresholding denoising on each layer: ;
[0051] The denoised signal is obtained by reconstruction. ;
[0052] from Extract features such as Mel-frequency cepstral coefficients (MFCC) and spectral centroid to construct a frequency domain matrix:
[0053] .
[0054] Discrete signal Decompose to level J:
[0055] ;
[0056] in, For wavelet basis functions, It is the first Discrete-time sampling signal from the microphone. This represents the total number of wavelet decomposition layers. for and location wavelet detail coefficients;
[0057] Soft thresholding:
[0058] In the formula For scale j and position Wavelet coefficients after thresholding For adaptive threshold, This refers to the number of sampling points per frame. ;
[0059] S23. Screen reliable frequency domain bandwidths using characteristic stability indices;
[0060] Characteristic stability index:
[0061] In the formula For frequency stability index, for The conjugate of complex numbers, For minute frequency offsets, values range from 0 to 1;
[0062] Reliable bandwidth selection:
[0063] ;
[0064] in , , In the formula These are the reliable characteristic frequency bands after screening.
[0065] S3. Based on the initial estimation result of the sound source location, control the bionic head to rotate, so that the microphone array forms a changing positioning baseline relative to the sound source, and record the head posture information and multi-frame arrival time difference data during the rotation process, specifically including:
[0066] S31, GCC-PHAT cross-correlation delay estimation
[0067] For any microphone pair calculate:
[0068] ;
[0069] The arrival time difference is the delay corresponding to the peak value of the cross-correlation. ;
[0070] Computational geometric delay:
[0071] ;
[0072] in, Indicates the sound source. Let m be the Cartesian coordinates of the m-th microphone. For the speed of sound, This represents the time required for the sound source to reach the m-th microphone. As the sound source Euclidean distance from the m-th microphone.
[0073] Calculate the time delay difference for each pair of microphones based on the time difference of arrival:
[0074] ;
[0075] Then GCC-PHAT:
[0076] ;
[0077] In the formula, This is represented as the Fourier transform of the time-domain signal of the m-th microphone signal. Expressed as angular frequency, GCC-PHAT serves as the core calculation tool for estimating the time difference of arrival for each pair of microphones. For time delay variables, It is a complex exponential function. It is the imaginary unit.
[0078] SRP-PHAT spatial spectrum:
[0079] ;
[0080] In the formula, is the spatial spectral amplitude of the candidate sound source position s, and M is the total number of microphones;
[0081] Based on the calculation results of GCC-PHAT, find the most likely location of the sound source in space;
[0082] S32: The effective peak value is determined by local peak statistics, sliding window fitting and abnormal peak suppression methods;
[0083] Traditional peak positioning: ;
[0084] Sliding peak positioning: ;
[0085] In the formula, The smoothing weighting coefficient determines whether response or accuracy is prioritized during tracking and localization; n is the frame number. For spatial search networks;
[0086] Auricular structure correction model: Due to the difference in refraction paths in the auricle, this invention introduces a bionic auricular path compensation term: ,in, It was obtained by biomimetic three-dimensional structure simulation of the auricle;
[0087] S33: A robust estimate is obtained by establishing a modified arrival time difference equation by combining the spatial transfer function (HRTF) of the auricle structure;
[0088] Geometric delay of incorporating biomimetic structures:
[0089] ;
[0090] in, , Let represent the frequency domain transfer function from the sound source location s to the m-th microphone. for phase response, Additional phase delay compensation for auricular asymmetry;
[0091] The GCC function after incorporating the biomimetic structure:
[0092] ;
[0093] For the corrected cross-power spectrum, For frequency structure weighting function, The asymmetric coefficient, , It is a frequency-weighted function. , , It is a very small constant to avoid dividing the denominator by zero;
[0094] The corrected spatial spectrum C-SRP_PHAT is then:
[0095] ;
[0096] In the formula, The function is the biomimetic modified GCC-PHAT cross-correlation function, where s is the coordinates of the spatial candidate sound source. This is the geometrically theoretical time delay difference;
[0097] The location of the sound source was ultimately determined using the sliding peak localization method.
[0098] Preliminary azimuth angle solution: sound source direction vector satisfy ;
[0099] Construct an overdetermined system of linear equations: ;
[0100] The preliminary horizontal angle is obtained by using least squares. With pitch angle .
[0101] S4. Input the multi-frame time-of-arrival data, frequency domain features, head posture information, and asymmetric acoustic transmission features into the sound source localization neural network model, and output the azimuth angle estimation result, pitch angle estimation result, and location confidence of the sound source, specifically including:
[0102] S41. Calculate the current rotation direction and rotation amplitude based on the estimated azimuth angle value at the previous moment;
[0103] Rotation strategy:
[0104] Let the estimated azimuth angle at time t be . The pitch angle estimate is ;
[0105] ;
[0106] ;
[0107] In the formula, / The current mechanical posture angle of the bionic head. / For each rotation angle required, / This is the maximum limit angle for a single rotation. Determine whether the direction of rotation is positive or negative;
[0108] Active rotation model: The bionic head rotation posture is represented by a rotation matrix. ;
[0109] in, Let yaw, pitch, and roll be the angles of the kth active rotation.
[0110] S42. Utilize an adaptive rotational gain control strategy to gradually reduce the rotational amplitude at near-field sound sources;
[0111] Distance estimation:
[0112] ;
[0113] In the formula, The median time delay is used to prevent outliers. Microphone spacing;
[0114] Adaptive gain function:
[0115] ;
[0116] In the formula, As the reference gain, Hyperbolic tangent function;
[0117] Update rotation amount:
[0118] ;
[0119] ;
[0120] Dynamic baseline enhancement: The microphone positions of the array are updated after rotation as follows: ;
[0121] Multi-pose accumulation can significantly expand the discriminable space, making the time difference of arrival solution more stable.
[0122] S43, Multi-pose Fusion Time Difference of Arrival Dataset
[0123] Composition sequence: .
[0124] S5. The extended Kalman filter algorithm is used to fuse the azimuth estimation results, elevation estimation results, positional confidence, and time series observation data to obtain the three-dimensional position and motion trajectory of the dynamic sound source, specifically including:
[0125] S51. Use CNN technology to extract local structures from the two-dimensional feature map of frequency domain and time difference of arrival;
[0126] CNN Feature Extraction:
[0127] Convolutional layers:
[0128] ;
[0129] In the formula, This refers to the sequence number of the network's convolutional layer. For convolution kernel weights, For bias, For activation function, The input feature map is t, where t is the time dimension. For frequency dimension, / The length and width of the convolution kernel;
[0130] Multi-scale fusion convolution:
[0131] ;
[0132] In the formula, For channel splicing, / / After extracting fine, medium, and coarse-grained acoustic features separately through multi-scale convolution, they are fused together.
[0133] CNN-BiLSTM orientation estimation model: Input consists of multi-frame time difference of arrival features and frequency domain features. ;
[0134] CNN extracts local features: ;
[0135] BiLSTM modeling of time series: ;
[0136] Output azimuth estimate: ;
[0137] S52. Combining BiLSTM to model the temporal relationship between multiple frames of acoustic observations;
[0138] Forward LSTM: The time sequence is from front to back;
[0139] Backward LSTM: The time sequence is from back to front;
[0140] Two-way integration: Forward and backward feature splicing, fusion of vertical and future frame temporal information;
[0141] Extended Kalman Filter (EKF) 3D Trajectory Tracking: State Vector: ;
[0142] The observation model maps the estimated angle as follows:
[0143] ;
[0144] EKF executes according to the forecast-update process:
[0145] ;
[0146] ;
[0147] Final output: 3D coordinates of the sound source ;
[0148] S53: The network outputs the azimuth, elevation, and confidence of the sound source, and uses a combination of cross-entropy loss and angle regression loss for training.
[0149] S6. The construction of the extended Kalman filter includes:
[0150] S61: Construct a three-dimensional motion model of the sound source;
[0151] Motion model:
[0152] Let the state vector be: , For three-dimensional position, For speed, For acceleration;
[0153] Equations of state:
[0154] ;
[0155] in , To accelerate the decay factor, It is a third-order identity matrix. For frame interval, The motion attenuation coefficient, The acceleration decay factor, Here is the motion state transition matrix;
[0156] S62: Construct the observation equation based on the CNN-BiLSTM output;
[0157] Observation vector: ;
[0158] Nonlinear observation equations:
[0159] , For the observation vector, It is a nonlinear geometric mapping function. To observe noise;
[0160] in ;
[0161] In the formula, To convert the horizontal azimuth angle from coordinates, To convert the pitch angle, The straight-line distance between the sound source and the bionic head;
[0162] S63: Dynamic sound source trajectory estimation is achieved using covariance propagation, Jacobian linearization, and iterative updates;
[0163] Prediction steps:
[0164] The current location of the sound source is predicted using the state of the previous moment;
[0165] Predicting covariance For process noise covariance;
[0166] Jacobian matrix:
[0167] To observe the first-order partial derivative of dust control at the prediction point, the nonlinear equation is linearized.
[0168] Update steps:
[0169] Kalman gain To observe the noise covariance;
[0170] Correct the state to obtain the optimal three-dimensional coordinates;
[0171] The updated covariance.
[0172] The three-dimensional positioning in S6 includes:
[0173] The three-dimensional positioning equation is: ;
[0174] in, To extend the global coordinates estimated by the Kalman filter, This is the head pose rotation matrix. It is the inverse matrix. To reposition the biomimetic head, The local coordinates of the sound source are obtained by the extended Kalman filter solution.
[0175] By incorporating the asymmetrical auricle structure of a biomimetic owl and an active head rotation mechanism, the spatial resolution capability is significantly enhanced under small array conditions, and the accuracy of joint estimation of azimuth and pitch angles is greatly improved, breaking through the physical constraints of traditional uniform arrays when size is limited. Through auricle-guided multi-scale wavelet denoising and spatial adaptive filtering, robust positioning performance is maintained in complex reverberation and noise environments, without the need for additional auxiliary sound sources or knowledge of the sound field environment throughout the process.
[0176] By fusing frequency domain-time difference of arrival hybrid features with a CNN-BiLSTM deep learning model, an end-to-end mapping from acoustic signals to three-dimensional spatial coordinates is achieved, significantly reducing the localization error compared to the traditional GCC-PHAT method. Combined with temporal trajectory tracking using extended Kalman filtering, smooth and continuous motion estimation is achieved in dynamic sound source scenarios.
[0177] This method is compatible with mobile platforms such as drones and robots, and does not require the deployment of fixed microphone arrays or a dedicated anechoic environment. It can be directly integrated into intelligent security, low-altitude detection and robot hearing systems to meet the real-time deployment needs under complex field conditions.
[0178] See Figure 1 It fully demonstrates the entire process of signal acquisition, preprocessing, TDOA solving, bionic rotating head baseline expansion, CNN-BiLSTM angle measurement, EKF trajectory filtering, and 3D coordinate output. It connects hardware bionic structure, multi-layer signal processing, deep learning, and filtering algorithms. Its solution has a closed-loop logic, realizing the collaborative optimization of structural design and signal algorithm, and performing layer-by-layer noise reduction and correction from signal source localization to output.
[0179] See Figure 2The horizontal axis represents the gradual change from quiet to -5dB severe noise, and the vertical axis represents the time difference to reach the target. As can be seen from the figure, under the same noise conditions, the TDOA delay error of the present invention is much smaller than that of the traditional algorithm. When the noise worsens, the error of the traditional algorithm increases sharply, while the error of the present invention increases gradually.
[0180] See Figure 3 The horizontal axis represents time, and the vertical axis represents the equivalent baseline length. As shown in the figure, when the fixed baseline length remains constant, the bionic dynamic baseline continues to increase as the head rotates autonomously.
[0181] See Figure 4 In the figure, TDOA represents the Time Difference of Arrival algorithm, TDOA+EKF represents the Time Difference of Arrival algorithm and Extended Kalman Filter algorithm, CNN-BiLSTM represents the Convolutional Bidirectional Long Short-Term Memory Network algorithm, and CNN-BiLSTM+EKF represents the Convolutional Bidirectional Long Short-Term Memory Network algorithm and Extended Kalman Filter algorithm. As can be seen from the figure, CNN extracts multi-dimensional features and EKF temporal smoothing complement each other, the error of dynamic sound source three-dimensional trajectory is significantly reduced, single-point localization jump is achieved, and the tracking accuracy of moving sound source is improved.
[0182] See Figure 5 This figure shows a comparative bar chart of positioning accuracy under five scenarios: open indoor space, indoor multipath reverberation, windy outdoor environment, mobile sound source, and outdoor mechanical noise. As can be seen from the figure, compared with the traditional array, the labeling accuracy is improved by 62%-65%, meeting the usage requirements of scenarios such as low-altitude anti-drone, industrial inspection, and security monitoring.
[0183] Example 2:
[0184] This invention also provides a sound source localization system with a biomimetic owl-like asymmetrical ear structure, including a biomimetic ear acquisition module, a signal preprocessing module, a time difference of arrival estimation module, an active rotation control module, a neural network localization module, and a trajectory fusion module. The biomimetic ear acquisition module includes a left ear component and a right ear component with asymmetrical acoustic transmission characteristics. Microphone arrays are respectively installed within the left and right ear components, and the microphone arrays are used to acquire multi-channel acoustic signals. The signal preprocessing module is used to perform noise suppression, spatial filtering, and frequency domain feature extraction on the multi-channel acoustic signals. The time difference of arrival estimation module is used to calculate the time difference of arrival between microphone channels based on the preprocessed multi-channel acoustic signals. The system combines the asymmetric acoustic transmission characteristics of the left and right ear components to generate an initial estimation result of the sound source's azimuth. The active rotation control module controls the rotation of the bionic head based on the initial sound source azimuth estimation result and simultaneously records head posture information and multi-frame arrival time difference data. The neural network localization module outputs the azimuth angle estimation result, pitch angle estimation result, and location confidence of the sound source based on multi-frame arrival time difference data, frequency domain features, head posture information, and asymmetric acoustic transmission characteristics. The trajectory fusion module fuses the azimuth angle estimation result, pitch angle estimation result, location confidence, and time series observation data using an extended Kalman filter algorithm to output the three-dimensional position and motion trajectory of the dynamic sound source.
[0185] Example 3:
[0186] This invention also provides a bionic sound source localization device, including a bionic head, a drive mechanism, a left ear component, a right ear component, a microphone array, a processor, and a memory. The left and right ear components are disposed in the bionic head and have asymmetric acoustic transmission characteristics. The microphone array is disposed within the left and right ear components and is used to collect multi-channel acoustic signals. The drive mechanism is used to drive the bionic head to rotate relative to the sound source. The memory stores program instructions executable by the processor. When the processor executes the program instructions, it preprocesses the multi-channel acoustic signals and estimates the time difference of arrival. Based on the initial estimation result of the sound source orientation, it controls the drive mechanism to adjust the posture of the bionic head. It inputs multi-frame time difference of arrival data, frequency domain features, head posture information, and asymmetric acoustic transmission characteristics into a sound source localization neural network model to obtain a sound source angle estimation result. The sound source angle estimation result and time series observation data are fused using an extended Kalman filter algorithm to output the three-dimensional position and motion trajectory of the dynamic sound source.
[0187] Example 4:
[0188] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the following steps:
[0189] Acquire multi-channel acoustic signals from microphone arrays within the left and right ear components, which have asymmetric acoustic transmission characteristics;
[0190] The multi-channel acoustic signal is subjected to noise suppression, spatial filtering, and frequency domain feature extraction; the arrival time difference between microphone channels is calculated based on the preprocessed multi-channel acoustic signal, and the initial estimation result of the sound source location is obtained by combining the asymmetric acoustic transmission characteristics;
[0191] Based on the initial estimation result of the sound source location, a bionic head rotation control quantity is generated, and head posture information and multi-frame arrival time difference data are recorded.
[0192] The multi-frame arrival time difference data, frequency domain features, head posture information, and asymmetric acoustic transmission features are input into the sound source localization neural network model to obtain the sound source angle estimation result and the location confidence.
[0193] By fusing the sound source angle estimation results, location confidence, and time series observation data using the extended Kalman filter algorithm, the three-dimensional position and motion trajectory of the dynamic sound source are output.
[0194] This invention is not limited to the embodiments described above. Any changes in shape or structure shall fall within the protection scope of this invention. The protection scope of this invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of this invention, but all such changes and modifications shall fall within the protection scope of this invention.
Claims
1. A method for sound source localization using a biomimetic owl-inspired asymmetrical ear structure, characterized in that, The method includes the following steps: S1. Construct a left ear assembly and a right ear assembly with asymmetric acoustic transmission characteristics, arrange microphone arrays in the left ear assembly and the right ear assembly respectively, and perform amplitude calibration, delay compensation and time synchronization on the sampling data of each microphone channel; S2. Perform noise suppression, spatial filtering, and frequency domain feature extraction on the multi-channel acoustic signals acquired by the microphone array to obtain preprocessed acoustic signals; S3. Calculate the arrival time difference between the left ear component and the right ear component and between the microphone channels based on the preprocessed acoustic signal, and obtain the initial estimation result of the sound source location by combining the asymmetric acoustic transmission characteristics of the left ear component and the right ear component. S4. Based on the initial estimation result of the sound source orientation, control the bionic head to rotate, so that the microphone array forms a changing positioning baseline relative to the sound source, and record the head posture information and multi-frame arrival time difference data during the rotation process. S5. Input the multi-frame arrival time difference data, frequency domain features, head posture information and asymmetric acoustic transmission features into the sound source localization neural network model, and output the azimuth angle estimation result, pitch angle estimation result and location confidence of the sound source. S6. The extended Kalman filter algorithm is used to fuse the azimuth estimation results, pitch estimation results, position confidence and time series observation data to obtain the three-dimensional position and motion trajectory of the dynamic sound source.
2. The sound source localization method for a biomimetic owl asymmetrical ear structure as described in claim 1, characterized in that, The asymmetric acoustic transmission characteristics of the left and right ear components are formed by at least two of the following: vertical height difference of the ear canal, difference of auricular opening direction, difference of ear canal length, difference of auricular cavity curvature, and difference of auricular edge refraction angle.
3. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 2, characterized in that, The asymmetric acoustic transfer characteristics include amplitude-frequency difference characteristics, phase-frequency difference characteristics, and additional phase delay characteristics obtained from the acoustic transfer functions of the left and right ear components.
4. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 1, characterized in that, The microphone array described in S1 includes multiple microphones respectively embedded in the left ear assembly and the right ear assembly. The multiple microphones are arranged at intervals along the circumferential, arcuate, or annular path of the auricle cavity to form a spatial sampling position that matches the asymmetrical structure of the auricle.
5. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 1, characterized in that, S2 specifically involves converting the time-domain acoustic signals acquired by each microphone channel into frequency-domain acoustic signals; The received signal covariance matrix of the microphone array is calculated based on the frequency domain acoustic signal, and the noise covariance matrix is estimated based on noise segments or low signal-to-noise ratio segments. Generate array steering vectors based on the spatial position of the microphone array and the directions of candidate sound sources; The spatial filtering weights for each microphone channel are determined based on the noise covariance matrix and the array steering vector. The frequency domain acoustic signals of each microphone channel are weighted and superimposed according to the spatial domain filtering weights to obtain the spatial domain filtered signal. The spectral amplitude features, channel phase difference features, and frequency band energy features are extracted from the spatial domain filtered signal and used as the frequency domain features.
6. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 5, characterized in that, The noise suppression includes performing multi-scale wavelet packet decomposition on the time-domain sampled signal of each microphone channel to obtain wavelet coefficients at different scales, performing adaptive soft thresholding on the wavelet coefficients, and reconstructing the denoised channel signal based on the processed wavelet coefficients.
7. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 1, characterized in that, The calculation of the arrival time difference in S3 includes: calculating the phase transformation generalized cross-correlation function for different microphone channel pairs, obtaining the candidate delay peaks of the channel pairs, and determining the effective arrival time difference from the candidate delay peaks through local peak statistics, sliding window fitting, and abnormal peak suppression.
8. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 7, characterized in that, Obtaining the initial estimation result of the sound source location includes: constructing a spatial response power spectrum based on the effective arrival time difference of each channel pair, searching for the peak value of the response power spectrum in the candidate spatial location, and converting the asymmetric acoustic transmission characteristics of the left ear component and the right ear component into an additional phase delay compensation amount to correct the spatial response power spectrum.
9. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 1, characterized in that, S4 controls the bionic head to perform rotation by: determining the rotation direction and rotation amplitude based on the angular deviation between the initial estimation result of the sound source orientation and the current head posture, and determining the adaptive rotation gain based on the sound source distance estimate, positioning confidence or arrival time difference stability, so that the bionic head forms a multi-posture positioning baseline during rotation.
10. The sound source localization method for the biomimetic owl asymmetrical ear structure according to claim 1, characterized in that, The sound source localization neural network model includes a convolutional neural network and a bidirectional long short-term memory network. The convolutional neural network is used to extract local acoustic features from the feature map formed by multi-frame time difference of arrival data, frequency domain features, head posture information, and asymmetric acoustic transmission features. The bidirectional long short-term memory network is used to extract temporal features between consecutive frame acoustic observations. The extended Kalman filter algorithm adjusts the observation noise covariance according to the location confidence and outputs the three-dimensional position and motion trajectory of the dynamic sound source.