A method, device, and medium for sagittal auditory localization based on Bayesian inference

By employing Bayesian inference methods and combining personalized head-related transfer functions and localization response data, the ill-posed problem of polar angle estimation in complex acoustic scenarios for auditory localization models was solved, achieving more accurate and stable polar angle localization.

CN120468776BActive Publication Date: 2025-10-31SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510955617.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-31
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing auditory localization models suffer from ill-posed polar angle estimation in complex acoustic scenarios, especially when the sound source spectrum is dynamic, resulting in a significant decrease in localization performance.

Method used

A Bayesian inference-based approach is adopted. By acquiring personalized head-related transfer function and localization response data, and combining the cross-correlation calculation of spectral features and feature templates, along with the perceptual likelihood function and spatial prior distribution, the maximum a posteriori estimation criterion is used to predict polar angles. The model is then optimized through parameter fitting to improve estimation accuracy.

Benefits of technology

The stability and accuracy of polar angle estimation have been improved, making the model's localization results in complex acoustic environments closer to the user's actual auditory experience, thus enhancing the model's robustness and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120468776B_ABST
    Figure CN120468776B_ABST
Patent Text Reader

Abstract

This invention relates to the field of auditory modeling technology, specifically to a method, device, and medium for sagittal auditory localization based on Bayesian inference. In this invention, perceptual evidence and spatial prior information are integrated based on Bayesian criteria to simulate human auditory localization behavior in the sagittal plane. A perceptual likelihood function is constructed by calculating the spectral cross-correlation between input spectral features and feature templates, and an adaptive correlation-similarity mapping mechanism is introduced to alleviate the ill-posed problem in dynamic spectral sound source localization tasks, thereby improving the stability and accuracy of polar angle estimation. The model parameters are fitted using the user's personalized head-related transfer function and localization response data, making the polar angle estimation results in complex listening scenarios closer to the user's actual localization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of auditory modeling technology, specifically to a method, device, and medium for sagittal auditory localization based on Bayesian inference. Background Technology

[0002] The human auditory system uses monoaural spectral cues to estimate the polar direction of sound sources in the sagittal plane. These cues are primarily formed by the directional filtering of sound by the trunk, head, and especially the auricle. Although some progress has been made in elucidating these mechanisms, the neural processing mechanisms of spectral cues in auditory localization are not fully understood, especially in complex acoustic scenarios.

[0003] Computational auditory models qualitatively describe and explain the human auditory perception mechanism by simulating the auditory processing stages of the auditory system. Traditional auditory localization models exhibit good performance when the sound source spectrum is relatively flat, generating human-like localization predictions; however, their localization effectiveness significantly decreases when the sound source has dynamic spectral characteristics. These models typically employ template matching criteria based on the L1 norm or Gaussian likelihood. When both the sound source spectrum and head-related transfer functions (HRTFs) are unknown, polar angle estimation leads to mathematically ill-posed problems. Summary of the Invention

[0004] In view of this, the present invention provides a method, device and medium for sagittal plane auditory localization based on Bayesian inference, so as to solve the ill-posed problem of polar angle estimation in the prior art.

[0005] In a first aspect, the present invention provides a sagittal plane auditory localization method based on Bayesian inference. The method includes: acquiring a user's personalized head-related transfer function and localization response data; extracting spectral features of the left and right ears from the binaural input signals respectively; performing cross-correlation calculations on the spectral features and feature templates in each polar angle direction on the target sagittal plane according to the Bayesian criterion to determine the perceptual likelihood function in each polar angle direction, and combining the perceptual likelihood function with the spatial prior distribution to determine the posterior probability distribution of each polar angle, wherein the feature templates are determined by the personalized head-related transfer function; selecting the polar angle prediction value of the model from the posterior probability distribution based on the maximum a posteriori estimation criterion, wherein the model is a model constructed based on spectral feature extraction and Bayesian inference; fitting model parameters to minimize the deviation between the model prediction and the actual localization response under a preset threshold; and performing polar angle localization on the binaural input signals based on the personalized fitting model, wherein the personalized fitting model is a model after parameter fitting.

[0006] In this invention, perceptual evidence and spatial prior information are integrated based on Bayesian criteria to simulate human auditory localization behavior in the sagittal plane. A perceptual likelihood function is constructed by calculating the spectral cross-correlation between input spectral features and feature templates to alleviate the ill-posed problem in dynamic spectral sound source localization tasks, thereby improving the stability and accuracy of polar angle estimation. The model parameters are fitted using the user's personalized head-related transfer function and localization response data, making the polar angle estimation results in complex listening scenarios closer to the user's actual localization performance.

[0007] In one alternative implementation, acquiring the user's personalized head-related transfer function and localization response data includes: obtaining the user's personalized head-related transfer function based on personalized measurements or physiological parameter matching; performing a psychoacoustic localization experiment to collect the user's localization response data in all acquisition directions on a spherical grid, and resampling the user's personalized head-related transfer function to match the spatial resolution of the localization response data.

[0008] In this invention, the head-related transfer function (HRTF) is obtained through personalized measurements or physiological parameter matching. Localization response data is collected in conjunction with psychoacoustic localization experiments, and the spatial resolution of the HRTF data is resampled. This accurately adapts to the user's physiological characteristics, enhances the personalization of the HRTF, and ensures that the localization response data matches the spatial characteristics of the HRTF. This provides more realistic and intuitive data for subsequent tasks such as polar angle estimation, improving the accuracy and robustness of the model's predictions.

[0009] In one optional implementation, the spectral features of the left and right ears are extracted from the binaural input signals, respectively, including: extracting the initial feature vector of the binaural input signals based on a gamma-pass filter bank of approximate basilar membrane filtering and half-wave rectification of hair cell conduction; determining the spectral amplitude profile feature based on the root mean square amplitude of the initial feature vector; and determining the noisy feature by combining the spectral amplitude profile feature and additive Gaussian noise, with the noisy feature serving as the spectral feature.

[0010] In this invention, a gamma-pass filter bank and half-wave rectification are used to extract initial features, which closely align with physiological mechanisms; the root mean square amplitude determines the spectral amplitude profile, focusing on key energy features; and additive Gaussian noise is used to simulate perceived noise. Thus, the processing of this simulated auditory system makes the extracted spectral features closer to human auditory perception, providing realistic physiological feature inputs for polar angle estimation and other functions, thereby improving the model's robustness and perceptual simulation accuracy in complex acoustic environments.

[0011] In one optional implementation, based on the Bayesian criterion, cross-correlation calculations are performed between spectral features and feature templates in each polar angle direction on the target sagittal plane to determine the perceptual likelihood function in each polar angle direction. The posterior probability distribution of each polar angle is then determined by combining the perceptual likelihood function with the spatial prior distribution. This includes: calculating cross-correlation coefficients based on spectral features and feature templates in each polar angle direction on the target sagittal plane, retaining positive correlation coefficients and setting non-positive correlation coefficients to zero, obtaining one-sided correlation vectors related to each polar angle direction, where the one-sided correlation vectors include left-side and right-side correlation vectors; mapping the one-sided correlation vectors to similarity vectors related to each polar angle direction using a psychometric function, where the similarity vectors include left-side and right-side similarity vectors, and the perceptual selectivity parameter and perceptual sensitivity parameter of the psychometric function are either fixed values ​​or adaptively adjusted using an adaptive correlation-similarity mapping mechanism; weighting and normalizing the left-side and right-side similarity vectors using a binaural weighted function to obtain the perceptual likelihood function related to each polar angle direction; and determining the posterior probability distribution of each polar angle by combining the perceptual likelihood function with the spatial prior distribution.

[0012] In this invention, the auditory perception process is accurately simulated by using Bayesian criteria, combined with cross-correlation calculation, psychometric functions, and binaural weighting functions: positive correlation is retained to focus effective matching, psychometric functions map subjective perception, binaural weighting integrates binaural information, and spatial prior information is combined to make Bayesian inference more consistent with the human auditory localization mechanism. This can effectively solve the ill-posed problem of dynamic sound source polar angle estimation and improve the accuracy and reliability of polar angle estimation.

[0013] In one optional implementation, when the perceptual selectivity parameter and the perceptual sensitivity parameter are adaptively adjusted using an adaptive correlation-similarity mapping mechanism, a psychometric function is used to map the one-sided correlation vector into similarity vectors related to each polar angle direction. This includes: arranging the one-sided correlation vectors in each polar angle direction in descending order, determining the maximum value and the second peak value of the one-sided correlation vector; performing linear regression on a preset psychometric function based on the maximum value of the one-sided correlation vector and the preset maximum value of the similarity vector, as well as the second peak value of the one-sided correlation vector and the preset minimum value of the similarity vector, to determine the adaptive parameters; and mapping the one-sided correlation vector into similarity vectors related to each polar angle direction based on the psychometric function determined by the adaptive parameters.

[0014] In this invention, an adaptive correlation-to-similarity mapping (ACSM) mechanism can be introduced into the psychometric function to simulate human active directional exploration behavior when perceptual confidence is low.

[0015] In one alternative implementation, selecting the polar angle prediction value of the model from the posterior probability distribution based on the maximum a posteriori estimation criterion includes: selecting the direction with the highest probability in the posterior distribution as the polar angle prediction value of the target sound source according to the maximum a posteriori estimation criterion, and introducing response noise, wherein the response noise follows a von Mises-Fischer distribution.

[0016] In this invention, response noise distributed by von Mises-Fischer is added in the direction of the polar angle with the highest probability to simulate real auditory noise interference.

[0017] In one optional implementation, fitting model parameters to minimize the deviation between model prediction and true positioning response under a preset threshold includes: using the polar angle prediction containing response noise as the predicted value and the polar angle direction in the positioning response data as the true value; adjusting preset parameters so that the deviation between the predicted value and the true value meets preset requirements, the preset parameters including the spectral feature extraction stage, the Bayesian inference stage, and free parameters in the response noise.

[0018] In this invention, parameters are adjusted based on predicted and true values ​​to optimize the model's adaptability to complex acoustic environments, improve the accuracy and robustness of polar angle estimation, and make the model more consistent with the perceptual behavior pattern of human auditory localization.

[0019] In one optional implementation, adjusting preset parameters to ensure the deviation between the predicted and true values ​​meets preset requirements includes: calculating polarity positioning performance indices corresponding to the predicted and true values, whereby the polarity positioning performance indices include local polarity error and quadrant error rate, and the predicted value is the model prediction value based on the current model parameters; determining the prediction deviation between the predicted and true values ​​based on the polarity positioning performance indices, whereby the prediction deviation includes the prediction deviations corresponding to the local polarity error and quadrant error rate, respectively; determining the joint prediction deviation based on the prediction deviations corresponding to the local polarity error and quadrant error rate, respectively; adjusting the initial parameters based on the relationship between the joint prediction deviation and the joint prediction deviation threshold, and repeating the process of determining the polarity positioning performance indices, prediction deviation, and joint prediction deviation, as well as adjusting the parameters, until the relationship between the joint prediction deviation and the joint prediction deviation threshold meets preset requirements, whereby the joint prediction deviation threshold is determined based on the threshold of the polarity positioning performance indices.

[0020] In this invention, the prediction bias and joint bias are determined by calculating the polarity positioning indices (local polarity error, quadrant error rate) between the predicted and actual values, and the parameters are iteratively adjusted based on their relationship with a threshold. This allows for precise quantification of the difference between the model's prediction and actual performance, and the parameter optimization is constrained by the joint bias threshold, enabling the model parameters to adapt to the user's auditory characteristics, improving the accuracy of polar angle estimation, and enhancing the model's adaptability and robustness to complex acoustic scenarios.

[0021] Secondly, this invention provides a sagittal plane auditory localization device based on Bayesian inference. The device includes: a data acquisition module for acquiring a user's personalized head-related transfer function and localization response data; a feature extraction module for extracting spectral features of the left and right ears from the binaural input signals; a Bayesian inference module for cross-correlation calculation of the spectral features with feature templates in each polar angle direction on the target sagittal plane according to the Bayesian criterion, determining the perceptual likelihood function in each polar angle direction, and combining the perceptual likelihood function with the spatial prior distribution to determine the posterior probability distribution of each polar angle, wherein the feature templates are determined by the personalized head-related transfer function; a perceptual decision module for selecting the polar angle prediction value of the model from the posterior probability distribution based on the maximum a posteriori estimation criterion, wherein the model is a model constructed based on spectral feature extraction and Bayesian inference; a parameter fitting module for fitting model parameters to minimize the deviation between the model prediction and the actual localization response under a preset threshold; and a direction prediction module for performing polar angle direction localization on the binaural input signals based on the personalized fitting model, wherein the personalized fitting model is the model after parameter fitting.

[0022] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the sagittal plane auditory localization method based on Bayesian inference described in the first aspect or any corresponding embodiment thereof.

[0023] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the Bayesian inference-based sagittal auditory localization method described in the first aspect or any corresponding embodiment thereof.

[0024] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the Bayesian inference-based sagittal auditory localization method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0025] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating the sagittal plane auditory localization method based on Bayesian inference according to an embodiment of the present invention.

[0027] Figure 2 This is a flowchart illustrating another method for sagittal auditory localization based on Bayesian inference according to an embodiment of the present invention.

[0028] Figure 3 This is a block diagram of a sagittal plane auditory localization model based on Bayesian inference, according to an embodiment of the present invention.

[0029] Figure 4 This is a schematic diagram illustrating the effect of high-frequency speech attenuation on polarity positioning performance according to an embodiment of the present invention.

[0030] Figure 5 This is a schematic diagram illustrating the impact of a non-personalized HRTF on polarity positioning performance according to an embodiment of the present invention.

[0031] Figure 6 This is a schematic diagram illustrating the effect of spectral ripple on polarity positioning performance according to an embodiment of the present invention.

[0032] Figure 7 This is a schematic diagram of the polar angle localization response of a typical digital listener under different SNRs according to an embodiment of the present invention;

[0033] Figure 8 This is a schematic diagram illustrating the polarity localization performance of five digital listeners under different SNRs according to an embodiment of the present invention;

[0034] Figure 9 This is a schematic diagram illustrating the polarity localization performance of speech stimuli under high-frequency information loss according to an embodiment of the present invention.

[0035] Figure 10 This is a structural block diagram of a sagittal plane auditory localization device based on Bayesian inference according to an embodiment of the present invention.

[0036] Figure 11 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0037] As described in the background section, in the human auditory system, the physiological structures of the auricle (outer ear), external auditory canal, and head produce a filtering effect on sounds incident from different directions. Specifically, when a sound source is located at a certain polar angle in the sagittal plane, the path of the sound entering the auricle is altered by the folds and curves of the auricle, causing interference, reflection, or absorption of sound waves of different frequencies, forming specific spectral characteristics. By memorizing or learning the correspondence between these spectral characteristics and polar angles, the brain can determine the vertical direction of the sound source based solely on monocular input.

[0038] The head-related transfer function (HRTF) describes the spectral changes of sound from a free field to the eardrum. For example, the HRTF at different polar angles can be pre-measured to create a "polar angle-spectral feature" template library. The perceived spectrum is then matched with templates in this library to determine the polar angle of the sound source. Matching is more effective when the sound source spectrum is flat; however, if the sound source spectrum is dynamic, the time-varying spectrum is coupled with the HRTF features, making it impossible to separate the "sound source characteristics" from the "directional characteristics" in template matching, thus reducing the accuracy of polar angle estimation.

[0039] Based on this, the use of a perceptual likelihood function based on spectral cross-correlation in Bayesian inference in this embodiment can effectively solve the ill-posed problem of auditory localization models when predicting the polar angle direction of dynamic spectral signals.

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] According to an embodiment of the present invention, a method for sagittal plane auditory localization based on Bayesian inference is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0042] This embodiment provides a sagittal plane auditory localization method based on Bayesian inference, which can be used in electronic devices such as computers, mobile phones, and tablets. Figure 1 This is a flowchart of a sagittal plane auditory localization method based on Bayesian inference according to an embodiment of the present invention, as shown below. Figure 1 As shown, the process includes the following steps:

[0043] Step S101: Obtain the user's personalized head-related transfer function (HRT) and location response data. The HRT describes the spectral modulation characteristics of sound as it propagates from free space to the ear canal entrance, caused by scattering, reflection, and diffraction from physiological structures such as the head, auricle, and torso. Specifically, it can be determined by the spatial direction (lateral and polar angles) of the sound source and the sound frequency; that is, the HRT differs for sound sources from different directions. Due to individual differences in human acoustic structure, different people have different HRTs. Therefore, this embodiment determines personalized HRTs for different users. Personalized HRTs can be determined through measurement or modeling, which will not be elaborated further here. The location response data represents the user's directional judgment of sound sources from different directions, reflecting the user's actual perception of the sound source's direction.

[0044] Step S102: Extract the spectral features of the left and right ears from the binaural input signals. The extraction process of these spectral features simulates the processing of the human auditory system; that is, the extracted spectral features are used as a representation of internal auditory perception. The binaural input signals can be determined by collecting sound signals from two microphones at the entrances of the left and right ear canals in a real environment or laboratory acoustic setting (such as an anechoic chamber or reverberation chamber).

[0045] Step S103: According to the Bayesian criterion, cross-correlation calculation is performed between the spectral features and the feature templates in each polar angle direction on the target sagittal plane to determine the perceptual likelihood function in each polar angle direction. The posterior probability distribution of each polar angle is determined by combining the perceptual likelihood function with the spatial prior distribution. The feature template is determined by the personalized head correlation transfer function.

[0046] In anatomy, the sagittal plane is a vertical plane that divides an organism (such as the head) into two symmetrical parts (similar to the plane that cuts the head from the tip of the nose to the back of the head). In auditory localization, sound source localization in the sagittal plane primarily involves determining vertical directions (such as above, in front, below, and behind), distinct from localization in the horizontal plane (left-right direction). The polar angle, on the other hand, refers to the angular deviation from the reference direction (such as directly in front) to the target direction in a polar coordinate system.

[0047] In this embodiment, the polar angle is defined in a binaural polar coordinate system, and it, together with the lateral angle, determines the spatial orientation of the sound source relative to the user. The horizontal angle range is from (Right side) to 90° (left side), while the polar angle range is from (Above and to the front) to 270° (below and to the rear). The target sagittal plane responds to the lateral angle. Decide.

[0048] Specifically, the feature template is obtained by uniformly resampling the DTF (Directional Transfer Function) data of the entire user space to obtain a fixed number of directional DTF data. By removing the direction-independent Common Transfer Function (CTF) part from the HRTF, the direction-dependent Directional Transfer Function (DTF) data can be obtained. Finally, after the above-mentioned spectral feature extraction stage, the spectral feature template library is obtained.

[0049] After determining the feature template, the extracted input spectral features are cross-correlated with the feature template to measure the feature matching degree. Then, the cross-correlation result is converted into a perceptual likelihood function to reflect the probability that "the spectral features come from a certain polar angle direction". Then, combined with spatial priors (such as the statistical law of the probability of the occurrence of polar angles), the likelihood and priors are fused through Bayes' theorem to obtain the posterior probability distribution of each polar angle, realizing the probability inference of the polar angle direction, so that the model fits the user's auditory physiology and spatial statistical characteristics.

[0050] Step S104: Select the polar angle prediction value of the model from the posterior probability distribution based on the maximum a posteriori estimation criterion. The model is a model constructed based on spectral feature extraction and Bayesian inference.

[0051] Specifically, after determining the posterior probability distribution, the polar angle direction corresponding to the maximum probability is determined from the posterior probability distribution based on the maximum a posteriori estimation criterion, which is the predicted value.

[0052] Step S105: Fit the model parameters to minimize the deviation between the model prediction and the actual localization response under a preset threshold.

[0053] Specifically, during the model parameter fitting process, the obtained predicted values ​​and the positioning response data are compared. If the deviation between the two does not meet the preset conditions, the parameter fitting process continues until the deviation between the two meets the preset conditions.

[0054] Step S106 involves locating the polar angle direction of the binaural input signal based on a personalized fitting model, where the personalized fitting model is the model after parameter fitting. Specifically, based on the fitting parameters from the previous step, the polar angle direction of any binaural input signal can be estimated, ultimately determining the direction of the sound source in the sagittal plane. That is, steps S102 to S104 are performed on any binaural input signal to ultimately determine the estimated polar angle direction of the sound source.

[0055] The sagittal auditory localization method based on Bayesian inference provided in this invention integrates perceptual evidence and spatial prior information based on Bayesian criteria to simulate human auditory localization behavior in the sagittal plane. A perceptual likelihood function is constructed by calculating the spectral cross-correlation between input spectral features and feature templates to alleviate the ill-posed problem in dynamic spectral sound source localization tasks, thereby improving the stability and accuracy of polar angle estimation. The model parameters are fitted using the user's personalized head-related transfer function and localization response data, making the polar angle estimation results in complex listening scenarios closer to the user's actual localization performance.

[0056] This embodiment provides a sagittal plane auditory localization method based on Bayesian inference, the process of which includes the following steps:

[0057] Step S201: Obtain the user's personalized header-related transmission function and positioning response data.

[0058] Specifically, step S201 includes:

[0059] Step S2011: Obtain the user's personalized head-related transfer function (HRTF) based on personalized measurements or physiological parameter matching. Specifically, when obtaining this function using personalized measurements, professional acoustic equipment can be used in an anechoic chamber or similar environment to transmit sound signals (such as broadband noise or swept-frequency signals) of known spectrum into the ear canal from different spatial directions (covering common lateral and polar angles). Simultaneously, a miniature microphone is placed at the ear canal entrance (or a simulated ear canal position) to collect the sound signals filtered by physiological structures such as the head and auricle. For example, a multi-speaker array can be used to surround the subject, playing signals sequentially at preset angles while simultaneously collecting the received signals in the ear canal. By comparing the spectral differences between the free field signal (without head obstruction) and the ear canal-collected signal, the HRTFs corresponding to different directions can be calculated, reflecting the filtering characteristics of a specific user's head and auricle for sounds from different directions.

[0060] When determining this function using physiological parameter matching, a database can be established based on a large amount of measured HRTFs data. Physiological parameters strongly correlated with HRTFs (such as head size, three-dimensional ear shape, ear canal length and diameter, etc.) can be extracted. After measuring these physiological parameters for new users, the HRTFs of individuals with the most similar physiological characteristics are matched in the database as an approximation of the user's personalized HRTFs. For example, parameters such as ear shape and head circumference can be obtained through 3D scanning and input into a machine learning model. The model then filters out sample HRTFs with high matching degrees to the physiological parameters from the database, and assigns them to new users after interpolation and optimization.

[0061] Step S2012 involves performing a psychoacoustic localization experiment to collect the user's localization response data across all acquisition directions on a spherical grid, and resampling the user's personalized head-related transfer function to match the spatial resolution of the localization response data. Specifically, a psychoacoustic localization experiment refers to an experiment conducted in a controlled acoustic environment (such as an anechoic chamber or virtual acoustic environment) where users subjectively determine the orientation of sound sources at different spatial locations. The aim is to obtain behavioral data on the user's actual perception of the sound source direction. A spherical grid refers to dividing a sphere centered on the user in three-dimensional space into discrete "grid points" according to certain rules. Each grid point corresponds to a spatial direction (using a polar angle). Horizontal angle (Interaural polar coordinates description).

[0062] After determining the positioning response data for all acquisition directions on the spherical grid through psychoacoustic positioning experiments, it is necessary to further determine whether the spatial resolution of the personalized head-related transfer function (HRTF) and the spatial resolution of the positioning response data are the same. If they are inconsistent, such as HRTFs being sampled at a polar angle of 5° and an azimuth angle of 10°, while the positioning response data is sampled at a polar angle of 10° and an azimuth angle of 15°, then the HRTFs need to be resampled.

[0063] Specifically, if the HRTFs sampling density is higher than that of the positioning response data, for sparser grid points in the positioning response data, the HRTFs value corresponding to the sparse point is calculated using data from surrounding HRTFs sampling points through interpolation (such as spherical interpolation, linear interpolation, etc.). For example, given the spectral data of several adjacent HRTFs sampling points, the HRTFs of the target sparse point are fitted using a spherical interpolation algorithm. If the HRTFs sampling density is lower than that of the positioning response data, the point closest (in spatial direction) to the grid point in the positioning response data is selected from the HRTFs sampling points as the matching point; or the HRTFs sampling data is downsampled through filtering, decimation, or other operations to make its spatial resolution consistent with that of the positioning response data.

[0064] Step S202: Extract the spectral features of the left and right ears from the binaural input signals respectively.

[0065] Specifically, step S202 includes:

[0066] Step S2021: Extract the initial feature vector of the binaural input signal based on the gamma-pass filter bank of approximate basilar membrane filtering and the half-wave rectification of hair cell conduction.

[0067] The basilar membrane in the human cochlea exhibits frequency selectivity, responding differently to different frequencies of sound at different locations, much like a set of filters that decompose sound into frequencies. A gammatone filterbank is a mathematical model that simulates this frequency decomposition characteristic of the basilar membrane. Specifically, each gammatone filter has a specific center frequency, bandpass filtering the input signal and outputting signal components within the corresponding frequency range. In this embodiment, the filterbank divides the 0.7-18kHz frequency range into 28 frequency bands, each corresponding to a filter. This decomposes the complex sound signal into 28 channels according to frequency, simulating the basilar membrane's response to different frequencies. The hair cells in the cochlea convert the mechanical signals of the basilar membrane vibration into electrical signals, and their response is nonlinear. Half-wave rectification is a simple method to simulate this nonlinear characteristic.

[0068] Based on this, a 28-dimensional (frequency band) initial feature vector of the signal for each ear is obtained using a gammatone filter bank that approximates the basilar membrane and half-wave rectification of hair cell conduction. The initial feature vector of each frequency band is expressed by the following formula:

[0069] (1)

[0070] In the formula, Represents a gammatone filter, subscript For frequency band indexing, For time indexing, Indicates stimulus (source) signal Impulse response associated with unilateral head The monoaural input signal obtained after linear convolution is superscripted. Refers to the left and right ears. To approximate the spectral resolution of the human cochlea, the gammatone filter bank is set to an equivalent rectangular bandwidth to divide the signal into 28 frequency bands in the frequency range of 0.7 to 18 kHz; the max function is used to preserve the non-negative part, realize half-wave rectification, and simulate the nonlinearity of hair cell conduction.

[0071] Step S2022: Determine the spectral magnitude profile features based on the root mean square amplitude of the initial feature vector; specifically, collect the root mean square amplitude over the entire time range from the initial feature vector to obtain the spectral magnitude profile (SMP) features, which are represented by the following formula:

[0072] (2)

[0073] Furthermore, since the initial feature vector extracted above includes vectors from 28 frequency bands, the SMP features from these 28 frequency bands are combined into a single-ear spectral feature vector, expressed by the following formula:

[0074] (3)

[0075] Step S2023: Combine the spectral amplitude profile features and additive Gaussian noise to determine the noisy features, which are then used as spectral features. Specifically, after determining the spectral amplitude profile, additive Gaussian noise is added to simulate the limited perceptual accuracy of the feature extraction stage. Therefore, the noisy features after adding additive Gaussian noise are expressed by the following formula:

[0076] (4)

[0077] Step S203: According to the Bayesian criterion, cross-correlation calculation is performed between the spectral features and the feature templates in each polar angle direction on the target sagittal plane to determine the perceptual likelihood function in each polar angle direction. The posterior probability distribution of each polar angle is determined by combining the perceptual likelihood function with the spatial prior distribution. The feature template is determined by the personalized head correlation transfer function.

[0078] Specifically, step S203 includes:

[0079] Step S2031: Calculate the cross-correlation coefficients based on the spectral features and the feature templates in each polar angle direction on the target sagittal plane. Retain the positive correlation coefficients and set the non-positive correlation coefficients to zero to obtain the one-sided correlation vectors related to each polar angle direction. The one-sided correlation vectors include the left-side correlation vector and the right-side correlation vector. Specifically, calculate the cross-correlation coefficients for each spectral feature and the feature templates in each polar angle direction on the target sagittal plane one by one. The one-sided correlation vectors are then determined using the following formula:

[0080] (5)

[0081] In the formula, the feature template Calculated based on the user's personalized HRTFs, which includes each polar angle on the target sagittal plane in equation (4). The noise-free characteristics. The target sagittal plane is determined by the lateral angular response. The decision is made based on the fact that the extracted spectral features include those of the left and right ears. Therefore, the spectral features from both sides are cross-correlated with the feature templates, resulting in two correlation vectors: a left-side correlation vector and a right-side correlation vector.

[0082] Step S2032 involves using a psychometric function to map the one-sided correlation vector into similarity vectors related to each polar angle. The similarity vectors include a left-side similarity vector and a right-side similarity vector. The perceptual selectivity parameter and perceptual sensitivity parameter of the psychometric function are either fixed values ​​or adaptively adjusted using an adaptive correlation-similarity mapping mechanism. Specifically, human auditory perception is not a simple "feature matching" process; it involves subjective nonlinearity and sensitivity differences. The psychometric function simulates this nonlinear mapping from "features to perceptual similarity," converting the obtained correlation vector into a similarity vector that better reflects human subjective perception. Specifically, the similarity vector is calculated using the following formula:

[0083] (6)

[0084] In the formula, and These represent the perception selectivity parameter and the perception sensitivity parameter, respectively.

[0085] In one alternative implementation, the parameters in the psychometric function can be fixed values. Furthermore, to better align with human auditory perception patterns in uncertain scenarios, an adaptive correlation-to-similarity mapping (ACSM) mechanism is introduced to simulate proactive directional exploration behavior in humans when perceptual confidence is low. Based on this, the parameters... and The following method is used to determine:

[0086] Step a1: Arrange the one-sided correlation vectors in descending order for each polar angle direction, and determine the maximum value and the second peak value of the one-sided correlation vectors.

[0087] Step a2 involves performing linear regression on a preset psychometric function based on the maximum value of the one-sided correlation vector, the maximum value of the preset similarity vector, the second peak of the one-sided correlation vector, and the minimum value of the preset similarity vector to determine the adaptive parameters. The preset maximum and minimum similarity values ​​can be predetermined. These two values ​​limit the range of similarity values, preventing the model from failing to characterize "low confidence" scenarios due to extreme values, thereby simulating the ambiguity of human perception. In this embodiment, the preset minimum similarity value... Similarity_min = 0.01, Preset maximum similarity Similarity_max = 0.99.

[0088] Therefore, based on the maximum value of the one-sided correlation vector C_max The maximum value of the preset similarity vector and the second peak of the one-sided correlation vector. C_2ndmax The two points determined by the minimum value of the preset similarity vector ( C_max , Similarity_max )and( C_2ndmax , Similarity_min Perform linear regression to solve for the adaptive parameters. and .

[0089] Step a3: Based on the psychometric function determined by the adaptive parameters, the one-sided correlation vector is mapped to a similarity vector related to each polar angle direction. Specifically, after determining the adaptive parameters, they are substituted into the psychometric function, and the left and right correlation vectors determined in step S2031 are substituted to obtain the left similarity vector. Similarity vector on the right .

[0090] Step S2033: The left and right similarity vectors are weighted, summed, and normalized using a binaural weighting function to obtain the perceptual likelihood function related to each polar angle direction; specifically, the determined perceptual likelihood function is expressed by the following formula:

[0091] (7)

[0092] In the formula, For the weighting function, the weighting coefficients for both ears are... In this model, the lateral angle response is assumed. Known.

[0093] Step S2034: Determine the posterior probability distribution of each polar angle by combining the perceived likelihood function and the spatial prior distribution. The spatial prior distribution is defined as a finite-variance Gaussian distribution centered on the horizontal plane in the vertical dimension.

[0094] (8)

[0095] In the formula, the variance of the posterior hemisphere The variance greater than that of the first hemisphere .

[0096] Based on the prior distribution of this space, the posterior probability distribution is expressed by the following formula:

[0097] (9)

[0098] Step S204: Select the polar angle prediction value of the model from the posterior probability distribution based on the maximum a posteriori estimation criterion.

[0099] Specifically, step S204 includes:

[0100] Step S2041: According to the maximum a posteriori estimation criterion, the direction with the highest probability in the posterior distribution is selected as the polar angle prediction value of the target sound source, and response noise is introduced. This response noise follows a von Mises-Fischer distribution. Specifically, according to the maximum a posteriori estimation criterion, the direction with the highest probability in the posterior distribution is selected as the polar angle prediction value of the target sound source. Then, response noise is introduced to simulate the inherent uncertainty in the motion response or decision-making process, thereby obtaining the final prediction result. The predicted value after adding noise is expressed by the following formula:

[0101] (10)

[0102] In the formula, Follows a mean of zero and a concentration parameter of The concentration of the von Mises-Fischer distribution can be converted to standard deviation. .

[0103] Step S205: Fit the model parameters to minimize the deviation between the model prediction and the actual localization response under a preset threshold.

[0104] Specifically, step S205 includes:

[0105] Step S2051: Use the polar angle prediction containing response noise as the predicted value, and locate the polar angle direction in the response data as the true value.

[0106] Step S2052 involves adjusting preset parameters to ensure that the deviation between the predicted and actual values ​​meets preset requirements. These preset parameters include parameters from the spectral feature extraction stage, the Bayesian inference stage, and free parameters in the response noise. Specifically, in this embodiment, the adjusted parameters include the variance of the anterior hemisphere. Variance of the posterior hemisphere Parameters of additive Gaussian noise Concentration standard deviation Perceived selectivity parameter Sensing sensitivity parameters .

[0107] In one optional implementation, adjusting preset parameters to ensure that the deviation between the predicted value and the actual value meets preset requirements includes the following steps:

[0108] Step b1: Calculate the polarity positioning performance index corresponding to the predicted value and the true value respectively. The polarity positioning performance index includes local polarity error and quadrant error rate. The predicted value is the model prediction value based on the current model parameters.

[0109] Specifically, the initial value and range of the preset parameters can be set as follows: The initial value is 11.5°, and the range is [5°, 25°]. The initial value is 11.5°, and the range is [5°, 40°]. The initial value is 3.5°, and the range is [0.5°, 10°]. The initial value is 17°, and the range is [5°, 25°]. The initial value is 6 The range is [1, 40]. ; The initial value is 0.7, and the range is [0.1, 1].

[0110] The polarity positioning performance index can be calculated based on the predicted values ​​determined by the initial parameters and the actual values ​​in the positioning response data, within lateral angle intervals of 0°, ±20°, ±40°, and ±60°. This index includes local polar errors (PE) and quadrant error rate (QE). When calculating the polarity positioning performance index, the calculation is repeated a preset number of times, such as K times, for each target direction, based on the initial parameters. Specifically, the two indices are expressed by the following formulas:

[0111] (11)

[0112] (12)

[0113] In the formula, when calculating the polarity positioning performance index of the true value, This represents the actual polar angle (true value) perceived by humans in the k-th trial. When calculating the polarity positioning performance index of the true value, The predicted value representing the polar angle; This represents the target polar angle (theoretical value) in the k-th trial. This represents the local polar response where the absolute error between the polar angle and the target polar angle is less than 90 degrees. The function limits the angle difference to [ Within the range of [180°]. PE represents the root mean square error of trials with polar angle errors less than 90 degrees, and QE represents the percentage of trials with polar angle errors exceeding 90 degrees.

[0114] Step b2: Determine the prediction deviation between the predicted value and the true value based on the polarity positioning performance index. The prediction deviation includes the prediction deviations corresponding to the local polarity error and the quadrant error rate, respectively. Specifically, the prediction deviation is expressed by the following formula:

[0115] (13)

[0116] In the formula, This represents the interval index of the horizontal angle. Indicates the number of horizontal angular intervals. This represents the frequency of sound sources appearing in each lateral angular interval. Furthermore, The representative metrics are PE or QE, with the superscripts A and P indicating actual (real person) performance and predictive performance, respectively.

[0117] Step b3: Determine the joint prediction bias based on the prediction biases corresponding to the local polarity error and quadrant error rates, respectively; specifically, the joint prediction bias is expressed by the following formula:

[0118] (14)

[0119] In the formula, the probability ( and This indicates the listener's positioning performance when they are completely guessing the location of the target.

[0120] Step b4: Adjust the initial parameters based on the relationship between the joint prediction deviation and the joint prediction deviation threshold, and repeat the process of determining the polarity positioning performance index, prediction deviation, joint prediction deviation, and parameter adjustment until the relationship between the joint prediction deviation and the joint prediction deviation threshold meets the preset requirements. The joint prediction deviation threshold is determined based on the threshold of the polarity positioning performance index.

[0121] For the polarity positioning performance index, the real-person positioning performance is determined using the following formula. and predictive positioning performance Is the relative deviation between them less than the target threshold? :

[0122] (15)

[0123] In the formula, the target threshold Including the thresholds corresponding to the two performance metrics respectively and ,in, , Combining these two thresholds, along with equations (14) and (15), we can obtain the joint prediction bias threshold. Then, the joint prediction bias determined based on equation (14) is compared with the joint prediction bias threshold. Convergence to threshold If the following conditions are met, the preset requirements can be fixed; otherwise, the parameters need to be adjusted, and steps S202 to S204 above should be repeated until the joint prediction bias d converges to the threshold. The parameters obtained below will be used as the final parameters.

[0124] Additionally, it should be noted that the ACSM mechanism was not introduced during the parameter adjustment process, as model calibration relies on localization data acquired based on broadband noise stimuli.

[0125] Step 206: Polar direction localization of the binaural input signals is performed based on a personalized fitting model, where the personalized fitting model is the model after parameter fitting. For details, please refer to [link to details]. Figure 1 Step S106 of the illustrated embodiment will not be described again here.

[0126] In this invention, the perceptual likelihood function based on spectral cross-correlation can effectively solve the ill-posed problem of auditory localization models when predicting the direction of dynamic spectral signals. Furthermore, the scheme introduces an adaptive correlation-to-similarity mapping (ACSM) mechanism to simulate human active orientation exploration behavior when perceptual confidence is low.

[0127] As a specific application embodiment of the present invention, such as Figure 2 As shown, the sagittal plane auditory localization method based on Bayesian inference is implemented using the following process:

[0128] Step S1: Obtain the user's personalized head-related transfer functions (HRTFs) and location response data;

[0129] Step S2: Extract the spectral features of the left and right ears from the binaural input signals as representations of internal auditory perception.

[0130] Step S3: According to the Bayesian criterion, the extracted spectral features are cross-correlated with the user's feature templates in the target sagittal plane direction one by one to obtain the perceptual likelihood function in the candidate polar angle direction. Then, it is combined with the spatial prior distribution to form the posterior probability distribution of the target polar angle.

[0131] Step S4: Based on the maximum a posteriori estimation criterion, the direction with the highest probability in the posterior distribution is selected as the polar angle prediction value of the target sound source. Then, response noise is introduced to simulate the inherent uncertainty in the motion response or decision-making process, thereby obtaining the final prediction result of the model.

[0132] Step S5: Search for a set of optimal model parameters within the preset parameter range so that the average polarity positioning error between the model prediction value and the true value is lower than the target threshold, and use this set of parameters as the user's personalized model parameters.

[0133] Step S6: Based on this calibrated personalized model, the polar angle direction of any binaural input signal can be estimated, and the direction of the sound source in the sagittal plane can be determined.

[0134] Taking a specific user as an example, the sagittal auditory localization method based on Bayesian inference includes:

[0135] Localization response data were collected from a hearing-normal subject. The target sound source was uniformly distributed on the surface of a virtual sphere centered on the listener. The target location was described using lateral and polar angles in a binaural polar coordinate system, where the lateral angle ranged from... (Right side) to 90° (left side), the polar angle range is from (Lower front) to 210° (lower back). Local polar errors (PE) and quadrant error rate (QE) of the real-person positioning response data are calculated as the true positioning performance values ​​in seven lateral angle intervals centered at 0°, ±20°, ±40°, and ±60°, spaced 20° apart. Each subject has at least 750 measurement points, and the measurement scale affects the parameter fitting quality. In this embodiment, the sound source signal is 500 ms broadband Gaussian noise.

[0136] Set initial model parameters ( , , , , , The initial values ​​of ) are as follows: , , , , , . Figure 3 A block diagram of a sagittal plane auditory localization model based on Bayesian inference provided in an embodiment of the present invention, as shown below. Figure 3 As shown, head-related impulse responses in different directions are convolved with broadband white noise as binaural input signals, and monoaural spectral feature vectors are calculated according to equations (1) to (3). Additive Gaussian noise (Equation (4)) is then considered to simulate monophonic spectral features that include perceptual ambiguity of the auditory system. Based on this feature calculation process, noiseless monoaural spectral feature vectors are extracted from the head-related transfer functions (HRTFs) of the subject in different directions as matching templates. The spectral features extracted from the binaural input signals are cross-correlated with the user's feature templates on the target sagittal plane one by one (Equation (5)). Finally, the perceptual likelihood function about the candidate polar angle direction is obtained through correlation-similarity mapping (Equation (6)) and binaural weighting (Equation (7)). According to the Bayesian criterion of Equation (8), the posterior distribution about the target polar angle is calculated by combining the spatial prior distribution of Equation (9) with the likelihood function. In the decision stage (Equation (10)), the maximum value of the posterior distribution is searched by the maximum a posteriori estimation method as the polar angle estimate. Finally, the motion response representing the inherent uncertainty of human auditory perception is added to obtain the polar angle prediction value. .

[0137] In this embodiment, the existing HRTF dataset of the subjects is resampled to 1500 directions to be compatible with the non-uniform HRTF acquisition grid, and 300 simulations are repeated for the 1500 directions to simulate the inherent randomness in model estimation. In this embodiment, the lateral angle of the target sound source is assumed to be known during the model prediction process to focus on the polar angle prediction performance of the model. In this embodiment, the correlation-similarity mapping of equation (6) can adopt an adaptive version. After enabling this mechanism, the model can dynamically adjust the parameters according to the cross-correlation results of the real-time spectrum. and This enhances the focus on the peak region of correlation. This mechanism helps improve the model's performance in complex auditory environments, making it closer to human auditory behavior of actively exploring the environment, and enhancing the consistency and rationality between the model and the actual perception process.

[0138] Based on real-person positioning data and model predictions, PE (Equation (11)) and QE (Equation (12)) within different lateral angle intervals are calculated as the true and predicted values ​​of positioning performance, respectively. The prediction deviation between the model predictions and human response data under the two performance indices is calculated according to Equation (13). Subsequently, the joint prediction bias is obtained by weighting and integrating the performance indicators according to their respective occurrence probabilities. (Equation (14)). Based on the relative deviation thresholds corresponding to the two performance indicators. (Equation (15)) calculates the joint prediction residual threshold Search for a set of optimal model parameters within a given parameter range, such that the joint prediction bias is below a target threshold. This set of parameters is then used as the personalized model parameters for the user. In this embodiment, to minimize prediction bias... The fmincon function with interior point method in Matlab was adopted, and a multi-starting point optimization strategy was applied to run local optimization algorithms from multiple initial points to find the global optimal solution.

[0139] Five models calibrated with personalized parameters were used as "digital listeners" in five localization experiments to evaluate the impact of different factors on model localization performance at an overall level. Experiments 1–3 replicated the localization experimental design from previous studies, evaluating the effects of high-frequency speech attenuation, non-personalized HRTF, and spectral ripple on localization performance, and obtaining corresponding real-person data for comparison with the prediction results of the localization model proposed in this invention. Experiments 4–5 aimed to explore the impact of the adaptive correlation-similarity mapping mechanism on model localization performance in complex acoustic environments, specifically including different signal-to-noise ratio (SNR) conditions and the absence of high-frequency cues. Listener-specific fitting parameters were used in the first three experiments, while parameters related to similarity mapping were used in experiments 4–5. , It will be dynamically adjusted according to the Adaptive Correlation-to-Similarity Mapping (ACSM) mechanism.

[0140] Figure 4 This study demonstrates the impact of varying degrees of high-frequency speech attenuation on the model's localization performance. The experimental stimulus consisted of 260 monosyllabic speech samples selected from a wideband corpus, with an average duration of 710 ms and a frequency range of 0.3–16 kHz. The speech was passed through an 8 kHz cutoff frequency with 0 dB stopband attenuation. dB and The stimulus was low-pass filtered to dB, and broadband Gaussian white noise was used as the baseline stimulus for comparison. Auditory localization experiments were conducted by uniformly placing the stimulus signals at 76 random locations on a virtual sphere around the listener's head, with each target direction repeated 50 times. The mean absolute polar angle error (MAPE) and QE (Equation (12)) were used as performance metrics to evaluate the model.

[0141]

[0142] Where the arctangent function By considering and The sign is used to ensure correct calculations in the four quadrants. The final experimental results show that the model's predictive performance exhibits good consistency with real-world data, specifically that the polarity positioning error gradually increases with the degree of high-frequency attenuation.

[0143] Figure 5 This study demonstrates the impact of non-personalized HRTFs on model localization performance. Experiments simulated the localization performance (PE and QE) of five digital listeners using their own (Own) and others' (Other) HRTFs within a ±30° lateral angle range on 250ms Gaussian white noise. Experimental results show that using non-personalized virtual auditory playback significantly increases polarity localization error.

[0144] Figure 6 This study demonstrates the impact of spectral ripple sources on polar localization performance. Experiments simulated the localization performance of five digital listeners when locating sound source signals with different ripple densities and depths. Under baseline conditions, the model's localization performance was slightly lower than that of human data. The polar error rate (PER) was used to evaluate the model's sagittal localization performance, achieved through statistical analysis of model responses with linear predictions deviating from baseline regression by more than 45 degrees. Under ripple density conditions, as density increased, the PER increase for human listeners reached its maximum near one ripple per octave. Subsequently, the PER increase gradually decreased, approaching zero at eight ripples per octave.

[0145] Figure 7 This figure shows the speech localization response of a typical subject under seven different SNR conditions, with and without ACSM enabled. Each column in the figure represents the localization prediction results under different SNR conditions. In the polar angle localization scatter plot, dots located on or near the positive diagonal represent correct or near-correct polar angle localization responses to the target sound source direction, while dots distributed near the negative diagonal represent responses with confusion. At lower SNRs, the polar angle response without ACSM mainly clusters around 180°, while the response with ACSM enabled clusters around... Around 180°, the regression fitting bias is shifted to 90°. These results indicate that ACSM helps prevent the polar response from being pulled towards the hemisphere with larger prior variance, making it closer to human orientation exploration behavior under low perceptual confidence. The average speech localization results from five digital listeners are... Figure 8 The diagram shows three sub-figures illustrating the impact of SNR on the model's localization performance. The model's localization performance is determined by the localization gain after linear regression analysis (…). Positioning deviation ( and positioning accuracy (determination coefficient) )Evaluate:

[0146] (17)

[0147] Among them, response Representative target Predicted polar angle, regression parameters and These represent positioning gain and bias, respectively, and can be obtained by analyzing all... - The data is estimated using the least squares method. The coefficient of determination... This can be approximated by calculating the square of the Pearson correlation coefficient between the target and the response. It can be observed that the model's localization performance gradually improves with increasing SNR after enabling ACSM, which aligns with human auditory perception patterns.

[0148] Figure 9 The simulation results, compiled from five digital listeners, show low-pass speech localization with and without ACSM enabled. The simulation results indicate that the model prediction performance with ACSM enabled is close to that under lower SNR conditions. The prediction results under the condition that the ACSM mechanism can simulate human orientation exploration behavior in scenarios with low perceptual confidence are shown.

[0149] This embodiment also provides a sagittal plane auditory localization device based on Bayesian inference, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0150] This embodiment provides a sagittal plane auditory localization device based on Bayesian inference, such as... Figure 10 As shown, it includes:

[0151] Data acquisition module 11 is used to acquire the user's personalized header-related transmission functions and positioning response data;

[0152] Feature extraction module 12 is used to extract the spectral features of the left and right ears from the binaural input signals respectively;

[0153] Bayesian inference module 13 is used to perform cross-correlation calculation between spectral features and feature templates in each polar angle direction on the target sagittal plane according to Bayesian criteria, determine the perceptual likelihood function in each polar angle direction, and combine the perceptual likelihood function with the spatial prior distribution to determine the posterior probability distribution of each polar angle. The feature template is determined by the personalized head correlation transfer function.

[0154] The perception and decision module 14 is used to select the polar angle prediction value of the model from the posterior probability distribution based on the maximum a posteriori estimation criterion. The model is a model constructed based on spectral feature extraction and Bayesian inference.

[0155] The parameter fitting module 15 is used to fit the model parameters to minimize the deviation between the model prediction and the actual positioning response under a preset threshold.

[0156] The direction prediction module 16 is used to perform polar angle direction positioning on the binaural input signal based on a personalized fitting model, wherein the personalized fitting model is a model after parameter fitting.

[0157] Further functional descriptions of the above modules are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0158] In one alternative implementation, the data acquisition module is specifically used to: obtain the user's personalized head-related transfer function based on personalized measurements or physiological parameters; perform a psychoacoustic localization experiment to collect the user's localization response data in all acquisition directions on a spherical grid, and resample the user's personalized head-related transfer function to match the spatial resolution of the localization response data.

[0159] In one optional implementation, the feature extraction module is specifically used to: extract the initial feature vector of the binaural input signal based on the gamma-pass filter bank of the approximate basilar membrane filter and the half-wave rectification of the hair cell conduction; determine the spectral amplitude profile feature based on the root mean square amplitude of the initial feature vector; and determine the noisy feature by combining the spectral amplitude profile feature and additive Gaussian noise, with the noisy feature serving as the spectral feature.

[0160] In one alternative implementation, the Bayesian inference module includes:

[0161] The correlation vector determination module is used to calculate the cross-correlation coefficient based on the spectral features and the feature templates of each polar angle direction on the target sagittal plane. The positive correlation coefficient is retained and the non-positive correlation coefficient is set to zero to obtain the one-sided correlation vectors related to each polar angle direction. The one-sided correlation vectors include the left correlation vector and the right correlation vector.

[0162] The similarity vector determination module is used to map a one-sided correlation vector into a similarity vector related to each polar angle direction using a psychometric function. The similarity vector includes a left similarity vector and a right similarity vector. The perceptual selectivity parameter and perceptual sensitivity parameter of the psychometric function are either fixed values ​​or adaptively adjusted using an adaptive correlation-similarity mapping mechanism.

[0163] The perceptual likelihood function determination module is used to perform weighted summation and normalization of the left and right similarity vectors using a binaural weighting function to obtain the perceptual likelihood function related to each polar angle direction.

[0164] The posterior distribution determination module is used to determine the posterior probability distribution of each polar angle by combining the perceptual likelihood function and the spatial prior distribution.

[0165] In an optional implementation, when the perceptual selectivity parameter and the perceptual sensitivity parameter are adaptively adjusted using an adaptive correlation-similarity mapping mechanism, the similarity vector determination module is specifically used to: sort the one-sided correlation vectors in each polar angle direction in descending order, and determine the maximum value and the second peak value of the one-sided correlation vector; perform linear regression on a preset psychometric function based on the maximum value of the one-sided correlation vector and the preset maximum value of the similarity vector, as well as the second peak value of the one-sided correlation vector and the preset minimum value of the similarity vector, to determine the adaptive parameters; and map the one-sided correlation vectors to similarity vectors related to each polar angle direction based on the psychometric function determined by the adaptive parameters.

[0166] In one alternative implementation, the perception decision module is specifically used to: select the direction with the highest probability in the posterior distribution as the polar angle prediction value of the target sound source according to the maximum a posteriori estimation criterion, and introduce response noise, which follows a von Mises-Fischer distribution.

[0167] In one alternative implementation, the parameter fitting module includes:

[0168] The prediction and true value determination module is used to take the polar angle prediction containing response noise as the prediction value and locate the polar angle direction in the response data as the true value.

[0169] The adjustment submodule is used to adjust the preset parameters so that the deviation between the predicted value and the true value meets the preset requirements. The preset parameters include the spectral feature extraction stage, the Bayesian inference stage, and the free parameters in the response noise.

[0170] In one optional implementation, the adjustment submodule is specifically used for: calculating the polarity positioning performance index corresponding to the predicted value and the true value respectively, the polarity positioning performance index including local polarity error and quadrant error rate, the predicted value being the model prediction value based on the current model parameters; determining the prediction deviation between the predicted value and the true value based on the polarity positioning performance index, the prediction deviation including the prediction deviation corresponding to the local polarity error and quadrant error rate respectively; determining the joint prediction deviation based on the prediction deviation corresponding to the local polarity error and quadrant error rate respectively; adjusting the initial parameters based on the relationship between the joint prediction deviation and the joint prediction deviation threshold, and repeating the process of determining the polarity positioning performance index, prediction deviation, joint prediction deviation, and parameter adjustment until the relationship between the joint prediction deviation and the joint prediction deviation threshold meets the preset requirements, the joint prediction deviation threshold being determined based on the threshold of the polarity positioning performance index.

[0171] This invention also provides a computer device having the above-described features. Figure 10 The device shown is a sagittal plane auditory localization device based on Bayesian inference.

[0172] Please see Figure 11, Figure 11 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 11 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 11 Take a processor 10 as an example.

[0173] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0174] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0175] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0176] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0177] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0178] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0179] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0180] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A sagittal plane auditory localization method based on Bayesian inference, characterized in that, The method includes: Obtain the user's personalized header-related transmission functions and location response data; Extract the spectral features of the left and right ears from the binaural input signals respectively; According to the Bayesian criterion, the spectral features are cross-correlated with the feature templates in each polar angle direction on the target sagittal plane to determine the perceptual likelihood function in each polar angle direction. The posterior probability distribution of each polar angle is determined by combining the perceptual likelihood function with the spatial prior distribution. The feature template is determined by the personalized head-related transfer function. The polar angle prediction value of the model is selected from the posterior probability distribution based on the maximum a posteriori estimation criterion. The model is a model constructed based on spectral feature extraction and Bayesian inference. Fit the model parameters to minimize the deviation between the model prediction and the actual localization response under a preset threshold; Polar direction localization of binaural input signals is performed based on a personalized fitting model, wherein the personalized fitting model is a model after parameter fitting.

2. The method according to claim 1, characterized in that, Obtain the user's personalized header-related transmission functions and location response data, including: The user's personalized head-related transfer function is obtained based on personalized measurements or physiological parameter matching. A psychoacoustic localization experiment was conducted to collect localization response data of the user in all acquisition directions on a spherical grid, and the user's personalized head-related transfer function was resampled to match the spatial resolution of the localization response data.

3. The method according to claim 1, characterized in that, The spectral features of the left and right ears are extracted from the binaural input signals, including: The initial feature vector of the binaural input signal is extracted based on the gamma-tongue filter bank of approximate basilar membrane filtering and the half-wave rectification of hair cell conduction. The spectral amplitude profile features are determined based on the root mean square amplitude of the initial feature vector. The noisy features are determined by combining the spectral amplitude profile features and additive Gaussian noise, and the noisy features are used as spectral features.

4. The method according to claim 1, characterized in that, According to the Bayesian criterion, the spectral features are cross-correlated with feature templates in each polar angle direction on the target sagittal plane to determine the perceptual likelihood function in each polar angle direction. The posterior probability distribution of each polar angle is then determined by combining the perceptual likelihood function with the spatial prior distribution, including: Based on the spectral features and the feature templates of each polar angle direction on the target sagittal plane, the cross-correlation coefficients are calculated, positive correlation coefficients are retained, and non-positive correlation coefficients are set to zero, to obtain the one-sided correlation vectors related to each polar angle direction. The one-sided correlation vectors include the left correlation vector and the right correlation vector. The unilateral correlation vector is mapped to a similarity vector related to each polar angle using a psychometric function. The similarity vector includes a left similarity vector and a right similarity vector. The perceptual selectivity parameter and perceptual sensitivity parameter of the psychometric function are either fixed values ​​or adaptively adjusted using an adaptive correlation-similarity mapping mechanism. The left and right similarity vectors are weighted and summed using a binaural weighting function and then normalized to obtain the perceptual likelihood function related to each polar angle direction. The posterior probability distribution of each polar angle is determined by combining the perceived likelihood function and the spatial prior distribution.

5. The method according to claim 4, characterized in that, When the perceptual selectivity parameter and perceptual sensitivity parameter are adaptively adjusted using an adaptive correlation-similarity mapping mechanism, a psychometric function is used to map the one-sided correlation vector into similarity vectors related to each polar angle direction, including: Arrange the one-sided correlation vectors in descending order for each polar angle direction, and determine the maximum value and the second peak value of the one-sided correlation vectors; Based on the maximum value of the one-sided correlation vector, the maximum value of the preset similarity vector, the second peak value of the one-sided correlation vector, and the minimum value of the preset similarity vector, a linear regression is performed on the preset psychometric function to determine the adaptive parameters, which include the perceptual selectivity parameter and the perceptual sensitivity parameter. The one-sided correlation vector is mapped to a similarity vector related to each polar angle direction based on the psychometric function determined by adaptive parameters.

6. The method according to claim 1, characterized in that, The polar angle predictions of the model are selected from the posterior probability distribution based on the maximum a posteriori estimation criterion, including: According to the maximum a posteriori estimation criterion, the direction with the highest probability in the posterior distribution is selected as the polar angle prediction value of the target sound source, and response noise is introduced, which follows the von Mises-Fischer distribution.

7. The method according to claim 6, characterized in that, Fitting model parameters to minimize the deviation between model predictions and actual localization responses within a preset threshold includes: The polar angle prediction, which includes response noise, is used as the predicted value, and the polar angle direction in the response data is used as the true value. The preset parameters are adjusted so that the deviation between the predicted value and the true value meets the preset requirements. The preset parameters include the spectral feature extraction stage, the Bayesian inference stage, and the free parameters in the response noise.

8. The method according to claim 7, characterized in that, Adjust the preset parameters so that the deviation between the predicted value and the actual value meets the preset requirements, including: Calculate the polarity positioning performance index corresponding to the predicted value and the true value respectively. The polarity positioning performance index includes local polarity error and quadrant error rate. The predicted value is the model prediction value corresponding to the current model parameters with preset parameters. The prediction deviation between the predicted value and the true value is determined based on the polarity positioning performance index. The prediction deviation includes the prediction deviations corresponding to the local polarity error and the quadrant error rate, respectively. The joint prediction bias is determined based on the prediction biases corresponding to local polarity error and quadrant error rate, respectively. The initial parameters are adjusted based on the relationship between the joint prediction deviation and the joint prediction deviation threshold. The process of determining the polarity positioning performance index, prediction deviation, joint prediction deviation, and parameter adjustment is repeated until the relationship between the joint prediction deviation and the joint prediction deviation threshold meets the preset requirements. The joint prediction deviation threshold is determined based on the threshold of the polarity positioning performance index.

9. A sagittal plane auditory localization device based on Bayesian inference, characterized in that, The device includes: The data acquisition module is used to acquire the user's personalized header-related transmission functions and location response data; The feature extraction module is used to extract the spectral features of the left and right ears from the binaural input signals, respectively. The Bayesian inference module is used to perform cross-correlation calculations between the spectral features and the feature templates in each polar angle direction on the target sagittal plane according to the Bayesian criterion, determine the perceptual likelihood function in each polar angle direction, and combine the perceptual likelihood function with the spatial prior distribution to determine the posterior probability distribution of each polar angle. The feature templates are determined by the personalized head-related transfer function. The perception and decision module is used to select the polar angle prediction value of the model from the posterior probability distribution based on the maximum a posteriori estimation criterion. The model is a model constructed based on spectral feature extraction and Bayesian inference. The parameter fitting module is used to fit the model parameters to minimize the deviation between the model prediction and the actual localization response under a preset threshold. The direction prediction module is used to locate the polar angle direction of the binaural input signal based on a personalized fitting model, wherein the personalized fitting model is a model after parameter fitting.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the sagittal auditory localization method based on Bayesian inference as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Transformer sound source localization imaging method, system and device based on sparse Bayesian learning model and storage medium

    CN119044891A

  • Sound source position estimation device, sound source position estimation method, and program

    JP2018063200A