Adaptive hearing screening optimization method based on acoustic feature analysis
By constructing multimodal auditory feature vectors and adopting deep neural network models, the challenges in the existing technology in noise suppression, spatial auditory ability quantification and personalized hearing screening are solved, and personalized listening evaluation is achieved, and the accuracy of hearing screening results is improved.
Patent Information
- Application Number
- CN202510431366.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art has great challenges in noise suppression, spatial auditory ability quantification and personalized hearing screening, especially when facing the differences in hearing status and hearing ability of different users, individual differences cannot be fully considered, resulting in limited accuracy of evaluation results.
By constructing an audio library, generating audio data with different signal-to-noise ratios, combining probe microphones to measure the acoustic transfer function of the user's ear canal, combining EEG data to extract time domain and time frequency features, building a multimodal auditory feature vector, and using a multi-head attention mechanism and deep neural network model for hearing threshold prediction, and finally predicting the user's listening age through a timing Transformer encoder.
It achieves personalized and accurate listening evaluation in a changing listening environment, overcomes the defects of ignoring individual differences and complex environmental noise in traditional methods, and improves the accuracy of hearing screening results.
Smart Images

Figure CN119924826A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of hearing test, and more specifically, to an adaptive hearing screening optimization method based on acoustic feature analysis. Background Art
[0002] Hearing impairment has become an increasingly serious public health problem worldwide, especially among the elderly, where hearing loss has a significant impact on quality of life. According to statistics from the World Health Organization, about 400 million people worldwide suffer from hearing impairment to varying degrees. With the aging of the population, the demand for screening and early intervention for hearing impairment is increasing. Accurate hearing screening not only helps to detect hearing problems early, but also provides data support for related intervention measures, thereby effectively slowing down or delaying the process of hearing loss.
[0003] Among traditional hearing screening methods, the most common is pure tone audiometry, which evaluates an individual's hearing threshold at different frequencies by playing pure tones of standard frequencies. Although this method is simple and widely used, it relies on the subjective judgment of hearing experts and is easily affected by factors such as environmental noise and equipment accuracy, resulting in limited accuracy and reliability of test results. In addition, pure tone audiometry usually ignores an individual's spatial hearing ability, the impact of environmental noise, and the complex interactions between different frequencies. Therefore, its accuracy is greatly reduced when faced with noisy environments and complex hearing damage.
[0004] In recent years, with the advancement of computer technology and artificial intelligence, hearing screening methods based on acoustic feature analysis have gradually gained attention. These methods analyze the user's auditory response by using different acoustic signals (such as pure tones, speech, and environmental noise) combined with advanced signal processing technology.
[0005] Although modern technology has made certain progress in time domain feature extraction, time-frequency feature analysis, and multimodal signal fusion, existing technologies still face great challenges in noise suppression, spatial hearing ability quantification, and personalized screening. Especially in practical applications, the impact of environmental noise on hearing test results cannot be ignored, and the accuracy of existing technologies in effectively distinguishing noise and signals and integrating personalized hearing characteristics is often not ideal. In addition, most existing technologies still find it difficult to achieve real-time, dynamic, and personalized hearing assessments, especially when there are large differences in the hearing conditions and hearing abilities of different users. Individual differences cannot be fully considered, resulting in limited accuracy of the assessment results. Summary of the invention
[0006] The present invention provides an adaptive hearing screening optimization method based on acoustic feature analysis, which is intended to solve the current technical problem that when the hearing conditions and hearing abilities of different users vary greatly, individual differences cannot be fully considered, resulting in limited accuracy of evaluation results.
[0007] The adaptive hearing screening optimization method based on acoustic feature analysis includes the following steps: Step 1: Based on the constructed audio library, pure tones from 125 Hz to 8 kHz, standard speech materials, and various types of environmental noise are generated, and the noise and target signal are superimposed by setting different signal-to-noise ratios to form audio data with various signal-to-noise ratios; Step 2: Use a probe microphone to measure the acoustic transfer function of the user's ear canal, and generate a personalized acoustic transfer function based on the head-related transfer function database; Step 3: Play the calibrated audio through the binaural simulation microphones equipped on the artificial head, and record the binaural audio signals. At the same time, synchronize the audio playback and the EEG device through the optical coupling hardware trigger to obtain EEG data, and pre-process the EEG data to extract time domain features and time-frequency features based on the pre-processed EEG data; Step 4: Processing binaural audio signals based on filters to obtain time domain envelopes of each frequency band, and then calculating the dynamic range compression ratio of the time domain envelopes of each frequency band; calculating the time difference and energy ratio of binaural signals based on the cross-correlation method, and using personalized acoustic transfer functions and standard acoustic transfer function libraries to perform similarity calculations, quantifying the user's spatial positioning ability, and obtaining spatial auditory features; constructing a unified multimodal auditory feature vector based on the time domain features, time-frequency features, and spatial auditory features; Step 5: Based on the multimodal auditory feature vector, a multi-head attention mechanism is used to weight the multimodal feature vector to obtain a weighted multimodal feature vector; Step 6: Based on the weighted multimodal feature vector and frequency / intensity parameters, the hearing threshold is predicted through a deep neural network model, the threshold deviation of each frequency point is output, and then the standard threshold is obtained by combining it with the traditional pure tone audiometry method; Step 7: Based on the multimodal feature vector, the actual age of the user and the standard threshold, a temporal Transformer encoder is used to predict the hearing age of the user, and a hearing decline assessment is performed on the user based on the predicted hearing age.
[0008] The present invention generates audio data with different signal-to-noise ratios based on a constructed audio library, and combines environmental noise and target signals to ensure that the hearing test needs under various noise conditions can be covered; secondly, a probe microphone is used to measure the acoustic transfer function of the user's ear canal, and a personalized acoustic transfer function is generated in combination with the head-related transfer function, so as to accurately reflect the individual's ear canal acoustic characteristics and avoid the one-size-fits-all assumption in traditional methods; then, the time domain and time-frequency features are extracted with the help of EEG data, and the spatial positioning features of the binaural simulation microphone are combined to further quantify the user's auditory response ability and spatial hearing ability; through the construction of a multimodal auditory feature vector, the features are weighted in combination with a multi-head attention mechanism, so as to more accurately predict the hearing threshold; finally, the user's hearing age is predicted by a time series Transformer encoder in combination with the user's age and the standard threshold, and the hearing loss is evaluated based on the prediction results; this enables the method to accurately adapt to the personalized differences of each user, overcome the defects of ignoring individual differences and complex environmental noise in traditional hearing screening methods, thereby improving the accuracy of screening results and ensuring that effective hearing assessment can be provided in a changing hearing environment.
[0009] Preferably, step 2 comprises the following steps: Measuring pure tone signals: Playing multi-band pure tone stimulation through an external speaker, and inserting a probe microphone into the ear canal to record the sound wave response of the pure tone stimulation after passing through the ear canal to obtain a measurement signal; Calculate the acoustic transfer function of the ear canal: Based on the input signal and the measured output signal, the acoustic transfer function of the ear canal is calculated in the frequency domain: ; Where: is the frequency response of the ear canal, i.e., the acoustic transfer function; and Represent the frequency domain of the input signal and the output signal respectively; User HRTF acquisition: inferring the applicable HRTF function from the standard HRTF database based on the user's individual physical signs; Synthesize personalized acoustic transfer function: The personalized acoustic transfer function of the user is obtained based on the product of the inferred HRTF function and the calculated acoustic transfer function of the ear canal.
[0010] Preferably, after designing the personalized ear canal transfer function in step 2, using an FIR filter to reversely compensate the headphone frequency response includes the following steps: Measure the frequency response of headphones: Play a standard pure tone signal through the headphones, use a probe microphone to record the signal output by the headphones, and calculate the frequency response of the headphones based on the signal output by the headphones: Where: represents the frequency domain representation of the headphone output signal; represents the frequency domain representation of the input signal; Indicates the frequency response of the headphones; FIR filter: Compensate the headphone output through the FIR filter: ; Where: represents the frequency response of the filter; Obtain FIR filter coefficients based on frequency response using discrete Fourier transform , the output signal is inversely compensated based on the filter coefficients: ; Where: Represents the convolution operation; represents the time domain representation of the input signal; represents the time domain representation of the output signal.
[0011] Preferably, the steps of acquiring and preprocessing EEG data in step 3 are as follows: Binaural audio signal playback: The calibrated audio signal is played through an artificial binaural simulation microphone. The audio signal includes multiple pure tones, standard speech materials, and multiple types of environmental noise. Different signal-to-noise ratios are set to form different test scenarios. During the test, the binaural audio signal is synchronized with the EEG device through an optically coupled hardware trigger. The EEG device records the brain's neural response to the binaural audio signal, while ensuring the time synchronization between the audio playback and the EEG recording. EEG signal preprocessing: perform multi-scale wavelet transform on the EEG signal to decompose the signal into sub-signals of multiple scales; perform soft threshold denoising on the sub-signal of each scale to remove high-frequency noise; perform adaptive filtering on the denoised signal to remove noise introduced by artifacts; reconstruct the processed signal through inverse wavelet transform to obtain the preprocessed EEG signal.
[0012] Preferably, the extracting of time domain features and time-frequency features comprises the following steps: Time domain feature extraction: The pre-processed EEG signal is subjected to time domain analysis to extract the average potential, amplitude and waveform complexity features to obtain the extracted time domain features; Time-frequency feature extraction: Wavelet transform is used to perform time-frequency analysis on the preprocessed EEG signal to extract the features of each frequency band.
[0013] Preferably, step 4 comprises the following steps: Filter processing of binaural audio signals: using a bandpass filter to decompose binaural audio signals into multiple frequency bands, each of which contains the energy distribution of the audio signal in that frequency band; Time domain envelope extraction and dynamic range compression ratio: In each frequency band, the time domain envelope of the signal is extracted by the envelope detector. When extracting, the Hilbert transform is used to obtain the time domain envelope; the dynamic range compression ratio of the time domain included in each frequency band is calculated; Cross-correlation calculates the time difference and energy ratio of binaural signals: the time difference is estimated by calculating the delay between binaural signals; the energy ratio is obtained by calculating the mean square value of binaural signals; Similarity calculation: The similarity between the personalized acoustic transfer function and the transfer function in the standardized acoustic transfer function library is calculated; the similarity uses the mutual normalized correlation as the similarity measure; Construct a multimodal auditory feature vector: Concatenate all features extracted from time domain features, time-frequency features, and spatial auditory features to construct a unified multimodal auditory feature vector.
[0014] Preferably, when the multi-head attention mechanism weights the multimodal feature vector, feature importance weights are introduced to adjust the attention weights, including the following steps: Feature importance calculation: The feature importance is calculated based on the variance of each sub-feature in the input vector in the multi-head attention mechanism, the correlation of the threshold, and the robustness under different noise conditions: ; ; ; Where: , , Respectively represent the variance of time domain features, time-frequency features, and spatial auditory features; , , They represent the correlation between time domain features, time-frequency features, and spatial auditory features respectively; , , They represent the robustness of time domain features, time-frequency features, and spatial auditory features respectively; , , represents the hyperparameters used to adjust the contribution of variance, correlation, and robustness to the feature importance calculation; Based on the importance of the calculated time domain features, time-frequency features, and spatial auditory features, normalization is performed to obtain the feature importance weights of each modality; Combining feature importance with attention weight: When weighting the query and key of each modality, feature importance weight is introduced to obtain attention weight: ; Where: Represents the feature importance weight of each modality; represents the inner product of the query vector and the key vector; represents the scaling factor; express Function, the inner product result is used Function that ensures that all attention weights are positive and sum to 1.
[0015] Preferably, step 6 comprises the following steps: Feature construction: concatenate the weighted multimodal feature vector and the frequency / intensity parameter to obtain a concatenated feature vector; Hearing threshold deviation prediction: using the concatenated feature vector as the input of a trained deep neural network model, and predicting the hearing threshold deviation through the deep neural network model; Pure tone audiometry threshold: Hearing threshold is obtained based on pure tone audiometry; Threshold fusion: The hearing threshold and hearing threshold deviation obtained by pure tone audiometry are summed, and the summed value is weighted and fused with the hearing threshold obtained by pure tone audiometry to obtain the standard threshold. The weights in weighted fusion are obtained based on Bayesian optimization, and the objective function of Bayesian optimization is as follows: ; Where: represents the actual hearing threshold of the ith frequency point; Represents the hearing threshold of the ith frequency point after fusion.
[0016] Preferably, the temporal Transformer encoder comprises an input embedding layer, a position encoding layer, a self-attention mechanism layer, a feedforward neural network layer and an output layer; The features fused by the weighted multimodal feature vector and the frequency / intensity parameters are input to the input embedding layer, mapped through a fully connected layer to obtain the feature representation of each time step; and then the sequence information is introduced through the position encoding layer; The features processed by the position encoding layer are used to capture the dependencies between time steps in the time series through the self-attention mechanism layer, and the weighted relationship between each time step and other time steps is calculated to obtain the attention weight of each time step. Based on the attention weight of each time step, the weighted representation of each time step is obtained; the weighted representation of each time step is input into the feedforward neural network layer for nonlinear change, and the feature representation after nonlinear change enters the output layer and is mapped to the predicted hearing age through a fully connected layer.
[0017] The beneficial effects of the present invention include: The present invention generates audio data with different signal-to-noise ratios based on a constructed audio library, and combines environmental noise and target signals to ensure that the hearing test needs under various noise conditions can be covered; secondly, a probe microphone is used to measure the acoustic transfer function of the user's ear canal, and a personalized acoustic transfer function is generated in combination with the head-related transfer function, so as to accurately reflect the individual's ear canal acoustic characteristics and avoid the one-size-fits-all assumption in traditional methods; then, the time domain and time-frequency features are extracted with the help of EEG data, and the spatial positioning features of the binaural simulation microphone are combined to further quantify the user's auditory response ability and spatial hearing ability; through the construction of a multimodal auditory feature vector, the features are weighted in combination with a multi-head attention mechanism, so as to more accurately predict the hearing threshold; finally, the user's hearing age is predicted by a time series Transformer encoder in combination with the user's age and the standard threshold, and the hearing loss is evaluated based on the prediction results; this enables the method to accurately adapt to the personalized differences of each user, overcome the defects of ignoring individual differences and complex environmental noise in traditional hearing screening methods, thereby improving the accuracy of screening results and ensuring that effective hearing assessment can be provided in a changing hearing environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0019] Figure 1 An overall step block diagram provided for an embodiment of the present invention.
[0020] Figure 2 A schematic block diagram of the overall data processing logic provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] See also Figure 1 and Figure 2 As shown, the adaptive hearing screening optimization method based on acoustic feature analysis includes the following steps: Step 1: Based on the constructed audio library, pure tones from 125 Hz to 8 kHz, standard speech materials, and various types of environmental noise are generated, and the noise and target signal are superimposed by setting different signal-to-noise ratios to form audio data with various signal-to-noise ratios; Exemplarily, the audio library contains three types of audio content: Pure Tone: Select pure tones with a frequency range from 125Hz to 8kHz (including low frequency, medium frequency and high frequency). The duration of each pure tone signal is selected to be 500ms or greater to ensure accurate frequency response. Standard speech materials: Choose standard language speech materials, such as letters, numbers, common words or short sentences; Environmental noise: such as city noise, home background noise, office noise, etc.; Signal-to-noise ratio (SNR) settings: By setting different signal-to-noise ratios, the hearing challenges in different environments are simulated. The SNR settings range from low signal-to-noise ratio to high signal-to-noise ratio. The calculation formula for the signal-to-noise ratio is as follows: ; Where: Indicates the power of the signal; represents the power of the noise; represents the signal-to-noise ratio, expressed in decibels; In this embodiment, in order to cover a variety of listening scenarios, we can select multiple SNR values, such as -10dB, 0dB, 10dB, 20dB, etc., to form audio samples with different noise interference intensities; Audio signal synthesis: Different types of environmental noise are superimposed with pure sound and standard speech materials to form the final audio sample. The specific steps are as follows: Pure tone and noise synthesis: select a pure tone signal of a frequency (e.g. 1000Hz) and iterate it with the ambient noise according to the selected SNR value; for example, at -dB SNR, the power of the noise will be much higher than that of the pure tone signal, while at 30dB SNR, the pure tone signal will be much stronger than the noise; Standard speech and noise synthesis: Similar to the synthesis of pure tone and noise, the standard speech material is superimposed with different noise signals to generate speech signals with different noise interferences; The above-mentioned pure tone and noise synthesis, standard speech and noise synthesis can be mixed according to the following formula: ; Where: Represents the mixed audio signal; represents the original target signal (pure tone or speech signal); represents a noise signal; Represents the noise scaling factor, which is used to control the ratio of noise intensity to signal intensity and is calculated based on the selected SNR: ; Based on the above process, a multi-SNR, multi-class audio dataset is generated; the audio dataset includes different noise types, pure tone frequency ranges, speech content, and different SNR settings.
[0023] Step 2: Use a probe microphone to measure the acoustic transfer function of the user's ear canal, and generate a personalized acoustic transfer function in combination with a head-related transfer function database; The step 2 comprises the following steps: Measuring pure tone signals: Play pure tone stimulation of multiple frequencies (e.g. 125Hz, 250Hz, etc.) through an external speaker, and insert a probe microphone into the ear canal to record the sound wave response of the pure tone stimulation after passing through the ear canal to obtain the measurement signal. ; Calculate the acoustic transfer function of the ear canal: Assume and The input signal and the output signal are represented in the frequency domain. Based on the input signal and the measured output signal, the acoustic transfer function of the ear canal is calculated in the frequency domain: ; Where: is the frequency response of the ear canal, i.e., the acoustic transfer function; and Represent the frequency domain of the input signal and the output signal respectively; The corresponding time domain signal is calculated as: ; in: represents inverse Fourier transform; User HRTF acquisition: inferring the applicable HRTF function from the standard HRTF database based on the user's individual physical signs; Exemplarily, the individual characteristics of the user include the length and width of the head, the length of the ear canal, and the height of the auricle; Calculate the Euclidean distance between the feature vector formed by the user's individual features and the feature vector of each user in the database to obtain the similarity, and select the N individual data with the highest similarity; Perform weighted averaging on the N individual data to obtain the inferred HRTF function ; Synthesize personalized acoustic transfer function: The user's personalized acoustic transfer function is obtained by multiplying the inferred HRTF function and the calculated acoustic transfer function of the ear canal: ; Where: represents the inferred HRTF function; As a further implementation of this embodiment, after designing the personalized ear canal transfer function in step 2, using an FIR filter to reversely compensate for the headphone frequency response includes the following steps: Measure the frequency response of headphones: Play a standard pure tone signal through the headphones, use a probe microphone to record the signal output by the headphones, and calculate the frequency response of the headphones based on the signal output by the headphones: ; Where: represents the frequency domain representation of the headphone output signal; represents the frequency domain representation of the input signal; Indicates the frequency response of the headphones; FIR filter: Compensate the headphone output through the FIR filter: ; Where: represents the frequency response of the filter; Obtain FIR filter coefficients based on frequency response using discrete Fourier transform , the output signal is inversely compensated based on the filter coefficients: ; Where: Represents the convolution operation; represents the time domain representation of the input signal; represents the time domain representation of the output signal.
[0024] Step 3: Play the calibrated audio through the binaural simulation microphones equipped on the artificial head, and record the binaural audio signals. At the same time, synchronize the audio playback and EEG equipment through the optical coupling hardware trigger to obtain electroencephalogram (EEG) data, and pre-process the EEG data to extract time domain features and time-frequency features based on the pre-processed EEG data; The steps of obtaining and preprocessing EEG data in step 3 are as follows: Binaural audio signal playback: The calibrated audio signal is played through an artificial binaural simulation microphone. The audio signal includes multiple pure tones, standard speech materials, and multiple types of environmental noise. Different signal-to-noise ratios are set to form different test scenarios. During the test, the binaural audio signal is synchronized with the EEG device through an optically coupled hardware trigger. The EEG device records the brain's neural response to the binaural audio signal, while ensuring the time synchronization between the audio playback and the EEG recording. EEG signal preprocessing: multi-scale wavelet transform is performed on the EEG signal to decompose the signal into sub-signals of multiple scales; soft threshold denoising is performed on the sub-signal of each scale to remove high-frequency noise; adaptive filtering is performed on the denoised signal to remove the noise introduced by artifacts; the processed signal is reconstructed by inverse wavelet transform to obtain the preprocessed EEG signal. The specific technical solution is as follows: ; Where: represents the wavelet transform coefficients at scale a and displacement b; represents the original EEG signal; represents the mother wavelet function; a represents the scale factor; b represents the translation factor; Perform threshold processing on each scale coefficient after wavelet transformation: ; Where: is the coefficient after threshold; Indicates the threshold value; After threshold processing, the noise coefficient in the signal is weakened or removed, retaining the important EEG signal components; Adaptive filtering: Assume that the denoised input signal is , the noise signal is , the target signal is The output of the filter is : ; Where: Indicates input signal The feature vector of ; w represents the weight of the filter; represents mean square error; According to the LMS algorithm, the filter weights are updated at each time step as: ; Where: represents the learning step length; After adaptive filtering, the signal is obtained Signal-based Perform inverse wavelet transform reconstruction to obtain the final denoised and filtered EEG signal: The extracting of time domain features and time-frequency features comprises the following steps: Time domain feature extraction: The pre-processed EEG signal is subjected to time domain analysis to extract the average potential, amplitude and waveform complexity features to obtain the extracted time domain features; the average point position quantifies the global potential level of the EEG signal by calculating the mean of the signal within a given time window; the amplitude measures the strength of the EEG signal by calculating the difference between the maximum and minimum values of the signal; the waveform complexity measures the complexity of the signal through sample entropy, reflecting the nonlinearity and unpredictability of the EEG signal; Time-frequency feature extraction: Use wavelet transform to perform time-frequency analysis on the preprocessed EEG signal and extract the characteristics of each frequency band: ; Where: represents the EEG signal; represents the wavelet function; t and f represent time and frequency respectively; Wavelet transform provides instantaneous frequency information of the signal, calculates the energy density of each frequency band, and reflects the response of the user's brain at different frequencies.
[0025] Step 4: Processing binaural audio signals based on filters to obtain time domain envelopes of each frequency band, and then calculating the dynamic range compression ratio of the time domain envelopes of each frequency band; calculating the time difference and energy ratio of binaural signals based on the cross-correlation method, and using personalized acoustic transfer functions and standard acoustic transfer function libraries to perform similarity calculations, quantifying the user's spatial positioning ability, and obtaining spatial auditory features; constructing a unified multimodal auditory feature vector based on the time domain features, time-frequency features, and spatial auditory features; The step 4 comprises the following steps: Filter processing binaural audio signal: Use bandpass filter to decompose binaural audio signal into multiple frequency bands, each frequency band contains the energy distribution of audio signal in this frequency band: Assume that signal After passing through a bandpass filter, the filtered output signal of frequency band i is obtained , the transfer function of its bandpass filter is , then the mathematical expression of the filtering process is: ; Where: represents inverse Fourier transform; represents Fourier transform; Based on this, the signal is decomposed into multiple frequency bands through a bandpass filter, each of which contains the energy distribution of the audio signal in that frequency band; Time domain envelope extraction and dynamic range compression ratio: In each frequency band, the time domain envelope of the signal is extracted by the envelope detector. When extracting, the Hilbert transform is used to obtain the time domain envelope; the dynamic range compression ratio of the time domain included in each frequency band is calculated: The calculation formula of the time domain envelope is as follows: ;Where: represents the time domain envelope of the i-th frequency band; Indicates the absolute value of the signal; In this embodiment, in order to reduce the influence of low-frequency components, Hilport transform is used to obtain the envelope: ; Where: represents the Hilport transform; Then calculate the dynamic range compression ratio of the time domain envelope of each frequency band, where the calculation formula for dynamic range compression is as follows: ; Where: represents the maximum value of the envelope signal of frequency band i; represents the minimum value of the envelope signal of frequency band i; represents the mean value of the envelope signal of frequency band i; In this embodiment, a high value corresponding to the dynamic range compression ratio indicates that the signal changes more drastically in this frequency band, and a low value indicates that the signal changes more steadily; Cross-correlation calculates the time difference and energy ratio of binaural signals: the time difference is estimated by calculating the delay between binaural signals; the energy ratio is obtained by calculating the mean square value of binaural signals; Time difference calculation: Assume that the left ear signal is , the right ear signal is , then the cross-correlation function is defined as: ; Where: represents the time delay; the maximum cross-correlation position The corresponding time difference is: ;In this embodiment, the time difference reflects the relative time difference between binaural signals reaching the brain, affecting spatial positioning; Energy ratio calculation: The energy ratio is the energy ratio of the left ear signal and the right ear signal, which is obtained by calculating the mean square value of the two signals. Assume that the energy of the left ear signal and the right ear signal are and , then the energy ratio is: ; In this embodiment, the energy ratio reflects the difference in sound intensity between the left and right ears, affecting the accuracy of spatial positioning.
[0026] Similarity calculation: The similarity between the personalized acoustic transfer function and the transfer function in the standardized acoustic transfer function library is calculated; the similarity uses the mutual normalized correlation as the similarity measure; Based on the personalized acoustic transfer function obtained in the previous process Compute the cross-normalized correlation with the transfer function from the standardized acoustic transfer function library: ; Where: represents the standard transfer function; represents the normalized similarity; In this embodiment, the similarity calculation further refines the quantification of spatial positioning and provides a personalized spatial positioning capability evaluation for each user.
[0027] Construct a multimodal auditory feature vector: All features extracted from time domain features, time-frequency features, and spatial auditory features are concatenated to construct a unified multimodal auditory feature vector, where the spatial auditory features include dynamic range compression ratio, time difference, energy ratio, and acoustic transfer function similarity.
[0028] Step 5: Based on the multimodal auditory feature vector, a multi-head attention mechanism is used to weight the multimodal feature vector to obtain a weighted multimodal feature vector; When the multi-head attention mechanism weights the multimodal feature vector, feature importance weights are introduced to adjust the attention weights, including the following steps: Feature importance calculation: The feature importance is calculated based on the variance of each sub-feature in the input vector in the multi-head attention mechanism, the correlation of the threshold, and the robustness under different noise conditions: ; ; ; Where: , , Respectively represent the variance of time domain features, time-frequency features, and spatial auditory features; , , They represent the correlation between time domain features, time-frequency features, and spatial auditory features respectively; , , They represent the robustness of time domain features, time-frequency features, and spatial auditory features respectively; , , represents the hyperparameters used to adjust the contribution of variance, correlation, and robustness to the feature importance calculation; The correlation of time domain features, time-frequency features and spatial auditory features is obtained by calculating the Pearson correlation coefficient between each feature and the target, and the correlation of each modal feature is obtained; The robustness of time domain features, time-frequency features, and spatial auditory features is evaluated by simulating different noise environments, adding Gaussian noise, and then calculating the feature changes after the noise is added. The feature changes are evaluated by calculating the Euclidean distance between the original features and the features after adding Gaussian noise. If a feature with high stability shows small changes under noise, it has good robustness.
[0029] Based on the importance of the calculated time domain features, time-frequency features, and spatial auditory features, normalization is performed to obtain the feature importance weights of each modality; Combining feature importance with attention weight: When weighting the query and key of each modality, feature importance weight is introduced to obtain attention weight: ; Where: Represents the feature importance weight of each modality; represents the inner product of the query vector and the key vector; represents the scaling factor; express Function, the inner product result is used Function, ensuring that all attention weights are positive and sum to 1; Weight each feature based on the calculated attention weights; In this embodiment, the importance of features is dynamically evaluated by taking into account multiple factors in total, rather than relying on the traditional relevance-based attention mechanism, which can better adapt to hearing screening tasks for different users and in different environments.
[0030] Step 6: Based on the weighted multimodal feature vector and frequency / intensity parameters, the hearing threshold is predicted through a deep neural network model, the threshold deviation of each frequency point is output, and then the standard threshold is obtained by combining it with the traditional pure tone audiometry method; The step 6 comprises the following steps: Feature construction: concatenate the weighted multimodal feature vector and the frequency / intensity parameter to obtain a concatenated feature vector; Hearing threshold deviation prediction: using the concatenated feature vector as the input of a trained deep neural network model, and predicting the hearing threshold deviation through the deep neural network model; Exemplarily, the deep neural network model includes an input layer, multiple fully connected layers and an output layer, each of the fully connected layers uses a ReLU activation function to capture nonlinear relationships; the loss function uses a mean square error loss function; and during the training process, an Adam optimizer is used to minimize the loss function.
[0031] Pure tone audiometry threshold: The hearing threshold is obtained based on pure tone audiometry; it should be noted that pure tone audiometry is a common hearing test method that determines the minimum intensity of sound or the lowest audible tone that an individual can hear by measuring pure tone audio signals.
[0032] Threshold fusion: The hearing threshold and hearing threshold deviation obtained by pure tone audiometry are summed, and the summed value is weighted and fused with the hearing threshold obtained by pure tone audiometry to obtain the standard threshold. ; Where: represents the hearing threshold of the ith frequency point after fusion; and Represents the weight parameter to be optimized, and the sum is 1; represents the true threshold value measured by the traditional method; Represents the hearing threshold deviation predicted by the deep neural network model.
[0033] The weights in weighted fusion are obtained based on Bayesian optimization, and the objective function of Bayesian optimization is as follows: ; Where: represents the actual hearing threshold of the ith frequency point; Represents the hearing threshold of the ith frequency point after fusion.
[0034] The process of optimization based on the Bayesian approach is as follows: Initial sampling: randomly select some initial points Perform sampling and calculate the objective function value; Gaussian process modeling: Use the selected initial points to build a Gaussian process model as a proxy model for the objective function; Select the next sampling point: Find the next sampling point on the surrogate model and select it by maximizing the expected improvement function; Update model: calculate the objective function value at the new sampling point and update the Gaussian process model; Iteration: Repeat the steps of selecting sampling points and updating the model until convergence or the preset number of iterations is reached.
[0035] The optimal weight is obtained through Bayesian optimization, and then the weight is used for weighted fusion in practical applications to obtain the standard threshold of each frequency point; Step 7: Based on the multimodal feature vector, the actual age of the user and the standard threshold, a temporal Transformer encoder is used to predict the hearing age of the user, and a hearing decline assessment is performed on the user based on the predicted hearing age; In this embodiment, the input of the temporal Transformer encoder includes a multimodal auditory feature vector, the actual age of the user, and a standard threshold. We combine the input features to obtain an input feature vector of the temporal Transformer encoder. An exemplary input feature vector is as follows: ; in: represents the standard threshold; Indicates the user's age; represents a multimodal auditory feature vector; Input embedding: First, the input features Embed and get the representation of each time step , mapped through a fully connected layer: ; Positional encoding: The order information of the sequence is introduced through positional encoding, which is generated by sine and cosine functions: ; Where: t represents the time step; i represents the dimension index; d represents the dimension of the feature vector; self-attention mechanism: the self-attention mechanism is used to capture the dependencies between the time steps in the time series, calculate the weighted relationship between each time step and other time steps, and convert the position-encoded features into As input, the self-attention mechanism is calculated as follows: ; ; ; ; Where: , , Matrices representing queries, keys, and values, respectively; , , They represent weight matrices respectively; d represents feature dimension; Feedforward neural network: After self-attention weighting, a layer of feedforward neural network is used to add nonlinear transformation: ; Where: , , represents the weight matrix; and All represent bias terms; Represents the input feature representation after weighting by the attention mechanism; After the temporal Transformer encoder outputs, a representation containing historical event steps is obtained This representation combines the features of each time step and is mapped to the predicted hearing age through a fully connected layer. ; After the hearing prediction is completed, the decline assessment is to compare the predicted hearing age with the actual age to obtain the degree of hearing decline; ; Where: Indicates actual age; like Significantly higher than , it indicates that the user has a more serious hearing loss; if Lower than This indicates that the user's hearing function is good.
[0036] In step 5 of this embodiment, the features of different modalities are weighted and fused through the multi-head attention mechanism, focusing on the weighting at the feature level. In step 7, the temporal Transformer attention mechanism focuses on the time dimension, captures the dependencies between different time steps, models the time information in the time series data, and helps the model learn how to predict future states based on historical information. The two attention mechanisms complement each other and enhance the expressiveness of the model.
[0037] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. An adaptive hearing screening optimization method based on acoustic feature analysis, characterized in that: The following steps are involved: Step 1: Based on the constructed audio library, pure tones from 125 Hz to 8 kHz, standard speech materials, and various types of environmental noise are generated, and the noise and target signal are superimposed by setting different signal-to-noise ratios to form audio data with various signal-to-noise ratios; Step 2: Use a probe microphone to measure the acoustic transfer function of the user's ear canal, and generate a personalized acoustic transfer function based on the head-related transfer function database; Step 3: Play the calibrated audio through the binaural simulation microphones equipped on the artificial head, and record the binaural audio signals. At the same time, synchronize the audio playback and the EEG device through the optical coupling hardware trigger to obtain EEG data, and pre-process the EEG data to extract time domain features and time-frequency features based on the pre-processed EEG data; Step 4: Processing binaural audio signals based on filters to obtain time domain envelopes of each frequency band, and then calculating the dynamic range compression ratio of the time domain envelopes of each frequency band; calculating the time difference and energy ratio of binaural signals based on the cross-correlation method, and using personalized acoustic transfer functions and standard acoustic transfer function libraries to perform similarity calculations, quantifying the user's spatial positioning ability, and obtaining spatial auditory features; constructing a unified multimodal auditory feature vector based on the time domain features, time-frequency features, and spatial auditory features; Step 5: Based on the multimodal auditory feature vector, a multi-head attention mechanism is used to weight the multimodal feature vector to obtain a weighted multimodal feature vector; Step 6: Based on the weighted multimodal feature vector and frequency / intensity parameters, the hearing threshold is predicted through a deep neural network model, the threshold deviation of each frequency point is output, and then the standard threshold is obtained by combining it with the traditional pure tone audiometry method; Step 7: Based on the multimodal feature vector, the actual age of the user and the standard threshold, a temporal Transformer encoder is used to predict the hearing age of the user, and a hearing decline assessment is performed on the user based on the predicted hearing age.
2. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The step 2 comprises the following steps: Measuring pure tone signals: Playing multi-band pure tone stimulation through an external speaker, and inserting a probe microphone into the ear canal to record the sound wave response of the pure tone stimulation after passing through the ear canal to obtain a measurement signal; Calculate the acoustic transfer function of the ear canal: Based on the input signal and the measured output signal, the acoustic transfer function of the ear canal is calculated in the frequency domain: ; Where: is the frequency response of the ear canal, i.e., the acoustic transfer function; and Representation of the frequency domain of the input signal and the output signal respectively; User HRTF acquisition: inferring the applicable HRTF function from the standard HRTF database based on the user's individual physical signs; Synthesize personalized acoustic transfer function: The personalized acoustic transfer function of the user is obtained based on the product of the inferred HRTF function and the calculated acoustic transfer function of the ear canal.
3. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: After designing the personalized acoustic transfer function in step 2, the FIR filter is used to reversely compensate the headphone frequency response, including the following steps: Measure the frequency response of headphones: Play a standard pure tone signal through the headphones, use a probe microphone to record the signal output by the headphones, and calculate the frequency response of the headphones based on the signal output by the headphones: ; Where: represents the frequency domain representation of the headphone output signal; represents the frequency domain representation of the input signal; Indicates the frequency response of the headphones; FIR filter: The output of the headphones is compensated by the FIR filter: ; Where: represents the frequency response of the filter; Obtain FIR filter coefficients based on frequency response using discrete Fourier transform , the output signal is inversely compensated based on the filter coefficients: ; Where: Represents the convolution operation; represents the time domain representation of the input signal; represents the time domain representation of the output signal.
4. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The steps of obtaining and preprocessing EEG data in step 3 are as follows: Binaural audio signal playback: The calibrated audio signal is played through an artificial binaural simulation microphone. The audio signal includes multiple pure tones, standard speech materials, and multiple types of environmental noise. Different signal-to-noise ratios are set to form different test scenarios. During the test, the binaural audio signal is synchronized with the EEG device through an optically coupled hardware trigger. The EEG device records the brain's neural response to the binaural audio signal, while ensuring the time synchronization between the audio playback and the EEG recording. EEG signal preprocessing: Perform multi-scale wavelet transform on EEG signals to decompose the signals into sub-signals of multiple scales; Perform soft threshold denoising on the sub-signals of each scale to remove high-frequency noise; Adaptively filter the denoised signal to remove the noise introduced by the artifacts; The processed signal is reconstructed by inverse wavelet transform to obtain the preprocessed EEG signal.
5. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: Extracting time domain features and time-frequency features The following steps are involved: Time domain feature extraction: The pre-processed EEG signal is subjected to time domain analysis to extract the average potential, amplitude and waveform complexity features to obtain the extracted time domain features; Time-frequency feature extraction: Wavelet transform is used to perform time-frequency analysis on the preprocessed EEG signal to extract the features of each frequency band.
6. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The step 4 comprises the following steps: Filter processing of binaural audio signals: using a bandpass filter to decompose binaural audio signals into multiple frequency bands, each of which contains the energy distribution of the audio signal in that frequency band; Time domain envelope extraction and dynamic range compression ratio: In each frequency band, the time domain envelope of the signal is extracted by the envelope detector. When extracting, the Hilbert transform is used to obtain the time domain envelope; the dynamic range compression ratio of the time domain included in each frequency band is calculated; Cross-correlation calculates the time difference and energy ratio of binaural signals: the time difference is estimated by calculating the delay between binaural signals; the energy ratio is obtained by calculating the mean square value of binaural signals; Similarity calculation: The similarity between the personalized acoustic transfer function and the transfer function in the standardized acoustic transfer function library is calculated; the similarity uses the mutual normalized correlation as the similarity measure; Construct a multimodal auditory feature vector: Concatenate all features extracted from time domain features, time-frequency features, and spatial auditory features to construct a unified multimodal auditory feature vector.
7. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: When the multi-head attention mechanism weights the multimodal feature vector, the feature importance weight is introduced to adjust the attention weight. The following steps are involved: Feature importance calculation: The feature importance is calculated based on the variance of each sub-feature in the input vector in the multi-head attention mechanism, the correlation of the threshold, and the robustness under different noise conditions: ; ; ; Where: , , Respectively represent the variance of time domain features, time-frequency features, and spatial auditory features; , , They represent the correlation between time domain features, time-frequency features, and spatial auditory features respectively; , , They represent the robustness of time domain features, time-frequency features, and spatial auditory features respectively; , , represents the hyperparameters used to adjust the contribution of variance, correlation, and robustness to the feature importance calculation; Based on the importance of the calculated time domain features, time-frequency features, and spatial auditory features, normalization is performed to obtain the feature importance weights of each modality; Combining feature importance with attention weight: When weighting the query and key of each modality, feature importance weight is introduced to obtain attention weight: ; Where: Represents the feature importance weight of each modality; represents the inner product of the query vector and the key vector; represents the scaling factor; express Function, the inner product result is used Function that ensures that all attention weights are positive and sum to 1.
8. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The step 6 comprises the following steps: Feature construction: concatenate the weighted multimodal feature vector and the frequency / intensity parameter to obtain a concatenated feature vector; Hearing threshold deviation prediction: using the concatenated feature vector as the input of a trained deep neural network model, and predicting the hearing threshold deviation through the deep neural network model; Pure tone audiometry threshold: Hearing threshold is obtained based on pure tone audiometry; Threshold fusion: The hearing threshold and hearing threshold deviation obtained by pure tone audiometry are summed, and the summed value is weighted and fused with the hearing threshold obtained by pure tone audiometry to obtain the standard threshold. The weights in weighted fusion are obtained based on Bayesian optimization, and the objective function of Bayesian optimization is as follows: ; Where: represents the actual hearing threshold of the ith frequency point; Represents the hearing threshold of the ith frequency point after fusion.
9. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The temporal Transformer encoder includes an input embedding layer, a position encoding layer, a self-attention mechanism layer, a feedforward neural network layer, and an output layer; The features fused by the weighted multimodal feature vector and the frequency / intensity parameters are input to the input embedding layer, mapped through a fully connected layer to obtain the feature representation of each time step; and then the sequence information is introduced through the position encoding layer; The features processed by the position encoding layer are used to capture the dependencies between time steps in the time series through the self-attention mechanism layer, and the weighted relationship between each time step and other time steps is calculated to obtain the attention weight of each time step. Based on the attention weight of each time step, the weighted representation of each time step is obtained; the weighted representation of each time step is input into the feedforward neural network layer for nonlinear change, and the feature representation after nonlinear change enters the output layer and is mapped to the predicted hearing age through a fully connected layer.
Citation Information
Patent Citations
Hearing test method and hearing screening instrument for automatically correcting influence of environmental noise
CN109480859A
System and method for personalization of auditory stimulus
CN111586513A
Multifunctional hearing evaluation earphone and evaluation method thereof
CN112315462A
Methods, devices and system for a compensated hearing test
CN113164102A
Auto-calibration intelligent hearing screening method and device
CN114305403A
Cited By
Auditory sense detection method and system based on ASSR result feedback
CN120585319A
Spatial audio testing method
CN120602882A
A method of testing spatial audio
CN120602882B
Listening training system for English teaching
CN120954287A