Adaptive Hearing Screening Optimization Method Based on Acoustic Feature Analysis

By constructing multimodal auditory feature vectors and using deep neural network models for hearing threshold prediction, combined with the timing Transformer encoder to predict hearing age, the challenges of noise suppression, spatial auditory ability quantification and personalized hearing screening in the prior art are solved, and personalized listening evaluation with high accuracy is achieved.

CN119924826BActive Publication Date: 2025-06-10WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510431366.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-06-10
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The prior art has great challenges in noise suppression, spatial auditory capacity quantification and personalized hearing screening, especially when considering individual differences between different users, resulting in limited accuracy of evaluation results.

Method used

By constructing an audio library to generate audio data with different signal-to-noise ratios, combining the personalized acoustic transfer function measured by the probe microphone and the time domain and time frequency characteristics extracted by EEG data, multimodal auditory feature vectors are constructed, and hearing threshold prediction is used using multi-head attention mechanism and deep neural network model, and hearing age is predicted in combination with the timing Transformer encoder.

Benefits of technology

It achieves accurate personalized listening evaluation in a variable listening environment, overcomes the defects of ignoring individual differences and complex environmental noise in traditional methods, and improves the accuracy of hearing screening results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119924826B_ABST
    Figure CN119924826B_ABST
Patent Text Reader

Abstract

This application belongs to the technical field of hearing tests, and relates to an optimized method for adaptive hearing screening based on acoustic feature analysis. By constructing an audio library, audio data with different signal-to-noise ratios is generated. A probe microphone is used to measure the acoustic transfer function of the user's ear canal, and combined with the head-related transfer function to generate a personalized acoustic transfer function. Time domain and time-frequency features are extracted from EEG data, combined with the spatial localization features of binaural simulation microphones to quantify the user's auditory response ability and spatial auditory ability. Through the construction of multi-modal auditory feature vectors, the features are weighted by combining the multi-head attention mechanism, combined with the user's age and standard threshold, and the hearing age of the user is predicted by the temporal Transformer encoder, and the hearing decline is evaluated according to the prediction result. It enables accurate adaptation to the individual differences of each user, overcomes the defects of ignoring individual differences and complex environmental noise in traditional hearing screening methods, thereby improving the accuracy of screening results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of hearing tests. More specifically, it relates to an optimized method for adaptive hearing screening based on acoustic feature analysis. Background Art

[0002] Hearing impairment has become an increasingly serious public health problem globally, especially among the elderly population. The impact of hearing decline on the quality of life is extremely significant. According to the statistics of the World Health Organization, approximately 400 million people worldwide are troubled by hearing impairments of varying degrees. With the intensification of population aging, the need for hearing impairment screening and early intervention is increasing day by day. Accurate hearing screening not only helps to detect hearing problems early but also provides data support for relevant intervention measures, thereby effectively slowing down or delaying the process of hearing decline.

[0003] In traditional hearing screening methods, the most common one is pure-tone audiometry, which evaluates the hearing thresholds of individuals at different frequencies by playing pure tones of standard frequencies. Although this method is simple and widely used, it relies on the subjective judgment of hearing experts and is easily affected by factors such as environmental noise and equipment accuracy, resulting in limited accuracy and reliability of test results. In addition, pure-tone audiometry usually ignores the individual's spatial hearing ability, the influence of environmental noise, and the complex interactions between different frequencies. Therefore, in the face of noisy environments and complex hearing impairments, the accuracy is greatly reduced.

[0004] In recent years, with the progress of computer technology and artificial intelligence, hearing screening methods based on acoustic feature analysis have gradually received attention. These methods analyze the user's auditory response by using different acoustic signals (such as pure tones, speech, and environmental noise) and combining advanced signal processing techniques.

[0005] Although modern technologies have made certain progress in time-domain feature extraction, time-frequency feature analysis, and multi-modal signal fusion, there are still significant challenges in noise suppression, quantification of spatial hearing ability, and personalized screening in the existing technologies. Especially in practical applications, the impact of environmental noise on hearing test results cannot be ignored. When the existing technologies effectively distinguish between noise and signals and fuse personalized auditory characteristics, the accuracy is often not ideal. In addition, most existing technologies are also difficult to achieve real-time, dynamic, and personalized hearing assessments. Especially when there are significant differences in the hearing conditions and auditory abilities of different users, individual differences cannot be fully considered, resulting in limited accuracy of the assessment results. Summary of the Invention

[0006] The present invention provides an optimized method for adaptive hearing screening based on acoustic feature analysis, aiming to solve the technical problem that when there are significant differences in the hearing conditions and auditory abilities of different users, individual differences cannot be fully considered, resulting in limitations in the accuracy of evaluation results.

[0007] The optimized method for adaptive hearing screening based on acoustic feature analysis includes the following steps:

[0008] Step 1: Based on the constructed audio library, generate pure tones ranging from 125 Hz to 8 kHz, standard speech materials, and various types of environmental noises, and superimpose the noises and target signals by setting different signal-to-noise ratios to form audio data with multiple signal-to-noise ratios.

[0009] Step 2: Use a probe microphone to measure the acoustic transfer function of the user's ear canal, and combine it with the head-related transfer function database to generate a personalized acoustic transfer function.

[0010] Step 3: Play the calibrated audio through an artificial head equipped with binaural simulation microphones, record the binaural audio signals, and simultaneously synchronize the audio playback and EEG device through an optical coupling hardware trigger to obtain EEG data, and preprocess the EEG data. Extract time-domain features and time-frequency features based on the preprocessed EEG data.

[0011] Step 4: Based on filter processing of the binaural audio signals, obtain the time-domain envelope of each frequency band, and then calculate the dynamic range compression ratio of the time-domain envelope of each frequency band; calculate the time difference and energy ratio of the binaural signals based on the cross-correlation method, and use the personalized acoustic transfer function and the standard acoustic transfer function library to calculate the similarity, quantify the user's spatial localization ability, and obtain the spatial auditory features; construct a unified multi-modal auditory feature vector based on the time-domain features, time-frequency features, and spatial auditory features.

[0012] Step 5: Based on the multi-modal auditory feature vector, use the multi-head attention mechanism to weight the multi-modal feature vector to obtain the weighted multi-modal feature vector.

[0013] Step 6: Based on the weighted multi-modal feature vector and frequency / intensity parameters, predict the hearing threshold through a deep neural network model, output the threshold deviation of each frequency point, and then combine it with the traditional pure tone audiometry method to obtain the standard threshold.

[0014] Step 7: Based on the multi-modal feature vector, the user's actual age, and the standard threshold, use a temporal Transformer encoder to predict the user's hearing age, and evaluate the hearing decline of the user based on the predicted hearing age.

[0015] The present invention generates audio data with different signal-to-noise ratios based on a constructed audio library, combines environmental noise and target signals to ensure that the hearing test requirements under various noise conditions can be covered; secondly, uses a probe microphone to measure the acoustic transfer function of the user's ear canal, and combines the head-related transfer function to generate a personalized acoustic transfer function, so as to accurately reflect the individual ear canal acoustic characteristics and avoid the one-size-fits-all assumption in traditional methods; then, extracts time-domain and time-frequency features from EEG data, combines the spatial localization features of binaural simulation microphones, and further quantifies the user's auditory response ability and spatial auditory ability; through the construction of multi-modal auditory feature vectors, combines the multi-head attention mechanism to weight the features, so as to more accurately predict the hearing threshold; finally, combines the user's age and standard threshold, predicts the user's hearing age through a temporal Transformer encoder, and evaluates hearing decline according to the prediction result; enabling it to accurately adapt to the individual differences of each user, overcome the defects of ignoring individual differences and complex environmental noise in traditional hearing screening methods, thereby improving the accuracy of screening results and ensuring that effective hearing assessments can be provided in a changing hearing environment.

[0016] Preferably, step 2 includes the following steps:

[0017] Measure pure tone signal: Play multi-band pure tone stimuli through an external speaker, and at the same time insert a probe microphone into the ear canal to record the acoustic wave response after the pure tone stimulus passes through the ear canal, obtaining a measurement signal;

[0018] Calculate the acoustic transfer function of the ear canal: Based on the input signal and the measured output signal, calculate the acoustic transfer function of the ear canal in the frequency domain: ; where: is the frequency response of the ear canal, that is, the acoustic transfer function; and respectively represent the frequency-domain representations of the input signal and the output signal;

[0019] User HRTF acquisition: Infer the applicable HRTF function from the standard HRTF database based on the user's individual physical signs;

[0020] Synthesize personalized acoustic transfer function: Obtain the user's personalized acoustic transfer function based on the product of the inferred HRTF function and the calculated acoustic transfer function of the ear canal.

[0021] Preferably, after designing a personalized ear canal transfer function in step 2, use a FIR filter to reverse-compensate the headphone frequency response, including the following steps:

[0022] Measure headphone frequency response: Play a standard pure tone signal through the headphone, and use a probe microphone to record the signal output by the headphone, and calculate the frequency response of the headphone based on the signal output by the headphone: Wherein: represents the frequency-domain representation of the headphone output signal; represents the frequency-domain representation of the input signal; represents the frequency response of the headphone;

[0023] FIR filter: The output of the headphone is compensated by the FIR filter: ; wherein: represents the frequency response of the filter;

[0024] The FIR filter coefficients are obtained by discrete Fourier transform based on the frequency response , and the output signal is inversely compensated based on the coefficients of the filter: ; wherein: represents the convolution operation; represents the time-domain representation of the input signal; represents the time-domain representation of the output signal.

[0025] Preferably, the steps of obtaining and preprocessing the EEG data in step 3 are as follows:

[0026] Binaural audio signal playback: Calibrated audio signals are played through artificially mounted binaural simulation microphones, where the audio signals include multiple segments of pure tones, standard speech materials, and multiple types of environmental noises, and different signal-to-noise ratios are set to form different test scenarios; during the test, the binaurally played audio signals are synchronized with the EEG device through an optical coupling hardware trigger, and the EEG device records the neural responses of the brain to the binaural audio signals, while ensuring the time synchronization between audio playback and EEG recording;

[0027] EEG signal preprocessing: The EEG signal is subjected to multi-scale wavelet transform to decompose the signal into sub-signals of multiple scales; soft threshold denoising is performed on the sub-signals of each scale to remove high-frequency noise; adaptive filtering is performed on the denoised signal to remove the noise introduced by artifacts; the processed signal is reconstructed through inverse wavelet transform to obtain the preprocessed EEG signal.

[0028] Preferably, the extraction of time-domain features and time-frequency features includes the following steps:

[0029] Time-domain feature extraction: The preprocessed EEG signal is subjected to time-domain analysis to extract features such as average potential, amplitude, and waveform complexity, to obtain the extracted time-domain features;

[0030] Time-frequency feature extraction: Wavelet transform is used to perform time-frequency analysis on the preprocessed EEG signal to extract the features of each frequency band.

[0031] Preferably, step 4 includes the following steps:

[0032] The filter processes the binaural audio signal: A band-pass filter is used to decompose the binaural audio signal into multiple frequency bands, and each frequency band contains the energy distribution of the audio signal in that frequency band;

[0033] Time-domain envelope extraction and dynamic range compression ratio: On each frequency band, the time-domain envelope of the signal is extracted through an envelope detector. When extracting, the Hilbert transform is used to obtain the time-domain envelope; the dynamic range compression ratio of the time-domain envelope of each frequency band is calculated;

[0034] Cross-correlation calculates the time difference and energy ratio of the binaural signals: The time difference is estimated by calculating the delay between the binaural signals; the energy ratio is obtained by calculating the mean square value of the binaural signals;

[0035] Similarity calculation: The similarity between two functions is calculated using a personalized acoustic transfer function and the transfer functions in a standardized acoustic transfer function library; among them, the cross-normalized correlation is used as the similarity metric;

[0036] Constructing a multi-modal auditory feature vector: All the features extracted from the time-domain features, time-frequency features, and spatial auditory features are concatenated to construct a unified multi-modal auditory feature vector.

[0037] Preferably, when the multi-head attention mechanism weights the multi-modal feature vector, a feature importance weight is introduced to adjust the attention weight, including the following steps:

[0038] Feature importance calculation: Based on the variance, correlation with the threshold, and robustness under different noise conditions of each sub-feature in the input vector in the multi-head attention mechanism, the feature importance is calculated: ; ; ; where: 、 、 represent the variances of the time-domain feature, time-frequency feature, and spatial auditory feature respectively; 、 、 represent the correlations of the time-domain feature, time-frequency feature, and spatial auditory feature respectively; 、 、 represent the robustness of the time-domain feature, time-frequency feature, and spatial auditory feature respectively; 、 、 represent hyperparameters used to adjust the contributions of variance, correlation, and robustness to the feature importance calculation;

[0039] Based on the calculated importance of the time-domain feature, time-frequency feature, and spatial auditory feature, normalization is performed to obtain the feature importance weights of each modality;

[0040] Combination of feature importance and attention weights: When weighting the queries and keys of each modality, introduce feature importance weights to obtain attention weights: where: represents the feature importance weight of each modality; represents the inner product of the query vector and the key vector; represents the scaling factor; represents function, applying the function to the inner product result to ensure that all attention weights are positive and their sum is 1.

[0041] Preferably, step 6 includes the following steps:

[0042] Feature construction: Concatenate the weighted multi-modal feature vectors and the frequency / intensity parameters to obtain the concatenated feature vectors;

[0043] Hearing threshold deviation prediction: Use the concatenated feature vectors as the input to a trained deep neural network model, and predict the hearing threshold deviation through the deep neural network model;

[0044] Pure tone audiometry threshold: Obtain the hearing threshold based on pure tone audiometry;

[0045] Threshold fusion: Sum the hearing threshold obtained by pure tone audiometry and the hearing threshold deviation, and perform weighted fusion based on the sum value and the hearing threshold obtained by pure tone audiometry to obtain the standard threshold; where the weights during weighted fusion are obtained based on Bayesian optimization, and the objective function of Bayesian optimization is as follows: ;

[0046] where: represents the true hearing threshold at the i-th frequency point; represents the fused hearing threshold at the i-th frequency point.

[0047] Preferably, the temporal Transformer encoder includes an input embedding layer, a positional encoding layer, a self-attention mechanism layer, a feed-forward neural network layer, and an output layer;

[0048] The features fused from the weighted multi-modal feature vectors and the frequency / intensity parameters are input into the input embedding layer, mapped through a fully connected layer to obtain the feature representation at each time step; then introduce the sequential information of the sequence through the positional encoding layer;

[0049] The features processed by the position encoding layer capture the dependencies between time steps in the time series through the self-attention mechanism layer, calculate the weighted relationships between each time step and other time steps, obtain the attention weights for each time step, and get the weighted representation for each time step based on the attention weights of each time step; the weighted representation of each time step is input into the feed-forward neural network layer for non-linear transformation, and the feature representation after non-linear transformation enters the output layer and is mapped to the predicted hearing age through a fully connected layer.

[0050] The beneficial effects of the present invention include:

[0051] Based on the constructed audio library, the present invention generates audio data with different signal-to-noise ratios, combines environmental noise and target signals to ensure that the hearing test requirements under various noise conditions can be covered; secondly, a probe microphone is used to measure the acoustic transfer function of the user's ear canal, and a personalized acoustic transfer function is generated in combination with the head-related transfer function, so as to accurately reflect the individual ear canal acoustic characteristics and avoid the one-size-fits-all assumption in traditional methods; then, time-domain and time-frequency features are extracted from EEG data, combined with the spatial localization features of binaural simulation microphones, to further quantify the user's auditory response ability and spatial auditory ability; through the construction of multi-modal auditory feature vectors and the weighting of features by the multi-head attention mechanism, the hearing threshold can be predicted more accurately; finally, in combination with the user's age and standard threshold, the hearing age of the user is predicted through the temporal Transformer encoder, and the hearing decline is evaluated according to the prediction result; enabling accurate adaptation to the individual differences of each user, overcoming the defects of ignoring individual differences and complex environmental noise in traditional hearing screening methods, thereby improving the accuracy of screening results and ensuring effective hearing assessment in a changing hearing environment. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is the overall step block diagram provided by the embodiment of the present invention.

[0054] Figure 2 It is the flow schematic block diagram of the overall data processing logic provided by the embodiment of the present invention. Detailed Embodiments

[0055] To make the technical problems, technical solutions, and beneficial effects to be solved by this application more clearly understood, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not used to limit this application.

[0056] See Figure 1 and Figure 2 As shown, the adaptive hearing screening optimization method based on acoustic feature analysis includes the following steps:

[0057] Step 1: Based on the constructed audio library, generate pure tones ranging from 125 Hz to 8 kHz, standard speech materials, and various types of environmental noises, and superimpose the noise and the target signal by setting different signal-to-noise ratios to form audio data with various signal-to-noise ratios.

[0058] Exemplarily, the audio library contains three categories of audio content:

[0059] Pure tones: Select pure tones with a frequency range from 125 Hz to 8 kHz (including low frequency, medium frequency, and high frequency). The duration of each pure tone signal is selected as 500 ms or greater than 500 ms to ensure accurate frequency response.

[0060] Standard speech materials: Select standard language speech materials, such as letters, numbers, common words, or short sentences.

[0061] Environmental noises: Such as city noise, home background noise, office noise, etc.

[0062] Setting of the signal-to-noise ratio (SNR): By setting different signal-to-noise ratios, simulate the hearing challenges in different environments. The setting of SNR covers from low signal-to-noise ratio to high signal-to-noise ratio. The calculation formula of the signal-to-noise ratio is as follows: ; where: represents the power of the signal; represents the power of the noise; represents the signal-to-noise ratio, expressed in decibels;

[0063] In this embodiment, in order to cover a variety of hearing scenarios, we can select multiple SNR values, such as: -10 dB, 0 dB, 10 dB, 20 dB, etc., to form audio samples with different noise interference intensities.

[0064] Audio signal synthesis: Superimpose different types of environmental noises with pure tones and standard speech materials to form the final audio samples. The specific steps are as follows:

[0065] Pure tone and noise synthesis: Select a pure tone signal of a certain frequency (e.g., 1000 Hz), and iterate it with environmental noise according to the selected SNR value; Exemplarily, at an SNR of -dB, the power of the noise will be much higher than that of the pure tone signal, while at an SNR of 30 dB, the pure tone signal is much stronger than the noise;

[0066] Standard speech and noise synthesis: Similar to the synthesis of pure tone and noise, superimpose standard speech materials with different noise signals to generate speech signals with different noise interferences;

[0067] The above pure tone and noise synthesis, standard speech and noise synthesis can be mixed according to the following formula: ; In the formula: represents the mixed audio signal; represents the original target signal (pure tone or speech signal); represents the noise signal; represents the scaling factor of the noise, which is used to control the ratio of the noise intensity to the signal intensity and is calculated according to the selected SNR: ;

[0068] Generate an audio dataset with multiple SNRs and multiple types based on the above process; The audio dataset includes different noise types, pure tone frequency ranges, speech contents, and different SNR settings.

[0069] Step 2: Measure the acoustic transfer function of the user's ear canal using a probe microphone, and generate a personalized acoustic transfer function in combination with the head-related transfer function database;

[0070] The said Step 2 includes the following steps:

[0071] Measure the pure tone signal: Play multi-band pure tone stimuli (e.g., 125 Hz, 250 Hz, etc.) through an external speaker, and at the same time insert a probe microphone into the ear canal to record the acoustic wave response of the pure tone stimulus after passing through the ear canal to obtain the measurement signal ;

[0072] Calculate the acoustic transfer function of the ear canal: Let and be the representations of the input signal and the output signal in the frequency domain. Based on the input signal and the measured output signal, calculate the acoustic transfer function of the ear canal in the frequency domain: ;

[0073] In the formula: is the frequency response of the ear canal, i.e., the acoustic transfer function; and respectively represent the representations of the input signal and the output signal in the frequency domain;

[0074] The corresponding time-domain signal is calculated as follows: ;

[0075] where: represents the inverse Fourier transform;

[0076] User HRTF acquisition: Infer the applicable HRTF function from the standard HRTF database based on the user's individual physical signs;

[0077] Exemplarily, the user's individual characteristics include the length and width of the head, as well as the length of the ear canal and the height of the auricle;

[0078] Calculate the Euclidean distance between the feature vector formed by the user's individual characteristics and the feature vectors of each user in the database to obtain the similarity, and select the N individual data with the highest similarity;

[0079] Perform weighted averaging on the N individual data to obtain the inferred HRTF function ;

[0080] Synthesize the personalized acoustic transfer function: Perform a product operation based on the inferred HRTF function and the calculated acoustic transfer function of the ear canal to obtain the user's personalized acoustic transfer function: ;

[0081] In the formula: represents the inferred HRTF function;

[0082] As a further implementation manner of this embodiment, after designing the personalized ear canal transfer function in step 2, use an FIR filter to reverse-compensate the headphone frequency response, including the following steps:

[0083] Measure the headphone frequency response: Play a standard pure-tone signal through the headphone, record the signal output by the headphone using a probe microphone, and calculate the frequency response of the headphone based on the signal output by the headphone: ; In the formula: represents the frequency-domain representation of the headphone output signal; represents the frequency-domain representation of the input signal; represents the frequency response of the headphone;

[0084] FIR filter: Compensate the output of the headphone through the FIR filter: ; In the formula: represents the frequency response of the filter;

[0085] Obtain the FIR filter coefficients based on the frequency response using the discrete Fourier transform , and perform reverse compensation on the output signal based on the coefficients of the filter: ;

[0086] In the formula: represents a convolution operation; represents the time-domain representation of the input signal; represents the time-domain representation of the output signal.

[0087] Step 3: Play the calibrated audio through the artificial head equipped with binaural simulation microphones, and record the binaural audio signals. At the same time, synchronize the audio playback and the EEG device through an optical coupling hardware trigger to obtain electroencephalogram (EEG) data, and preprocess the EEG data. Extract time-domain features and time-frequency features based on the preprocessed EEG data;

[0088] The steps for obtaining and preprocessing the EEG data in Step 3 are as follows:

[0089] Binaural audio signal playback: Use the calibrated audio signal and play it through the artificial head equipped with binaural simulation microphones. The audio signal includes multiple segments of pure tones, standard speech materials, and multiple types of environmental noises, and different signal-to-noise ratios are set to form different test scenarios; during the test, the binaural-played audio signal synchronizes the EEG device through an optical coupling hardware trigger, and the EEG device records the neural response of the brain to the binaural audio signal, while ensuring the time synchronization between the audio playback and the EEG recording;

[0090] EEG signal preprocessing: Perform multi-scale wavelet transform on the EEG signal to decompose the signal into sub-signals of multiple scales; perform soft threshold denoising on the sub-signals of each scale to remove high-frequency noise; perform adaptive filtering on the denoised signal to remove the noise introduced by artifacts; the processed signal is reconstructed through inverse wavelet transform to obtain the preprocessed EEG signal. The specific technical solution is as follows: ;

[0091] In the formula: represents the wavelet transform coefficient at scale a and displacement b; represents the original EEG signal; represents the mother wavelet function; a represents the scale factor; b represents the translation factor;

[0092] Perform threshold processing on the coefficients of each scale after wavelet transform: ;

[0093] In the formula: is the coefficient after thresholding; represents the threshold;

[0094] After threshold processing, weaken or remove the noise coefficients in the signal and retain the important electroencephalogram signal components;

[0095] Adaptive filtering: Let the input signal after denoising be , and the noise signal be , the target signal is , and the output of the filter is : ;

[0096] In the formula: represents the input signal 's eigenvector; w represents the weight of the filter; represents the mean square error;

[0097] According to the LMS algorithm, the filter weight is updated at each time step as: ; In the formula: represents the learning step size;

[0098] After adaptive filtering processing, the signal Based on the signal perform inverse wavelet transform reconstruction to obtain the final denoised and filtered EEG signal:

[0099] The extraction of time-domain features and time-frequency features includes the following steps:

[0100] Time-domain feature extraction: The preprocessed EEG signal is subjected to time-domain analysis to extract the average potential, amplitude, and waveform complexity features, obtaining the extracted time-domain features; among them, the average potential quantifies the global potential level of the EEG signal by calculating the mean value of the signal within a given time window; the amplitude measures the intensity of the EEG signal by calculating the difference between the maximum and minimum values of the signal; the waveform complexity measures the complexity of the signal through sample entropy, reflecting the nonlinearity and unpredictability of the EEG signal;

[0101] Time-frequency feature extraction: Use wavelet transform to perform time-frequency analysis on the preprocessed EEG signal to extract the features of each frequency band: ;

[0102] In the formula: represents the EEG signal; represents the wavelet function; t and f represent time and frequency respectively;

[0103] The wavelet transform provides the instantaneous frequency information of the signal, calculates the energy density of each frequency band, and reflects the response of the user's brain at different frequencies.

[0104] Step 4: Process the binaural audio signal based on the filter to obtain the time-domain envelope of each frequency band, and then calculate the dynamic range compression ratio of the time-domain envelope of each frequency band; calculate the time difference and energy ratio of the binaural signals based on the cross-correlation method, and use the personalized acoustic transfer function and the standard acoustic transfer function library to perform similarity calculation to quantify the user's spatial localization ability, obtaining the spatial auditory features; construct a unified multi-modal auditory feature vector based on the time-domain features, time-frequency features, and spatial auditory features;

[0105] Step 4 includes the following steps:

[0106] Filter processing of the binaural audio signal: The binaural audio signal is decomposed into multiple frequency bands by a band-pass filter, and each frequency band contains the energy distribution of the audio signal in that frequency band. Assume the signal After passing through a band-pass filter, the filtered output signal of frequency band i is obtained , and the transfer function of the band-pass filter is , then the mathematical expression of the filtering process is: ;

[0107] In the formula: represents the inverse Fourier transform; represents the Fourier transform;

[0108] Based on this, the signal is decomposed into multiple frequency bands by a band-pass filter, and each frequency band contains the energy distribution of the audio signal in that frequency band;

[0109] Time-domain envelope extraction and dynamic range compression ratio: On each frequency band, the time-domain envelope of the signal is extracted through an envelope detector. When extracting, the Hilbert transform is used to obtain the time-domain envelope; when calculating the dynamic range compression ratio of the time-domain envelope of each frequency band:

[0110] The calculation formula of the time-domain envelope is as follows: ; In the formula: represents the time-domain envelope of the i-th frequency band; represents the absolute value of the signal;

[0111] In this embodiment, in order to reduce the influence of low-frequency components, the Hilbert transform is used to obtain the envelope: ;

[0112] In the formula: represents the Hilbert transform;

[0113] Then calculate the dynamic range compression ratio of the time-domain envelope of each frequency band. The calculation formula of the dynamic range compression is as follows: ;

[0114] In the formula: represents the maximum value of the envelope signal of frequency band i; represents the minimum value of the envelope signal of frequency band i; represents the mean value of the envelope signal of frequency band i;

[0115] In this embodiment, a high value of the dynamic range compression ratio corresponds to a relatively drastic change in the signal in that frequency band, and a low value indicates a relatively stable signal change;

[0116] Calculate the time difference and energy ratio of binaural signals through cross-correlation: estimate the time difference by calculating the delay between binaural signals; then obtain the energy ratio by calculating the mean square values of binaural signals;

[0117] Calculation of time difference: Assume the left-ear signal is , and the right-ear signal is , then the cross-correlation function is defined as: ;

[0118] In the formula: represents the time delay; the time difference corresponding to the maximum cross-correlation position is: ; In this embodiment, the time difference reflects the relative time difference of binaural signals reaching the brain and affects spatial localization;

[0119] Calculation of energy ratio: The energy ratio is the energy ratio of the left-ear and right-ear signals, obtained by calculating the mean square values of the two signals. Let the energies of the left-ear signal and the right-ear signal be and respectively, then the energy ratio is: ;

[0120] In this embodiment, the energy ratio reflects the sound intensity difference between the left and right ears and affects the accuracy of spatial localization.

[0121] Calculation of similarity: Use the personalized acoustic transfer function and the transfer functions in the standardized acoustic transfer function library to calculate the similarity between the two functions; among them, the mutual normalized correlation is used as the similarity metric;

[0122] Based on the personalized acoustic transfer function obtained in the previous process and the transfer functions in the standardized acoustic transfer function library to calculate the mutual normalized correlation: ;

[0123] In the formula: represents the standard transfer function; represents the normalized similarity;

[0124] In this embodiment, the similarity calculation further refines the quantization of spatial localization and provides an individualized spatial localization ability assessment for each user.

[0125] Construct a multi-modal auditory feature vector: Concatenate all the features extracted from time-domain features, time-frequency features, and spatial auditory features to construct a unified multi-modal auditory feature vector, where the spatial auditory features include the dynamic range compression ratio, time difference, energy ratio, and acoustic transfer function similarity.

[0126] Step 5: Use the multi-head attention mechanism to weight the multi-modal feature vectors based on the multi-modal auditory feature vectors to obtain the weighted multi-modal feature vectors;

[0127] When the multi-head attention mechanism weights the multi-modal feature vectors, introduce feature importance weights to adjust the attention weights, including the following steps:

[0128] Feature importance calculation: Calculate the feature importance based on the variance, correlation with the threshold, and robustness under different noise conditions of each sub-feature in the input vector of the multi-head attention mechanism: ; ; ;

[0129] In the formula: , , respectively represent the variances of time-domain features, time-frequency features, and spatial auditory features; , , respectively represent the correlations of time-domain features, time-frequency features, and spatial auditory features; , , respectively represent the robustness of time-domain features, time-frequency features, and spatial auditory features; , , represent hyperparameters used to adjust the contributions of variance, correlation, and robustness to the calculation of feature importance;

[0130] Among them, the correlations of time-domain features, time-frequency features, and spatial auditory features are obtained by calculating the Pearson correlation coefficient of each feature with the target to obtain the correlation of each modal feature;

[0131] The robustness of time-domain features, time-frequency features, and spatial auditory features is evaluated by simulating different noise environments, adding Gaussian noise, and then calculating the feature change after adding the noise to evaluate the stability. Among them, the feature change is evaluated by calculating the Euclidean distance between the original feature and the feature after adding Gaussian noise; if the feature with high stability shows less change under noise, its robustness is better.

[0132] Normalize the calculated importance of time-domain features, time-frequency features, and spatial auditory features to obtain the feature importance weights of each modality;

[0133] Combination of feature importance and attention weights: Introduce feature importance weights when weighting the query and key of each modality to obtain the attention weights: ;

[0134] In the formula: Represents the feature importance weights for each modality; Represents the inner product of the query vector and the key vector; Represents the scaling factor; Represents Function, which applies Function to ensure that all attention weights are positive and sum to 1;

[0135] Weight each feature based on the calculated attention weights;

[0136] In this embodiment, the importance of features is dynamically evaluated by considering multiple factors in total, rather than relying on traditional correlation-based attention mechanisms, which can better adapt to the hearing screening tasks of different users and different environments.

[0137] Step 6: Based on the weighted multi-modal feature vectors and frequency / intensity parameters, perform hearing threshold prediction through a deep neural network model, output the threshold deviation for each frequency point, and then combine it with the traditional pure tone audiometry method to obtain the standard threshold;

[0138] The said step 6 includes the following steps:

[0139] Feature construction: Concatenate the weighted multi-modal feature vectors and frequency / intensity parameters to obtain the concatenated feature vectors;

[0140] Hearing threshold deviation prediction: Use the concatenated feature vectors as the input of the trained deep neural network model, and predict the hearing threshold deviation through the deep neural network model;

[0141] Exemplarily, the deep neural network model includes an input layer, multiple fully connected layers, and an output layer. Each fully connected layer uses the ReLU activation function to capture non-linear relationships; the loss function uses the mean square error loss function; during the training process, the Adam optimizer is used to optimize to minimize the loss function.

[0142] Pure tone audiometry threshold: Obtain the hearing threshold based on pure tone audiometry; it should be noted that pure tone audiometry is a currently common hearing test method, which determines the lowest intensity or lowest audible tone of the sound that an individual can hear by measuring pure tone audio signals.

[0143] Threshold fusion: Sum the hearing threshold obtained by pure tone audiometry and the hearing threshold deviation, and perform weighted fusion based on the sum value and the hearing threshold obtained by pure tone audiometry to obtain the standard threshold; ;

[0144] In the formula: Represents the hearing threshold of the i-th frequency point after fusion; And Represents the weight parameter to be optimized, and the sum is 1; Represents the true threshold measured by the traditional method; Represents the deviation of the hearing threshold predicted by the deep neural network model.

[0145] Among them, the weights during weighted fusion are obtained based on Bayesian optimization, and the objective function of Bayesian optimization is as follows: ;

[0146] In the formula: Represents the hearing threshold at the i-th true frequency point; Represents the hearing threshold at the i-th frequency point after fusion.

[0147] The process of optimization based on the said Bayesian is as follows:

[0148] Initial sampling: Randomly select some initial points Perform sampling and calculate the objective function value;

[0149] Gaussian process modeling: Use the selected initial points to construct a Gaussian process model as a surrogate model for the objective function;

[0150] Select the next sampling point: Find the next sampling point on the surrogate model by maximizing the expected improvement function;

[0151] Update the model: Calculate the objective function value at the new sampling point and update the Gaussian process model;

[0152] Iteration: Repeat the steps of selecting sampling points and updating the model until convergence or reaching the preset number of iterations.

[0153] Obtain the optimal weights through Bayesian optimization, and then use the weights for weighted fusion in actual applications to obtain the standard threshold for each frequency point;

[0154] Step 7: Based on the said multi-modal feature vector, the actual age of the user, and the standard threshold, use the temporal Transformer encoder to predict the hearing age of the user, and evaluate the hearing decline of the user based on the predicted hearing age;

[0155] In this embodiment, the input of the temporal Transformer encoder includes the multi-modal auditory feature vector, the actual age of the user, and the standard threshold. We combine the input features to obtain the input feature vector for the input temporal Transformer encoder. An exemplary input feature vector is as follows: ;

[0156] Among them: Represents the standard threshold; Represents the user's age; Represents the multi-modal auditory feature vector;

[0157] Input embedding: First, embed the input features to obtain the representation at each time step , and map it through a fully connected layer: ;

[0158] Position encoding: Introduce the sequential information of the sequence through position encoding, which is generated by sine and cosine functions: ;

[0159] In the formula: t represents the time step; i represents the dimension index; d represents the dimension of the feature vector; Self-attention mechanism: Capture the dependencies between time steps in the time series through the self-attention mechanism, calculate the weighted relationship between each time step and other time steps, and use the features after position encoding as the input. The calculation process of the self-attention mechanism is as follows: ; ; ; ;

[0160] In the formula: , , represent the matrices of query, key, and value respectively; , , represent the weight matrices respectively; d represents the feature dimension;

[0161] Feed-forward neural network: After self-attention weighting, pass through a feed-forward neural network to increase the non-linear transformation: ;

[0162] In the formula: , represent the weight matrices; and both represent the bias terms; represents the input feature representation weighted by the attention mechanism;

[0163] After the output of the temporal Transformer encoder, obtain a representation containing historical event steps This representation combines the features of each time step, and maps this representation to the predicted hearing age through a fully connected layer ;

[0164] After the hearing prediction is completed, the decline assessment is to compare the predicted hearing age with the actual age to obtain the degree of hearing decline; ;

[0165] Wherein: represents the actual age;

[0166] If is significantly higher than , it indicates that the user has a relatively serious hearing decline; if is lower than it indicates that the user has good hearing function.

[0167] In step 5 of this embodiment, the multi-head attention mechanism is used to weight and fuse the features of different modalities, focusing on the weighting at the feature level. In step 7, the temporal Transformer attention mechanism focuses on the time dimension, captures the dependencies between different time steps, models the time information in the time series data, and helps the model learn how to predict the future state based on historical information. The two attention mechanisms complement each other and enhance the expression ability of the model.

[0168] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An adaptive hearing screening optimization method based on acoustic feature analysis, characterized in that: The following steps are involved: Step 1: Based on the constructed audio library, pure tones from 125 Hz to 8 kHz, standard speech materials, and various types of environmental noise are generated, and the noise and target signal are superimposed by setting different signal-to-noise ratios to form audio data with various signal-to-noise ratios; Step 2: Use a probe microphone to measure the acoustic transfer function of the user's ear canal, and generate a personalized acoustic transfer function based on the head-related transfer function database; The step 2 comprises the following steps: Measuring pure tone signals: Playing multi-band pure tone stimulation through an external speaker, and inserting a probe microphone into the ear canal to record the sound wave response of the pure tone stimulation after passing through the ear canal to obtain a measurement signal; Calculate the acoustic transfer function of the ear canal: Based on the input signal and the measured output signal, the acoustic transfer function of the ear canal is calculated in the frequency domain: ; Where: is the frequency response of the ear canal, i.e., the acoustic transfer function; and Represent the frequency domain of the input signal and the output signal respectively; User HRTF acquisition: inferring the applicable HRTF function from the standard HRTF database based on the user's individual physical signs; Synthesizing personalized acoustic transfer function: obtaining the user's personalized acoustic transfer function based on the product of the inferred HRTF function and the calculated acoustic transfer function of the ear canal; Step 3: Play the calibrated audio through the binaural simulation microphones equipped on the artificial head, and record the binaural audio signals. At the same time, synchronize the audio playback and the EEG device through the optical coupling hardware trigger to obtain EEG data, and pre-process the EEG data to extract time domain features and time-frequency features based on the pre-processed EEG data; Step 4: Processing binaural audio signals based on filters to obtain time domain envelopes of each frequency band, and then calculating the dynamic range compression ratio of the time domain envelopes of each frequency band; calculating the time difference and energy ratio of binaural signals based on the cross-correlation method, and using personalized acoustic transfer functions and standard acoustic transfer function libraries to perform similarity calculations, quantifying the user's spatial positioning ability, and obtaining spatial auditory features; constructing a unified multimodal auditory feature vector based on the time domain features, time-frequency features, and spatial auditory features; Step 5: Based on the multimodal auditory feature vector, a multi-head attention mechanism is used to weight the multimodal feature vector to obtain a weighted multimodal feature vector; Step 6: Based on the weighted multimodal feature vector and frequency / intensity parameters, the hearing threshold is predicted through a deep neural network model, the threshold deviation of each frequency point is output, and then the standard threshold is obtained by combining it with the traditional pure tone audiometry method; Step 7: Based on the multimodal feature vector, the actual age of the user and the standard threshold, a temporal Transformer encoder is used to predict the hearing age of the user, and a hearing decline assessment is performed on the user based on the predicted hearing age.

2. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: After designing the personalized acoustic transfer function in step 2, the FIR filter is used to reversely compensate the headphone frequency response, including the following steps: Measure the frequency response of headphones: Play a standard pure tone signal through the headphones, use a probe microphone to record the signal output by the headphones, and calculate the frequency response of the headphones based on the signal output by the headphones: ; Where: represents the frequency domain representation of the headphone output signal; represents the frequency domain representation of the input signal; Indicates the frequency response of the headphones; FIR filter: Compensate the headphone output through the FIR filter: ; Where: represents the frequency response of the filter; Obtain FIR filter coefficients based on frequency response using discrete Fourier transform , the output signal is inversely compensated based on the filter coefficients: Where: Represents the convolution operation; represents the time domain representation of the input signal; represents the time domain representation of the output signal.

3. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The steps of obtaining and preprocessing EEG data in step 3 are as follows: Binaural audio signal playback: The calibrated audio signal is played through an artificial binaural simulation microphone. The audio signal includes multiple pure tones, standard speech materials, and multiple types of environmental noise. Different signal-to-noise ratios are set to form different test scenarios. During the test, the binaural audio signal is synchronized with the EEG device through an optically coupled hardware trigger. The EEG device records the brain's neural response to the binaural audio signal, while ensuring the time synchronization between the audio playback and the EEG recording. EEG signal preprocessing: Perform multi-scale wavelet transform on EEG signals to decompose the signals into sub-signals of multiple scales; Perform soft threshold denoising on the sub-signals of each scale to remove high-frequency noise; Adaptively filter the denoised signal to remove the noise introduced by the artifacts; The processed signal is reconstructed by inverse wavelet transform to obtain the preprocessed EEG signal.

4. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: Extracting time domain features and time-frequency features The following steps are involved: Time domain feature extraction: The pre-processed EEG signal is subjected to time domain analysis to extract the average potential, amplitude and waveform complexity features to obtain the extracted time domain features; Time-frequency feature extraction: Wavelet transform is used to perform time-frequency analysis on the preprocessed EEG signal to extract the features of each frequency band.

5. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The step 4 comprises the following steps: Filter processing of binaural audio signals: using a bandpass filter to decompose binaural audio signals into multiple frequency bands, each of which contains the energy distribution of the audio signal in that frequency band; Time domain envelope extraction and dynamic range compression ratio: In each frequency band, the time domain envelope of the signal is extracted by the envelope detector. When extracting, the Hilbert transform is used to obtain the time domain envelope; the dynamic range compression ratio of the time domain included in each frequency band is calculated; Cross-correlation calculates the time difference and energy ratio of binaural signals: the time difference is estimated by calculating the delay between binaural signals; the energy ratio is obtained by calculating the mean square value of binaural signals; Similarity calculation: The similarity between the personalized acoustic transfer function and the transfer function in the standardized acoustic transfer function library is calculated; the similarity uses the mutual normalized correlation as the similarity measure; Construct a multimodal auditory feature vector: Concatenate all features extracted from time domain features, time-frequency features, and spatial auditory features to construct a unified multimodal auditory feature vector.

6. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: When the multi-head attention mechanism weights the multimodal feature vector, the feature importance weight is introduced to adjust the attention weight. The following steps are involved: Feature importance calculation: The feature importance is calculated based on the variance of each sub-feature in the input vector in the multi-head attention mechanism, the correlation of the threshold, and the robustness under different noise conditions: ; ; ; Where: , , Respectively represent the variance of time domain features, time-frequency features, and spatial auditory features; , , They represent the correlation between time domain features, time-frequency features, and spatial auditory features respectively; , , They represent the robustness of time domain features, time-frequency features, and spatial auditory features respectively; , represents the hyperparameters used to adjust the contribution of variance, correlation, and robustness to the feature importance calculation; Based on the importance of the calculated time domain features, time-frequency features, and spatial auditory features, normalization is performed to obtain the feature importance weights of each modality; Combining feature importance with attention weight: When weighting the query and key of each modality, feature importance weight is introduced to obtain attention weight: ; Where: Represents the feature importance weight of each modality; represents the inner product of the query vector and the key vector; represents the scaling factor; express Function, the inner product result is used Function that ensures that all attention weights are positive and sum to 1.

7. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The step 6 comprises the following steps: Feature construction: concatenate the weighted multimodal feature vector and the frequency / intensity parameter to obtain a concatenated feature vector; Hearing threshold deviation prediction: using the concatenated feature vector as the input of a trained deep neural network model, and predicting the hearing threshold deviation through the deep neural network model; Pure tone audiometry threshold: Hearing threshold is obtained based on pure tone audiometry; Threshold fusion: The hearing threshold and hearing threshold deviation obtained by pure tone audiometry are summed, and the summed value is weighted and fused with the hearing threshold obtained by pure tone audiometry to obtain the standard threshold. The weights in weighted fusion are obtained based on Bayesian optimization, and the objective function of Bayesian optimization is as follows: ; Where: represents the actual hearing threshold of the ith frequency point; Represents the hearing threshold of the ith frequency point after fusion.

8. The adaptive hearing screening optimization method based on acoustic feature analysis according to claim 1, characterized in that: The temporal Transformer encoder includes an input embedding layer, a position encoding layer, a self-attention mechanism layer, a feedforward neural network layer, and an output layer; The features fused by the weighted multimodal feature vector and the frequency / intensity parameters are input to the input embedding layer, mapped through a fully connected layer to obtain the feature representation of each time step; and then the sequence information is introduced through the position encoding layer; The features processed by the position encoding layer are used to capture the dependencies between time steps in the time series through the self-attention mechanism layer, and the weighted relationship between each time step and other time steps is calculated to obtain the attention weight of each time step. Based on the attention weight of each time step, the weighted representation of each time step is obtained; the weighted representation of each time step is input into the feedforward neural network layer for nonlinear change, and the feature representation after nonlinear change enters the output layer and is mapped to the predicted hearing age through a fully connected layer.

Citation Information

Patent Citations

  • Hearing test method and hearing screening instrument for automatically correcting influence of environmental noise

    CN109480859A

  • Multifunctional hearing evaluation earphone and evaluation method thereof

    CN112315462A