Hearing threshold detection method based on multi-mode auditory electrophysiological signal fusion

By constructing a multi-task deep learning model that integrates swept-frequency OAE, ABR, and middle ear reflex signals, high-precision quantitative prediction of hearing thresholds and classification of hearing loss types across the entire auditory pathway are achieved. This solves the problem of existing technologies being unable to systematically integrate multi-dimensional physiological signals and provides an end-to-end intelligent hearing assessment solution.

CN121817871APending Publication Date: 2026-04-10SHANGHAI UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-10
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies lack a systematic integration of multidimensional heterogeneous physiological signals reflecting the cochlea, auditory center, and external and middle ear conduction pathways, making it impossible to achieve end-to-end, high-precision quantitative prediction of full-frequency hearing thresholds and accurate localization of hearing loss lesions.

Method used

A multi-task deep learning model integrating frequency-sweep OAE, ABR, and middle ear reflex is constructed. Multimodal auditory electrophysiological signals are collected through a unified frequency-sweep stimulus signal, and signal preprocessing and feature extraction are performed to train the deep learning model to achieve quantitative prediction of hearing threshold and classification of hearing loss type.

Benefits of technology

It achieves high-precision, wide dynamic range quantitative prediction of hearing thresholds across the entire auditory pathway, and simultaneously performs automatic classification and discrimination of the nature of hearing loss, providing comprehensive diagnostic information, reducing reliance on professionals, and is suitable for large-scale screening and populations with difficulties in cooperation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121817871A_ABST
    Figure CN121817871A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of crossing of medical instruments and artificial intelligence, and provides a hearing threshold detection method based on multi-mode hearing electrophysiological signal fusion. Comprising stimulation frequency otoacoustic emission and auditory brainstem response induced by sweep frequency stimulation sound, distortion product otoacoustic emission induced by dual-tone stimulation and middle ear sound immittance / reflection signals, and a multi-task deep learning model is constructed to perform fusion analysis on the signals. According to the method, accurate quantitative regression prediction of hearing thresholds with any frequency resolution and automatic classification and discrimination of hearing loss properties are synchronously realized, the problems of single evaluation dimension and insufficient complex hearing loss analysis ability in the prior art are effectively solved, the comprehensiveness, accuracy and clinical operation efficiency of objective hearing evaluation are remarkably improved, and the method is suitable for popularization and application. And an innovative system solution is provided for fine diagnosis of the auditory function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical devices and artificial intelligence, and particularly relates to a method for classifying the whole-pathway hearing loss state of an auditory system and objectively and quantitatively detecting a hearing threshold, based on a convolutional neural network and using multi-modal auditory electrophysiological signals, including stimulus-frequency otoacoustic emissions (SFOAEs), distortion-product otoacoustic emissions (DPOAEs), middle ear reflexes, and auditory brainstem responses (ABRs). BACKGROUND

[0002] Otoacoustic emissions (OAEs) are a byproduct of the cochlea's active process, which is a weak audio energy generated in the cochlea, conducted through the ossicular chain and tympanic membrane, and released into the external auditory canal. Its generation depends on the functional normal cochlear amplifier and the complete and healthy out-hair cells (OHCs). When OHCs are damaged, the cochlea's active process is weakened or disappears, accompanied by the same change in OAEs. This correlation makes OAEs represent the change in absolute hearing threshold in sensorineural hearing loss (SNHL) dominated by OHCs, and thus can be used as a non-invasive means for objectively and non-invasively reflecting the state of the cochlea.

[0003] OAEs can be divided into spontaneous otoacoustic emissions (SOAEs) and evoked otoacoustic emissions (EOAEs) according to different stimulation signals. According to the different stimulation signals used to induce OAEs, EOAEs can be divided into various types, including distortion-product otoacoustic emissions (DPOAEs), transient-evoked otoacoustic emissions (TEOAEs), and stimulus-frequency otoacoustic emissions (SFOAEs).

[0004] Currently, the primary clinical method for assessing the degree of hearing loss is pure-tone audiometry (PTA), with the absolute hearing threshold (ATH) measured by PTA being established as the "gold standard." However, the accuracy of hearing thresholds obtained from subjective feedback by the subjects is easily affected by subjective factors such as their attention, cooperation, and mental state. This is especially true in individuals with poor cooperation, such as infants, young children, patients with mental disorders, and those suspected of pseudodeafness, making it difficult to obtain reliable and objective results. To address the limitations of PTA in terms of subjective dependence and localization ability, transient evoked otoacoustic emissions (TEOAEs) and distortion product otoacoustic emissions (DPOAEs) are used clinically for qualitative screening, but these methods lack quantitative results for measuring hearing thresholds. Furthermore, electrophysiological testing techniques, such as the auditory brainstem response (ABR), have also been used for the objective assessment of hearing thresholds and lesion localization. However, their detection thresholds can only indirectly reflect hearing status and cannot provide quantitative diagnostic results for hearing thresholds. Moreover, the analysis of ABR waveforms relies on the subjective experience of experts. In summary, existing objective testing techniques generally suffer from insufficient sensitivity, complex operation, or lack of standardization, especially in their quantitative assessment capabilities for early and mild hearing loss, which are far inferior to those of PTA. Therefore, the research and development of novel objective ATH testing methods that are non-invasive, highly accurate, and easy to operate remains a bottleneck that urgently needs to be overcome in clinical audiology.

[0005] Objective assessment of auditory function primarily follows a development path from "signal acquisition" to "feature analysis" and then to "intelligent prediction." Early research focused on the reliable acquisition of otoacoustic emissions (OAEs). In recent years, OAE detection technology has developed to a highly integrated stage. For example, publication number CN218247217U discloses a device for screening TEOAEs and DPOAEs. Publication number CN120477757A realizes a comprehensive research-grade OAE instrument with customizable test parameters for rapid detection of multiple test parameters, including stimulation frequency OAEs (SFOAEs) and distortion product OAEs (DPOAEs). It includes fine structure detection in the frequency dimension and I / O function detection in the intensity dimension, making it suitable for various needs in OAE research and applications. However, the functionality of such advanced devices remains limited to "data presentation." How to translate these rich features into clinically usable diagnostic conclusions is a crucial aspect that has not yet been addressed. Furthermore, scholars have attempted to use single-type signals to predict or classify hearing status through statistical or machine learning models. For example, early scholars often derived thresholds from the growth function of distortion product otoacoustic emissions (DPOAEs) to estimate hearing thresholds. However, this method has a large error between the predicted hearing threshold and the actual value, and its applicability is limited. Publication number CN120658984A predicts cochlear hearing loss based on the fine structure of swept-frequency SFOAE. These works have gradually verified the feasibility of extracting information from physiological signals to assess hearing, but the information bottleneck of a single signal source always exists, and its prediction accuracy and robustness are often limited under complex auditory pathology conditions.

[0006] To overcome this bottleneck, technological development naturally leads to the exploration of multi-signal fusion. For example, some studies have combined ABR with DPOAE for assessing ototoxic drug damage in experimental animals (publication number CN120938422A). However, this approach is deeply tied to a specific scenario, and its goals and methods differ significantly from the quantitative prediction needs of human clinical hearing. More recently, a more direct advancement is seen in (CN120477757A), which for the first time integrates fine structural information from multi-frequency, multi-intensity SFOAE and DPOAE to construct an input feature matrix for prediction based on deep learning, pushing the predictive power of single modalities to a new level. However, it only uses OAE signals, limiting prediction results to 0 dBHL~65 dBHL, and does not consider the close relationship between hearing status and middle ear conduction and brainstem nerves. Single OAE signals mainly reflect the function of cochlear outer hair cells, making it difficult to effectively distinguish between conductive, sensorineural, and mixed hearing losses. For example, when middle ear lesions are present, OAE signals may completely disappear or become abnormal, but the degree of middle ear conduction loss cannot be directly quantified. Furthermore, the lack of assessment of neural pathways means that auditory echocardiograms (OAEs) cannot evaluate the neurological function of the auditory pathway behind the cochlea. For diseases such as auditory neuropathy spectrum disorder (ANSD), relying solely on OAEs to predict hearing thresholds may lead to serious biases. In contrast, the middle ear reflex (acoustic reflex) is an authoritative method for assessing the mechanical state of the middle ear conduction system (as the basis of sound conduction) and the integrity of the reflex arc formed by the auditory nerve -> brainstem -> facial nerve. It has irreplaceable value in differentiating conductive hearing loss, identifying loudness recruitment (cochlear hearing loss), and screening for large retrocochlear lesions (such as acoustic neuroma). Tone-Burst ABR (short pure-tone auditory brainstem response) directly records the electrical activity from the auditory nerve to the brainstem, reflecting the integrity and synchronization ability of the entire peripheral auditory pathway (from the cochlea to the brainstem). It is particularly irreplaceable in cases of severe hearing loss and neurological problems.

[0007] In summary, there is currently a lack of an objective quantitative detection technology for hearing thresholds across the entire auditory system. There is also a lack of a systematic integration of multi-dimensional heterogeneous physiological signals reflecting the conduction pathways of the cochlea, auditory center, and outer and middle ear, and a complete technical solution that utilizes deep learning models to achieve end-to-end, high-precision, full-frequency quantitative prediction of hearing thresholds and accurately locate the location of hearing loss lesions. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a hearing threshold detection method based on the fusion of multimodal auditory electrophysiological signals. It constructs a multi-task deep learning model that integrates swept-frequency OAEs, ABR, and middle ear reflexes, simultaneously achieving quantitative prediction of hearing thresholds at arbitrary frequency resolutions and classification of hearing loss types. It adopts the swept-frequency OAEs signal acquisition and preprocessing paradigm (i.e., using a uniformly patterned swept-frequency stimulus signal to extract SPOAEs and distortion product otoacoustic emissions (DPOAEs) signals), and makes key improvements and additions to enhance prediction accuracy, improve model robustness, and provide a more comprehensive clinical assessment of auditory function.

[0009] To achieve the above objectives, the present invention provides the following technical solution:

[0010] A hearing threshold detection method based on multimodal auditory electrophysiological signal fusion includes the following steps:

[0011] Step S110, Signal Acquisition: Using a unified frequency sweep stimulation signal paradigm, the time-domain signals of otoacoustic emissions and auditory brainstem response induced by the same frequency sweep pure tone stimulation sequence are acquired simultaneously; and based on a unified hardware clock triggering mechanism, the distortion product otoacoustic emission signals and the original time-domain response signals of middle ear acoustic impedance induced by the probe tone are acquired sequentially.

[0012] Step S120, Signal preprocessing and feature construction: The stimulation frequency otoacoustic emission signal, distortion product otoacoustic emission signal, auditory brainstem response signal and middle ear acoustic impedance signal are preprocessed, time-domain aligned and feature extracted respectively to construct a multimodal feature set;

[0013] Step S130, Model Training: Construct a dataset based on the multimodal feature set and its corresponding hearing threshold labels and hearing loss type labels, and train a deep learning model to obtain a trained hearing threshold prediction model and hearing loss classification model.

[0014] Step S140, Functional evaluation: For the sample to be tested, call the pre-trained model, input its multimodal physiological signal features, and output its predicted hearing threshold value at each target frequency point and the discrimination result of hearing loss type.

[0015] Preferably, in step S120, the constructed multimodal feature set takes one of the following two forms:

[0016] a) A normalized waveform tensor directly composed of the time-domain aligned multi-channel time series;

[0017] b) Feature vectors extracted from each modal signal, including amplitude, latency, signal-to-noise ratio, and acoustic parameters; where missing or unextracted feature components are filled with preset special coding values.

[0018] Preferably, the deep learning model is an end-to-end multi-task learning model, comprising:

[0019] A shared feature extraction layer is used to learn a shared high-dimensional feature representation from the input features;

[0020] A regression output layer is used to map the shared feature representation to continuous hearing threshold prediction values ​​for each target frequency point;

[0021] A classification output layer is used to map the shared feature representation to a classification probability distribution of hearing loss type;

[0022] The model is trained by jointly minimizing the regression loss and the classification loss.

[0023] Preferably, the shared feature extraction layer is implemented using a one-dimensional convolutional neural network; the regression output layer and the classification output layer are both implemented using fully connected networks.

[0024] Preferably, the joint minimization of regression loss and classification loss specifically means that the total loss function of the model is a weighted sum of the mean squared error loss of the regression task and the cross-entropy loss of the classification task.

[0025] Preferably, in step S130, K-fold cross-validation combined with an early stopping strategy is used to train and evaluate the model, and multiple fixed random seeds are used in the outer layer to repeat the experiment to evaluate the robustness of the model performance.

[0026] Preferably, in step S140, the target frequency point output includes one or more of 0.5, 0.75, 1, 1.5, 2, 3, 4, 6 and 8 kHz; the output hearing loss type includes normal, conductive, sensorineural and mixed.

[0027] A hearing threshold detection device based on multimodal auditory electrophysiological signal fusion, used to implement a hearing threshold detection method, comprising:

[0028] The signal acquisition module is used to control the acoustic stimulation generation unit, the electrophysiological signal acquisition unit, and the middle ear acoustic impedance testing unit, and to perform signal acquisition according to a unified master clock triggering sequence.

[0029] The signal processing and feature construction module is used to receive the raw time-domain signal from the signal acquisition module, and to preprocess, align and construct features of the raw time-domain signal to output a standardized multimodal feature set.

[0030] The model training module is used to manage training datasets, configure deep learning model architecture, loss function and optimization strategy, perform model training and evaluation, and generate and store trained model parameters.

[0031] The hearing function assessment module is used to load the model generated by the model training module, receive new sample features from the signal processing and feature construction module, perform functional assessment, and output an assessment report containing the predicted hearing threshold values ​​and hearing loss type discrimination results at each frequency point.

[0032] Preferably, the signal acquisition module specifically includes:

[0033] The stimulus generation and output unit is used to generate four-segment sweep frequency signals, a single-segment dual-tone sweep frequency signal, and a wideband / single-frequency probe tone, and outputs them through corresponding independent acoustic channels.

[0034] The synchronous acquisition unit is used to synchronously acquire the sound pressure signal in the ear canal and the electroencephalogram signal of the scalp electrode during the stimulation of the four-segment sweep frequency signal, and to trigger the acquisition of the distortion product otoacoustic emission signal and the middle ear acoustic impedance signal in a preset time sequence based on a unified hardware clock triggering mechanism.

[0035] The clock and trigger management unit provides a unified hardware clock reference and coordinates the precise timing of each stimulus output and signal acquisition action.

[0036] A computer program product includes a computer program that, when executed by a processor, implements the steps of a hearing threshold detection method; wherein the computer program product is a computer-readable storage medium storing the computer program, or an electronic device loaded with the computer program.

[0037] This invention provides a hearing threshold detection method based on the fusion of multimodal auditory electrophysiological signals. It has the following beneficial effects:

[0038] This system integrates frequency-of-stimulation otoacoustic emissions (SFOAE), distortion product otoacoustic emissions (DPOAE), auditory brainstem response (ABR), and middle ear acoustic impedance testing onto a single hardware platform. By designing a unified frequency-sweeping stimulation paradigm and a master clock triggering system, it achieves one-stop, synchronous or quasi-synchronous acquisition of key physiological signals reflecting the entire auditory pathway (cochlea, auditory nerve / brainstem, and middle ear). This fundamentally solves the problems of long testing times, cumbersome operation, difficulty in data spatiotemporal alignment, and subject state fluctuations caused by multi-device, multi-stage testing, laying a solid hardware foundation for obtaining high-quality, highly consistent multimodal fusion data.

[0039] This invention breaks through the information bottleneck of a single signal source by systematically integrating multidimensional heterogeneous signals covering key aspects of peripheral and central hearing. Through a multi-task deep learning model, it performs deep fusion analysis on SFOAE (reflecting cochlear nonlinearity), DPOAE (reflecting cochlear active mechanisms), ABR (reflecting neural pathway synchronicity), and middle ear acoustic impedance (reflecting conduction pathway status). The model can automatically learn complementary and correlated features across modalities. This not only achieves high-precision, wide dynamic range (e.g., covering hearing loss from normal to severe) objective quantitative prediction of hearing thresholds at various target frequencies, but also simultaneously enables automatic classification and discrimination of the nature of hearing loss (normal, conductive, sensorineural, mixed), providing comprehensive diagnostic information far exceeding that of a single modality.

[0040] This invention provides an end-to-end intelligent solution from raw signals to diagnostic conclusions. Utilizing a pre-trained multi-task model, it can rapidly analyze multimodal signals from new subjects, outputting a complete auditory function assessment report including quantitative hearing thresholds and qualitative classifications in a single output. This process significantly reduces the need for complex waveform interpretation and experience-based expertise from professionals, substantially improving the automation level, diagnostic efficiency, and consistency of objective hearing assessments. It is particularly suitable for large-scale screenings, populations with difficulties in cooperation, and primary healthcare settings.

[0041] The multimodal fusion framework and model architecture proposed in this invention have good versatility. The frequency sweep stimulation paradigm can efficiently acquire wideband continuous responses; the deep learning model can adapt to different forms of feature input (such as raw waveform tensors or artificial feature vectors). This scheme reserves flexible technical expansion space for future integration of more types of physiological signals (such as cortical auditory evoked potentials) or combination with other clinical indicators to further improve assessment performance. Attached Figure Description

[0042] Figure 1 This is one of the process diagrams provided by the present invention;

[0043] Figure 2 The test signal paradigm diagrams for swept frequency distortion product otoacoustic emissions (DPOAEs) and stimulation frequency otoacoustic emissions (SFOAEs) provided by this invention;

[0044] Figure 3 This is a schematic diagram of the structure of a multi-task deep learning model provided in an embodiment of the present invention;

[0045] Figure 4 A schematic diagram of the structure of the method and apparatus provided by the present invention;

[0046] Figure 5 This is a structural block diagram provided for an embodiment of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] See attached document Figures 1-2 As shown, in one embodiment, a hearing threshold detection method based on multimodal auditory electrophysiological signal fusion includes the following steps:

[0049] Step S110, Signal Acquisition: Using a unified frequency sweep stimulation signal paradigm, the time-domain signals of stimulation frequency otoacoustic emissions (SFOAE) and auditory brainstem response (Tone-Burst ABR) induced by the same frequency sweep pure tone stimulation sequence are acquired simultaneously; and based on a unified hardware clock triggering mechanism, the distortion product otoacoustic emissions (DPOAE) signal and the original time-domain response signal of middle ear acoustic impedance induced by the probe tone are acquired sequentially.

[0050] The SFOAE and ABR signals acquired simultaneously are obtained using a four-segment frequency sweep signal stimulation paradigm, wherein the stimulation sound exists in all four segments and the phases alternate, and the inhibition sound exists in the last two segments.

[0051] The sequentially acquired DPOAE signals are obtained using a single-segment sweep frequency signal stimulation paradigm and are induced by two fundamental sweep frequency signals with a fixed frequency ratio.

[0052] The middle ear acoustic impedance signal is induced by a broadband or single-frequency probe tone;

[0053] Specifically, the signals are acquired on the same hardware device, using a unified sweep frequency stimulation signal paradigm, and simultaneously acquiring the SFOAE time-domain signal and the auditory brainstem response (Tone-Burst ABR) time-domain EEG signal induced by the same sweep frequency pure tone stimulation sequence. The distortion product otoacoustic emission (DPOAE) signal and the middle ear acoustic impedance primitive time-domain response signal induced by probe tones (e.g., broadband probe tones or single-frequency probe tones) are acquired sequentially.

[0054] Simultaneous acquisition of SFOAE and ABR signals: using a four-segment sweep frequency signal stimulation paradigm (segments A, B, C, and D);

[0055] The stimulus sound exists in each segment of the four frequency sweep signals A, B, C, and D, and presents an initial phase that alternates between positive and negative (for example, positive phase in segment A, negative phase in segment B, positive phase in segment C, and negative phase in segment D); its frequency sweep range is 0.5~8kHz.

[0056] The suppressor sound exists only in segments C and D with the same initial phase; the instantaneous frequency ratio of the suppressor sound to the stimulus sound remains at N, with N ranging from 0.9 to 1.1 (e.g., Figure 2 In the example, N is 1.05.

[0057] The SFOAE signal is obtained by calculating the residuals of the four-segment re-collection signals; the calculation formula is: ,in , , , These are the time-domain sound pressure signals corresponding to the four stimuli; the stimuli and inhibition sounds can be transmitted to the human ear through two independent sound tubes.

[0058] While applying four-segment frequency-sweeping stimulation sounds, raw time-domain EEG signals of auditory brainstem response (Tone-Burst ABR) are simultaneously acquired via scalp electrodes.

[0059] Sequential acquisition of DPOAE signals: A single-segment sweep frequency signal stimulation paradigm (single-segment stimulation paradigm) is adopted, and the segment duration can be, for example, 1 second.

[0060] The stimulus sound consists of a first fundamental tone (f1) sweep signal and a second fundamental tone (f2) sweep signal. The instantaneous frequency ratio between the two remains M, i.e., f2 / f1 = M, where M takes the value between 1.0 and 1.4 (e.g., Figure 2 In the example, M is 1.22); the sweep frequency range of the second fundamental tone (f2) is 0.5~8kHz. The first and second fundamental tones are transmitted to the ear through two separate acoustic tubes.

[0061] The instantaneous frequency of the DPOAE signal derived from this is 2f1-f2. This signal is included in the total re-sampling signal with pitch re-sampling artifacts and needs to be extracted by subsequent filtering.

[0062] Acquiring middle ear acoustic impedance signals: induced by a wideband probe tone or a 226Hz single-frequency probe tone;

[0063] The raw time-domain response signals of the middle ear acoustic impedance, such as tympanograms and acoustic reflex thresholds, were collected from the subjects to obtain middle ear functional parameters.

[0064] The original time-domain waveforms of the distortion product otoacoustic emission and the original time-domain waveforms of the middle ear acoustic impedance were independently and sequentially acquired based on the same hardware clock triggering mechanism within a time period different from the simultaneous recording of the SFOAE / ABR. All original signals (SFOAE, ABR, DPOAE, middle ear acoustic impedance) were time-domain aligned using the unified time reference during post-processing.

[0065] Step S120, Signal preprocessing and feature construction: The stimulation frequency otoacoustic emission signal, distortion product otoacoustic emission signal, auditory brainstem response signal and middle ear acoustic impedance signal are preprocessed, time-domain aligned and feature extracted respectively to construct a multimodal feature set;

[0066] The preprocessing includes: using a tracking filter with dynamic zero-pole changes to purify the original time-domain response signals of SFOAE and DPOAE; the time-domain alignment includes: resampling all signals to the same frequency based on a unified clock timestamp, aligning the time zeros with the stimulus start point as a reference, extracting a fixed time window, and forming a time-domain aligned multi-channel time series.

[0067] Specifically, the signal preprocessing and feature construction process the SFOAE signal, DPOAE signal, auditory brainstem response signal, and middle ear acoustic impedance signal to construct a multimodal feature set for characterizing the function of the auditory system.

[0068] Signal preprocessing uses a tracking filter with dynamically changing zeros and poles to filter out interference from the acquired single-shot DPOAE or SFOAE raw time-domain response signal. For specific methods, please refer to the prior patent CN120670992A.

[0069] For the DPOAE signal, the dynamic poles of the tracking filter are set to the DPOAE extraction frequency (2f1-f2).

[0070] For the SFOAE signal, the residual signal is first calculated, and then the dynamic poles of the tracking filter are set to the stimulation frequency and the dynamic zeros are set to the suppression frequency, so as to extract the purified SFOAE time-domain waveform from the residual.

[0071] Time-domain alignment: Based on the acquisition timestamp recorded by a unified master clock, the purified swept-frequency OAE waveform, the synchronously recorded auditory brainstem response raw time-domain EEG signal, and the synchronously / sequentially acquired middle ear acoustic impedance raw time-domain waveform are uniformly resampled to the same sampling frequency.

[0072] Time zero-point alignment is performed using the start emission time of each stimulus sequence (SFOAE / ABR stimulus, DPOAE stimulus, probe tone) as the absolute time reference.

[0073] By extracting fixed time windows of the same length, a multi-channel time series that is strictly aligned in the time domain is obtained.

[0074] Step S130, Model Training: Construct a dataset based on the multimodal feature set and its corresponding hearing threshold labels and hearing loss type labels, and train a deep learning model to obtain a trained hearing threshold prediction model and hearing loss classification model.

[0075] Specifically, the dataset includes a label set and a feature set;

[0076] The feature set contains the original time-domain waveforms of SFOAE signals and DPOAE signals under multiple test intensities, the original time-domain EEG signals of auditory brainstem response, and the original time-domain waveforms of middle ear acoustic impedance, forming a multi-channel waveform tensor or feature vector.

[0077] The label set (dual-task label) includes a subset of regression labels and a subset of classification labels;

[0078] The regression label subset consists of the continuous hearing thresholds (dBHL) of each target frequency point to be predicted. During construction, regularization rules can be applied: label values ​​below the preset minimum hearing threshold (e.g., -10dBHL) are set as the minimum hearing threshold, and label values ​​above the preset maximum hearing threshold (e.g., 120dBHL) are set as the maximum hearing threshold. At the same time, all hearing threshold labels are quantized to a preset fixed step size (e.g., 5dB).

[0079] The classification label subset consists of discrete hearing loss types determined according to clinical diagnostic criteria, including at least: normal hearing loss, conductive hearing loss, sensorineural hearing loss, and mixed hearing loss;

[0080] The regression label subset and the classification label subset together form the complete label set for multi-task learning.

[0081] like Figure 3 As shown, the construction of the deep learning (DL) model method specifically includes three parts: dataset, model architecture, and model evaluation. The dataset is constructed as follows: a waveform tensor built from the original multi-channel time series. The label set contains two parts: first, the hearing thresholds (continuous values) for multiple frequencies of interest to be predicted; second, the corresponding hearing loss types (discrete categories). The model architecture is an end-to-end multi-task learning architecture, including an input layer, a shared feature extraction layer, and a dual-task output layer: the input layer receives the waveform tensor; the shared feature extraction layer uses a one-dimensional convolutional neural network (1D-CNN) framework to automatically learn the deep time-domain representation of the signal; the dual-task output layer uses parallel fully connected layers to respectively realize the regression prediction of hearing thresholds and the classification of loss types. Cross-validation is used for model evaluation, and the model performance is comprehensively evaluated by metrics such as the mean absolute error (MAE) of the regression task and the accuracy of the classification task.

[0082] In the specific process of dataset construction, the waveform tensor (feature set) is organized by samples from the multi-channel time series processed in step S120. The tensor dimension of each sample (i.e., a test record) is [C,T], where C is the total number of channels, determined by different signal sources (SFOAE, ABR, DPOAE, middle ear signal) and their sub-channels at various test intensities; T is the unified number of time points after alignment and truncation. The construction of the label set includes dual-task labels: regression labels are the hearing thresholds of each target frequency after normalization (e.g., setting upper and lower limits); classification labels are the hearing loss types determined according to clinical diagnosis. The model is trained end-to-end with the goal of minimizing the joint loss function (e.g., the weighted sum of regression mean square error and classification cross-entropy).

[0083] In the model construction process, to fully leverage the potential of multimodal auditory electrophysiological signals for hearing threshold prediction and pathological discrimination, an end-to-end multi-task deep learning architecture was chosen. This architecture takes multi-channel raw time-series tensors as input and consists of a shared feature extraction backbone and two parallel task-specific outputs. The shared feature extraction backbone is mainly composed of a one-dimensional convolutional neural network (1D-CNN) used to automatically learn temporal and cross-channel abstract features in the waveform. Specifically, a multi-layer 1D-CNN is used to process the input tensor. For example, the first layer can use multiple one-dimensional convolutional kernels of size k to extract features along the time dimension; subsequent layers can be stacked, with pooling layers intermittently inserted to compress the feature dimension. Batch normalization, ReLU activation, and Dropout regularization can be performed after each convolutional layer. Finally, a global pooling layer aggregates the temporal features into a fixed-dimensional shared high-dimensional feature vector.

[0084] The dual-task output layer receives the shared feature vector and processes it in two branches: For the regression part, it consists of one or more fully connected layers, ultimately mapping the features to a scalar output representing the predicted hearing threshold value at the target frequency. Constraints can be applied to the output to limit the predicted value to a reasonable physiological range (e.g., -10 to 120 dBHL). For the classification part, it consists of one or more fully connected layers and a Softmax activation function, mapping the features to a probability distribution vector representing the predicted probability that the input sample belongs to each type of hearing loss (e.g., normal, conductive, sensorineural, mixed). The model is trained end-to-end using a joint loss function, which is a weighted sum of the regression task loss (e.g., mean squared error MSE) and the classification task loss (e.g., cross-entropy loss), thereby simultaneously optimizing both tasks. The hyperparameters in the above network structure, such as the number of layers, kernel size, and number of neurons, can be adaptively adjusted according to the actual data scale and task requirements.

[0085] During model evaluation, 6-fold cross-validation was used during training. The dataset was stratified and divided into approximately equal numbers of 6 folds. One fold subset was used as the test set in rotation, and another fold was randomly stratified and used as the validation set from the remaining 5 folds. The remaining 4 folds served as the training set. The training set was used to optimize model parameters to fit the relationship between the data and labels. The validation set was used to determine the optimal combination of hyperparameters during model training. The test set was used to provide the model's generalization prediction performance. Training used the Adam optimizer with a learning rate of 0.001, a batch size of 32, a maximum training epoch of 200, and an early stopping mechanism (patience value of 15 epochs). The mean absolute error (MAO) loss function was used. A 20-cycle fixed random seed was used in the outer layer to evaluate robustness. The generalization error of each cross-validation was the average of the errors from the 6-fold test set, and the overall generalization error was calculated as the average result of 20 cross-validations with the fixed random seed setting.

[0086] Step S140, Functional evaluation: For the sample to be tested, call the pre-trained model, input its multimodal physiological signal features, and output its predicted hearing threshold value at each target frequency point and the discrimination result of hearing loss type.

[0087] Specifically, the functional assessment steps are as follows: A pre-trained hearing threshold prediction model and loss type discrimination model are invoked to predict and classify the multimodal physiological signals of the new subject; after obtaining the trained model, the multimodal physiological signal features (waveform tensors or feature vectors) obtained from processing the test sample are input into the multi-task learning model; after model processing and analysis, two sets of results are output simultaneously:

[0088] Predicted hearing threshold values ​​for each desired target frequency point (such as 0.5, 0.75, 1, 1.5, 2, 3, 4, 6 and 8 kHz).

[0089] The test ear determines the type of hearing loss (e.g., normal, conductive, sensorineural, mixed) or its probability distribution.

[0090] This allows for the output of a complete auditory function assessment report that includes both quantitative predictions and qualitative categories in a single step.

[0091] In step S120, the constructed multimodal feature set is one of the following two forms: a) a normalized waveform tensor directly composed of the time-domain aligned multi-channel time series;

[0092] b) Feature vectors extracted from each modal signal, including amplitude, latency, signal-to-noise ratio, and acoustic parameters; where missing or unextracted feature components are filled with preset special coding values.

[0093] Specifically, feature set construction: The constructed multimodal feature set can take one of the following two forms:

[0094] Form 1 (Waveform Tensor): A strictly time-domain aligned multi-channel time series is concatenated and standardized along the channel dimension (e.g., Z-score normalization) to form a standardized multimodal raw waveform feature set (i.e., waveform tensor), which serves as the direct input to the deep learning model. The waveform tensor dimension of each sample (i.e., a test record) is [C, T], where C is the total number of channels, determined by different signal sources (SFOAE, ABR, DPOAE, middle ear signal) and their sub-channels at various test intensities; T is the number of unified time points after alignment and truncation.

[0095] Form 2 (Feature Vector): Features are manually extracted from each modal signal to form a feature vector, including:

[0096] Extract spectral features such as amplitude and signal-to-noise ratio at each frequency point from SFOAE and DPOAE signals.

[0097] Electrophysiological parameters such as latency and amplitude of waves I, III, and V were extracted from the auditory brainstem response (ABR) signal.

[0098] Acoustic parameters such as peak pressure, gradient, static acoustic compliance, and acoustic reflection threshold of the 226Hz tympanogram were extracted from the middle ear acoustic impedance signal.

[0099] Special coding rules: For missing or unextracted feature components, pre-defined special coding values ​​are used to fill them in. For example, a first specific coding value (such as -1.0) is used to indicate that the feature is missing because the test was not performed; a second specific coding value (such as -100.0) is used to indicate that the physiological signal corresponding to the feature was tested but not extracted.

[0100] The deep learning model is an end-to-end multi-task learning model, including:

[0101] A shared feature extraction layer is used to learn a shared high-dimensional feature representation from the input features;

[0102] A regression output layer is used to map the shared feature representation to continuous hearing threshold prediction values ​​for each target frequency point;

[0103] A classification output layer is used to map the shared feature representation to a classification probability distribution of hearing loss type;

[0104] The model is trained by jointly minimizing the regression loss and the classification loss.

[0105] Specifically, the deep learning model is an end-to-end multi-task learning model.

[0106] The shared feature extraction layer employs a convolutional neural network framework to automatically learn and output a high-dimensional shared feature representation from the input multi-channel waveform tensor or feature vector. Specifically, it can be implemented using a one-dimensional convolutional neural network (1D-CNN). For example, using a multi-layer 1D-CNN, batch normalization can be performed after the convolutional layers, the Leaky ReLU activation function can be used, and pooling layers (such as global average pooling) can be intermittently inserted to compress the feature dimension, ultimately aggregating into a fixed-dimensional shared high-dimensional feature vector.

[0107] The dual-task output layer includes regression output and classification output;

[0108] The regression output maps the shared feature vector to continuous hearing threshold prediction values ​​for each target frequency point through a fully connected network; constraints can be applied after the output to limit the prediction values ​​to a reasonable physiological range (e.g., -10 to 120 dBHL); target frequency points may include, but are not limited to: 0.5 kHz, 0.75 kHz, 1 kHz, 1.5 kHz, 2 kHz, 3 kHz, 4 kHz, 6 kHz, 8 kHz.

[0109] The classification output transforms the shared feature vector into a predicted probability distribution for various types of hearing loss through a fully connected network and a Softmax activation function.

[0110] The shared feature extraction layer is implemented using a one-dimensional convolutional neural network (1D-CNN); the regression output layer and the classification output layer are both implemented using fully connected networks.

[0111] The joint minimization of regression loss and classification loss is specifically defined as follows: the total loss function of the model is a weighted sum of the mean squared error loss of the regression task and the cross-entropy loss of the classification task.

[0112] In step S130, K-fold cross-validation combined with an early stopping strategy is used to train and evaluate the model, and multiple fixed random seeds are used in the outer layer to repeat the experiment to evaluate the robustness of the model performance.

[0113] Specifically, a K-fold cross-validation strategy combined with an early stopping strategy is employed. For example, the dataset is stratified and divided into approximately equal numbers of 6 folds. 6-fold cross-validation is used, with one fold serving as the test set, one fold from the remaining five folds serving as the validation set, and the remaining four folds serving as the training set. Training uses the Adam optimizer with a learning rate of 0.001, a batch size of 32, a maximum training epoch of 200, and an early stopping mechanism (stopping if the validation set loss does not decrease for 15 consecutive epochs).

[0114] Loss function: The model is optimized by minimizing a joint loss function, which is a weighted sum of the loss of the regression task (such as mean squared error, MSE) and the loss of the classification task (such as cross-entropy), thereby optimizing both tasks simultaneously.

[0115] The cross-validation training and evaluation process described above is repeated in the outer layer using multiple different fixed random seeds (e.g., 20) to assess the robustness of the model performance. Model performance can be comprehensively evaluated using metrics such as mean absolute error (MAE) for regression tasks and accuracy for classification tasks.

[0116] In step S140, the target frequency points output include one or more of 0.5, 0.75, 1, 1.5, 2, 3, 4, 6, and 8 kHz; the output hearing loss types include normal, conductive, sensorineural, and mixed.

[0117] like Figure 4 As shown, in one embodiment, a hearing threshold detection device based on multimodal auditory electrophysiological signal fusion is used to implement a hearing threshold detection method, including:

[0118] The signal acquisition module 810 is used to control the acoustic stimulation generation unit, the electrophysiological signal acquisition unit, and the middle ear acoustic impedance testing unit, and to perform signal acquisition according to a unified master clock trigger sequence.

[0119] The signal processing and feature construction module 820 is used to receive the raw time-domain signal from the signal acquisition module 810, and to preprocess, align and construct features of the raw time-domain signal to output a standardized multimodal feature set.

[0120] The model training module 830 is used to manage training datasets, configure deep learning model architecture, loss function and optimization strategy, perform model training and evaluation, and generate and store trained model parameters.

[0121] The hearing function assessment module 840 is used to load the model generated by the model training module 830, receive new sample features from the signal processing and feature construction module 820, perform functional assessment, and output an assessment report containing the predicted hearing threshold values ​​and hearing loss type discrimination results at each frequency point.

[0122] The signal acquisition module 810 specifically includes:

[0123] Stimulus generation and output unit: used to generate the four-segment sweep frequency signal, single-segment dual-pitch sweep frequency signal and wideband / single-frequency probe tone, and output them through the corresponding independent acoustic channels.

[0124] Synchronous acquisition unit: used to synchronously acquire the sound pressure signal (SFOAE) in the ear canal and the electroencephalogram (ABR) of the scalp electrodes during the stimulation of the four-segment sweep frequency signals, and based on the unified hardware clock triggering mechanism, to trigger the acquisition of the distortion product otoacoustic emission signal and the middle ear acoustic impedance signal in a preset time sequence.

[0125] Clock and Trigger Management Unit: Used to provide a unified hardware clock reference and coordinate the precise timing of each stimulus output and signal acquisition action.

[0126] like Figure 5 As shown, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of a hearing threshold detection method; wherein, the computer program product is a computer-readable storage medium storing the computer program, or an electronic device loaded with the computer program.

[0127] Specifically, the electronic device includes a processor, a memory (including non-volatile storage media and internal memory), and a network interface connected via a system bus. The non-volatile storage media stores an operating system and a computer program. When the processor executes the computer program, it implements the steps of a hearing threshold detection method; specifically, this includes: controlling an integrated testing device to complete integrated signal acquisition; performing synchronous alignment, resampling, and standardization processing on the multimodal raw waveforms; constructing a dataset and training a multi-task deep learning model; calling the trained model to analyze new samples and outputting a hearing function assessment report.

[0128] Specifically, a computer-readable storage medium (such as ROM, PROM, EEPROM, flash memory, RAM, hard disk, etc.) stores a computer program (instructions); when the computer program is executed by a processor, it implements the steps of the hearing threshold detection method.

[0129] By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0130] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A hearing threshold detection method based on multimodal auditory electrophysiological signal fusion, characterized in that, Includes the following steps: Step S110, Signal Acquisition: Using a unified frequency sweep stimulation signal paradigm, the time-domain signals of otoacoustic emissions and auditory brainstem response induced by the same frequency sweep pure tone stimulation sequence are acquired simultaneously; and based on a unified hardware clock triggering mechanism, the distortion product otoacoustic emission signals and the original time-domain response signals of middle ear acoustic impedance induced by the probe tone are acquired sequentially. Step S120, Signal preprocessing and feature construction: The stimulation frequency otoacoustic emission signal, distortion product otoacoustic emission signal, auditory brainstem response signal and middle ear acoustic impedance signal are preprocessed, time-domain aligned and feature extracted respectively to construct a multimodal feature set; Step S130, Model Training: Construct a dataset based on the multimodal feature set and its corresponding hearing threshold labels and hearing loss type labels, and train a deep learning model to obtain a trained hearing threshold prediction model and hearing loss classification model. Step S140, Functional evaluation: For the sample to be tested, call the pre-trained model, input its multimodal physiological signal features, and output its predicted hearing threshold value at each target frequency point and the discrimination result of hearing loss type.

2. The method according to claim 1, characterized in that, In step S120, the constructed multimodal feature set takes one of the following two forms: a) A normalized waveform tensor directly composed of the time-domain aligned multi-channel time series; b) Feature vectors extracted from each modal signal, including amplitude, latency, signal-to-noise ratio, and acoustic parameters; where missing or unextracted feature components are filled with preset special coding values.

3. The method according to claim 1, characterized in that, The deep learning model is an end-to-end multi-task learning model, including: A shared feature extraction layer is used to learn a shared high-dimensional feature representation from the input features; A regression output layer is used to map the shared feature representation to continuous hearing threshold prediction values ​​for each target frequency point; A classification output layer is used to map the shared feature representation to a classification probability distribution of hearing loss type; The model is trained by jointly minimizing the regression loss and the classification loss.

4. The method according to claim 3, characterized in that, The shared feature extraction layer is implemented using a one-dimensional convolutional neural network; the regression output layer and the classification output layer are both implemented using fully connected networks.

5. The method according to claim 3 or 4, characterized in that, The joint minimization of regression loss and classification loss is specifically defined as follows: the total loss function of the model is a weighted sum of the mean squared error loss of the regression task and the cross-entropy loss of the classification task.

6. The method according to claim 1, characterized in that, In step S130, K-fold cross-validation combined with an early stopping strategy is used to train and evaluate the model, and multiple fixed random seeds are used in the outer layer to repeat the experiment to evaluate the robustness of the model performance.

7. The method according to claim 1, characterized in that, In step S140, the target frequency points output include one or more of 0.5, 0.75, 1, 1.5, 2, 3, 4, 6, and 8 kHz; the output hearing loss types include normal, conductive, sensorineural, and mixed.

8. A hearing threshold detection device based on multimodal auditory electrophysiological signal fusion, used to implement the hearing threshold detection method according to any one of claims 1-7, characterized in that, include: The signal acquisition module is used to control the acoustic stimulation generation unit, the electrophysiological signal acquisition unit, and the middle ear acoustic impedance testing unit, and to perform signal acquisition according to a unified master clock triggering sequence. The signal processing and feature construction module is used to receive the raw time-domain signal from the signal acquisition module, and to preprocess, align and construct features of the raw time-domain signal to output a standardized multimodal feature set. The model training module is used to manage training datasets, configure deep learning model architecture, loss function and optimization strategy, perform model training and evaluation, and generate and store trained model parameters. The hearing function assessment module is used to load the model generated by the model training module, receive new sample features from the signal processing and feature construction module, perform functional assessment, and output an assessment report containing the predicted hearing threshold values ​​and hearing loss type discrimination results at each frequency point.

9. The detection device according to claim 8, characterized in that, The signal acquisition module specifically includes: The stimulus generation and output unit is used to generate four-segment sweep frequency signals, a single-segment dual-tone sweep frequency signal, and a wideband / single-frequency probe tone, and outputs them through corresponding independent acoustic channels. The synchronous acquisition unit is used to synchronously acquire the sound pressure signal in the ear canal and the electroencephalogram signal of the scalp electrode during the stimulation of the four-segment sweep frequency signal, and to trigger the acquisition of the distortion product otoacoustic emission signal and the middle ear acoustic impedance signal in a preset time sequence based on a unified hardware clock triggering mechanism. The clock and trigger management unit provides a unified hardware clock reference and coordinates the precise timing of each stimulus output and signal acquisition action.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the hearing threshold detection method according to any one of claims 1 to 7; wherein the computer program product is a computer-readable storage medium storing the computer program, or an electronic device loaded with the computer program.

Citation Information

Patent Citations

  • Otoacoustic emission signal detection system and method

    CN120477757A

  • Hearing threshold prediction system and method based on sweep frequency SFOAE fine structure

    CN120658984A

  • Hearing threshold prediction method based on frequency sweep OAEs and deep learning model

    CN120670992A

  • Method and system for collecting and analyzing auditory electrophysiological data of experimental animal

    CN120938422A

  • Device for TEOAE and DPOAE otoacoustic emission screening

    CN218247217U

Cited By

  • Data set construction method and device for prevention and treatment of occupational hearing loss

    CN122177500A