Online self-adaptive ABR threshold value automatic identification method and system
By combining online adaptive ABR threshold automatic identification method with single-intensity and multi-intensity models, and adjusting the stimulus intensity in real time, the subjectivity and real-time issues of ABR threshold identification in existing technologies are solved, and efficient and accurate ABR threshold detection is achieved.
Patent Information
- Application Number
- CN202511478985.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-01-23
AI Technical Summary
Existing ABR threshold recognition technologies rely on manual interpretation, which is highly subjective and inefficient. Traditional machine learning methods have weak real-time performance, making it difficult to achieve online adaptive detection, and they lack robustness in low signal-to-noise ratio scenarios.
An online adaptive ABR threshold automatic identification method is adopted, which combines a single-intensity ABR detection model and a multi-intensity ABR sequence threshold discrimination model. Through confidence score discrimination and repeated detection strategy, the stimulus intensity is adjusted in real time to realize a detection process that combines coarse search and fine search.
It enables real-time, reliable, and accurate determination of ABR thresholds during clinical testing, reduces invalid intensity point testing, improves testing efficiency and diagnostic accuracy, and adapts to different signal-to-noise ratio environments.
Smart Images

Figure CN121370153A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of medical signal processing, artificial intelligence and auditory nerve diagnosis, and particularly relates to an online adaptive ABR threshold automatic identification method and system. BACKGROUND
[0002] Auditory brainstem response (ABR) is a bioelectric signal induced by short sound or short pure tone auditory stimulation, and is widely used in clinical hearing function detection scenes such as newborn hearing screening, nervous system disease diagnosis and anesthesia depth assessment. ABR is usually recorded by a non-invasive electrode placed on the scalp, and its typical characteristics are I to VII series of wave peaks, each of which corresponds to the neural activity of different anatomical structures in the auditory pathway. In clinical practice, a series of progressively decreasing stimulus intensities (such as sound pressure level, SPL, or normalized hearing level, nHL) are set to obtain single-ear multi-intensity ABR stacked waveforms, so as to realize objective hearing evaluation. The ABR threshold refers to the lowest stimulus intensity level at which the response can be detected in the multi-level stacked waveforms. This index is particularly important for special groups who cannot complete behavioral audiometry (such as infants, developmentally delayed patients, etc.), therefore, accurate identification of the ABR threshold is an important basis for carrying out objective hearing diagnosis. The existing ABR threshold identification technology mainly includes:
[0003] (1) Manual interpretation: The clinical ABR threshold identification technology mainly relies on the manual interpretation of audiologists, and the waveforms are collected under a series of stimulus intensities, and the presence or absence of V wave is determined by the experience of audiologists. This method is highly dependent on the experience and technical level of audiologists, and is highly subjective and low in efficiency, making it difficult to promote in primary medical institutions where experienced audiologists are lacking;
[0004] (2) Statistical / classical signal methods (Fmp, Fsp, Hotelling's T2, cross-correlation, correlation coefficient, cross-covariance, etc.), which are mostly post-batch analysis; real-time performance is weak, and is unstable in low signal-to-noise ratio scenarios;
[0005] (3) Traditional machine learning and deep learning: focusing on the identification of “presence / absence of response” of a single waveform, ignoring the dependence between multiple waveforms in morphology and timing, making it difficult to meet the actual needs of clinicians for accurate and robust threshold determination. At the same time, most of the current methods are offline threshold estimation, and most methods cannot judge and adaptively adjust the next stimulation during the clinical collection process, resulting in a long process time. SUMMARY
[0006] To solve the above problems, the application provides an online adaptive ABR threshold automatic identification method and system.
[0007] In a first aspect, the application provides an online adaptive ABR threshold automatic identification method, comprising:
[0008] S0. Set the maximum stimulation intensity L max , the minimum stimulation intensity L min , the minimum step size S min , the confidence threshold alpha, and the maximum number of repeated detections N max ; initialize the step size S and the stimulation intensity L;
[0009] S1. Collect the ABR waveform signal at the stimulation intensity L and perform preprocessing to obtain a preprocessed waveform;
[0010] S2. Input the preprocessed waveform into a single-intensity ABR detection model to output a discrimination result and a confidence score, the discrimination result is used to display whether there is a V wave; determine whether the confidence score is not less than the confidence threshold alpha, if yes, execute step S4, if not, execute step S3;
[0011] S3. Determine whether the maximum number of repeated detections N max is reached, if yes, use the result of the last detection to execute step S4, if not, execute step S1;
[0012] S4. Determine whether S=S min is true, if yes, execute step S6; if not, check whether the discrimination result of the current preprocessed waveform shows that there is a V wave, if yes, execute step S5, if not, first update S=S / 2, then update L=L+S, and execute step S1 again;
[0013] S5. Determine whether L – S ≥L min is true, if yes, update L=L-S and execute step S1, if not, first update S=S / 2, then update L=L-S, and execute step S1 again;
[0014] S6. Determine whether the termination test condition is met, if yes, output the preprocessed waveforms under different stimulation intensities as ABR stacked waveforms, and execute step S7, if not, update L=L+S, and execute step S1 again;
[0015] S7. Input the ABR stacked waveforms into a multi-intensity ABR sequence threshold discrimination model to output the ABR threshold.
[0016] In a second aspect, based on the method of the first aspect, the application provides an online adaptive ABR threshold automatic identification system, comprising:
[0017] a parameter setting module, configured to initialize system parameters; the system parameters include maximum stimulation intensity L max , minimum stimulation intensity L min , minimum step size S min , confidence threshold α, maximum number of repeated detections N max ; initialize step size S and stimulation intensity L;
[0018] a preprocessing module, configured to collect an ABR waveform signal under stimulation intensity L and perform preprocessing to obtain a preprocessed waveform;
[0019] a single-intensity detection module, configured to detect the preprocessed waveform according to a single-intensity ABR detection model to obtain a discrimination result and a confidence score;
[0020] an adaptive judgment module, configured to judge whether to terminate the single-intensity test according to an output of the single-intensity detection module, if yes, output an ABR stacked waveform, otherwise, adjust the stimulation intensity and call the preprocessing module again;
[0021] an ABR threshold recognition module, configured to process the ABR stacked waveform according to a multi-intensity ABR sequence threshold discrimination model to output an ABR threshold.
[0022] Advantages of the present application:
[0023] Real-time online: the present application can realize a clinical detection process of side-by-side collection and judgment, and can dynamically adjust the stimulation intensity according to the discrimination result and the confidence score of each time in the detection process, adaptively determine the termination condition, and effectively avoid the lag of the traditional method which needs post-batch analysis.
[0024] Trustworthy and robust: the present application adopts confidence score discrimination and repeated detection strategy, and simultaneously introduces minimum step size and boundary condition constraint, to ensure the reliability and consistency of the detection result, and reduce the misjudgment in low signal-to-noise ratio environment.
[0025] Efficient and time-saving: the present application adopts an adaptive detection mechanism of coarse search and fine search, first narrows down the threshold interval with a step size of 20 dB, and then finely searches with a step size of 10 dB, greatly reduces the test of invalid intensity points, shortens the measurement time, and improves the clinical work efficiency.
[0026] Accurate and objective: the present application uses a single-intensity ABR detection model to determine whether there is a V wave in the collection process in real time, to avoid excessive accumulation of invalid data; finally, the morphological features and latency information of the cross-intensity waveform are fused through a multi-intensity ABR sequence model, and the threshold is comprehensively output, taking into account the real-time performance and robustness, and significantly improving the diagnostic accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1It is a kind of ABR threshold automatic detection identification method framework schematic diagram for the embodiment of the application.
[0028] Figure 2 It is a kind of ABR threshold automatic detection identification method overall flow chart for the embodiment of the application.
[0029] Figure 3 It is a kind of single intensity ABR detection model architecture diagram for the embodiment of the application.
[0030] Figure 4 It is a kind of multi-intensity ABR sequence threshold discrimination model structure diagram for the embodiment of the application.
[0031] Figure 5 It is a kind of processing flow schematic diagram of multi-head DTW attention mechanism module for the embodiment of the application. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0033] The existing clinical ABR threshold identification technology has defects such as dependence on manual work, low efficiency, large subjective error and the like, and cannot realize online adaptive detection. Therefore, the application provides an ABR threshold automatic detection identification method. Specifically, first, a single intensity ABR detection model is used to automatically judge whether an ABR waveform has a response, and online detection combining coarse search and fine search is realized by adaptively adjusting the stimulation intensity; then, ABR waveforms collected under multiple stimulation intensities are input into a multi-intensity ABR sequence threshold discrimination model to determine the ABR threshold. The method minimizes manual intervention, and can quickly and accurately determine the ABR hearing threshold of a subject. Compared with the prior art which can only perform offline analysis, the application can guide the selection of stimulation intensity in real time and output the final threshold, greatly improving the detection efficiency and objectivity.
[0034] Figure 1 It is a kind of online adaptive ABR threshold automatic identification method framework schematic diagram according to some embodiments of the application. Figure 2 It is a kind of online adaptive ABR threshold automatic identification method overall flow chart according to some embodiments of the application.
[0035] Some embodiments of the application provide an online adaptive ABR threshold automatic identification method, as shown in Figure 1 , Figure 2
[0036] S0. Set the maximum stimulation intensity L max , the minimum stimulation intensity L min , the minimum step size S min , the confidence threshold α, the maximum number of repeated detections N max ; initialize the step size S and the stimulation intensity L;
[0037] S1. Collect the ABR waveform signal at the stimulation intensity L and pre-process it to obtain a pre-processed waveform;
[0038] S2. Input the pre-processed waveform into the single-intensity ABR detection model to output a discrimination result and a confidence score, the discrimination result is used to display whether there is a V wave; determine whether the confidence score is not less than the confidence threshold α, if yes, execute step S4, if not, execute step S3;
[0039] S3. Determine whether the maximum number of repeated detections N max has been reached, if yes, use the result of the latest detection to execute step S4, if not, execute step S1;
[0040] S4. Determine whether S=S min is true, if yes, execute step S6; if not, check whether the discrimination result of the current pre-processed waveform shows the presence of a V wave, if yes, execute step S5, if not, first update S=S / 2, then update L=L+S, and execute step S1 again;
[0041] S5. Determine whether L – S ≥ L min is true, if yes, update L=L-S and execute step S1, if not, first update S=S / 2, then update L=L-S, and execute step S1 again;
[0042] S6. Determine whether the termination test condition is met, if yes, output the pre-processed waveforms at different stimulation intensities as ABR stacked waveforms, and execute step S7, if not, update L=L+S, and execute step S1 again;
[0043] S7. Input the ABR stacked waveforms into the multi-intensity ABR sequence threshold discrimination model to output the ABR threshold.
[0044] In some embodiments, the ear of the subject is stimulated with sound and ABR waveform signals are collected in an electrically shielded room using a conventional clinical evoked potential test system. The stimulation parameters can be set as: 100 µs click sound or tone burst sound as the stimulation type, alternating polarity as the stimulation polarity, 20.1 times / s, 31.5 times / s or 44.4 times / s as the stimulation rate. In addition, a four-electrode configuration is used during the test, i.e., a recording electrode is set at Cz (vertex), a ground electrode is set at the midpoint of the forehead, and reference electrodes are set on both mastoids.
[0045] Specifically, the conventional clinical evoked potential test system can be, for example, Eclipse evoked potential test system, Interacoustics Inc., Denmark, etc.
[0046] In some embodiments, the ABR waveform signals collected at a predetermined stimulation intensity are pre-processed, including: band-pass filtering and resampling the ABR waveform signals.
[0047] Specifically, the band-pass filtering range is adjustable within 0.033-3 kHz, and the sampling rate of resampling can be set to 30 kHz.
[0048] In some embodiments, the output of the pre-processed waveform after being input into the single-intensity ABR detection model can be divided into the following three cases:
[0049] 1) If the discrimination result shows that there is a V wave, and the confidence score ≥ confidence threshold α, it is determined that there is an effective ABR response at the current stimulation intensity;
[0050] 2) If the discrimination result shows that there is no V wave, and the confidence score ≥ confidence threshold α, it is determined that there is no effective ABR response at the current stimulation intensity;
[0051] 3) If the confidence score < confidence threshold α, it is determined that the detection result is uncertain, and repeated detection at the same stimulation intensity is needed.
[0052] In some embodiments, when the output of the single-intensity ABR detection model belongs to the first case, adaptive coarse search is performed:
[0053] The stimulation intensity L is reduced by one step S, i.e., L is updated to L-S, and then steps S1-S2 are re-executed with the updated stimulation intensity.
[0054] If the output of the single-intensity ABR detection model changes (i.e. no valid ABR response is detected, a transition from "V-wave present" to "V-wave absent" is achieved), it indicates that the ABR threshold has been crossed, which is located between the "last V-wave present stimulus intensity" and the "first V-wave absent stimulus intensity". The stimulus intensity at this time is recorded, and then the adaptive fine search second stage is performed.
[0055] If the output of the single-intensity ABR detection model still belongs to the first kind, but the current stimulus intensity has reached the minimum allowed measurement intensity (i.e. L-S < L min ), the adaptive fine search first stage is performed.
[0056] In some embodiments, when the output of the single-intensity ABR detection model belongs to the second kind, the adaptive fine search second stage is directly performed.
[0057] In some embodiments, as shown in Figure 2 , the adaptive fine search is divided into two stages:
[0058] In the first stage, the output of the single-intensity ABR detection model still determines the presence of V-wave, but the current stimulus intensity has reached the minimum allowed measurement intensity (i.e. L-S < L min ), at this time the step size is halved, i.e. S = S / 2, and then the stimulus intensity L is reduced by one step S, i.e. L = L-S, and then the step S1 and subsequent operations are re-executed with the updated stimulus intensity.
[0059] In the second stage, the output of the single-intensity ABR detection model determines the absence of V-wave, and the step size is directly halved, i.e. S = S / 2, and then the stimulus intensity L is updated as L = L+S with the updated step size, so as to approach the boundary between the "last V-wave present stimulus intensity" and the "first V-wave absent stimulus intensity", and then the step S1 and subsequent operations are re-executed with the updated stimulus intensity.
[0060] In some embodiments, the termination test conditions include:
[0061] When L+S > L max , or L-S < L min , or the stimulus intensity corresponding to the value of L+S has been detected, the test is terminated.
[0062] In some embodiments, a one-dimensional residual neural network (Residual Network, ResNet) is used as the backbone architecture of the single-intensity ABR detection model, as shown in Figure 3As shown, it comprises 1x3 convolutional layers, maximum pooling layers, first residual blocks, second residual blocks, average pooling layers, and fully connected layers in sequence; the first residual blocks and the second residual blocks have the same structure and each comprises two 1x3 convolutional layers. The input waveform is subjected to 1x3 convolutional layers to extract basic features and subjected to maximum pooling, the maximum pooling result is subjected to deep feature extraction through multiple residual blocks, and the deep features are compressed through global average pooling and output binary classification logits through the fully connected layer.
[0063] Specifically, the output channel number of the initial 1x3 convolutional layer is 64, and the output channel numbers of the first residual block and the second residual block are 128 and 256 respectively.
[0064] In some embodiments, to ensure the reliability of the detection result in the clinical process, the application converts the binary classification logits output by the model into confidence, introduces a Softmax function after the fully connected layer, converts the binary classification logits into probability values, and respectively corresponds to the confidence of "V wave" and "no V wave".
[0065] In some embodiments, to avoid overconfidence of the model or deviation of the confidence distribution, when introducing the Softmax function after the fully connected layer, a temperature coefficient T>0 is further introduced, and the probability calculation formula after temperature scaling is:
[0066]
[0067] wherein, p i represents the probability value of the class i, z i represents the logit corresponding to the class i in the output of the fully connected layer. By adjusting T, the confidence distribution can be balanced, and the robustness of the model under different intensity and noise conditions can be improved.
[0068] In some embodiments, the backbone structure of the single-intensity ABR detection model can also use CNN, Transformer, CNN+LSTM, MobileNet-1D, and other lightweight networks.
[0069] In some embodiments, the application proposes a multi-intensity ABR sequence threshold discrimination model combining dynamic time warping (DTW) and series-temporal Transformer. The processing process of the multi-intensity ABR sequence threshold discrimination model is as shown in Figure 3 As shown, it comprises:
[0070] The ABR stacked waveform is subjected to DTW similarity perception sequence Transformer to obtain a DTW processed sequence;
[0071] To better perform timing modeling, embodiments of the present application introduce absolute position encoding of stimulus intensity, that is, the sequence of stimulus intensity corresponding to the ABR stacked waveform is passed through an embedding layer, and then the DTW processed sequence is added to the output of the embedding layer and input into the hierarchical multi-scale time Transformer to obtain multi-scale time features.
[0072] The multi-scale time features are passed through a max-pooling layer and a classifier to obtain a discrimination result.
[0073] In some embodiments, as shown in Figure 4 The DTW similarity-aware sequence Transformer includes 3 residual connection series encoders, the ABR stacked waveform is passed through the first series encoder to obtain a first output sequence, the first output sequence is added to the ABR stacked waveform and input into the second series encoder to obtain a second output sequence, and the second output sequence is added to the ABR stacked waveform and input into the third series encoder to obtain the DTW processed sequence.
[0074] Specifically, the processing process of each series encoder includes:
[0075] S11. Calculate the query matrix Q, the value matrix V, and the key matrix K according to the input sequence; the calculation formula is as follows:
[0076]
[0077]
[0078]
[0079] where x i represents the i-th sequence in the input sequence, q i , k i , v i represents the query vector, the key vector, and the value vector corresponding to x i , W q , W k , W v represents a trainable linear projection matrix, d k represents the encoding dimension.
[0080] S12. Perform stimulus level-informed RoPE (Stimulus level-informed RoPE) on the query matrix Q and the key matrix K to embed the stimulus intensity information into the feature space to ensure the semantic correspondence between the waveform and the intensity, to obtain the query matrix Q' and the key matrix K'; the stimulus level-informed RoPE is represented as follows:
[0081]
[0082]
[0083]
[0084] in, x represents i The corresponding query vector encoded by the horizontal rotation position of the stimulus. x represents j The corresponding key vector, f, encoded by the horizontal rotation position of the stimulus. q ( ) and f k ( ) represent the rotation position encoding functions for the query vector and the key vector, respectively, m i x represents i The corresponding stimulus level encoding value, This indicates that predefined parameters are included. The rotation position matrix, with predefined parameters. θ i This represents the rotation frequency parameter of the i-th dimension pair.
[0085] Specifically, the stimulus level encoding values are shown in Table 1:
[0086] Threshold category Stimulus level (dB nHL) Encoding value NT / 10 100 100 9 90 90 8 80 80 7 70 70 6 60 60 5 50 50 4 40 40 3 30 30 2 20 20 1 10 10 0 / 0 -1 / Zero padding 11
[0087] S13. Input the query matrix Q', key matrix K', and value matrix V into the Multi-Head DTW Attention module to obtain the attention features;
[0088] Unlike traditional self-attention mechanisms, multi-head DTW attention mechanisms can perform waveform alignment and context dependency modeling, maintaining alignment even when waveforms exhibit latency shifts. Each head DTW attention mechanism can be represented as:
[0089]
[0090] Where, Q'= K'= N represents the number of waveforms in the ABR stacked waveform; DTWscores represents the DTW similarity matrix of the similarity between waveforms across stimulus intensities. It is mainly used to represent the waveform association across stimulus intensities and reflects the overall morphological similarity between ABR waveforms in a single ear. It is obtained by calculating the DTW similarity between every two waveforms in the ABR stacked waveform. This represents the Hadama product.
[0091] Query vector and key vector The inner product between them is:
[0092]
[0093]
[0094] As Figure 5 shown, the multi-head DTW attention mechanism can be represented as:
[0095]
[0096] wherein MHDTWA(Q’,K’,V) represents the output of the multi-head DTW attention mechanism module, Concat() represents a concatenation operation, head h represents the h-th=1,2,…,H S attention head, H S represents the number of attention heads, W o represents an output projection matrix, and A represents the number of sampling points of a pre-processed waveform in the ABR stacked waveform.
[0097] S14. After connecting the attention features with the input sequence residual and normalizing, the normalized result is fed into a feed-forward network (FFN), and then connected with the normalized result residual and normalized to obtain the output sequence , N is the number of waveforms in the multi-intensity ABR sequence.
[0098] In some embodiments, the hierarchical multi-scale time Transformer adopts a hierarchical design, which is composed of three sequential encoders in series, aiming to capture multi-scale time information. Each sequential encoder includes a one-dimensional convolution layer, a normalization layer, a position encoding module, a multi-head temporal reduction attention (MHTRA) module, an FFN, and an Add & Norm layer.
[0099] The processing process of the l-th=1,2,3 sequential encoder includes:
[0100] S21. The input is processed by a one-dimensional convolution layer and a normalization layer to obtain a convolution output. The output obtained by processing the input by a position encoding module is added to the convolution output to obtain an addition result, which is represented as:
[0101]
[0102] wherein C (l) represents the addition result obtained by adding the convolution output and the position encoding module output in the l-th sequential encoder, G l( ) represents the normalization layer in the l-th sequential encoder, Conv(l) represents a one-dimensional convolution layer in the first time sequence encoder, the kernel size of which is k (l) , the filter number is f (l) , and the step size is s (l) ; Z l-1 represents the input of the first time sequence encoder, represents the position information in the vector for maintaining the time sequence.
[0103] S22. Calculate the query matrix, the key matrix and the value matrix according to the addition result, input the three matrices into the multi-head time sequence dimension reduction attention module, and obtain the time sequence attention feature.
[0104] Specifically, in the multi-head time sequence dimension reduction attention module, in order to reduce the storage cost, the time dimension of the key and value vectors is first compressed by one-dimensional convolution and layer normalization, and the operation process can be represented as:
[0105]
[0106] Wherein, TR(x) represents time sequence dimension reduction on the input vector x, the input vector x is the query vector or the key vector, R (l) represents the time sequence dimension reduction ratio factor, Reshape( ) represents the vector shape remodeling, and LayerNorm( ) represents the layer normalization.
[0107] Subsequently, the multi-head attention mechanism is used to process the query matrix, the dimension-reduced key matrix and the value matrix, which are input, to model the long-range dependency relationship on the time feature, and the output formula is:
[0108]
[0109]
[0110] Wherein, is a linear projection matrix, represents the h1=1,2,…,H T th attention head in the first time sequence encoder, H T is the number of attention heads, and MHTRA(Q (l) ,K (l) ,V (l) ) represents the output of the multi-head time sequence dimension reduction attention module, W o1 is the output projection matrix. Attention( ) represents the attention mechanism, Q (l) , K (l) , V (l) represent the query matrix, the key matrix and the value matrix of the first time sequence encoder.
[0111] S23. The time sequence attention feature is connected with the addition result residual error and then normalized. The normalized result is connected with the normalized result residual error after passing through the feedforward neural network, and then normalized to obtain the output.
[0112] In some embodiments, nine-fold cross-validation is used for hyperparameter optimization when training the single-intensity ABR detection model and the multi-intensity ABR sequence threshold discrimination model. After the hyperparameters are determined, the optimal hyperparameter configuration is used to train the final model on the entire training set for testing. In the model optimization process, the cross-entropy loss function is used, and the Adam optimizer is selected. In the parameter update process, a multi-step learning rate decay strategy is combined, and the learning rate is gradually reduced according to the number of iterations to achieve more robust convergence. In the testing stage, the test set data is preprocessed and input into the trained model, and the recognition result is output. The evaluation index of the single-intensity ABR detection model uses the prediction accuracy; the evaluation index of the multi-intensity ABR sequence threshold discrimination model includes the mean absolute error (MAE), the prediction accuracy, the accuracy within the 10 dB nHL error allowed range, and the overall length of the online detection process.
[0113] The training data set includes a main data set and a cross-center external data set, wherein:
[0114] The main data set contains 8350 subjects, including 499 normal hearing subjects, 4612 hearing loss patients, and 3239 subjects whose hearing status is unknown. The subjects are aged 0-95 years, covering different populations: 1118 infants (0-6 months), 3961 children (6 months-18 years), 2409 adults (18-60 years), and 862 elderly people (over 60 years old). This data set covers samples of different age groups and different hearing levels, ensuring the generality and universality of the model.
[0115] The cross-center external data set contains a total of 136 subjects, including 9 normal hearing subjects, 67 hearing impaired subjects, and 60 hearing unknown subjects. The age range is 1-73 years, covering 14 infants, 85 children, 30 adults, and 7 elderly people, which is used to verify the cross-center robustness of the model.
[0116] When the data set is divided, the "subject" is taken as the minimum unit to ensure that all waveforms of the same subject are completely allocated to the training set or the test set, avoiding information leakage. The main data set is divided into training set: test set = 9: 1, about 90% of the data is used for model training, and 10% of the data is used for testing. The cross-center data set is used as an independent external test set.
[0117] In some embodiments, the maximum stimulation intensity L ma x=100 dB nHL, the minimum stimulation intensity L min= 10 dB nHL, minimum step size S min = 10 dB nHL, confidence threshold a = 0.7, maximum number of repeated detections N max = 3; initialize step size S = 20 dB nHL, stimulus intensity L = 80 dB nHL. Wherein, the threshold a can be adaptively learned.
[0118] Exemplarily, the process of detecting a certain subject under this parameter is as follows:
[0119] 1) The first ABR waveform signal acquisition is performed at L = 80 dB nHL, the preprocessed waveform is input into the single-intensity ABR detection model, the discrimination result shows that there is V wave, and the confidence score is 0.92 (i.e. the model output probability value); the coarse search condition is met, and L is updated to 60 dB nHL;
[0120] 2) The second ABR waveform signal acquisition is performed at L = 60 dB nHL, the preprocessed waveform is input into the single-intensity ABR detection model, the discrimination result shows that there is V wave, and the confidence score is 0.88; the coarse search condition is met, and L is updated to 40 dB nHL;
[0121] 3) The third ABR waveform signal acquisition is performed at L = 40 dB nHL, the preprocessed waveform is input into the single-intensity ABR detection model, the discrimination result shows that there is V wave, and the confidence score is 0.81; the coarse search condition is met, and L is updated to 20 dB nHL;
[0122] 4) The fourth ABR waveform signal acquisition is performed at L = 20 dB nHL, the preprocessed waveform is input into the single-intensity ABR detection model, the discrimination result shows that there is no V wave, and the confidence score is 0.85; the coarse search condition is not met, the adaptive fine search second stage is executed, S is updated to 10 dB nHL, and L is updated to 30 dB nHL;
[0123] 5) The fifth ABR waveform signal acquisition is performed at L = 30 dB nHL, the preprocessed waveform is input into the single-intensity ABR detection model, the discrimination result shows that there is V wave, and the confidence score is 0.78; this indicates that the ABR threshold is between 20 dB nHL and 30 dB nHL. At this time, the condition that the current step size = minimum step size is met, whether the termination condition is met is judged, since the preprocessed waveform of the stimulus intensity L + S = 30 + 10 = 40 dB nHL has been detected, the termination condition is met; the preprocessed waveforms collected at the stimulus intensities of 80, 60, 40, 30 and 20 are stacked together to form an ABR stacked waveform;
[0124] 6) The ABR stacked waveform input multi-intensity ABR sequence threshold discrimination model outputs the ABR threshold value as 30 dB nHL.
[0125] In some embodiments, the step size S and the minimum step size S min A combination of 20 dB nHL and 10 dB nHL can be used, or a combination of 15 dB nHL and 5 dB nHL, or a combination of 10 dB nHL and 5 dB nHL, etc.
[0126] In some embodiments, when the confidence score is less than the threshold value a for repeated detection, if the confidence score is still less than the confidence threshold value a after reaching the maximum number of repeated detections, a lower weight is given to the sample when inputting the multi-intensity ABR sequence threshold discrimination model.
[0127] In some embodiments, if the transition from "V wave present" to "V wave absent" or the transition from "V wave absent" to "V wave present" has not been achieved when the termination test condition is met, a "recommend retest / other test" prompt is given. Specifically, for example, the discrimination result of the preprocessed waveform under the condition of L = 10 dB nHL still shows that the V wave exists, at this time L - S min <L min , the detection result of the ABR threshold value ≤ 10 dB nHL is directly output, and a "recommend retest / other test" prompt is given. Also for example, the discrimination result of the preprocessed waveform under the condition of L = 100 dB nHL still shows that the V wave does not exist, at this time L + S min >L max , the detection result of the ABR threshold value ≥ 100 dB nHL is directly output, and a "recommend retest / other test" prompt is given.
[0128] In some embodiments, the present application provides an online adaptive ABR threshold automatic identification system, comprising:
[0129] A parameter setting module for initializing system parameters; the system parameters include the maximum stimulation intensity L max , the minimum stimulation intensity L min , the minimum step size S min , the confidence threshold value a, the maximum number of repeated detections N max ; the step size S and the stimulation intensity L are initialized.
[0130] A preprocessing module for acquiring ABR waveform signals at the stimulation intensity L and preprocessing to obtain preprocessed waveforms;
[0131] A single-intensity detection module for detecting the preprocessed waveforms according to the single-intensity ABR detection model to obtain a discrimination result and a confidence score;
[0132] An adaptive judgment module is configured to judge whether to terminate the single-intensity test according to the output of the single-intensity detection module, and if yes, output the ABR stacked waveform, otherwise, adjust the stimulation intensity and call the preprocessing module again.
[0133] An ABR threshold identification module is configured to process the ABR stacked waveform according to a multi-intensity ABR sequence threshold value discrimination model, and output an ABR threshold value.
[0134] The method of the present application is highly consistent with the current ABR detection process in the clinic, the model used is a lightweight structure, which supports rapid deployment on conventional hardware devices (such as ARM processors, embedded platforms), and has good application feasibility.
[0135] In the present application, unless otherwise explicitly specified and limited, the terms "mounting", "setting", "connecting", "fixing", "rotating" and the like should be understood in a broad sense, for example, can be fixedly connected, or can be detachably connected, or can be integrated; can be mechanically connected, or can be electrically connected; can be directly connected, or can be indirectly connected through an intermediate medium; can be the internal communication of two elements or the interaction relationship between two elements, unless otherwise explicitly limited, the above-mentioned terms in the present application can be understood according to the specific meaning in the specific situation by the ordinary skilled in the art.
[0136] Although the embodiments of the present application have been shown and described, it can be understood by those skilled in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application, the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. An online adaptive ABR threshold automatic identification method, characterized in that, include: S0. Set the maximum stimulation intensity L max Minimum stimulus intensity L min Minimum step size S min Confidence threshold α, maximum number of repeated tests N max Initialize step size S and stimulus intensity L; S1. Acquire ABR waveform signals at stimulus intensity L and preprocess them to obtain preprocessed waveforms; S2. Input the preprocessed waveform into the single-intensity ABR detection model to output the discrimination result and confidence score. The discrimination result shows whether V wave exists. Determine whether the confidence score is not less than the confidence threshold α. If yes, proceed to step S4. If not, proceed to step S3. S3. Determine if the maximum number of repeated checks N has been reached. max If yes, then use the result of the most recent detection to execute step S4; otherwise, execute step S1. S4. Determine if S=S min If the condition is met, proceed to step S6; otherwise, check if the current preprocessed waveform discrimination result shows the presence of a V wave. If it is present, proceed to step S5; otherwise, first update S=S / 2, then update L=L+S, and proceed to step S1 again. S5. Determine if L – S ≥ L min If the condition is true, update L=LS and execute step S1; otherwise, first update S=S / 2, then update L=LS, and execute step S1 again. S6. Determine whether the test termination condition is met. If yes, output the preprocessed waveforms under different stimulus intensities as ABR stacked waveforms and execute step S7. If not, update L=L+S and execute step S1 again. S7. Input the ABR stacked waveform into the multi-intensity ABR sequence threshold discrimination model and output the ABR threshold.
2. The online adaptive ABR threshold automatic identification method according to claim 1, characterized in that, Set the maximum stimulation intensity L max =100dB nHL, minimum stimulus intensity L min =10dB nHL, minimum step size S min =10dB nHL, confidence threshold α=0.7, maximum number of repeated tests N max =3; Initial step size S=20dB nHL, stimulus intensity L=80dB nHL.
3. The online adaptive ABR threshold automatic identification method according to claim 1, characterized in that, The test is terminated when L+S min >L max , or LS min <L min If the value of L+S has already been checked, terminate the test.
4. The online adaptive ABR threshold automatic identification method according to claim 1, characterized in that, The single-intensity ABR detection model consists of a cascaded 1×3 convolutional layer, a max pooling layer, a first residual block, a second residual block, an average pooling layer, and a fully connected layer. The first and second residual blocks have the same structure, each consisting of two 1×3 convolutional layers.
5. The online adaptive ABR threshold automatic identification method according to claim 4, characterized in that, A temperature coefficient T and a softmax function are introduced after the fully connected layer to convert the output of the fully connected layer into probability values, expressed as: Where, p i z represents the probability value of category i. i This represents the logit of category i in the output of the fully connected layer.
6. The online adaptive ABR threshold automatic identification method according to claim 1, characterized in that, The processing steps of ABR stacked waveforms in the multi-intensity ABR sequence threshold discrimination model include: The ABR stacked waveform is processed by the DTW similarity-aware sequence Transformer to obtain the DTW processed sequence; the DTW similarity-aware sequence Transformer includes three residual-connected sequence encoders. The stimulus intensity sequence corresponding to the ABR stacked waveform is passed through an embedding layer. The DTW processed sequence is added to the output of the embedding layer and then input into a hierarchical multi-scale time transformer to obtain multi-scale time features. The hierarchical multi-scale time transformer includes three cascaded time encoders. The multi-scale temporal features are passed through a max pooling layer and a classifier to obtain the discrimination result.
7. The online adaptive ABR threshold automatic identification method according to claim 6, characterized in that, The processing steps for each sequence encoder include: S11. Calculate the query matrix Q, value matrix V, and key matrix K based on the input sequence; S12. Encode the stimulus horizontal rotation position on the query matrix Q and the key matrix K to obtain the query matrix Q' and the key matrix K'; S13. Input the query matrix Q', key matrix K', and value matrix V into the multi-head DTW attention mechanism module to obtain attention features; the processing of the multi-head DTW attention mechanism module is expressed as follows: In the formula, MHDTWA(Q',K',V) represents the output of the multi-head DTW attention mechanism module, Concat() represents the concatenation operation, and head h Represents the h-th digits as h=1,2,…,H S One point of attention, H S W represents the number of heads of attention. o This represents the output projection matrix, and DTWscores represents the waveform similarity across stimulus intensities. Represents the Hadama product; S14. After concatenating the attention features with the input sequence residuals, normalize them. Then, pass the normalized result through a feedforward neural network and concatenate it with the normalized result residuals again to obtain the output sequence.
8. The online adaptive ABR threshold automatic identification method according to claim 6, characterized in that, The processing steps for the l=1, 2, and 3rd timing encoders include: S21. Pass the input through a one-dimensional convolutional layer and a normalization layer to obtain the convolutional output. Add the output obtained by the position encoding module to the convolutional output to obtain the sum. S22. Calculate the query matrix, key matrix, and value matrix based on the summation result, and input the three matrices into the multi-head temporal dimensionality reduction attention module to obtain the temporal attention features; S23. After concatenating the temporal attention features with the summation result residual, normalize the result. Then, pass the normalized result through a feedforward neural network and concatenate it with the normalized result residual, and normalize it again to obtain the output.
9. The online adaptive ABR threshold automatic identification method according to claim 8, characterized in that, The processing steps of the multi-head temporal dimensionality reduction attention module include: In the formula, MHTRA(Q (l) ,K (l) V (l) The ) represents the output of the multi-head temporal dimensionality reduction attention module, and Concat() represents the concatenation operation. This represents the h1=1,2,…,H sequence encoder. T One point of attention, H T W represents the number of heads of attention. o1 The output projection matrix is represented by Q, Attention() represents the attention mechanism, and Q represents the Q-matrix. (l) K (l) V (l) Represents the query matrix, key matrix, and value matrix of the l-th time encoder; Let R denote the linear projection matrix, TR(x) represent temporal dimensionality reduction of the input vector x, where the input vector x is a query vector or a key vector. (l) Represents the temporal dimensionality reduction scaling factor, Reshape() represents vector shape reshaping, and LayerNorm() represents layer normalization.
10. An online adaptive ABR threshold automatic identification system, characterized in that, The online adaptive ABR threshold automatic identification method as described in any one of claims 1-9 includes: The parameter setting module is used to initialize system parameters; the system parameters include the maximum stimulus intensity L. max Minimum stimulus intensity L min Minimum step size S min Confidence threshold α, maximum number of repeated tests N max Initialize step size S and stimulus intensity L; The preprocessing module is used to acquire ABR waveform signals at stimulus intensity L and perform preprocessing to obtain preprocessed waveforms; The single intensity detection module is used to detect the preprocessed waveform according to the single intensity ABR detection model to obtain the discrimination result and confidence score; The adaptive judgment module is used to determine whether to terminate the single intensity test based on the output of the single intensity detection module. If it is terminated, it outputs the ABR stacked waveform; otherwise, it adjusts the stimulus intensity and calls the preprocessing module again. The ABR threshold recognition module is used to process ABR stacked waveforms according to the multi-intensity ABR sequence threshold discrimination model and output the ABR threshold.