An aural auxiliary evaluation method based on a brain-computer interface
Patent Information
- Application Number
- CN202610928865.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-15
Smart Images

Figure CN122744784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brain-computer interface technology, specifically to a hearing-assisted assessment method based on a brain-computer interface. Background Technology
[0002] Objective assessment of auditory processing ability is a core requirement of audiological diagnosis and treatment. It is widely used in individualized fitting of hearing aids and cochlear implants, hearing screening of children and people who cannot cooperate with behavioral audiometry, and assessment of speech comprehension ability in complex scenarios. Brain-computer interface-based auditory assessment, which collects multi-channel EEG signals during the subject's listening process and objectively infers the level of speech intelligibility under the condition of natural speech as stimulation, has become one of the development directions in this field.
[0003] The current mainstream approach is to extract the temporal envelope from speech, establish a linear mapping between it and multi-channel EEG signals through a time response function, and then use the correlation coefficient between the reconstructed envelope and the original envelope as the output for intelligibility assessment.
[0004] However, in the medium signal-to-noise ratio range, the intelligibility score of the same subject shows a non-monotonic jump relative to the signal-to-noise ratio. For example, the assessment results of the same subject at different signal-to-noise ratio levels sometimes show an anomalous ranking where the high signal-to-noise ratio score is lower than the low signal-to-noise ratio score; the assessment results are inconsistent under different noise types, and the assessment results of different subjects show unexplained individual differences.
[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a hearing-assisted assessment method based on a brain-computer interface.
[0007] To achieve the above objectives, the technical solution of the present invention is as follows:
[0008] In a first aspect, the present invention discloses a hearing-assisted assessment method based on a brain-computer interface, comprising the following steps:
[0009] Acquire multi-channel EEG signals and synchronized speech stimulation signals during the subject's listening process;
[0010] Envelope feature sequence and fine structure feature sequence are extracted from the speech stimulus signal. The envelope feature sequence reflects the low-frequency modulation component of the speech envelope, and the fine structure feature sequence is composed of the instantaneous phase of multiple sub-bands. The upper frequency limit of the envelope feature sequence is lower than the lower frequency limit of any sub-band in the fine structure feature sequence.
[0011] The envelope reconstruction sequence corresponding to the envelope feature sequence is reconstructed from the multi-channel EEG signal using the first time window length, and the fine structure reconstruction sequence corresponding to the fine structure feature sequence is reconstructed from the multi-channel EEG signal using the second time window length. The fine structure reconstruction sequence is composed of the instantaneous phase estimates corresponding to multiple sub-bands. The first time window length is greater than the second time window length.
[0012] Based on the envelope reconstruction sequence and the fine structure reconstruction sequence, the instantaneous phase difference sequence between the two is calculated. Within the preset evaluation time window, the time rate of change of the phase distribution concentration of the instantaneous phase difference sequence is determined as the phase difference diffusion slope.
[0013] Based on the phase difference diffusion slope, the envelope path weight and fine structure path weight are determined. The envelope path weight is positively correlated with the phase difference diffusion slope, while the fine structure path weight is negatively correlated with the phase difference diffusion slope.
[0014] The envelope reconstruction score is determined based on the correlation between the envelope reconstruction sequence and the envelope feature sequence, and the fine structure reconstruction score is determined based on the phase consistency between the fine structure reconstruction sequence and the fine structure feature sequence.
[0015] The envelope reconstruction score and the fine structure reconstruction score are weighted and summed using the envelope pathway weight and the fine structure pathway weight to obtain the auditory intelligibility score and output it.
[0016] Secondly, this invention discloses a hearing-assisted assessment system based on a brain-computer interface, comprising:
[0017] The signal acquisition module is used to acquire multi-channel EEG signals and synchronous speech stimulation signals during the subject's listening process;
[0018] The feature extraction module is used to extract envelope feature sequences and fine structure feature sequences from speech stimulus signals. The envelope feature sequence reflects the low-frequency modulation components of the speech envelope, and the fine structure feature sequence is composed of the instantaneous phases of multiple sub-bands. The upper frequency limit of the envelope feature sequence is lower than the lower frequency limit of any sub-band in the fine structure feature sequence.
[0019] The signal reconstruction module is used to reconstruct the envelope reconstruction sequence corresponding to the envelope feature sequence from the multi-channel EEG signal using a first time window length, and to reconstruct the fine structure reconstruction sequence corresponding to the fine structure feature sequence from the multi-channel EEG signal using a second time window length. The fine structure reconstruction sequence consists of the instantaneous phase estimates corresponding to multiple sub-bands. The first time window length is greater than the second time window length.
[0020] The diffusion slope determination module is used to calculate the instantaneous phase difference sequence between the envelope reconstruction sequence and the fine structure reconstruction sequence. Within a preset evaluation time window, it determines the rate of change of the phase distribution concentration of the instantaneous phase difference sequence as the phase difference diffusion slope.
[0021] The weight determination module is used to determine the envelope path weight and the fine structure path weight based on the phase difference diffusion slope. The envelope path weight is positively correlated with the phase difference diffusion slope, while the fine structure path weight is negatively correlated with the phase difference diffusion slope.
[0022] The reconstruction score determination module is used to determine the envelope reconstruction score based on the correlation between the envelope reconstruction sequence and the envelope feature sequence, and to determine the fine structure reconstruction score based on the phase consistency between the fine structure reconstruction sequence and the fine structure feature sequence.
[0023] The scoring output module is used to perform a weighted summation of the envelope reconstruction score and the fine structure reconstruction score using the envelope path weight and the fine structure path weight, so as to obtain the auditory intelligibility score and output it.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] 1. Two independent and observable tracking pathways were established by simultaneously extracting envelope feature sequences and fine structure feature sequences from speech stimulus signals and intrinsically separating their upper frequency limit from the lower frequency limit of their subbands. The two tracking estimates were reconstructed independently from multi-channel EEG signals using the lengths of the first and second time windows, ensuring that each pathway was reliably reconstructed at its intrinsic frequency scale. The rate of change of the concentration of phase distribution of the instantaneous phase difference sequence between the two reconstruction results within the evaluation time window was used as the phase difference diffusion slope to dynamically allocate the weights of the envelope pathway and the fine structure pathway. This ensured that the final auditory intelligibility score could still provide monotonous, stable, and cross-individual consistent results when the two pathways switched in a coordinated state.
[0026] 2. This invention automatically triggers subband configuration updates when there is a severe imbalance in the decoding quality of two paths by setting a comparison mechanism between path misalignment degree and misalignment threshold. By reducing the bandwidth of subbands whose phase-locked loop values fall into a preset abnormal range or inserting neighboring subbands, the invention provides targeted refinement for the subbands with the most severe noise pollution, enabling fine-structured paths to recover some tracking capabilities without increasing hardware costs. At the same time, through a preset maximum re-extraction limit and an output mechanism for evaluating unstable alarm flags, the solution can still output a comprehensibility score with confidence annotations even in extreme cases. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is an overall block diagram of the method in Embodiment 1 of the present invention;
[0029] Figure 2 This is a timing diagram of the weight redistribution driven by the phase difference diffusion slope in Embodiment 1 of the present invention;
[0030] Figure 3 This is an overall block diagram of the system in Embodiment 2 of the present invention. Detailed Implementation
[0031] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Example 1:
[0033] like Figures 1-2 As shown, a brain-computer interface-based auditory-assisted assessment method includes the following steps:
[0034] To facilitate understanding, the following example illustrates the objective intelligibility assessment of an adult subject with sensorineural hearing loss after wearing a hearing aid. During the assessment, the subject sat upright in a soundproof room and listened to a 30-second continuous natural Chinese speech segment under free-field conditions via a calibrated speaker, with background noise from multiple conversations superimposed. The segment was divided into two periods: the signal-to-noise ratio (SNR) remained at +5 dB in the first half (0-15 seconds) and deteriorated to -5 dB in the second half (15-30 seconds). A 64-channel EEG acquisition cap was fitted to the subject's scalp according to the international 10-20 system, with electrode impedance controlled below 5 kΩ.
[0035] In step S1, multi-channel EEG signals and synchronous speech stimulation signals are acquired during the subject's listening process.
[0036] Specifically, the stimulus synthesis system plays speech stimulus signals. This signal is a single-channel time-series signal with a sampling rate of 16kHz, representing the acoustic input heard by the subject; multi-channel EEG signals Simultaneous acquisition using 64-channel scalp electrodes at a sampling rate of 2kHz characterizes the electrophysiological responses of the auditory cortex and related brain regions during auditory processing. To ensure millisecond-level alignment between the speech stimulus signal and the multi-channel EEG signal on the time axis, a standard experimental psychology triggering device (such as Cedrus or BIOPAC series triggers) is used to generate a synchronous trigger stream. The trigger pulses are connected to the event output port of the speech playback system and the event tag input port of the EEG amplifier, respectively.
[0037] After acquiring multichannel EEG signals, basic preprocessing is performed to obtain clean, strictly synchronized multichannel temporal signals from the raw scalp EEG data, which contains a large amount of noise and artifacts. Basic preprocessing includes the following steps:
[0038] Power frequency notch filtering: for multi-channel EEG signals A second-order IIR notch filter is used at 50Hz to filter out power frequency interference from the power grid.
[0039] Bandpass filtering: A zero-phase bidirectional Butterworth bandpass filter is used to filter the signal after power frequency notch filtering. The lower limit of the passband is 1Hz and the upper limit is 500Hz, so as to simultaneously cover the low frequency band corresponding to the envelope path and the high frequency band corresponding to the fine structure path.
[0040] Bad channel detection and interpolation: Bad channels are identified based on indicators such as channel variance and correlation with adjacent channels. In this embodiment, channels with "channel variance exceeding three times the standard deviation of the mean variance of all channels" or "average correlation coefficient with the eight nearest neighbor channels being less than 0.4" are identified as bad channels and reconstructed using spherical spline interpolation of adjacent channels.
[0041] Artifact Removal: Independent Component Analysis (ICCA) is performed on the bandpass-filtered multichannel EEG signal to decompose the multichannel signal into several independent source components. Then, an automatic classifier such as ICLabel is used to identify artifact source components corresponding to blinking, electromyography (EMG), electrocardiography (ECG), and poor electrode contact. In this embodiment, source components with an artifact probability greater than 0.8 (or a brain component probability less than 0.3) output by ICLabel are identified as artifact source components. These artifact source components are set to zero and then projected back into the multichannel space.
[0042] Reference electrode reset: The reference electrodes of the multi-channel EEG signal are reset to the average reference, that is, the average value of the instantaneous voltage of all channels is used as the new reference voltage.
[0043] Segmentation by stimulus event: Based on the event timestamps provided by the synchronized trigger stream, continuous multi-channel EEG signals are segmented. The segments are divided into segments anchored by the starting point of each stimulus event. For continuous natural speech segments in this embodiment, they are segmented according to sentence boundaries.
[0044] After the above preprocessing, a multi-channel EEG signal that is strictly synchronized with the speech stimulus signal and has had major artifacts removed is obtained. In subsequent steps, directly use This refers to the preprocessing result, through which the synchronous data foundation required for subsequent dual-path feature extraction and decoding is obtained.
[0045] In step S2, envelope feature sequence and fine structure feature sequence are extracted from the speech stimulus signal. The envelope feature sequence reflects the low-frequency modulation component of the speech envelope, and the fine structure feature sequence is composed of the instantaneous phase of multiple sub-bands. The upper frequency limit of the envelope feature sequence is lower than the lower frequency limit of any sub-band in the fine structure feature sequence.
[0046] Traditional methods extract only the temporal envelope from speech as a single feature stream, neglecting the parallel tracking pathways of the auditory cortex for the fine structure of speech. The key design of this method lies in simultaneously extracting two intrinsically separable feature streams on the frequency scale, corresponding to two independent tracking pathways in the auditory cortex: the envelope pathway tracks slow modulation components, representing the rhythm, intensity, prosody, and syllable hierarchy of speech; the fine structure pathway tracks fast oscillation components, representing detailed acoustic information such as pitch and formant fine structure. The advantage of this dual-path extraction is that, compared to traditional methods that only reflect the subject's ability to track slow modulation information, this method can simultaneously reflect the state of two independent tracking pathways, providing a physical basis for subsequent dynamic weight adjustments based on the coordinated state of the two pathways.
[0047] Specifically, envelope feature sequences and fine structure feature sequences are extracted from speech stimulus signals, including two parallel execution paths.
[0048] The first path is used to extract the envelope feature sequence: perform Hilbert transform on the speech stimulus signal to obtain the envelope analytic signal, take the magnitude of the envelope analytic signal as the instantaneous amplitude, perform low-pass filtering on the instantaneous amplitude to obtain the envelope feature sequence, and the upper passband of the low-pass filter is the upper frequency limit of the envelope feature sequence.
[0049] The form of the envelope analytic signal is ,in, Speech stimulation signal The corresponding envelope resolution signal; For Hilbert transform operators; The unit is the imaginary unit. The instantaneous amplitude is... This characterizes the instantaneous intensity of the speech waveform. The low-pass filter employs a zero-phase bidirectional Butterworth filter; in this embodiment, the upper passband limit is... Taking 40Hz, the resulting low-pass filtered time series is the envelope feature sequence. It represents the slow modulation information of speech. The value is based on the fact that the energy of the sentence rhythm, word rhythm, and syllable level rhythm of speech is mainly concentrated in the 0 to 30 Hz frequency band. Taking 40 Hz can cover the complete slow modulation information while leaving a certain transition band margin for subsequent processing.
[0050] The second path is used to extract the fine structure feature sequence: the speech stimulus signal is decomposed into multiple sub-band signals through a Gammatone filter bank or an equivalent rectangular bandwidth sub-band filter bank, and a Hilbert transform is performed on each sub-band signal to obtain the sub-band analytic signal. The argument of the sub-band analytic signal is taken as the instantaneous phase sequence of the sub-band, and the instantaneous phase sequences of each sub-band are aligned and summarized along a unified time axis to form a fine structure feature sequence.
[0051] It should be noted that, unlike the subband energy analysis commonly used in existing technologies, this scheme retains the instantaneous phase-locked information of each subband rather than the subband energy information. This is because the intrinsic property of fine structure pathway tracking in the cortex is phase-locked accuracy.
[0052] The Gammatone filter bank simulates the frequency analysis characteristics of the basilar membrane in the human ear. Specifically, it analyzes the speech stimulus signal... pass Several parallel bandpass filters, each with a center frequency of... Uniformly distributed along the ERB scale, bandwidth parameters Incrementing according to the ERB scale. In this embodiment With a frequency of 16, the sub-band center frequency covers the range of 80Hz to 8kHz. Among them, This represents the total number of fine-structured subbands; For the first The center frequency of each sub-band; For the first The bandwidth parameters of each sub-band. The output of each sub-band is denoted as... , of which Subband signals of each subband It represents the fast oscillation component of speech within this frequency band.
[0053] For each sub-band signal Perform the Hilbert transform to obtain the subband analytic signal. The argument of the sub-band analyzed signal is taken as the instantaneous phase timing of that sub-band, i.e., the first... Instantaneous phase timing signal of each sub-band ,in, The first in the speech stimulus signal Sub-band analysis signals corresponding to each sub-band signal; The first in the speech stimulus signal The instantaneous phase timing of each sub-band characterizes the instantaneous phase state of the rapid oscillation of that sub-band; For complex argument operations. Put all The instantaneous phase time series of each sub-band are aligned and summarized along a unified time axis to obtain the fine structure feature sequence. .
[0054] Furthermore, the envelope feature sequence and the fine structure feature sequence must satisfy the intrinsic separation condition on the frequency scale, namely:
[0055] ;
[0056] in, The frequency bandwidth occupied by the envelope feature sequence; The frequency bandwidth occupied by fine-structure feature sequences; This represents the upper limit of the frequency of the envelope feature sequence; , These are the first two numbers in the fine structure feature sequence. The lower and upper frequency limits of each sub-band; This represents the total number of fine-structure subbands. The core significance of this condition is that the upper frequency limit of the envelope feature sequence is much smaller than the lower frequency limit of any subband in the fine-structure feature sequence, and the frequency bandwidths occupied by the two feature streams are objectively separated and do not overlap. This intrinsic separation condition is the physical basis for the design of the differentiated time window in the subsequent step S3.
[0057] Following the aforementioned clinical assessment example, after performing this step on the 30-second speech segment, the envelope feature sequence of the single-channel temporal morphology is obtained. (Frequency range 0 to 40 Hz) and a fine-structure feature sequence consisting of 16 sub-band instantaneous phase timings. (Frequency range 80Hz to 8kHz).
[0058] In step S3, the envelope reconstruction sequence corresponding to the envelope feature sequence is reconstructed from the multi-channel EEG signal using the first time window length, and the fine structure reconstruction sequence corresponding to the fine structure feature sequence is reconstructed from the multi-channel EEG signal using the second time window length. The fine structure reconstruction sequence consists of the instantaneous phase estimates corresponding to multiple sub-bands, and the first time window length is greater than the second time window length.
[0059] The key design element of this step is that the decoding of the two paths is performed independently with significantly different time window lengths. The envelope path uses a coarse-grained long window, while the fine-structure path uses a fine-grained short window. This differentiated window length design directly stems from the intrinsic separation of the dual-path frequency scale established in step S2: the coarse-grained long window ensures that the slow modulation and slow rhythm information of the envelope path can be reliably estimated, while the fine-grained short window ensures that the fine-structure path can accommodate the complete fast phase-locked loop cycle. The purpose of the differentiated time window design is to avoid decoding distortion caused by traditional single time windows that are too short in the low-frequency band or too long in the high-frequency band, allowing the tracking estimation of each path to achieve optimal resolution at its intrinsic frequency scale.
[0060] Specifically, the length of the first time window is not less than the minimum signal period corresponding to the upper frequency limit of the envelope feature sequence, and the length of the second time window is not less than the minimum signal period corresponding to the lower frequency limit of the lowest subband in the fine structure feature sequence. The above intrinsic constraints can be expressed as:
[0061] ;
[0062] in, The length of the first time window; The length of the second time window; This represents the upper limit of the frequency of the envelope feature sequence; This represents the lower frequency limit of the lowest subband in the fine-structure feature sequence. The minimum signal period corresponding to the upper limit of the frequency of the envelope feature sequence; This represents the minimum signal period corresponding to the lowest subband frequency limit in the fine-structure feature sequence. Its core significance lies in the fact that the decoding time window must accommodate at least one complete corresponding frequency signal period; since the upper frequency limit of the envelope path is much lower than the lower frequency limit of the fine-structure path, therefore... and There are natural intrinsic magnitude differences between them. It must be much greater than Following the aforementioned clinical assessment example, due to Taking 40Hz as the minimum frequency and the lowest sub-band frequency lower limit as approximately 80Hz, the length of the first time window in this embodiment is... Take 500ms (much greater) (lower limit of ms), second time window length Take 20ms (greater than) (lower limit of ms).
[0063] The calculation process of the envelope reconstruction sequence includes: expanding the samples of the multi-channel EEG signal within the first time window to obtain the time delay feature vector; fitting the linear mapping from the time delay feature vector to the envelope feature sequence through ridge regression to obtain the envelope reconstruction mapping; and performing inverse estimation of the time delay feature vector based on the envelope reconstruction mapping to obtain the envelope reconstruction sequence.
[0064] Specifically, for each moment The length of the first time window in the past The multi-channel EEG signal samples within a time period are unfolded into a high-dimensional feature vector, the dimension of which is equal to the product of the number of channels and the number of samples within that window length. In this embodiment, that is... The dimension is 1000 (where 1000 = 500 ms × 2 kHz). Ridge regression fits this high-dimensional feature vector to the envelope feature sequence by minimizing the sum of the L2 norm of the reconstruction error and the L2 norm regularization term of the weights. The linear mapping has a closed-form solution as follows:
[0065] ;
[0066] in, The envelope reconstruction mapping is the linear mapping coefficient vector obtained by fitting from the time-delay feature vector to the envelope feature sequence. This is the training design matrix obtained by stacking time-delay feature vectors in time order; This is the envelope feature sequence value vector corresponding to the training time step; The ridge regression regularization coefficient; It is the identity matrix; This indicates the matrix transpose.
[0067] It should be noted that envelope reconstruction mapping Training is completed in the pre-calibration phase: a training set is constructed using the speech-EEG synchronization data collected in the pre-calibration phase, and leave-one-out cross-validation is performed in time blocks from a log-scale uniformly distributed set of candidate values. The value that minimizes the reconstruction error under leave-one-out cross-validation is selected as the minimum value. The final value, corresponding This refers to the envelope reconstruction mapping of the subject. The online evaluation phase directly uses the training data... Multichannel EEG signals used in assessment periods By back-estimating the time-delay feature vector at each time step, the envelope reconstruction sequence can be obtained. This characterizes the tracking status of the subject's auditory cortex for slow-modulated speech information.
[0068] The calculation process of the fine structure reconstruction sequence includes: decomposing the multi-channel EEG signal into multiple brain electronic band responses according to the same sub-band configuration as the fine structure feature sequence; performing Hilbert transform on each brain electronic band response within the second time window and taking the argument to obtain the instantaneous phase estimate of each sub-band; and aligning and summarizing the instantaneous phase estimates of each sub-band according to a unified time axis to form the fine structure reconstruction sequence.
[0069] Specifically, the preprocessed multichannel EEG signals Using the same Gammatone subband configuration as in step S2 (the same 16 subbands, the same center frequency and bandwidth parameters), multi-subband decomposition was performed to obtain the brain electron band response corresponding to each subband. , of which Individual brain electronic band response This is a multi-channel time-series matrix, characterizing the electrophysiological response of each electrode on the subject's scalp within this sub-band frequency range. The response for each brain electron band is also shown. Each channel undergoes a Hilbert transform to obtain the analytic signal of that channel in that subband, and the argument is taken to obtain the instantaneous phase timing of that channel in that subband. ,in For channel index, , Total number of channels (in this embodiment) ).
[0070] To summarize the instantaneous phase timing data of the 64 channels in each sub-band into a representative instantaneous phase estimate for that sub-band, this embodiment employs channel weighting. Weighted circular average method:
[0071] ;
[0072] in, For the first in the fine structure reconstruction sequence The instantaneous phase estimation of each subband characterizes the overall tracking state of the subject's auditory cortex for the rapid oscillation of that subband; For channel In the Channel weights on each sub-band; For channel In the Instantaneous phase timing on each sub-band; Total number of channels; The imaginary unit; This involves complex argument operations. Channel weights. During the pre-calibration stage, it is determined that: the channel is relative to the first... Individual with voice instant phase The phase-locked value is used as the weight of the channel, and the weights of all channels are normalized so that their sum is 1. The principle of this design is that the auditory cortex projection areas (such as temporal lobe electrodes T7 / T8 / TP7 / TP8, etc.) have higher phase-locked capabilities than other areas, and giving them higher weights can improve the reliability of representative instantaneous phase estimation.
[0073] Furthermore, within a sliding window of the second time window length, the phase-locked relationship between the two can be estimated by taking the circular mean modulus of the difference sequence between the instantaneous phase of the brain electronic band response and the instantaneous phase of the corresponding speech subband. The instantaneous phase estimates of each subband are then aligned and summarized along a unified time axis to obtain the fine structure reconstruction sequence. .
[0074] Following the aforementioned clinical assessment example, this step outputs the envelope reconstruction sequence. (Single-channel timing signal) and fine structure reconstruction sequence (A multi-subband timing matrix consisting of 16 subband instantaneous phase estimates).
[0075] In step S4, based on the envelope reconstruction sequence and the fine structure reconstruction sequence, the instantaneous phase difference sequence between the two is calculated. Within a preset evaluation time window, the time rate of change of the phase distribution concentration of the instantaneous phase difference sequence is determined as the phase difference diffusion slope.
[0076] The calculation process of the instantaneous phase difference sequence includes: resampling and aligning the envelope reconstruction sequence and the fine structure reconstruction sequence to a unified time axis; performing a Hilbert transform on the resampled and aligned envelope reconstruction sequence and taking the argument to obtain the instantaneous phase of the envelope pathway; calculating the time-averaged energy of the sub-band signal of each sub-band in the speech stimulus signal within the evaluation time window to obtain the sub-band energy of each sub-band; using the sub-band energy of each sub-band as the weight, performing a circular weighted average on the instantaneous phase estimates of each sub-band in the resampled and aligned fine structure reconstruction sequence to obtain the instantaneous phase of the fine structure pathway; and reducing the difference between the instantaneous phase of the envelope pathway and the instantaneous phase of the fine structure pathway to... The instantaneous phase difference sequence is obtained by dividing the interval.
[0077] Specifically, in step S3, the envelope pathway and the fine structure pathway use different window lengths. and Since the sampling grids of the two reconstruction results may not be completely consistent, the coarser one of the two time axes (usually the envelope path time axis) is used as the reference time axis, and the other path is resampled and aligned to the reference time axis using a linear interpolation method.
[0078] Reconstructed envelope sequence after resampling and alignment Performing the Hilbert transform yields its analytic signal, and the argument of this analytic signal is taken as the instantaneous phase of the envelope path. This represents the instantaneous phase state of the envelope pathway reconstruction result. The subband energy of each subband is calculated using the following formula:
[0079] ;
[0080] in, To evaluate the first time window The subband energy of each subband represents the energy level of the subband carrying speech information within the evaluation time window; The first in the speech stimulus signal Subband signals of each subband; To evaluate the length of the time window; This is the current assessment moment.
[0081] The subband energy of each subband As weights, the instantaneous phase estimation of each sub-band The instantaneous phase of the fine-structured pathway is obtained by performing a circular weighted average using the following formula:
[0082] ;
[0083] in, The instantaneous phase of the fine structure pathway represents the instantaneous phase state after the fine structure pathway reconstruction results are summarized. For the first Subband energy of each subband; For the first in the fine structure reconstruction sequence Instantaneous phase estimation of each sub-band; This represents the total number of sub-bands. The imaginary unit; To map the instantaneous phase of each sub-band onto a complex exponent on the unit circle; This is for complex argument calculations. The argument is calculated using a complex exponential mapping instead of directly taking the arithmetic mean of the angles because the phase variable... When there are loops and jumps in the interval, a direct arithmetic mean will produce an incorrect average result due to the jump points. However, the circular weighted mean can naturally handle the loops in the phase and is the standard operation in directional statistics.
[0084] Instantaneous phase of the envelope path Instantaneous phase with fine structure pathway The difference is reduced to The interval is used to obtain the instantaneous phase difference sequence:
[0085] ;
[0086] in, It is an instantaneous phase difference sequence, representing the phase coordination state of the two tracking paths at each instant; The instantaneous phase of the envelope path; For the instantaneous phase of the fine structure path; To reduce the results to Standard modulo operation of the interval. The sequence is stable within the window when the phase lock of the two paths is good, and its values are scattered within the window when the phase relationship between the two paths begins to spread.
[0087] To avoid the situation where the two-way phase estimation itself has high noise. A sharp jump can occur, interfering with subsequent circular variance calculations. After complex exponential mapping, a moving average filter is applied in the time direction (in this embodiment, a moving average of 5 sampling points is used), and then the argument is taken to restore the phase sequence, which is used as the input for subsequent steps.
[0088] Next, the process of determining the phase difference diffusion slope includes: constructing a circular statistical measure applicable to the phase variable within the evaluation time window for the instantaneous phase difference sequence to characterize the phase distribution concentration of the instantaneous phase difference sequence within the evaluation time window, and obtaining the phase lock value of the evaluation time window; inverting the phase lock value to the phase diffusion degree, so that the phase diffusion degree is negatively correlated with the phase distribution concentration, and the sum of the phase diffusion degree and the phase lock value is a unit constant; calculating the rate of change of the phase diffusion degree with respect to time as the phase difference diffusion slope, the rate of change reflecting the speed at which the instantaneous phase difference sequence transitions from the phase lock state to the phase diffusion state.
[0089] The key design here is that instead of directly using the instantaneous or absolute values of the instantaneous phase difference sequence (because these values are significantly affected by individual differences), it uses the rate of change of its diffusion degree over time within the evaluation time window as the decision signal. This diffusion trend statistic is stable across individuals, sensitive to different noise types, and conforms to the objective physical characteristics of the dual-path coordination mechanism. The advantage of this design is that it overcomes the problem in traditional schemes where the absolute phase value is affected by individual differences, leading to uninterpretable evaluation results, and makes the decision signal more stable across subjects and noise types.
[0090] Evaluation Time Window The setting principles are as follows: First, it should fully cover the sentence-level rhythm of speech, with the lower limit of the sentence-level rhythm period being approximately several hundred milliseconds to 1 second; second, it should satisfy the statistical stability of the phase difference diffusion degree, requiring a sufficient number of phase difference samples within the window to stably estimate the circular statistic, with the sample size being at least several dozen. Based on the above principles, the evaluation time window in this embodiment... Take 2 seconds. When the evaluation period is shorter than At that time, the sliding window is used to fill or shorten forward at the start time. The number of samples must match the available time (but must meet the minimum sample number requirement of "no less than 30 samples within the window").
[0091] A circular statistical measure applicable to the phase variable is constructed for the instantaneous phase difference sequence within the evaluation time window. Specifically, for the evaluation time window... Each phase difference sample within Map it to the complex number on the unit circle Calculate the time average for all complex exponents within the window. The phase-locked value of the evaluation time window is obtained by taking the modulus of the average value over that time period.
[0092] ;
[0093] in, To evaluate the time window at time The phase-locked value (for the instantaneous phase difference sequence of two paths) characterizes the concentration of phase distribution of the instantaneous phase difference sequence within the evaluation time window; To evaluate the time window The instantaneous phase difference sequence within; To evaluate the length of the time window; Calculate the time average within the window; This is a complex modulo operation. The range of values for the phase-locked loop (PLL) value is... When all phase differences within the window are the same (perfect phase-locked loop), the phase-locked value is 1; when the phase difference is within... The interval is uniformly distributed (completely diffused), and the phase-locked value approaches 0.
[0094] Invert the phase-locked value to the phase spread:
[0095] ;
[0096] in, Phase spread is the degree of discretization of the phase coordination state of the two tracking paths; To evaluate the phase-locked value of the time window, the phase spread is negatively correlated with the phase distribution concentration, and the sum of the phase spread and the phase-locked value remains a unit constant of 1.
[0097] Calculate the rate of change of phase spread with respect to time, and use it as the phase difference spread slope:
[0098] ;
[0099] in, The phase difference diffusion slope represents the speed at which the phase coordination state of the two tracking paths transitions from lock-in to diffusion. Phase diffusivity; To evaluate the phase-locked loop value of the time window; To obtain the partial derivative with respect to time, in engineering implementation, this partial derivative is approximated using discrete difference methods (central difference or forward / backward difference) known in the art:
[0100] ;
[0101] in, The interval between adjacent evaluation moments, i.e., the evaluation time window. The sliding step size on the time axis is set to 100ms in this embodiment (i.e., calculation is performed every 100ms). and for adjacent Taking the difference yields ).
[0102] The phase difference diffusion slope reflects the speed at which the instantaneous phase difference sequence transitions from a phase-locked state to a phase-diffused state: when the phase diffusion increases rapidly (i.e., the instantaneous phase difference sequence quickly transitions from a locked state to a diffused state), the phase difference diffusion slope is significantly positive; when the phase diffusion is stable (either the locked state or the diffused state is maintained), the phase difference diffusion slope is close to zero; when the phase diffusion decreases rapidly (i.e., the instantaneous phase difference sequence returns from diffused to locked), the phase difference diffusion slope is significantly negative.
[0103] To ensure the robustness of the phase difference diffusion slope estimation under normal noise conditions, this step further performs correlation statistics on the phase-locked residuals of different sub-bands obtained independently in the fine structure reconstruction sequence: if the Pearson correlation coefficient between the residuals of different sub-bands is higher than 0.5 (indicating that the residuals are consistent), then the phase difference diffusion slope estimation is robust; if there is significant inconsistency between the residuals of different sub-bands (indicating that the phase estimation of some sub-bands may be contaminated by noise), then a sub-band residual consistency flag is reported to subsequent steps as an auxiliary criterion for triggering subsequent sub-band configuration updates.
[0104] Using the aforementioned clinical assessment example, in the first half (signal-to-noise ratio +5dB, 0 to 15 seconds), the phase-locked value... The phase spread stabilizes at approximately 0.82. The phase difference diffusion slope stabilizes at approximately 0.18. During the latter half of the transition period (when the signal-to-noise ratio deteriorates from +5dB to -5dB, approximately 15 to 17 seconds), the phase-locked value... The phase spread decreased from 0.82 to 0.40. The phase difference diffusion slope increases from 0.18 to 0.60. It shows a significant positive value.
[0105] In step S5, based on the phase difference diffusion slope, the envelope path weight and the fine structure path weight are determined. The envelope path weight is positively correlated with the phase difference diffusion slope, and the fine structure path weight is negatively correlated with the phase difference diffusion slope.
[0106] The core causal logic of weight reallocation lies in the following: when the two paths are well-locked, the phase difference diffusion changes slowly, the decoding results of both paths are reliable, and the weights of the two paths should be balanced; when the phase difference diffusion intensifies, the phase lock of the fine structure path begins to collapse, and the reconstruction result of this path is no longer reliable, so the weight of the fine structure path should be actively reduced and the weight of the envelope path increased; when the phase difference recovers from the diffusion state, the fine structure path becomes reliable again, and the weight of the fine structure path should be restored. The beneficial effect of this dynamic weight adjustment mechanism is that it prevents the evaluation result from being contaminated by the unreliable reconstruction result of the failed path when the coordination state of the two paths switches, thereby eliminating the problem of non-monotonic jumps in the scoring of the medium signal-to-noise ratio range.
[0107] Specifically, the process of determining the envelope path weight and fine structure path weight based on the phase difference diffusion slope includes: comparing the absolute value of the phase difference diffusion slope with a preset inflection point neighborhood discrimination threshold; if the absolute value is not greater than the inflection point neighborhood discrimination threshold, then setting the envelope path weight and fine structure path weight to be constant and not changing with the phase difference diffusion slope, and the sum of the two to be a unit constant; if the absolute value is greater than the inflection point neighborhood discrimination threshold, then obtaining the envelope path weight by applying a monotonically saturated mapping with a preset coupling coefficient to the phase difference diffusion slope, and keeping the sum of the envelope path weight and fine structure path weight a unit constant; the monotonically saturated mapping makes the envelope path weight monotonically approach the saturation value as the phase difference diffusion slope increases.
[0108] Specifically, this step has two working modes:
[0109] Mode 1 is the locked mode: when the absolute value of the phase difference diffusion slope is not greater than the inflection point neighborhood discrimination threshold (i.e. When the envelope path weights are applied, corresponding to the stable phase-locked region or the stable diffusion region, the contributions of the two paths are similar. With fine structure pathway weights Both are set to 0.5, and their sum remains a constant of 1, and neither changes with the specific value of the phase difference diffusion slope.
[0110] Mode 2 is the dominant weight redistribution mode: when the absolute value of the phase difference diffusion slope is greater than the inflection point neighborhood discrimination threshold (i.e. When ), the weights of the two paths are calculated using the following formula:
[0111] ;
[0112] in, For the envelope pathway at time The weights represent the proportions allocated to the envelope pathway reconstruction score during fusion; For fine-structured pathways at time The weights represent the proportions allocated to the fine-structure pathway reconstruction score during fusion; The phase difference diffusion slope; Let be the coupling coefficient of the weighting function, so that The value falls within the effective response range of the sigmoid function; The sigmoid function is defined as follows: , is a well-known monotonically saturated function in this field.
[0113] The boundary behavior of this monotonically saturated mapping is as follows: when the phase difference diffusion slope approaches 0, the envelope path weight approaches 0. The two pathways are fused in a balanced manner. When the phase difference diffusion slope increases significantly in the positive direction (i.e., diffusion intensifies rapidly), the weight of the envelope pathway approaches 1, and the weight of the fine structure pathway approaches 0. The fusion result is completely biased towards the envelope pathway, avoiding contamination by the reconstruction result of the unreliable fine structure pathway. When the phase difference diffusion slope increases significantly in the negative direction (i.e., diffusion quickly recovers and locks), the weight of the envelope pathway approaches 0, and the weight of the fine structure pathway approaches 1. The fusion result is biased towards the fine structure pathway (because this pathway becomes reliable again). Therefore, the weight of the envelope pathway is positively correlated with the phase difference diffusion slope, while the weight of the fine structure pathway is negatively correlated with the phase difference diffusion slope.
[0114] Inflection point neighborhood discrimination threshold Pre-calibrate using the following method:
[0115] Under a pre-defined calibration corpus, which contains multiple speech samples with different signal-to-noise ratio (SNR) levels (SNR is the power ratio of the speech signal to the background noise in a speech sample), multi-channel EEG signals are collected from the subjects or a pre-defined subject group at multiple SNR levels. A reference scoring sequence is obtained based on the correlation between the multi-channel EEG signals at each SNR level and the envelope feature sequences of the corresponding speech samples. Based on the results obtained by determining the phase difference diffusion slope at each SNR level, a sample sequence of phase difference diffusion slope relative to SNR is obtained. The reference scoring sequence is arranged in ascending order of SNR, and the lowest SNR level that satisfies the condition that the score corresponding to the next SNR is not greater than the score corresponding to the previous SNR is identified. The absolute value of the phase difference diffusion slope corresponding to the lowest SNR level is determined as the inflection point neighborhood discrimination threshold.
[0116] The design principle of the above calibration method is as follows: Under the known corpus where the monotonically increasing signal-to-noise ratio corresponds to the monotonically increasing intelligibility, the critical signal-to-noise ratio at which the reference scoring sequence exhibits a non-monotonic jump is the critical point at which the coordination state of the two pathways of the subject begins to collapse; the absolute value of the phase difference diffusion slope corresponding to this critical point is the reasonable threshold for triggering weight redistribution in subsequent evaluation.
[0117] The specific specifications of the calibration corpus are as follows: The calibration corpus contains 5 to 7 signal-to-noise ratio (SNR) levels (typically distributed in 5 dB intervals from −10 dB to +15 dB), with at least 2 minutes of speech-EEG synchronized data collected at each SNR level; the reference scoring sequence is calculated using the same method as the envelope reconstruction score in step S6 of this evaluation method, i.e., directly using the envelope reconstruction score at each SNR level (based on the already trained envelope reconstruction mapping). The EEG data collected at each signal-to-noise ratio level were reconstructed and compared with the envelope feature sequence of the corresponding speech sample (Pearson correlation) as a reference score.
[0118] Following the aforementioned clinical assessment example, EEG signals were collected from the subject at five signal-to-noise ratio levels: −5dB, 0dB, +5dB, +10dB, and +15dB. The corresponding reference scoring sequences were 0.32, 0.40, 0.55, 0.51, and 0.68, respectively, with corresponding absolute values of phase difference diffusion slopes of 0.30, 0.22, 0.18, 0.15, and 0.10, respectively. After sorting by signal-to-noise ratio in ascending order, it was found that the score corresponding to +10dB (0.51) was lower than the score corresponding to +5dB (0.55), violating monotonicity. Therefore, +10dB was identified as the critical signal-to-noise ratio level for this subject. The absolute value of the phase difference diffusion slope acquired at +10dB (0.15) was then used. The threshold for discriminating the inflection point neighborhood of this subject is determined as follows: This threshold can be used as an individual-specific threshold for subsequent evaluation of the subject, or it can be used as the average of the values obtained after repeating the above calibration on a group of subjects.
[0119] Weighting function coupling coefficient Threshold for discrimination in the neighborhood of the inflection point After calibration, determine the following formula: This ensures that, under the phase difference diffusion slope value corresponding to the critical signal-to-noise ratio,... The preset "significant bias" level is reached (0.9 in this embodiment), that is: ;
[0120] in, The coupling coefficients of the weighting function are such that the weighting function is coupled in... cross It immediately enters the effective bias zone; The absolute value of the phase difference spread slope corresponding to the critical signal-to-noise ratio level, and the threshold for discrimination in the neighborhood of the inflection point. Take the same value; for The value of the sigmoid input at that time. Continuing with the previous clinical assessment example, .
[0121] Using the aforementioned clinical assessment example, in the first half (signal-to-noise ratio +5dB). Enter locked mode. During the latter half of the transition period (the signal-to-noise ratio deteriorates from +5dB to -5dB), It has entered a weight redistribution-dominated mode. , The weights are significantly biased towards the envelope pathway.
[0122] In step S6, the envelope reconstruction score is determined based on the correlation between the envelope reconstruction sequence and the envelope feature sequence, and the fine structure reconstruction score is determined based on the phase consistency between the fine structure reconstruction sequence and the fine structure feature sequence.
[0123] Specifically, the envelope reconstruction score is the Pearson correlation coefficient between the envelope reconstruction sequence and the envelope feature sequence:
[0124] ;
[0125] in, The envelope reconstruction score characterizes the tracking quality of the envelope pathway for slow-modulated speech information. For envelope reconstruction sequence; Envelope feature sequence; The time mean of the envelope reconstruction sequence within the evaluation time window; This represents the time mean of the envelope feature sequence within the evaluation time window. The Pearson correlation coefficient ranges from [value missing]. The closer the value is to 1, the better the reconstruction quality.
[0126] The process of determining the fine structure reconstruction score includes: for each sub-band in the fine structure reconstruction sequence, calculating the phase-locked value between the instantaneous phase estimate of the sub-band and the instantaneous phase of the corresponding sub-band in the fine structure feature sequence to obtain the phase-locked value of each sub-band; calculating the time-averaged energy of the sub-band signal of each sub-band in the speech stimulus signal within the evaluation time window to obtain the sub-band energy of each sub-band; using the sub-band energy of each sub-band as the weight, performing a weighted summation of the phase-locked values of each sub-band and normalizing according to the sum of the sub-band energies of each sub-band to obtain the fine structure reconstruction score.
[0127] Specifically, for each sub-band The phase-locked value of this sub-band is calculated using the following formula:
[0128] ;
[0129] in, For the first The phase-locking value of each sub-band represents the degree of phase-locking between the EEG phase estimation and the speech phase of that sub-band. For the first in the fine structure reconstruction sequence Height belt at all times Instantaneous phase estimation; For the fine structure feature sequence, the first Height belt at all times The instantaneous phase; To assess the total number of samples within the time window; The unit is the imaginary unit. This calculation method is a standard implementation of phase-locked loop analysis (proposed by Lachaux et al. in 1999), which is well-known in the field.
[0130] Subband energy of each subband Using the sub-band energy defined in step S4, the phase-locked loop values of each sub-band are weighted and summed, then normalized according to the sum of the sub-band energies to obtain the fine structure reconstruction score.
[0131] ;
[0132] in, The fine structure reconstruction score characterizes the quality of the fine structure pathway in tracking fast oscillation information in speech. For the first Subband energy of each subband; For the first Phase-locked values for each sub-band; This represents the total number of sub-bands. The range of values for the fine structure reconstruction score is... The closer the value is to 1, the better the reconstruction quality. The design principle of using subband energy as the weight is that subbands with higher energy carry more speech information, and their phase-locked loop results can better reflect the subject's ability to track the fine structure of that frequency band. Therefore, they should have a greater weight in the fine structure reconstruction score.
[0133] Using the aforementioned clinical assessment example, in the first half of the assessment period (signal-to-noise ratio +5dB), the envelope reconstruction score... Fine structure reconstruction score In the latter half of the evaluation period (signal-to-noise ratio −5dB), the envelope reconstruction score... The fine structure reconstruction score is obtained by weighting and summing the phase-locked loop values of the 16 sub-bands using sub-band energy. .
[0134] In step S7, the envelope reconstruction score and the fine structure reconstruction score are weighted and summed using the envelope path weight and the fine structure path weight to obtain the auditory intelligibility score and output it.
[0135] Specifically, auditory intelligibility scores are calculated using the following formula:
[0136] ;
[0137] in, Auditory intelligibility scores characterize the objective level of speech intelligibility of the subjects during the assessment period; For envelope path weights; For fine-structured pathway weights; The envelope reconstruction score; The score is given for fine structure reconstruction.
[0138] The formula is formally a weighted summation operation known in the field, but its weights... , The weighting function, derived from the diffusion slope driven in step S5, is the core feature of this scheme. Its function is to replace the fixed fusion weights with dynamic weights driven by the dual-path coordination state, enabling the final score to automatically adapt to real-time changes in the decoding quality of the two paths. This allows for stable and interpretable evaluation results across different noise types, signal-to-noise ratio levels, and individual paths.
[0139] If a single comprehensive score is required clinically, rather than a temporal score, auditory intelligibility can be scored. The time average is summed over the entire assessment period to obtain the final single-value auditory intelligibility score.
[0140] Following the aforementioned clinical assessment example, in the first half of the assessment period (locked mode), The auditory intelligibility score is: .
[0141] Furthermore, the fine structure feature sequence is composed of the instantaneous phases of each sub-band obtained after sub-band decomposing the speech stimulus signal according to a preset number of sub-bands and preset sub-band bandwidth parameters for each sub-band. After obtaining the envelope path weights and fine structure path weights, and before obtaining and outputting the auditory intelligibility score, the method also includes the following correction mechanism:
[0142] The absolute value of the difference between the envelope path weights and the fine-structure path weights is used to obtain the path miscoordination:
[0143] ;
[0144] in, The path misalignment degree characterizes the relative imbalance in decoding quality between the two tracking paths; For envelope path weights; For fine-structured pathway weights.
[0145] If the path miscoordination is not greater than the preset miscoordination threshold (the miscoordination threshold in this embodiment) If we take 0.8, then we continue with the steps of weighted summation to obtain the auditory intelligibility score and output it.
[0146] If the path misalignment is greater than the preset misalignment threshold and the number of re-extractions is less than the preset maximum number of re-extractions (in this embodiment, the maximum number of re-extractions is 2), then the subband bandwidth parameter is reduced or the number of subbands is increased. The steps from extracting the fine structure feature sequence to obtaining the envelope path weight and fine structure path weight are re-executed with the updated subband configuration to obtain the updated envelope path weight and fine structure path weight. The updated path misalignment is then returned to the discrimination step for further discrimination. The number of re-extractions is the cumulative number of times the subband configuration has been updated since the start of this evaluation.
[0147] It should be noted that after each subband configuration update, the fine structural feature extraction in step S2 and the fine structural reconstruction in step S3 (including the corresponding brain electron band response decomposition, instantaneous phase extraction of each channel, and circular weighted summarization by channel weight, etc.) need to be re-executed according to the updated subband configuration. The reconstruction mapping of the envelope pathway... Unaffected and requiring no retraining, the computational overhead of backtracking correction is mainly concentrated on the fine structure pathway side.
[0148] If the pathway discoordination exceeds a preset discoordination threshold and the number of re-extractions has reached the preset maximum number of re-extractions, the weighted summation to obtain the auditory intelligibility score continues. The auditory intelligibility score, along with an assessment instability alarm flag, is output, allowing clinicians to obtain an assessment result with a reliability rating. The beneficial effect of this alarm mechanism is that, even in extreme cases where the coordination between the two pathways is severely disrupted and the self-correction mechanism cannot recover, an intelligibility score can still be provided for clinical reference. Simultaneously, the alarm flag clarifies the reliability limitations of the score for clinicians, preventing misjudgment.
[0149] Furthermore, the process of reducing the subband bandwidth parameter or increasing the number of subbands includes: selecting several subbands as target subbands from multiple subbands, and updating the subband configuration of the target subbands using at least one of the following methods: reducing the subband bandwidth parameter of the target subband; adding supplementary subbands in the frequency band neighborhood of the target subband, wherein the frequency band of the supplementary subband is located within the frequency band of the target subband or is adjacent to the frequency band of the target subband.
[0150] The process of selecting several subbands as target subbands from multiple subbands includes: for each subband in the fine structure reconstruction sequence, obtaining the phase-locked value of each subband based on the phase-locked value between the instantaneous phase estimate of that subband and the instantaneous phase of the corresponding subband in the fine structure feature sequence; and identifying the subbands whose phase-locked values fall within a preset abnormal interval as target subbands. The preset abnormal interval is usually set as the interval where the phase-locked value is significantly lower than the population average (in this embodiment, it is taken as...). (Interval), and sub-bands falling into this interval usually correspond to sub-bands that are contaminated by noise or have failed phase estimation.
[0151] Following the aforementioned clinical assessment example, in the latter half of the assessment period (signal-to-noise ratio −5dB), the envelope pathway weights... Fine-structured path weights Pathway miscoordination If the phase lock value is greater than the loss of coordination threshold of 0.8, and the current re-extraction count is 0 (less than the maximum re-extraction count of 2), then a backtracking correction is triggered: the phase lock values of the 16 sub-bands are checked, and the phase lock values of the 3rd, 5th, and 11th sub-bands are identified as 0.12, 0.08, and 0.15, respectively, all falling within the preset abnormal range. Within this process, these three sub-bands are identified as target sub-bands. The sub-band configuration is updated by "adding supplementary sub-bands in the frequency band neighborhood of the target sub-bands". That is, one supplementary sub-band is inserted in the frequency band of each target sub-band, increasing the total number of sub-bands from 16 to 19. The fine structure feature extraction in step S2 and the fine structure reconstruction in step S3 are re-executed with the updated sub-band configuration, and then steps S4 to S5 are re-executed to obtain the updated envelope path weights and fine structure path weights. The number of re-extractions is accumulated to 1.
[0152] Assuming the updated pathway misalignment decreases to 0.44 (less than the misalignment threshold of 0.8), the corresponding... , The updated fine structure reconstruction score changed from 0.22 to 0.25 (because the noise-contaminated subbands were refined, partially restoring the tracking ability of the fine structure pathways); then the weighted summation to obtain the auditory intelligibility score and its output continues, with the auditory intelligibility score being... .
[0153] If the path discrepancy is still greater than the discrepancy threshold after one backtracking, the number of re-extractions is further accumulated to 2 and the above update process is repeated; if the path discrepancy is still greater than the discrepancy threshold after two updates, the step of weighted summation to obtain the auditory intelligibility score is continued, and the auditory intelligibility score and the assessment instability alarm flag are output together.
[0154] In summary, this invention extracts the envelope feature sequence of slow modulation and the fine structural feature sequence of the instantaneous phase morphology of multiple subbands, and independently reconstructs two tracking estimates with different time window lengths. Then, it uses the rate of change of the diffusion degree of the instantaneous phase difference between the two tracking estimates within the evaluation time window as the decision signal to dynamically adjust the fusion weight of the two paths. When the two paths are severely out of sync, a backtracking correction mechanism for subband configuration updates is triggered. Thus, even when the coordination state of the two tracking paths switches, a stable and interpretable auditory intelligibility score can still be given. This solves the problems of non-monotonic jumps in the score in the medium signal-to-noise ratio range, inconsistencies across noise types, and difficulty in interpreting cross-individual differences in the existing technology.
[0155] Example 2:
[0156] like Figure 3 As shown, a brain-computer interface-based hearing-assisted assessment system includes:
[0157] The signal acquisition module is used to acquire multi-channel EEG signals and synchronous speech stimulation signals during the subject's listening process;
[0158] The feature extraction module is used to extract envelope feature sequences and fine structure feature sequences from speech stimulus signals. The envelope feature sequence reflects the low-frequency modulation components of the speech envelope, and the fine structure feature sequence is composed of the instantaneous phases of multiple sub-bands. The upper frequency limit of the envelope feature sequence is lower than the lower frequency limit of any sub-band in the fine structure feature sequence.
[0159] The signal reconstruction module is used to reconstruct the envelope reconstruction sequence corresponding to the envelope feature sequence from the multi-channel EEG signal using a first time window length, and to reconstruct the fine structure reconstruction sequence corresponding to the fine structure feature sequence from the multi-channel EEG signal using a second time window length. The fine structure reconstruction sequence consists of the instantaneous phase estimates corresponding to multiple sub-bands. The first time window length is greater than the second time window length.
[0160] The diffusion slope determination module is used to calculate the instantaneous phase difference sequence between the envelope reconstruction sequence and the fine structure reconstruction sequence. Within a preset evaluation time window, it determines the rate of change of the phase distribution concentration of the instantaneous phase difference sequence as the phase difference diffusion slope.
[0161] The weight determination module is used to determine the envelope path weight and the fine structure path weight based on the phase difference diffusion slope. The envelope path weight is positively correlated with the phase difference diffusion slope, while the fine structure path weight is negatively correlated with the phase difference diffusion slope.
[0162] The reconstruction score determination module is used to determine the envelope reconstruction score based on the correlation between the envelope reconstruction sequence and the envelope feature sequence, and to determine the fine structure reconstruction score based on the phase consistency between the fine structure reconstruction sequence and the fine structure feature sequence.
[0163] The scoring output module is used to perform a weighted summation of the envelope reconstruction score and the fine structure reconstruction score using the envelope path weight and the fine structure path weight, so as to obtain the auditory intelligibility score and output it.
[0164] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.
[0165] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0166] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A brain-computer interface-based hearing assistance evaluation method, characterized in that, Includes the following steps: Acquire multi-channel EEG signals and synchronized speech stimulation signals during the subject's listening process; An envelope feature sequence and a fine structure feature sequence are extracted from the speech stimulus signal. The envelope feature sequence reflects the low-frequency modulation component of the speech envelope. The fine structure feature sequence is composed of the instantaneous phases of multiple sub-bands, and the upper frequency limit of the envelope feature sequence is lower than the lower frequency limit of any sub-band in the fine structure feature sequence. The envelope reconstruction sequence corresponding to the envelope feature sequence is reconstructed from the multi-channel EEG signal using a first time window length, and the fine structure reconstruction sequence corresponding to the fine structure feature sequence is reconstructed from the multi-channel EEG signal using a second time window length. The fine structure reconstruction sequence is composed of instantaneous phase estimates corresponding to multiple sub-bands, and the first time window length is greater than the second time window length. Based on the envelope reconstruction sequence and the fine structure reconstruction sequence, the instantaneous phase difference sequence between the two is calculated. Within a preset evaluation time window, the time rate of change of the phase distribution concentration of the instantaneous phase difference sequence is determined as the phase difference diffusion slope. Based on the phase difference diffusion slope, the envelope path weight and the fine structure path weight are determined. The envelope path weight is positively correlated with the phase difference diffusion slope, and the fine structure path weight is negatively correlated with the phase difference diffusion slope. The envelope reconstruction score is determined based on the correlation between the envelope reconstruction sequence and the envelope feature sequence, and the fine structure reconstruction score is determined based on the phase consistency between the fine structure reconstruction sequence and the fine structure feature sequence. The envelope reconstruction score and the fine structure reconstruction score are weighted and summed using the envelope path weight and the fine structure path weight to obtain the auditory intelligibility score and output it.
2. The auditory-assisted assessment method based on brain-computer interface according to claim 1, characterized in that: Extracting envelope feature sequences and fine structure feature sequences from the speech stimulus signal includes: The speech stimulus signal is subjected to Hilbert transform to obtain an envelope analytic signal. The magnitude of the envelope analytic signal is taken as the instantaneous amplitude. The instantaneous amplitude is subjected to low-pass filtering to obtain the envelope feature sequence. The upper passband of the low-pass filter is the upper frequency limit of the envelope feature sequence. The speech stimulus signal is decomposed into multiple subband signals by a Gammatone filter bank or an equivalent rectangular bandwidth subband filter bank. A Hilbert transform is performed on each subband signal to obtain the subband analytical signal. The argument of the subband analytical signal is taken as the instantaneous phase sequence of the subband. The instantaneous phase sequences of each subband are aligned and summarized along a unified time axis to form the fine structure feature sequence.
3. The auditory-assisted assessment method based on brain-computer interface according to claim 1, characterized in that: The length of the first time window is not less than the minimum signal period corresponding to the upper frequency limit of the envelope feature sequence, and the length of the second time window is not less than the minimum signal period corresponding to the lower frequency limit of the lowest sub-band in the fine structure feature sequence. The calculation process of the envelope reconstruction sequence includes: expanding the samples of the multi-channel EEG signal within the first time window to obtain a time-delay feature vector; fitting a linear mapping from the time-delay feature vector to the envelope feature sequence using ridge regression to obtain an envelope reconstruction mapping; and performing inverse estimation of the time-delay feature vector based on the envelope reconstruction mapping to obtain the envelope reconstruction sequence. The calculation process of the fine structure reconstruction sequence includes: decomposing the multi-channel EEG signal into multiple brain electronic band responses according to the same sub-band configuration as the fine structure feature sequence; performing a Hilbert transform on each brain electronic band response within the second time window and taking the argument to obtain the instantaneous phase estimate of each sub-band; and aligning and summarizing the instantaneous phase estimates of each sub-band according to a unified time axis to form the fine structure reconstruction sequence.
4. The auditory-assisted assessment method based on brain-computer interface according to claim 1, characterized in that: The envelope reconstruction score is the Pearson correlation coefficient between the envelope reconstruction sequence and the envelope feature sequence; The process of determining the fine structure reconstruction score includes: for each sub-band in the fine structure reconstruction sequence, calculating the phase-locked value between the instantaneous phase estimate of the sub-band and the instantaneous phase of the corresponding sub-band in the fine structure feature sequence, and obtaining the phase-locked value of each sub-band; The time-averaged energy of each sub-band signal in the speech stimulus signal is calculated within the evaluation time window to obtain the sub-band energy of each sub-band. Using the sub-band energy of each sub-band as a weight, the phase-locked loop values of each sub-band are weighted and summed, and then normalized according to the sum of the sub-band energies of each sub-band to obtain the fine structure reconstruction score.
5. The auditory-assisted assessment method based on brain-computer interface according to claim 1, characterized in that: The calculation process of the instantaneous phase difference sequence includes: The envelope reconstruction sequence and the fine structure reconstruction sequence are resampled and aligned to a unified time axis; Perform a Hilbert transform on the resampled and aligned envelope reconstruction sequence and take the argument to obtain the instantaneous phase of the envelope path; The time-averaged energy of each sub-band signal in the speech stimulus signal is calculated within the evaluation time window to obtain the sub-band energy of each sub-band. Using the subband energy of each subband as a weight, a circular weighted average is performed on the instantaneous phase estimates of each subband in the resampled and aligned fine structure reconstruction sequence to obtain the instantaneous phase of the fine structure path; The difference between the instantaneous phase of the envelope path and the instantaneous phase of the fine structure path is reduced to... The instantaneous phase difference sequence is obtained by dividing the interval.
6. The auditory-assisted assessment method based on brain-computer interface according to claim 1, characterized in that: The process of determining the phase difference diffusion slope includes: A circular statistical measure suitable for phase variables is constructed for the instantaneous phase difference sequence within the evaluation time window to characterize the phase distribution concentration of the instantaneous phase difference sequence within the evaluation time window, thereby obtaining the phase-locked value of the evaluation time window; The phase-locked value is inverted to the phase spread, so that the phase spread is negatively correlated with the phase distribution concentration, and the sum of the phase spread and the phase-locked value is a unit constant. The rate of change of the phase spread with respect to time is calculated as the phase difference spread slope, which reflects the speed at which the instantaneous phase difference sequence transitions from a phase-locked state to a phase-spreading state.
7. The auditory-assisted assessment method based on brain-computer interface according to claim 1, characterized in that: The process of determining the envelope path weights and fine structure path weights based on the phase difference diffusion slope includes: The absolute value of the phase difference diffusion slope is compared with a preset inflection point neighborhood discrimination threshold. If the absolute value is not greater than the inflection point neighborhood discrimination threshold, then the envelope path weight and the fine structure path weight are set to be constant and do not change with the phase difference diffusion slope, and the sum of the two is a unit constant. If the absolute value is greater than the inflection point neighborhood discrimination threshold, then the envelope path weight is obtained by applying a preset coupling coefficient monotonic saturation mapping to the phase difference diffusion slope, and the sum of the envelope path weight and the fine structure path weight is kept as a unit constant; the monotonic saturation mapping makes the envelope path weight monotonically approach the saturation value as the phase difference diffusion slope increases; The threshold for discriminating the neighborhood of the inflection point is pre-calibrated in the following manner: Under a preset calibration corpus set, the calibration corpus set contains multiple speech samples with different signal-to-noise ratio levels, where the signal-to-noise ratio is the power ratio of the speech signal to the background noise in the speech sample; The multi-channel EEG signals were collected from the subjects or the pre-defined subject group at the multiple signal-to-noise ratio levels. A reference scoring sequence is obtained based on the correlation between the multi-channel EEG signals and the envelope feature sequences of the corresponding speech samples at each signal-to-noise ratio level. Based on the results obtained from the steps of determining the phase difference diffusion slope at each signal-to-noise ratio level, a sample sequence of the phase difference diffusion slope relative to the signal-to-noise ratio is obtained; The reference score sequence is arranged in ascending order of the signal-to-noise ratio (SNR). The lowest SNR level that satisfies the condition that the score corresponding to the next SNR is not greater than the score corresponding to the previous SNR is identified. The absolute value of the phase difference diffusion slope corresponding to the lowest SNR level is determined as the inflection point neighborhood discrimination threshold.
8. The auditory-assisted assessment method based on brain-computer interface according to claim 1, characterized in that: The fine structural feature sequence is composed of the instantaneous phase of each sub-band obtained by sub-band decomposing the speech stimulus signal according to the preset number of sub-bands and the preset sub-band bandwidth parameters of each sub-band; After obtaining the envelope pathway weights and fine structure pathway weights, and before obtaining and outputting the auditory intelligibility score, the method further includes: The absolute value of the difference between the envelope path weight and the fine structure path weight is calculated to obtain the path miscoordination. If the degree of path discrepancy is not greater than the preset discrepancy threshold, then continue to execute the step of weighted summation to obtain auditory intelligibility score and output it; If the path misalignment is greater than the preset misalignment threshold and the number of re-extractions is less than the preset maximum number of re-extractions, then the sub-band bandwidth parameter is reduced or the number of sub-bands is increased. The steps from extracting the fine structure feature sequence to obtaining the envelope path weight and fine structure path weight are re-executed with the updated sub-band configuration to obtain the updated envelope path weight and fine structure path weight. The updated path misalignment is then returned to the discrimination step for further discrimination. The number of re-extractions is the cumulative number of times the sub-band configuration has been updated since the start of this evaluation. If the path discrepancy is greater than the preset discrepancy threshold and the number of re-extractions has reached the preset maximum number of re-extractions, then the step of weighted summation to obtain the auditory intelligibility score continues, and the auditory intelligibility score and the evaluation instability alarm flag are output together.
9. The auditory-assisted assessment method based on a brain-computer interface according to claim 8, characterized in that: The process of reducing the subband bandwidth parameter or increasing the number of subbands includes: Select several subbands from multiple subbands as target subbands, and update the subband configuration of the target subbands using at least one of the following methods: Reduce the subband bandwidth parameter of the target subband; A supplementary sub-band is added within the frequency band neighborhood of the target sub-band, wherein the frequency band of the supplementary sub-band is located within or adjacent to the frequency band of the target sub-band; The process of selecting several subbands as target subbands from multiple subbands includes: for each subband in the fine structure reconstruction sequence, obtaining the phase lock value of each subband based on the phase lock value between the instantaneous phase estimate of the subband and the instantaneous phase of the corresponding subband in the fine structure feature sequence; and determining the subband whose phase lock value of each subband is within a preset abnormal range as the target subband.
10. A hearing-assisted assessment system based on a brain-computer interface, characterized in that, include: The signal acquisition module is used to acquire multi-channel EEG signals and synchronous speech stimulation signals during the subject's listening process; The feature extraction module is used to extract an envelope feature sequence and a fine structure feature sequence from the speech stimulus signal. The envelope feature sequence reflects the low-frequency modulation component of the speech envelope, and the fine structure feature sequence is composed of the instantaneous phases of multiple sub-bands. The upper frequency limit of the envelope feature sequence is lower than the lower frequency limit of any sub-band in the fine structure feature sequence. The signal reconstruction module is used to reconstruct the envelope reconstruction sequence corresponding to the envelope feature sequence from the multi-channel EEG signal using a first time window length, and to reconstruct the fine structure reconstruction sequence corresponding to the fine structure feature sequence from the multi-channel EEG signal using a second time window length. The fine structure reconstruction sequence is composed of instantaneous phase estimates corresponding to multiple sub-bands. The first time window length is greater than the second time window length. The diffusion slope determination module is used to calculate the instantaneous phase difference sequence between the envelope reconstruction sequence and the fine structure reconstruction sequence, and to determine the rate of change of the phase distribution concentration of the instantaneous phase difference sequence within a preset evaluation time window, as the phase difference diffusion slope. The weight determination module is used to determine the envelope path weight and the fine structure path weight based on the phase difference diffusion slope. The envelope path weight is positively correlated with the phase difference diffusion slope, and the fine structure path weight is negatively correlated with the phase difference diffusion slope. The reconstruction score determination module is used to determine the envelope reconstruction score based on the correlation between the envelope reconstruction sequence and the envelope feature sequence, and to determine the fine structure reconstruction score based on the phase consistency between the fine structure reconstruction sequence and the fine structure feature sequence. The scoring output module is used to perform a weighted summation of the envelope reconstruction score and the fine structure reconstruction score using the envelope path weight and the fine structure path weight to obtain and output the auditory intelligibility score.