Visual lip-aided assessment method for speech perception in noise environment based on electroencephalogram
By collecting and analyzing EEG signals and combining multimodal data fusion technology, the neural coding mechanism of lip information in noisy environments is quantified. This solves the problem that traditional methods are difficult to evaluate the auxiliary role of visual lip information in noisy environments, and achieves accurate quantification of the auxiliary role of visual lip information. This provides theoretical support for the rehabilitation of hearing-impaired patients and the optimization of hearing aids.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to effectively assess the neural mechanisms of visual lip information in speech perception under noisy conditions, and traditional methods lack quantitative analysis of neural oscillations, cross-frequency coupling, and dynamic connections in brain networks, failing to meet the assessment needs in complex acoustic environments.
By collecting EEG signals and combining multimodal data fusion analysis, the relative time-frequency power spectral density, phase-locking index, and power coherence between brain regions are extracted. The speech envelope is reconstructed using a multivariate time response function, and the neural coding mechanism of lip information in a noisy environment is quantified.
This study reveals the dynamic regulation mechanism of lip information on the auditory cortex and dorsal pathway, and achieves precise quantification of the auxiliary effect of visual lip information, providing theoretical support for the rehabilitation of hearing-impaired patients and the optimization of hearing aids.
Smart Images

Figure CN121412622B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brain-computer interface technology, and in particular to a visual lip-assisted assessment method for speech perception in noisy environments based on electroencephalography (EEG). The method simultaneously collects speech perception ability and EEG signals, and combines multimodal data fusion analysis to realize the assessment of the auxiliary role of visual lip information in complex noisy environments. Background Technology
[0002] In quiet environments, people can easily perceive speech. However, in daily verbal communication, speech is often interfered with by noise, making speech perception relatively difficult. In addition to auditory speech information, visual information is an important means and method to assist in the speech perception process. This is because visual information contains information about the timing and content of the incoming speech signal, including the amplitude envelope related to acoustic information and the location and movement of the sound, so as to indicate to the listener that the speaker has begun to speak and improve the accuracy of speech recognition.
[0003] Behavioral tests of speech perception (such as recognition accuracy and reaction time) are commonly used methods to evaluate the auxiliary role of visual lip information in complex acoustic environments. Under noise interference conditions, researchers quantify the gain effect of visual assistance on auditory perception by comparing the difference in speech recognition accuracy with and without visual lip information. However, such methods have significant limitations: First, behavioral indicators only reflect the perceptual results and cannot reveal the specific mechanisms by which lip information dynamically regulates the neural encoding of the auditory cortex (such as the superior temporal gyrus STG) and dorsal pathways (such as the inferior frontal gyrus IFG); second, the subjective strategies of the subjects (such as differences in attention allocation and lip-reading experience) can introduce significant individual biases, especially in the elderly or hearing-impaired populations, where their behavioral responses may be disconnected from their actual neural responses; furthermore, traditional behavioral paradigms struggle to capture the millisecond-level temporal dynamics of audiovisual integration. Therefore, there is an urgent need to integrate multimodal neurophysiological indicators (such as relative time-frequency power spectral density of EEG bands, cross-frequency coupling, and functional connectivity) to construct a more accurate evaluation system for audiovisual video recording and integration performance.
[0004] When people listen to each other's speech and perform speech recognition, the auditory primary cortex (Brodmann partitions: 41, 42) is involved in the initial processing of information transmitted from the outside world to the central nervous system. The superior temporal gyrus (STG, Brodmann partition: 21) and the superior temporal sulcus (STS) play important roles in speech processing. Therefore, in addition to simple behavioral assessments of the auxiliary role of visual-lip information, neural function indicators at the central nervous system level also have important reference value for assessing the auxiliary role of visual-lip information.
[0005] Electroencephalography (EEG) signals are signals generated by the brain that reflect human brain activity and response characteristics, as well as electrophysiological and pathological information. They offer advantages such as portable instruments, low testing costs, and high temporal resolution, and have been widely used in cognitive neuroscience and psychiatry. They are particularly suitable for providing a neurophysiological supplement to the assessment of visual-lip information assistance. Recording and analyzing EEG signals can lead to a better understanding of the neural mechanisms and functional characteristics of visual-lip information assistance, thus aiding in the optimization of hearing aids.
[0006] In natural communication scenarios, speech perception relies on the integration of auditory and visual information. We perceive external information and integrate visual and auditory information to form a unified perception, thereby enhancing speech perception. Among the visual information, the most important is lip information, which includes information such as amplitude envelope related to acoustic information and the position and movement of articulation. However, the assessment methods for the auxiliary role of lip information have the following limitations: 1. Single modality dependence: Existing technologies are mostly based on simple audiovisual stimuli (such as the McGurk effect), lacking effective utilization of dynamic lip information in natural language communication. 2. Poor noise adaptability: Existing EEG analysis is mostly for quiet environments and has not systematically studied the impact mechanism of noise type (such as white noise, speech noise) and signal-to-noise ratio (SNR) on neural coding. 3. Single assessment dimension: Traditional methods rely on behavioral indicators (such as recognition accuracy) and lack quantitative analysis of neural mechanisms such as neural oscillations, cross-frequency coupling, and dynamic connectivity of brain networks. Currently, some studies have begun to combine EEG technology with lip-based speech recognition. However, these research methods mostly focus on EEG data collection and analysis under simple conditions (such as the McGurk effect, simple initials, finals, and short words). Assessment schemes for the lip-based speech recognition effect in long, simple sentences within natural communication scenarios are still lacking. To address these issues, there is an urgent need for an alternative physiological indicator to design a more comprehensive assessment scheme for the lip-based speech recognition effect, to reveal the neural mechanisms of lip-based speech recognition in noisy environments, and to provide theoretical support for the rehabilitation of hearing-impaired individuals and the optimization of hearing aids. Summary of the Invention
[0007] The purpose of this invention is to address the technical deficiencies in the existing technology by providing a visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG).
[0008] The technical solution adopted to achieve the purpose of this invention is:
[0009] A visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) includes the following steps:
[0010] Step 1, conduct stimulation experiments to collect EEG data: During each experiment, the subject listens to the target speech audio, pays attention to the lip movements of the target speaker in the video, collects the subject's EEG data, judges whether the subject correctly repeats the audio and obtains the speech recognition accuracy rate.
[0011] Step 2: Preprocess the EEG data collected in Step 1. Select the preprocessed EEG data from 5 brain regions and extract them for a fixed duration. Group the data according to noise type, signal-to-noise ratio level, and visual conditions to obtain the EEG dataset. The 5 brain regions are STS pSTG, pIFG, MP, A, and V. The noise type includes white noise and speech noise. The visual conditions are stationary lip and moving lip.
[0012] Step 3: Extract EEG features from the EEG datasets of different groups obtained in Step 2: The EEG features include the relative time-frequency power spectral density of each electrode in each brain region at different time points. Phase lock-in index (PLV) between brain regions, power coherence (COH) between brain regions, and Pearson correlation between the reconstructed speech envelope and the original speech envelope at lag time ;
[0013] Step 4: Combine the EEG characteristics obtained in Step 3 under different stimulus conditions. PLV, COH The correlation between visual lip information and the speech recognition accuracy obtained in step 1 was analyzed to assess the auxiliary function of visual lip information in speech perception.
[0014] In the above technical solution, in step 1, the stimulation conditions of the stimulation experiment include five categories, namely, original audio, still lip video, and AV. s Still lip video with added white noise at different signal-to-noise ratios (A) w V s Audio with added speech noise and still lip video with different signal-to-noise ratios s V s Audio with added white noise and lip movement video with different signal-to-noise ratios w V c Audio with added speech noise and lip movement video with different signal-to-noise ratios s V c ;
[0015] Each stimulus experiment consists of 18 modules. First, press A. w V s A s V s A w V c A s V cThe four modules cycle four times, with the signal-to-noise ratio gradually increasing during each cycle. The last two modules are AV. s .
[0016] In the above technical solution, the preprocessing in step 2 includes noise reduction and artifact removal, filtering, and bad conduction processing. Step 2 involves extracting EEG data within 2 seconds of the target speaker beginning to speak to obtain a first EEG dataset. Step 3 uses this first EEG dataset to calculate the PLV and COH between different brain regions. ;
[0017] In step 2, EEG data is extracted 0.3 seconds before and 2 seconds after the target speaker begins speaking to obtain a second EEG dataset. In step 3, the second EEG dataset is used to calculate... .
[0018] In the above technical solution, in step 3, the relative time-frequency power spectral density The calculation includes the following steps:
[0019] Step s31: Based on the EEG data collected in the second EEG dataset, calculate the time-frequency power spectral density of each electrode at different time points and frequencies. ;
[0020] Step s32, based on the results obtained in step s31 Divide the power spectral density at each frequency point by the sum of the power spectral densities at all frequencies within 1-45Hz to obtain the relative time-frequency power spectral density of each electrode at different time points. ;
[0021] Step s33, for the results obtained in step s32 Baseline correction was performed to obtain the relative time-frequency power spectral density of each electrode at different time points after correction. ;
[0022] Step s34, based on the results obtained in step s33 Calculate the relative time-frequency power spectral density of each brain region under different stimulus conditions. The corresponding different frequency points under each frequency band The relative time-frequency power spectral density of each brain region at different frequency bands is obtained by summing the values. Time-frequency diagrams of each brain region under different stimulation conditions are then plotted to visually observe the changes in the relative power spectral density of each frequency band over time.
[0023] In the above technical solution, in step s31, The calculation method is as follows:
[0024] ;
[0025] yes c Electrode in k Frequency points corresponding to window time points f The time-frequency power spectral density, It represents c The first electrode m segment of EEG signal, n For a point in time, M The number of sampling points for each data segment. j Represents the imaginary unit , For window functions; .
[0026] In the above technical solution, in step s32, The calculation method is as follows:
[0027] ;
[0028] It is the relative time-frequency power spectral density of electrode c at the frequency point f corresponding to the k-window time point.
[0029] In the above technical solution, in step s33, The calculation method is as follows:
[0030] ;
[0031] After baseline correction was performed c Electrode in k Window time point f The relative time-frequency power spectral density at a given frequency point yes c The frequency point of the electrode at the baseline correction time point f The relative time-frequency power spectral density, This is the baseline correction time point.
[0032] In the above technical solution, in step s34, The calculation method is as follows:
[0033] ;
[0034] For each brain region in k Window time point f The relative time-frequency power spectral density at frequency points, where the Cortex is an STS pSTG, pIFG, MP, A, or V brain region. The electrodes represent the corresponding brain regions. The electrodes for the STS pSTG brain region include CP5, CP3, CP1, P5, P3, CP6, CP4, CP2, P6, and P4; the electrodes for the pIFG brain region include AF3, F5, F3, FC5, FC3, FC1, AF4, F6, F4, FC6, FC4, and FC2; the electrodes for the MP brain region include FZ, F1, F2, FC1, FC2, and CZ; the electrodes for the A brain region include FT7, TP7, T7, FT8, TP8, and T8; and the electrodes for the V brain region include PO3, PO5, PO7, O1, PO4, PO6, PO8, and O2.
[0035] In the above technical solution, the phase-locked index PLV is calculated in step 3 as follows:
[0036] ;
[0037] Representative at Channel within a time period x With channel y In time t Phase-locked index PLV, Represents the first trial under a single trial Within a time period, the time obtained using the Hilbert transform t passage x phase Phase with channel y The difference.
[0038] In the above technical solution, the calculation method for power coherence (COH) in step 3 is as follows:
[0039] ;
[0040] For the first Channel within a time period x and channels y Power coherence at frequency f, , and Representing channels x(m,t) With channel y(m,t) Power spectral density and channels x(m,t) Power spectral density, channels y(m,t) The power spectral density, For the first A time period t For a point in time.
[0041] In the above technical solution, in step 3, the multivariate time response function (mTRF) model is used as the decoder for calculation. The calculation method is as follows:
[0042] ;
[0043] ;
[0044] ;
[0045] ;
[0046] The original speech envelope, Representative at The reconstructed speech envelope under lag time, Electrode n exist EEG data at specific time points, Represents error, lag time The timeframe is -100ms to 500ms.
[0047] For decoder, , I It is the identity matrix. λ It is a smoothing constant. R It is the lagged time series of the response matrix r;
[0048] To reconstruct the speech envelope.
[0049] In the above technical solution, in step 4, the result obtained in step 3... PLV, COH and For each noise type and under different signal-to-noise ratio conditions, the static lip condition and the moving lip condition of the added noise frequency are compared. PLV, COH and The differences were compared to analyze whether visual lip information under noisy conditions affected the accuracy of speech perception.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] 1. This invention reveals the dynamic regulatory mechanism of lip information on the neural encoding of the auditory cortex (e.g., STG) and dorsal pathways (e.g., IFG / SMG) by simultaneously acquiring EEG signals and dynamic lip visual features of natural speech perception in a noisy environment, including... Phase locking values, power coherence, and relative time-frequency power spectral density enhancement in the frequency band (1-4Hz) and the thet band (4-8Hz);
[0052] 2. Multimodal quantization model: This invention uses multivariate time response function (mTRF) to reconstruct speech envelope and spectrogram from 64-lead EEG signals, and quantifies the contribution of lip movement information to the decoding of speech information under noise masking;
[0053] 3. This invention provides biomarkers for personalized rehabilitation training for hearing-impaired patients and multimodal noise reduction algorithms for hearing aids. By integrating high temporal resolution EEG features with dynamic lip movement analysis, this invention overcomes the limitations of traditional behavioral testing, achieving an upgrade in assessment from "description of perceptual results" to "analysis of neural mechanisms," and providing an innovative methodology for the precise quantification of the auxiliary role of visual lip information in complex acoustic environments. Attached Figure Description
[0054] Figure 1 This is a schematic diagram illustrating the experimental scheme of the present invention, wherein (a) is a material example and (b) is a schematic diagram of the experimental material types.
[0055] Figure 2 This is a schematic diagram of the experimental procedure of the present invention.
[0056] Figure 3 This is a brain example of the relative time-frequency power spectral density of the experimental scheme of this invention (stimulation condition is A with a signal-to-noise ratio of -12dB). w V s and A w V c And A with a signal-to-noise ratio of -6dB s V s and A s V c ).
[0057] Figure 4 This is an example diagram showing the changes of PLV and COH over time among different brain regions in the experimental scheme of this invention.
[0058] Figure 5 This is an example diagram showing the Pearson correlation between the reconstructed speech envelope and the original speech envelope of the EEG signal in the experimental scheme of this invention, as the time lag changes.
[0059] Figure 6 This is an example diagram comparing the Pearson correlation between the reconstructed speech envelope and the original speech envelope of a specific brain region after the addition of lip information in the experimental scheme of this invention. Detailed Implementation
[0060] The present invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0061] A visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) includes the following steps.
[0062] Step 1: Conduct stimulation experiments to collect EEG data:
[0063] (1) Video recording and calibration:
[0064] Speech resource library construction: The experiment used Mandarin Speech Perception (MSP) sentences ( Figure 1 (a)). Each sentence consists of 7 Chinese characters, totaling 50 sentences. All sentences were recorded in a soundproof room by a trained male at a normal speaking speed (sampling rate 48kHz), with simultaneous acquisition of the speaker's pronunciation video (resolution 1920×1080, 30fps). The recording equipment was a Nikon Z5 SLR camera. Each sentence was recorded multiple times during the recording process. The material with the clearest pronunciation and the best video quality was selected as the final experimental material.
[0065] The acquired video and audio were processed separately. The video was edited using Adobe Premiere Pro 2022. At 1920×1080 resolution, all areas except the central 400×200 area were set to black. The central area only showed the speaker's lips and surrounding area (below the nose, above the chin, and between the cheeks). The audio was noise-reduced using Adobe Audition 2020. Finally, the audio and video were combined and the video length was edited. The video duration from the start of the video to the speaker opening their mouth was 1.5 seconds, and the duration from the speaker finishing a sentence to the end of the video was 1 second. The video showing the speaker's moving lips is referred to as V. c (Congruent Video) refers to a video that shows the speaker's lips at rest. s (StaticVideo);
[0066] (2) Noise synthesis and classification:
[0067] The noise types used in the experiment included white noise and speech noise. White noise is a random signal with a constant power spectral density and uniform energy distribution across frequencies. White noise with a frequency range of 250Hz-8000Hz was generated using Matlab. Speech noise consisted of four people reading different excerpts from the Chinese classic novel *Robinson Crusoe* in a soundproof room. After noise reduction, the audio was mixed and edited using Adobe Audition 2020. The noise and the recorded sentence audio were superimposed; the original audio was referred to as A (Original Audio), and the white noise superimposed on the original audio was referred to as A (Original Audio). w(Audio with White Noise) refers to superimposing speech noise onto the original audio and calling it A. s (Audio with Speech Noise)( Figure 1 (b)
[0068] (3) Experimental conditions:
[0069] Visual conditions: Dynamic lip video (V c ), static lip video (V s );
[0070] Auditory conditions: Original audio (A), white noise superposition (A) w ), speech noise superposition (A) s );
[0071] (4) Experimental materials:
[0072] Original audio still lip video (AV) s ) is achieved through editing and mixing V s Together with A, 50 sentences were generated;
[0073] There are two types of still lip videos with added noise audio (A) w V s and A s V s This is achieved by mixing edited sentence audio with white noise or speech noise, and then synthesizing it with still lip video. White noise is mixed with sentences 1-25, and speech noise is mixed with sentences 26-50.
[0074] Noisy audio lip-motion videos also fall into two categories (A) w V c and A s V c The process involves mixing the edited sentence audio with white noise or speech noise, and then synthesizing it with the corresponding sentence's dynamic lip video.
[0075] White noise signal-to-noise ratio (SNR) wn The signal-to-noise ratio (SNR) of speech noise is set to four gradients: -15dB, -12dB, -9dB, and -7dB. sn The signal-to-noise ratio (SNR) is set to four levels: -9dB, -6dB, -4dB, and 0dB. Because speech noise is more disruptive than white noise, the SNR of speech noise is set higher than that of white noise to ensure that the degree to which they affect speech perception is similar.
[0076] The total experimental material consisted of 450 videos, including 200 videos of still lips with added noise (50 sentences * 4 signal-to-noise ratios), 200 videos of moving lips with added noise (50 sentences * 4 signal-to-noise ratios), and 50 videos of still lips with original audio.
[0077] The overall experimental procedure is as follows: Figure 2 As shown, it includes the following core steps.
[0078] (1) Subject preparation and equipment wearing:
[0079] EEG cap installation: Following the international 10-20 system, a NeuroScan device was used to record EEG data from 64 electrodes at a sampling rate of 1000Hz. All data were taken from the average value of all scalp channels, and all electrode impedances were kept below 10kΩ.
[0080] Stimulus standardization: Visual stimuli were presented on a monitor with a resolution of 1920×1080 and a refresh rate of 60Hz. Auditory stimuli were output using two JBL PEBBLES Mini BT2 speakers at a volume of 50dB SPL, and depending on the experimental content, white noise and speech noise were mixed at volumes of 50, 54, 56, 57, 59, 62, and 65dB SPL.
[0081] Standardized positioning: Subjects sit in the center of the soundproof room (1m away from the monitor), with their heads fixed to the support, keeping their horizontal line of sight aligned with the center of the screen.
[0082] (2) Experimental stimulus:
[0083] Stimulus presentation: Divided into 18 blocks, press A w V s →A s V s →A w V c →A s V c The blocks are presented in a block order, with the signal-to-noise ratio gradually increasing from low to high. The last two blocks are AV. s Each block contains 25 trials ( Figure 2 );
[0084] Participants were instructed to listen carefully to the target speech and pay attention to the speaker's lip movements in the video. After each sentence in the video was played, participants were asked to repeat the sentence they had just heard as much as possible. The experiment was recorded. The percentage of correct answers was calculated as the speech recognition accuracy rate. Participants were allowed to guess if they were unsure of the answer, but they could not give the same response for all sentences. If a participant could not hear clearly, they could answer that they did not hear clearly and proceed to the next trial.
[0085] Before data collection, the experimenters must communicate with the subjects to confirm and record their basic information, including age and gender, and complete a lip-reading test to ensure that subjects cannot directly obtain speech information through lip movements and to ensure that subjects meet the inclusion criteria. Before the experiment begins, the experimenter must explain the experimental procedure and the noise interference present in the experiment to the subjects and obtain their consent. Subjects are reminded to avoid unnecessary body and head movements during data collection to reduce artifacts in the EEG data.
[0086] Trial procedure:
[0087] 1. Fixation on the cross (1000ms) → 2. Audiovisual stimulation (4-5s dynamic / static lip video + noisy speech) → 3. Speech repetition and recording (maximum response time 10s), with A w V c For example ( Figure 2 (c)
[0088] Interval control:
[0089] Interval between trials (1.5s), rest for 2 minutes between blocks to avoid fatigue effects;
[0090] After completing the experiments on all blocks, the experiment was terminated, and the EEG data acquisition device worn by the subjects was removed to collect EEG data.
[0091] Step 2: Preprocess the EEG data collected in Step 1. The preprocessing includes noise reduction and artifact removal, filtering, and bad conduction processing. Select the preprocessed EEG data from 5 brain regions and extract them for a fixed duration to obtain the EEG dataset. The 5 brain regions are STS pSTG, pIFG, MP, A, and V.
[0092] Although unnecessary interference is minimized during EEG signal acquisition, the extremely weak EEG signals and low signal-to-noise ratio, coupled with susceptibility to various physiological and non-physiological factors such as electromyography (EMG), electrooculography (EOG), power line interference, and poor electrode contact, result in a significant amount of non-EEG signals in the acquired data. These signals are termed artifacts. To extract specific features from EEG signals, rigorous noise reduction and artifact removal processing is essential to eliminate or reduce the influence of these artifacts as much as possible, preserving the most authentic and original EEG signals.
[0093] This invention uses the open-source toolkit EEGLAB based on Matlab (R2021a) to process the recorded raw EEG data. First, since the EEG data in the experiment was acquired by the NeuroScan device, its raw data format already contains the position information of each electrode, so there is no need to perform electrode lead localization during data processing.
[0094] In the experiment, the sampling rate for acquiring EEG data was 1000Hz. Then, a 45Hz low-pass filter, a 1Hz high-pass filter, and a 48-52Hz band-stop filter were used to filter all data to reduce muscle artifacts and power line interference. The EEG data was then downsampled to 500Hz. According to the Nyquist sampling theorem, downsampling the data to 500Hz not only meets the accuracy requirements of data analysis but also increases processing speed and reduces unnecessary computation. This processing method better preserves the information of the EEG signal while eliminating or reducing the influence of interference noise, allowing for more accurate analysis of EEG features. In this invention, the original EEG data, after downsampling, can be used for subsequent data analysis and mining.
[0095] If an electrode experiences poor scalp contact due to the subject's movements during the experiment, resulting in increased impedance or significant signal drift, this electrode is classified as a bad lead. Although some EEG preprocessing tutorials directly delete bad leads, the number and location of bad leads can vary between different subjects. Directly deleting bad leads would lead to inconsistencies in electrode counts among different subjects, potentially increasing variables when comparing EEG responses and lacking rigor. Therefore, for bad lead processing, this invention uses the spherical curve method for interpolation calculation, that is, replacing the electrode with other electrodes of better signal quality after calculation. In actual processing, if the number of bad leads is greater than five, these data need to be excluded.
[0096] To better assess the participants' speech perception, the length of EEG data under all experimental conditions will be further processed. Data from each trial will be truncated, excluding the time segment during which the target speaker speaks (i.e., the first 1.5 seconds and the last 1 second) before further analysis. Thanks to the high temporal resolution of scalp EEG, EEG rhythm response analysis allows for in-depth understanding of the participants' brain response characteristics under various experimental conditions. Therefore, this protocol performs EEG rhythm response analysis under each experimental condition of speech perception.
[0097] According to noise type (white noise A) w Speech noise A s Signal-to-noise ratio (A) w -15dB to -7dB; A s -9dB to 0dB) and visual conditions (dynamic lip V) cStatic lip V s The EEG data were grouped. Then, the data was further segmented for analysis of relative time-frequency power spectral density. In each trial, EEG data from 0.3 seconds before and 2 seconds after the target speaker begins speaking (a total of 2.3 seconds) will be selected as the EEG dataset (i.e., truncated from 1.2 seconds to 3.5 seconds). The analysis will focus on phase lock value (PLV), power coherence (COH), and envelope reconstruction correlation. Pearson correlation between the reconstructed speech envelope and the original speech envelope at lag time When performing the experiment, the EEG data within 2 seconds after the target speaker begins to speak (i.e., 1.5s to 3.5s) is selected as the EEG dataset for each trial.
[0098] To better investigate the neural electrical activity of various brain regions along the speech perception pathway and the connectivity between different brain regions, five brain regions (STS pSTG, pIFG, MP, A, and V) will be selected for analysis according to the 10-10 standard lead system. Specifically, STS pSTG (superior temporal sulcus / posterior superior temporal gyrus) includes: CP5, CP3, CP1, P5, P3, CP6, CP4, CP2, P6, P4; pIFG (posterior inferior frontal gyrus) includes: AF3, F5, F3, FC5, FC3, FC1, AF4, F6, F4, FC6, FC4, FC2; MP (motor cortex) includes: FZ, F1, F2, FC1, FC2, CZ; A (auditory cortex) includes: FT7, TP7, T7, FT8, TP8, T8; and V (visual cortex) includes: PO3, PO5, PO7, O1, PO4, PO6, PO8, O2.
[0099] Step 3: Extract EEG features under different noise types, signal-to-noise ratio levels, and visual conditions based on the EEG dataset obtained in Step 2. The EEG features include the relative time-frequency power spectral density of each electrode in each brain region at different time points. Phase lock-in index (PLV) and power coherence (COH) between different brain regions Pearson correlation between the reconstructed speech envelope and the original speech envelope at lag time .
[0100] (a) Calculate the relative time-frequency power spectral density of each electrode in each brain region at different time points. .
[0101] To investigate neural responses during speech processing, time-frequency analysis can be used to convert continuous EEG signals to the time-frequency domain for measurement (example figure shown). Figure 3 The relative time-frequency power spectral density of each electrode at different time points was calculated using the short-time Fourier transform (STFT) method. :
[0102] Taking the time-frequency power spectral density of EEG data from a single electrode as an example, this method uses EEG signals of length 1150 (2.3s * 500 sampling rate, totaling 1150 sampling points) as an example. The power spectral density of each electrode at different time points and frequencies was calculated. , k For the time point of the window, n For the time point of the data, f For frequency points ( k, n =1,2,....1150; f =1,2,....45). The principle is as follows: a window length of 200ms (100 sampling points) and a step size of 2ms (1 sampling point) are selected. To allow the STFT to slide completely 1150 times, the EEG signal... Add 50 zero values at the beginning and end, and then add the EEG signal. It was divided into 1150 segments. Window function. Using a Hamming window, the time-frequency power spectral density of each electrode at different time points and frequencies is obtained by performing a short-time Fourier transform on 100 data points within each time window and then squaring the results. The specific calculation formulas are shown in Formulas 1-3:
[0103] (Formula 1);
[0104] (Formula 2);
[0105] In formula 1 w(n) For window functions, M The number of sampling points for each data segment is 100. n For each time point in the data, in Formula 2, for c Electrode in Frequency points corresponding to window time points f The power spectral density, w(n) For window functions, The representative is the first segment of EEG signal, j Represents the imaginary unit ;
[0106] (Formula 3);
[0107] Formula 3 is the time-frequency calculation formula extended to multiple leads, where Representative c Electrode in The power spectral density at frequency point f corresponding to the window time point. It representsc The first electrode Segment of EEG signals.
[0108] To reduce potential variability among subjects, the energy of each EEG signal needs to be normalized. To determine the relative time-frequency power spectral density of the EEG rhythm at each stage, for each subject at each time point, the power spectral density at each frequency point is divided by the sum of the power spectral densities at all frequencies within 1-45 Hz to calculate the relative time-frequency power spectral density;
[0109] (Formula 5);
[0110] In formula 5 Is the c electrode in The relative time-frequency power spectral density at the frequency point f corresponding to the window time point.
[0111] To eliminate time-related low-frequency drift or trend terms in the signal and avoid interference with time-frequency feature analysis, baseline correction is calculated using 0.2 seconds before the speaker begins speaking (baseline: 1.3s~1.5s). Data from 1.2s to 1.3s is not used for baseline correction because it involves zero-padding.
[0112] (Formula 6);
[0113] In formula 6 The c electrode represents the position after baseline correction. The relative time-frequency power spectral density at frequency f under the window time point It is the c electrode at the baseline correction time point The relative time-frequency power spectral density at the corresponding frequency point f. This represents a time range of 1.3s to 1.5s.
[0114] Calculated using MATLAB Then, further calculations were performed on each brain region. Window time point f Relative time-frequency power spectral density at frequency points The Cortex is located in the STS pSTG, pIFG, MP, A, or V brain regions;
[0115] ;
[0116] For each brain region in The relative time-frequency power spectral density at frequency f under the window time point, where the Cortex is an STS pSTG, pIFG, MP, A, or V brain region. Electrodes representing the corresponding brain regions.
[0117] by Taking brain regions as an example:
[0118] (Formula 7);
[0119] In formula 7 Representing the STS pSTG brain region The relative time-frequency power spectral density at the f-frequency point corresponding to the window time point, where cortex represents the corresponding electrode selected for the STS pSTG brain region.
[0120] Different frequency points corresponding to each frequency band By summing the values, the relative time-frequency power spectral density of different frequency bands in each brain region is obtained. By plotting the time-frequency diagrams of each brain region under different conditions, the changes in the relative power spectral density of each frequency band over time can be observed intuitively. Figure 3 As shown.
[0121] (ii) Phase Locking Index (PLV) and Power Coherence (COH) between Brain Regions
[0122] Phase lock-in index (PLV) and power coherence (COH) were used to calculate the PLV and COH values between brain regions in the delta, theta, and beta bands. To quantify connectivity changes more precisely and to correlate them with relative time-frequency power spectral density results, the EEG data from each trial within 2 seconds of the target speaker beginning to speak were divided into ten segments, each 200 ms long. The PLV and COH indices between brain regions were calculated for each segment. Figure 4 (As shown).
[0123] Phase synchronization between paired regions is measured using PLV (e.g., V vs. STS pSTG), where 0 ≤ PLV ≤ 1;
[0124] (Formula 8);
[0125] In formula 8 This represents the PLV values of channel x and channel y at time t over m time intervals. This represents the phase of channel x at time t, obtained using the Hilbert transform during the m-th time interval in a single trial. Phase with channel y The difference. N represents the number of trials.
[0126] The power coherence at different frequency bands was evaluated using COH.
[0127] (Formula 9);
[0128] (Formula 10);
[0129] In formula 9 This represents the cross-power spectral density of a certain frequency band f between channel x and channel y in the m-th time period under certain conditions. , and Let represent the power spectral density between channel x(m,t) and channel y(m,t), respectively, where m is the m-th time interval and t is the time point. In Formula 10... Represents the power coherence of channels x and y at frequency f during the m-th time interval, and is the coherence function. The square of the absolute value.
[0130] (III) Calculation Pearson correlation between the reconstructed speech envelope and the original speech envelope at lag time :
[0131] The effect of lip movement features on enhancing the Pearson correlation between the reconstructed speech envelope and the original speech envelope was quantified using a multivariate time response function (mTRF) model (e.g., ...). Figure 5 , Figure 6 This step will be completed using the mTRF toolbox in Matlab;
[0132] (Formula 11);
[0133] In Formula 11, the Decoder model generates speech envelope data by decoding the stimulus features EEG from neural responses;
[0134] (Formula 12);
[0135] In Formula 12, the decoder (i.e., the Decoder model) represents the neural response (i.e., EEG) Return to the original speech envelope A linear mapping of (i.e., Envelope) The lag time τ represents the error, ranging from -100ms to 500ms. Here, the decoder integrates the neural response within the specified lag time range τ. Ideally, these lag time windows will capture the neural data's response to the speech envelope, thereby optimizing the reconstruction of the speech envelope.
[0136] (Formula 13);
[0137] In Formula 13, the decoder By minimizing the original speech envelope and reconstructing the speech envelope It is estimated using the MSE between them.
[0138] In the actual calculation process, the decoder will easily calculate it using formula 14;
[0139] (Formula 14);
[0140] Referring to the Tikhonov regularization / ridge regression method, a smoothing term is added to reduce the variance of the estimates and prevent overfitting;
[0141] (Formula 15);
[0142] Where I is the identity matrix, and λ is the smoothing constant or "ridge parameter". Cross-validation can be used to adjust the ridge parameter;
[0143] ;
[0144] in R It is the lagged time series of the response matrix r. For simplicity, a single neural channel response system will be defined as follows: R :value min and max These represent the minimum and maximum time lags, respectively. R In the diagram, each time lag is arranged in a column, with non-zero lags padded with zeros to ensure causality. The time window for calculating the reconstructed envelope is defined as... window = max - min Therefore, the dimension of R is T× window In order to include a constant term (y-intercept) in the regression model, in R The left side is connected to a column of 1s. The variable r is a matrix containing all neural response data, with channels arranged column-wise (i.e., a T×N matrix; the example above shows a single neural channel). R ).
[0145] To scale to N channels, the method is to replace each column of R with N columns (each column representing a single channel). For N channels, the size of R is T×(N*). window The constant term is included by concatenating N columns 1 to the left side of R. The final generated decoder d is N*. window The vector.
[0146] Finally, the calculated decoder d is compared with the neural responses of the corresponding trials in the decoded set to obtain a T× window The matrix, where each column represents the corresponding... The reconstructed speech envelope is obtained by lag time. Then, the Pearson correlation between the reconstructed speech envelope and the original envelope is calculated, which quantifies the effect of lip movement features on improving the Pearson correlation between the reconstructed speech envelope and the original envelope.
[0147] (Formula 16);
[0148] (Formula 17);
[0149] In formula 16 Representative at The reconstructed speech envelope under lag time, Electrode n exist EEG data at specific time points, For the decoder. In Formula 17 In order to be in Pearson correlation between the reconstructed speech envelope and the original speech envelope under lag time. This represents the original speech envelope.
[0150] In actual calculation of P During the process, we use cross-validation to calculate the trials within each block, that is, the first trial is used as the test set, and the remaining trials from the second to the twenty-fifth are used as the training set, to obtain... and calculate the corresponding Then, the second trial is used as the test set, and the remaining trials are used as the training set. Similarly, the results are calculated. Finally, the results of twenty-five cross-calculations were... Average, take The maximum value in the value is used as the correlation between the reconstructed speech envelope of the block and the original envelope.
[0151] Step 4: Analyze the relative time-frequency power spectral density of each electrode at different time points under different stimulation conditions obtained in Step 3. Phase lock-in index (PLV) and power coherence (COH) between different brain regions Pearson correlation between the reconstructed speech envelope and the original speech envelope at lag time The correlation between the accuracy rate obtained in step 1 and visual lip information was analyzed to assess its auxiliary function in speech perception. A mapping relationship between neurophysiological indicators and perceptual efficacy was established, leading to the following conclusions:
[0152] (1) The relative time-frequency power spectral density of delta and theta of STS pSTG: A w V c and A s V c Compared to A w V s and A s V s The delta-theta relative time-frequency power spectral density is enhanced in the range of 200–500 ms;
[0153] (2) Connectivity of STS pSTG and V: A w V c and A s V c Compared to A w V s and A s V s Within 0ms to 400ms, the delta-theta PLV and COH values between STS pSTG and V increase;
[0154] (3) Correlation between the reconstructed speech envelope of STS pSTG and the original speech envelope: A w V c and A s V c Compared to A w V s and A s V s The maximum value of the Pearson correlation between the reconstructed speech envelope and the original speech envelope increases within the range of 0ms to 400ms.
[0155] If the subject's EEG data shows a loss or delay in the relative time-frequency power spectral density enhancement of the delta-theta in the STS pSTG brain region between 200ms and 500ms, or a loss or delay in the PLV and COH values of the STS pSTG brain region's connectivity with V between 0ms and 400ms, it indicates a weakened or untimely response to visual information, meaning the subject failed to effectively integrate visual information in a timely manner. Considering that this subject could not effectively utilize and integrate visual lip information to assist speech perception in a complex noisy environment, this suggests a weakened or untimely response to visual information.
[0156] If the subject is in A w V c and A s V cThe correlation between the reconstructed speech envelope and the original speech envelope in the STS pSTG and V brain regions under these conditions was compared with that in A. w V s and A s If the visual lip movement (V) is not effectively enhanced, it indicates that the subject failed to effectively track the target lip movement and acquire visual time and rhythmic information. Considering that subjects may not be able to effectively acquire visual lip information to aid speech perception in complex noisy environments, this could be a contributing factor.
[0157] The above description is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG), characterized in that, Includes the following steps: Step 1, conduct stimulation experiments to collect EEG data: During each experiment, the subject listens to the target speech audio, pays attention to the lip movements of the target speaker in the video, collects the subject's EEG data, judges whether the subject correctly repeats the audio and obtains the speech recognition accuracy rate. Step 2: Preprocess the EEG data collected in Step 1. Select the preprocessed EEG data from 5 brain regions and extract them for a fixed duration. Group the data according to noise type, signal-to-noise ratio level, and visual conditions to obtain the EEG dataset. The 5 brain regions are STS pSTG, pIFG, MP, A, and V. The noise type includes white noise and speech noise. The visual conditions are stationary lip and moving lip. Step 3: Extract EEG features from the EEG datasets of different groups obtained in Step 2: The EEG features include the relative time-frequency power spectral density of each electrode in each brain region at different time points. Phase lock-in index (PLV) between brain regions, power coherence (COH) between brain regions, and Pearson correlation between the reconstructed speech envelope and the original speech envelope at lag time ; In step 3, the multivariate time response function (mTRF) model is used as the decoder for computation. The calculation method is as follows: ; ; ; ; The original speech envelope, Representative at The reconstructed speech envelope under lag time, Electrode n exist EEG data at specific time points, Represents error, lag time The timeframe is -100ms to 500ms. For decoder, , I It is the identity matrix. λ It is a smoothing constant. R It is the lagged time series of the response matrix r; To reconstruct the speech envelope; Step 4: Combine the EEG characteristics obtained in Step 3 under different stimulus conditions. PLV, COH The correlation between visual lip information and the speech recognition accuracy obtained in step 1 was analyzed to assess the auxiliary function of visual lip information in speech perception.
2. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 1, characterized in that, Includes the following steps: In step 1, the stimulation conditions of the stimulation experiment include five categories, namely, original audio, still lip video, and AV. s Still lip video with added white noise at different signal-to-noise ratios (A) w V s Audio with added speech noise and still lip video with different signal-to-noise ratios s V s Audio with added white noise and lip movement video with different signal-to-noise ratios w V c Audio with added speech noise and lip movement video with different signal-to-noise ratios s V c ; Each stimulus experiment consists of 18 modules. First, press A. w V s A s V s A w V c A s V c The four modules cycle four times, with the signal-to-noise ratio gradually increasing during each cycle. The last two modules are AV. s .
3. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 1, characterized in that... The preprocessing in step 2 includes noise reduction and artifact removal, filtering, and bad conduction processing. Step 2 involves extracting EEG data within 2 seconds of the target speaker beginning to speak, obtaining the first EEG dataset. Step 3 uses this first EEG dataset to calculate the PLV and COH between different brain regions. ; In step 2, EEG data is extracted 0.3 seconds before and 2 seconds after the target speaker begins speaking to obtain a second EEG dataset. In step 3, the second EEG dataset is used to calculate... .
4. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 3, characterized in that, In step 3, the relative time-frequency power spectral density The calculation includes the following steps: Step s31: Based on the EEG data in the second EEG dataset, calculate the time-frequency power spectral density of each electrode at different time points and frequencies. ; Step s32, based on the results obtained in step s31 Divide the power spectral density at each frequency point by the sum of the power spectral densities at all frequencies within 1-45Hz to obtain the relative time-frequency power spectral density of each electrode at different time points. ; Step s33, for the results obtained in step s32 Baseline correction was performed to obtain the relative time-frequency power spectral density of each electrode at different time points after correction. ; Step s34, based on the results obtained in step s33 Calculate the relative time-frequency power spectral density of each brain region under different stimulus conditions. The corresponding different frequency points under each frequency band The relative time-frequency power spectral density of each brain region at different frequency bands is obtained by summing the values. Time-frequency diagrams of each brain region under different stimulation conditions are then plotted to visually observe the changes in the relative power spectral density of each frequency band over time.
5. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 4, characterized in that... In step s31 The calculation method is as follows: ; yes c Electrode in k Frequency points corresponding to window time points f The time-frequency power spectral density, It represents c The first electrode m segment of EEG signal, n For the time point of the data, M The number of sampling points for each data segment. j Represents the imaginary unit , For window functions; ; In step s32 The calculation method is as follows: ; It is the relative time-frequency power spectral density of electrode c at the frequency point f corresponding to the k-window time point.
6. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 4, characterized in that, In step s33 The calculation method is as follows: ; This represents the relative time-frequency power spectral density of electrode c at frequency f within the k-window time point after baseline correction. It is the relative time-frequency power spectral density of electrode c at the frequency point f corresponding to the baseline correction time point. It is the baseline correction time point; In step s34 The calculation method is as follows: ; The relative time-frequency power spectral density of each brain region at the f-frequency point under the k-window time point, where Cortex represents the STS pSTG, pIFG, MP, A, or V brain region. The electrodes represent the corresponding brain regions. The electrodes for the STS pSTG brain region include CP5, CP3, CP1, P5, P3, CP6, CP4, CP2, P6, and P4; the electrodes for the pIFG brain region include AF3, F5, F3, FC5, FC3, FC1, AF4, F6, F4, FC6, FC4, and FC2; the electrodes for the MP brain region include FZ, F1, F2, FC1, FC2, and CZ; the electrodes for the A brain region include FT7, TP7, T7, FT8, TP8, and T8; and the electrodes for the V brain region include PO3, PO5, PO7, O1, PO4, PO6, PO8, and O2.
7. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 1, characterized in that, In step 3, the phase-locked index PLV is calculated as follows: ; Representative at Channel within a time period x With channel y The phase-locked index PLV at time t, Represents the first trial under a single trial m The phase of channel x at time t obtained by Hilbert transform within a time interval. and channels y phase The difference.
8. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 1, characterized in that, In step 3, the power coherence (COH) is calculated as follows: ; For the first m Channel within a time period x and channels y Power coherence at frequency f, , and Representing channels x(m,t) With channel y(m,t) Power spectral density and channels x(m,t) Power spectral density, channels y(m,t) The power spectral density, m For the first m A time period t For a point in time.
9. The visual lip-assisted assessment method for speech perception in a noisy environment based on electroencephalography (EEG) as described in claim 1, characterized in that, In step 4, the result obtained in step 3... PLV, COH and For each noise type and under different signal-to-noise ratio conditions, the static lip condition and the moving lip condition of the added noise frequency are compared. PLV, COH and The differences were compared to analyze whether visual lip information under noisy conditions affected the accuracy of speech perception.
Citation Information
Patent Citations
Continuous voice envelope neural entrainment extraction method based on brain power supply imaging
CN113143293A
Multi-task auditory attention detection method and device based on electroencephalogram signals
CN118709119A