Speech decoding analysis method and system based on electroencephalogram signals and computer equipment
By using power ratio as a biomarker in brain-computer interface technology and utilizing frequency domain features to perform fine decoding of EEG signals, the limitations of existing natural language processing technologies have been overcome. This has enabled accurate identification and differentiation of word categories, improved the accuracy and reliability of language decoding, and provided a richer understanding of language processing mechanisms and a basis for rehabilitation training.
Patent Information
- Application Number
- CN202512041247.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-02-24
AI Technical Summary
Existing brain-computer interface technologies have significant limitations in processing natural language, especially in the field of language decoding. Traditional experimental studies cannot fully simulate and capture real language processing situations, resulting in limited understanding of the neural mechanisms of the brain in natural language processing and difficulty in explaining and distinguishing isolated neurophysiological signals under natural language stimulation.
Using power ratio as a biomarker, multi-channel EEG signals are divided into short-time EEG signals corresponding to individual words in the corpus, converted into frequency domain power spectral density, and the power ratio is calculated. The absolute power of the δ, θ, α, β, and γ sub-bands is calculated using Kaiser window combined with discrete Fourier transform. Semantic category classification is then performed using support vector machine, random forest, or neural network.
This technology enables precise decoding of different word classes under natural stimuli, improving the accuracy and reliability of brain-computer interface language decoding. It can more accurately identify and distinguish the EEG characteristics of different word classes, deeply explore the brain's language processing mechanism, provide precise rehabilitation training programs for patients with disabilities, and enhance their language communication abilities.
Smart Images

Figure CN121561106A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electroencephalogram (EEG) signal processing technology, and particularly relates to a speech decoding and analysis method, system, and computer device based on EEG signals. Background Technology
[0002] Brain-computer interface (BCI) technology, as an emerging medical communication method, aims to establish a direct information exchange channel between the brain and external devices. In the medical field, BCI technology offers paralyzed patients the possibility of regaining motor function and provides assistive communication tools for patients with speech disorders.
[0003] Language processing is one of the core functions of human cognition. As a tool for communication in society, language requires close cooperation and integration among the distributed neural networks in the brain. The concept of language is essentially knowledge about specific categories, and words, as the basic building blocks of language, are the fundamental units for constructing complex linguistic expressions and conveying meaning. Analysis at the individual word level is of great value in revealing the underlying neural mechanisms of patients with brain disorders. Different neural systems support different kinds of words; how to quantify this difference is one of the key questions in revealing the mechanisms of language processing.
[0004] However, current brain-computer interface (BCI) technology still has significant limitations in processing natural language. Particularly in the field of language decoding, traditional experimental research often relies on artificially designed language tasks. These tasks may not adequately simulate and capture real-world language processing scenarios, limiting a comprehensive understanding of the neural mechanisms involved in natural language processing. More importantly, the neural activity triggered by natural language stimuli often overlaps over time. This complexity makes interpreting and distinguishing isolated neurophysiological signals in natural stimulus environments an extremely challenging technical problem, requiring further research and innovation to overcome existing technological bottlenecks. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method that uses power ratio as a biomarker to achieve precise decoding of different word classes of words under natural stimuli based on the frequency domain features of EEG signals.
[0006] The technical solution of this invention is:
[0007] This invention provides a speech decoding and analysis method based on electroencephalogram (EEG) signals, comprising the following steps:
[0008] A corpus is selected as the natural stimulus, and the corpus is a continuous natural language streaming media;
[0009] Acquire multichannel EEG signals from subjects under the natural stimuli;
[0010] Based on timestamps, the multi-channel EEG signals are divided into short-time EEG signals corresponding to individual words in the corpus;
[0011] The short-time EEG signal is converted into a frequency domain power spectral density, and the power ratio of the short-time EEG signal is calculated based on the power spectral density.
[0012] Based on the power ratio, the words in the corpus are classified into semantic categories to achieve language decoding.
[0013] Further, the method for converting the short-time EEG signal into a frequency domain power spectral density and calculating the power ratio of the short-time EEG signal based on the power spectral density includes:
[0014] For each short-time EEG signal segment, a part-of-speech-signal length dual-linkage adaptive Kaiser window combined with discrete Fourier transform is used to estimate the power spectral density and calculate the absolute power of the signal in the δ, θ, α, β, and γ sub-bands.
[0015] Based on the absolute power, at least two types of power ratios are calculated as biomarkers, including gamma band power ratio, α / β band power ratio, and θ / β band power ratio.
[0016] Furthermore, the shape parameter β of the Kaiser window is dynamically adjusted according to the part of speech of the word and the length of the short-time EEG signal.
[0017] Furthermore, the shape parameter β of the Kaiser window is set based on the initial parameter setting of the part-of-speech category and dynamically corrected based on the short-time EEG signal length.
[0018] Furthermore, the part of speech of the words includes nouns, verbs, concrete words, abstract words, or emotional words, and the power ratio includes gamma band power ratio, α / β band power ratio, and θ / β band power ratio; wherein, the power ratio and the semantic category of the words have the following mapping relationship:
[0019] The gamma band power ratio is applicable to distinguishing between nouns, verbs, abstract words, and concrete words;
[0020] The θ / β band power ratio is suitable for distinguishing between nouns and verbs;
[0021] The α / β band power ratio is used to distinguish between abstract and concrete words.
[0022] Furthermore, the method for semantically classifying words in the corpus based on the power ratio to achieve language decoding includes:
[0023] At least two power ratios are used as joint feature vectors and input into a classification model to classify the semantic categories of words in the corpus. The semantic categories include at least one of part-of-speech category, specificity category, or emotional valence category, so as to achieve decoding from EEG signals to word semantic categories. The classification model is at least one of support vector machine, random forest, or neural network.
[0024] Furthermore, the corpus includes at least one of audiobooks, podcasts, audio, video, and text.
[0025] Furthermore, the method includes the following steps: analyzing and comparing the power ratios of different brain regions separately to analyze the differences in the response of different brain regions to word part-of-speech categories.
[0026] This invention also provides a speech decoding system based on electroencephalogram (EEG) signals, which applies the above-mentioned speech decoding and analysis method based on EEG signals, including:
[0027] The corpus module is used to provide continuous natural language streaming as natural stimuli;
[0028] Multi-channel EEG device: used to collect multi-channel EEG signals generated by the subject when receiving the natural stimulation;
[0029] The signal segmentation module is connected to the corpus module and the multi-channel EEG device respectively. It is used to receive the EEG signals collected by the multi-channel EEG device and segment the EEG signals into short-time EEG signals corresponding to a single word in the corpus based on the timestamp.
[0030] A frequency domain conversion module is used to receive the short-time EEG signal and convert the short-time EEG signal into a frequency domain power spectral density.
[0031] The calculation and analysis module is used to calculate the power ratio of the short-time EEG signal based on the power spectral density, and to perform semantic category classification of the words in the corpus based on the power ratio, and output the language decoding results.
[0032] The present invention also provides a computer device that is programmed or configured to perform the above-described speech decoding and analysis method based on electroencephalogram (EEG) signals, or the computer device has a computer program stored in its memory that is programmed or configured to perform the above-described speech decoding and analysis method based on EEG signals.
[0033] The beneficial technical effects of this invention are:
[0034] This invention divides EEG signals into short-time EEG signals corresponding to individual words in a corpus, converting these short-time EEG signals within the word context into frequency domain power spectral density. Based on this power spectral density, the power ratio of the short-time EEG signals is calculated, and this power ratio is used as a biomarker to semantically classify words in the corpus, quantifying part-of-speech differences to achieve language decoding. This solves the challenge of extracting EEG signals from individual word classes under natural stimuli and enables more accurate identification and differentiation of EEG features corresponding to different word classes. It allows for in-depth exploration of neural activity patterns in the brain when processing different word classes, providing richer information and deeper insights into understanding the brain's language processing mechanisms. This leads to more effective brain-computer interface language decoding, improving the accuracy and reliability of decoding. For patients with brain disorders, this analytical method helps to accurately locate neural abnormalities in their language processing, providing a basis for developing targeted rehabilitation training programs, thereby improving patients' language communication abilities, enhancing their quality of life, and increasing their social participation. Furthermore, this invention, based on a corpus of natural stimuli, enables the brain-computer interface system to be trained and applied in scenarios that more closely resemble real language use, improving the system's adaptability and flexibility to natural language. Through interdisciplinary methods and signal processing technology innovations in linguistics, neuroscience, and natural language processing, this invention achieves precise decoding of different parts of speech of words in natural language environments using short-time EEG signal frequency domain features. Attached Figure Description
[0035] Figure 1 This is a flowchart of the speech decoding and analysis method based on electroencephalogram (EEG) signals of the present invention;
[0036] Figure 2 This is a schematic diagram showing the position of the multi-channel electrodes according to a preferred embodiment of the present invention;
[0037] Figure 3 This is a diagram showing the brainwave segments corresponding to different emotion words;
[0038] Figure 4 This is a diagram showing the EEG segments corresponding to nouns and verbs;
[0039] Figure 5 It is the ratio of the gamma-band power of the noun to the verb;
[0040] Figure 6 It is the ratio of gamma-band power between concrete words and abstract words;
[0041] Figure 7 It is the ratio of gamma-band power of positive emotion words to neutral emotion words;
[0042] Figure 8 It is the ratio of gamma-band power of negative emotion words to neutral emotion words;
[0043] Figure 9 This is a schematic diagram showing the percentage of gamma band power corresponding to different brain regions;
[0044] Figure 10 This is a schematic diagram showing the power ratio of the theta / β bands corresponding to different brain regions;
[0045] Figure 11 This is a schematic diagram showing the power ratio of the α / β bands corresponding to different brain regions;
[0046] Figure 12 It is an evaluation of the accuracy of power ratio analysis and SWP analysis for classifying different word categories;
[0047] Figure 13 It is an evaluation of the F1 score for different word categories using power ratio analysis and SWP analysis. Detailed Implementation
[0048] In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0049] Please see Figures 1 to 13 As shown, this invention provides a speech decoding and analysis method based on electroencephalogram (EEG) signals, comprising the following steps:
[0050] S1: Select a corpus as a natural stimulus, wherein the corpus is a continuous natural language streaming media.
[0051] The corpus is a continuous natural language streaming media that meets the quantitative standards of semantic coverage ≥85%, part-of-speech balance ≥0.7, and emotion word intensity labeling error ≤5%, ensuring that it covers more than 95% of commonly used words in natural language (duration distribution 80-250ms).
[0052] Furthermore, the corpus includes at least one of audiobooks, podcasts, audio, video, and text.
[0053] In this embodiment, the corpus selected is the first chapter of the audiobook *Alice's Adventures in Wonderland* (continuous audio streaming, without manual interruption or deliberately designed language tasks). The *Alice's Adventures in Wonderland* corpus contains 84 consecutive sentences and 2129 words. After filtering using natural language processing techniques, it exhibits 88% semantic coverage, 0.75 part-of-speech balance, and a 3.2% error in emotion word intensity labeling, fully meeting the requirements for continuous natural language streaming.
[0054] Natural Language Processing (NLP) techniques are used to perform part-of-speech tagging on words in the corpus, providing a benchmark for subsequent classification. Specifically, the Natural Language Toolkit (NLTK) is used to retrieve nouns and verbs, a process that includes text segmentation, part-of-speech tagging, and word classification. Nouns and verbs are then identified and isolated based on their unique linguistic attributes. Words for the entire stimulus are classified according to the dataset's concreteness rating and a dictionary. SenticNet is used as a sentiment dictionary to annotate different sentiment words. Based on these unique semantic categories, the corpus is divided into concrete words, abstract words, nouns, verbs, and positive, negative, and neutral sentiment words.
[0055] For the aforementioned natural stimuli from *Alice's Adventures in Wonderland*, all words were multi-labeled using natural language processing techniques (i.e., a word can belong to multiple semantic categories simultaneously). After processing, the following independent and non-overlapping subsets of analysis were obtained for subsequent classification task validation:
[0056] Part-of-speech category task set: contains 354 pure noun samples and 434 pure verb samples (excluding words that have both noun and verb parts of speech).
[0057] The specificity category task set contains 257 words with high specificity scores (concrete words) and 1763 words with low specificity scores (abstract words).
[0058] The emotional valence task set, based on the SenticNet dictionary annotations, includes 136 negative emotion words, 206 positive emotion words, and 1107 neutral emotion words.
[0059] The sum of the sample sizes of the above subsets exceeds the total number of words because some words belong to multiple research dimensions (e.g., a word may be both a noun and a concrete word). However, in each specific classification task, we use mutually exclusive sample subsets to ensure the clarity of the evaluation.
[0060] Subjects (this study included 33 individuals proficient in English as participants. Subjects were native English speakers and demonstrated sufficient understanding of the stimulus material, *Alice's Adventures in Wonderland*.) were passively exposed to auditory stimulation. The audio playback speed was adjusted to 85% of normal speech speed (to optimize natural language understanding) and standardized (uniform sampling rate of 44.1 kHz, volume normalization). Throughout the natural stimulation process, subjects remained seated to avoid affecting the accuracy of EEG data. Multiple subjects, wearing multi-channel EEG devices, were simultaneously exposed to auditory stimulation.
[0061] It is important to understand that this invention studies the understanding of word semantics under natural stimuli, rather than auditory testing, thus providing a new perspective for speech decoding in brain-computer interfaces.
[0062] S2: Acquire multichannel EEG signals from the subject under the natural stimulus.
[0063] This invention is based on a corpus of natural stimuli, enabling brain-computer interface systems to be trained and applied in scenarios that more closely resemble real language use, thereby improving the system's adaptability and flexibility to natural language.
[0064] like Figure 2 As shown, this invention uses EEG data recorded by 61 active electrodes located using the international 10-10 system. The impedance is maintained below 20 kΩ. These signals are sampled at 500 Hz, and the EEG data from the scalp are bandpass filtered between 0.1 and 200 Hz.
[0065] Simultaneously, the device collects the acoustic features (fundamental frequency, duration, and energy peak) of each word in the corpus, providing data support for subsequent timestamp generation and signal segmentation. After collection, the raw EEG signals undergo preliminary quality screening to remove abnormal segments caused by blinking or electromyography interference (retaining subject data with a signal quality compliance rate ≥90%).
[0066] S3: Based on timestamps, the multi-channel EEG signals are divided into short-time EEG signals corresponding to individual words in the corpus. By timestamping the words in the corpus, the subject's EEG signals are recorded while receiving natural stimuli, thus constructing time-locked EEG signals. That is, based on timestamps, short-time EEG signals corresponding to individual words are generated.
[0067] The specific timestamp generation process is as follows:
[0068] Audio stimuli: Mel frequency cepstral coefficients (MFCC) are used to extract the acoustic features of words, and dynamic time warping (DTW) algorithm is used to detect abrupt changes in acoustic features to achieve word boundary localization with a localization error of ≤ ±5ms;
[0069] Video / text stimuli: timestamps are marked based on the start / end time of word presentation (video subtitle timeline, text scrolling timer), with a positioning error of ≤±3ms.
[0070] By extracting the start and end times of each word from the stimulus, a unique timestamp can be provided for each word. Therefore, the time series of a single short-time EEG signal corresponding to each word can be extracted, so that the short-time EEG signal can form a time mapping with a single word, thereby solving the problem of extracting EEG signals of a single word class under natural stimuli.
[0071] Please see Figure 3 As shown, Figure 3This is a schematic diagram of EEG segments corresponding to different emotion words, that is, a schematic diagram of the time series of short-time EEG signals corresponding to the time series of different emotion words; please refer to [link / reference]. Figure 4 As shown, Figure 4 This is a schematic diagram of the EEG segments corresponding to nouns and verbs, that is, a schematic diagram of the time series of nouns and verbs corresponding to the time series of short-time EEG signals. Figure 3 and Figure 4 It is known that by providing a separate timestamp for each word, EEG signals can be more accurately associated with words, thereby improving the accuracy and reliability of the research. This invention can overcome the limitations of extracting single-word EEG signals under traditional laboratory conditions, making it possible to study brain activity under more natural stimulation conditions.
[0072] S4: Convert the short-time EEG signal into a frequency domain power spectral density, and calculate the power ratio of the short-time EEG signal based on the power spectral density. This conversion transforms the time-domain EEG signal into a frequency-domain power spectral density to extract more discernible features.
[0073] The EEG signal length corresponding to a single word under natural language stimulation is 80 to 250 sampling points, which is a short-time non-stationary signal. Directly using conventional window functions such as rectangular window and Hanning window will lead to serious spectral leakage. However, the Kaiser window has adjustable sidelobe suppression capability and can flexibly balance time-frequency resolution, making it the preferred window function in this step.
[0074] Furthermore, this invention utilizes a part-of-speech-signal length dual-linkage adaptive Kaiser window combined with Discrete Fourier Transform (DFT) for power spectral density estimation. This allows for a more flexible balance between the time resolution and frequency resolution in short-time EEG signal spectrum analysis, helping to obtain more accurate power spectral estimates with less leakage, thereby improving the discriminative power of subsequent power ratio features.
[0075] Furthermore, the shape parameter β of the Kaiser window is dynamically adjusted according to the part of speech of the word and the length of the short-time EEG signal to maximize the final word classification accuracy.
[0076] Furthermore, the shape parameter β of the Kaiser window is set based on the initial parameter setting of the part-of-speech category and dynamically corrected based on the short-time EEG signal length to ensure complete coverage of the signal segment corresponding to each word.
[0077] Specifically, the initial parameter settings are based on part-of-speech categories: the system pre-defines different shape parameter β propensity values for different word semantic categories. This setting is based on the neural oscillation characteristics and cognitive depth involved in the brain's processing of different parts of speech. When processing nouns, the system tends to use a relatively large β value. This is because nouns usually represent concrete entities and are expected to induce more stable and continuous neural activity patterns; using a larger β value can provide better frequency resolution to accurately capture their spectral features. When processing verbs, the system tends to use a medium β value to achieve a balance between temporal and frequency resolution, adapting to the dynamic processing of action concepts. When processing adjectives, adverbs, emotion words, or abstract words, the system tends to use a relatively small β value. The processing of these words often relies more on context or involves diffuse neural integration; using a smaller β value can enhance the sidelobe suppression capability of the window function, thereby obtaining a more stable and less leaky spectral estimate, which is crucial for extracting reliable band power features.
[0078] After determining the initial tendency based on part-of-speech tags, the system will perform dynamic correction according to the actual length of the current short-time EEG signal to address the impact of signal duration variations on the quality of spectrum estimation.
[0079] Specifically, the system presets a reference duration. The current signal length is compared to this reference duration. If the current signal length is shorter than the reference duration, based on the principle that "the shorter the signal, the stronger the required sidelobe suppression," the β value is adjusted upwards by a preset increment based on the part-of-speech tendency value obtained in the first step. The greater the difference between the current signal length and the reference length, the larger the upward adjustment. If the current signal length is longer than the reference duration, the β value is finely adjusted downwards by a preset decrease based on the part-of-speech tendency value to moderately improve frequency resolution while maintaining basic spectral quality. If the current signal length is equal to the reference duration, the β value mainly follows the part-of-speech tendency value.
[0080] The shape parameter β of the Kaiser window is dynamically adjusted by the part-of-speech tag and the length of the short-time EEG signal to form a Kaiser window shape parameter β that adapts to the part-of-speech attribute and time scale. Subsequently, the Kaiser window shape parameter β is used for the subsequent Discrete Fourier Transform. This setting significantly improves classification performance and maximizes the final word classification accuracy.
[0081] Furthermore, the spectrum estimation and power ratio calculation process is as follows:
[0082] First, determine the β parameter and window length of the Kaiser window based on the part-of-speech tags and actual length of the short-time EEG signal;
[0083] After applying an adaptive Kaiser window to the signal, a 512-point DFT transform is performed.
[0084] The absolute power of the sub-bands δ (0.5~4Hz), θ (4~8Hz), α (8~13Hz), β (13~30Hz), and γ (30~45Hz) is calculated based on the transformation results;
[0085] Ultimately, at least two power ratios were selected as biomarkers, including the gamma band power ratio (the ratio of the γ band to the total power of the entire frequency band), the α / β band power ratio (the ratio of the α band to the β band power), and the θ / β band power ratio (the ratio of the θ band to the β band power).
[0086] This setup significantly improves the specificity and accuracy of the power ratio, laying a core foundation for subsequent fine-grained part-of-speech classification.
[0087] S5: Based on the power ratio, perform semantic category classification on the words in the corpus to achieve language decoding, thereby realizing the mapping and decoding from neural features to semantic information.
[0088] Furthermore, the part of speech of the words includes nouns, verbs, concrete words, abstract words, or emotional words, wherein the emotional words include positive emotional words, negative emotional words, and neutral emotional words.
[0089] Furthermore, the power ratio includes the gamma band power ratio, the α / β band power ratio, and the θ / β band power ratio.
[0090] Specifically, the method for semantically classifying words in the corpus based on the power ratio to achieve language decoding includes:
[0091] S51: Feature Construction and Model Input: Combine the multiple power ratios calculated in S4 (e.g., the gamma band power ratio and θ / β band power ratio of a word) into a joint feature vector. For example, when classifying nouns / verbs, it is preferable to use the θ / β band power ratio and gamma band power ratio to form a two-dimensional feature vector; when classifying concrete words / abstract words, it is preferable to use the α / β band power ratio and gamma band power ratio to form a two-dimensional feature vector; when classifying sentiment valence, the gamma band power ratio, θ / β band power ratio, and α / β band power ratio can be used together to form a three-dimensional feature vector. Combine the selected power ratios calculated from the short-time EEG signal corresponding to each word to form the feature vector of that word.
[0092] S52: Classification Decoding: Input the feature vector into a pre-trained classification model (Support Vector Machine (SVM) is used in this embodiment, but random forest or neural network can also be used). The task of the model is to classify the semantic category of words, and the classification target includes at least one of the following: part-of-speech category (noun / verb), concreteness category (concrete word / abstract word), or sentiment valence category (positive / negative / neutral).
[0093] Specifically, please refer to Figures 5 to 8 As shown, where, Figure 5 It is the ratio of the gamma-band power of the noun to that of the verb. Figure 6 It is the ratio of gamma-band power between concrete words and abstract words. Figure 7 It is the ratio of gamma-band power of positive emotion words to neutral emotion words. Figure 8 It is the gamma-band power ratio of negative emotion words to neutral emotion words.
[0094] Depend on Figures 5 to 8 It is known that the gamma-ray power ratio (GBR) is an effective tool for distinguishing different word categories, especially when processing nouns, verbs, concrete words, abstract words, and emotional words. When comparing nouns and verbs, or concrete words and abstract words, most recording sites on the EEG electrodes showed significant differences, indicating that the GBR can effectively capture the differences in neural activity between these categories. However, in comparisons involving emotional words, such as positive and neutral emotional words, or negative and neutral emotional words, the differences between the PO3 and PO4 sites were not significant. This may be because these emotional words have similar neural activity characteristics in specific brain regions, making it difficult for the GBR to distinguish them. Nevertheless, the GBR still exhibits high discriminative power in other comparisons, providing important evidence for language decoding.
[0095] The θ / β band power ratio exhibits varying degrees of sensitivity and discriminative ability across different aspects of language processing. It demonstrates a significant advantage in distinguishing between nouns and verbs, as well as between positive and neutral emotional words, but performs poorly in differentiating between concrete and abstract words, and between negative and neutral emotional words.
[0096] Distinguishing between nouns and verbs: The theta / beta band power ratio demonstrated high sensitivity and effectiveness in distinguishing between nouns and verbs. Significant differences were observed at 55 out of 61 recorded sites, indicating that the theta / beta band power ratio can significantly differentiate between nouns and verbs. This significance may stem from the differences in language function and brain processing mechanisms between nouns and verbs, which the theta / beta band power ratio effectively captures.
[0097] Distinguishing between concrete and abstract words: The θ / β band power ratio showed low sensitivity in distinguishing between concrete and abstract words. Only 7 out of 61 recording sites showed significant differences. This suggests that the θ / β band power ratio is not well-suited for distinguishing between concrete and abstract words, possibly because the neural activity characteristics of concrete and abstract words in the θ / β band are quite similar.
[0098] Processing of Emotional Words: In processing emotional words, the theta / β band power ratio showed varying discriminative power across different emotion categories. When comparing negative and neutral emotional words, no significant differences were found in the theta / β band power ratio at 44 out of 61 recording sites, indicating no significant difference between negative and neutral emotional words at these sites. However, when comparing positive and neutral emotional words, significant differences were found in the theta / β band power ratio at 51 out of 61 recording sites, suggesting a significant difference between positive and neutral emotional words in this regard. This difference may be related to the more unique or stronger neural activity patterns elicited by positive emotional words in the brain.
[0099] When distinguishing between concrete and abstract words, 42 out of 61 recording sites showed significant differences, indicating that the α / β band power ratio has high sensitivity and effectiveness in distinguishing between concrete and abstract words, effectively differentiating these two types of vocabulary. Concrete words are generally associated with sensory experiences and concrete objects, while abstract words involve more concepts and ideas. The significant difference in the α / β band power ratio when the brain processes these two types of words may be because the brain utilizes different neural resources and cognitive processes when processing concrete and abstract words.
[0100] When distinguishing between verbs and nouns, only 15 out of 61 recording sites showed significant differences, indicating that the α / β band power ratio is relatively insensitive in differentiating between verbs and nouns. Verbs and nouns represent actions and things in language, respectively, and they have different functions in syntax and semantics. Although the α / β band power ratio can distinguish between these two types of words at some sites, its overall effect is not as significant as the distinction between concrete and abstract words. This may be because the neural representations of verbs and nouns in the brain have more overlap or similarity, making it difficult to accurately distinguish them using only the α / β band power ratio.
[0101] Regarding the differentiation of emotion words, the α / β band power ratio performed as follows: When comparing positive and negative emotion words, most of the 61 recorded sites showed no significant difference, indicating that the α / β band power ratio has low sensitivity in directly distinguishing between positive and negative emotion words. When comparing negative and neutral emotion words, only 10 of the 61 recorded sites showed significant differences, indicating that the α / β band power ratio has weak specific recognition ability for negative emotion words. However, when comparing positive and neutral emotion words, 47 of the 61 recorded sites showed significant differences, indicating that the α / β band power ratio has a better effect in distinguishing between positive and neutral emotion words. This may be because positive emotion words trigger more unique or stronger neural activity patterns in the brain, and these patterns can be captured better in the α / β band. Overall, the performance of the α / β band power ratio in emotion word classification shows that it has a certain ability to distinguish between different emotion categories, but the effect varies depending on the combination of emotion word types.
[0102] Experimental results show that different power ratio features exhibit varying sensitivities in distinguishing different semantic categories. Specifically:
[0103] The gamma band power ratio demonstrates broad discriminative value in various classification tasks (noun / verb, concrete / abstract, and emotion words). The θ / β band power ratio exhibits high sensitivity and classification accuracy in distinguishing nouns from verbs, but its effectiveness is limited in distinguishing concrete from abstract words. The α / β band power ratio significantly contributes to distinguishing concrete from abstract words, but its sensitivity is relatively low in distinguishing nouns from verbs and certain emotion word categories (such as negative and neutral). Therefore, in practical decoding, different power ratio features can be selected or combined to construct a joint feature vector based on the target classification task (such as part of speech, specificity, and emotion) to improve decoding performance.
[0104] By dividing EEG signals into short-time EEG signals corresponding to individual words in a corpus, the short-time EEG signals in the context of a word are converted into frequency domain power spectral density. Based on the power spectral density, the power ratio of the short-time EEG signals is calculated. Using the power ratio as a biomarker, words in the corpus are semantically categorized, and part-of-speech differences are quantified to achieve language decoding. This solves the problem of extracting EEG signals of individual word classes under natural stimuli and can more accurately identify and distinguish the EEG features corresponding to different word classes. It allows for in-depth exploration of the neural activity patterns of the brain when processing different word classes, providing richer information and deeper insights into understanding the brain's language processing mechanisms. This leads to more effective brain-computer interface language decoding, improving the accuracy and reliability of decoding. Furthermore, this invention achieves fine-grained decoding of different parts of speech of words in a natural language environment through interdisciplinary methods and signal processing technology innovations in linguistics, neuroscience, and natural language processing.
[0105] S6: Power ratio response analysis of different brain regions.
[0106] To further explore the brain's language processing mechanisms and optimize decoding performance, this invention also includes separate analysis and comparison steps of the power ratio response characteristics of different brain regions:
[0107] Specifically, according to the electrode layout, a total of 61 electrodes cover the entire brain, with each electrode number corresponding to a marked brain region (e.g., ...). Figure 2 (As shown). Different brain regions exhibit different neural differences when comparing nouns, verbs, concrete words, abstract words, and emotional words, such as... Figures 9 to 11 As shown.
[0108] Figure 9 These represent the gamma band power ratios of different brain regions. Subgraphs (a) and (b) correspond to CP3 (e16), (c) and (d) correspond to CP4 (e12), (e) and (f) correspond to P3 (e29), (g) and (h) correspond to P4 (e26), (i) and (j) correspond to PO3 (e28), and (k) and (l) correspond to PO4 (e27). Subgraphs (a), (c), (e), (g), (i), and (k) represent the gamma band power ratios of nouns, verbs, concrete words, and abstract words, while subgraphs (b), (d), (f), (h), (j), and (l) represent the gamma band power ratios of negative, neutral, and positive emotion words.
[0109] Depend on Figure 9As can be seen, subplots (j) and (l) show that the differences in emotion words at PO3 and PO4 loci are not significant. Furthermore, when comparing positive and negative emotion words, 20% of the recorded loci showed no significant differences, suggesting that gamma band power ratios may be less sensitive to emotion words in specific brain regions.
[0110] Figure 10 These represent the theta / beta band power ratios for different brain regions. Subplots (a) and (b) correspond to FC5 (e32), (c) and (d) to PO7 (e45), (e) and (f) to TP7 (e46), (g) and (h) to T7 (e47), (i) and (j) to FT7 (e48), and (k) and (l) to F7 (e49). Subplots (a), (c), (e), (g), (i), and (k) represent the gamma band power ratios for nouns, verbs, concrete words, and abstract words, while subplots (b), (d), (f), (h), (j), and (l) represent the gamma band power ratios for negative, neutral, and positive emotion words.
[0111] Depend on Figure 10 It was found that the theta / β band power ratio exhibited significant differences across different brain regions, particularly in distinguishing between nouns and verbs, and between concrete and abstract words. In these comparisons, the theta / β band power ratios in multiple brain regions showed distinct patterns, reflecting differences in neural activity when processing these word classes. Particularly at the T7 (e47) and FC5 (e32) sites, the differences in the theta / β band power ratios were particularly significant when comparing words with negative, neutral, and positive emotional valences. This indicates that the theta / β band power ratio is not only effective in part-of-speech classification but can also, to some extent, distinguish words with different emotional valences, providing important evidence for language decoding and emotion recognition.
[0112] Figure 11The α / β band power ratios were shown across different brain regions. The α / β band power ratio exhibited significant differences in distinguishing between nouns and verbs, as well as between concrete and abstract words, particularly in the T7 (e47) and FT7 (e48) brain regions, where the variations were especially pronounced. However, the α / β band power ratio was relatively weaker in distinguishing emotion words. Especially for negative and neutral emotion words, the significant differences across fewer recording sites made these two categories more difficult to differentiate using the α / β band power ratio. Overall, both the θ / β and α / β band power ratios were effective in distinguishing different word categories, but their effectiveness was not significant across different brain regions when dealing with emotion words, especially in differentiating between neutral and positive emotion words. This suggests that these two band power ratios have different sensitivities and applicability in different aspects of language processing, and may require the combination of other neural markers or methods to further improve the accuracy of emotion word classification.
[0113] By analyzing and comparing the power ratios of different brain regions individually, we can analyze the differences in how different brain regions respond to word part-of-speech categories. This can explain the significant neural activity characteristics of brain regions when processing specific novel words, and construct a more refined and accurate language decoding model to better capture the dynamic processes of different brain regions in language processing, thereby improving the accuracy and efficiency of language decoding.
[0114] Furthermore, the differences between brain regions in comparisons of nouns, verbs, concrete words, and abstract words were greater than the differences in comparisons of emotional words.
[0115] Furthermore, a speech decoding and analysis method based on electroencephalogram (EEG) signals also includes the following steps:
[0116] Part-of-speech tagging of words in a corpus is performed using natural language processing techniques;
[0117] The accuracy of the power comparison word part-of-speech classification is calculated based on the part-of-speech tagging.
[0118] In addition, the results of the power comparison-based word part-of-speech classification of the present invention are compared with the classification results of Small World Propensity (SWP) to analyze the difference in accuracy between the two.
[0119] Part-of-speech (POS) classification of words was performed based on power ratio analysis and support vector machines. During the classification process, a word's POS category was selected as the target category at each step, and a comparison category was chosen from the remaining word POS categories. To evaluate the classification performance of the power ratio, the small-world bias (SWP) algorithm was used for comparison, as it can detect global cooperation between brain regions.
[0120] like Figure 12 and Figure 13 As shown, in terms of overall classification accuracy and F1 score, power ratio analysis outperforms the SWP algorithm in classifying all word part-of-speech categories, especially sentiment words. Although the accuracy of part-of-speech classification varies under different power ratio metrics, negative sentiment words show the highest classification accuracy across all three power ratio metrics. Figure 12 As shown, abstract words have the lowest classification accuracy in the θ / β band power ratio, while neutral emotion words show the lowest classification accuracy in both the α / β band power ratio and the gamma band power ratio.
[0121] SWP analysis showed that the classification accuracy for all word parts of speech categories was low for all subjects. Figure 13 This further highlights the superiority of power ratio analysis, which achieves a higher F1 score compared to the SWP algorithm, especially for nouns and verbs. These results demonstrate that power ratio analysis is more reliable than the SWP algorithm in distinguishing different word categories, particularly in terms of sentiment words and abstract words.
[0122] The results showed that the power ratio analysis consistently outperformed the SWP analysis, particularly in classifying negative emotion words. For example, the power ratio achieved an accuracy of over 75% in classifying negative emotion words, while SWP only reached approximately 52%. However, both methods faced challenges with neutral emotion words. The power ratio demonstrated moderate success (65%–70% accuracy) across most word categories, compared to SWP's accuracy of less than 60% across all word categories for all subjects.
[0123] This invention further performs statistical analysis on the performance of different power ratios in various word classification tasks:
[0124] Gamma-band power ratio: In the noun vs. verb and concrete vs. abstract word tasks, more than 85% of the 61 electrode sites showed significant differences (p<0.01), indicating that it is suitable for distinguishing nouns, verbs, concrete words, and abstract words;
[0125] θ / β band power ratio: 55 sites were significant in the noun vs. verb task (p<0.05), but only 7 sites were significant in the concrete word vs. abstract word task, which verifies that it is more suitable for distinguishing between nouns and verbs;
[0126] α / β band power ratio: 42 sites were significant in the concrete word vs. abstract word task (p<0.01), while only 15 sites were significant in the noun vs. verb task, indicating that it is more suitable for distinguishing between concreteness and abstractness.
[0127] Therefore, power ratio analysis is a powerful biomarker for distinguishing various word part-of-speech categories. It is particularly effective in classifying negative emotion words, nouns, and verbs. Although the SWP algorithm shows promise, it performs worse than power ratio analysis for emotion words and abstract words. These findings highlight power ratio analysis as a powerful tool for understanding brain language processing, providing important insights for optimizing language decoding models and improving brain-computer interface systems.
[0128] This invention also provides a speech decoding system based on electroencephalogram (EEG) signals, which applies the above-mentioned speech decoding and analysis method based on EEG signals, including:
[0129] The corpus module is used to provide continuous natural language streaming as natural stimuli;
[0130] Multi-channel EEG device: used to collect multi-channel EEG signals generated by the subject when receiving the natural stimulation;
[0131] The signal segmentation module is connected to the corpus module and the multi-channel EEG device respectively. It is used to receive the EEG signals collected by the multi-channel EEG device and segment the EEG signals into short-time EEG signals corresponding to a single word in the corpus based on the timestamp.
[0132] A frequency domain conversion module is used to receive the short-time EEG signal and convert the short-time EEG signal into a frequency domain power spectral density.
[0133] The calculation and analysis module is used to calculate the power ratio of the short-time EEG signal based on the power spectral density, and to perform semantic category classification of the words in the corpus based on the power ratio, and output the language decoding results.
[0134] This invention also provides an application of a speech decoding and analysis method based on electroencephalogram (EEG) signals in brain-computer interfaces (BCIs). In BCI applications, by identifying brain regions associated with specific part-of-speech categories, more effective signal processing and classification algorithms can be designed, thereby improving the system's understanding and response to the user's linguistic intent. Combining power ratio analysis of different brain regions with other neuroimaging techniques (such as fMRI and MEG) enables the fusion of multimodal data, providing a more comprehensive understanding of the neural mechanisms of language processing. This offers new tools and perspectives for studying the functional division and network collaboration of the brain in language processing, contributing to a deeper understanding of the neural basis of language cognition.
[0135] The present invention also provides a computer device that is programmed or configured to perform the above-described speech decoding and analysis method based on electroencephalogram (EEG) signals, or the computer device has a computer program stored in its memory that is programmed or configured to perform the above-described speech decoding and analysis method based on EEG signals.
[0136] In summary, EEG signals induced by different types of words show significant differences in power ratio. The α / β band power ratio, θ / β band power ratio, and γ band power ratio (the ratio of γ band power to the total EEG signal power) can effectively distinguish between nouns and verbs, concrete words and abstract words, and also show differences between negative, positive, and neutral emotional words. Based on the α / β band power ratio, the classification accuracy for negative emotional words is the highest (72%), while the accuracy for neutral emotional words is the lowest (only 54%). Power ratio analysis shows significantly higher classification accuracy for concrete words than for abstract words (p<0.05), while nouns are easier to distinguish than verbs in terms of γ power ratio. From an emotional perspective, the classification accuracy for negative emotional words is the highest, the classification accuracy for positive emotional words is between 60% and 70%, and the classification accuracy for neutral emotional words is the lowest. This invention utilizes time-frequency conversion analysis to analyze the power ratio characteristics of short-time EEG time series corresponding to individual words, transforming EEG-related information from the time domain to the frequency domain. This provides new neural markers for language cognitive processing mechanisms. Future research should further optimize the emotion word classification model and explore multimodal feature fusion strategies to promote the development and application of speech brain-computer interfaces. Through power ratio analysis, isolated short EEG signals reflect significant differences between different word part-of-speech categories. Furthermore, by combining power ratio analysis with support vector machines, word classification was performed for each subject, achieving a maximum accuracy of 80%, demonstrating the potential of power ratio analysis in future speech brain-computer interfaces. Power ratio provides a new analytical tool for quantifying future language processing. That is, power ratio can serve as a reliable biomarker to distinguish different word part-of-speech categories, such as nouns and verbs, concrete and abstract words, and different emotion words (positive, negative, and neutral). By applying power ratio analysis to short EEG signals under natural stimuli, the neural characteristics of different word part-of-speech categories can be revealed, overcoming the shortcomings of traditional experimental settings using artificial words or phrases. Using power ratio as a biomarker, this method categorizes words in a corpus semantically, quantifies part-of-speech differences, and enables language decoding. This solves the challenge of extracting EEG signals from individual word classes under natural stimuli and allows for more accurate identification and differentiation of EEG features corresponding to different word classes. It also allows for in-depth exploration of neural activity patterns in the brain when processing different word classes, providing richer information and deeper insights into the brain's language processing mechanisms. This leads to more effective brain-computer interface language decoding, improving the accuracy and reliability of decoding. For patients with brain disorders, this analytical method helps to accurately locate neural abnormalities in their language processing, providing a basis for developing targeted rehabilitation training programs, thereby improving patients' language communication abilities, enhancing their quality of life, and increasing their social participation.
[0137] Furthermore, this invention, based on a corpus of natural stimuli, enables the brain-computer interface system to be trained and applied in scenarios more closely resembling real language use, improving the system's adaptability and flexibility to natural language. Through interdisciplinary methods and signal processing technology innovations in linguistics, neuroscience, and natural language processing, this invention achieves precise decoding of different parts of speech of words in natural language environments using frequency domain features of short EEG signals.
[0138] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A speech decoding and analysis method based on electroencephalogram (EEG) signals, characterized in that, Includes the following steps: A corpus is selected as the natural stimulus, and the corpus is a continuous natural language streaming media; Acquire multichannel EEG signals from subjects under the natural stimuli; Based on timestamps, the multi-channel EEG signals are divided into short-time EEG signals corresponding to individual words in the corpus; The short-time EEG signal is converted into a frequency domain power spectral density, and the power ratio of the short-time EEG signal is calculated based on the power spectral density. Based on the power ratio, the words in the corpus are classified into semantic categories to achieve language decoding.
2. The speech decoding and analysis method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, The method for converting the short-time EEG signal into a frequency domain power spectral density and calculating the power ratio of the short-time EEG signal based on the power spectral density includes: For each short-time EEG signal segment, a part-of-speech-signal length dual-linkage adaptive Kaiser window combined with discrete Fourier transform is used to estimate the power spectral density and calculate the absolute power of the signal in the δ, θ, α, β, and γ sub-bands. Based on the absolute power, the power ratio is calculated as a biomarker, and the power ratio includes at least one of the gamma band power ratio, α / β band power ratio, and θ / β band power ratio.
3. The speech decoding and analysis method based on electroencephalogram (EEG) signals according to claim 2, characterized in that, The shape parameter β of the Kaiser window is dynamically adjusted according to the part of speech of the word and the length of the short-time EEG signal.
4. The speech decoding and analysis method based on electroencephalogram (EEG) signals according to claim 3, characterized in that, The shape parameter β of the Kaiser window is set based on the initial parameter setting of the part-of-speech category and dynamically corrected based on the short-time EEG signal length.
5. The speech decoding and analysis method based on electroencephalogram (EEG) signals according to claim 2, characterized in that, The parts of speech of the words include nouns, verbs, concrete words, abstract words, or emotional words; wherein, the power ratio and the semantic category of the words have the following mapping relationship: The gamma band power ratio is applicable to distinguishing between nouns, verbs, abstract words, and concrete words; The θ / β band power ratio is suitable for distinguishing between nouns and verbs; The α / β band power ratio is used to distinguish between abstract and concrete words.
6. The speech decoding and analysis method based on electroencephalogram (EEG) signals according to claim 2, characterized in that, The method for semantically classifying words in the corpus based on the power ratio to achieve language decoding includes: At least two power ratios are used as joint feature vectors and input into a classification model to classify the semantic categories of words in the corpus. The semantic categories include at least one of part-of-speech category, specificity category, or emotional valence category, so as to achieve decoding from EEG signals to word semantic categories. The classification model is at least one of support vector machine, random forest, or neural network.
7. The speech decoding and analysis method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, The corpus includes at least one of audiobooks, podcasts, audio, video, and text.
8. The speech decoding and analysis method based on electroencephalogram (EEG) signals according to claim 1, characterized in that, It also includes the following steps: We analyzed and compared the power ratios of different brain regions separately to analyze the differences in the response of different brain regions to word part-of-speech categories.
9. A speech decoding system based on electroencephalogram (EEG) signals, employing the speech decoding and analysis method based on EEG signals as described in any one of claims 1 to 8, characterized in that, include: The corpus module is used to provide continuous natural language streaming as natural stimuli; Multi-channel EEG device: used to collect multi-channel EEG signals generated by the subject when receiving the natural stimulation; The signal segmentation module is connected to the corpus module and the multi-channel EEG device respectively. It is used to receive the EEG signals collected by the multi-channel EEG device and segment the EEG signals into short-time EEG signals corresponding to a single word in the corpus based on the timestamp. A frequency domain conversion module is used to receive the short-time EEG signal and convert the short-time EEG signal into a frequency domain power spectral density. The calculation and analysis module is used to calculate the power ratio of the short-time EEG signal based on the power spectral density, and to perform semantic category classification of the words in the corpus based on the power ratio, and output the language decoding results.
10. A computer device, characterized in that, The computer device is programmed or configured to perform the speech decoding and analysis method based on electroencephalogram (EEG) signals as described in any one of claims 1 to 8, or the computer device has a computer program stored in its memory that is programmed or configured to perform the speech decoding and analysis method based on EEG signals as described in any one of claims 1 to 8.