A Non-invasive EEG Imaginary Language Decoding Method and System Based on Multidimensional Optical Flow Vectors
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]但现有研究仍存在明显的局限性:(1)多数研究仅简单套用图像领域的光流估计流程,光流估计的核心假设与EEG数据的适配性未得到系统验证,导致提取的光流特征缺乏生理合理性;(2)现有研究仅利用光流矢量数据作可视化分析,未将想象语言过程中脑电活动的动态传播与聚集规律,应用到具体的解码任务中;(3)现有研究未充分利用多频段EEG的互补特性,简单复用现有EEG特征集或单一频段,缺乏跨节律动力学特征的融合与构建,导致特征的鲁棒性与判别能力不足,难以满足想象语言精准解码的实际需求
首次系统完成光流三大假设与 EEG 数据的生理适配性验证,建立可量化判定体系,为光流方法在脑电分析中的应用提供理论与实验支撑;
Smart Images

Figure CN122413170B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of non-invasive EEG signal processing and imaginative language decoding technology, specifically involving a non-invasive EEG imaginative language decoding method and system based on multi-dimensional optical flow vectors. Background Technology
[0002] Imaginary language is the process by which an individual performs semantic representation, vocabulary activation, and grammatical construction within the brain without external language output. It holds significant theoretical value and application prospects in fields such as brain-computer interfaces (BCI), neurorehabilitation, and cognitive science. Non-invasive electroencephalography (EEG), with its advantages of high temporal resolution, ease of operation, and non-invasiveness, has become the mainstream technique for capturing brain electrical activity related to imaginary language and decoding cognitive states. It has been widely applied in feature extraction and recognition research of imaginary language.
[0003] Current research on EEG-based imaginary language decoding largely focuses on traditional feature extraction methods, primarily revolving around the time, frequency, and spatial characteristics of EEG signals, such as power spectral density, wavelet coefficients, and entropy. Statistical tests, including t-tests, ANOVA, mutual information, principal component analysis, and linear discriminant analysis, are used to explore the correlation between these static features and the cognitive process of imaginary language, thus distinguishing different categories of imaginary language. However, the EEG activity corresponding to imaginary language is a dynamic evolutionary process with significant spatiotemporal continuity and propagation characteristics. Traditional static features can only capture local energy distributions or single-dimensional information of EEG signals, failing to comprehensively depict the dynamic spatiotemporal evolution of EEG activity during imaginary language processing. This results in decoding accuracy and robustness that cannot meet practical application requirements, becoming a core bottleneck restricting the development of imaginary language decoding technology.
[0004] To overcome the limitations of traditional static features and fully capture the dynamic spatiotemporal characteristics of imaginative language EEG activity, researchers have attempted to transfer optical flow estimation methods from computer vision to EEG signal processing, conducting preliminary explorations. Existing research primarily applies optical flow estimation to dynamic EEG feature extraction, mostly focusing on characterizing motion information in a single-band EEG amplitude field. By calculating optical flow vectors between adjacent frames, these are applied to feature enhancement for EEG-related tasks such as motor imagery and emotion recognition. Some studies also explore the correlation between extracting optical flow features from EEG activity and cognitive brain activity.
[0005] However, existing research still has obvious limitations: (1) Most studies simply apply the optical flow estimation process in the field of image processing. The core assumptions of optical flow estimation and the compatibility with EEG data have not been systematically verified, resulting in the lack of physiological rationality of the extracted optical flow features; (2) Existing research only uses optical flow vector data for visualization analysis, without applying the dynamic propagation and aggregation laws of EEG activity in the process of imaginary language to specific decoding tasks; (3) Existing research does not make full use of the complementary characteristics of multi-band EEG, simply reuses existing EEG feature sets or single bands, lacks the fusion and construction of cross-rhythmic dynamic features, resulting in insufficient robustness and discriminative ability of features, making it difficult to meet the actual needs of accurate decoding of imaginary language. Summary of the Invention
[0006] To overcome the shortcomings of existing technologies, such as the lack of spatiotemporal dynamic features of EEG, the lack of physiological adaptability verification of optical flow methods, and the inefficient use of multi-rhythm information, this invention provides a non-invasive EEG-based method for decoding imaginary language based on multi-dimensional optical flow vectors. Through optical flow adaptability verification, EEG current field construction, source-sink identification, trajectory tracking, dynamic rhythm screening, and random forest classification, this method achieves high-precision decoding of imaginary language.
[0007] The method is specifically as follows: S1. Acquire non-invasive EEG signals, filter the signals, perform Hilbert transform, sliding window amplitude extraction and frequency band specific Kriging interpolation to generate multi-rhythm EEG amplitude field sequences; S2. Quantitatively verify the three core assumptions of constant brightness, spatial consistency and temporal continuity in optical flow estimation, and determine the compatibility between the EEG amplitude field and the optical flow method. S3. Based on the adaptability results, the CLG optical flow algorithm is used to calculate the optical flow vector field of the EEG amplitude field, and the source and sink points are identified by divergence, determinant and Poincaré index; S4. Track the source and sink points to extract 18-dimensional single-rhythmic optical flow features, including optical flow velocity and direction, source and sink point ratio, and trajectory length and offset, to form multi-rhythmic optical flow features; S5. Dynamic weight allocation is performed on the multi-rhythmic optical flow features through the task-aware gating network TAG-Net, and the optimal rhythm combination is selected and fused to obtain multi-dimensional optical flow features. S6. Input the multidimensional optical flow features into a random forest classifier that has undergone hyperparameter optimization to complete the classification and decoding of the imagined language.
[0008] Furthermore, in step S2, the quantification verification includes: The constant brightness assumption is verified by phase continuity, the spatial consistency assumption is verified by local stationarity, the temporal continuity assumption is verified by inter-frame velocity variation, and a comprehensive fit score is calculated. When the overall fit score is ≥0.7, the EEG data is deemed to meet the application conditions for optical flow estimation.
[0009] Furthermore, the source and sink points include the source point and the sink point. In step S3, the source and sink point identification satisfies the following conditions: Valid source points: divergence div > 0, determinant det > 0, Poincaré exponent close to 1; Valid sinks: divergence < 0, determinant det > 0, Poincaré exponent close to -1; Furthermore, noise fluctuations are filtered out using a divergence threshold to improve recognition robustness.
[0010] Furthermore, in step S4, the trajectory tracking employs a four-fold boundary constraint strategy, including: Initial point boundary constraints, velocity query boundary constraints, ODE solver hard boundary constraints, and trajectory result secondary truncation constraints ensure that the trajectory coordinates do not exceed the limits.
[0011] Furthermore, in step S4, the 18-dimensional optical flow features include: Average velocity, peak velocity, velocity standard deviation, velocity median, directional mode, directional standard deviation, directional absolute mean, principal direction proportion, average trajectory length, maximum trajectory length, average trajectory offset, source center normalized coordinates, sink center normalized coordinates, source proportion, sink proportion, velocity spatiotemporal average, directional spatiotemporal average, and optical flow activity.
[0012] Furthermore, in step S5, the task-aware gating network TAG-Net includes: Global average pooling is performed on the 18-dimensional optical flow features of each rhythm, and the features are concatenated to obtain cross-rhythm features; Task conditional encoding is introduced, and the weights of each rhythm are obtained through a fully connected layer and the Softmax function; The top-k rhythms are selected as the optimal combination based on their weights to achieve weighted fusion of multiple rhythm features.
[0013] Furthermore, in step S6, the random forest uses grid search and 5-fold hierarchical cross-validation to optimize hyperparameters, including: The parameters include the number of decision trees, the maximum depth of the decision tree, the minimum number of split samples per node, and the class weights.
[0014] Another aspect of the present invention provides a non-invasive EEG imagery language decoding system based on multidimensional optical flow vectors, the system comprising: Data preprocessing module: Acquires non-invasive EEG signals, performs filtering, Hilbert transform, sliding window amplitude extraction, and frequency band-specific Kriging interpolation on the signals to generate multi-rhythm EEG amplitude field sequences; Optical flow adaptability verification module: Quantitatively verifies the three core assumptions of constant brightness, spatial consistency and temporal continuity of optical flow estimation, and determines the adaptability of EEG amplitude field to optical flow method; Optical flow vector field calculation module: Based on the adaptation results, the CLG optical flow algorithm is used to calculate the optical flow vector field of the EEG amplitude field, and the source and sink points are identified by divergence, determinant and Poincaré index; Multi-rhythmic optical flow feature extraction module: Tracks the source and sink points to extract 18-dimensional single-rhythmic optical flow features, including optical flow velocity and direction, source and sink point ratio, and trajectory length and offset, to form multi-rhythmic optical flow features; Dynamic rhythm selection module: The task-aware gating network TAG-Net dynamically assigns weights to multi-rhythmic optical flow features, selects the optimal rhythm combination, and fuses them to obtain multi-dimensional optical flow features; Classification and Decoding Module: Inputs multidimensional optical flow features into a random forest classifier that has undergone hyperparameter optimization to complete the classification and decoding of the imagined language.
[0015] The beneficial effects of the method described in this invention are as follows: For the first time, the physiological compatibility of the three major hypotheses of optical flow with EEG data was verified systematically, and a quantifiable judgment system was established, providing theoretical and experimental support for the application of optical flow methods in EEG analysis. Breaking through the bottleneck of traditional static EEG features, it extracts multi-dimensional spatiotemporal dynamic features such as velocity, direction, source and sink, and trajectory, and realizes a refined quantitative characterization of the propagation, convergence, and evolution of EEG activity. A multi-rhythmic optical flow feature engineering system with 18 dimensions in a single frequency band and 54 dimensions across frequency bands was constructed. Combined with TAG-Net to dynamically select the optimal rhythm combination, the feature discrimination power and decoding robustness were significantly improved. Achieving state-of-the-art (SOTA) performance on the InnerSpeech2021 and Bimodal InnerSpeech2023 datasets, with average accuracy of 80.04% and 70.84% respectively on the Inner task, it can provide key technical support for aphasia rehabilitation and practical brain-computer interfaces. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method described in an embodiment of the present invention. Detailed Implementation
[0017] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0018] This embodiment provides a non-invasive EEG-based method for decoding imaginary language using multidimensional optical flow vectors. This method addresses core issues in current EEG-based imaginary language decoding research, such as the inability of traditional static features to characterize the spatiotemporal dynamic evolution of EEG activity, the lack of physiological adaptability and deep feature mining in optical flow estimation methods, and the insufficient utilization of multi-band information. First, the adaptability of the optical flow estimation hypothesis to EEG data is verified by modeling the EEG amplitude field as a describable, flowing spatiotemporal process. Then, an EEG optical flow vector field is constructed, and multidimensional dynamic features such as source-sink points, propagation trajectories, and velocity directions are extracted. Simultaneously, multi-rhythmic characteristics of Delta, Theta, Alpha, Beta, and Gamma are combined to form a cross-band multidimensional optical flow feature fusion system. Finally, the aforementioned multi-rhythmic optical flow vector features are input into the decoding model to achieve high-precision recognition of imaginary language. The overall technical route is as follows: Figure 1 As shown: a. Optical flow estimation hypothesis and EEG compatibility verification: Based on the CLG model, the three hypotheses of optical flow estimation are verified, and optical flow vector fields between adjacent EEG amplitude fields under different rhythms are generated. b. Source and sink identification: Based on the divergence, determinant, and Poincaré exponent features of the optical flow field, source and sink points in the EEG optical flow field are identified, characterizing the spatial aggregation and diffusion patterns of EEG activity. c. Dynamic trajectory tracking: The dynamic evolution path of EEG activity is tracked through the time series of optical flow vectors to capture its spatiotemporal continuity characteristics. d. Multidimensional optical flow vector representation: Five types of dynamic features are extracted from the dimensions of velocity, direction, trajectory, source and sink structure, and spatial distribution to construct multi-scale optical flow feature vectors across rhythms. e. Imaginary language decoding: Based on the TAG-Net model selected by dynamic rhythm and a random forest classifier, accurate decoding of EEG signals under the imaginary language task is achieved.
[0019] 1. Dataset and Preprocessing: (1) Dataset: The experimental datasets in this embodiment use two public datasets, InnerSpeech2021 and BimodalInnerSpeech2023 (abbreviated as BInnerSpeech2023), as experimental data support.
[0020] InnerSpeech 2021. Data were acquired using a Biosemi 128-channel system, with electrode placement following a standard 10⁻⁵ system and a sampling rate of 254 Hz. The dataset included 10 subjects (all without a history of neurological disease or language impairment, and who signed informed consent forms before the experiment; the experimental procedure complied with ethical review requirements) who completed EEG acquisition tasks under three core task conditions: “Pronounced Speech (Pron)”, “Imagined Speech (Inner)”, and “Visualized Condition (Vis)”. Each task condition was further subdivided into four directional stimuli (“Up”, “Down”, “Left”, and “Right”). The total duration of a single session was 4 seconds (1 second fixed + 2.5 seconds task + 1 second rest). Referring to the original text, after removing incorrect trials, Pron had 25 trials per word category, Inner had 25 trials per word category, and Vis had 25 trials per word category, for a total of 10 people × (25+25+25) trials × 4 words = 3000 data points.
[0021] Bimodal InnerSpeech 2023. Data was acquired using a Biosemi 64-channel system, with electrode placement following a standard 10⁻⁵ system and a sampling rate of 512 Hz. The dataset included 4 subjects (all without a history of neurological disease or language impairment, and who signed informed consent forms before the experiment; the experimental procedure complied with ethical review requirements), who completed EEG and fMRI acquisition tasks. This embodiment only used the EEG dataset, with 8 words in two categories of stimuli under the Inner task conditions: social (child, daughter, father, wife) and numerical (four, three, ten, six)). The total duration of a single session trial was 4 seconds (1 second fixed + 2 seconds task + 1 second rest), with 25 trials per word, totaling 4 subjects × 25 trials × 8 words = 800 data points.
[0022] (2) Pretreatment: The raw InnerSpeech2021 EEG data were preprocessed. The preprocessing followed the data collection parameter settings set by the dataset authors to ensure data originality and comparability. The difference lay in extracting event-related potentials (ERPs) under different conditions (e.g., Pron, Inner, Vis) and different categories (e.g., up, down, left, right) for each participant. Each trial's data was then extracted into a time window, from 1 second to 3.5 seconds after stimulus presentation (i.e.,...). t min = 1, t max= 3.5). Based on the findings of neuroscience research, five core functional rhythms were selected. b ={δ wave (1–4 Hz); θ wave (4–8 Hz); α wave (8–13 Hz); β wave (13–30 Hz); γ wave (30–60 Hz)}.
[0023] The raw BInnerSpeech2023 EEG data underwent preprocessing. The preprocessing followed the data collection parameter settings set by the dataset authors to ensure data originality and comparability. The difference lay in the extraction of event-related potentials (ERPs) for each participant under different categories (social and numerical) and different words (e.g., child, daughter, father, wife, four, three, ten, six). Each trial's data was then segmented into a time window, from 1 second to 3 seconds after stimulus presentation. t min = 1, t max = 3). Based on the research findings and characteristics of data in neuroscience, five core functional rhythms were selected: delta wave (1–4 Hz); theta wave (4–8 Hz); alpha wave (8–13 Hz); beta wave (13–30 Hz); and gamma wave (30–40 Hz).
[0024] For each rhythm in the two datasets mentioned above, bandpass filtering was performed using the `mne.filter.filter_data` function from the MNE-Python library. An FIR (Finite Impulse Response) filter was then used to extract amplitude information from the filtered signal. A Hilbert transform was applied to the signal of each rhythm, converting a real-valued signal into a complex-valued analytic signal. The amplitude of the analytic signal corresponds to the instantaneous amplitude of the original signal. The phase value corresponds to the instantaneous phase of the original signal. Using a sliding window of length 20 ms (win_length = 0.02) in the time dimension, the average amplitude of the signal within each window is calculated, transforming the continuous time series into a series of discrete time frames, each representing the average state of brain activity within that window.
[0025] To ensure the interpolation results better reflect the actual brain shape and the needs of data analysis, this embodiment uses the international 10-20 system to set the two-dimensional coordinates of each electrode as follows: Frequency band specific Kriging interpolation was performed on the circular EEG topography map. The interpolated result was then processed by rectangular masking. A 64×64 rectangular region was defined based on the distribution range of the electrodes, and the interpolation result outside the region was set to zero.
[0026] Finally, to avoid a sudden drop in amplitude at the edges of the rectangle, the distance from each point outside the edge to the nearest edge point was calculated, and the amplitude was linearly attenuated based on this distance. The discrete... and Convert to multi-rhythm amplitude field and phase field ( For frame numbering, Two-dimensional spatial coordinates of the reconstructed spatial field after interpolation , The EEG electrode channels (original discrete acquisition points) form the input basis for subsequent optical flow calculations.
[0027] 2. Compatibility verification of optical flow estimation assumptions with EEG data Optical flow estimation, a classic method in computer vision for characterizing spatiotemporal motion sequences, is based on three core assumptions: brightness constancy assumption (BCA), spatial consistency, and temporal continuity (micromotion). However, EEG amplitude topography, reflecting spatiotemporal sequences of brain activity, differs significantly from traditional visual image data in signal nature, evolutionary patterns, and noise characteristics. To verify the suitability of the classic assumptions of optical flow estimation in EEG multi-rhythm amplitude field analysis, a quantifiable EEG-optical flow assumption suitability assessment system is established. Simultaneously, this provides a theoretical basis and engineering verification scheme for the rational application of optical flow methods in EEG feature decoding of imaginary language.
[0028] (1) Phase continuity verifies the constant brightness assumption The core of optical flow estimation is the brightness constancy assumption, which states that the gray values at the same spatial location in an image sequence remain stable over time. The grayscale features at this location do not change abruptly. The optical flow velocity at a point in space. Spatial location in an image sequence Place (Gray value at time). In the process of applying optical flow estimation methods to EEG signal analysis, this embodiment uses... The spatiotemporal continuity corresponds to the stability characteristics of gray values in visual images. By ensuring the spatiotemporal smoothness of EEG activity through the continuous evolution characteristics of phase, the assumption of constant brightness can be adapted and verified in the EEG signal analysis scenario.
[0029] First, the original phase sequence Normalize to [0, 2π] to eliminate dimensional effects. ( For frame index, For pixel coordinates, The normalized phase value (unit: rad). Represents the original phase field The maximum value across all pixels and all frames, that is, the global maximum amplitude of the original phase in the entire set of data.
[0030] Then, to eliminate the pseudo-jumps caused by phase periodicity, the phase difference between adjacent frames is calculated and wrapped in [ The inter-frame phase jump value Δ is obtained from the interval [π, π]. (Unit: rad): ; In the formula, Δ This represents the inter-frame phase jump value (unit: rad), and mod represents the modulo operation.
[0031] Next, the phase discontinuity (PD) is quantized. PD is the sum of the absolute differences between the phase values of all pixels within a single frame and the mean of their 8-neighborhood, divided by the total number of pixels, representing the global phase discontinuity. ; In the formula, Represents pixels The 8-neighborhood phase mean, This represents the total number of pixels. This represents the vertical pixel height of a single frame image. This represents the horizontal pixel width of a single frame of an image. This refers to the PD value of a single frame. The average PD value across all frames. This indicates the total number of frames.
[0032] Finally, define the formula for determining whether the constant brightness assumption is satisfied: ; in, Indicates rhythm The corresponding global average phase discontinuity, Indicates rhythm The absolute value of the phase change between frames. (Subscript) It is used to distinguish different brainwave rhythms (δ wave (1–4 Hz); θ wave (4–8 Hz); α wave (8–13 Hz); β wave (13–30 Hz); γ wave (30–40 Hz)) to realize the assumption of constant brightness by determining the rhythm separately.
[0033] In the determination formula, when the phase jump threshold is 0.9π and the percentage of pixels with phase jumps exceeding the threshold is less than 5%, the constant brightness assumption is satisfied. The phase synchronization of the visual cortical neuron group to optical flow stimulation is manifested as continuous phase in scalp EEG, corresponding to "stable brightness pattern" in the optical flow method; phase decoupling corresponds to "abrupt brightness," directly reflecting the effectiveness of the constant brightness assumption in EEG signal analysis.
[0034] (2) Local stationarity verifies the spatial consistency hypothesis The spatial consistency assumption is that pixels within a local neighborhood of an image (any pixel) and They have the same motion trend (consistent optical flow velocity). ( For the local neighborhood of the image, , (where is the optical flow velocity between two points in the neighborhood), satisfying ( (The speed difference is a very small positive number.)
[0035] The spatial consistency assumption of the optical flow algorithm is transformed into a local stationarity verification in EEG data. This embodiment reshapes the EEG amplitude sequence into... Dimensions ( For frame number, (Total number of pixels per frame), set the default window size to 20 frames and the step size to 5 frames, and perform sliding window truncation on the amplitude sequence of each channel.
[0036] First, the EEG topography map was divided into 8 local regions covering the whole brain: (0,H / 2,0,W / 2), (0,H / 2,W / 2,W), (H / 2,H,0,W / 2), (H / 2,H,W / 2,W), (H / 4,3H / 4,0,W / 3), (H / 4,3H / 4,W / 3,2W / 3), (0,H / 3,W / 4,3W / 4), and (2H / 3,H,W / 4,3W / 4).
[0037] To eliminate the influence of the amplitude mean, the normalized variance of 8 local regions is calculated for each time window (default 20 frames): ; In the formula, For EEG amplitude sequences, For the first A local area; For the first The first window, the Normalized variance of each region.
[0038] Then, the spatial stationarity (LS) of the variance of each local region is obtained; the closer it is to 1, the better the spatial stationarity. ; ; In the formula, The total number of time windows. For average local correlation, It is a minimal constant ( To prevent the denominator from being 0 and avoid division by zero errors.
[0039] Finally, we define a formula for determining whether the spatial consistency assumption is satisfied: ; In the determination formula, when Percentage <1%, and Furthermore, the spatial consistency assumption is satisfied.
[0040] (3) Verification of the temporal continuity assumption by the change in velocity between frames The temporal continuity assumption in optical flow estimation requires that pixel motion displacement be small within the time interval between adjacent frames, ensuring that grayscale value changes can be approximated by Taylor expansion. This embodiment verifies this assumption by calculating the motion velocity quantization of the EEG amplitude sequence and determining the proportion of small motions.
[0041] The horizontal velocity field of each frame is calculated using the CLG (Combined Local-Global) optical flow algorithm. ) and vertical velocity field ( The average change in the velocity field between adjacent frames is characterized by calculating RVC (Relative Velocity Change). ; ; In the formula, This represents the total number of pixels in a single frame. For frame number, Indicates the first Frame and the The average change in the velocity field between frames This represents the average change in velocity field across all adjacent frame pairs.
[0042] Finally, define the formula for determining whether the time continuity assumption is satisfied: ; In the determination formula, when Pixels per frame, and percentage of pixels with small motion. A percentage exceeding 70% satisfies the temporal continuity assumption. The percentage of pixels with small motion is obtained as follows: The optical flow vector magnitude is calculated pixel-by-pixel. Pixels with a displacement of less than 2.0 pixels within a single frame are counted as small motion pixels. The percentage of small motion pixels in a single frame is obtained by dividing the number of small motion pixels in a single frame by the total number of pixels in that frame. The average of this percentage is then calculated over all frames. .
[0043] (4) Calculation of comprehensive suitability score This embodiment transforms the verification results of three assumptions—constant brightness, spatial smoothness (spatial consistency), and small motion (temporal continuity)—into a score ranging from 0 to 1. The average of these scores is used to obtain the final comprehensive fitness score (AS), and a threshold is set for fitness determination. Constant brightness score: Spatial consistency score: Time continuity score: Overall score The final overall fit score threshold was set at 0.7, establishing a standardized criterion for experimental verification.
[0044] 3. Source-sink point identification in electroencephalogram (EEG) optical flow field The essence of optical flow estimation is to track the movement trend of feature values in a spatiotemporal sequence. In EEG analysis, this embodiment uses the spatiotemporal sequence of amplitude topographic map as the input of optical flow algorithm, and compares the amplitude to the gray value of visual image in traditional optical flow applications. Then, it uses divergence and Poincaré index to identify the source and sink points in the EEG optical flow field.
[0045] First, CLG optical flow estimation is used as the feature analysis module. The Combined Local-Global (CLG) algorithm combines the advantages of local and global optical flow, and updates the speed by combining local Gaussian filtering and global Gaussian smoothing. The core formula is: , ; In the formula, , The velocity is the locally filtered velocity along the x-axis and y-axis. , The speed after global smoothing; The smoothing coefficient (default 0.04); These are the spatial and temporal partial derivatives, respectively.
[0046] Source and sink points are core features reflecting the "divergence" and "convergence" of brain electrical activity. They are determined by the divergence of the optical flow vector field. The mathematical expression for divergence is: ; In the formula, This is the optical flow velocity vector. A positive divergence indicates that the location is a source point (divergence of brain electrical activity), and a negative divergence indicates that the location is a sink point (convergence of brain electrical activity).
[0047] Further filtering of effective source and sink points is achieved by combining the determinant of the vector field. The formula for calculating the determinant is: ; Valid source points must meet the following requirements , A valid remittance point must meet the following requirements: , .
[0048] To verify the directional consistency of the vector field, the Poincaré index is introduced. The type of vector field is determined by calculating the total change in vector angles along a local closed loop. The mathematical expression is: ; ; In the formula, This represents the number of pixels within the local window. The range of values for the direction angle of motion (unit: °) is [ 180°, 180°], This represents the vector angle difference between adjacent pixels. The Poincaré index is approximately 1 at the source (divergent field), approximately -1 at the sink (convergent field), approximately ±1 at the vortex (spiral field), and approximately 0 at the saddle (hyperbolic field). In an ideal, continuous, noise-free field, the Poincaré index is a strictly integer {0, ±1}, and the average Poincaré index... It reflects the numerical deviation caused by noise and discrete sampling in the actual EEG data, and the value fluctuates around the corresponding integer.
[0049] Combining divergence, determinant, and Poincaré index as triple criteria, the source and sink points in the optical flow vector field are identified. The core judgment logic is as follows: ; In the formula, The formula for determining the source point is as follows: This is the formula for determining the sink; 1 indicates a source / sink, and 0 indicates a non-source / sink. The threshold (default 0.01) filters out minor divergence fluctuations caused by noise, improving the robustness of the discrimination. This is the tolerance threshold for the Poincaré index, which tolerates topological index deviations caused by data noise and discrete sampling errors.
[0050] 4. Dynamic trajectory tracking based on electroencephalogram (EEG) Optical flow analysis only extracts and performs basic statistics (velocity, direction, etc.) of the optical flow vectors in the EEG amplitude field, resulting in discrete vector field data that cannot reflect the spatiotemporal continuity and propagation evolution characteristics of EEG activity. This section uses Ordinary Differential Equation (ODE) initial value problem modeling and RK45 numerical solution to transform the discrete optical flow vector field into a continuous dynamic EEG trajectory, converting static vector data into dynamic trajectory evolution information.
[0051] The optical flow vector field is treated as a velocity field, and the dynamic trajectory of the EEG is tracked using ordinary differential equations with boundary constraints, with the source / sink point as the initial point. The trajectory tracking is transformed into an ODE initialization problem, which is solved using the RK45 numerical method. A velocity field mapping function is constructed, with time as its input. with spatial coordinates The output is the optical flow velocity vector at the corresponding position: ; In the formula, for Pixel coordinates at time [time] This is the optical flow velocity vector at that location.
[0052] To address the coordinate out-of-bounds problem during trajectory solving, a four-fold boundary constraint strategy is proposed, setting the time span to [missing information]. Trajectory sampling was performed at 25 equally spaced evaluation points, using a four-fold boundary constraint strategy: Initial point boundary constraints: The initial point (source point or field center) of trajectory tracking is constrained to the effective coordinate range by a truncation function before being input into the ODE solver; Velocity query boundary restriction: Before querying the optical flow velocity field value, the query coordinates are truncated to the valid area to avoid array index out of bounds; ODE solver boundary constraints: When using solve_ivp to solve initial value problems, hard coordinate boundaries are set directly through the bounds parameter to limit the trajectory range at the solver level; `solve_ivp` is a general-purpose tool provided by the SciPy library in Python, specifically designed to solve time-varying differential equations based on initial states and calculate the trajectory. In this example, it is used for numerical integration when calculating the trajectory.
[0053] `bounds` is not a native parameter of ODE, but a custom boundary constraint parameter that stores the upper and lower limits of the effective spatial coordinates. Hard boundary constraints on trajectory coordinates are achieved by relying on the ODE solver.
[0054] Secondary constraint on trajectory results: The complete trajectory sequence output by the ODE solution is truncated again to achieve double protection of boundary constraints.
[0055] The effective coordinate range is defined as follows: , ( , (These represent the spatial height and width of the EEG amplitude field, respectively). Through the above four constraints, the coordinate out-of-bounds problem in trajectory tracking can be systematically avoided, ensuring the stable solution and physical validity of the spatiotemporal trajectory of EEG activity.
[0056] 5. Multidimensional Optical Flow Vector Characterization of EEG Using multi-rhythm EEG data processed by optical flow analysis as input, the system sequentially goes through two core steps: multi-rhythm feature extraction and multi-dimensional feature fusion, ultimately achieving the classification and recognition task related to EEG signals, ensuring the robustness and effectiveness of the features throughout the process.
[0057] (1) Calculation of optical flow velocity and direction Optical flow analysis outputs the x-direction velocity sequence between adjacent frames. and y-direction velocity sequence Based on this, pixel-level motion velocity and motion direction angle are calculated, serving as the basis for feature extraction. Global velocity set. and global direction angle set ( (where is the total number of pixels), and all subsequent statistical features are calculated based on these two sets.
[0058] The motion rate of a single pixel is obtained by calculating the vector magnitude of the horizontal and vertical velocities. The sum of all pixel rates yields the global velocity set. .
[0059] From the arctangent function Converting motion direction angle The global orientation angle set is obtained by summing all pixel orientation angles. .
[0060] Pair the sequences useq and vseq pixel by pixel, calculate the velocity magnitude and motion azimuth angle of each pixel, and store them in S and Θ respectively.
[0061] (2) Calculation of the ratio of source and sink points Source and sink points are the diverging / converging cores of EEG activity in the optical flow vector field, as indicated by source point masking. and remittance mask Calculate the source point ratio and sink point ratio to characterize the spatial clustering characteristics of EEG activity: ; ; In the formula, H and W represent the height and width of the EEG topographic map. , The value range is [0,1].
[0062] (3) Calculation of trajectory length and offset Pixel trajectory is a set of coordinates over time. , for The pixel coordinates at each moment are used to calculate the cumulative trajectory length and total trajectory offset, characterizing the spatiotemporal propagation characteristics of EEG activity. ; ; In the formula, The cumulative length of the trajectory (reflecting the total distance of the motion). This represents the total trajectory offset (reflecting the final displacement of the motion).
[0063] The system comprises 18 dimensions across five categories: velocity statistics, direction statistics, trajectory features, source-sink features, and distribution features. All features undergo outlier handling (null and infinite values are set to 0) to ensure their effectiveness and stability.
[0064] The optical flow vector feature dimensions are shown in Table 1: Table 1:
[0065] 6. Selection of dynamic rhythms Given the significant differences in neural activation rhythms corresponding to different EEG decoding tasks (imagine language, vocal language, and visual imagination), traditional fixed frequency band selection is difficult to adapt to task-specific cognitive patterns. This embodiment constructs a Task-Aware Gating Network (TAG-Net) to achieve adaptive weight allocation and dynamic selection of optimal rhythm combinations across rhythm features, thereby enhancing the task-specificity and discriminative power of decoding features.
[0066] The experiment iterates through the EEG data after optical flow analysis in step 5, recording the five rhythms of α, δ, θ, β, and γ waves in each trial. The characteristic structure of a single rhythm is (T is the number of time frames; H×W is the two-dimensional spatial grid on the flattened cortex; D=18 is the dynamic feature dimension).
[0067] To avoid dimensionality explosion and overfitting caused by directly performing cross-rhythm attention calculations, global average pooling (GAP) is first performed on each rhythm to compress the spatiotemporal features into a fixed-length vector. Then, the global features of the five rhythms are concatenated to obtain the cross-rhythm basic features. A task type condition vector C is introduced to encode the three tasks Pron, Inner, and Vis using one-hot encoding. Task semantic information is injected into the gating network to guide the model to focus on rhythms strongly correlated with the current task. (C represents the EEG decoding task). The conditional vector is fused with cross-rhythm features to obtain the gated input. .Will Two fully connected layers and an activation function are fed into the signal to evaluate the rhythm importance. The first fully connected layer and sigmoid activation function output the original rhythm weights. The second layer of Softmax normalization makes the weights non-negative and sums to 1, resulting in the final rhythm weights. By weight Sort from highest to lowest, and select... Each rhythm is selected as the optimal rhythm combination, and low-contribution and noisy rhythms are discarded. The 18-dimensional optical flow features of the selected rhythms are then weighted and fused to obtain the final multi-rhythm optical flow features used for decoding. , here These refer to the frequency band sequence numbers for delta / theta / alpha / beta / gamma, etc., and have the same meaning as above.
[0068] 7. Imaginary Language Decoding Based on the optimal rhythm group Z obtained in step 6 Top-k Count n, obtain The input data consists of an n×18 dimensional fused feature vector. To eliminate the dimensional differences between dynamic features of optical flow in different dimensions and avoid the interference of inconsistent dimensions on the training efficiency and classification accuracy of the subsequent random forest classifier, a normalizer is used to normalize all input data to a standard normal distribution with a mean of 0 and a standard deviation of 1. To further improve the accuracy and robustness of the imagined language decoding model, a grid search with 5-fold cross-validation (GridSearchCV) is used to optimize the core hyperparameters of the random forest classifier. The optimal parameters are selected based on the classification accuracy, with the highest average cross-validation accuracy as the scoring criterion. Combining the characteristics of small samples and high dimensions of EEG signals with the requirements of the imagined language decoding task, the core hyperparameter search space for the InnerSpeech2021 dataset is pre-defined as follows: number of decision trees (n_estimators): [50, 100, 200]; maximum depth of decision trees (max_depth): [5, 10, None]; minimum number of samples per node split (min_samples_split): [2, 5]; class weight parameter (class_weight) [None, balanced]. For the BInnerSpeech2023 dataset, the core hyperparameter search space is pre-defined as follows: number of decision trees (n_estimators): [50, 100, 200]; maximum depth of decision trees (max_depth): [8, 15, None]; minimum number of samples per node split (min_samples_split): [2, 5]; class weight parameter (class_weight) [None, balanced]. Finally, the test set data that was not used in training is input into this optimal model, and the model outputs the classification result to determine the category to which the test sample belongs.
[0069] 8. Results This embodiment systematically evaluates the proposed method on the InnerSpeech2021 dataset, performing modeling and evaluation based on three tasks (Inner, Pron, and Vis), and comparing it with related research in recent years.
[0070] Table 2 shows that among traditional machine learning methods, the method proposed by Diego Lopez-Bernal et al. (2024) uses KNN and SVM classifiers, with an accuracy of only 56.08% and 59.55%, respectively; the method proposed by Merola et al. (2023) based on quadratic kernel SVM (Support Vector Machine) has an accuracy of only 33.90%; the method proposed by Gasparini et al. (2022) explores various models such as SVM (Support Vector Machine), XGBoost (Extreme Gradient Boosting Tree), LSTM (Long Short-Term Memory Network), and BiLSTM (Bidirectional Long Short-Term Memory Network), but the overall performance is still in the range of 26.20%~31.30%, failing to effectively capture the complex patterns of EEG signals. In terms of deep learning methods, although the 3D CNN (three-dimensional convolutional neural network) model (the method proposed by YFWang et al., (2026)) has improved compared with traditional methods, achieving accuracy of 62.67%, 61.79% and 53.25% in Pron, Inner and Vis tasks respectively, it is still limited by static feature modeling, and does not make full use of the spatiotemporal dynamic information of EEG signals, and its feature expression ability is insufficient.
[0071] The paper title and source corresponding to the method proposed by Diego Lopez-Bernal et al. (2024) are as follows: Lopez-Bernal, D., Balderas, D., Ponce, P.,&Molina, A. (2024a). Exploring inter-trial coherence for inner speech classification in EEG-basedbrain–computer interface. Journal of Neural Engineering, 21(2), Article026048. The paper title and source corresponding to the method proposed by Merola et al. (2023) are as follows: Merola, NR, Venkataswamy, NG,&Imtiaz, MH (2023). Can machinelearning algorithms classify inner speech from EEG brain signals?. In 2023IEEE World AI IoT Congress (AIIoT) (pp. 0466–0470). IEEE. The paper title and source corresponding to the method proposed by Gasparini et al. (2022) are as follows: Gasparini, F., Cazzaniga, E., & Saibene, A. (2022). Inner speech recognition through electroencephalographic signals. CEUR workshopproceedings, 3368, 48–61. The paper title and source corresponding to the method proposed by YF Wang et al. (2026) are as follows: Wang Y, Li Y, Wang J, et al. Sparse optical flow outliers elimination method based on Borda stochastic neighborhood graph. Machine Learning: Science and Technology, 2024, 5(1): 015022. In contrast, the method proposed in this embodiment employs a Random Forest model. Among 10 subjects (sub1~sub10), the average decoding accuracy for the Pron, Inner, and Vis tasks reached 80.04%, 76.26%, and 77.59%, respectively, representing an improvement of over 15% compared to the existing best methods, with some tasks showing improvements exceeding 50%. Even with a mixed modeling setting using data from 10 subjects, the accuracy for all three tasks remained within the range of 69.76%~72.94%, all superior to existing comparative methods. The comparative results demonstrate that modeling EEG signals as a spatiotemporal dynamic field and extracting multidimensional optical flow statistical features can more fully uncover the differences in EEG patterns across different imaginary language tasks, and the combination with the Random Forest model achieves stronger discriminative capabilities.
[0072] Table 2:
[0073] To further verify the generalization ability of the proposed method, this embodiment conducts comparative experiments on the Bimodal InnerSpeech 2023 dataset and systematically compares it with existing research. This dataset includes two semantic tasks: social category and numeric category. Compared to the InnerSpeech 2021 dataset, it has a higher level of semantic abstraction and more complex differences between categories, placing higher demands on the ability to represent EEG features. Therefore, this embodiment also uses classification accuracy as a unified evaluation metric to ensure the comparability between different methods.
[0074] As shown in Table 3, the method proposed by Jiahao Qin et al. (2024) is based on Markov, Random Forest (RF), and Support Vector Machine (SVM) models, respectively. The highest accuracy in the social category task is 47.20%, while it is only 36.50% in the numeric category task, indicating a relatively low overall performance. This suggests that traditional shallow models struggle to fully extract discriminative information from EEG signals in complex semantic classification tasks. The method proposed by Hussna Elnoor M. Abdalla et al. (2026) combines recurrent neural networks (RNNs) and deep neural networks (DNNs), achieving accuracy of 52.19% and 50.87% in the two categories, respectively, showing improvement over traditional methods. However, it is still limited by temporal modeling capabilities and feature representation methods, resulting in limited performance improvement. Furthermore, the method proposed by Diego Lopez-Bernal et al. (2024) (k-NN (K-Nearest Neighbors) + SVM) achieves an average accuracy of approximately 59.55%, obtaining the best value. In comparison, the method proposed in this embodiment, based on EEG optical flow statistical features and a random forest model, achieves 75.00% accuracy in the social category task and 66.67% accuracy in the numeric category task. The overall performance is 18.95% higher than the best existing method proposed by Diego Lopez-Bernal et al. (2024).
[0075] The paper title and source corresponding to the method proposed by Jiahao Qin et al. (2024) are as follows: Qin J, Zong L, Liu F. Exploring inner speech recognition via cross-perception approach in EEG and fMRI. Applied Sciences, 2024, 14(17): 7720. The paper title and source corresponding to the method proposed by Hussna Elnoor M. Abdalla et al. (2026) are as follows: Elnoor M. Abdalla H, Basri H, Aris I, et al. EEG-based inner speechdecoding using phase-locking and spatial features with a dual-branch deeplearning model. Open Research Africa, 2026, 8: 26. The paper title and source corresponding to the method proposed by Diego Lopez-Bernal et al. (2024) are as follows: Lopez-Bernal D, Balderas D, Ponce P, et al. Exploring inter-trialcoherence for inner speech classification in EEG-based brain–computer interface. Journal of Neural Engineering, 2024, 21(2):14. Table 3:
Claims
1. A non-invasive EEG imagery language decoding method based on multidimensional optical flow vectors, characterized in that, The method includes the following steps: S1. Acquire non-invasive EEG signals, filter the signals, perform Hilbert transform, sliding window amplitude extraction and frequency band specific Kriging interpolation to generate multi-rhythm EEG amplitude field sequences; S2. Quantitatively verify the three core assumptions of constant brightness, spatial consistency and temporal continuity in optical flow estimation, and determine the compatibility between the EEG amplitude field and the optical flow method. Quantitative verification includes: The constant brightness assumption is verified by phase continuity, the spatial consistency assumption is verified by local stationarity, the temporal continuity assumption is verified by inter-frame velocity variation, and a comprehensive fit score is calculated. When the overall fit score is ≥0.7, the EEG data is deemed to meet the application conditions for optical flow estimation. S3. Based on the adaptability results, the CLG optical flow algorithm is used to calculate the optical flow vector field of the EEG amplitude field, and the source and sink points are identified by divergence, determinant and Poincaré index; S4. Track the source and sink points to extract 18-dimensional single-rhythmic optical flow features, including optical flow velocity and direction, source and sink point ratio, and trajectory length and offset, to form multi-rhythmic optical flow features; The trajectory tracking employs a four-fold boundary constraint strategy, including: Initial point boundary constraints, velocity query boundary constraints, ODE solver hard boundary constraints, and trajectory result secondary truncation constraints ensure that the trajectory coordinates do not exceed the limits; S5. Dynamic weight allocation is performed on the multi-rhythmic optical flow features through the task-aware gating network TAG-Net, and the optimal rhythm combination is selected and fused to obtain multi-dimensional optical flow features. S6. Input the multidimensional optical flow features into a random forest classifier that has undergone hyperparameter optimization to complete the classification and decoding of the imagined language.
2. The non-invasive EEG imagery language decoding method according to claim 1, characterized in that, The source and sink points include the source point and the sink point. In step S3, the source and sink point identification meets the following conditions: Valid source points: divergence div > 0, determinant det > 0, Poincaré exponent close to 1; Valid sinks: divergence < 0, determinant det > 0, Poincaré exponent close to -1; Furthermore, noise fluctuations are filtered out using a divergence threshold to improve recognition robustness.
3. The non-invasive EEG imagery language decoding method according to claim 2, characterized in that, In step S4, the 18-dimensional optical flow features include: Average velocity, peak velocity, velocity standard deviation, velocity median, directional mode, directional standard deviation, directional absolute mean, principal direction proportion, average trajectory length, maximum trajectory length, average trajectory offset, source center normalized coordinates, sink center normalized coordinates, source proportion, sink proportion, velocity spatiotemporal average, directional spatiotemporal average, and optical flow activity.
4. The non-invasive EEG imagery language decoding method according to claim 3, characterized in that, In step S5, the task-aware gating network TAG-Net includes: Global average pooling is performed on the 18-dimensional optical flow features of each rhythm, and the features are concatenated to obtain cross-rhythm features; Task conditional encoding is introduced, and the weights of each rhythm are obtained through a fully connected layer and the Softmax function; The top-k rhythms are selected as the optimal combination based on their weights to achieve weighted fusion of multiple rhythm features.
5. The non-invasive EEG imagery language decoding method according to claim 4, characterized in that, In step S6, the random forest uses grid search and 5-fold hierarchical cross-validation to optimize hyperparameters, which include: The parameters include the number of decision trees, the maximum depth of the decision tree, the minimum number of split samples per node, and the class weights.
6. A non-invasive EEG imagery language decoding system based on multidimensional optical flow vectors, characterized in that, The system includes: Data preprocessing module: Acquires non-invasive EEG signals, performs filtering, Hilbert transform, sliding window amplitude extraction, and frequency band-specific Kriging interpolation on the signals to generate multi-rhythm EEG amplitude field sequences; Optical flow adaptability verification module: Quantitatively verifies the three core assumptions of constant brightness, spatial consistency and temporal continuity of optical flow estimation, and determines the adaptability of EEG amplitude field to optical flow method; Quantitative verification includes: The constant brightness assumption is verified by phase continuity, the spatial consistency assumption is verified by local stationarity, the temporal continuity assumption is verified by inter-frame velocity variation, and a comprehensive fit score is calculated. When the overall fit score is ≥0.7, the EEG data is deemed to meet the application conditions for optical flow estimation. Optical flow vector field calculation module: Based on the adaptation results, the CLG optical flow algorithm is used to calculate the optical flow vector field of the EEG amplitude field, and the source and sink points are identified by divergence, determinant and Poincaré index; Multi-rhythmic optical flow feature extraction module: Tracks the source and sink points to extract 18-dimensional single-rhythmic optical flow features, including optical flow velocity and direction, source and sink point ratio, and trajectory length and offset, to form multi-rhythmic optical flow features; The trajectory tracking employs a four-fold boundary constraint strategy, including: Initial point boundary constraints, velocity query boundary constraints, ODE solver hard boundary constraints, and trajectory result secondary truncation constraints ensure that the trajectory coordinates do not exceed the limits; Dynamic rhythm selection module: The task-aware gating network TAG-Net dynamically assigns weights to multi-rhythmic optical flow features, selects the optimal rhythm combination, and fuses them to obtain multi-dimensional optical flow features; Classification and Decoding Module: Inputs multidimensional optical flow features into a random forest classifier that has undergone hyperparameter optimization to complete the classification and decoding of the imagined language.
Citation Information
Patent Citations
Three-dimensional convolution method for decoding imaginary language electroencephalogram topographic map
CN120995220A
KR1016758750000B1