A transparence modulation brain-computer interface control system and control method based on cross-modal visual fusion
Patent Information
- Application Number
- CN202610811770.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-06
- Publication Date
- 2026-08-28
AI Technical Summary
遮挡与视觉疲劳,高对比度周期性闪烁刺激与真实场景叠加时,对环境图像产生遮挡干扰,并因持续高频闪烁易引发视觉疲劳,限制系统长期使用;
本发明通过环境感知与视觉刺激的融合,避免传统SSVEP脑控系统中双屏或界面切换方式带来的认知负担,实现更加直观一致的人机交互;通过透明度调制方式嵌入刺激信号,在不遮挡环境信息的前提下实现有效诱发,提高视觉信息利用率并降低视觉干扰;通过引入基于通道统计建模的解码方法及时间窗融合机制,在弱诱发信号条件下提升识别稳定性与抗干扰能力;通过事件驱动与置信度门控机制,提高系统控制的可靠性与安全性。综合来看,本发明在交互模式、信号诱发方式及系统稳定性方面相较现有技术具有明显优势,适用于多类移动平台的脑机接口控制场景。
Smart Images

Figure CN122653431A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of brain-computer interface (BCI) technology, and relates to non-invasive EEG signal processing, human-computer interaction and intelligent control technology. Specifically, it relates to a transparency modulation brain-computer interface control method and system that integrates environmental perception information and visual evoked mechanisms. Background Technology
[0002] With the development of Brain-Computer Interface (BCI) technology, non-invasive BCIs based on steady-state visual evoked potentials (SPVs) have been widely used in intelligent control, human-computer interaction, and auxiliary control systems due to their advantages such as no need for complex training, high signal stability, and superior information transmission efficiency. However, in practical applications on mobile platforms such as drones and vehicles, as well as in complex dynamic environments, existing SSVEP BCI systems still have significant shortcomings in terms of natural interaction, environmental adaptability, and robust decoding capabilities for weak signals, making it difficult to meet the needs of efficient and low-burden human-computer interaction in real-world scenarios. Therefore, it is necessary to improve existing technologies to enhance their practicality and stability in complex scenarios.
[0003] The most recent brain-computer interface technologies typically employ immersive visual stimulation systems based on virtual reality (VR) or augmented reality (AR). These systems overlay visual stimuli onto a real or simulated environment to enhance the immersive experience and intuitive operation. Such systems generally include an immersive display module, a visual stimulus generation module, an EEG acquisition module, a signal processing and decoding module, and an execution control module. In these methods, the system first presents a three-dimensional or augmented reality environment to the user through VR or AR devices. Simultaneously, multiple visual stimulus regions with different frequency characteristics are overlaid on the environment, such as icons, buttons, or spatially anchored targets floating in the scene. These stimulus regions typically employ periodic flashing, brightness modulation, or pattern changes to induce steady-state visual evoked potential (SSVEP) signals.
[0004] The basic steps in its implementation include: 1) Construct virtual or augmented reality scenes and load environmental visual information; 2) Overlay multiple preset stimulus target areas in the scene and assign them different stimulus frequencies or encoding patterns; 3) Users generate corresponding SSVEP responses by gazing at specific stimulus targets in VR / AR scenes; 4) The EEG acquisition module acquires multi-channel EEG signals and performs preprocessing; 5) Use canonical correlation analysis (CCA) or other improved correlation methods to perform frequency matching and classification of the signal; 6) Output the recognition results and drive external devices to execute corresponding control commands.
[0005] Existing research (such as Riechmann et al., 2015) has explored embedding SSVEP stimuli into natural scenes in immersive visual environments to enhance interactive immersion through VR environments; some subsequent works have also attempted to overlay stimulus targets with the real environment in AR interfaces to enhance the intuitiveness of task execution. However, these methods still achieve visual stimulus presentation through "explicit stimulus target overlay".
[0006] From the perspective of working mechanism, existing technologies are based on "fixed visual stimulation driving EEG response". The visual stimulation and environmental perception information are independent of each other, which leads to the following three types of fundamental problems: High cognitive load means users need to switch their attention between environmental perception and target selection, increasing the burden of visual attention resource allocation and reducing the naturalness of interaction and operational efficiency. Occlusion and visual fatigue: When high-contrast periodic flickering stimulation is superimposed on real scenes, it causes occlusion interference to environmental images, and the continuous high-frequency flickering can easily cause visual fatigue, limiting the long-term use of the system. Insufficient robustness in complex scenes: Factors such as changes in illumination, motion artifacts, and limited stimulus intensity cause SSVEP signals to have a low signal-to-noise ratio. Existing decoding methods (such as FBCCA) have an increased false positive rate and insufficient stability in complex scenes.
[0007] The aforementioned problems all stem from the inherent defects in the existing SSVEP brain-computer interface system regarding the generation of visual stimuli and the fusion mechanism of environmental information. Summary of the Invention
[0008] This invention proposes a transparency modulation brain-computer interface control system and method based on cross-modal visual fusion. It drives the generation of visual stimuli through environmental perception and combines a stable decoding method for weakly evoked signals to achieve highly reliable brain-computer interface control in complex scenarios.
[0009] The technical solution adopted in this invention is as follows: This invention first proposes a transparency modulation brain-computer interface control system based on cross-modal visual fusion, which consists of a visual perception and stimulation module, a signal processing and intent parsing module, and a control execution module. I. Visual Perception and Stimulus Module, including the following sub-modules: (1) Environmental perception module: real-time scene images are acquired through camera, and target area information, including target location, scale and category, is obtained by target detection or region segmentation method; (2) Stimulus generation module: constructs a visual stimulus region based on the target region information, and applies periodic transparency modulation to the region to implicitly embed the steady-state visual evoked potential (SSVEP) evoked stimulus signal into the original scene image without changing the original image content and structure, without occlusion interference, so that the stimulus signal is embedded into the original image without changing its content and structure. The transparency modulation can be implemented using periodic functions, including sine functions, square wave functions, or piecewise functions, and its general expression is:
[0010] in Let a0 be the transparency at time t, a0 be the base transparency, and A be the modulation amplitude. For modulation frequency, As the initial phase, different control commands correspond to different frequencies or modulation modes, enabling independent encoding of multiple commands and SSVEP induction.
[0011] (3) EEG acquisition module: acquires multi-channel EEG signals from the stimulus generation module through EEG acquisition equipment, and completes signal amplification and digitization through analog front end and data acquisition unit.
[0012] II. Signal Processing and Intent Parsing Module, including the following sub-modules: (1) Preprocessing module, which preprocesses the multi-channel EEG signals transmitted from the EEG acquisition module, including bandpass filtering, power frequency interference suppression and artifact removal; (2) Feature extraction module: Based on the processed multi-channel EEG signals, features for transparency modulation evoked signals are constructed; (3) Mahalanobis spatial correlation analysis module: In the data feature discrimination stage, a correlation analysis method based on Mahalanobis spatial modeling is introduced. By modeling the statistical characteristics of multi-channel signals, stable identification of weak induced signals is achieved. (4). Channel contribution weighting module: Channel feature contribution calculation, contribution weighting calculation, further combined with channel quality assessment results, weighting the contribution of each channel feature to improve overall robustness; (5) Sliding time window decision module: In order to improve the stability of real-time control, this invention introduces a time-series decision fusion mechanism based on sliding time window, which accumulates and analyzes the discrimination results within the continuous time window. The time window length can be set to 1–3s, and the sliding step size can be 1–1.5s. The final identification result is output through majority voting or weighted fusion strategy, thereby reducing the control instability problem caused by instantaneous misjudgment. III. Control Execution Module, which includes the following sub-modules: (1) The confidence asynchronous control module receives control instructions and confidence information from the sliding time window decision module. It adopts an event-driven asynchronous control mechanism to decouple signal decoding from execution control. The control instruction is triggered only when the recognition result meets the preset confidence threshold. (2) Remote communication module: The communication module transmits images through a low-latency communication protocol, and the EEG signals are converted into control commands through an algorithm to confirm the completion of command transmission; (3) Actuator module: The actuator (such as a car or drone) performs corresponding actions according to the instructions received from the remote communication module, and can achieve closed-loop regulation through feedback information.
[0013] This invention also proposes a control method for a transparency modulation brain-computer interface control system based on cross-modal visual fusion, comprising the following steps: (1). Environmental perception steps: real-time scene images are acquired through cameras, and target area information in the scene is obtained by using target detection or region segmentation methods; (2). Stimulus generation step: Construct a visual stimulus region based on the target region information, apply a periodic transparency modulation signal to the region, embed the transparency modulation signal into the original scene image, and the transparency modulation signal is generated by a periodic function; (3). EEG acquisition steps: acquire multi-channel EEG signals generated when the user gazes at the visual stimulation area through an EEG acquisition device; (4). Signal processing and intent parsing steps: The multi-channel EEG signal is preprocessed, and features for transparency modulation evoked signals are constructed based on the processed multi-channel EEG signal. The processed EEG signal is decoded using a canonical correlation analysis method based on regularized Mahalanobis distance to identify the user's control intent. (5) Weighted channel contribution, calculation of channel feature contribution, weighted contribution calculation, further combined with the channel quality assessment results, weighted the feature contribution of each channel to improve overall robustness; (6) Sliding time window decision: In order to improve the stability of real-time control, a time-series decision fusion mechanism based on sliding time window is introduced. The discrimination results within the continuous time window are accumulated and analyzed. The time window length can be set to 1–3 s, and the sliding step size can be 0.1–0.5 s. The final identification result is output through majority voting or weighted fusion strategy, thereby reducing the control instability caused by instantaneous misjudgment. (7). Control command output steps: It accepts control commands from the sliding time window decision module, adopts an event-driven asynchronous control mechanism to decouple signal decoding from execution control, and triggers control commands only when the recognition result meets the preset confidence threshold. The communication module confirms the completion of command transmission via the low-latency image transmission protocol TCP. The actuator performs corresponding actions based on the instructions received from the remote communication module, and can achieve closed-loop regulation through feedback information.
[0014] Furthermore, in the stimulation generation step, the periodic function of the periodic transparency modulation signal is a sine function, a square wave function, or a piecewise function, and its expression is: α(t) = α0 + A•sin (2πft + φ), where α(t) is the transparency at time t, α0 is the basic transparency, A is the modulation amplitude, f is the modulation frequency, and φ is the initial phase.
[0015] Furthermore, the canonical correlation analysis method based on regularized Mahalanobis distance includes the following steps: Data processing of EEG signals and background noise; Solve for the covariance matrix of the background noise; The covariance matrix is then subjected to Tikhonov regularization; Calculate the Mahalanobis distance between the feature data to be identified and each classification target template; The classification target is determined based on the minimum Mahalanobis distance criterion.
[0016] Furthermore, canonical correlation analysis based on regularized Mahalanobis distance extracts the spatial projection component of the response in different trials and obtains the spatially filtered signal as a feature:
[0017] in For spatial filters, Raw EEG data, T represents the extracted feature vector, and T represents the sum of the features extracted from the vector. The transpose of the spatial filter matrix; After dimensionality reduction and segmentation of the EEG signal, Where Nc is the number of channels and S represents the number of sampling points. Let represent the sampling rate, T represent the time window length, and R be the set of rational numbers. The main expression is that x is a matrix of size Nc*S, and the key point is that the number of rows and columns of the matrix is Nc and S. One-dimensional features are obtained after filtering with a spatial filter:
[0018] Where k represents the number of classification targets, for classification target k, construct the feature template:
[0019] Where represents the classification template of classification target k, and represents the number of training data corresponding to classification target k; Similarly, feature extraction is performed on the noise signal to obtain features:
[0020] Before calculating the background noise covariance matrix, it is necessary to ensure that the mean is 0; therefore, the noise matrix is centered.
[0021] in The mean of the background noise:
[0022] From this, we can obtain the average covariance matrix of the background noise:
[0023] Since Mahalanobis distance requires inverting the covariance matrix, and considering that the amount of real-world driving scenario data is relatively small and the background noise is strong, the covariance matrix is close to a singular matrix, leading to instability in the inversion process, Tikhonov regularization is applied to the covariance matrix.
[0024] Where Tr() represents the trace of the matrix. express The length of I is The identity matrix; For the new feature data, calculate its Mahalanobis distance with each classification target template:
[0025] The classification target is determined according to the minimum Mahalanobis distance: .
[0026] The beneficial effects of this invention are: This invention, by fusing environmental perception and visual stimulation, avoids the cognitive burden caused by dual-screen or interface switching methods in traditional SSVEP brain-computer interface systems, achieving a more intuitive and consistent human-computer interaction. It embeds stimulus signals through transparency modulation, achieving effective induction without obscuring environmental information, improving visual information utilization and reducing visual interference. By introducing a decoding method based on channel statistical modeling and a time window fusion mechanism, it enhances recognition stability and anti-interference capabilities under weak induction signal conditions. Event-driven and confidence-gating mechanisms improve the reliability and security of system control. In summary, this invention has significant advantages over existing technologies in terms of interaction mode, signal induction method, and system stability, and is suitable for brain-computer interface control scenarios on various mobile platforms. Attached Figure Description
[0027] Figure 1 This is a system structure diagram of the present invention; Figure 2 This is an effect diagram of the present invention constructing a visual stimulation region based on target region information and applying a periodic transparency modulation signal to the region, wherein (1) stimulation region construction; (2) transparency time change sequence; (3) stimulation square superposition effect with visual perception; (4) comparison diagram with traditional scheme; Figure 3 This invention is a sliding time window fusion weighted decision graph; Figure 4 This is an analysis diagram of the experimental results of an embodiment of the present invention; Figure 5 In this embodiment, obstacles, passage boundaries, and target driving areas are arranged in a complex indoor lighting environment to simulate the autonomous assisted perception and brain-controlled motion control process of an intelligent vehicle in a real scene. Detailed Implementation
[0028] The present invention will be further described in detail below with reference to specific embodiments. The present invention includes, but is not limited to, the following embodiments. Any simple modifications or equivalent substitutions made under the spirit and principles of the present invention should be covered within the protection scope of the present invention.
[0029] Example 1: Brain-Controlled Experiment of an Intelligent Car under Complex Indoor Lighting like Figure 1 As shown in the figure, the transparency modulation brain-computer interface control system based on cross-modal visual fusion provided in this embodiment consists of a visual perception and stimulation module, a signal processing and intent parsing module, and a control execution module. These have been described in detail in the technical solution section and will not be repeated here. The specific hardware configuration and control method are described below.
[0030] 1. Experimental hardware configuration This embodiment employs an 8-channel non-invasive dry electrode EEG acquisition device. The electrodes are arranged according to the international 10-20 standard, and channels O1, O2, Oz, Pz, PO3, PO4, P3, and P4 are selected to acquire visual cortical EEG signals; the sampling rate is set to 250Hz. The environmental perception module uses a 1080P industrial camera mounted on the robotic arm support, operating at 30fps, to acquire images of the environment in front of and around the vehicle. The actuator is a four-wheeled intelligent vehicle equipped with a robotic arm, whose chassis supports four types of movement commands: forward, backward, left turn, and right turn. In this embodiment, the robotic arm mainly serves as a visual perception carrier, providing environmental image acquisition with an adjustable viewing angle, and does not perform grasping or handling operations. The host computer uses an industrial control computer to complete EEG decoding, visual perception, and control command generation. The wireless communication module uses a 2.4G wireless pass-through module with a communication latency of less than 20ms.
[0031] 2. Environmental perception and transparency stimulus generation like Figure 5 As shown, this embodiment arranges obstacles, passage boundaries, and target driving areas in a complex indoor lighting environment to simulate the autonomous assisted perception and brain-controlled motion control process of an intelligent vehicle in a real-world scenario. A robotic arm equipped with a camera collects real-time images of the road ahead and the surrounding environment. The fixed or preset observation posture of the robotic arm expands the image acquisition field of view, improving the perception of obstacles, drivable areas, and path boundaries. The host computer processes the collected images in real-time, using lightweight target detection or image segmentation algorithms to identify obstacle areas, drivable areas, and the central interactive area of the screen. It automatically selects unobstructed areas with suitable contrast as visual stimulus embedding locations.
[0032] In this embodiment, the transparency modulation function uses sinusoidal modulation, with the base transparency a0 set to 0.2 and the modulation amplitude A set to 0.8. To implement the encoding of four types of motion commands, the modulation frequencies are set to 8Hz, 10Hz, 12Hz, and 14Hz, corresponding to the forward, backward, left-turn, and right-turn actions of the vehicle. The modulation phases are set to 0, 0.5π, 1π, and 1.5π. The visual stimulus is superimposed on the original environmental image captured by the robotic arm's camera in a way that changes the transparency. The stimulus area blends naturally with the actual road image, without any extra white flashing squares or obvious interface obstruction, thereby reducing visual interference and operational burden in complex environments.
[0033] 3. Preprocessing of EEG signals The raw EEG signals were subjected to a 5–30 Hz bandpass filter to remove DC drift and high-frequency electromyographic noise; a 50 Hz power frequency notch filter was used to suppress power frequency interference from indoor lighting and electrical equipment. After preprocessing, steady-state visual evoked potential signals related to transparency-modulated visual stimuli were retained to provide low-noise EEG data for subsequent brain-controlled command recognition.
[0034] 4. Decoding process of RMCCA algorithm based on Regularized Mahalanobis Distance CCA.
[0035] In this embodiment, the time window length is set to 2 seconds and the sliding step size is set to 1 second. First, steady-state gaze data of the subjects under four types of frequency stimuli are collected to construct four command templates: forward, backward, left turn, and right turn. The real-time acquired EEG data is centered, and the noise covariance matrix is solved. A Tikhonov regularization term is introduced to prevent matrix singularity and improve decoding stability. Then, the Mahalanobis distance between the EEG features to be identified and the four command templates is calculated, and the category with the smallest distance is selected as the preliminary recognition result. Simultaneously, the channel weights are adaptively adjusted based on the real-time signal-to-noise ratio of each channel. The weights of high-noise channels are reduced, while the weights of high-quality visually relevant channels such as O1, O2, Oz, Pz, PO3, PO4, P3, and P4 are increased, completing the optimization of EEG features and command discrimination.
[0036] 5. Timing-based decision-making and vehicle control In this embodiment, three consecutive sliding window judgment results are retained, and a majority voting strategy is used to output the final control command. The confidence threshold is set to 0.85; motion control commands are only issued to the intelligent vehicle chassis when the algorithm outputs a confidence score greater than the threshold. The vehicle executes forward, backward, left, or right turns based on the recognition results. During the movement, the robotic arm maintains the camera in a preset observation posture, continuously acquiring images of the surrounding environment and feeding the image information back to the host computer for updating visual information, obstacle positions, etc.
[0037] After the vehicle performs an action, the host computer provides visual feedback to the user based on the environmental images captured by the robotic arm's camera, thus forming a closed-loop control process of "environmental perception - EEG decoding - vehicle control - status feedback". Implementation effect
[0038] like Figure 5 As shown, in an experimental scenario with complex indoor lighting, obstacle obstruction, and changing road conditions, this embodiment can achieve real-time perception of the vehicle's surrounding environment using a robotic arm and camera. By inducing stable SSVEP EEG signals through transparency-modulated visual stimuli, it achieves reliable control of the chassis movement of the robotic arm-equipped intelligent vehicle. Experimental results show that the method in this embodiment achieves an average recognition accuracy of 93.89% under four types of motion commands; compared with the traditional fixed-interface flashing stimulus control scheme, the recognition accuracy is improved by approximately 3.8%; because the visual stimulus is integrated with the real-world environment, the subject's subjective visual fatigue is reduced, the burden of attention switching is decreased, and the intelligent vehicle's movement can be continuously and stably controlled for over 30 minutes.
[0039] In this embodiment, the robotic arm is mainly used to carry the camera and assist in environmental perception, without involving end-effector grasping, gripping, or transporting motion control. This implementation verifies the applicability of the brain-controlled method in environmental perception and chassis motion control scenarios of a "mobile platform with a robotic arm," providing a foundation for subsequent expansion to robotic arm collaborative operation, target grasping, and complex task execution.
[0040] Example 2: Experimental Example of Eight-Direction Brain-Controlled Displacement of a UAV Based on a Gantry Restraint Platform 1. Experimental Scenarios and Hardware This embodiment employs an indoor UAV constraint experimental platform, including a gantry suspension mechanism, a quadcopter drone, a ground-scaled map, an EEG acquisition device, a host computer control system, and a communication module. The drone is fixed to the top of the gantry via suspension and does not engage in free flight; instead, the drone's movement above the map is simulated by suspension displacement. A scaled map is placed on the ground to simulate outdoor terrain, roads, or target areas. The drone's onboard camera or the platform's visual acquisition module acquires images of the map below in real time and transmits them to the host computer. EEG acquisition uses a 24-channel dry electrode EEG device, focusing on acquiring visually evoked EEG signals from the occipital and parieto-occipital regions, with a sampling rate set to 250Hz. Control commands are set to five categories: start, forward, backward, left, and right.
[0041] 2. Stimulation modulation parameter settings like Figure 2 As shown, the host computer overlays five semi-transparent visual stimulus areas onto the zoomed map image, corresponding to five types of commands: start, forward, backward, left, and right. The start stimulus area is placed in the center of the screen or in a position that is easy to look at.
[0042] The RGB values of each stimulus region are determined based on the average color of the underlying map image at its coverage location, allowing the stimulus region to blend naturally with the map background and avoiding the occlusion of map information caused by traditional high-contrast flashing squares. This embodiment uses transparency modulation to induce SSVEP signals. The five stimulus regions correspond to modulation frequencies of 8, 9, 10, 11, and 12 Hz, and correspond to five types of control commands.
[0043] 3. Algorithm adaptation and optimization The acquired EEG signals were subjected to bandpass filtering (0.5–45 Hz) and power frequency notch filtering (50 Hz), and artifacts such as blinking and eye movements were removed. To address potential mechanical vibrations and environmental noise during suspension operation, signal quality was assessed for each channel, reducing the weight of high-noise channels and increasing the weight of high-quality visual channels in the occipital region.
[0044] like Figure 3As shown, this embodiment uses the RMCCA algorithm for five types of command recognition. The time window length is set to 2 seconds, the sliding step size is set to 1.5 seconds, and the step size is 0.1 seconds. By calculating the correlation and Mahalanobis distance between real-time EEG features and five frequency templates, the category with the highest matching degree is selected as the candidate control command.
[0045] The system continuously retains the recognition results of six sliding windows and outputs the final command using a majority voting strategy. Control commands are only sent to the gantry suspension displacement mechanism when the recognition confidence level is higher than a preset threshold.
[0046] When a start command is recognized, the system enters brain-controlled displacement control mode; when a forward, backward, leftward, or rightward command is recognized, the suspension drives the drone to undergo restricted displacement along the corresponding direction of the zoomed map. The suspension position is fed back to the host computer via displacement sensors and mapped to the coordinates of the zoomed ground map, realizing closed-loop control of "visual stimulus selection - EEG command recognition - constrained displacement execution - map position feedback".
[0047] 4. Experimental Results This embodiment, while ensuring experimental safety, realizes the start-up and four-directional displacement control (forward, backward, left, and right) of a drone based on EEG signals. Because the drone is constrained by the gantry suspension, the risks of falling, collisions, and attitude instability during free flight are avoided, making it suitable for verifying brain-controlled drone control algorithms.
[0048] like Figure 4 As shown, experimental results indicate that the average recognition accuracy of the five types of control commands using the transparency modulation visual stimulation and RMCCA decoding method described in this embodiment reaches 93.02%, an improvement of 5.08% compared to traditional methods. Compared with traditional high-contrast flashing stimulation methods, this embodiment has lower visual fatigue and a lower false trigger rate, enabling safe, continuous, and controllable displacement of the UAV on zoomed maps.
[0049] Example 3: Performance Comparison Experiment of Five-Class Asynchronous Brain-Computer Interface Algorithms To verify the effectiveness of the RMCCA algorithm of this invention in brain-controlled command recognition, this embodiment compares and tests feature extraction algorithms such as CCA, eCCA, TRCA, TDCA, and RMCCA based on 5-class asynchronous brain-computer interface data. The experiments statistically analyze the classification accuracy, information transmission rate, and response time of each algorithm under different time windows.
[0050] like Figure 4As shown in the experimental results, the RMCCA algorithm used in this invention achieves a good balance between classification accuracy, information transmission rate, and response time. Within a 0.9s time window, the RMCCA algorithm achieves a classification accuracy of 92.22%, a maximum information transmission rate of 118.13 bits / min, and a response time between eCCA and TRCA, demonstrating superior overall performance compared to other algorithms. The experimental results indicate that the RMCCA algorithm can achieve stable and efficient brain-computer interface command recognition within a short time window, making it suitable for real-time brain-computer interface control scenarios.
[0051] System General Implementation Instructions This invention adopts a modular system architecture, which can flexibly adapt hardware devices and controlled objects according to different application scenarios. The EEG acquisition end is compatible with dry electrodes, wet electrodes and portable head-mounted EEG devices; the visual acquisition end can select vehicle-mounted cameras, airborne cameras, wide-angle cameras or panoramic cameras according to the needs of the scenario; the execution end can be used not only for smart cars and drones, but also can be extended to intelligent carriers such as robotic arms, rehabilitation exoskeletons, smart wheelchairs, virtual reality interactive devices and other intelligent carriers.
[0052] In practice, parameters such as transparency modulation, number of stimulation regions, frequency encoding method, time window length, sliding step size, and regularization coefficient can all be adjusted according to lighting conditions, number of control commands, subject status, and hardware platform performance. The system as a whole does not depend on specific equipment models and can be replaced and combined among different EEG acquisition devices, visual perception modules, and actuators, demonstrating good engineering adaptability and application value.
Claims
1. A transparency modulation brain-computer interface control system based on cross-modal visual fusion, characterized in that... It consists of a visual perception and stimulus module, a signal processing and intent parsing module, and a control execution module. (1) The visual perception and stimulation module includes the following sub-modules in sequence: ①. Environmental perception module: acquires real-time scene images through camera and uses target detection or region segmentation methods to obtain target region information, including target location, scale and category parameters; ②. Stimulus generation module: Constructs a visual stimulus region based on target region information, applies periodic transparency modulation to the region, implicitly embeds the steady-state visual evoked potential (SSVEP) evoked stimulus signal into the original scene image, and generates multi-channel EEG signals. ③. EEG acquisition module: acquires multi-channel EEG signals from the stimulus generation module through EEG acquisition equipment, and completes signal amplification and digitization through analog front-end and data acquisition unit; (2) The signal processing and intent parsing module includes the following sub-modules: ①. Preprocessing module: preprocesses the multi-channel EEG signals from the EEG acquisition module, including bandpass filtering, power line interference suppression, and artifact removal; ②. Feature extraction module: Constructs features for transparency modulation evoked signals based on the processed multi-channel EEG signals; ③. Mahalanobis spatial correlation analysis module: Based on the correlation analysis method of Mahalanobis spatial modeling, it achieves stable identification of weak induced signals by modeling the feature data constructed by the feature extraction module. ④. Channel contribution weighting module: This module weights the contribution of each channel feature and further combines it with the channel quality assessment results to improve overall robustness; ⑤. Sliding time window decision module: To improve the stability of real-time control, a time-series decision fusion mechanism based on sliding time window is introduced. The discrimination results within the continuous time window are accumulated and analyzed. The time window length can be set to 1-3s, and the sliding step size can be 1-1.5s. The final recognition result is output through majority voting or weighted fusion strategy, thereby reducing the control instability problem caused by instantaneous misjudgment. (3) The control execution module includes the following sub-modules: ①. The confidence asynchronous control module receives control commands and confidence information from the sliding time window decision module. It adopts an event-driven asynchronous control mechanism and triggers control commands only when the recognition result meets the preset confidence threshold. ②. Remote communication module: The communication module transmits images through a low-latency communication protocol, and the EEG signals are converted into control commands through an algorithm to confirm the completion of command transmission; ③. Actuator module: The actuator performs corresponding actions according to the instructions received from the remote communication module, and can achieve closed-loop regulation through feedback information.
2. The control method for a transparency modulation brain-computer interface control system based on cross-modal visual fusion according to claim 1, characterized in that... Includes the following steps: (1). Environmental perception steps: real-time scene images are acquired through cameras, and target area information in the scene is obtained by using target detection or region segmentation methods; (2). Stimulus generation step: Construct a visual stimulus region based on the target region information, apply a periodic transparency modulation signal to the region, embed the transparency modulation signal into the original scene image, and the transparency modulation signal is generated by a periodic function; (3). EEG acquisition steps: acquire multi-channel EEG signals generated when the user gazes at the visual stimulation area through an EEG acquisition device; (4). Signal processing and intent parsing steps: The multi-channel EEG signal is preprocessed, and features for transparency modulation evoked signals are constructed based on the processed multi-channel EEG signal. The processed EEG signal is decoded using a canonical correlation analysis method based on regularized Mahalanobis distance to identify the user's control intent. (5) Channel contribution weighting: calculate the contribution of each channel feature and the contribution weighting calculation, and further combine the channel quality assessment results to improve the overall robustness; (6) Sliding time window decision: In order to improve the stability of real-time control, a time-series decision fusion mechanism based on sliding time window is introduced. The discrimination results within the continuous time window are accumulated and analyzed. The time window length can be set to 1–3s, and the sliding step size can be 1–1.5s. The final identification result is output through majority voting or weighted fusion strategy, thereby reducing the control instability caused by instantaneous misjudgment. (7). Control command output steps: It accepts control commands from the sliding time window decision module, adopts an event-driven asynchronous control mechanism to decouple signal decoding from execution control, and triggers control commands only when the recognition result meets the preset confidence threshold. The communication module confirms the completion of command transmission via the low-latency image transmission control protocol TCP command. The actuator performs corresponding actions based on the instructions received from the remote communication module, and can achieve closed-loop regulation through feedback information.
3. The method for controlling a brain-computer interface based on cross-modal visual fusion with transparency modulation according to claim 2, characterized in that, In the stimulation generation step, the periodic function of the periodic transparency modulation signal is a sine function, a square wave function, or a piecewise function, and its expression is: α(t)=α0+A•sin(2πft+φ), where α(t) is the transparency at time t, α0 is the basic transparency, A is the modulation amplitude, f is the modulation frequency, and φ is the initial phase.
4. The method for controlling a brain-computer interface based on cross-modal visual fusion with transparency modulation according to claim 2, characterized in that, The canonical correlation analysis method based on regularized Mahalanobis distance includes the following steps: Data processing of EEG signals and background noise; Solve for the covariance matrix of the background noise; The covariance matrix is then subjected to Tikhonov regularization; Calculate the Mahalanobis distance between the feature data to be identified and each classification target template; The classification target is determined based on the minimum Mahalanobis distance criterion.
5. The method for controlling a brain-computer interface based on cross-modal visual fusion with transparency modulation according to claim 4, characterized in that, Based on the canonical correlation analysis method using regularized Mahalanobis distance, the spatial projection component of the response is extracted in different trials, and the spatially filtered signal is obtained as a feature: ; in For spatial filters, Raw EEG data, T represents the extracted feature vector, and T represents the sum of the features extracted from the vector. The transpose of the spatial filter matrix; After dimensionality reduction and segmentation of the EEG signal, Where Nc is the number of channels and S represents the number of sampling points. Let represent the sampling rate, T represent the time window length, and R be the set of rational numbers. The main expression is that x is a matrix of size Nc*S, and the key point is that the number of rows and columns of the matrix is Nc and S. One-dimensional features are obtained after filtering with a spatial filter: ; Where k represents the number of classification targets, for classification target k, construct the feature template: ; Where represents the classification template of classification target k, and represents the number of training data corresponding to classification target k; Similarly, feature extraction is performed on the noise signal to obtain features: ; Before calculating the background noise covariance matrix, it is necessary to ensure that the mean is 0; therefore, the noise matrix is centered. ; in The mean of the background noise: ; From this, we can obtain the average covariance matrix of the background noise: ; Since Mahalanobis distance requires inverting the covariance matrix, and considering that the amount of real-world driving scenario data is relatively small and the background noise is strong, the covariance matrix is close to a singular matrix, leading to instability in the inversion process, Tikhonov regularization is applied to the covariance matrix. ; Where Tr() represents the trace of the matrix. express Length, I is The identity matrix; For the new feature data, calculate its Mahalanobis distance with each classification target template: ; The classification target is determined according to the minimum Mahalanobis distance: 。