A method for realizing computer control using multi-modal motor imagery technique
By using multimodal motion imagery technology, combined with EEG and MEG signals, and employing a weighted average Bayesian fusion method and a triangular incremental method, the problems of inaccurate operation feedback and imprecise positioning of virtual mice in existing technologies have been solved. This has enabled the association of multiple operation logics and improved the accuracy and flexibility of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2023-11-16
- Publication Date
- 2026-04-21
AI Technical Summary
Existing virtual mouse methods that combine eye trackers and motion imagery technology suffer from low accuracy in operation feedback, inaccurate mouse positioning, and overly simplistic operation.
Using multimodal motor imagery technology, combined with high-density EEG and MEG signals, a classifier is constructed using a weighted average Bayesian fusion method and a triangular incremental method to associate multiple operational logics with behavioral patterns, and remote control is achieved by combining it with an eye tracker.
It improved the accuracy of operation feedback, ensured the system response speed, optimized mouse positioning, realized multiple operation functions, and enhanced the robustness and user-friendliness of the system.
Smart Images

Figure CN117930965B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of brain science and cognitive science, and in particular to a method for implementing a virtual mouse using multimodal motion imagination technology. Background Technology
[0002] Brain-computer interfaces (BCIs) are interactive systems built between humans and machines. Signal acquisition in BCIs can generally be categorized into non-invasive, semi-invasive, and invasive methods. Non-invasive methods, because they do not cause harm to the human body, are likely to be the first to be applied in real-world situations. Based on whether the brainwave signals are triggered by external stimuli, BCI technologies can be divided into evoked and spontaneous methods. Common evoked BCI technologies include steady-state visual evoked potentials (SSVEPs) and event-related potentials (ERPs). SSVEPs primarily involve evoking harmonic signals of corresponding frequencies in the brain when the subject fixates on a visual stimulus that flashes black and white at a fixed frequency, thereby determining the subject's intention. ERPs are long-latency evoked potentials that mainly reflect neurophysiological changes in the brain during cognitive processes. Classical ERPs can be categorized into four main components based on latency: P1, N1, P2, N2, and P3. P or N represents positive or negative, and 1, 2, or 3 represent peaks appearing approximately 100ms, 200ms, or 300ms after visual stimulation, respectively. The most classic application is the P300 Speller proposed by Farewell and Donchin et al. in 1988. Its stimulation paradigm mainly consists of random flashing of target and non-target stimuli. Approximately 300ms after the appearance of the low-probability target stimulus, a positive peak, called the P300 signal, can be observed in the subject's EEG. By detecting this signal, the subject's intention can be determined. Evoked brain-computer interface (BCI) technologies are typically limited by stimulation paradigms. Furthermore, both SSVEP and ERP require prolonged flashing stimulation of the subject's eyes, easily causing visual fatigue and not conforming to human operating habits. Therefore, this patent primarily employs spontaneous brain-computer interface technology, motor imagery (MI).
[0003] MI (Mind-Induced) brain-computer interface (BI) technology is independent of external stimuli and relies on endogenous, spontaneous brainwave signals. It only requires the subject to imagine a specific motor task, without requiring actual movement, to detect specific waveforms in different brain regions, thus identifying the subject's intention. When a subject imagines movement of a limb, it triggers changes in different rhythms (sensorimotor rhythms, SMRs) in the brain's sensorimotor cortex. The SMRs induced by imagining movement of different limbs exhibit different spatiotemporal distribution characteristics, and their EEG (electroencephalogram) has high discriminative power. Therefore, MI patterns in the EEG can be decoded using pattern recognition methods and converted into control commands to manipulate external devices. However, MI technology is limited by its ability to classify a limited number of objects. Currently, it can classify hands, feet, and tongue relatively well, meaning it can only output five categories at a time, making it difficult to directly output more commands to achieve more complex tasks. Evoked brain-computer interface (BCI) technology can achieve a greater number of target recognitions by arranging multiple tasks on a display. Spontaneous BCIs are more flexible and better suited to human operating habits. This prompted us to think about how to build a transformation bridge between the limited output categories and task arrangements of MI.
[0004] Magnetoencephalography (MEG) offers complementary information in terms of source depth and conductivity sensitivity, but also provides radial / tangential dipole detection. While previous studies have demonstrated the feasibility of Bronchocardiography (BC) and neural feedback based on MEG activity, the potential benefits of combining it with electroencephalography (EEG) signals have not been fully explored. The development of portable magnetometers based on optically pumped magnetometers could have practical implications for this integration. To address this knowledge gap, we considered a cohort of healthy subjects simultaneously recording high-density EEG and MEG signals during a MI-based brain-computer interface task.
[0005] Eye trackers are important instruments in basic psychological research, typically used to record the eye movement patterns of a person processing visual information. They are widely used in research on attention, visual perception, and reading. Many laptops now include eye trackers as a selling point; for example, the Alienware 17R5 features a Tobii eye tracker. However, because they can only track the gaze path of the human eye and lack confirmation functions, their practical use in daily life is quite limited, usually being used for gaming or professional psychological analysis.
[0006] The current traditional method of implementing a virtual mouse by combining eye trackers and asynchronous motion imagination technology still has many problems, such as low accuracy of operation feedback, inaccurate mouse positioning, and complex de-jitter algorithms with overly simplistic operation.
[0007] To address the aforementioned issues, this solution proposes several optimization methods. First, it improves the accuracy of operation feedback by using multimodal motion visualization technology. Second, it ensures a high system response speed; although multiple behavior classifiers are introduced, a highly complex model is not used. Third, it replaces the conventional de-jitter algorithm with a minimum circle covering algorithm based on the triangular incremental method, solving the problems of inaccurate mouse positioning and high algorithm time complexity. Finally, it uses a multi-classifier to distinguish more behavior patterns and associates six operation logics with behavior patterns, remotely implementing functions such as clicking the left and right mouse buttons, scrolling the mouse wheel, and turning the soft keyboard on / off. Summary of the Invention
[0008] The purpose of this invention is to solve the problems of low accuracy of operation feedback, inaccurate mouse positioning, and overly simplistic operation based on eye trackers and EEG motor imagery technology.
[0009] To achieve the above objectives, the present invention employs the following technical means:
[0010] This invention provides a method for computer control using multimodal motion visualization technology, comprising the following steps:
[0011] Step S1: Use a multimodal signal acquisition device to record motion image signals;
[0012] Step S2: Preprocess and extract features from the extracted EEG, MAG, and GRAD multimodal signals using the algorithm;
[0013] Step S3: Construct a classifier by using a weighted average-based Bayesian fusion method to assign higher weights to the patterns that best classify the data.
[0014] Step S4: Using the classification results, associate the motion imagery signals with the left and right mouse buttons, scroll wheel function, and soft keyboard on / off function, and combine with the eye tracker to achieve remote control of the computer.
[0015] In the above technical solution, the specific steps of S1 are as follows:
[0016] S11: Using multimodal devices means simultaneously recording high-density EEG and magnetoencephalography (MEG) signals in a motor imagery (MI) task. The system uses a 74-channel EEG system and the F1ekta Neuromag TRIUX machine to simultaneously record EEG and MEG data. The electrode positions on the scalp follow the 10-10 montage standard, and the reference electrode is located on the left scapula.
[0017] S12: Simultaneously consider the activity of electroencephalography (EEG) and magnetoencephalography (MEG), where the MEG consists of signals from a magnetometer and a gradient meter.
[0018] In the above technical solution, the specific steps of S2 are as follows:
[0019] S21: Locate the EEG, MAG, and GRAD signal electrodes and remove useless electrodes;
[0020] S22: Perform filtering, segmentation, noise reduction, and artifact removal operations on the raw EEG signal in sequence;
[0021] S22: Use MaxFilter to perform temporal-domain signal spatial separation on the raw magnetoencephalogram (MEG) signal to remove environmental noise;
[0022] S22: All signals from the magnetoencephalogram (MEG) are downsampled to 250 Hz and divided into 5-second cycles corresponding to the target cycle;
[0023] S23: Experts visually inspect the magnetoencephalogram (MEG) to remove artifacts. To better simulate the online scenario, no other artifact removal algorithms are used. After verification, all available frequency bands are retained.
[0024] S24: Using discrete oblate spheroid sequences as orthogonal window functions, a multi-window frequency transformation method is implemented to smooth the spectrum. This algorithm is implemented using the Fieldtrip toolbox.
[0025] S25: For each acquired signal period, extract the corresponding feature matrix M from the EEG, MAG, and GRAD power spectra. i EEG, MAG, and GRAD extract M i The dimensions are 74×36, 102×36 and 204×36, respectively;
[0026] S26: Use a feature extraction algorithm to extract features from matrix M i The most relevant features were extracted, focusing on the sensor signals of the contralateral motion region. The feature matrix sizes for EEG, MAG, and GRAD were 11×36, 15×36, and 30×36, respectively. Then, in M... i Nonparametric clustering-based permutation t-tests were performed between the dynamic and resting spectra to correct and eliminate low-quality signals and reduce the error rate. To this end, a statistical test threshold of p < 0.05 was set, and multiple comparisons were made to correct errors.
[0027] Finally, N was extracted within the standard frequency bands b = θ (4-7Hz), α (8-13Hz), β (14-29Hz), and γ (30-40Hz). f The most discriminative features are as follows:
[0028]
[0029] Where ζ i,b Represents the eigenvector, N f =1...10, with 10 as the maximum limit for feature dimensions, conforming to montage standards. It is the i-th selected feature, which is calculated and selected from the signals of each mode and each frequency band.
[0030] In the above technical solution, the specific steps of S3 are as follows:
[0031] S31: After feature extraction, classifiers are trained on the collected EEG signals, MAG magnetometer signals and GRAD gradient meter signals respectively. The motion imagination (MI) classifier for each modal signal is trained using the LDA algorithm with five-fold cross-validation. The three classifiers are then voted on to obtain the final classification result, which is one of the following states: left hand, right hand, both hands grasping, both hands open, both legs open, both legs together, or idle state.
[0032] In the LDA algorithm, there are N classes, the total number of samples is m, and the set of examples of the i-th class is X. i The number of samples in the corresponding category is m. i Define the global scatter matrix S t as follows:
[0033]
[0034] Where μ represents the mean vector of all samples, S w S is the intra-class scatter matrix b The inter-class scatter matrix;
[0035] Intraclass scatter matrix S w Represented as:
[0036]
[0037] S w It is defined as the sum of the scatter matrices of each category, where μ i S represents the mean vector of category i, from which we can obtain S b The expression, that is:
[0038]
[0039] Optimize objective function J:
[0040] where w∈R d×(N-1) tr(.) represents the trace of the matrix;
[0041] When solving the problem, let the above expression tr(W) T S w W) = 1, then the Lagrange multiplier method can be used, where the larger J is, the more it conforms to the goal of LDA. In the formula, W and W T These are the projection matrix and the transpose, respectively.
[0042] S32: Use Bayesian fusion to integrate information from different modalities to obtain the classification probability of each classifier, and then obtain the behavior classification result;
[0043] A weighted average-based Bayesian fusion method was used to integrate information from different modalities, and the posterior probability p was linearly combined. i p i Obtained by the classification of each mode i, with parameter λ i The weighted average is expressed by the following formula:
[0044]
[0045] Where P EEG P MAG P GRAR The classification probabilities obtained using EEG, MAG, and GRAD signals are represented respectively, and the multimodal classification result is as follows:
[0046]
[0047] Where p i It is obtained from the classification of each modality i, where P is the multimodal fusion classification probability;
[0048] By assigning higher weights to the patterns that best classify the data, multimodal motion imagery technology is achieved, which significantly improves the accuracy of motion imagery behavior classification.
[0049] In the above technical solution, the specific steps of S4 are as follows:
[0050] S41: Extract the coordinate sequence of the eye tracker for the current time period online, with the sequence duration set to 1 second;
[0051] S42: Use the triangular incremental method to find the minimum circle coverage for the coordinate sequence within the current 1 second. Set the center of the minimum circle to the virtual cursor position and use the Python pynput toolkit to control the cursor movement.
[0052] S42: Obtain a continuous sequence of motion-imagined behaviors based on the classification results of the classifier;
[0053] S43: Based on the association between motion imagination behavior and specific keyboard and mouse operations, common computer functions are realized.
[0054] In the above technical solution, the specific keyboard and mouse operation associations refer to the following: imagining your left hand making a fist corresponds to clicking the left mouse button; imagining your right hand making a fist corresponds to clicking the right mouse button; imagining both hands making fists corresponds to scrolling the scroll wheel upwards; imagining both hands opening corresponds to scrolling the scroll wheel downwards; imagining both feet opening corresponds to opening the on-screen keyboard; and imagining both feet closing corresponds to closing the on-screen keyboard.
[0055] Because the present invention employs the above-mentioned technical means, it has the following beneficial effects:
[0056] I. Improve the accuracy of operational feedback by optimizing the motor imagery behavior classification model. The above fusion method utilizes the complementarity of EEG and MEG to better identify ERD mechanisms.
[0057] In all frequency bands, mode type significantly affects AUC values (analysis of variance p < 10). -3 The number of features had no significant impact (p>0.05). The AUC value obtained by the fusion algorithm was significantly higher than that of any other mode, and the Tukey Kramer variance post-hoc test yielded p<0.016. The results showed that the best AUC was obtained in the α- and β-bands. In this example, the fused AUC value was significantly higher than the AUC obtained by EEG, MAG, and GRAD respectively. If the fusion effect is compared with the effect of the best single mode, the average improvement is 12.8%.
[0058] Second, the system maintains a high response speed, with an average response latency of 22ms as verified. Although multiple behavior classifiers are introduced, highly complex models are not used, keeping the time complexity low and ensuring the algorithm's high response speed.
[0059] Third, optimize the mouse positioning algorithm to improve system robustness and solve the problem of inaccurate mouse positioning. Use the triangular incremental method to solve for the minimum circle coverage within the current time period to achieve cursor debouncing and prevent users from accidentally clicking the cursor.
[0060] IV. Because the intensity of the brain's magnetic field outside the scalp is on the order of 10-100 fT, approximately one hundred millionth of the Earth's magnetic field, the requirements for magnetoencephalography (MEG) detection equipment are extremely high, demanding both good temporal and spatial resolution. The Elekta Neuromag TRIUX used for detecting MAG and GRAD, which includes a MEG shielding chamber and a cryogenic superconducting system, can acquire whole-brain MEG signals with high precision. The GRAD signal was measured by 204 gradient meters.
[0061] Introducing the tangential component GRAD to characterize MI features not only improves the spatial resolution of the signal but also facilitates the search for the most discriminative features, and better extracts the high-dimensional representation of the MI classifier, such as three-dimensional spatial features.
[0062] Gradient alignment is used when processing GRAD signals. Given eigenvalue multiplicity and eigenvector sign uncertainty, gradients calculated for two or more datasets (e.g., left and right brain regions) may not be directly comparable due to different eigenvector orders. Gradient alignment improves similarity and correspondence, thereby enhancing model classification accuracy.
[0063] Gradients can be aligned using Procrustes analysis or implicitly through joint embedding. It's important to note that if the manifold spaces differ significantly, alignment may not provide a reasonable output. In such cases, the algorithm's attention will be biased towards EEG and MAG to correct the bias.
[0064] Fifth, the technology incorporates multiple operation options to make it more user-friendly. These include: clicking the left and right mouse buttons, scrolling up and down, and a soft keyboard on / off switch.
[0065] VI. By using multimodal signal features, the classifier can distinguish and classify more behavioral patterns (such as hands and feet).
[0066] VII. More comprehensive utilization of human brain signals, opening up new directions for multimodal applications. EEG signals are prone to interference and noise due to the human body and environment. MEG, on the other hand, has higher spatial resolution, its electromagnetic field measurements are unaffected by the medium, and it can also perform three-dimensional spatial localization of the direction, location, and intensity of current sources, showing a trend towards replacing invasive examinations in this area. Attached Figure Description
[0067] Figure 1 This is a classification diagram of human brain multimodal motion imagery signals, representing a preferred embodiment of a method for computer control using multimodal motion imagery technology according to the present invention.
[0068] Figure 2 This is a schematic diagram of a preferred embodiment of a method for computer control using multimodal motion imagination technology according to the present invention.
[0069] Figure 3 This is a schematic diagram of a soft keyboard interface according to a preferred embodiment of a method for computer control using multimodal motion imagination technology according to the present invention.
[0070] Figure 4 This is a schematic diagram showing the classification contribution ratio of each mode in a preferred embodiment of a method for computer control using multimodal motion imagination technology according to the present invention. Detailed Implementation
[0071] The embodiments of the present invention will be described in detail below. Although the present invention will be described and illustrated in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, any modifications or equivalent substitutions made to the present invention should be covered within the scope of the claims of the present invention.
[0072] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art will understand that the present invention can be practiced without these specific details.
[0073] This invention addresses the problems of existing virtual mice and keyboards combined with eye trackers by providing a multimodal motion imagery method to improve the accuracy of motion imagery behavior classification while ensuring the system's robustness and high response speed.
[0074] We imagine the action as the left hand grasping the mouse as the left mouse button, the right hand grasping as the right mouse button, both hands grasping as scrolling the mouse wheel upwards, both hands open as scrolling the mouse wheel downwards, both legs open as opening the soft keyboard, and both legs together as closing the soft keyboard.
[0075] like Figure 1 The diagram shown is a classification diagram of human brain multimodal motor imagery signals, representing a preferred embodiment of a computer control method combining eye tracker and multimodal motor imagery technology according to the present invention.
[0076] Eye tracking, EEG, MAG, and GRAD represent different modalities of information. Eye tracking signals primarily enable mouse movement; that is, the mouse pointer moves wherever the eyes look. Multimodal motion imagery integrates EEG, MAG, and GRAD information, enabling left and right mouse clicks, vertical scrolling, on / off soft keyboard access, and text input. This method significantly improves the accuracy of motion imagery classification and enhances user efficiency through the use of multiple behavioral logics.
[0077] like Figure 2 The diagram shown is a preferred embodiment of a method for computer control using multimodal motion visualization technology according to the present invention.
[0078] Figure 3 A schematic diagram of the soft keyboard interface.
[0079] Figure 4 This is a multimodal contribution distribution map for a set of behavior classifiers, where the pie chart shows the λ of each modality obtained through the fusion method. i The value of .
[0080] For example, to open a file:
[0081] Method 1: Step 1: The user stares at the file, and the eye tracker aligns the mouse pointer with the file; Step 2: The user imagines the movement of their left hand twice; Step 3: Through multimodal motion imagery technology, it is determined that the user double-clicked the left mouse button twice, thus opening the file.
[0082] Method 2: Step 1 is the same as above; Step 2, the user imagines the movement of their right hand, and the multimodal motion imagination technology analyzes the user's intention to right-click on the file to open the function menu; Step 3, the user moves their gaze to the "Open" function, and the eye tracker aligns the mouse pointer with "Open"; Step 4, the user imagines the movement of their left hand to select the "Open" function.
[0083] Method 3: Step 1 is the same as above; Step 2, the user imagines their feet spreading apart, and multimodal motion imagery analysis shows that the user intends to open the soft keyboard; Step 3, the user moves their gaze to the "Enter" key on the soft keyboard, and the eye tracker aligns the mouse pointer with the "Enter" key; Step 4, the user imagines their left hand moving, thus selecting the "Enter" function.
[0084] To enable browsing web pages or PDF documents, a commonly used scrollbar function is needed. Step 1: The user places their gaze on the web page or PDF document; Step 2: The user imagines grasping the page with both hands, which is equivalent to scrolling upwards; Step 3: The user continues to imagine grasping the page with both hands, and the page continues to scroll until the desired position is reached, at which point the user stops imagining grasping the page with both hands.
[0085] In the above embodiment, the computer uses EEG and MEG detection equipment to detect the user's EEG, MAG, and GRAD and extracts a multimodal signal with a window length of 2000ms every 100ms. Then, it processes the extracted signal using a multi-window frequency transformation method based on discrete slepian sequences. After feature extraction using a permutation t-test algorithm based on clustering, the sample score is obtained by linear discriminant analysis. Subsequently, the motion imagery classification results of the three modalities are fused using a Bayesian fusion method based on weighted average, thereby improving the accuracy of behavior classification.
[0086] The above algorithm is used to train a classifier for motion imagination (MI) of each modal signal. Three classifiers are obtained and the final classification result is obtained by voting. The result is one of the following states: left hand, right hand, both hands grasping, both hands open, both legs open, and both legs together or idle.
[0087] Example 1
[0088] This invention provides a method for computer control using multimodal motion visualization technology, comprising the following steps:
[0089] Step S1: Use a multimodal signal acquisition device to record motion image signals;
[0090] Step S2: Preprocess and extract features from the extracted EEG, MAG, and GRAD multimodal signals using the algorithm;
[0091] Step S3: Construct a classifier by using a weighted average-based Bayesian fusion method to assign higher weights to the patterns that best classify the data.
[0092] Step S4: Using the classification results, associate the motion imagery signals with the left and right mouse buttons, scroll wheel function, and soft keyboard on / off function, and combine with the eye tracker to achieve remote control of the computer.
[0093] In the above technical solution, the specific steps of S1 are as follows:
[0094] S11: Using multimodal devices means simultaneously recording high-density EEG and magnetoencephalography (MEG) signals in a motor imagery (MI) task. The EEG and MEG data are recorded simultaneously using a 74-channel EEG system and an Elekta Neuromag TRIUX machine. The electrode positions on the scalp follow the 10-10 montage standard, and the reference electrode is located on the left scapula.
[0095] S12: Simultaneously consider the activity of electroencephalography (EEG) and magnetoencephalography (MEG), where the MEG consists of signals from a magnetometer and a gradient meter.
[0096] In the above technical solution, the specific steps of S2 are as follows:
[0097] S21: Locate the EEG, MAG, and GRAD signal electrodes and remove useless electrodes;
[0098] S22: Perform filtering, segmentation, noise reduction, and artifact removal operations on the raw EEG signal in sequence;
[0099] S22: Use MaxFilter to perform temporal-domain signal spatial separation on the raw magnetoencephalogram (MEG) signal to remove environmental noise;
[0100] S22: All signals from the magnetoencephalogram (MEG) are downsampled to 250 Hz and divided into 5-second cycles corresponding to the target cycle;
[0101] S23: Experts visually inspect the magnetoencephalogram (MEG) to remove artifacts. To better simulate the online scenario, no other artifact removal algorithms are used. After verification, all available frequency bands are retained.
[0102] S24: Using discrete oblate spheroid sequences as orthogonal window functions, a multi-window frequency transformation method is implemented to smooth the spectrum. This algorithm is implemented using the Fieldtrip toolbox.
[0103] S25: For each acquired signal period, extract the corresponding feature matrix M from the EEG, MAG, and GRAD power spectra. i EEG, MAG, and GRAD extract M i The dimensions are 74×36, 102×36 and 204×36, respectively;
[0104] S26: Use a feature extraction algorithm to extract features from matrix M i The most relevant features were extracted, focusing on the sensor signals of the contralateral motion region. The feature matrix sizes for EEG, MAG, and GRAD were 11×36, 15×36, and 30×36, respectively. Then, in M... i Nonparametric clustering-based permutation t-tests were performed between the dynamic and resting spectra to correct and eliminate low-quality signals and reduce the error rate. To this end, a statistical test threshold of p < 0.05 was set, and multiple comparisons were made to correct errors.
[0105] Finally, N was extracted within the standard frequency bands b = θ (4-7Hz), α (8-13Hz), β (14-29Hz), and γ (30-40Hz). f The most discriminative features are as follows:
[0106]
[0107] Where ζ i,b Represents the eigenvector, N f =1…10, with 10 as the maximum limit for feature dimensions, conforming to montage standards. It is the i-th selected feature, which is calculated and selected from the signals of each mode and each frequency band.
[0108] In the above technical solution, the specific steps of S3 are as follows:
[0109] S31: After feature extraction, classifiers are trained on the collected EEG signals, MAG magnetometer signals and GRAD gradient meter signals respectively. The motion imagination (MI) classifier for each modal signal is trained using the LDA algorithm with five-fold cross-validation. The three classifiers are then voted on to obtain the final classification result, which is one of the following states: left hand, right hand, both hands grasping, both hands open, both legs open, both legs together, or idle state.
[0110] In the LDA algorithm, there are N classes, the total number of samples is m, and the set of examples of the i-th class is X. i The number of samples in the corresponding category is m. i Define the global scatter matrix S t as follows:
[0111]
[0112] Where μ represents the mean vector of all samples, S w S is the intra-class scatter matrix b The inter-class scatter matrix;
[0113] Intraclass scatter matrix S wRepresented as:
[0114]
[0115] S w It is defined as the sum of the scatter matrices of each category, where μ i S represents the mean vector of category i, from which we can obtain S b The expression, that is:
[0116]
[0117] Optimize objective function J:
[0118] where w∈R d×(N-1) tr(.) represents the trace of the matrix;
[0119] When solving the problem, let the above expression tr(W) T S w W) = 1, then the Lagrange multiplier method can be used, where the larger J is, the more it conforms to the goal of LDA. In the formula, W and W T These are the projection matrix and the transpose, respectively.
[0120] S32: Use Bayesian fusion to integrate information from different modalities to obtain the classification probability of each classifier, and then obtain the behavior classification result;
[0121] A weighted average-based Bayesian fusion method was used to integrate information from different modalities, and the posterior probability p was linearly combined. i p i Obtained by the classification of each mode i, with parameter λ i The weighted average is expressed by the following formula:
[0122]
[0123] Where P EEG P MAG P GRAR The classification probabilities obtained using EEG, MAG, and GRAD signals are represented respectively, and the multimodal classification result is as follows:
[0124]
[0125] Where p i It is obtained from the classification of each modality i, where P is the multimodal fusion classification probability;
[0126] By assigning higher weights to the patterns that best classify the data, multimodal motion imagery technology is achieved, which significantly improves the accuracy of motion imagery behavior classification.
[0127] In the above technical solution, the specific steps of S4 are as follows:
[0128] S41: Extract the coordinate sequence of the eye tracker for the current time period online, with the sequence duration set to 1 second;
[0129] S42: Use the triangular incremental method to find the minimum circle coverage for the coordinate sequence within the current 1 second. Set the center of the minimum circle to the virtual cursor position and use the Python pynput toolkit to control the cursor movement.
[0130] S42: Obtain a continuous sequence of motion-imagined behaviors based on the classification results of the classifier;
[0131] S43: Based on the association between motion imagination behavior and specific keyboard and mouse operations, common computer functions are realized.
[0132] In the above technical solution, the specific keyboard and mouse operation associations refer to the following: imagining your left hand making a fist corresponds to clicking the left mouse button; imagining your right hand making a fist corresponds to clicking the right mouse button; imagining both hands making fists corresponds to scrolling the scroll wheel upwards; imagining both hands opening corresponds to scrolling the scroll wheel downwards; imagining both feet opening corresponds to opening the on-screen keyboard; and imagining both feet closing corresponds to closing the on-screen keyboard.
[0133] This invention achieves the following beneficial effects: It combines eye-tracking and multimodal motor imagery technology to enable remote computer control, allowing users to remotely control the computer using only their eyes and brain without taking any physical action. The use of multimodal motor imagery technology significantly improves the classification accuracy of user motor imagery and greatly reduces the occurrence of repeated motor imagery, making the technology more mature. This technology means that assistive devices for people with movement disorders are no longer limited to specific character spelling, which has significant practical implications.
[0134] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for achieving computer control using multimodal motion visualization technology, characterized in that... It includes the following steps: Step S1: Use a multimodal signal acquisition device to record motion image signals; Step S2: Use the algorithm to preprocess and extract features from the multimodal signals of the extracted EEG signal, MAG magnetometer signal, and GRAD gradient meter signal; Step S3: Construct a classifier by using a weighted average-based Bayesian fusion method to assign higher weights to the patterns that best classify the data. The specific steps for S3 are as follows: S31: After feature extraction, classifiers are trained on the collected EEG signals, MAG magnetometer signals and GRAD gradient meter signals respectively. The motion imagination (MI) classifier for each modal signal is trained using the LDA algorithm with five-fold cross-validation. The three classifiers are then voted on to obtain the final classification result, which is one of the following states: left hand, right hand, both hands grasping, both hands open, both legs open, both legs together, or idle state. In the LDA algorithm, there exists There are classes, and the total number of samples is . The first in the sample The class example collection is The number of samples in the corresponding category is Define the global scatter matrix as follows: in This represents the mean vector of all samples. The scatter matrix is the intra-class scatter matrix. The inter-class scatter matrix; Intraclass scatter matrix Represented as: It is defined as the sum of the scatter matrices for each category, where Indicate category The mean vector, from which we can obtain... The expression, that is: Optimize objective function : in , Represents the trace of a matrix; When solving the problem, let the above equation Then, the Lagrange multiplier method can be used, where... The larger the value, the more it aligns with the goals of LDA, as shown in the formula. and These are the projection matrix and the transpose, respectively. S32: Use Bayesian fusion to integrate information from different modalities to obtain the classification probability of each classifier, and then obtain the behavior classification result; A weighted average-based Bayesian fusion method was used to integrate information from different modalities, and the posterior probabilities were linearly combined. , By each mode The classification is obtained by parameters. The weighted average is expressed by the following formula: in The classification probabilities obtained using EEG, MAG, and GRAD signals are represented respectively, and the multimodal classification result is as follows: in By each mode Obtained through classification, For multimodal fusion classification probabilities; By assigning higher weights to patterns that best classify the data, multimodal motion visualization technology can be achieved. Step S4: Using the classification results, associate the motion imagery signals with the left and right mouse buttons, scroll wheel function, and soft keyboard on / off function, and combine with the eye tracker to achieve remote control of the computer.
2. The method for computer control using multimodal motion visualization technology according to claim 1, characterized in that, The specific steps for S1 are as follows: S11: Using multimodal devices means simultaneously recording high-density EEG and magnetoencephalography (MEG) signals in a motor imagery (MI) task. The EEG and MEG data are recorded simultaneously using a 74-channel EEG system and an Elekta Neuromag TRIUX machine. The electrode positions on the scalp follow the 10-10 montage standard, and the reference electrode is located on the left scapula. S12: Simultaneously consider the activity of electroencephalography (EEG) and magnetoencephalography (MEG), where the MEG consists of signals from a magnetometer and a gradient meter.
3. The method for computer control using multimodal motion visualization technology according to claim 1, characterized in that, The specific steps for S2 are as follows: S21: Locate the EEG, MAG, and GRAD signal electrodes and remove useless electrodes; S22: Perform filtering, segmentation, noise reduction, and artifact removal operations on the raw EEG signal in sequence; S22: Use MaxFilter to perform temporal-domain signal spatial separation on the raw magnetoencephalogram (MEG) signal to remove environmental noise; S22: All signals from the magnetoencephalogram (MEG) are downsampled to 250 Hz and divided into 5-second cycles corresponding to the target cycle; S23: Experts visually inspect the magnetoencephalogram (MEG) to remove artifacts. To better simulate the online scenario, no other artifact removal algorithms are used. After verification, all available frequency bands are retained. S24: Using discrete oblate spheroid sequences as orthogonal window functions, a multi-window frequency transformation method is implemented to smooth the spectrum. This algorithm is implemented using the Fieldtrip toolbox. S25: For each acquired signal period, extract the corresponding feature matrix from the EEG, MAG, and GRAD power spectra. EEG, MAG and GRAD extraction The dimensions are 74×36, 102×36 and 204×36, respectively; S26: Employing feature extraction algorithms from matrices Extracting the most relevant features, focusing on the sensor signals of the contralateral motion region, the feature matrix sizes of EEG, MAG, and GRAD are 11×36, 15×36, and 30×36, respectively. Then, in... A nonparametric, cluster-based permutation t-test is performed between the dynamic and resting spectra to correct and eliminate low-quality signals and reduce the error rate. To this end, a statistical test threshold is set. Errors were corrected through multiple comparisons; Ultimately in the standard frequency band , , , Extracted from the inside The most discriminative features are as follows: in Represents the eigenvector. The feature dimension was set to a maximum of 10, which conforms to the montage standard. The first choice Each feature is selected by calculating the signal for each mode and each frequency band.
4. A method for computer control using multimodal motion visualization technology according to claim 1, characterized in that, The specific steps for S4 are as follows: S41: Extract the coordinate sequence of the eye tracker for the current time period online, with the sequence duration set to 1 second; S42: Use the triangular incremental method to find the minimum circle coverage for the coordinate sequence within the current 1 second. Set the center of the minimum circle to the virtual cursor position and use the Python pynput toolkit to control the cursor movement. S42: Obtain a continuous sequence of motion-imagined behaviors based on the classification results of the classifier; S43: Based on the association between motion imagination behavior and specific keyboard and mouse operations, common computer functions are realized.
5. A method for computer control using multimodal motion visualization technology according to claim 4, characterized in that, The specific keyboard and mouse operation associations include imagining your left hand making a fist and clicking the left mouse button; imagining your right hand making a fist and clicking the right mouse button; imagining both hands making fists and scrolling the scroll wheel up; imagining both hands opening and scrolling the scroll wheel down; imagining both feet opening and opening the on-screen keyboard; and imagining both feet closing and closing the on-screen keyboard.
Citation Information
Patent Citations
Brain computer interface systems and methods of use thereof
CN110248697A
Method for realizing virtual mouse by combining eye tracker and asynchronous motor imagery technology
CN112764544A