Multi-modal intelligent wheelchair control method and system based on in-ear brain-computer interface
By collecting EEG, EMG, and eye movement signals from the ear canal through an in-ear brain-computer interface and combining them with multimodal signal fusion technology, the problem of insufficient operation complexity and precision of existing brain-controlled wheelchairs has been solved, realizing precise control of intelligent wheelchairs and humanized travel assistance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI YANGZHI REHABILITATION HOSPITAL
- Filing Date
- 2025-08-11
- Publication Date
- 2026-05-19
AI Technical Summary
Existing brain-controlled wheelchairs based on non-invasive brain-computer interfaces have shortcomings in terms of ease of operation, comfort, cost, accuracy, and stability. Furthermore, traditional scalp EEG devices suffer from low signal-to-noise ratio and limitations.
Using an in-ear brain-computer interface, the system collects electroencephalogram (EEG) signals in the user's ear canal, combines them with electromyography (EMG) and eye movement (EMG) signals, and uses multimodal signal fusion technology to decode intentions, thereby achieving precise control of the wheelchair.
It improves the sensitivity and accuracy of wheelchair control, provides highly intelligent and humanized travel assistance, adapts to the needs of patients in different conditions, and has the advantages of being portable, comfortable to wear, and easy to operate.
Smart Images

Figure CN120983220B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equipment control technology, specifically to a multimodal intelligent wheelchair control method and system based on an in-ear brain-computer interface. Background Technology
[0002] Electroencephalogram (EEG) and electromyogram (EMG) signals are common bio-information sources available for the human body, reflecting physiological states and intentions. Therefore, brain-controlled wheelchairs based on Brain-Computer Interface (BCI) technology can detect and decode EEG signals generated by brain activity, converting them into control commands to directly control the wheelchair. This is suitable for individuals with limb paralysis due to neuromuscular diseases or spinal cord injuries, significantly improving their freedom of movement and quality of life.
[0003] Brain-computer interfaces (BCIs) can be divided into invasive and non-invasive BCIs. Invasive BCIs can obtain higher resolution neural signals, but electrode implantation is invasive and increases the risk of complications, facing ethical and safety challenges. Non-invasive scalp EEG is a widely used non-invasive technology that uses a head-mounted electrode array to acquire signals, and has the advantages of low cost and overall high efficiency.
[0004] However, brain-controlled wheelchairs based on non-invasive brain-computer interfaces still face many challenges. User experience factors such as ease of operation and comfort affect product acceptance, while high-precision hardware manufacturing and complex software development lead to high costs. Furthermore, there are technical issues such as a lack of standardized testing, the need for extensive training of motor imagery paradigms with stability affected by various factors, the lack of intuitive control over evoked paradigms and their tendency to degrade performance, and limitations of traditional scalp EEG devices (such as electrode-related problems and low signal-to-noise ratio). Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal intelligent wheelchair control method and system based on an in-ear brain-computer interface. By using an in-ear signal acquisition method, it is closer to the user's brain signal source, thereby enabling the acquisition of more refined brain electrical activity and improving the quality of the acquired brain electrical signals. At the same time, it can also acquire richer, more delicate, and more comprehensive bioelectrical signals caused by subtle physiological movements such as electromyography and eye movements. In other words, the acquired ear canal electrical signals integrate multiple types of bioelectrical signal information, such as brain muscle activity information, to more comprehensively reflect the physiological activity characteristics of brain muscle fusion, providing a more expressive signal foundation for downstream decoding and control.
[0006] It facilitates improved sensitivity and accuracy of wheelchair control, enabling users with mobility impairments to achieve precise control of their wheelchairs through thought and physical activity; by automating and precisely classifying intentions such as starting, stopping, and turning of the wheelchair, it achieves precise movement control of intelligent wheelchairs in various complex scenarios, providing a highly intelligent and humanized travel assistance solution for people with mobility impairments.
[0007] Furthermore, based on the EEG signals from both ear canals and the mixed electrooculography and electromyography features therein, multiple control intentions can be decoded, and a brain-muscle fusion control method can be used to ensure the naturalness and flexibility of the user's wheelchair control. It can realize a variety of operation commands, such as start, stop, turn left, turn right, etc., to meet the needs of patients in different conditions.
[0008] To achieve the above objectives, the present invention provides a multimodal intelligent wheelchair control method based on an in-ear brain-computer interface, comprising: identifying a classification result representing the user's intention to control the wheelchair based on electroencephalogram (EEG) signals collected from the user's ear canal; and driving the wheelchair device to perform corresponding movement actions based on the classification result.
[0009] The present invention also provides a wheelchair control system, comprising: an in-ear signal acquisition device, a main control unit, and a drive system; the in-ear signal acquisition device is used to acquire electroencephalogram (EEG) signals in the user's ear canal and send them to the main control unit; the main control unit is used to execute the above-described multimodal intelligent wheelchair control method based on an in-ear brain-computer interface.
[0010] The present invention also provides a computer-readable storage medium, which is a non-volatile or non-transient storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface as described above.
[0011] In one embodiment, based on electroencephalogram (EEG) signals collected from the user's ear canal, a classification result characterizing the user's intention to control the wheelchair is identified, including:
[0012] The left and right ear canal EEG signals were collected from the user's left and right ear canals, respectively.
[0013] The left ear canal EEG signal and the right ear canal EEG signal are input into a preset intention classification model to obtain the classification result representing the user's intention to control the wheelchair.
[0014] In one embodiment, the preset intent classification model includes: a left ear feature extraction module, a right ear feature extraction module, a left ear feature fusion module, a right ear feature fusion module, a feature splicing module, and a fully connected classification module;
[0015] The left ear canal EEG signal and the right ear canal EEG signal are input into a preset intention classification model to obtain the classification result representing the user's intention to control the wheelchair, as output by the preset intention classification model, including:
[0016] The left ear canal EEG signal is input to the left ear feature extraction module, and the left ear feature extraction module extracts the left ear EEG features from the left ear canal EEG signal and inputs them to the left ear feature fusion module and the right ear feature fusion module respectively.
[0017] The right ear canal EEG signal is input to the right ear feature extraction module, and the right ear feature extraction module inputs the right ear EEG features extracted from the right ear canal EEG signal to the left ear feature fusion module and the right ear feature fusion module respectively;
[0018] The left ear feature fusion module fuses the left ear EEG features and the right ear EEG features to obtain a first fused signal feature, and the right ear feature fusion module fuses the left ear EEG features and the right ear EEG features to obtain a second fused signal feature;
[0019] The feature splicing module splices the first fused signal feature and the second fused signal feature to obtain global features of the left and right ear canals, and inputs them into the fully connected classification module.
[0020] The fully connected classification module obtains classification results that characterize the user's intention to control the wheelchair based on the global features of the left and right ear canal signals.
[0021] In one embodiment, the left ear feature fusion module includes: a first cross-attention module and a first feedforward mapping module connected in sequence; the right ear feature fusion module includes: a second cross-attention module and a second feedforward mapping module connected in sequence.
[0022] The left ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a first fused signal feature, and the right ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a second fused signal feature, including:
[0023] The first cross-attention module performs cross-attention calculation on the left ear EEG features and the right ear EEG features to obtain the first cross feature. The first feedforward mapping module mixes the first cross feature with channel dimension information to obtain the first fused signal feature.
[0024] The second cross-attention module performs cross-attention calculation on the left ear EEG features and the right ear EEG features to obtain the second cross features. The second feedforward mapping module mixes the second cross features with channel dimension information to obtain the second fused signal features.
[0025] In one embodiment, the wheelchair device is equipped with a display device and an image acquisition device; the method further includes:
[0026] The image acquisition device acquires the user's gaze position on the display device, and the target position selected by the user is determined based on the gaze position.
[0027] When in navigation mode, a navigation path is planned and the wheelchair device is controlled to move based on the current location and the target location.
[0028] In one embodiment, the display area of the display device is divided into multiple sub-display areas;
[0029] Based on the gaze position, the target position selected by the user is determined, including:
[0030] Determine the sub-display area to which the gaze position belongs, and take the preset position corresponding to the sub-display area to which the gaze position belongs as the target position.
[0031] In one embodiment, acquiring the user's gaze position on the display device via the image acquisition device includes:
[0032] The image acquisition device acquires a user's facial image and inputs the facial image into a preset gaze position recognition model to obtain the user's gaze position on the display device, as output by the gaze position recognition model.
[0033] In one embodiment, the gaze position recognition model includes: a convolutional neural network, a feature fusion module, a multilayer perceptron, a fully connected layer, and a nonlinear activation function;
[0034] The facial image is input into a preset gaze position recognition model to obtain the user's gaze position on the display device, as output by the gaze position recognition model, including:
[0035] An eye region image containing the user's eye features is obtained from the facial image, and the eye features are extracted from the eye region image using the convolutional neural network;
[0036] The feature fusion module is used to fuse the eye features with the non-eye features extracted from the facial image, and the fused features are then input into the multilayer perceptron.
[0037] The output of the multilayer perceptron is passed through the fully connected layer and the nonlinear activation function to obtain the user's gaze position on the display device.
[0038] In one embodiment, the complexity of the task to be performed is determined based on the current location and the target location;
[0039] When the task complexity is high, enter navigation mode; when the task complexity is low, enter brain-computer interface control mode.
[0040] In one embodiment, the user's alertness level is determined based on the collected EEG signals of the user;
[0041] If the user's alertness level is lower than the preset alertness level, the system enters navigation mode; if the user's alertness level is higher than the preset alertness level, the system enters brain-computer interface control mode.
[0042] In one embodiment, the in-ear signal acquisition device includes: a left ear acquisition unit and a right ear acquisition unit; the left ear acquisition unit is used to acquire the left ear canal electrical signal of the user's left ear canal;
[0043] The right ear acquisition unit is used to acquire the electrical signal of the user's right ear canal. Attached Figure Description
[0044] Figure 1 This is a detailed flowchart of the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to the first embodiment of the present invention;
[0045] Figure 2 yes Figure 1 A detailed flowchart of step 101 in the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface;
[0046] Figure 3 This is a schematic diagram of a preset intent classification model according to the first embodiment of the present invention;
[0047] Figure 4 This is a flowchart illustrating the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to the second embodiment of the present invention.
[0048] Figure 5 This is a schematic diagram of a wheelchair control system according to a third embodiment of the present invention. Detailed Implementation
[0049] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings to provide a clearer understanding of the purpose, features, and advantages of the present invention. It should be understood that the embodiments shown in the drawings are not intended to limit the scope of the present invention, but are merely illustrative of the essential spirit of the technical solution of the present invention.
[0050] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known apparatuses, structures, and techniques associated with this application may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.
[0051] Unless the context requires otherwise, throughout the specification and claims, the word “comprising” and its variations, such as “including” and “having”, shall be understood to have an open, inclusive meaning, that is, to be interpreted as “including, but not limited to”.
[0052] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.
[0053] The singular forms “a” and “the” used in this specification and the appended claims include plural references unless otherwise expressly stated herein. It should be noted that the term “or” is generally used to include the meaning of “or / and” unless otherwise expressly stated herein.
[0054] In the following description, in order to clearly demonstrate the structure and working method of the present invention, a number of directional terms will be used. However, terms such as "front", "back", "left", "right", "outer", "inner", "outer", "inner", "up", and "down" should be understood as convenient terms and not as limiting terms.
[0055] The first embodiment of this invention relates to a multimodal intelligent wheelchair control method based on an in-ear brain-computer interface. This method is applied to a main control unit for controlling a wheelchair device. The main control unit is, for example, a controller or processor. The main control unit is connected to an in-ear signal acquisition device, such as an earphone, which includes a flexible electrode that can be inserted into the user's ear canal. This flexible electrode is biocompatible and is placed at a designated location in the user's ear canal. For example, the flexible electrode is attached to a soft, elastic cylindrical body, so that when the cylindrical body is inserted into the user's ear canal, the flexible electrode can adhere to the ear canal wall. Furthermore, the wheelchair device includes a drive system and a rechargeable and / or replaceable power supply. The drive system and power supply are mounted on the wheelchair body, and the power supply powers the devices on the wheelchair device. It should also be noted that the wheelchair device further includes other components of a conventional wheelchair, such as footrests, backrest, drive unit, and wheels, which are similar to those conventionally used in the art and will not be described in detail here. Any wheelchair device used in the art that does not conflict with the technical solution of this application can be used in this application.
[0056] The specific process of the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface in this embodiment is as follows: Figure 1 As shown, when a user first uses the wheelchair device, the device can collect the user's basic physiological signal characteristics and update the user's characteristic data in real time during the user's use of the wheelchair device in order to dynamically adjust the control parameters.
[0057] Step 101: Based on the EEG signals collected from the user's ear canal, identify the classification results that characterize the user's intention to control the wheelchair.
[0058] Specifically, the in-ear signal acquisition device can connect wirelessly (via Bluetooth, Wi-Fi, etc.) to the main control unit. When the user is using the wheelchair device, the in-ear signal acquisition device can collect the user's ear canal electrical signals in real time through flexible electrodes inserted into the user's ear canal. These ear canal electrical signals include pure electroencephalogram (EEG) signals, as well as electrooculogram (EOG) signals caused by eye movements and electromyogram (EMG) signals generated by subtle facial muscle movements such as teeth clenching, thus integrating multiple types of physiological and electrical information. The in-ear signal acquisition device sends the EEG signals collected from the ear canal to the main control unit. The main control unit converts the EEG signals from the ear canal into digital signals and then identifies the classification results representing the user's intention to control the wheelchair from them.
[0059] In one example, after receiving the ear canal EEG signal sent by the in-ear signal acquisition device, the main control unit first preprocesses the ear canal EEG signal. The preprocessing methods include: low-noise amplification and / or bandpass filtering. The low-noise amplifier amplifies the weak signal in the ear canal EEG signal to reduce noise interference. The bandpass filtering uses a bandpass filter with a preset band to remove high-frequency noise and low-frequency drift in the ear canal EEG signal. The preset band is, for example, 0.5-60Hz.
[0060] In one example, such as Figure 2 As shown, step 101 includes the following sub-steps:
[0061] Sub-step 1011: Acquire the EEG signals from the left and right ear canals of the user, respectively.
[0062] Sub-step 1012: Input the left ear canal EEG signal and the right ear canal EEG signal into a preset intention classification model to obtain the classification result representing the user's intention to control the wheelchair output by the preset intention classification model.
[0063] Specifically, the in-ear signal acquisition device includes a left ear signal acquisition device and a right ear signal acquisition device. The left ear signal acquisition device is worn by the user in the left ear and can acquire the electrical signals of the left ear canal. The right ear signal acquisition device is worn by the user in the right ear and can acquire the electroencephalogram (EEG) signals of the right ear canal. Both the left and right ear canal EEG signals contain mixed EOG signals caused by eye movements and EMG signals generated by subtle facial muscle movements such as teeth clenching.
[0064] The changes in neural activity caused by different psychological or physiological states such as relaxation, tension, eye movement to the left or right, and teeth clenching can be reflected in the electroencephalogram (EEG) signals collected from the left and right ear canals, manifesting as different electrophysiological characteristic patterns.
[0065] Both left and right ear canal EEG signals were sent to the main control unit, which contained a pre-defined intent classification model, an end-to-end network EECAF; please refer to [reference needed]. Figure 3 The preset intent classification model includes: a left ear feature extraction module, a right ear feature extraction module, a left ear feature fusion module, a right ear feature fusion module, a feature splicing module, and a fully connected classification module.
[0066] Specifically, the left ear canal EEG signal is input to the left ear feature extraction module, and the left ear feature extraction module extracts the left ear EEG features from the left ear canal EEG signal and inputs them to the left ear feature fusion module and the right ear feature fusion module respectively; wherein, the left ear feature extraction module is a one-dimensional convolutional neural network (or other neural network structure capable of extracting temporal signal features), which can extract temporal and spatial features from the left ear canal EEG signal.
[0067] The right ear feature extraction module is input to the right ear feature extraction module, and the right ear feature extraction module extracts the right ear EEG features from the right ear canal EEG signal and inputs them to the left ear feature fusion module and the right ear feature fusion module respectively; wherein, the right ear feature extraction module is a one-dimensional convolutional neural network (or other neural network structure that can extract temporal signal features), which can extract temporal and spatial features from the right ear EEG signal.
[0068] The left ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a first fused signal feature, and the right ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a second fused signal feature.
[0069] Specifically, the left ear feature fusion module includes a first cross-attention module and a first feedforward mapping module connected in sequence; the right ear feature fusion module includes a second cross-attention module and a second feedforward mapping module connected in sequence. Further, a residual connection and a layer normalization module (Add&Norm) are connected between the first cross-attention module and the first feedforward mapping module, and a residual connection and a layer normalization module (Add&Norm) are connected between the first feedforward mapping module and the feature splicing module.
[0070] The left ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a first fused signal feature, and the right ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a second fused signal feature, including:
[0071] The first cross-attention module performs cross-attention calculation on the left ear EEG features and the right ear EEG features to obtain a first cross-feature. The first feedforward mapping module (i.e., the feedforward network) mixes the first cross-feature with channel dimension information to obtain a first fused signal feature. Specifically, the first cross-attention module is a multi-head cross-attention module, used to perform deep fusion processing on the left and right ear EEG signals to fully explore and integrate the key features in the signals. In the first cross-attention module, the left ear EEG features corresponding to the left ear canal EEG signal are used as the Q query mode, and the right ear EEG features corresponding to the right ear EEG signal are used as the K key and V value. The multi-head cross-attention mechanism is used to update and fuse the temporal and spatial features in the left and right ear EEG features to obtain a first fused signal feature containing deep correlation information between the left and right ear EEG signals.
[0072] The second cross-attention module performs cross-attention calculations on the left ear EEG features and the right ear EEG features to obtain a second cross-feature. The second feedforward mapping module (i.e., the feedforward network) mixes the second cross-feature with channel dimension information to obtain a second fused signal feature. Similarly, the second cross-attention module is a multi-head cross-attention module used to perform deep fusion processing on the left and right ear EEG signals to fully mine and integrate key features in the signals. In the second cross-attention module, the right ear EEG features corresponding to the right ear canal EEG signal are used as the Q-query mode, and the left ear EEG features corresponding to the left ear EEG signal are used as the K-key and V-value. Using the multi-head cross-attention mechanism, the temporal and spatial features in the left and right ear EEG features are updated and fused, which can also obtain a second fused signal feature containing deep correlation information between the left and right ear EEG signals.
[0073] As shown above, the EEG signals from the left and right ear canals are respectively fed into two parallel cross-modal feature fusion modules (i.e., Cross-Modal Attention Blocks, specifically the left ear feature fusion module CAB1 and the right ear feature fusion module CAB2). Both CAB1 and CAB2 can model the correlation between the EEG signals from the left and right ear canals, achieving deep fusion of multi-channel brain signals. Both CAB1 and CAB2 include cross-modal attention mechanisms, residual connections and layer normalization, as well as feedforward mapping structures, maintaining key signal features of the current channel while fusing information from the contralateral side. In this way, the system can effectively improve its ability to recognize user intentions and enhance its robustness to changes in multi-source signals.
[0074] The feature splicing module performs a fully connected splicing of the first fused signal features and the second fused signal features to obtain global features of the left and right ear canal EEG signals, and inputs them into the fully connected classification module. Then, the fully connected classification module obtains a classification result representing the user's intention to control the wheelchair based on the global features of the left and right ear canal signals. The first and second fusion signal features are combined by a feature splicing module to obtain a global feature that simultaneously includes EEG signals from both ear canals. This global feature reflects the EEG, electromyography (EMG), and electrooculography (EOG) signals in the left and right ear canals, achieving effective fusion of these signal features. The global feature can characterize the user's voluntary thought activities, eye movements, or asymmetrical body movements such as teeth clenching. The wheelchair device is pre-set with control intentions corresponding to different action paradigms (i.e., biological actions). By classifying the global feature and the action paradigms corresponding to the preset global features, the main control unit presets the action paradigms corresponding to the rhythmic changes in EEG signals caused by voluntary thought activities, which correspond to the control of the wheelchair device's start, stop, and turn operations. Thus, the ear canal EEG signals are identified to obtain the classification results representing their corresponding EEG signals. Based on these classification results, the user's start / stop control intentions for controlling the wheelchair device's start, stop, left turn, and right turn can be determined.
[0075] As described above, the fused features from the corresponding channels of the left and right ear canals are merged through the feature splicing module to form a unified global representation. This representation is further input into the fully connected classification module (FC) to perform multi-class recognition of the user's movement intentions. The system can output four types of wheelchair control commands, including forward, left turn, right turn, and stop, realizing non-invasive, real-time, intention-driven human-computer interaction control functions.
[0076] In practical applications, the feature extraction module and the feature fusion module can be flexibly configured according to specific usage scenarios. For example, different parameters such as network depth, feature dimension or number of attention heads can be set to adapt to the brain signal feature distribution and signal quality of different users, thereby improving the individual adaptability and control accuracy of the system.
[0077] In practical applications, the feature encoder and CAB module can be flexibly configured according to specific usage scenarios. For example, different parameters such as network depth, feature dimension or number of attention heads can be set to adapt to the brain signal feature distribution and signal quality of different users, thereby improving the individual adaptability and control accuracy of the system.
[0078] Step 102: Based on the classification results, drive the wheelchair device to perform the corresponding movement actions.
[0079] Specifically, the classification results of user action paradigms characterize the user's start / stop and / or steering control intentions. Based on these classification results, corresponding control commands can be generated and sent to the wheelchair's drive system. The drive system then executes these commands to drive the wheelchair to perform the corresponding movement. For example, if the classification result indicates a left turn intention, the main control unit will control the drive system to turn the wheelchair to the left. This system can autonomously decode and generate commands based on the user's natural brainwave changes, thereby achieving high-precision movement control of the intelligent wheelchair in complex environments and improving the user's independent living ability and travel efficiency.
[0080] In this embodiment, the in-ear signal acquisition method is closer to the user's brain signal source, thereby enabling the acquisition of more refined brain electrical activity and improving the quality of the acquired brain electrical signals. At the same time, it can also acquire richer, more delicate and comprehensive bioelectrical signals caused by subtle physiological movements such as electromyography and eye movements. In other words, the acquired ear canal electrical signals integrate multiple types of bioelectrical signal information, such as brain muscle activity information, to more comprehensively reflect the physiological activity characteristics of brain muscle fusion, providing a more expressive signal foundation for downstream decoding and control.
[0081] It facilitates improved sensitivity and accuracy of wheelchair control, enabling users with mobility impairments to achieve precise control of their wheelchairs through thought and physical activity; by automating and precisely classifying intentions such as starting, stopping, and turning of the wheelchair, it achieves precise movement control of intelligent wheelchairs in various complex scenarios, providing a highly intelligent and humanized travel assistance solution for people with mobility impairments.
[0082] Furthermore, based on the decoding of multiple control intentions using EEG signals from both ear canals and the brain-muscle fusion control method, the system ensures the naturalness and flexibility of wheelchair control for users, enabling various operation commands such as start, stop, left turn, and right turn to meet the needs of patients in different conditions.
[0083] Among them, the wheelchair control is achieved by collecting electrical signals in the ear canal through in-ear electrodes. It has the advantages of being portable, comfortable to wear, and easy to operate, making it suitable for long-term wear; it is also more suitable for people with ALS and similar mobility impairments.
[0084] The second embodiment of the present invention relates to a multimodal intelligent wheelchair control method based on an in-ear brain-computer interface. Compared with the first embodiment, this embodiment provides a navigation method for a wheelchair device.
[0085] The specific process of the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface in this embodiment is as follows: Figure 4As shown; in this embodiment, the wheelchair device is equipped with a user-facing display device, which displays content to the user; the wheelchair device is also equipped with an image acquisition device (e.g., a camera), which can capture the user's facial image.
[0086] Step 201: Obtain the user's gaze position on the display device through the image acquisition device, and determine the target position selected by the user based on the gaze position.
[0087] Specifically, when a user triggers target location selection, the user can be prompted to determine the target location by looking at the display device. At this time, the image acquisition device first obtains the user's gaze position on the display device; specifically including:
[0088] The image acquisition device acquires a user's facial image and inputs it into a preset gaze position recognition model to obtain the user's gaze position on the display device, output by the gaze position recognition model. Specifically, the gaze position recognition model is a pre-trained model, which can be a deep learning model, a neural network model, etc. The model training process is as follows:
[0089] First, a training set consisting of multiple sample data is constructed. Each sample data includes: a corresponding facial image and a sample gaze position on the display device. Each sample data is obtained by: capturing the user's facial image through a camera when the user is sitting in the wheelchair device; at the same time, the coordinates of the position the user wants to gaze at (i.e., the sample gaze position) are marked on the display area to obtain a sample data.
[0090] Build an initial recognition model, which can be a deep learning model, a neural network model, etc.; its input is a facial image and its output is the gaze position.
[0091] The initial recognition model is trained using the training set mentioned above until the number of iterations reaches the set number or the error is less than the preset value. The training is then considered complete, and the initial recognition model obtained is used as the gaze position recognition model.
[0092] The gaze position recognition model is then configured into the wheelchair device, so that facial images can be input into the gaze position recognition model, and the position information output by the gaze position recognition model is the user's gaze position on the display device.
[0093] For example, the gaze position recognition model includes: a convolutional neural network, a feature fusion module, a multilayer perceptron, a fully connected layer, and a nonlinear activation function; during training, the gaze position recognition model adopts an end-to-end training method, and the parameters of all modules are jointly optimized through the backpropagation algorithm.
[0094] The facial image is input into a preset gaze position recognition model to obtain the user's gaze position on the display device, as output by the gaze position recognition model, including:
[0095] An eye region image containing the user's eye features is obtained from the facial image, and the eye features are extracted from the eye region image using the convolutional neural network. Specifically, by detecting facial key points, especially feature points around the eyes, a rectangular region of a fixed size containing both eyes can be accurately cropped out, and this rectangular region image is the eye region image.
[0096] The feature fusion module is used to fuse the eye features with the non-eye features extracted from the facial image, and the fused features are then input into the multilayer perceptron.
[0097] The output of the multilayer perceptron is passed through the fully connected layer and the nonlinear activation function to obtain the user's gaze position on the display device, for example, represented by the two-dimensional coordinates of the gaze point in the display device coordinate system.
[0098] Furthermore, when it is detected that a user gazes at a certain location for more than a set time threshold (such as 2 seconds), that location is determined to be the gaze location.
[0099] Subsequently, based on the gaze position, the target position selected by the user is determined; specifically as follows:
[0100] The display area of the display device is divided into multiple sub-display areas, each corresponding to a preset location. These preset locations may include, for example, the living room location, kitchen location, bedroom location, refrigerator location, or doorway location. For instance, some or all of the display area may be divided into multiple rectangular sub-display areas, each corresponding to a preset location. In navigation mode, prompt text indicating the preset location can be displayed within each sub-display area.
[0101] After obtaining the user's gaze position on the display area, the sub-display area to which the gaze position belongs is determined, and the preset position corresponding to the sub-display area to which the gaze position belongs is taken as the target position; the main control unit enters navigation mode.
[0102] Step 202: When in navigation mode, plan a navigation path and control the wheelchair device to move based on the current location and the target location.
[0103] Specifically, when in navigation mode, the wheelchair device enters navigation mode and prompts the user to determine the target location by looking at the display device.
[0104] After determining the target location, the current location of the wheelchair device can be obtained, and then at least one navigation path from the current location to the target location can be planned according to the preset navigation algorithm.
[0105] Specifically, the wheelchair device integrates multiple sensors such as LiDAR and cameras, and employs the RTAB-Map (Real-Time Appearance-Based Mapping) algorithm for real-time localization and mapping. In practical use, the system first completes the mapping process in the target environment, generating a two-dimensional grid map for navigation planning. The RTAB-Map algorithm fuses visual and LiDAR information to construct a sparse graph structure for loop closure detection and pose optimization, thereby obtaining and storing a high-precision environmental map. After determining the target location, based on the current and target locations, it can plan at least one navigation path from the current location to the target location within the constructed environmental map. This allows the system to autonomously plan the optimal path and then navigate according to the planned path, intelligently avoiding obstacles during the journey.
[0106] In this embodiment, the user can select either a navigation mode or a brain-computer interface control mode. For example, when the user does not perform any operation on the wheelchair device, the wheelchair device is in a resting state. In the resting state, the wheelchair device will not accept user control. The user can change the state of the wheelchair device through preset actions, such as a teeth-clenching action to switch to a stopped state. In the stopped state, the wheelchair is in brain-computer interface control mode. At this time, the main control unit can identify the classification result representing the user's intention to control the wheelchair based on the EEG signals collected from the user's ear canal; and drive the wheelchair device to perform the corresponding movement action according to the classification result. This part is similar to steps 101 and 102 in the first embodiment, and will not be described in detail here.
[0107] Users can also trigger the exit from the brain-computer interface control mode through preset actions. At this time, users can switch the state of the wheelchair device through other actions, such as switching to navigation mode or to resting state through eye movements. Thus, users can actively choose between navigation mode and brain-computer interface control mode.
[0108] Furthermore, during the process of the wheelchair device selecting the optimal path for navigation according to the planned navigation route, if the navigation is completed, navigation fails (e.g., the wheelchair is trapped), or the user triggers a preset action (e.g., a long clenching of teeth), the main control unit controls the wheelchair device to enter a resting state.
[0109] In this embodiment, under brain-computer interface control mode, attention monitoring can be incorporated during movement. For example, based on the collected ear canal electrical signals, it can be determined whether the user is fatigued or distracted. Specifically, the characteristics and energy ratios of the delta, theta, alpha, and beta bands in the ear canal electrical signals are analyzed to assess the patient's alertness in real time. Theta waves increase significantly when fatigued or distracted, while alpha waves dominate in a relaxed state with eyes closed, but exhibit "alpha inhibition" when performing tasks or concentrating. Beta waves reflect increased alertness and cognitive load, typically increasing significantly during focused tasks. Therefore, in a conscious and focused state, beta wave energy is high, theta wave energy is low, and alpha wave activity is suppressed. The main control unit can use this to determine whether the user is fatigued or distracted. If fatigue or distraction is detected during real-time control, the main control unit will automatically switch the wheelchair to a resting state and stop movement. Thus, the introduction of the attention monitoring mechanism greatly enhances the wheelchair's active safety protection capabilities in complex application scenarios.
[0110] Furthermore, during wheelchair movement, a forward-facing depth camera can detect the relative position of obstacles (including static and dynamic obstacles) to the wheelchair device. If the distance between the wheelchair device and the obstacle is detected to be lower than a safety threshold, the main control unit will automatically switch the wheelchair to a resting state and stop moving to avoid collision accidents.
[0111] This embodiment realizes multimodal fusion control, that is, the combination of brain-computer interface control mode and navigation mode; in brain-computer interface control mode, users can directly control the movement of the wheelchair device by closing their eyes or in a relaxed state, clenching their teeth, or moving their eyes, which facilitates short-distance fine-tuning of the wheelchair device.
[0112] In navigation mode, users can also select target locations by gazing, thereby achieving automatic navigation. This is suitable for patients who cannot input information through voice or manual triggering, enabling these patients to independently complete long-distance travel through navigation.
[0113] This multimodal interaction method ensures that the wheelchair can meet the needs of precise control and provide automated driving capabilities, thus improving the overall user experience.
[0114] The third embodiment of the present invention relates to a wheelchair control system, please refer to... Figure 5 The wheelchair control system includes: an in-ear signal acquisition device 1, a main control unit 2, and a wheelchair device 3.
[0115] The in-ear signal acquisition device 1, for example, is in the form of an earphone. It includes a flexible electrode that can be inserted into the user's ear canal. This flexible electrode is biocompatible and is placed in a designated position in the user's ear canal. For example, the flexible electrode is attached to a soft, elastic cylindrical body, so that when the cylindrical body is inserted into the user's ear canal, the flexible electrode can adhere to the ear canal wall. Additionally, the wheelchair device includes a drive system and a rechargeable and / or replaceable power supply. The drive system and power supply are mounted on the wheelchair body, and the power supply powers the devices on the drive system. It should also be noted that the drive system also includes other components of a conventional wheelchair, such as footrests and a backrest, which are similar to those conventionally used in the art and will not be described further here.
[0116] The in-ear signal acquisition device 1 is used to acquire the electrical signals of the user's ear canal and send them to the main control unit 2. The in-ear signal acquisition device includes a left ear acquisition unit and a right ear acquisition unit; the left ear acquisition unit is used to acquire the electrical signals of the user's left ear canal; and the right ear acquisition unit is used to acquire the electrical signals of the user's right ear canal.
[0117] The main control unit 2 is used to execute the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface described in the first embodiment.
[0118] The drive system is used to drive the drive system to perform movement actions, including but not limited to: forward, backward, left turn, right turn, etc.
[0119] Since the first and second embodiments correspond to this embodiment, this embodiment can be implemented in conjunction with the first and second embodiments. The relevant technical details mentioned in the first and second embodiments remain valid in this embodiment, and the technical effects achievable in the first embodiment can also be achieved in this embodiment. To reduce repetition, they will not be repeated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.
[0120] The fourth embodiment of the present invention relates to a computer-readable storage medium, which is a non-volatile or non-transient storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface as described in the first embodiment.
[0121] The preferred embodiments of the present invention have been described in detail above, but it should be understood that, if necessary, aspects of the embodiments can be modified to utilize aspects, features, and concepts from various patents, applications, and publications to provide other embodiments.
[0122] In light of the detailed description above, these and other changes can be made to the embodiments. Generally, the terminology used in the claims should not be considered as limiting to the specific embodiments disclosed in the specification and claims, but should be understood to include all possible embodiments together with the full scope of equivalents enjoyed by these claims.
Claims
1. A multimodal intelligent wheelchair control method based on an in-ear brain-computer interface, characterized in that, include: The left and right ear canal EEG signals were collected from the user's left and right ear canals, respectively. The left ear canal EEG signal and the right ear canal EEG signal are input into a preset intention classification model to obtain the classification result representing the user's intention to control the wheelchair output by the preset intention classification model. Based on the classification results, the wheelchair device is driven to perform corresponding movement actions; The preset intent classification model includes: a left ear feature extraction module, a right ear feature extraction module, a left ear feature fusion module, a right ear feature fusion module, a feature splicing module, and a fully connected classification module; The left ear canal EEG signal and the right ear canal EEG signal are input into a preset intention classification model to obtain the classification result representing the user's intention to control the wheelchair, as output by the preset intention classification model, including: The left ear canal EEG signal is input to the left ear feature extraction module, and the left ear feature extraction module extracts the left ear EEG features from the left ear canal EEG signal and inputs them to the left ear feature fusion module and the right ear feature fusion module respectively. The right ear canal EEG signal is input to the right ear feature extraction module, and the right ear feature extraction module inputs the right ear EEG features extracted from the right ear canal EEG signal to the left ear feature fusion module and the right ear feature fusion module respectively; The left ear feature fusion module uses the left ear EEG features as the Q query mode and the right ear EEG features as the K key and V value. It adopts a multi-head cross-attention mechanism to update and fuse the temporal and spatial features of the left ear EEG features and the right ear EEG features to obtain the first fused signal features. The right ear feature fusion module uses the right ear EEG features as the Q query mode and the left ear EEG features as the K key and V value. It adopts a multi-head cross-attention mechanism to update and fuse the temporal and spatial features of the left ear EEG features and the right ear EEG features to obtain the second fused signal features. The feature splicing module splices the first fused signal feature and the second fused signal feature to obtain global features of the left and right ear canal EEG signals, and inputs them into the fully connected classification module; The fully connected classification module obtains classification results that characterize the user's intention to control the wheelchair based on the global features of the left and right ear canal EEG signals.
2. The multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to claim 1, characterized in that, The left ear feature fusion module includes: a first cross-attention module and a first feedforward mapping module connected in sequence; the right ear feature fusion module includes: a second cross-attention module and a second feedforward mapping module connected in sequence. The left ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a first fused signal feature, and the right ear feature fusion module fuses the left ear EEG features with the right ear EEG features to obtain a second fused signal feature, including: The first cross-attention module performs cross-attention calculation on the left ear EEG features and the right ear EEG features to obtain the first cross feature. The first feedforward mapping module mixes the first cross feature with channel dimension information to obtain the first fused signal feature. The second cross-attention module performs cross-attention calculation on the left ear EEG features and the right ear EEG features to obtain the second cross features. The second feedforward mapping module mixes the second cross features with channel dimension information to obtain the second fused signal features.
3. The multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to claim 1, characterized in that, The wheelchair device is equipped with a display device and an image acquisition device; the method further includes: The image acquisition device acquires the user's gaze position on the display device, and the target position selected by the user is determined based on the gaze position. When in navigation mode, a navigation path is planned and the wheelchair device is controlled to move based on the current location and the target location.
4. The multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to claim 3, characterized in that, The display area of the display device is divided into multiple sub-display areas; Based on the gaze position, the target position selected by the user is determined, including: Determine the sub-display area to which the gaze position belongs, and take the preset position corresponding to the sub-display area to which the gaze position belongs as the target position.
5. The multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to claim 3, characterized in that, Obtaining the user's gaze position on the display device through the image acquisition device includes: The image acquisition device acquires a user's facial image and inputs the facial image into a preset gaze position recognition model to obtain the user's gaze position on the display device, as output by the gaze position recognition model.
6. The multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to claim 5, characterized in that, The gaze position recognition model includes: a convolutional neural network, a feature fusion module, a multilayer perceptron, a fully connected layer, and a nonlinear activation function; The facial image is input into a preset gaze position recognition model to obtain the user's gaze position on the display device, as output by the gaze position recognition model, including: An eye region image containing the user's eye features is obtained from the facial image, and the eye features are extracted from the eye region image using the convolutional neural network; The feature fusion module is used to fuse the eye features with the non-eye features extracted from the facial image, and the fused features are then input into the multilayer perceptron. The output of the multilayer perceptron is passed through the fully connected layer and the nonlinear activation function to obtain the user's gaze position on the display device.
7. The multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to claim 4, characterized in that, Based on the current location and the target location, determine the complexity of the task to be performed; When the task complexity is high, enter navigation mode; when the task complexity is low, enter brain-computer interface control mode.
8. The multimodal intelligent wheelchair control method based on an in-ear brain-computer interface according to claim 3, characterized in that, The user's alertness level is determined based on the collected EEG signals; If the user's alertness level is lower than the preset alertness level, the system enters navigation mode; if the user's alertness level is higher than the preset alertness level, the system enters brain-computer interface control mode.
9. A wheelchair control system, characterized in that, include: Wheelchair-mounted device, in-ear signal acquisition device, and main control unit; The in-ear signal acquisition device is used to acquire the ear canal electrical signal of the user and send it to the main control unit; The main control unit is used to execute the multimodal intelligent wheelchair control method based on an in-ear brain-computer interface as described in any one of claims 1 to 8, so as to drive the wheelchair device.
10. The wheelchair control system according to claim 9, characterized in that, The in-ear signal acquisition device includes: a left ear acquisition unit and a right ear acquisition unit; The left ear acquisition unit is used to acquire the electrical signal of the user's left ear canal; The right ear acquisition unit is used to acquire the electrical signal of the user's right ear canal.