Multi-mode fusion intention recognition system and method

The multimodal intention recognition system, which utilizes electromyography (EMG) and electroencephalography (EEG) to coordinate perception, uses EMG signals to control EEG acquisition and visual stimulation. Combined with a deep learning model, it solves the problems of low accuracy and insufficient stability in intention recognition among the elderly and disabled, and achieves high-precision, low-power intention recognition.

CN122018672APending Publication Date: 2026-05-12SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2025-12-17
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing intent recognition methods suffer from low accuracy, poor adaptability, and insufficient stability among the elderly and disabled populations. They are particularly difficult to use reliably under conditions of weak electromyography, and single-modal recognition systems are prone to misjudgment or recognition failure when subjected to noise interference.

Method used

A multimodal fusion intention recognition system based on EEG-EMG-physical sensing is adopted. The EEG acquisition module is started and stopped by EMG signals. The separability of EEG signals is improved by combining visual stimulation patterns. Deep learning models are used for feature extraction and classification to build a multimodal collaborative intention recognition mechanism with low misjudgment and high robustness.

Benefits of technology

Stable output is achieved under conditions of weak electromyography and low signal-to-noise ratio, which improves the user's autonomous interaction capability and user experience, reduces power consumption, and improves recognition accuracy and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018672A_ABST
    Figure CN122018672A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode fusion intention recognition system and method based on electroencephalogram-myoelectricity-physical sensing. The problems that a traditional single myoelectricity recognition mode is low in recognition precision, poor in adaptability, insufficient in stability and the like in old people or disabled people are solved. An electromyographic switch control unit of the system is used for monitoring the muscle activity condition of masseter muscles on the face of a user and quantizing electromyographic signals, and the system judges whether to start an electroencephalogram collection unit or not according to the quantized value; the electroencephalogram acquisition unit is used for acquiring electroneurographic signals generated by the brain under visual stimulation; the electroencephalogram processing and recognition unit comprises a signal preprocessing module and a neural network classification module, the signal preprocessing module preprocesses collected electroencephalogram data, preprocessed electroencephalogram signals are input into the neural network classification module, feature extraction and classification are carried out, and accurate recognition of movement or operation intentions of a user is achieved; and the communication control unit outputs the identification result to external control equipment in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multimodal fusion intention recognition system and method based on electroencephalography (EEG), electromyography (EMG), and physical sensing. Background Technology

[0002] With the accelerating aging of the population and the increasing number of people with congenital or acquired disabilities, the elderly and disabled will face limitations in their mobility, significantly impacting their quality of life and social participation. To assist them in performing basic movements such as walking, standing, and grasping, intelligent devices such as exoskeleton robots and mobility aids have become important technological means. However, in practical applications, accurately, stably, and in real-time recognizing the user's movement intentions remains a major technological bottleneck for current exoskeleton control systems.

[0003] Most existing intent recognition methods rely on single-modal physiological or behavioral signals, such as surface electromyography (sEMG), electroencephalography (EEG), inertial sensor data, or voice commands. Existing intent recognition technologies suffer from the following main drawbacks in the field of assistive device control: (1) Relying on electromyography (EMG) signals, the recognition rate of users with weak EMG is low. Existing methods usually use EMG as the main source of intent, but the amplitude of EMG in the elderly and disabled is weak and the stability is poor. Traditional EMG methods based on threshold or pattern recognition are prone to missed triggers and false triggers, making it difficult to use reliably under weak EMG conditions.

[0004] (2) Monomodal recognition leads to insufficient utilization of EEG signals. Many systems rely solely on electromyography, pressure, or inertial sensors, while failing to fully utilize the evokedness and separability of EEG signals. This results in an inability to compensate when electromyography is unreliable or unavailable, thus limiting the ability to express intent.

[0005] (3) The system has poor fault tolerance and insufficient robustness. When encountering noise interference, loose sensors, or temporary signal interruption, single-modal recognition lacks a collaborative judgment mechanism, which can easily lead to misjudgment, delayed judgment, or recognition failure, making it difficult to meet the requirements of stability and real-time performance in practical applications.

[0006] (4) Lack of effective coordination mechanism between modalities. Existing technologies lack a control logic that effectively links electromyographic triggering with electroencephalography classification, and cannot achieve low-power, continuous and consistent coordinated control in the triggering, acquisition and decision stages. Summary of the Invention

[0007] To address the issues of low accuracy, poor adaptability, and insufficient stability in traditional single-mode EMG (electromyography) recognition among the elderly and disabled, this invention proposes a multimodal fusion intention recognition system and method based on EEG, EMG, and physical sensing. It utilizes EMG signals as a low-cost, low-power intention trigger source to control the start and stop of the EEG acquisition module, achieving reliable triggering logic even under weak EMG conditions. During the EEG acquisition stage, more elicitable visual stimulation patterns are designed to improve the separability and stability of EEG signals. Subsequently, a deep learning model is used to extract and classify EEG features. When EEG confidence is insufficient, the triggered EMG state is combined for auxiliary decision-making, thereby constructing a low-false-judgment, highly robust multimodal collaborative intention recognition mechanism. This system can maintain stable output under weak EMG, low signal-to-noise ratio EEG, and complex environments, and is suitable for various assistive devices such as exoskeletons and wheelchairs, effectively improving users' autonomous interaction capabilities and user experience.

[0008] The multimodal fusion intention recognition system based on EEG-EMG-physical sensing proposed in this invention is unique in that it includes an EMG switch control unit, an EEG acquisition unit, an EEG stimulation unit, an EEG processing and recognition unit, and a communication control unit.

[0009] The electromyography (EMG) switch control unit is used to monitor the muscle activity of the user's masseter muscle and quantify the EMG signal. The system determines whether to activate the EEG acquisition unit based on the magnitude of the quantified value. The EEG stimulation unit is used to provide visual stimulation to the user, and the EEG acquisition unit is used to collect the neural electrical signals generated by the brain under visual stimulation. The EEG processing and recognition unit includes a signal preprocessing module and a neural network classification module. The signal preprocessing module preprocesses the collected EEG data, and the preprocessed EEG signal is input into the neural network classification module for feature extraction and classification to achieve accurate recognition of the user's movement or operation intention. The communication control unit outputs the recognition results to external control devices in real time.

[0010] Furthermore, the electromyography switch control unit monitors the muscle activity of the user's masseter muscle via an electromyography sensor; The electromyography (EMG) switch control unit quantifies the EMG signal, and the system determines whether to activate the EEG acquisition unit based on the magnitude of the quantized value. Specifically, this includes: Electromyographic signals of the masseter muscle In duration of Integrating within the sliding time window yields the electromyographic integral value. Its definition is:

[0011] When the integral value Exceeding the set threshold When the system determines that the user has actively triggered an intent recognition request, it activates the EEG acquisition module; when the score remains below the threshold... If the condition is not triggered, the EEG acquisition module is shut down.

[0012] Furthermore, the EEG acquisition unit includes a wearable EEG acquisition headband, which has multiple high-sensitivity electrodes built in and arranged in specific areas of the user's scalp (such as the occipital lobe and its surrounding visual cortex) to acquire neural electrical signals generated by the brain under visual stimulation.

[0013] Furthermore, the EEG stimulation unit is a transparent display screen based on mixed reality technology. The display screen has a black and white spiral polar coordinate chessboard visual stimulation pattern. The pattern is driven by a sine function to periodically contract and expand, forming a continuous and smooth visual motion stimulation.

[0014] Furthermore, the visual stimulus pattern on the display screen is specifically defined as follows: with the center of the display screen as the origin of polar coordinates, for each pixel... First, convert it to polar coordinates: , Stimulating patterns are concentric rings and It consists of several sector-shaped regions, and the nominal width of each annulus under static conditions is... ,in To achieve the maximum stimulation radius, and to realize gentle contraction and expansion animation, a sine function is used to control the radial offset of the ring, and the normalized radial offset is defined as: (Range 0 to 1) in Given the stimulation frequency, the actual radius offset is: , Then the first The rings at time The dynamic inner radius and outer radius are respectively: , , Combined with the protection radius of the center of vision and maximum display radius Constraints are imposed on the above radius: , , like If the ring is completely within the central protection area, it will not be displayed; if If the ring is outside the display area, it will be ignored for any pixel. If satisfied , and , Then the pixel is assigned to the first... The ring, the first Each sector-shaped unit employs an alternating odd-even chessboard principle: when Fill with black when Fill with white when it is in use; if If it is, then it will be retained as the background color.

[0015] Furthermore, the signal preprocessing module preprocesses the acquired EEG data. The preprocessing operations include: bandpass filtering, sliding window segmentation, and baseline correction of the raw EEG signal.

[0016] Furthermore, the signal preprocessing module preprocesses the acquired EEG data, including performing artifact removal operations to improve signal consistency and signal-to-noise ratio.

[0017] Furthermore, the neural network classification module consists of a multi-layer convolutional structure and a channel attention mechanism, and can be combined with a time series modeling network to extract deep temporal and spatial features of the signal, thereby achieving high-precision classification and recognition of user intent.

[0018] In addition, this invention also proposes a multimodal fusion intent recognition method, based on the above-mentioned multimodal fusion intent recognition system, comprising the following steps: Step 1: Monitor the muscle activity of the user's facial masseter muscle through the electromyography switch control unit, quantify the electromyography signal, and determine whether to activate the EEG acquisition unit based on the magnitude of the quantized value. Step 2: When the quantification value is greater than the set value, the EEG stimulation unit is activated to provide visual stimulation to the user, and the EEG acquisition unit is activated to collect the user's raw EEG signals using the EEG sensor. Step 3: Use the signal preprocessing module to preprocess the acquired EEG data; Step 4: The preprocessed EEG signal is input into the deep neural network module for feature extraction and classification, so as to achieve accurate recognition of the user's movement or operation intention; Step 5: Output the identification results to the external control device in real time through the communication control unit.

[0019] Furthermore, in step 1, the electromyographic signal is quantized, and the system determines whether to activate the electroencephalogram (EEG) acquisition unit based on the magnitude of the quantized value. Specifically: Electromyographic signals of the masseter muscle In duration of Integrating within the sliding time window yields the electromyographic integral value. Its definition is:

[0020] When the integral value Exceeding the set threshold When the system determines that the user has actively triggered an intent recognition request, it activates the EEG acquisition module; when the score remains below the threshold... If the condition is not triggered, the EEG acquisition module is shut down. In step 2, the EEG stimulation unit provides visual stimulation to the user by using a black and white spiral polar coordinate chessboard visual stimulation pattern set on the display screen. The pattern is driven by a sine function to periodically contract and expand, forming a continuous and smooth visual motion stimulation. In step 3, the preprocessing operations include: bandpass filtering, sliding window segmentation, baseline correction, and artifact removal of the raw EEG signal.

[0021] Advantages of this invention: Compared to traditional methods relying on electromyography (EMG) recognition, this invention offers greater applicability and stability, particularly suitable for elderly individuals with weak EMG signals and insufficient masseter muscle strength, as well as patients with myasthenia gravis. The system integrates and judges the masseter muscle EMG signal, using it only to initiate or halt EEG acquisition, allowing the EEG acquisition unit to operate on demand. This significantly reduces power consumption, enhances the wearable device's battery life, and improves trigger response speed. The visual stimulation component uses a periodic contraction-expansion black-and-white spiral checkerboard pattern instead of high-frequency flashing graphics, resulting in less visual burden and higher comfort. Simultaneously, it induces phase-stable and highly separable EEG responses, enhancing signal quality. Combined with a deep neural network featuring multi-layer convolution and temporal modeling structures, along with a channel attention mechanism, this invention accurately extracts the spatial and temporal features of EEG, achieving high-precision intent recognition under low signal-to-noise ratio conditions, significantly improving the overall system reliability and recognition performance. Attached Figure Description

[0022] Figure 1 : The overall structural block diagram of the multi-modal fusion intent recognition method proposed in this invention; Figure 2 : A schematic diagram of the checkerboard stimulation paradigm in the EEG stimulation unit proposed in this invention; Figure 3 : Schematic diagram of the neural network structure of this invention. Detailed Implementation

[0023] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments: This invention proposes a multimodal intention recognition system based on electromyography (EMG) and electroencephalography (EEG) collaborative perception. Its core features include: utilizing EMG signals from the masseter muscle surface as control signals for EEG acquisition; automatically starting and stopping the EEG acquisition module by integrating the EMG signals within a sliding time window and comparing the integrated results with a set threshold; inducing phase-stable and separable EEG responses through visual stimulation to enhance signal quality; processing and classifying the EEG signals using a neural network; and combining gentle visual evoked stimulation to improve the recognizability of the EEG signals, ultimately achieving accurate recognition and assisted control of user intentions.

[0024] Specifically, the multimodal intention recognition system based on electromyography-electroencephalography (EMG) collaborative perception mainly includes an EMG switch control unit, an EEG acquisition unit, an EEG stimulation unit, an EEG processing and recognition unit, and a communication control unit.

[0025] See Figure 1 The method of using the multimodal fusion intent recognition system includes the following steps: Step 1: Monitor the muscle activity of the user's masseter muscle through the electromyography switch control unit, quantify the electromyography signal, and determine whether to activate the EEG acquisition unit based on the magnitude of the quantified value.

[0026] In practical use, when a user attempts to express an intention, such as raising their hand or looking at a target, the system first monitors the muscle activity of the masseter muscle in their face using an electromyography (EMG) sensor. The EMG signal does not directly participate in intention recognition but serves as a control switch for the EEG acquisition unit. To achieve a quantifiable and reproducible triggering mechanism, this invention monitors the EMG signal of the masseter muscle. In duration of Integrating within the sliding time window yields the electromyographic integral value. Its definition is:

[0027] When the integral value Exceeding the set threshold When the system determines that the user has actively triggered an intent recognition request, it "activates" the EEG acquisition unit; when the score remains below the threshold... If the signal is not triggered, the EEG acquisition unit is "shut down," thus reducing energy consumption and avoiding false recognition. Threshold It can be adaptively determined based on the statistical characteristics of electromyography integrals in the user's resting state, for example: ,in and These represent the mean and standard deviation of the resting phase electromyographic integrals, respectively. The adjustment factor is set to 2–3. This integral-threshold triggering logic can operate reliably under weak electromyographic conditions and adapt to individual differences among different users.

[0028] This invention utilizes electromyography (EMG) signals from the masseter muscle surface as control signals for EEG acquisition. By integrating the EMG signals within a sliding time window and comparing the integrated result with a set threshold, the EEG acquisition module can be automatically started and stopped. This mechanism offers advantages such as low power consumption, fast response, and independence from complex user actions. It can reliably initiate the EEG recognition process under weak EMG conditions, significantly improving the system's ease of use and trigger robustness.

[0029] Step 2: When the quantification value is greater than the set value, the EEG stimulation unit is turned on to provide visual stimulation to the user, and the EEG acquisition unit is started to collect the user's raw EEG signals using the EEG sensor.

[0030] The EEG acquisition unit is a wearable EEG acquisition headband device with multiple built-in high-sensitivity electrodes placed in specific areas of the user's scalp (such as the occipital lobe and its surrounding visual cortex) to acquire neural electrical signals generated by the brain under visual stimulation.

[0031] The EEG unit is a transparent display screen based on mixed reality technology that can be worn on the user's head and has visual see-through capabilities. This screen not only displays visual stimuli but also allows the user to clearly see the real-world scene, ensuring that the user's real field of vision is not obstructed during stimulation, facilitating practical operation or environmental perception.

[0032] The stimulation content displayed on the screen is generated by a specially developed EEG-induced program. With the center of the screen as the origin of polar coordinates, for each pixel... First, convert it to polar coordinates: .

[0033] Stimulating patterns are concentric rings and It consists of several sector-shaped regions, and the nominal width of each annulus under static conditions is... ,in The maximum stimulation radius is defined as follows: To achieve gentle contraction and expansion animation, this invention uses a sine function to control the radial offset of the ring, and the normalized radial offset is defined as: (Range 0 to 1) in The stimulus frequency is (e.g., f=8Hz). The actual radius offset is: .

[0034] Then the first The rings at time The dynamic inner radius and outer radius are respectively: , .

[0035] Combined with the protection radius of the center of vision and maximum display radius The present invention imposes constraints on the aforementioned radius: , .

[0036] like If the ring is completely within the central protection area, it will not be displayed; if If the ring is outside the display area, it will be ignored. For any pixel... If satisfied , and , Then the pixel is assigned to the first... The ring, the first A sector-shaped unit. This invention adopts the alternation principle of an odd-even chessboard: when Fill with black when Fill with white when it is in use; if The background color is retained. The spiral / polar coordinate checkerboard pattern exhibits smooth and continuous contraction and expansion over time, placing less burden on the visual system compared to traditional flashing stimuli, reducing fatigue, and making it more suitable for prolonged, continuous use. Simultaneously, it forms clear stimulation frequency components in the frequency domain, possessing the ability to induce specific neural responses under multi-target conditions, further enhancing the recognizability and stimulation efficiency of EEG signals. The stimulation paradigm is attached. Figure 2 As shown.

[0037] The black and white spiral polar coordinate chessboard visual stimulation pattern of this invention uses a sine function to drive periodic contraction and expansion, forming a continuous and smooth visual motion stimulus. This stimulation method is gentler and less visually burdensome than traditional flashing stimulation, and can induce steady-state visual evoked potentials (SSVEPs) with stable phase characteristics and high signal-to-noise ratio. It also has good multi-target discrimination ability, providing higher quality decodeable signals for subsequent EEG classification.

[0038] Step 3: Preprocess the acquired EEG data using the signal preprocessing module, specifically as follows: The acquired EEG signals first enter the signal preprocessing module. The system performs bandpass filtering on the raw EEG signals to retain effective neural activity components within specific frequency bands and suppress low-frequency drift and high-frequency noise interference. Let the original EEG signal be... Channel EEG signal is The impulse response of the bandpass filter is The filtered output is: , in This represents the filter order.

[0039] After filtering is complete, a sliding window operation is performed on the signal, i.e., the signal is processed according to a set window length. and step length The continuous EEG signal is divided into multiple segments of equal length. Each sub-segment can be represented as: .

[0040] The system then performs baseline correction on each segment to improve the comparability between segments. The correction formula is as follows:

[0041] Building upon the above, the system further performs artifact removal. Artifact removal can employ algorithms such as Independent Component Analysis (ICA) to represent multi-channel EEG signals as... ,in For the observed signal matrix, For independent component matrices, The mixture matrix is ​​obtained by estimating the separation matrix. get Based on the component characteristics, artifact components corresponding to eye movements, muscle activity, or power frequency interference are identified and set to zero to obtain the corrected component. Finally, the artifact-free signal is reconstructed. The combination of baseline correction and artifact removal can effectively improve signal consistency and signal-to-noise ratio, providing high-quality input for subsequent feature extraction and classification.

[0042] Step 4: The preprocessed EEG signal is input into the deep neural network module for feature extraction and classification, so as to achieve accurate recognition of the user's movement or operation intention.

[0043] After completing preprocessing step 3 above, each sub-segment signal is sequentially input into the neural network classification module. (See also...) Figure 3 This module consists of a multi-layer convolutional structure and a channel attention mechanism, and can be combined with a time series modeling network to extract deep temporal and spatial features of the signal. Taking a single preprocessed sample as an example, its data can be represented as a dimensional... tensor ,in For the number of channels, For time points. The first layer uses a size of A two-dimensional convolutional kernel is used to convolve along the time axis to extract local temporal features; then a kernel of size [missing information] is used. The convolutional kernels perform spatial convolution along the channel directions, enabling cross-channel spatial pattern modeling. The convolutional outputs, after batch normalization and ELU non-linear activation, are input into the Squeeze-and-Excitation (SE) channel attention module. The SE module first performs global average pooling on each channel to calculate the channel description vector. Its form is: , Channel weights are then generated through two layers of fully connected networks and Sigmoid activation. : , in For ReLU function, For the Sigmoid function, , These are learnable parameters. Finally, the original feature map undergoes channel recalibration: , This enhances the response of channels relevant to the target and suppresses invalid channels. The aforementioned convolution-attention-pooling structure can be stacked in multiple layers (e.g., with channel numbers of 8, 16, 32, 64, and 128 respectively), extracting higher-level and more abstract spatiotemporal features of EEG layer by layer. At the network's end, a flattening operation and a fully connected layer map the high-dimensional features to an intent category score vector y, and the probability of each category is obtained by combining this with a softmax function. , The final output corresponds to the intent category. This enables high-precision recognition of gaze targets or control commands.

[0044] This invention constructs a deep neural network for EEG intention recognition. This network combines a spatial convolutional structure (for extracting cross-channel spatial features) with a convolutional temporal network and integrates a channel attention mechanism to improve effective feature response. This model can extract multi-dimensional, highly discriminative spatiotemporal features from preprocessed EEG signals, achieving high-precision classification and recognition of user intentions, while maintaining high stability under weak signal and noise interference conditions. Compared to traditional EEGNet and similar structures, the network of this invention has the advantage of multi-layered SE modules strengthening key channels in weak EEG signals, and the overall structure is optimized for weak signals.

[0045] Step 5: Output the identification results to the external control device in real time through the communication control unit.

[0046] Ultimately, the system transmits the identified intent signals to external control devices, such as exoskeletons, wheelchairs, and smart prostheses, in real time via the communication module, enabling synchronized interaction between user intent and actual actions.

[0047] At the application level, in addition to intention recognition for the elderly and disabled, this invention can be extended to various scenarios such as medical rehabilitation training, barrier-free human-computer interaction, intelligent wearable control, and special operation assistance. In terms of structural design, the visual stimulus pattern is not limited to the spiral chessboard form, but can also be replaced with dynamic patterns of other geometric shapes, textures, or names, as long as the same periodic geometric transformation principle is adopted. The specific structure of the EEG processing and recognition unit can also be replaced by an equivalent neural network model, filtering method, or signal detection mechanism. The communication control unit is not limited to the current protocol and name, and can adopt other wireless or wired communication methods with equivalent logical mechanisms.

[0048] The above is just one example of the use of this invention patent, and the specific use is not limited to the above process.

Claims

1. A multimodal fusion intention recognition system based on EEG-EMG-physical sensing, characterized in that: It includes an electromyography (EMG) switch control unit, an EEG acquisition unit, an EEG stimulation unit, an EEG processing and recognition unit, and a communication control unit; The electromyography switch control unit is used to monitor the muscle activity of the user's facial masseter muscle and quantify the electromyography signal. The system determines whether to activate the electroencephalogram (EEG) acquisition unit based on the magnitude of the quantized value. The brain stimulation unit is used to provide visual stimulation to the user, and the brain acquisition unit is used to acquire the neural electrical signals generated by the brain under visual stimulation. The EEG processing and recognition unit includes a signal preprocessing module and a neural network classification module. The signal preprocessing module preprocesses the collected EEG data, and the preprocessed EEG signals are input into the neural network classification module for feature extraction and classification, so as to achieve accurate recognition of the user's movement or operation intentions. The communication control unit outputs the identification results to external control devices in real time.

2. The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to claim 1, characterized in that: The electromyography switch control unit monitors the muscle activity of the user's masseter muscle through an electromyography sensor; The electromyography (EMG) switch control unit quantifies the EMG signal, and the system determines whether to activate the EEG acquisition unit based on the magnitude of the quantized value. Specifically, this includes: Electromyographic signals of the masseter muscle In duration of Integrating within the sliding time window yields the electromyographic integral value. Its definition is: When the integral value Exceeding the set threshold When the system determines that the user has actively triggered an intent recognition request, it activates the EEG acquisition module; when the score remains below the threshold... If the condition is not triggered, the EEG acquisition module is shut down.

3. The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to claim 1, characterized in that: The EEG acquisition unit includes a wearable EEG acquisition headband, which has multiple high-sensitivity electrodes installed in a specific area of ​​the user's scalp to acquire neural electrical signals generated by the brain under visual stimulation.

4. The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to claim 1, characterized in that: The EEG stimulation unit includes a display screen with a black and white spiral polar coordinate chessboard visual stimulation pattern. The pattern is driven by a sine function to periodically contract and expand, forming a continuous and smooth visual motion stimulation.

5. The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to claim 4, characterized in that: The visual stimulus pattern on the display screen is specifically defined as follows: with the center of the display screen as the origin of polar coordinates, for each pixel... First, convert it to polar coordinates: , Stimulating patterns are concentric rings and It consists of several sector-shaped regions, and the nominal width of each annulus under static conditions is... ,in To achieve the maximum stimulation radius, and to realize gentle contraction and expansion animation, a sine function is used to control the radial offset of the ring, and the normalized radial offset is defined as: (Range 0 to 1) in Given the stimulation frequency, the actual radius offset is: , Then the first The rings at time The dynamic inner and outer radii are respectively: , , Combined with the protection radius of the center of vision and maximum display radius Constraints are imposed on the above radius: , , like If the ring is completely within the central protection area, it will not be displayed; if If the ring is outside the display area, it will be ignored for any pixel. If satisfied , and , Then the pixel is assigned to the first... The ring, the first Each sector-shaped unit employs an alternating odd-even chessboard principle: when Fill with black when Fill with white when it is in use; if If it is, then it will be retained as the background color.

6. The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to claim 1, characterized in that: The signal preprocessing module preprocesses the acquired EEG data. The preprocessing operations include: bandpass filtering, sliding window segmentation, and baseline correction of the raw EEG signal.

7. The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to claim 6, characterized in that: The signal preprocessing module preprocesses the acquired EEG data and also performs artifact removal operations to improve signal consistency and signal-to-noise ratio.

8. The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to claim 1, characterized in that: The neural network classification module consists of a multi-layer convolutional structure and a channel attention mechanism. It can be combined with a time series modeling network to extract deep temporal and spatial features of the signal, thereby achieving high-precision classification and recognition of user intent.

9. A multimodal fusion intent recognition method, characterized in that, The multimodal fusion intention recognition system based on EEG-EMG-physical sensing according to any one of claims 1-8 includes the following steps: Step 1: Monitor the muscle activity of the user's facial masseter muscle through the electromyography switch control unit, quantify the electromyography signal, and determine whether to activate the EEG acquisition unit based on the magnitude of the quantized value. Step 2: When the quantification value is greater than the set value, the EEG stimulation unit is activated to provide visual stimulation to the user, and the EEG acquisition unit is activated to collect the user's raw EEG signals using the EEG sensor. Step 3: Use the signal preprocessing module to preprocess the acquired EEG data; Step 4: The preprocessed EEG signal is input into the deep neural network module for feature extraction and classification, so as to achieve accurate recognition of the user's movement or operation intention; Step 5: Output the identification results to the external control device in real time through the communication control unit.

10. The multimodal fusion intent recognition method according to claim 9, characterized in that: In step 1, the electromyographic signal is quantified, and the system determines whether to activate the electroencephalogram (EEG) acquisition unit based on the magnitude of the quantized value. Specifically: Electromyographic signals of the masseter muscle In duration of Integrating within the sliding time window yields the electromyographic integral value. Its definition is: When the integral value Exceeding the set threshold When the system determines that the user has actively triggered an intent recognition request, it activates the EEG acquisition module; when the score remains below the threshold... If the condition is not triggered, the EEG acquisition module will be shut down. In step 2, the EEG stimulation unit provides visual stimulation to the user by using a black and white spiral polar coordinate chessboard visual stimulation pattern set on the display screen. The pattern is driven by a sine function to periodically contract and expand, forming a continuous and smooth visual motion stimulation. In step 3, the preprocessing operations include: bandpass filtering, sliding window segmentation, baseline correction, and artifact removal of the raw EEG signal.