Fusion intent prediction method, device, system, storage medium and program product

By employing a large-scale intention prediction model in brain-computer interface devices, and switching the prediction criteria based on the sampling rate of eye movement and electroencephalogram signals, dynamic intention prediction is performed using the TD3 module and intelligent completion module. This solves the problems of high misjudgment rate and system latency in traditional devices under dynamic environments, and achieves efficient and robust intention prediction.

CN121256665BActive Publication Date: 2026-07-24BEIJING JI MASCH TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING JI MASCH TECH CO LTD
Filing Date
2025-08-28
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional brain-computer interface devices cannot effectively handle dynamic environments, resulting in high intention misjudgment rates, system latency, and insufficient utilization of computing resources. In particular, existing fusion solutions cannot effectively suppress EEG signal noise and eye movement drift problems when eye movement signals are attenuated in patients with ALS.

Method used

A large-scale intent prediction model is adopted, which switches the prediction criteria based on the sampling rate thresholds of eye movement signals and EEG signals. The TD3 module and intelligent completion module are used for dynamic intent prediction. Combined with the feedback optimization model, it is ensured that eye movement signals are used as the main signal when the quality of eye movement signals is high, and EEG signals are used as the main signal when the quality is low.

Benefits of technology

It improves the accuracy of intent prediction and the robustness of the system, reduces the false positive rate, ensures long-term availability in cases such as ALS patients, and achieves efficient intent prediction and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256665B_ABST
    Figure CN121256665B_ABST
Patent Text Reader

Abstract

A fusion intention prediction method, device, system, storage medium and program product, comprising: collecting electroencephalogram signals and eye movement signals of a human body; obtaining an eye movement sampling rate of the eye movement signals and comparing the eye movement sampling rate with a preset sampling rate threshold; inputting the electroencephalogram signals and the eye movement signals into an intention prediction large model, the intention prediction large model performing intention prediction based on the electroencephalogram signals and the eye movement signals; outputting an intention prediction result; when the eye movement sampling rate is greater than or equal to the preset sampling rate threshold, performing fusion learning of the intention prediction large model; and when the eye movement sampling rate is less than the preset sampling rate threshold, closing the fusion learning of the intention prediction large model. The present disclosure solves the problems that the prior art uses a pre-set static model, cannot handle dynamic environments, has low intention prediction accuracy, and cannot be used for a long time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of brain-computer interfaces, and more particularly to a method, apparatus, system, storage medium, and program product for fusion intention prediction. Background Technology

[0002] Traditional brain-computer interface (BCI) devices often rely on users to actively imagine actions to generate control signals. For example, users can control a smart wheelchair by imagining forward or backward movements, or control voice and text input devices by imagining desired phrases. By collecting the body's electroencephalogram (EEG) signals through invasive or non-invasive BCI devices, desired control functions can be achieved, providing rehabilitation assistance, communication support, and mobility assistance for people with disabilities. However, traditional BCIs often fail to detect the user's cognitive details, resulting in insufficient accuracy and efficiency in control.

[0003] Scientific research has proven that for people with disabilities, the degeneration of eye muscles is significantly delayed compared to the degeneration of muscles in other parts of the body. Therefore, users can directly transmit control information by controlling their eye movements according to certain rules. For example, eye trackers can collect eye movements such as saccades, fixations, and blinks, and control external devices based on pre-set control logic.

[0004] Currently, there are some technical solutions that simultaneously collect EEG and eye movement (EMT) signals, fuse the data according to certain control logic, and then use them to control external execution devices. However, existing fusion solutions use pre-set static programs, which cannot handle dynamic environments, such as changes in user attention or attenuation of EMT signals. This leads to excessive physical load, requiring frequent system recalibration, increasing computational overhead and energy consumption. For disabled individuals, especially those with ALS (Amyotrophic Lateral Sclerosis), their EMT signals gradually attenuate. As the EMT signals attenuate, the error in intention fusion using the static model increases, eventually causing the model to fail. Furthermore, in static programs, noise in EEG signals cannot be effectively suppressed, and eye movement drift (such as eye movement degeneration in ALS patients) cannot be adequately addressed, resulting in a high intention misjudgment rate. A misjudgment rate that is too high (>20%) significantly increases system latency, leading to insufficient utilization of computational resources for intention prediction.

[0005] Therefore, a novel fusion intention prediction method is needed. When eye-tracking signals are available, intention prediction is primarily based on these signals, while a large-scale intention prediction model is trained to improve the accuracy of EEG signal control. When eye-tracking signals are unavailable, the large-scale intention prediction model relies primarily on EEG signals for intention prediction. Through a long-term, dynamic training process, the long-term usability of the fusion intention prediction model is ensured. Summary of the Invention

[0006] The embodiments of this disclosure provide a fusion intention prediction method that solves problems such as high intention misjudgment rate and system latency when collecting EEG signals and eye movement signals for intention prediction. In particular, existing solutions using pre-set static models cannot handle dynamic environments, excessive physical load, high system misjudgment rate, and poor robustness. A fusion intention prediction device, system, storage medium, and program product capable of implementing this method are also provided.

[0007] The first aspect of this disclosure provides a method for fusing intent prediction, comprising: acquiring electroencephalogram (EEG) signals and eye movement (EMG) signals of a human body; obtaining the eye movement sampling rate of the EMG signals and comparing it with a preset sampling rate threshold; inputting the EMG signals and the EMG signals into a large-scale intent prediction model, wherein the large-scale intent prediction model performs intent prediction based on the EMG signals and the EMG signals; outputting an intent prediction result; performing fusing learning of the large-scale intent prediction model when the EMG sampling rate is greater than or equal to the preset sampling rate threshold; and disabling fusing learning of the large-scale intent prediction model when the EMG sampling rate is less than the preset sampling rate threshold.

[0008] For example, in at least one embodiment, the intent prediction big model performs intent prediction based on the EEG signal and the eye movement signal, including: inputting the EEG signal and the eye movement signal into a TD3 module, the TD3 module performing preliminary intent prediction based on the EEG signal and the eye movement signal; inputting the preliminary intent prediction result into the intelligent completion module; the intelligent completion module analyzing the historical context, ranking the preliminary intent prediction result by confidence, and obtaining the intent prediction result based on the confidence ranking.

[0009] For example, in at least one embodiment, feedback on the intention prediction result is collected, and the feedback result is input into the TD3 module and the intelligent completion module, which perform model optimization based on the feedback result.

[0010] For example, in at least one embodiment, when the eye-tracking sampling rate is greater than or equal to the preset sampling rate threshold, the intent prediction big model uses the eye-tracking signal as the primary basis for intent prediction; the model state vector of the intent prediction big model includes EEG signal features, eye-tracking signal features, and historical intent context, which improves the confidence of the mapping of the EEG signal and the eye-tracking signal in the intent prediction big model while predicting intent; when the eye-tracking sampling rate is less than the preset sampling rate threshold, the intent prediction big model uses the EEG signal as the primary basis for intent prediction.

[0011] For example, in at least one embodiment, acquiring the EEG signal and the eye movement signal of a human body includes the following steps: acquiring the EEG signal of a human body, removing artifacts from the EEG signal, and extracting the EEG time-frequency features from the EEG signal; acquiring the eye movement signal of a human body, calculating the sliding standard deviation of the eye movement signal, obtaining the stable trajectory of the eye movement signal, and extracting the eye movement features from the eye movement signal; dynamically adjusting the EEG time-frequency features and the eye movement features to obtain an N-dimensional feature vector; inputting the EEG signal and the eye movement signal into the intent prediction model means inputting the N-dimensional feature vector into the intent prediction model.

[0012] A second aspect of this disclosure provides a fusion intention prediction device, including an EEG acquisition module, an eye-tracking acquisition module, a control module, an intention prediction big data model, and an output module; wherein, the EEG acquisition module is configured to acquire EEG signals of a human body; the eye-tracking acquisition module is configured to acquire eye-tracking signals of a human body; the control module is configured to acquire the eye-tracking sampling rate of the eye-tracking signals and compare it with a preset sampling rate threshold; when the eye-tracking sampling rate is greater than or equal to the preset sampling rate threshold, the intention prediction big data model training and update is activated; when the eye-tracking sampling rate is less than the preset sampling rate threshold, the intention prediction big data model training and update is deactivated; the intention prediction big data model is configured to perform intention prediction based on the EEG signals and the eye-tracking signals; and the output module is configured to output the intention prediction result.

[0013] For example, in at least one embodiment, a signal processing module is further included. This signal processing module is configured to perform artifact removal on the EEG signal, extract the EEG time-frequency features from the EEG signal, calculate the sliding standard deviation of the eye movement signal to obtain a stable trajectory of the eye movement signal, extract the eye movement features from the eye movement signal, and dynamically adjust the EEG time-frequency features and the eye movement features to obtain an N-dimensional feature vector. The N-dimensional feature vector is input into the intent prediction model. The intent prediction model includes a TD3 module, an intelligent completion module, and a feedback module. The TD3 module is configured to perform preliminary intent prediction based on the EEG signal and the eye movement signal, and input the preliminary intent prediction result into the intelligent completion module. The intelligent completion module is configured to analyze historical context, rank the preliminary intent prediction result by confidence level, and obtain the intent prediction result based on the confidence level ranking. The feedback module is configured to collect feedback on the intent prediction result, input the feedback result into the TD3 module and the intelligent completion module, and optimize the model based on the feedback result.

[0014] A third aspect of this disclosure provides a fusion intent prediction system, comprising: a memory for non-transitory storage of computer-executable instructions; and a processor for running the computer-executable instructions, characterized in that the computer-executable instructions are executed by the processor to perform the fusion intent prediction method according to any one of the preceding claims.

[0015] A fourth aspect of this disclosure provides a non-transitory storage medium for non-transitory storage of computer-executable instructions, characterized in that, when the computer-executable instructions are executed by a computer, the fusion intent prediction method according to any one of the preceding claims is performed.

[0016] The fifth aspect of this disclosure provides a computer program product, including a computer program, characterized in that, when executed by a processor, the computer program implements the fusion intent prediction method described in any of the preceding claims. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0018] Figure 1 This is a flowchart of a fusion intent prediction method according to an embodiment of the present disclosure.

[0019] Figure 2 This is a flowchart of data acquisition and processing according to an embodiment of the present disclosure.

[0020] Figure 3 This is a flowchart of the intent prediction large model processing according to embodiments of the present disclosure.

[0021] Figure 4 This is a schematic diagram of a fusion intent prediction device according to an embodiment of the present disclosure.

[0022] Figure 5 This is a schematic diagram illustrating the relationship between the fusion intent prediction modules according to embodiments of the present disclosure. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0024] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes. In this disclosure, “multiple” means two or more.

[0025] According to embodiments of this disclosure, Figure 1 A method for predicting fused intentions is disclosed, including step S1: acquiring electroencephalogram (EEG) signals and eye movement (EMG) signals from the human body. There are no limitations on the methods and devices used for acquiring EEG and EMG signals; existing brain-computer interface (BCI) devices and eye movement acquisition devices can be used. The acquisition of EEG and EMG signals is simultaneous and continuous. Simultaneous acquisition ensures that at the same point in time, the user generates a single control intention, at which point both the EEG and EMG signals simultaneously exhibit specific signal characteristics for this intention. Continuous acquisition ensures the continuity of the control intention, avoiding contradictory control intentions.

[0026] Step S2: Obtain the eye movement sampling rate of the eye movement signal and compare it with a preset sampling rate threshold. For example, if the preset sampling rate threshold is 5Hz, the control potential is set to high level when the eye movement sampling rate of the eye movement signal is greater than or equal to 5Hz, and the control potential is set to low level when the eye movement sampling rate of the eye movement signal is less than 5Hz. Using the sampling rate threshold comparison method can effectively distinguish the signal quality of the eye movement signal. The algorithm adopts different strategies for intent prediction depending on whether the eye movement signal quality is high or low.

[0027] Step S3: Input the EEG and eye movement (EMT) signals into the intent prediction model. The intent prediction model predicts intent based on the EEG and EMT signals. The prediction basis of the intent prediction model differs depending on the comparison result between the EMT sampling rate and the preset sampling rate threshold. When the control potential is set to a high level, it indicates that the EMT signal is of high quality and can directly express the user's intent; in this case, intent prediction is primarily based on the EMT signal. When the control potential is set to a low level, it indicates that the EMT signal is of low quality and can no longer clearly express the user's intent; in this case, intent prediction is primarily based on the EEG signal.

[0028] Step S4: Output the intent prediction result. After fusing EEG and eye-tracking signals using the above method to predict the intent, the intent prediction result is output to an external execution device. This disclosure does not limit the specific external execution device; the output intent prediction result can control any desired external execution device. For example, execution devices such as smart wheelchairs, game interfaces, robotic arms, and voice / text input devices can all be controlled through the intent prediction result.

[0029] Step S51: When the eye-tracking sampling rate is greater than or equal to the preset sampling rate threshold, perform fusion learning of the intent prediction large model. Alternatively, Step S52: When the eye-tracking sampling rate is less than the preset sampling rate threshold, disable fusion learning of the intent prediction large model. The intent prediction large model is continuously optimized during use. The overall optimization strategy is that when the eye-tracking signal quality is high, the control accuracy is high, and fusion learning optimization is performed on the large model to improve the accuracy of intent prediction. When the eye-tracking signal quality is low, to avoid the negative impact of invalid eye-tracking signals on the optimization of the large model, fusion learning of the intent prediction large model is disabled. Disabling fusion learning of the intent prediction large model means freezing the module that trains EEG signals based on eye-tracking signals for intent prediction. Intent prediction based on EEG signals is no longer trained and updated according to eye-tracking signals. The intent prediction large model can still continuously train and optimize based on the information fed back from the feedback module and its own training capabilities.

[0030] Preferred, such as Figure 2 As shown, after collecting the human body's electroencephalogram (EEG) signals and eye movement signals, the following steps are also included: Step S11: Collecting the human body's EEG signals. Step S11 involves using an EEG signal acquisition device to collect the raw EEG signals of the human body.

[0031] Step S12: Artifact removal from the raw EEG signal. The raw EEG signal contains artifacts. Step S12 removes artifacts from the raw EEG signal by performing 0.5-40Hz bandpass filtering for noise reduction.

[0032] Step S13: Extract the EEG time-frequency features from the EEG signal. EEG signal features cover multiple dimensions, including the time domain, frequency domain, time-frequency domain, and spatial domain. Each feature reflects different characteristics of the original signal from different perspectives. Extracting the EEG time-frequency features from the EEG signal is beneficial for time alignment with the eye movement signal in subsequent steps.

[0033] Step S14: Acquire human eye movement signals. Step S14 involves using an eye movement signal acquisition device to acquire raw human eye movement signals.

[0034] Step S15: Calculate the sliding standard deviation of the eye movement signal to obtain the stable trajectory of the eye movement signal. Step S15 calculates the original eye movement signal to obtain the stable trajectory of the eye movement signal. Based on the trajectory, eye movements, such as blinking, looking left, looking right, etc., can be identified. The position of eye gaze can also be analyzed to understand the user's intention.

[0035] Step S16: Extract eye movement features from the eye movement signal. The extracted eye movement features include a multi-dimensional vector of eye movement features, which includes the time dimension.

[0036] Step S17: Perform dynamic time adjustment on the EEG time-frequency features and eye-tracking features, align the time of the two features, and obtain an N-dimensional feature vector. For example... Figure 2 As shown, steps S11-S13 and steps S14-S16 are performed sequentially. Step S17 performs time alignment and fusion on the features extracted in steps S13 and S16, and the fused features are used by the algorithm model to obtain an N-dimensional feature vector.

[0037] Step S18: Input the N-dimensional feature vector into the large model intended for prediction. Figure 1 In step S3, the EEG signal and eye movement signal are input into the intention prediction big model. Preferably, step S18 is specifically executed to input the fused multidimensional feature vector into the big model for processing.

[0038] Preferred, such as Figure 3As shown, step S3 includes step S31: inputting EEG and eye-tracking signals into the TD3 module. The TD3 module performs preliminary intent prediction based on the EEG and eye-tracking signals. The TD3 module includes two critic networks, one actor network, and a delayed update layer. The TD3 module first runs the actor network to map the 28-dimensional feature vector to intent probabilities. The 28-dimensional feature vector is a state vector derived from eye-tracking signals, EEG signals, and historical context. Historical context refers to the sequential relationship between intents recorded by the model, such as intent B often following intent A, or intent C not immediately following intent A. Intent probability is the probability of an intent prediction value, such as controlling a wheelchair to move forward or controlling the input of a character. The critic network is responsible for evaluating the value of the action (Q-value), and its optimization objective is to minimize the Q-value prediction error. The two critic networks independently evaluate the Q-value, using the smaller value given by the two critic network functions as the Q-value target. The two critic networks significantly reduce the overestimation problem of a single critic during the update process, thereby improving the stability of the model. The TD3 module employs a delayed update strategy, where the critic network guides the actor network's updates. The critic network is trained twice, and the actor network is trained once. Experimental results show that training both the actor and critic networks simultaneously without using a delayed update layer leads to training instability. However, when only the actor network is fixed, the critic network often converges to the correct result. Therefore, the TD3 model updates the actor network at a lower frequency, updating it only once after the critic network has achieved a good result through two updates.

[0039] Preferably, the eye-tracking spatial localization formula in the TD3 module is:

[0040]

[0041] in, Γ is the output baseline action vector (the decision baseline in the early stages of training), and Γ is the spatial localization function (the coordinate transformation engine implemented by the FPGA). It is the eye-tracking coordinate vector at the current time t. It is the eye movement trajectory vector (dynamic memory) within the historical δ time window.

[0042] Preferably, the EEG supervised training formula in the TD3 module is:

[0043]

[0044] in, It is the supervised learning loss value, f ΦIt is an EEG intent recognition model (Φ is a trainable parameter). It is the EEG feature vector at the current time t. It is the eye-tracking reference action (as a supervisory signal), and λR(Φ) is the regularization term (L2 constraint, λ = 0.01).

[0045] Preferably, the integrated decision formula in the TD3 module is:

[0046]

[0047] in, It is the system's final decision action, π θ Policy function (θ is the Actor network parameter), It is the fused feature vector, and τ is the EEG confidence threshold (experimental value 0.15).

[0048] Preferably, the user feedback reward formula in the TD3 module is:

[0049]

[0050] Among them, R (t) It is the reinforcement learning reward value that is transformed by feedback. It is a reward transition function (three-state quantization: {-1, -0.3, +1}). It is the user feedback signal after the action is executed Δt (collected 200ms after the action). It is the system's final decision-making action.

[0051] Preferably, the objective formula for optimizing the critic strategy in the TD3 module is:

[0052]

[0053] Among them, y i R is the target Q-value (TD target) of the i-th sample in the critic network. i R is the user feedback reward for the i-th sample. (t) , π φ′ It is the target actor network (parameter φ′ is periodically soft-updated from the main network φ), γ is the discount factor (γ = 0.99), Q θ It is the critic value function (dual-network overestimation prevention). It is a state characteristic. It is the minimum Q-value among the two target critic networks (used to suppress overestimation).

[0054] Preferably, the optimization objective formula for the actor strategy in the TD3 module is:

[0055]

[0056] in, It is the policy gradient of the actor network with respect to parameter φ. It involves calculating the expected value of a batch of samples from experience replay (the actual calculation uses the batch mean). It is the gradient of the main critic network with respect to the action (the gradient from the action to the Q-value in the chain rule). It is the gradient of the actor network with respect to the parameter φ (the gradient of the policy output with respect to the parameter). State characteristics It is an action vector

[0057] Preferably, when the eye-tracking sampling rate is greater than or equal to a preset sampling rate threshold, the TD model is in the training phase. The TD3 module performs dynamic calculations, using eye-tracking signals as the true labels, which also serve as supervision signals in the TD3 module's algorithm. The TD3 module's state vector S = [EEG features, eye-tracking features, historical intent context], action vector A = intent prediction value (e.g., character selection), and reward R is based on accuracy (success +1, error -1). The TD3 module's calculations also need to incorporate speed metrics. According to the control conditions, the quality of eye-tracking signals is high at this stage, ensuring that the EEG signals learn a high-confidence mapping under TD3's guidance during training. After a period of training, the accuracy of intent prediction relying solely on EEG signals improves by >30%.

[0058] Preferably, when the eye-tracking sampling rate is greater than or equal to a preset sampling rate threshold, in intent prediction, the eye-tracking signal is the primary basis (with a weight of 80% in the algorithm), and the EEG signal is the secondary basis (with a weight of 20% in the algorithm).

[0059] Preferably, when the eye-tracking sampling rate is less than a preset sampling rate threshold, the system detects eye-tracking signal distortion (e.g., worsening of ALS, eye-tracking sampling rate <5Hz), and the TD3 module automatically switches to prediction mode. At this time, the fusion learning of the TD3 module (a module that optimizes the accuracy of EEG signals based on eye-tracking signals) is frozen, and the TD3 module calibrated during the training phase is used for intent prediction.

[0060] Preferably, in some embodiments, when the eye-tracking sampling rate is less than a preset sampling rate threshold, the eye-tracking signal is only used for comparison of the eye-tracking sampling rate, and the electroencephalogram (EEG) signal is used independently as the basis for intention prediction.

[0061] Preferably, in other embodiments, when the eye movement sampling rate is less than a preset sampling rate threshold, the eye movement signal is defined as a weak eye movement signal. In intent prediction, the EEG signal is the primary basis (with a weight of 80% in the algorithm), and the eye movement signal is the secondary basis (with a weight of 20% in the algorithm).

[0062] Preferred, such as Figure 3 As shown, step S3 includes step S32: inputting the preliminary intent prediction result into the intelligent completion module; the intelligent completion module analyzes the historical context, sorts the preliminary intent prediction results by confidence, and obtains the intent prediction result based on the confidence ranking. The preliminary intent prediction result output by the TD3 module still has some error, while in most control scenarios, the control logic sequence follows a pattern. For example, in text input, users have specific language habits, the language they use has specific grammatical rules, and the input characters follow a pattern. Similarly, in controlling a wheelchair, a forward command rarely transitions directly to a backward command without a stop command; in fixed scenarios, due to the repeated overlap of wheelchair trajectories, the control commands also follow a pattern. In step S32, the intelligent completion module, based on an NLP algorithm, analyzes the historical context within a certain window length (e.g., window length = 5), sorts the preliminary intent prediction results output by the TD3 module by confidence, generates the top 3 completion options, and completes the preliminary intent prediction result based on the probability of the completion options to obtain the intent prediction result. For example, when a user is inputting text, based on the user's language habits and grammar, the model identifies typos in the initial intent prediction results, and the large model makes a selection from the completion options to output the correct intent prediction results.

[0063] Preferably, the intelligent completion module performs intelligent completion based on the LSTM algorithm or the transformer algorithm.

[0064] Preferred, such as Figure 3 As shown, step S3 includes step S33: collecting feedback on the intent prediction results, inputting the feedback results into the TD3 module and the intelligent completion module, and optimizing the model based on the feedback results. Step S33 is the step where the user actively confirms the intent prediction results output in steps S31 and S32 and provides feedback on whether they are correct or not. For example, after outputting the intent prediction results, the user's microsaccade signals are analyzed, and the confirmation or rejection state is distinguished using a gaze position of <200ms. Another example is recognizing EEG signal feedback patterns: the user imagining a right fist = confirmation, the user imagining mathematical calculation = rejection. The feedback results are input into the TD3 module and the intelligent completion module to optimize their algorithms. After the TD3 module and the intelligent completion module have undergone complete training, the method disclosed herein can achieve a communication speed of ≥20 words / minute in the later stages of ALS patients, approaching the level of daily conversation.

[0065] According to embodiments of this disclosure, Figure 4A device for fusion intention prediction is disclosed, comprising an EEG acquisition module 1, an eye-tracking acquisition module 2, a control module 3, an intention prediction big data model 4, and an output module 5. The EEG acquisition module 1 is configured to acquire EEG signals from a human body, and the eye-tracking acquisition module 2 is configured to acquire eye-tracking signals from a human body. The control module 3 is configured to acquire the eye-tracking sampling rate of the eye-tracking signals and compare it with a preset sampling rate threshold. When the eye-tracking sampling rate is greater than or equal to the preset sampling rate threshold, the intention prediction big data model training and update is activated; when the eye-tracking sampling rate is less than the preset sampling rate threshold, the intention prediction big data model training and update is deactivated. The intention prediction big data model 4 is configured to perform intention prediction based on EEG signals and eye-tracking signals. The output module 5 is configured to output the intention prediction result and control an external execution device 6 based on the intention prediction result.

[0066] Preferably, the EEG acquisition module 1 is a multi-channel high-precision non-invasive brain-computer interface.

[0067] Preferably, the eye-tracking acquisition module 2 is a 60Hz infrared eye tracker, which can capture gaze coordinates with a measurement accuracy of 0.5°.

[0068] Preferably, the control module 3, the intent prediction large model 4, and the output module 5 are integrated into the same computing device.

[0069] Preferred, such as Figure 5 As shown, the fusion intention prediction device also includes a signal processing module. The signal processing module is configured to remove artifacts from the EEG signal, extract the EEG time-frequency features from the EEG signal, calculate the sliding standard deviation of the eye movement signal to obtain the stable trajectory of the eye movement signal, extract the eye movement features from the eye movement signal, and perform dynamic time adjustment on the EEG time-frequency features and eye movement features to obtain an N-dimensional feature vector. The N-dimensional feature vector is input into the large-scale intention prediction model.

[0070] Preferred, such as Figure 5 As shown, after a user generates an intention, the EEG acquisition module collects raw EEG signals, and the eye-tracking acquisition module collects raw eye-tracking signals. The raw EEG signals and raw eye-tracking signals are then input into the signal processing module for signal processing. Simultaneously, the control module performs threshold control based on the eye-tracking sampling rate, selects different intention prediction modes, and chooses whether to disable the fusion learning of the algorithm module. The main intention prediction model includes a TD3 module, an intelligent completion module, and a feedback module. The TD3 module is configured to perform preliminary intention prediction based on EEG and eye-tracking signals, and inputs the preliminary intention prediction results into the intelligent completion module. The intelligent completion module is configured to analyze historical context, rank the preliminary intention prediction results by confidence, and obtain the final intention prediction result based on the confidence ranking. Feedback on the intention prediction results is collected and input into the TD3 module and the intelligent completion module, which then optimize the model based on the feedback results.

[0071] According to an embodiment of this disclosure, a fusion intent prediction system is also provided, comprising: a memory for non-temporarily storing computer-executable instructions; and a processor for running the computer-executable instructions, characterized in that the computer-executable instructions are executed by the processor to perform the fusion intent prediction method according to any one of the above.

[0072] According to embodiments of this disclosure, a non-transitory storage medium is also provided for storing computer-executable instructions non-transitory, characterized in that when the computer-executable instructions are executed by a computer, the fusion intent prediction method according to any one of the above is executed.

[0073] According to an embodiment of this disclosure, a computer program product is also provided, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the fusion intent prediction method according to any one of the above.

[0074] The following points also need to be explained:

[0075] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure, and other structures can be referred to the general design.

[0076] (2) For clarity, the thickness of devices, layers, or regions is enlarged or reduced in the drawings used to describe embodiments of the present disclosure, i.e., these drawings are not drawn to scale. It will be understood that when an element such as a layer, film, region, or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.

[0077] (3) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0078] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure shall be determined by the scope of the claims.

Claims

1. A method for predicting fusion intent, comprising: Collecting electroencephalogram (EEG) and eye movement signals from the human body; The eye-tracking sampling rate of the eye-tracking signal is obtained and compared with a preset sampling rate threshold: When the eye-tracking sampling rate is greater than or equal to the preset sampling rate threshold, the intent prediction model uses the eye-tracking signal as the primary basis for intent prediction. When the eye-tracking sampling rate is less than the preset sampling rate threshold, the intent prediction big model uses the EEG signal as the main basis for intent prediction. The electroencephalogram (EEG) signal and the eye movement (EMG) signal are input into the intent prediction model. The intent prediction model performs intent prediction based on the EEG signal and the EMG signal, including: The electroencephalogram (EEG) signal and the eye movement (EMT) signal are input into the TD3 module, which performs preliminary intention prediction based on the EEG signal and the EMT signal. Input the preliminary intent prediction results into the intelligent completion module. The intelligent completion module analyzes the historical context, sorts the preliminary intent prediction results by confidence, and obtains the intent prediction results based on the confidence ranking. Output the intent prediction result; When the eye-tracking sampling rate is greater than or equal to the preset sampling rate threshold, the fusion learning of the intent prediction large model is performed; When the eye-tracking sampling rate is less than the preset sampling rate threshold, the fusion learning of the intent prediction large model is turned off; Feedback on the intent prediction results is collected, and the feedback results are input into the TD3 module and the intelligent completion module. The TD3 module and the intelligent completion module optimize the model based on the feedback results. The model state vector of the intent prediction big model includes EEG signal features, eye movement signal features, and historical intent context, which improves the confidence of the mapping of the EEG signals and the eye movement signals in the intent prediction big model while predicting intent.

2. The fusion intent prediction method according to claim 1, characterized in that, Collecting the electroencephalogram (EEG) signals and eye movement signals from a human body includes the following steps: Collect the electroencephalogram (EEG) signals of the human body, remove artifacts from the EEG signals, and extract the EEG time-frequency features from the EEG signals; The eye movement signals of the human body are collected, the sliding standard deviation of the eye movement signals is calculated, the stable trajectory of the eye movement signals is obtained, and the eye movement features in the eye movement signals are extracted. Dynamic time adjustment is performed on the EEG time-frequency features and the eye movement features to obtain an N-dimensional feature vector; Inputting the EEG signal and the eye movement signal into the intention prediction model means inputting the N-dimensional feature vector into the intention prediction model.

3. A device for fusion intention prediction, comprising an EEG acquisition module, an eye-tracking acquisition module, a control module, a large-scale intention prediction model, and an output module; wherein, The EEG acquisition module is configured to acquire EEG signals from the human body; The eye-tracking acquisition module is configured to acquire human eye-tracking signals; The control module is configured to acquire the eye movement sampling rate of the eye movement signal and compare it with a preset sampling rate threshold. When the eye movement sampling rate is greater than or equal to the preset sampling rate threshold, the intention prediction large model training update is activated. When the eye movement sampling rate is less than the preset sampling rate threshold, the intention prediction large model training update is deactivated. The intent prediction model is configured to predict intent based on the electroencephalogram (EEG) signals and the eye movement signals. The intent prediction model includes a TD3 module, an intelligent completion module, and a feedback module. The TD3 module is configured to perform preliminary intention prediction based on the EEG signal and the eye movement signal, and input the preliminary intention prediction result into the intelligent completion module. The intelligent completion module is configured to analyze historical context, sort the preliminary intent prediction results by confidence, and obtain the intent prediction results based on the confidence sort. The feedback module is configured to collect feedback on the intent prediction result, input the feedback result into the TD3 module and the intelligent completion module, and the TD3 module and the intelligent completion module optimize the model based on the feedback result; The output module is configured to output the intent prediction result.

4. The fusion intent prediction device according to claim 3, characterized in that, It also includes a signal processing module, which is configured to remove artifacts from the EEG signal, extract the EEG time-frequency features from the EEG signal, calculate the sliding standard deviation of the eye movement signal to obtain the stable trajectory of the eye movement signal, extract the eye movement features from the eye movement signal, and perform dynamic time adjustment on the EEG time-frequency features and the eye movement features to obtain an N-dimensional feature vector; the N-dimensional feature vector is input into the intent prediction large model.

5. A fusion intent prediction system, comprising: Memory is used to store non-temporary executable instructions for a computer. And a processor for running the computer-executable instructions, characterized in that the computer-executable instructions are executed by the processor to perform the fusion intent prediction method according to any one of claims 1-2.

6. A non-transitory storage medium for non-transitory storage of computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a computer, the fusion intent prediction method according to any one of claims 1-2 is performed.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the fusion intent prediction method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Fatigue detection and regulation method based on electroencephalogram-eye movement bimodal signal

    CN109009173A

  • Sensor-based training intervention

    US20230380740A1