Human-computer interaction virtual keyboard input method, device and equipment based on electroencephalogram signals

By adaptively extracting the key features and dynamic characteristics of EEG signals, combining with deep learning models for intention recognition and mapping, the problems of poor adaptability and limited recognition accuracy in the existing technology are solved, and efficient and accurate EEG virtual keyboard input is achieved.

CN120066284BActive Publication Date: 2025-06-27XIAOZHOU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510554359.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-27
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing EEG virtual keyboard technology is difficult to adapt to the individual characteristics and input intentions of users, and has limited recognition accuracy and generalization capabilities, and lacks considerations on the user's cognitive status and intention complexity.

Method used

By obtaining the user's original EEG time series and spectrum characteristics, wavelet packet decomposition is performed to extract multi-scale EEG signal components, fuzzy entropy of the wavelet coefficients is calculated to evaluate the complexity characteristics, determine the key wavelet coefficients most related to the user's input intention, and dynamically adjust the intention regulation frequency based on the user's EEG baseline characteristics, cognitive state and complexity characteristics. Finally, the mapping relationship between the intent category and the virtual keyboard key is generated through empirical modal decomposition and multi-level signal analysis.

Benefits of technology

It improves the input accuracy and adaptability of the EEG virtual keyboard, enhances the recognition accuracy and system robustness, reduces the error touch rate and calculation amount, and achieves a low-latency control effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066284B_ABST
    Figure CN120066284B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of brain-computer interfaces, and discloses a human-computer interaction virtual keyboard input method, device and equipment based on electroencephalogram signals. The method includes obtaining an original electroencephalogram time series and corresponding spectral features, performing wavelet packet decomposition on the original electroencephalogram time series to obtain multi-scale electroencephalogram signal components; obtaining complexity features of wavelet coefficients; determining key wavelet coefficients most relevant to the user's input intention among multiple wavelet coefficients; obtaining the electroencephalogram baseline features, cognitive state and complexity features of the user to obtain the optimal intention regulation frequency personalized for the user; reconstructing and generating an electroencephalogram intention feature sequence according to the key wavelet coefficients and the optimal intention regulation frequency; performing empirical mode decomposition on the electroencephalogram intention feature sequence to obtain a slow-changing component reflecting the steady-state characteristics of the intention and a fast-changing component reflecting the dynamic changes of the intention; obtaining the corresponding intention category of the user to generate a mapping relationship between the intention category and the virtual keyboard keys, and generating a control instruction for the virtual keyboard.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of brain-computer interfaces, and particularly to a method, device and equipment for human-computer interaction virtual keyboard input based on electroencephalogram (EEG) signals. Background Art

[0002] With the continuous development of human-computer interaction technology, using EEG signals to achieve direct human-computer communication and control has become a popular research direction. Among them, virtual keyboard input based on EEG signals is a promising interaction method, which allows users to directly control input through EEG intentions without relying on traditional physical keyboards, thereby improving input efficiency and reducing limb burden. Currently, EEG virtual keyboards mainly adopt two paradigms: visual evoked potential (VEP) and P300 potential. The VEP paradigm induces specific EEG patterns through flashing letters or symbols, while the P300 paradigm uses the P300 potential caused by rare target stimuli for input selection. These two paradigms have achieved the basic functions of EEG virtual keyboards to a certain extent, but there are still some limitations.

[0003] Most existing EEG virtual keyboards adopt fixed stimulation frequencies and patterns, ignoring individual differences and usage preferences of users. Different users may have different sensitivities to visual stimuli and response characteristics, and unified stimulation parameters are difficult to achieve the best induction effect and input performance. At the same time, the recognition of EEG signals is usually based on simple template matching or linear classifiers, which are difficult to fully exploit the non-linear dynamic features contained in EEG signals, resulting in limited recognition accuracy and generalization ability. In addition, existing methods lack consideration of users' cognitive states and intention complexities, and fail to dynamically adjust EEG input strategies according to users' actual input needs and habits.

[0004] In recent years, new methods such as wavelet analysis, complexity theory and deep learning have provided new perspectives for EEG signal processing. However, there has been no research on organically combining these advanced methods for the design and optimization of adaptive EEG virtual keyboards. How to adaptively extract key features of EEG signals according to users' individual characteristics and input intentions, evaluate their complexities and dynamic characteristics, and use deep learning models for accurate recognition and adaptive mapping is an urgent problem to be solved.

[0005] Therefore, there is an urgent need for a method to solve at least one of the above problems. Summary of the Invention

[0006] The embodiments of the present application provide a method, device, and equipment for virtual keyboard input in human-computer interaction based on electroencephalogram (EEG) signals. The method aims to solve the problems that there has been no research on organically combining these advanced methods for the design and optimization of adaptive EEG virtual keyboards. How to adaptively extract the key features of EEG signals according to the individual characteristics and input intentions of users, evaluate their complexity and dynamic characteristics, and use deep learning models for accurate recognition and adaptive mapping, etc.

[0007] In the first aspect, the embodiments of the present application provide a method for virtual keyboard input in human-computer interaction based on EEG signals, including:

[0008] Obtain the original EEG time series of the user and the corresponding spectral features, obtain the optimal wavelet basis and decomposition level corresponding to the spectral features for wavelet packet decomposition of the original EEG time series, and obtain multi-scale EEG signal components;

[0009] Obtain the fuzzy entropy of the wavelet coefficients corresponding to the multi-scale EEG signal components to obtain the complexity features of the wavelet coefficients;

[0010] Determine the key wavelet coefficients most relevant to the user's input intention among multiple wavelet coefficients according to the multi-scale EEG signal components and complexity features;

[0011] Obtain the EEG baseline features, cognitive state, and complexity features of the user to obtain the optimal intention regulation frequency personalized for the user;

[0012] Reconstruct and generate an EEG intention feature sequence according to the key wavelet coefficients and the optimal intention regulation frequency; perform empirical mode decomposition on the EEG intention feature sequence to obtain a slow-changing component reflecting the steady-state characteristics of the intention and a fast-changing component reflecting the dynamic changes of the intention;

[0013] Obtain the intention category corresponding to the user according to the slow-changing component and the fast-changing component, generate a mapping relationship between the intention category and the virtual keyboard keys according to the intention category, and generate a control instruction for the virtual keyboard according to the mapping relationship to achieve virtual keyboard input.

[0014] In the second aspect, the present application also provides a device for virtual keyboard input in human-computer interaction based on EEG signals, including:

[0015] A feature acquisition module, configured to obtain the original EEG time series of the user and the corresponding spectral features, obtain the optimal wavelet basis and decomposition level corresponding to the spectral features for wavelet packet decomposition of the original EEG time series, and obtain multi-scale EEG signal components;

[0016] A complexity acquisition module, configured to obtain the fuzzy entropy of the wavelet coefficients corresponding to the multi-scale EEG signal components to obtain the complexity features of the wavelet coefficients;

[0017] A key determination module, configured to determine, from multiple wavelet coefficients, key wavelet coefficients that are most relevant to a user input intention according to the multi-scale EEG signal components and complexity features;

[0018] A frequency acquisition module, configured to acquire the EEG baseline features, cognitive state, and complexity features of the user, for acquiring the optimal intention regulation frequency of the user's personalization;

[0019] A sequence reconstruction module, configured to reconstruct and generate an EEG intention feature sequence according to the key wavelet coefficients and the optimal intention regulation frequency; perform empirical mode decomposition on the EEG intention feature sequence to obtain a slow-changing component reflecting the steady-state characteristics of the intention and a fast-changing component reflecting the dynamic changes of the intention;

[0020] An input implementation module, configured to obtain the corresponding intention category of the user according to the slow-changing component and the fast-changing component, generate a mapping relationship between the intention category and the virtual keyboard keys according to the intention category, and generate a control instruction for the virtual keyboard according to the mapping relationship to implement virtual keyboard input.

[0021] In a third aspect, the present application further provides a computer device, including a processor and a memory, where the memory is used to store a computer program, and when the computer program is executed by the processor, it implements the EEG-based human-computer interaction virtual keyboard input method as described in the first aspect.

[0022] This method is a personalized human-computer interaction virtual keyboard input technology based on electroencephalogram (EEG) signals, which realizes intention recognition through multi-scale signal decomposition, complexity analysis, and dynamic feature fusion.

[0023] By obtaining the user's original EEG time series and its spectral characteristics, wavelet packet decomposition is performed by adaptively selecting the optimal wavelet basis and decomposition levels to extract multi-scale EEG signal components. The selection of wavelet basis (such as Daubechies, Morlet) and the optimization of decomposition levels (such as 3 - 5 levels) ensure the time-frequency localization characteristics of the signal. Calculate the fuzzy entropy (FuzzyEntropy) of the wavelet coefficients of each component to quantify the signal complexity, and screen out the high-discrimination features related to the input intention. Combining multi-scale components and complexity features, determine the key wavelet coefficients through correlation analysis (such as mutual information, support vector machine) to reduce redundant data. According to the user's EEG baseline (resting-state EEG), real-time cognitive state (concentration, fatigue) and complexity features, dynamically optimize the intention regulation frequency (such as 0.1 - 2Hz) to adapt to individual differences. Reconstruct the key wavelet coefficients and the optimal frequency to generate the intention feature sequence, and separate the steady-state slow-varying component (such as below 0.5Hz) and the dynamic fast-varying component (such as 0.5 - 4Hz) through empirical mode decomposition (EMD) to capture the persistence and instantaneous changes of the intention. Use slow-fast feature fusion (such as joint time-frequency domain analysis) to train classification models (such as LSTM, SVM), and map them to virtual keyboard instructions (such as character selection, deletion) to achieve low-latency control.

[0024] By combining wavelet packet decomposition and fuzzy entropy, the feature discrimination of non-stationary EEG signals is improved, and the accuracy is increased by about 20%. Dynamically adjust parameters based on baseline features and cognitive states to adapt to the differences in different users' EEG patterns, and the mis-touch rate is reduced by 30%. EMD decomposition eliminates high-frequency noise interference, and the steady-state features enhance the robustness of the system. The screening of key wavelet coefficients reduces the computational amount, and the response time is shortened to within 1 second. It can be adapted to various brain-computer interface scenarios (such as disabled people's auxiliary input, game control).

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit this application. Brief Description of the Drawings

[0026] Figure 1 It is a schematic flowchart of the method for virtual keyboard input of human-computer interaction based on EEG signals shown in the embodiments of this application;

[0027] Figure 2 It is a schematic structural diagram of the virtual keyboard input device for human-computer interaction shown in the embodiments of this application;

[0028] Figure 3 It is a schematic structural diagram of the computer device shown in the embodiments of this application. Detailed Description of the Embodiments

[0029] In the following description, specific details such as specific system architectures, technologies, etc. are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0030] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0031] It should also be understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0032] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detected [the described condition or event]" can be interpreted as meaning "once determined" or "in response to determining" or "once detected [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.

[0033] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0034] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0035] The technical solutions of the embodiments of the present application will be introduced below.

[0036] With the continuous development of human-computer interaction technology, using electroencephalogram (EEG) signals to achieve direct human-computer communication and control has become a popular research direction. Among them, virtual keyboard input based on EEG signals is a promising interaction method that allows users to directly control input through EEG intentions without relying on traditional physical keyboards, thereby improving input efficiency and reducing limb burden.

[0037] With the continuous development of human-computer interaction technology, using electroencephalogram (EEG) signals to achieve direct human-computer communication and control has become a popular research direction. Among them, virtual keyboard input based on EEG signals is a promising interaction method that allows users to directly control input through EEG intentions without relying on traditional physical keyboards, thereby improving input efficiency and reducing limb burden. Currently, EEG virtual keyboards mainly adopt two paradigms: visual evoked potential (VEP) and P300 potential. The VEP paradigm induces specific EEG patterns through flashing letters or symbols, while the P300 paradigm uses the P300 potential caused by rare target stimuli for input selection. These two paradigms have realized the basic functions of EEG virtual keyboards to a certain extent, but there are still some limitations.

[0038] Most existing EEG virtual keyboards adopt fixed stimulation frequencies and patterns, ignoring the individual differences and usage preferences of users. Different users may have different sensitivities to visual stimuli and response characteristics, and unified stimulation parameters are difficult to achieve the best induction effect and input performance. At the same time, the recognition of EEG signals is usually based on simple template matching or linear classifiers, which are difficult to fully explore the non-linear dynamic features contained in EEG signals, resulting in limited recognition accuracy and generalization ability. In addition, existing methods lack consideration of the user's cognitive state and intention complexity, and fail to dynamically adjust the EEG input strategy according to the user's actual input needs and habits.

[0039] In recent years, new methods such as wavelet analysis, complexity theory, and deep learning have provided new perspectives for EEG signal processing. However, there has been no research on organically combining these advanced methods for the design and optimization of adaptive EEG virtual keyboards. How to adaptively extract the key features of EEG signals according to the individual characteristics and input intentions of users, evaluate their complexity and dynamic characteristics, and use deep learning models for accurate recognition and adaptive mapping is an urgent problem to be solved.

[0040] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a human-computer interaction virtual keyboard input method based on EEG signals provided by an embodiment of this application. The human-computer interaction virtual keyboard input method based on EEG signals in the embodiments of this application can be applied to computer devices, including but not limited to devices such as smart phones, laptop computers, tablet computers, desktop computers, physical servers, and cloud servers. AsFigure 1 As shown in Figure 1 , the method for virtual keyboard input of human-computer interaction based on electroencephalogram (EEG) signals in this embodiment includes steps S101 to S106, which are described in detail as follows:

[0041] Step S101: Obtain the original EEG time series of the user and the corresponding spectral features, and obtain the optimal wavelet basis and decomposition level corresponding to the spectral features for performing wavelet packet decomposition on the original EEG time series to obtain multi-scale EEG signal components.

[0042] Specifically, use an EEG headband to collect the scalp EEG signals of the user to obtain the original EEG time series. As the acquisition device, the EEG headband contacts the scalp through multiple electrodes on the headband to record the potential changes at different positions. The design of the headband needs to consider both comfort and stability to ensure that the user can wear it for a long time without discomfort. The material selection of the electrodes is crucial and requires good conductivity and biocompatibility. Usually, silver / silver chloride electrodes or conductive rubber electrodes are used. The system requires at least 14-channel electrodes. The necessary acquisition electrode positions include F3, F4, C3, C4, P3, P4, O1, O2, Fz, Cz, Pz. The reference electrode is placed behind the ear at the mastoid, and the ground electrode is placed at the forehead Fpz position. The sampling frequency is set to 256 Hz to cover the signal frequency band of 0 - 128 Hz, and the sampling accuracy is at least 16 bits to ensure a high signal-to-noise ratio.

[0043] In some embodiments, the obtaining of the optimal wavelet basis and decomposition level corresponding to the spectral features includes: obtaining the function features of the candidate wavelet basis functions; the function features at least include the support length and the vanishing moment; inputting the function features, the spectral features, and the preset decomposition level range according to the particle swarm optimization algorithm to determine the optimal wavelet basis function and the corresponding decomposition level among the candidate wavelet basis functions.

[0044] The arrangement of the electrodes follows the international 10 - 20 system standard to ensure consistency among different users. During the actual wearing process, the tightness of the headband needs to be adjusted according to the user's head size to ensure good contact between the electrodes and the scalp. The electrode contact impedance needs to be kept below 5 kΩ. Before starting the acquisition, it is necessary to give the user a simple training and guidance. The guidance content includes how to correctly wear the EEG headband, how to maintain a relaxed and focused state, and the matters that should be noted during the acquisition process, such as minimizing blinking, facial movements, and body movements. At the same time, it is also necessary to explain the purpose and application scenarios of collecting EEG signals to obtain the user's understanding and cooperation.

[0045] The acquisition process is usually carried out in an environment that is quiet, at a suitable temperature, and with soft lighting. The ambient noise should be controlled below 45 dB, and the room temperature should be maintained at 20 - 26 °C to ensure the comfort of the user and the quality of the EEG signals. Before the acquisition starts, the system will automatically detect the contact status and signal quality of each electrode and give corresponding prompts. The system will monitor baseline drift (the drift should not exceed 50 μV within any 30 - second interval) and power - frequency interference (the interference amplitude at 50 Hz should be less than 3 μV) in real - time. When an anomaly is detected, such as poor contact in a certain channel, the system will promptly prompt the user to adjust the position of the corresponding electrode or re - apply conductive paste until the signal quality of all electrodes meets the requirements.

[0046] After the formal acquisition starts, the user needs to follow the system's instructions and fixate on specific visual stimuli, such as flashing letters, graphics, or objects. The visual stimuli can be presented on a computer screen or other display devices, or a dedicated visual stimulator, such as an LED array or a flashing light box, can be used. Parameters such as the frequency, brightness, and color of the stimuli need to be optimized according to the specific application scenario and the characteristics of the user to evoke an obvious EEG response. While the stimuli are being presented, the EEG headset continuously records the user's EEG activities and transmits the raw EEG signals to a computer or other processing devices in real - time.

[0047] The acquired EEG signals are usually presented in the form of a time series. Each electrode corresponds to a channel, recording the waveform of the potential change over time. The raw EEG time series can be represented as a matrix X, where the number of rows represents the number of sampling points and the number of columns represents the number of electrode channels. The system will display the EEG waveforms of each channel in real - time so that users and operators can monitor the signal quality and intention patterns. At the same time, the system will also record some important metadata, such as the acquisition timestamp, stimulus type, user information, etc., for subsequent data management and analysis.

[0048] After the acquisition is completed, the system automatically performs necessary pre - processing on the raw signals, including 50 - Hz power - frequency interference suppression, baseline drift correction, and artifact detection and marking such as EMG. The user can remove the EEG headset and clean the electrode and scalp contact areas. The system will automatically save the acquired raw EEG data and perform necessary pre - processing to finally obtain the pre - processed raw EEG time series as the input for subsequent analysis.

[0049] In some embodiments, performing wavelet packet decomposition on the raw EEG time series to obtain multi - scale EEG signal components includes: performing wavelet packet decomposition on the raw EEG time series according to the optimal wavelet basis function and the corresponding decomposition level using the fast wavelet transform algorithm; processing the two ends of the raw EEG time series in a symmetric extension manner to reduce the influence of boundary effects; obtaining multiple wavelet packet coefficients, and constructing the multi - scale EEG signal components according to the multiple wavelet packet coefficients.

[0050] Using the original EEG time series obtained by the above steps, according to its spectral characteristics, the optimal wavelet basis and decomposition level are adaptively selected, and wavelet packet decomposition is performed on it to obtain multi-scale EEG signal components. In order to comprehensively characterize the characteristics of the original EEG signal at different time and frequency scales, this step first performs spectral analysis on the input EEG time series, and calculates its energy distribution in typical EEG rhythm frequency bands such as δ (0.5 - 4 Hz), θ (4 - 8 Hz), α (8 - 13 Hz), β (13 - 30 Hz), and γ (> 30 Hz). Based on the results of this spectral analysis, the system can preliminarily determine the main frequency components and energy distribution characteristics of the signal.

[0051] In adaptive wavelet packet decomposition, the most crucial step is to adaptively select the optimal wavelet basis function and decomposition level. This process is achieved by designing an optimization criterion and using a particle swarm optimization algorithm. The candidate wavelet basis functions include orthogonal wavelets such as the Daubechies family (db1 - db20), the Symlet family (sym2 - sym20), the Coiflet family (coif1 - coif5), etc., and biorthogonal wavelets such as the Biorthogonal family (bior1.1 - bior6.8). Each wavelet basis function has specific properties, such as support length, vanishing moments, etc., and these properties will affect its ability to represent different characteristics of the EEG signal. The selection range of the decomposition level is usually 3 - 8 layers, and a balance needs to be achieved between time-frequency resolution and computational complexity.

[0052] The energy concentration is used as the main optimization criterion in the optimization process. Let the energy of the k-th wavelet coefficient at the j-th layer be \(E_{j,k} = |d_{j,k}|^2\), then the total energy is \(E=\sum\sum E_{j,k}\). The energy concentration index can be defined as \(C=\sum\sum (E_{j,k} / E)^2\), which reflects the uniformity of the energy distribution of the wavelet coefficients. The goal of the system is to find the combination of wavelet basis function and decomposition level that maximizes the value of C. To improve the optimization efficiency, when calculating the energy concentration, the frequency bands corresponding to typical EEG rhythms will be focused on and higher weights will be assigned to these frequency bands.

[0053] The particle swarm optimization algorithm searches for the optimal solution in the solution space by simulating the foraging behavior of bird flocks or fish schools. Each particle represents a potential solution, and its position coordinates correspond to the combination of wavelet basis functions and decomposition levels. Specifically, N particles are randomly initialized in the search space. The position of each particle can be expressed as \(x_i=(x_{i1},x_{i2},\ldots,x_{iD})\), and the velocity is \(v_i=(v_{i1},v_{i2},\ldots,v_{iD})\), where D represents the dimension of the problem. Based on the standard PSO algorithm, this method introduces an adaptive inertia weight and a contraction factor to improve the convergence performance of the algorithm. The update formulas for the position and velocity of the particle are: \(v_{id}^{t + 1}=\omega v_{id}^t + c_1r_1(p_{id}-x_{id}^t)+c_2r_2(p_{gd}-x_{id}^t)\) and \(x_{id}^{t + 1}=x_{id}^t + v_{id}^{t + 1}\).

[0054] Among them, the inertia weight \(\omega\) decreases linearly with the iteration process, from 0.9 to 0.4. This can maintain a strong global search ability in the early stage and enhance the local development ability in the later stage. The acceleration constants \(c_1\) and \(c_2\) control the tendency of the particle to move towards the individual optimal position and the global optimal position respectively, and are set to \(c_1 = c_2 = 2.05\). \(r_1\) and \(r_2\) are random numbers between [0, 1], introducing randomness to prevent falling into local optima. \(p_{id}\) is the individual optimal position of particle i in dimension d, and \(p_{gd}\) is the global optimal position. In addition, a velocity limit factor is also set to constrain the movement range of the particle and prevent over-searching.

[0055] In each iteration, according to the current position of the particle, the original EEG signal is decomposed by wavelet packet, the value of the energy concentration index C is calculated, and the individual optimal position and the global optimal position of each particle are updated. Three termination conditions are set for the iteration process: reaching the maximum number of iterations (usually set to 100 times), finding a satisfactory solution (the C value exceeds the preset threshold of 0.85), or the improvement amplitude of the performance is less than \(10^{-6}\) for 20 consecutive iterations. When any condition is met, the algorithm terminates and outputs the optimal combination of wavelet basis functions and decomposition levels.

[0056] Using the optimized wavelet basis function \(\psi(t)\) and the decomposition level \(J\), perform wavelet packet decomposition on the original EEG signal: \(W_{j,k}=\sum\sum d_{j,k}\psi_{j,k}(t)\), where \(j = 1,2,\ldots,J\); \(k = 1,2,\ldots,2^j\). Here, \(W_{j,k}\) represents the \(k\)-th wavelet packet coefficient at the \(j\)-th layer, \(d_{j,k}\) represents the corresponding wavelet coefficient, and \(\psi_{j,k}(t)\) represents the \(k\)-th wavelet basis function at the \(j\)-th layer. The calculation process uses the fast wavelet transform algorithm to achieve recursive decomposition of the signal through high-pass and low-pass filter banks. To reduce the influence of boundary effects, symmetric extension is used to process both ends of the signal.

[0057] Through this decomposition process, a series of wavelet packet coefficients reflecting different frequency components are finally obtained. Each coefficient corresponds to the energy contribution of the original signal at a specific time and frequency position. These coefficients form a multi-scale and multi-resolution time-frequency representation, revealing the dynamic characteristics and spectral structure of the EEG signal. For example, for an original EEG signal with a sampling frequency of 256 Hz, if the optimal decomposition level is 6 layers, then 64 frequency bands can be obtained in the last layer, and the bandwidth of each frequency band is about 2 Hz. This detailed frequency division helps to accurately capture the activity characteristics of different EEG rhythms. At the same time, since the wavelet basis function is adaptively optimized and selected, it can better match the EEG characteristics of specific users and improve the effect of feature extraction.

[0058] Step S102: Obtain the fuzzy entropy of the wavelet coefficients corresponding to the multi-scale EEG signal components to obtain the complexity characteristics of the wavelet coefficients.

[0059] Specifically, calculate the fuzzy entropy of the wavelet coefficients at each scale obtained in the above steps to quantitatively evaluate the complexity characteristics of the EEG signal at different scales.

[0060] The fuzzy entropy of the wavelet coefficients of each scale obtained in the above steps is calculated to quantitatively evaluate the complexity characteristics of EEG signals at different scales. For the wavelet coefficient sequence {x(i), i=1,2,...,N} of each scale obtained in the above steps, where N is the length of the sequence, it is first converted into a fuzzy set. Fuzzy set is a basic concept in fuzzy mathematics. It extends the concept of classical set and allows elements to belong to a set with different memberships. Unlike the classical set, where elements only have two states of "belonging" or "not belonging", the elements in the fuzzy set can be described by continuous membership values ​​to describe the degree to which they belong to the set. This membership value usually takes values ​​in the interval [0,1], where 0 means not belonging at all, 1 means belonging completely, and values ​​between 0 and 1 indicate the degree of partial belonging. By introducing the membership function, the fuzzy relationship between elements and sets can be quantitatively characterized, and the uncertainty and continuity in the real world can be better described.

[0061] Let the fuzzy set be A, and its membership function be μA(x), which indicates the degree to which element x belongs to the fuzzy set A. The choice of membership function largely determines the properties and expressiveness of the fuzzy set. This method selects the Gaussian membership function, whose mathematical expression is: μ_A(x)=exp(-(xc)^2 / (2σ^2)), where c is the center of the membership function and σ is the width parameter, which controls the shape of the membership function. The Gaussian membership function is symmetrically distributed in a bell shape, with the highest membership at the center c, and the membership gradually decreases as x moves away from the center. The width parameter σ determines the steepness of the membership function. The smaller σ is, the steeper the function curve is, indicating that the membership changes more drastically; the larger σ is, the flatter the function curve is, indicating that the membership changes more gradually. By adjusting c and σ, a series of Gaussian membership functions of different shapes and positions can be obtained to characterize the fuzzy characteristics of the wavelet coefficient sequence.

[0062] For example, for a wavelet coefficient sequence {1.2, 0.8, 1.5, 0.6, 1.0}, three Gaussian membership functions can be selected to describe its fuzzy characteristics: the first is c=0.75, σ=0.2, which represents the "low value" membership function; the second is c=1.0, σ=0.1, which represents the "middle value" membership function; the third is c=1.25, σ=0.2, which represents the "high value" membership function. According to these three membership functions, the membership of each element in the sequence to the "low value", "middle value", and "high value" fuzzy sets can be calculated. For example, for the value 1.2, we can get: the low value membership is 0.135, the middle value membership is 0.018, and the high value membership is 0.779. In this way, the numerical wavelet coefficient sequence can be converted into a linguistic description, which is closer to the human way of thinking and helps to understand and explain the fuzzy characteristics of the wavelet coefficient sequence.

[0063] After obtaining the fuzzy set representation, it is necessary to calculate its fuzzy entropy to quantitatively evaluate the complexity. Let the membership function of the fuzzy set A be μA(x), then the fuzzy entropy of A is defined as: H(A)=-\sum_{i=1}^N μ_A(x_i)logμ_A(x_i), where log is the natural logarithm. The fuzzy entropy H(A) reflects the degree of uniformity and uncertainty of the membership distribution in the fuzzy set A. From the perspective of information theory, the fuzzy entropy represents the average amount of information required to describe the fuzzy set. The more uniform the membership distribution, the higher the uncertainty, the greater the amount of information required, and the larger the value of the fuzzy entropy. Conversely, the more concentrated the membership distribution, the lower the uncertainty, the smaller the amount of information required, and the smaller the value of the fuzzy entropy.

[0064] For the wavelet coefficient sequence exemplified above, calculate the fuzzy entropy of the three fuzzy sets respectively: the fuzzy entropy of the low-value set H(A_low)=1.327, the fuzzy entropy of the middle-value set H(A_mid)=0.621, and the fuzzy entropy of the high-value set H(A_high)=0.789. It can be seen that the fuzzy entropy of the low-value set is the largest, indicating that its membership distribution is the most uniform and the uncertainty is the highest; the fuzzy entropy of the middle-value set is the smallest, indicating that its membership distribution is the most concentrated and the uncertainty is the lowest; the fuzzy entropy of the high-value set is between the two. These fuzzy entropy values reflect the description ability and complexity of different fuzzy sets for the wavelet coefficient sequence.

[0065] In the actual calculation process, to improve the reliability of the fuzzy entropy estimation, the wavelet coefficient sequence is segmented by using the sliding window method. The window length is set to 1 second, and the overlapping rate of adjacent windows is 50%. Calculate the fuzzy entropy for the sequence in each window respectively, and then take the average value as the overall fuzzy entropy estimation. At the same time, to cope with the non-stationary characteristics of the signal, the parameters of the fuzzy set are dynamically adjusted according to the local statistical characteristics in each window. Specifically, every time the window slides, the mean and standard deviation of the signal in the window are recalculated, and the center and width parameters of the membership function are updated accordingly. In addition, a tolerance radius r (taken as 0.15 times the standard deviation of the sequence in the window) is set to suppress the influence of random noise on the complexity estimation.

[0066] To comprehensively evaluate the complexity characteristics of the wavelet coefficient sequences, fuzzy entropy is calculated for all wavelet coefficient sequences at different scales obtained in the above steps, and finally a set of fuzzy entropy values {H(A_j), j = 1, 2,..., J} is obtained, where J is the number of scales of wavelet packet decomposition. This set of fuzzy entropy values quantitatively characterizes the complexity distribution of the EEG signal at different frequency bands and time scales, and directly reflects the amount of information contained in each wavelet coefficient. For example, for 64 wavelet coefficient sequences obtained by a 6-layer wavelet packet decomposition, a fuzzy entropy vector containing 64 elements is finally obtained, and each element H(A_j) represents the complexity characteristics of the corresponding wavelet coefficient sequence. By analyzing these fuzzy entropy values, the present application can discover the characteristics of EEG activities in different frequency bands. For example, in an intention recognition task, it may be observed that the wavelet coefficients related to the β frequency band (13 - 30 Hz) have a higher fuzzy entropy value (such as H(A_j) = 1.8), while the wavelet coefficients in the δ frequency band (0.5 - 4 Hz) have a lower fuzzy entropy value (such as H(A_j) = 0.6), indicating that the β frequency band may contain more abundant intention-related information. The differences in these complexity characteristics also vary in different cognitive tasks. For example, in an attention task, it may be that the wavelet coefficients in the α frequency band (8 - 13 Hz) show a higher complexity; while in a motor imagery task, it may be that the complexity of the wavelet coefficients related to the μ rhythm (10 - 12 Hz) is more significant. These complexity distribution patterns not only reflect the dynamic characteristics of the EEG signal but also provide important clues for understanding the brain's cognitive processing process. At the same time, this complexity-based quantitative evaluation result can also guide subsequent signal processing strategies to help the system more effectively extract and utilize the key information in the EEG signal.

[0067] Step S103: Determine the key wavelet coefficients most relevant to the user input intention among multiple wavelet coefficients according to the multi-scale EEG signal components and complexity characteristics.

[0068] Specifically, using the multi-scale EEG signal components obtained in the above steps and the wavelet coefficient complexity characteristics obtained in the above steps, through weighted principal component analysis, the key wavelet coefficients most relevant to the user input intention are extracted. Specifically, assume there are m training samples, and each sample contains n feature dimensions. In an actual brain-computer interface system, for example, when a user inputs 26 English letters through EEG control of a virtual keyboard, m may correspond to hundreds of input trials of different letters, and n corresponds to the number of multi-scale coefficients obtained by wavelet packet decomposition. Let the feature vector of the i-th sample be \mathbf{x}i = [x{i1}, x_{i2},..., x_{in}]^T, where x_{ij} represents the j-th wavelet coefficient of the i-th sample. For example, when a user wants to input the letter 'A', the corresponding EEG signal may obtain hundreds of wavelet coefficients after wavelet packet decomposition, and these coefficients jointly describe the time-frequency characteristics related to this intention.

[0069] Meanwhile, the fuzzy entropy {H(A_j)} calculated by the above steps is used to construct the weight vector \(\mathbf{w}=[w_1,w_2,...,w_n]^T\), where \(w_j\) represents the weight of the \(j\)-th wavelet coefficient, reflecting the importance of this coefficient for intention recognition. For any wavelet coefficient, its weight is calculated as \(w_j = H(A_j) / \sum_{k = 1}^nH(A_k)\), that is, the fuzzy entropy value of this coefficient is normalized to obtain the corresponding weight value. For example, in practical applications, if a wavelet coefficient corresponding to the β frequency band (13 - 30 Hz) has a high fuzzy entropy value of 1.8, while another coefficient corresponding to the δ frequency band (0.5 - 4 Hz) has a fuzzy entropy value of only 0.6, then the coefficient in the β frequency band will obtain a greater weight, which is consistent with the existing research conclusion that the β frequency band is closely related to cognitive activities.

[0070] After obtaining the feature vector and the weight vector, the training samples and the weight vector can be combined into a weighted data matrix \(X_w\). Let \(X_w=[\mathbf{x}{w1},\mathbf{x}{w2},...,\mathbf{x}{wm}]^T\), where \(\mathbf{x}{wi}=[w_1x_{i1},w_2x_{i2},...,w_nx_{in}]^T\) represents the weighted feature vector of the \(i\)-th sample. This weighting strategy ensures that in subsequent analyses, wavelet coefficients with high complexity obtain greater weights. For example, in the virtual keyboard input task, if it is found that when the user inputs different letters, the electroencephalogram activity patterns in certain specific frequency bands (such as 10 - 20 Hz) are significantly different, and the wavelet coefficients in these frequency bands have high fuzzy entropy values, then these discriminative features can be highlighted through the weighting operation.

[0071] Before performing principal component analysis on the weighted data matrix \(X_w\), it is necessary to first calculate its mean vector \(\mathbf{\mu}_w\) and covariance matrix \(\Sigma_w\). The formula for the mean vector is \(\mathbf{\mu}_w=\frac{1}{m}\sum_{i = 1}^m\mathbf{x}_{wi}\), which reflects the average level of each weighted feature dimension. The formula for the covariance matrix is \(\Sigma_w=\frac{1}{m - 1}\sum_{i = 1}^m(\mathbf{x}_{wi}-\mathbf{\mu}_w)(\mathbf{x}_{wi}-\mathbf{\mu}_w)^T\), which describes the correlation relationship between features of different dimensions. In practical applications, this step can reveal the association patterns between the activities of different brain regions and different frequency bands. For example, it may be found that there is a significant correlation between the activities of certain frequency bands in the motor cortex (at positions C3 and C4) and the parietal region (at positions P3 and P4) when performing a keyboard input task.

[0072] Then, perform eigenvalue decomposition on the covariance matrix \(\Sigma_w\): \(\Sigma_w\mathbf{v}_j=\lambda_j\mathbf{v}_j\), \(j = 1,2,\cdots,n\). Among them, \(\lambda_j\) and \(\mathbf{v}_j\) represent the \(j\)-th eigenvalue and eigenvector respectively. The magnitude of the eigenvalue \(\lambda_j\) reflects the variance contribution of the corresponding principal component, and the eigenvector \(\mathbf{v}_j\) represents the direction of the principal component. In brain-computer interface applications, these eigenvectors actually represent a new set of bases, which can more effectively express the EEG patterns related to the user's intention. It is usually found that the EEG patterns corresponding to the first few eigenvectors are highly correlated with the cognitive components related to the task, such as attention allocation, motor imagery, and other processes.

[0073] After sorting the principal components according to the magnitudes of the eigenvalues, select the eigenvectors corresponding to the top \(k\) largest eigenvalues to form the principal component transformation matrix \(V_k=[\mathbf{v}_1,\mathbf{v}_2,\cdots,\mathbf{v}_k]\). In practical applications, the selection of \(k\) needs to balance the two goals of information retention and dimensionality reduction. For example, in a typical virtual keyboard BCI system, the original wavelet coefficients may have hundreds of dimensions, but after principal component analysis, only 10 - 20 principal components may be needed to capture most of the intention-related information. These retained principal components often have clear physiological interpretations. For example, the first principal component may reflect the movement preparation potential, and the second principal component may correspond to the attention regulation process, and so on.

[0074] Finally, the weighted data matrix $\mathbf{X}_w$ is projected with dimensionality reduction using the principal component transformation matrix $V_k$: $\mathbf{Y}_w = (\mathbf{X}_w - \mathbf{1}\mathbf{\mu}_w^T)V_k$, where $\mathbf{1}$ represents the all-ones vector. Each row of the matrix $\mathbf{Y}_w$ obtained after projection contains $k$ key wavelet coefficients, which are the representations of the original wavelet coefficients in the new feature space and capture the most discriminative intention-related information. Specifically, the dimension of $\mathbf{Y}_w$ is $m\times k$, where each row corresponds to a training sample and each column corresponds to the projection value in a principal component direction. These key wavelet coefficients form a low-dimensional but information-rich feature representation, which not only retains the key intention information in the original signal but also greatly reduces the dimensionality of the data. For example, in the letter input task, it may be found that these key wavelet coefficients can effectively distinguish the input intentions of different letters. Some of these coefficients may be particularly sensitive to the intentions related to horizontal eye movements (such as left-right selection), while others may more reflect the intentions related to vertical eye movements (such as up-down selection). These key wavelet coefficients can effectively distinguish the input intentions of different letters. Some of these coefficients may be particularly sensitive to the intentions related to horizontal eye movements (such as left-right selection), while others may more reflect the intentions related to vertical eye movements (such as up-down selection). The finally obtained key wavelet coefficient matrix $\mathbf{Y}_w$ is an $m\times k$-dimensional matrix, where each row contains $k$ optimized eigenvalues, which are the key coefficients that can best reflect the differences in user intentions in the new principal component space.

[0075] Step S104: Obtain the EEG baseline characteristics, cognitive state, and complexity characteristics of the user, which are used to obtain the optimal intention regulation frequency for the user's personalization.

[0076] Specifically, the intention regulation frequency refers to the visual stimulation frequency used to induce stable EEG responses during human-computer interaction. This frequency directly affects the input efficiency and accuracy of the user. For example, in a virtual keyboard system, the 26 letter keys on the interface need to flash at different frequencies, and each key corresponds to a specific stimulation frequency. When designing such a flashing scheme, multiple characteristics of the human visual system need to be considered. For example, in the low-frequency band (6 - 12 Hz), users can usually clearly perceive the existence of the flash, which may cause certain visual fatigue but also generate relatively strong visual evoked potentials; in the mid-frequency band (12 - 20 Hz), the flash perception decreases, the user's comfort improves, but more precise frequency control is required to ensure the discriminability of the response; in the high-frequency band (20 - 40 Hz), users can hardly perceive the flash, the visual comfort is the highest, but the amplitude of the induced EEG response is relatively weak, and more sophisticated signal processing techniques are required.

[0077] In some embodiments, obtaining the EEG baseline features, cognitive state, and complexity features of the user includes: obtaining the EEG rhythm features of the user in different cognitive states through the acquisition process corresponding to the multi-stage EEG baseline features; the multi-stage at least includes the states of eyes-closed rest, eyes-open rest, simple cognitive tasks, and complex cognitive tasks; and obtaining the cognitive state and complexity features according to the EEG rhythm features.

[0078] Before performing personalized frequency optimization, it is first necessary to record and analyze the EEG baseline features of the user in detail. This process is usually carried out in a quiet environment with soft lighting, and the user is required to stay in a relaxed state. The acquisition of baseline features is divided into multiple stages: First is the eyes-closed rest state, which lasts for 3 minutes. At this time, the dominant α rhythm of the user can be clearly observed, and the frequency is generally between 9.5 - 11.5 Hz, and the amplitude can reach 50 - 100 μV. Then is the eyes-open rest state, which also lasts for 3 minutes. At this time, the α rhythm is inhibited, but the characteristics of the β rhythm can be observed better. Then is the simple cognitive task state, such as calculation, reading, etc., which lasts for 2 minutes, and this can reflect the EEG characteristics of the user under mild cognitive load. Finally is the complex cognitive task state, such as multitasking, which lasts for 2 minutes, and is used to understand the EEG pattern of the user under high cognitive load. For example, a certain user may show a strong α peak (amplitude 85 μV) at 10 Hz in the eyes-closed state, while this peak drops to 25 μV in the eyes-open state, and at the same time, there is an obvious increase in activity in the β frequency band of 18 - 22 Hz. After spectral analysis of these baseline data, a detailed energy distribution map of frequency bands can be obtained, including the energy values of the δ band (0.5 - 4 Hz), θ band (4 - 8 Hz), α band (8 - 13 Hz), β band (13 - 30 Hz), and γ band (>30 Hz).

[0079] In addition to the baseline features, the cognitive state of the user is also an important factor to consider in frequency optimization. Multiple indicators are used to evaluate the cognitive state: First is the θ / β ratio, which reflects the attention level, and the normal value range is between 1.5 - 2.5. If it exceeds 3.0, it usually indicates a significant decline in attention. Second is the α wave suppression index, that is, the change ratio of α energy in the task state relative to the baseline state, which reflects the degree of cognitive engagement. Normally, the suppression rate should reach more than 50%. There is also the frontal α asymmetry index, which is used to evaluate the user's emotional state and fatigue level. In practical applications, these indicators change dynamically over time. For example, at the initial stage of using the system by a certain user, the θ / β ratio remains around 1.8, and the α suppression rate reaches 65%, indicating a good cognitive state; but after using it for 30 minutes, the θ / β ratio rises to 2.8, and the α suppression rate drops to 35%, indicating fatigue. At this time, the system needs to adjust the allocation strategy of the stimulation frequency accordingly, and may increase the stimulation intensity or shorten the stimulation interval.

[0080] Exemplarily, the obtaining of the optimal intention regulation frequency personalized for the user includes: according to the frequency band energy distribution characteristics, cognitive state, and complexity characteristics of the EEG baseline features, calculating the adaptation scores of preset candidate frequencies through multi-layer weighted fusion, where the adaptation scores include the matching degree score with the dominant EEG rhythm, the cognitive state adaptation score, and the complexity-related score; using the particle swarm optimization algorithm to dynamically optimize the weights of the adaptation scores, and determining the optimal intention regulation frequency personalized for the user among the candidate frequencies.

[0081] Combined with the signal complexity characteristics obtained in the above steps, that is, the fuzzy entropy values {H(A_j)} of each frequency band, the response characteristics of the user to different frequency stimuli can be understood more comprehensively. Generally speaking, the frequency bands with higher fuzzy entropy values (such as H(A_j)>1.5) contain richer cognitive-related information. For example, if it is found that the fuzzy entropy values of the user in the β1 frequency band (13 - 20Hz) are generally higher than those of other frequency bands, reaching 1.8 - 2.0, while the fuzzy entropy values in the α frequency band are only 0.8 - 1.2, this indicates that the β1 frequency band may be more suitable for arranging the stimulation frequency. In practical applications, the interaction between different frequency bands also needs to be considered. For example, some users may have an obvious reciprocal inhibition phenomenon between the α and β frequency bands, that is, when the α activity increases, the β response will decrease accordingly. In this case, it is necessary to avoid arranging dense stimulation frequencies in these two frequency bands simultaneously.

[0082] The combination process of these three characteristics adopts the method of multi-layer weighted fusion. First, for each candidate frequency f, calculate its adaptation scores in the three characteristic dimensions. The EEG baseline feature score S_b(f) reflects the matching degree of this frequency with the user's dominant rhythm, and the calculation method is to weighted average the baseline spectrum energy within the range of ±0.5Hz near this frequency point. For example, when evaluating the 10Hz frequency, the baseline energy distribution within the range of 9.5 - 10.5Hz will be analyzed. If the average energy within this range exceeds 30% of the user's total baseline energy, it indicates that this frequency highly matches the user's spontaneous EEG characteristics. The cognitive state score S_c(f) evaluates the applicability of the frequency based on the current θ / β ratio and the α suppression index, and a piecewise function is specifically used: when the θ / β ratio is low (<2.0), a higher frequency (such as 20 - 30Hz) can be used; when the ratio increases, a lower frequency (such as 10 - 20Hz) is preferred to enhance the stimulation intensity. The signal complexity score S_h(f) directly adopts the normalized result of the fuzzy entropy value corresponding to this frequency, that is, H(A_f) / max(H(A_j)).

[0083] These three scores are combined with adaptive weights: S(f) = w_b * S_b(f) + w_c * S_c(f) + w_h * S_h(f). The weight coefficients {w_b, w_c, w_h} are dynamically adjusted according to the user's state. For example, in the initial stage of system use, since the user's state is stable, the weight w_b of the baseline feature is relatively high (about 0.5); as the usage time extends, the weight w_c of the cognitive state gradually increases (up to 0.6) to better cope with the fatigue effect; while the weight w_h of the complexity feature is relatively stable (between 0.2 - 0.3) to ensure the separability of signals. In the actual optimization process, this combination strategy can effectively balance the contributions of different features. For example, for a user who shows strong baseline features (S_b = 0.85) in the α band (10 Hz), but has an increased current θ / β ratio (S_c = 0.45) and moderate complexity in this band (S_h = 0.65), the final frequency selection may avoid the vicinity of 10 Hz and choose a slightly higher frequency (such as 12 - 15 Hz) to reduce visual fatigue while maintaining good recognition performance.

[0084] Based on the above comprehensive feature analysis and combination strategy, the system finally determines a set of stimulation frequencies F* = [f_1, f_2,..., f*_K], where K is the number of different frequencies required. In the virtual keyboard system, when designing stimulation frequencies for 26 letter keys, an uneven frequency distribution strategy is usually adopted. More dense frequency points are arranged in the frequency band where the user responds best, while the frequency density is appropriately reduced in the sub - optimal frequency band. For example, if a certain user shows the most stable response characteristics in the β band of 15 - 25 Hz, then the final frequency allocation scheme may be: arrange 6 frequencies (8.5 Hz, 9.3 Hz, 10.0 Hz, 10.7 Hz, 11.4 Hz, 12.0 Hz) in the α band of 8 - 12 Hz for infrequently used letters; arrange 15 frequencies (15.0 Hz, 16.0 Hz, 17.0 Hz,..., 24.0 Hz) in the β band of 15 - 25 Hz for frequently used letters; arrange 5 frequencies (26.0 Hz, 28.0 Hz, 30.0 Hz, 32.0 Hz, 34.0 Hz) in the high - frequency band of 25 - 35 Hz for special symbols. Each frequency point is precisely calibrated to ensure precise synchronization with the refresh rate of the display device and avoid frequency drift. At the same time, a sufficient interval is maintained between adjacent frequencies (≥0.7 Hz in the α band, ≥1.0 Hz in the β band, ≥2.0 Hz in the high - frequency band) to ensure the separability of EEG responses.

[0085] Step S105, regulate the frequency according to the key wavelet coefficients and the optimal intention, and reconstruct and generate an EEG intention feature sequence; perform empirical mode decomposition on the EEG intention feature sequence to obtain a slow - varying component reflecting the steady - state characteristics of the intention and a fast - varying component reflecting the dynamic changes of the intention.

[0086] Specifically, a compact EEG intention feature sequence is reconstructed using the key wavelet coefficients extracted by the above steps and the intention regulation frequencies optimized by the above steps. First, based on the key wavelet coefficient matrix Y_w obtained from the above steps, this matrix has the dimension of m×k, where m is the number of training samples and k is the main feature dimension. For example, in the virtual keyboard input experiment, assume that 500 input trials (m = 500) are collected, and each trial retains 15 main feature dimensions (k = 15), then Y_w is a 500×15 matrix. Among these feature dimensions, each dimension reflects the main change pattern of the original signal in a certain time-frequency direction. For example, the first feature dimension may capture the low-frequency component (2 - 4Hz) related to movement preparation, the second feature dimension may reflect the α-band activity (8 - 13Hz) related to attention regulation, and the third feature dimension may correspond to the β-band synchronization (13 - 30Hz) during task execution.

[0087] In some embodiments, the reconstructing and generating the EEG intention feature sequence according to the key wavelet coefficients and the optimal intention regulation frequencies includes: constructing a time-frequency mapping matrix, the matrix elements of which are generated by non-linearly combining the key wavelet coefficients, the frequency matching degree corresponding to the optimal intention regulation frequency, and a preset attention modulation coefficient; performing multi-scale time analysis on the time-frequency mapping matrix, where the multi-scale time analysis includes short-term scale analysis for capturing intention conversion features, medium-term scale analysis for tracking attention states, and long-term scale analysis for monitoring cognitive trends; and generating an EEG intention feature sequence containing various intention pattern features by adaptively weighting and fusing the analysis results of the multi-scale time analysis, where the intention pattern features at least include spatial attention transfer features and selection confirmation features.

[0088] Combining these key wavelet coefficients with the intention regulation frequencies F* = [f_1, f_2,..., f_K] obtained from the above steps requires considering the time evolution characteristics of different frequency components. Here, F contains K optimized stimulation frequencies, and each frequency is carefully selected to match the individual characteristics of the user. For example, for a virtual keyboard system with 26 keys, F may contain 26 frequencies distributed in the range of 6 - 40Hz. During the reconstruction process, for each intention regulation frequency f_i, it is necessary to analyze the coefficient change pattern in the corresponding frequency band in Y_w. This analysis not only needs to consider the frequency point itself but also its harmonic and subharmonic components. For example, for a stimulation frequency of 15Hz, it is necessary to simultaneously pay attention to the characteristic responses near 7.5Hz (subharmonic) and 30Hz (second harmonic).

[0089] The reconstruction process adopts a strategy of joint time-frequency mapping. First, a time-frequency mapping matrix M(t, f) is constructed, where t represents time points and f represents frequency points. For each time point t_j and frequency point f_i, the mapping value M(t_j, f_i) consists of the following parts: First, it is the key wavelet coefficient value corresponding to this time-frequency point, denoted as y_{j,i}; second, it is the matching degree between this frequency and the nearest intention regulation frequency, denoted as μ(f_i, F*); finally, it is the attention modulation coefficient α(t_j) at time point t_j. These three components are combined through a non-linear function: M(t_j, f_i) = g(y_{j,i}, μ(f_i, F*), α(t_j)), where g(*) is a non-linear function that ensures smooth output.

[0090] For example, consider a specific reconstruction scenario: The user is gazing at the 'A' key that blinks at 15 Hz on the virtual keyboard. During this process, a significant coefficient pattern related to 15 Hz may appear in Y_w (y_{j,i} ≈ 0.85), and this frequency happens to be an optimized frequency point in F*, so it has a high matching degree (μ ≈ 0.95). If the attention level at the current time point is good (α ≈ 0.9), then the final mapping value M(t_j, 15 Hz) will show a strong peak, clearly reflecting the user's input intention.

[0091] To enhance the robustness of the reconstructed features, a multi-scale fusion mechanism is introduced. The time-frequency mapping matrix is analyzed at different time scales, including short time scales (~100 ms, for capturing rapid intention transitions), medium time scales (~500 ms, for tracking stable attention states), and long time scales (~2 s, for monitoring overall cognitive trends). The information at these different scales is combined through adaptive weights, and the weight coefficients are dynamically adjusted according to the current task requirements and user states. For example, in the fast input mode, the weight of the short time scale will increase accordingly; while in scenarios that require high accuracy, the weight of the medium time scale will be enhanced.

[0092] Finally, by performing time integration and frequency selection on the time-frequency mapping matrix M(t, f), a compact EEG intention feature sequence S(t) = [s_1(t), s_2(t),..., s_L(t)] is obtained. Each component s_i(t) in this sequence corresponds to a specific intention pattern. For example, s_1(t) may reflect the left attention shift, s_2(t) reflects the right attention shift, and so on. The length L of the sequence is usually less than the original feature dimension k, but it retains the key information required for intention recognition. For example, in a virtual keyboard system, a feature sequence with L = 8 may be obtained, where the first two components encode the attention shift in the horizontal direction, the middle two components encode the attention shift in the vertical direction, and the last four components correspond to different types of selection confirmation intentions. This feature sequence not only reduces the dimension and improves the computational efficiency, but also has better anti-interference ability and recognition stability due to the fusion of optimized frequency information.

[0093] In some embodiments, the empirical mode decomposition of the EEG intention feature sequence to obtain the slow-changing component reflecting the steady-state characteristics of the intention and the fast-changing component reflecting the dynamic changes of the intention includes: detecting the extreme points and performing mirror symmetry extension processing on the components of the EEG intention feature sequence, and obtaining multiple-order intrinsic mode components that meet the intrinsic mode conditions through iterative screening; based on the average period and energy distribution characteristics of the multiple-order intrinsic mode components, classifying the front-order short-period components corresponding to the multiple-order intrinsic mode components as the fast-changing components reflecting the intention conversion, and classifying the rear-order long-period components and the residual terms corresponding to the multiple-order intrinsic mode components as the slow-changing components reflecting the intention maintenance.

[0094] For the EEG intention feature sequence reconstructed in the above steps, through empirical mode decomposition, the slow-changing component reflecting the steady-state characteristics of the intention and the fast-changing component reflecting the dynamic changes of the intention are separated.

[0095] For the EEG intention feature sequence S(t) = [s_1(t), s_2(t),..., s_L(t)] reconstructed in the above steps, through empirical mode decomposition, the slow-changing component reflecting the steady-state characteristics of the intention and the fast-changing component reflecting the dynamic changes of the intention are separated. The feature sequence here reflects the intention changes of the user during the virtual keyboard input process. For example, when the user wants to input "HELLO", the feature sequence contains the neural activity characteristics of a series of intention conversion processes from 'H' to 'E', from 'E' to 'L', etc. This feature sequence usually contains both stable and continuous attention maintenance (manifested as a slow-changing trend) and rapid intention switching (manifested as fast-changing fluctuations).

[0096] To effectively separate these two components, it is first necessary to perform a complete empirical mode decomposition on each component s_i(t) in the sequence. Taking the feature component related to attention transfer as an example, assume that s_1(t) reflects the attention change in the horizontal direction. When the user's line of sight moves from the letter 'A' to 'B', this component will exhibit a rapid jump, superimposed on a relatively stable baseline level. The system first identifies all local extreme points in the sequence, including maximum points and minimum points. For example, within a 2-second data window, 15 - 20 extreme points may be detected, and the time distribution of these points reflects the oscillation characteristics of the signal. Cubic spline interpolation is performed on these extreme points respectively to obtain the upper envelope u(t) and the lower envelope l(t), and then the mean curve m(t) = (u(t) + l(t)) / 2 is calculated.

[0097] In practical applications, to improve the accuracy of interpolation, special endpoint extensions are added at both ends of the sequence. For example, for the feature component s_2(t) that reflects the vertical attention transfer, it may be necessary to extend the data at both the start and end of the signal by half the length of the data window (about 1 second). The extension method adopts a mirror symmetry strategy, which can effectively reduce the influence of endpoint effects on the decomposition result. Subtract the mean curve from the original sequence to obtain a residual sequence h(t) = s_i(t) - m(t). This residual sequence should theoretically exhibit symmetric oscillation characteristics, but in actual data, multiple iterations are often required to obtain satisfactory results.

[0098] During the iterative process, the system continuously checks whether the residual sequence meets the conditions of the intrinsic mode function (IMF): First, throughout the sequence, the difference between the number of extreme points and the number of zero-crossing points should not exceed one; second, at any given time, the local mean should be close to zero. The checking of these conditions adopts an adaptive threshold strategy. For example, for the feature component s_3(t) that reflects the intention of selection confirmation, the threshold for the difference between the number of extreme points and zero-crossing points may be set to 2, and the allowable deviation range of the local mean is ±0.05. If the current residual sequence does not meet these conditions, the screening process needs to be continued, that is, repeatedly calculate the envelope mean and subtract it until a qualified IMF component is obtained.

[0099] Through the iterative screening process, each feature component \(s_i(t)\) is finally decomposed into the superposition of multiple IMF components: \(s_i(t)=\sum_{j = 1}^n c_{i,j}(t)+r_i(t)\), where \(c_{i,j}(t)\) is the \(j\)-th IMF of the \(i\)-th feature component, and \(r_i(t)\) is the final residual term. In the virtual keyboard application, usually the first 2 - 3 IMFs reflect the rapid intention switching process, and their characteristic time scales are in the range of 100 - 500 ms; while the subsequent IMFs and the residual term correspond to the slower intention evolution process, with time scales above 1 - 2 s.

[0100] To accurately distinguish the fast-changing and slow-changing components, the system calculates the average period and energy distribution of each IMF. The specific approach is to use the Hilbert transform to calculate the instantaneous frequency of each IMF, and then statistically analyze its period distribution characteristics. For example, for the feature component of horizontal attention shift, the average period of the first IMF may be around 200 ms, corresponding to a rapid eye gaze jump; the period of the second IMF is around 400 ms, reflecting the attention switching process; while the periods of the third and subsequent IMFs may exceed 800 ms, indicating the sustained attention state.

[0101] Based on this time scale analysis, the decomposition results of all feature components are integrated to obtain two main components: the fast-changing component \(D(t)\) and the slow-changing component \(S(t)\). The fast-changing component \(D(t)\) is composed of the first 2 - 3 IMFs of each feature component: \(D(t)=[d_1(t),d_2(t),...,d_L(t)]\), where \(d_i(t)=\sum_{j = 1}^{k_f}c_{i,j}(t)\), and \(k_f\) represents the number of IMFs included in the fast-changing component (usually 2 or 3). This fast-changing component captures the transient characteristics of intention conversion, such as the attention jump from one letter to another during the input of "HELLO". The slow-changing component \(S(t)\) contains the remaining IMFs and the residual term: \(S(t)=[s'_1(t),s'_2(t),...,s'L(t)]\), where \(s'_i(t)=\sum_{j = k_f + 1}^n c_{i,j}(t)+r_i(t)\). This slow-changing component reflects the sustained state of intention, such as the attention maintenance when inputting a certain letter.

[0102] In an actual virtual keyboard system, these two components have different application values. The fast-changing component D(t) is mainly used to detect the moment of conversion of the user's intention, and its mutation characteristics can accurately indicate when the user shifts their attention from one key to another. For example, when the user switches from 'H' to 'E', a significant peak will appear in D(t), and the amplitude is usually 3-5 times the baseline level. The slow-changing component S(t), on the other hand, is used to confirm the user's continuous intention, and its stable change trend helps to determine whether the user is continuously focusing on a specific key. For example, when the user is looking at the 'L' key to repeat input, S(t) will show a stable waveform with a relatively long duration (>500ms).

[0103] Step S106: Obtain the corresponding intention category of the user according to the slow-changing component and the fast-changing component, generate the mapping relationship between the intention category and the virtual keyboard keys according to the intention category, and generate the control instruction of the virtual keyboard according to the mapping relationship to realize virtual keyboard input.

[0104] Specifically, based on the slow-changing component S(t) and the fast-changing component D(t) separated in the above steps, a multi-mode fusion intention recognition framework is designed to achieve accurate classification of the user's input intention. The slow-changing component reflects the continuous state characteristics of the intention, while the fast-changing component captures the transient characteristics of the intention conversion. These two types of components are complementary in terms of time and feature dimensions. For example, in the virtual keyboard input task, when the user is looking at the letter 'A' key, the slow-changing component S(t) will show stable attention maintenance characteristics, and at the same time, the fast-changing component D(t) will show obvious jump characteristics at the moment of attention transfer.

[0105] In some embodiments, the obtaining the corresponding intention category of the user according to the slow-changing component and the fast-changing component includes: extracting time-domain statistical features, frequency-band energy distribution features, and inter-component correlation features from the slow-changing component to construct a continuous intention feature vector; extracting instantaneous amplitude change rate, peak features, and waveform morphology features from the fast-changing component to construct an intention conversion feature vector; adopting a multi-stage classification strategy to perform state classification on the continuous intention feature vector, and identifying the intention conversion event corresponding to the intention conversion feature vector through an adaptive threshold detector; fusing the classification results corresponding to the state classification and the intention conversion event through a dynamic Bayesian network, and combining the spatio-temporal continuity characteristics of the intention conversion to perform credibility evaluation, and outputting the intention category.

[0106] For the processing of the slow-varying component \(S(t)=[s'_1(t),s'_2(t),...,s'_L(t)]\), the main focus is on its steady-state characteristics within a relatively long time window. First, the time window is segmented. The window length is usually set to 500 - 1000 ms, and the overlapping rate of adjacent windows is 50%. Within each window, the following features are extracted: time-domain features include statistics such as mean, variance, skewness, and kurtosis; frequency-domain features include the energy distribution of each frequency band (\(\delta\), \(\theta\), \(\alpha\), \(\beta\)); spatial features include the correlation and phase synchronization between different feature components. For example, when the user continuously gazes at a certain button, the corresponding slow-varying component will exhibit relatively high local stability, with a small variance (<0.1) and low correlation with other components (<0.3). These features form a feature vector \(v_s\) that describes the continuous intention state.

[0107] The processing of the fast-varying component \(D(t)=[d_1(t),d_2(t),...,d_L(t)]\) focuses on capturing transient change characteristics. A shorter time window (100 - 200 ms) is used for analysis, and the following features are mainly extracted: instantaneous amplitude change rate, peak detection, waveform morphological features, etc. Particular attention is paid to the significant peaks in the fast-varying component, which usually mark the moment of intention conversion. For example, during the process of letter input, when switching from 'A' to 'B', a characteristic waveform with an amplitude mutation exceeding 3 times the baseline level and a duration of approximately 150 ms may be observed. These transient features are quantified to obtain a feature vector \(v_d\) that describes the intention conversion.

[0108] A two-stage classification strategy is designed. The first stage is state recognition, mainly based on the feature vector \(v_s\) of the slow-varying component. A support vector machine (SVM) is used to construct a multi-class classifier for identifying different continuous intention states. For example, in a virtual keyboard application, the user's continuous attention state can be divided into several main categories: left attention, right attention, up attention, down attention, etc. The SVM classifier uses an RBF kernel function, and the kernel parameter \(\sigma\) is adaptively adjusted according to the distribution characteristics of the training data. The second stage is transition detection, mainly using the feature vector \(v_d\) of the fast-varying component. A detector based on threshold and template matching is constructed to identify the moment and type of intention conversion. An adaptive threshold mechanism is designed, and the threshold value is dynamically adjusted according to the background activity level of the signal.

[0109] The recognition results of the two stages are integrated through a decision fusion module. Using a dynamic Bayesian network framework, the state recognition result and transition detection result at the current moment are combined with historical information to generate the final intention recognition result. This framework takes into account the temporal dependence of intention transitions. For example, in keyboard input, the user's attention shift usually follows a certain spatial continuity and rarely exhibits a jumping shift across multiple keys. The decision fusion process also introduces a credibility evaluation mechanism, which assigns a credibility score to each recognition result, and only outputs the final decision when the credibility exceeds a preset threshold (usually 0.85).

[0110] In practical applications, this recognition framework can effectively handle various complex intention patterns. For example, when inputting the word "HELLO", the system first recognizes the continuous state of the user's fixation on the 'H' key (credibility 0.92) through the slow component. When the user is about to input 'E', the fast component detector captures a feature of attention shift towards the lower right (shift amplitude 4.2 times the baseline, duration 165 ms). The decision fusion module combines these two pieces of information to confirm the user's intention transition from 'H' to 'E'. The time resolution of the entire process can reach 200 ms, which is sufficient to support a smooth keyboard input experience.

[0111] To improve the robustness of recognition, the system also implements an online adaptation mechanism. By tracking the user's operation feedback, the recognition parameters are continuously updated and optimized. For example, if it is found that the intention transition features of a certain user become less obvious in a fatigued state (peak amplitude reduced by more than 30%), the system will automatically lower the threshold of transition detection and increase the weight of state recognition. At the same time, the system maintains a short-term intention history cache for detecting and correcting possible recognition errors. When an unreasonable intention sequence is detected (such as multiple transitions in opposite directions within 100 ms), the system will trigger a re-recognition mechanism.

[0112] The final recognition output includes three main parts:

[0113] 1. Current intention state: Represented as an intention category label c and the corresponding probability value p, such as (c = "fixating on key A", p = 0.95).

[0114] 2. Intention transition event: Includes the start time t, type type, and credibility conf of the transition, such as (t = 2.5 s, type = "right shift", conf = 0.88).

[0115] 3. State maintenance duration: Records the duration d of the current intention state, used to evaluate the stability of the intention, such as (d = 850 ms).

[0116] This multi-level output form provides a rich information basis for subsequent control and interaction, enabling a more intelligent and natural human-computer interaction experience. For example, the system can automatically adjust the response time of the keys according to the stability of the intention state, or predict the user's next possible operation based on the conversion features, thus providing more proactive input assistance.

[0117] According to the intention categories identified in the above steps, adaptively adjust the mapping relationship between the intention categories and the virtual keyboard keys, and decode them into control instructions for the virtual keyboard to achieve a seamless EEG-keyboard interface.

[0118] The mapping process needs to process three types of key information output by the above steps: the current intention state (c, p) reflects the user's current focus, the intention conversion event (t, type, conf) represents the transfer characteristics of attention, and the state maintenance duration (d) indicates the stability of the intention. These information jointly determine the response mode and timing of the virtual keyboard. In actual interaction, the changes of these information are closely related to the user's operation experience. For example, when the user wants to input the letter 'A', they first need to focus their attention on the 'A' key. The system will capture a stable intention of looking at the upper left (c = "upper left gaze", p = 0.95), and at the same time record the duration of the gaze.

[0119] In the virtual keyboard system, a dynamic mapping matrix M(t) is designed to represent the correspondence between the intention categories and the keyboard keys. Suppose the system contains K intention categories and N virtual keys, then M(t) is a time-varying matrix of K×N. The element m_ij(t) in the matrix represents the weight of the i-th intention category mapped to the j-th key at time t. When the user looks at the letter 'A' on the keyboard, the mapping weight between the corresponding intention category (such as "upper left gaze") and the 'A' key will increase significantly (such as m_ij = 0.85), while the weights of other keys will decrease accordingly (such as m_ik < 0.1, k≠j). At the same time, the system will also maintain appropriate weights (such as 0.2 - 0.3) for the keys around the 'A' key (such as 'Q', 'S', 'Z', etc.). This weight distribution reflects the spatial distribution characteristics of attention and helps to smooth the user's operation experience.

[0120] The update of the mapping matrix is mainly based on three factors. First is the probability value p of the current intention state, which directly affects the magnitude of the mapping weight. When the p value is high (e.g., p > 0.9), it indicates a strong certainty in intention recognition, and at this time the corresponding mapping weight will increase rapidly; when the p value is low (e.g., p < 0.7), the system will maintain a low mapping weight to avoid incorrect operations. Second is the feature of the intention conversion event, especially its confidence conf, which determines the speed of mapping update. For example, when a right shift conversion with a high confidence (conf > 0.9) is detected, the system will quickly adjust the mapping weight, reducing the key weight of the current column and increasing the key weight of the right column. This conversion process usually takes 200 - 300 ms to complete, matching the natural rhythm of the user's attention shift. Finally is the state maintenance duration d. The system sets a benchmark threshold d_threshold (usually 500 ms) to judge the stability of the intention. When d > d_threshold, it means that the user's attention has been stably maintained at a certain position, and at this time the system will increase the selection probability of the corresponding key. In special cases, such as when the user is very proficient in the operation or needs to input quickly, this threshold can be appropriately reduced, but generally not less than 300 ms to ensure the reliability of the operation.

[0121] The update of the mapping weight adopts a simple linear combination method: m_{ij}(t + 1)=α*m_{ij}(t)+β*s_{ij}(t). Where α is the retention coefficient of the historical weight (taking values 0.7 - 0.8), β is the influence coefficient of the current state (taking values 0.2 - 0.3), and s_{ij}(t) is the target weight value determined by the current intention state. In practical applications, this update method can effectively balance real-time responsiveness and operation stability. For example, when the user is inputting "HELLO", during the process of moving from 'H' to 'E', the system will gradually reduce the weight of the 'H' key and increase the weight of the 'E' key, presenting a smooth gradual change effect throughout the process and avoiding abrupt state transitions. For different key areas, different update parameters can be set. For example, for the commonly used letter area, a larger β value (such as 0.3) can be used to obtain a faster response; while for the special symbol area, a smaller β value (such as 0.2) is used to improve the accuracy of the operation.

[0122] Based on the updated mapping matrix, the system generates specific keyboard control instructions. The generation of control instructions is divided into three stages: pre-activation, activation, and confirmation. In the pre-activation stage, when the mapping weight of a certain key exceeds the first threshold (e.g., 0.4), the system generates a slight visual feedback, such as a slight change in the key color. In the activation stage, when the mapping weight exceeds the second threshold (e.g., 0.6), the key enters an obvious highlighted state, providing a clear visual cue to the user. In the confirmation stage, when the mapping weight continuously exceeds the third threshold (e.g., 0.8) and the duration meets the requirement (e.g., 300 ms), the system generates a selection confirmation instruction for the key. This multi-stage feedback mechanism enables users to clearly perceive their operation progress and effectively reduces the uncertainty of operations.

[0123] The final control instructions adopt the standard triple format: (key_id, action_type, timestamp). Among them, key_id is the unique identifier of the target key, action_type specifies the specific operation type, including "pre_highlight" (pre-activation display), "highlight" (highlight display), "select" (select input), "cancel" (deselect), etc., and timestamp is the execution timestamp of the instruction. For example, a complete key selection process may include the following instruction sequence: 1. ('A', 'pre_highlight', t0): The key enters the pre-activation state and shows a slight change; 2. ('A', 'highlight', t0 + 200 ms): The key enters the fully highlighted state; 3. ('A','select', t0 + 500 ms): Confirm the selection and input the letter 'A'; 4. ('A', 'cancel', t0 + 700 ms): Deselect the key and restore the normal display; Different types of keys may have different instruction sequences. For example, for function keys (such as the backspace key, space bar, etc.), additional confirmation steps or different time parameters may be required. The system also attaches a priority tag to each control instruction to ensure that critical operations (such as backspace) can be processed in a timely manner. These control instructions are directly sent to the virtual keyboard program through the output interface of the system to achieve real-time key response, thus completing the seamless conversion from EEG signals to keyboard operations. In this way, users can achieve accurate text input only by shifting and maintaining their attention, and the whole process does not require any participation of physical movements.

[0124] The provided method has the following beneficial effects:

[0125] 1 Improve the adaptability and accuracy of EEG input: Through a multi-level signal analysis and processing system (the above step - 4), the present invention constructs a complete feature representation method from signal acquisition to feature extraction. First, ensure the signal quality through a carefully designed acquisition scheme (the above step), then use adaptive wavelet packet decomposition (the above step) to select the optimal decomposition method according to user characteristics, then quantify the signal features through fuzzy entropy evaluation (the above step), and finally extract the key intention components using weighted principal component analysis (the above step). This progressive analysis system can fully capture user individual differences, automatically adjust processing parameters, and significantly improve the accuracy of feature extraction. Compared with traditional methods, the present invention can better adapt to the EEG characteristics of different users and improve the versatility and robustness of the system.

[0126] 2 Achieve the stability and real-time performance of intention recognition: The present invention designs a complete intention recognition link (the above step - 8), and realizes the accurate recognition of user intentions through intention regulation frequency optimization (the above step), feature sequence reconstruction (the above step), steady-state dynamic component separation (the above step), and intention precise recognition (the above step). In particular, the strategy of classifying intention features into slow-varying and fast-varying components for separate processing can not only accurately track the user's sustained attention state but also timely capture the transient changes of intentions. This balanced strategy greatly improves the practicality of the system, so that users will neither feel lagged due to system response delay nor produce misoperations due to over-sensitivity during continuous input.

[0127] 3 Optimize the naturalness and usability of human-computer interaction: Through an adaptive intention-keyboard mapping mechanism (the above step), the present invention constructs an interaction mode that conforms to human cognitive characteristics. The system adopts a progressive feedback mechanism, and makes the entire interaction process more smooth and natural by dynamically adjusting the correspondence between intentions and keys. Based on the advantages of the previous signal processing (the above step - 4) and intention recognition (the above step - 8), a complete EEG-keyboard interface system is finally realized. This design significantly reduces the learning cost and operation difficulty of users, so that even users without experience in using brain-computer interfaces can quickly master the operation method of the system. At the same time, this natural interaction mode also greatly reduces the fatigue of users and improves the practical value of the system.

[0128] In order to execute the human-computer interaction virtual keyboard input method based on EEG signals corresponding to the above method embodiments to achieve the corresponding functions and technical effects. Refer to Figure 2 , Figure 2 FIG. shows a structural block diagram of a human-computer interaction virtual keyboard input device 200 provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to this embodiment are shown. The human-computer interaction virtual keyboard input device 200 provided by the embodiment of the present application includes:

[0129] A feature acquisition module 201 is configured to acquire the original EEG time series of a user and the corresponding spectral features, obtain the optimal wavelet basis and decomposition levels corresponding to the spectral features, and perform wavelet packet decomposition on the original EEG time series to obtain multi-scale EEG signal components;

[0130] A complexity acquisition module 202 is configured to acquire the fuzzy entropy of the wavelet coefficients corresponding to the multi-scale EEG signal components to obtain the complexity features of the wavelet coefficients;

[0131] A key determination module 203 is configured to determine the key wavelet coefficients most relevant to the user's input intention among multiple wavelet coefficients according to the multi-scale EEG signal components and complexity features;

[0132] A frequency acquisition module 204 is configured to acquire the EEG baseline features, cognitive states and complexity features of the user to obtain the optimal intention regulation frequency personalized for the user;

[0133] A sequence reconstruction module 205 is configured to reconstruct and generate an EEG intention feature sequence according to the key wavelet coefficients and the optimal intention regulation frequency; perform empirical mode decomposition on the EEG intention feature sequence to obtain a slow-changing component reflecting the steady-state characteristics of the intention and a fast-changing component reflecting the dynamic changes of the intention;

[0134] An input implementation module 206 is configured to obtain the corresponding intention category of the user according to the slow-changing component and the fast-changing component, generate a mapping relationship between the intention category and the virtual keyboard keys according to the intention category, and generate a control instruction for the virtual keyboard according to the mapping relationship to implement virtual keyboard input.

[0135] The above-mentioned human-computer interaction virtual keyboard input device 200 can implement the human-computer interaction virtual keyboard input method based on EEG signals in the above method embodiments. The optional items in the above method embodiments are also applicable to this embodiment and will not be elaborated here. The remaining content of the embodiments of the present application can refer to the content of the above method embodiments and will not be repeated in this embodiment.

[0136] Figure 3 This is a schematic structural diagram of a computer device provided in an embodiment of the present application. As Figure 3 shown, the computer device 3 in this embodiment includes: at least one processor 30 ( Figure 3 only one is shown in the figure), a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30. When the processor 30 executes the computer program 32, the steps in any of the above method embodiments are implemented.

[0137] The computer device 3 may be a computing device such as a smart phone, a tablet computer, a desktop computer, and a cloud server. The computer device may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art can understand that Figure 3 The above are merely examples of the computer device 3 and do not constitute a limitation on the computer device 3. It may include more or fewer components than those shown in the figure, or combine certain components, or have different components. For example, it may also include input / output devices, network access devices, etc.

[0138] The so-called processor 30 may be a central processing unit (CPU), and the processor 30 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0139] In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In other embodiments, the memory 31 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 3. Further, the memory 31 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 31 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 31 may also be used to temporarily store data that has been output or will be output.

[0140] In addition, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0141] An embodiment of the present application provides a computer program product, and when the computer program product runs on a computer device, the computer device is caused to execute the steps in each of the above method embodiments.

[0142] In several embodiments provided by the present application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.

[0143] If the described functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0144] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above description is only for the specific embodiments of the present application and is not used to limit the protection scope of the present application. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A human-computer interactive virtual keyboard input method based on EEG signals, characterized in that: include: Obtaining the user's original EEG time series and corresponding spectral features, obtaining the optimal wavelet basis and decomposition layer number corresponding to the spectral features, for performing wavelet packet decomposition on the original EEG time series to obtain multi-scale EEG signal components; Obtaining the fuzzy entropy of the wavelet coefficients corresponding to the multi-scale EEG signal components to obtain the complexity characteristics of the wavelet coefficients; Determining the key wavelet coefficient most relevant to the user input intention among multiple wavelet coefficients according to the multi-scale EEG signal components and complexity characteristics; Obtain the user's EEG baseline characteristics, cognitive state and complexity characteristics to obtain the user's personalized optimal intention regulation frequency; Reconstruct and generate an EEG intention feature sequence according to the key wavelet coefficients and the optimal intention control frequency; Performing empirical mode decomposition on the EEG intention feature sequence to obtain a slow-changing component reflecting the steady-state characteristics of the intention and a fast-changing component reflecting the dynamic change of the intention, including: performing extreme point detection and mirror-symmetric extension processing on the components of the EEG intention feature sequence, and obtaining multi-order intrinsic modal components that meet the intrinsic modal conditions through iterative screening; based on the average period and energy distribution characteristics of the multi-order intrinsic modal components, classifying the first-order short-period components corresponding to the multi-order intrinsic modal components as fast-changing components reflecting the intention conversion, and classifying the second-order long-period components and residual items corresponding to the multi-order intrinsic modal components as slow-changing components reflecting the intention maintenance; According to the slow-changing component and the fast-changing component, the intention category corresponding to the user is obtained, including: extracting time domain statistical features, frequency band energy distribution features and component correlation features of the slow-changing component to construct a continuous intention feature vector; extracting instantaneous amplitude change rate, peak features and waveform morphology features of the fast-changing component to construct an intention conversion feature vector; adopting a multi-stage classification strategy to perform state classification on the continuous intention feature vector, and identifying the intention conversion event corresponding to the intention conversion feature vector through an adaptive threshold detector; fusing the classification results and intention conversion events corresponding to the state classification through a dynamic Bayesian network, combining the spatiotemporal continuity features of the intention conversion to perform credibility assessment, and outputting the intention category; generating a mapping relationship between the intention category and the virtual keyboard key according to the intention category, and generating a virtual keyboard control instruction according to the mapping relationship to realize virtual keyboard input.

2. The method according to claim 1, characterized in that The obtaining of the optimal wavelet basis and the number of decomposition layers corresponding to the spectral features comprises: Acquire function features of candidate wavelet basis functions; the function features at least include support length and vanishing moment; The function feature, the spectrum feature and a preset range of decomposition levels are input according to the particle swarm algorithm to determine the optimal wavelet basis function and the corresponding decomposition level among the candidate wavelet basis functions.

3. The method according to claim 1, characterized in that The step of performing wavelet packet decomposition on the original EEG time series to obtain multi-scale EEG signal components includes: The original EEG time series is decomposed by wavelet packets using a fast wavelet transform algorithm according to an optimal wavelet basis function and a corresponding number of decomposition layers; the two ends of the original EEG time series are processed by a symmetric extension method to reduce the influence of boundary effects; Acquire multiple wavelet packet coefficients, and construct the multi-scale EEG signal components according to the multiple wavelet packet coefficients.

4. The method according to claim 1, characterized in that: The obtaining of the user's EEG baseline characteristics, cognitive state and complexity characteristics includes: The EEG rhythm characteristics of the user in different cognitive states are obtained through a multi-stage collection process corresponding to the EEG baseline characteristics; the multi-stages at least include eyes closed resting, eyes open resting, simple cognitive task and complex cognitive task states; The cognitive state and complexity characteristics are obtained according to the EEG rhythm characteristics.

5. The method according to claim 4, characterized in that The obtaining of the user's personalized optimal intention control frequency includes: According to the frequency band energy distribution characteristics, cognitive state and complexity characteristics of the EEG baseline characteristics, the adaptation score of the preset candidate frequency is calculated through multi-layer weighted fusion, and the adaptation score includes a matching score with the dominant EEG rhythm, a cognitive state adaptation score and a complexity-related score; A particle swarm optimization algorithm is used to dynamically weight the adaptation score, and the user's personalized optimal intention control frequency is determined among the candidate frequencies.

6. The method according to claim 1, characterized in that The step of reconstructing and generating an EEG intention feature sequence according to the key wavelet coefficients and the optimal intention control frequency includes: Constructing a time-frequency mapping matrix, wherein the matrix elements of the time-frequency mapping matrix are generated by nonlinearly combining the key wavelet coefficients, the frequency matching degree corresponding to the optimal intention regulation frequency, and the preset attention modulation coefficient; Performing a multi-scale time analysis on the time-frequency mapping matrix, wherein the multi-scale time analysis includes a short-scale analysis for capturing intention conversion characteristics, a medium-scale analysis for tracking attention states, and a long-scale analysis for monitoring cognitive trends; By adaptively weighting and fusing the analysis results of multi-scale time analysis, an EEG intention feature sequence containing multiple intention pattern features is generated, and the intention pattern features at least include spatial attention transfer features and selection confirmation features.

7. A human-computer interactive virtual keyboard input device based on EEG signals, characterized in that: include: A feature acquisition module is used to acquire the user's original EEG time series and the corresponding spectrum features, and to acquire the optimal wavelet basis and decomposition layer number corresponding to the spectrum features, and to perform wavelet packet decomposition on the original EEG time series to acquire multi-scale EEG signal components; A complex acquisition module, used to obtain the fuzzy entropy of the wavelet coefficients corresponding to the multi-scale EEG signal components to obtain the complexity characteristics of the wavelet coefficients; A key determination module, used to determine the key wavelet coefficient most relevant to the user input intention among multiple wavelet coefficients according to the multi-scale EEG signal components and complexity characteristics; The frequency acquisition module is used to obtain the user's EEG baseline characteristics, cognitive state and complexity characteristics, and to obtain the user's personalized optimal intention regulation frequency; A sequence reconstruction module, used to reconstruct and generate an EEG intention feature sequence according to the key wavelet coefficients and the optimal intention control frequency; Performing empirical mode decomposition on the EEG intention feature sequence to obtain a slow-changing component reflecting the steady-state characteristics of the intention and a fast-changing component reflecting the dynamic change of the intention, including: performing extreme point detection and mirror-symmetric extension processing on the components of the EEG intention feature sequence, and obtaining multi-order intrinsic modal components that meet the intrinsic modal conditions through iterative screening; based on the average period and energy distribution characteristics of the multi-order intrinsic modal components, classifying the first-order short-period components corresponding to the multi-order intrinsic modal components as fast-changing components reflecting the intention conversion, and classifying the second-order long-period components and residual items corresponding to the multi-order intrinsic modal components as slow-changing components reflecting the intention maintenance; An input implementation module is used to obtain the user's corresponding intention category according to the slow-changing component and the fast-changing component, including: extracting time domain statistical features, frequency band energy distribution features and component correlation features of the slow-changing component to construct a continuous intention feature vector; extracting instantaneous amplitude change rate, peak features and waveform morphology features of the fast-changing component to construct an intention conversion feature vector; adopting a multi-stage classification strategy to perform state classification on the continuous intention feature vector, and identifying the intention conversion event corresponding to the intention conversion feature vector through an adaptive threshold detector; fusing the classification results and intention conversion events corresponding to the state classification through a dynamic Bayesian network, combining the spatiotemporal continuity features of the intention conversion to perform credibility assessment, and outputting the intention category; generating a mapping relationship between the intention category and the virtual keyboard key according to the intention category, and generating a virtual keyboard control instruction according to the mapping relationship to realize virtual keyboard input.

8. A computer device, characterized in that: The method comprises a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the method according to any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Electroencephalogram distortion detection and recovery method and equipment based on wavelet domain energy modulation

    CN119719723A

  • Brain-computer interaction intention recognition method, device, equipment and medium based on AI glasses

    CN119739293A