Multimodal AI Glasses Visual-EEG Collaborative Control Method, Device and Equipment

By constructing a multimodal data acquisition framework and feature fusion technology, the problems of multimodal signal timing alignment and feature matching are solved, the coordinated control of visual and EEG signals is realized, and the accuracy and efficiency of interactive intention recognition are improved.

CN120085760BActive Publication Date: 2025-07-25XIAOZHOU TECH CO LTD

Patent Information

Application Number
CN202510545530.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-25
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing human-computer interaction methods lack the collaborative utilization of multimodal information, resulting in difficulty in timing alignment of visual and EEG signals, inaccurate feature extraction, and affecting the accuracy and efficiency of interactive intent recognition.

Method used

A multimodal data acquisition framework is built, and aligned data flow is generated through time stamp marking, visual scene analysis and time domain decomposition of EEG signals, attention maps and intention features are generated, feature matrices are built for modal alignment and feature fusion, unified representations are generated, timing segmentation and dynamic mode analysis are performed, and control sequences are generated.

Benefits of technology

It improves the real-time and robustness of multimodal data processing, enhances control reliability in complex scenarios, reduces user cognitive load, and is suitable for high-speed interactive scenarios such as AR navigation and emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120085760B_ABST
    Figure CN120085760B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of brain-computer interfaces, and particularly to a multi-modal AI glasses vision-EEG collaborative control method, device and equipment. The method includes: constructing a multi-modal data acquisition framework for collecting visual scene images and EEG signals, generating an aligned data stream to identify target features and generate an attention map; extracting temporal features and identifying intention features; constructing a feature matrix, performing modal alignment on the feature matrix, and performing feature fusion based on the alignment result of the modal alignment to generate a unified representation; performing temporal segmentation, extracting associated features, constructing a state sequence based on the associated features, using the state sequence to mark transition nodes, and generating a dynamic pattern according to the transition nodes; performing information analysis using the dynamic pattern, determining modal weights, performing feature selection based on the modal weights, and performing classification mapping on the feature selection result to generate a control sequence; generating an interaction instruction according to the control sequence to complete the multi-modal AI glasses vision-EEG collaborative control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of brain-computer interfaces, and in particular to a multi-modal AI glasses vision-electroencephalogram collaborative control method, device and equipment. Background Art

[0002] Multi-modal AI glasses can simultaneously acquire users' visual information and electroencephalogram signals by integrating visual sensors and electroencephalogram acquisition electrodes, providing new possibilities for realizing natural human-computer interaction. However, existing human-computer interaction methods often process visual or electroencephalogram signals separately, lacking the collaborative utilization of multi-modal information. This fragmented processing method has several key problems: due to the lack of a unified data acquisition and synchronization mechanism, it is impossible to ensure the temporal alignment of visual and electroencephalogram signals; single-modal feature extraction cannot capture complementary information between modalities, reducing the accuracy of interaction intention recognition; independent control strategies are difficult to achieve collaborative decision-making of visual attention and electroencephalogram intention, affecting the naturalness and efficiency of interaction.

[0003] There are many challenges in the collaborative processing of multi-modal signals in the prior art. There are significant differences in sampling characteristics and data structures between visual and electroencephalogram signals, and the temporal alignment and feature matching of this heterogeneous data become key problems. The signal quality of both modalities is easily affected by the external environment and user status, making it difficult to establish a stable and reliable signal quality evaluation mechanism. There is a complex spatio-temporal correlation between visual attention and electroencephalogram intention, and traditional feature extraction methods are difficult to effectively express this correlation relationship, resulting in insufficient accuracy of interaction intention recognition. Computational resource limitations in real-time interaction scenarios, as well as potential mutual interference between modalities, all pose great challenges to collaborative control. To address these problems, the present invention proposes an intelligent control method based on vision-electroencephalogram collaboration.

[0004] Therefore, there is an urgent need for a method to solve at least one of the above problems. Summary of the Invention

[0005] The embodiments of this application provide a multi-modal AI glasses vision-electroencephalogram collaborative control method, device and equipment. The method aims to solve the many challenges in the collaborative processing of multi-modal signals in the prior art. There are significant differences in sampling characteristics and data structures between visual and electroencephalogram signals, and the temporal alignment and feature matching of this heterogeneous data become key problems. The signal quality of both modalities is easily affected by the external environment and user status, making it difficult to establish a stable and reliable signal quality evaluation mechanism. There is a complex spatio-temporal correlation between visual attention and electroencephalogram intention, and traditional feature extraction methods are difficult to effectively express this correlation relationship, resulting in insufficient accuracy of interaction intention recognition. Computational resource limitations in real-time interaction scenarios, as well as potential mutual interference between modalities, all pose great challenges to collaborative control.

[0006] In a first aspect, an embodiment of the present application provides a method for visual - electroencephalogram collaborative control of a multi - modal AI glasses, including:

[0007] Construct a multi - modal data acquisition framework for collecting visual scene images and electroencephalogram signals, perform timestamp marking based on the collected visual scene images and electroencephalogram signals, and generate an aligned data stream;

[0008] Perform scene analysis on the visual part of the aligned data stream to identify target features, perform spatio - temporal positioning based on the target features to generate an attention map; perform time - domain decomposition on the electroencephalogram part of the aligned data stream to extract temporal features, perform spectral conversion based on the temporal features to obtain spectral features for identifying intention features output by the attention state;

[0009] Construct a feature matrix according to the attention map and intention features, perform modal alignment on the feature matrix, and perform feature fusion based on the alignment result of the modal alignment to generate a unified representation;

[0010] Perform time - series segmentation on the unified representation, extract correlation features, construct a state sequence based on the correlation features, use the state sequence to mark transition nodes, generate a dynamic pattern according to the transition nodes; perform information analysis using the dynamic pattern to determine modal weights, perform feature selection based on the modal weights, and perform classification mapping on the feature selection result to generate a control sequence;

[0011] Generate an interaction instruction according to the control sequence to complete the visual - electroencephalogram collaborative control of the multi - modal AI glasses.

[0012] In a second aspect, the present application further provides a device for visual - electroencephalogram collaborative control of a multi - modal AI glasses, including:

[0013] A data generation module for constructing a multi - modal data acquisition framework for collecting visual scene images and electroencephalogram signals, performing timestamp marking based on the collected visual scene images and electroencephalogram signals, and generating an aligned data stream;

[0014] A scene analysis module for performing scene analysis on the visual part of the aligned data stream to identify target features, performing spatio - temporal positioning based on the target features to generate an attention map; performing time - domain decomposition on the electroencephalogram part of the aligned data stream to extract temporal features, performing spectral conversion based on the temporal features to obtain spectral features for identifying intention features output by the attention state;

[0015] A matrix construction module for constructing a feature matrix according to the attention map and intention features, performing modal alignment on the feature matrix, and performing feature fusion based on the alignment result of the modal alignment to generate a unified representation;

[0016] A time series segmentation module, which is used to perform time series segmentation on the unified representation, extract associated features, construct a state sequence based on the associated features, mark transition nodes using the state sequence, and generate a dynamic pattern according to the transition nodes; perform information analysis using the dynamic pattern, determine modal weights, perform feature selection based on the modal weights, and perform classification mapping on the feature selection result to generate a control sequence;

[0017] An instruction generation module, which is used to generate interaction instructions according to the control sequence to complete the visual-electroencephalogram collaborative control of the multi-modal AI glasses.

[0018] In a third aspect, the present application also provides a computer device, including a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, it implements the multi-modal AI glasses visual-electroencephalogram collaborative control method as described in the first aspect.

[0019] This method is a multi-modal AI glasses collaborative control technology that combines visual scene images and electroencephalogram signals, and specifically includes the following core steps: synchronously collect visual scene images (through a camera) and electroencephalogram signals (through an EEG sensor), and perform timestamp marking to generate an aligned multi-modal data stream. Visual part: Perform scene analysis on the image to identify target features (such as objects, gestures, text), and generate an attention map (such as target position, movement trajectory) through spatio-temporal positioning. Electroencephalogram part: Perform time-domain decomposition (such as filtering, segmentation) on the electroencephalogram signal, extract time series features, and then extract spectral features through spectral conversion (such as Fourier transform, wavelet analysis) to identify the user's attention state and intention (such as focus, selection, rejection). Construct a cross-modal feature matrix from the visual attention map and the electroencephalogram intention features, and achieve feature fusion through modal alignment (such as attention mechanism, matrix projection) to generate a unified representation. Perform time series segmentation and correlation analysis on the fused data, construct a state sequence (such as user behavior pattern), adjust the modal weights through a dynamic pattern (such as Markov chain, reinforcement learning strategy), and optimize feature selection. Based on classification mapping (such as neural network, decision tree), convert the optimized features into a control sequence (such as menu selection, cursor movement), generate interaction instructions, and achieve the collaborative control of the AI glasses (such as voice broadcast, AR display).

[0020] The complementarity between vision and EEG signals (vision captures the external environment, and EEG reflects the user's intention) reduces the misjudgment of a single modality and enhances the control reliability in complex scenarios. Timestamp alignment and dynamic mode optimization ensure the real-time processing of multimodal data and are applicable to high-speed interaction scenarios (such as AR navigation and emergency response). By dynamically adjusting the modality weights, it adapts to different user habits and environmental changes (such as light interference and EEG noise), improving the system robustness. Combining the fixation point (vision) and intention (EEG) realizes seamless control of "seeing what you think", reduces the user's cognitive load, and is applicable to the assisted interaction of disabled persons.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic flowchart of the vision-EEG collaborative control method of the multimodal AI glasses shown in the embodiments of this application;

[0023] Figure 2 It is a schematic structural diagram of the vision-EEG collaborative control device of the multimodal AI glasses shown in the embodiments of this application;

[0024] Figure 3 It is a schematic structural diagram of the computer device shown in the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0026] It should be understood that when used in the specification and the appended claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0027] It should also be understood that the term "and / or" used in the specification and the appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0028] As used in the specification of this application and the appended claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".

[0029] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and shall not be construed as indicating or implying relative importance.

[0030] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0031] The technical solutions of the embodiments of this application will be introduced below.

[0032] By integrating a vision sensor and electroencephalogram (EEG) acquisition electrodes, the multimodal AI glasses can simultaneously obtain the visual information and EEG signals of the user, providing new possibilities for realizing natural human-computer interaction. However, existing human-computer interaction methods often process visual or EEG signals separately and lack the collaborative utilization of multimodal information. This fragmented processing method has multiple key problems: due to the lack of a unified data acquisition and synchronization mechanism, it is impossible to ensure the temporal alignment of visual and EEG signals; the feature extraction of a single modality cannot capture the complementary information between modalities, reducing the accuracy of interaction intention recognition; independent control strategies are difficult to achieve the collaborative decision-making of visual attention and EEG intention, affecting the naturalness and efficiency of interaction.

[0033] There are many challenges in the collaborative processing of multimodal signals in the prior art. There are significant differences in the sampling characteristics and data structures between visual and electroencephalogram (EEG) signals. The temporal alignment and feature matching of such heterogeneous data have become key problems. The signal quality of both modalities is vulnerable to external environments and user states, making it difficult to establish a stable and reliable signal quality evaluation mechanism. There is a complex spatio-temporal correlation between visual attention and EEG intention, and traditional feature extraction methods are difficult to effectively express this correlation, resulting in insufficient accuracy in the recognition of interaction intentions. The computational resource limitations in real-time interaction scenarios, as well as potential mutual interference between modalities, have posed great challenges to collaborative control. To address these problems, the present invention proposes an intelligent control method based on visual-EEG collaboration.

[0034] Please refer to Figure 1 , Figure 1 FIG. [FIGURE NUMBER] is a schematic flowchart of a visual-EEG collaborative control method for a multimodal AI glasses provided by an embodiment of the present application. The visual-EEG collaborative control method for the multimodal AI glasses in the embodiment of the present application can be applied to a computer device, which includes but is not limited to devices such as smart phones, laptop computers, tablet computers, desktop computers, physical servers, and cloud servers. As Figure 1 shown, the visual-EEG collaborative control method in this embodiment includes steps S101 to S105, which are described in detail as follows:

[0035] Step S101: Construct a multimodal data acquisition framework for acquiring visual scene images and EEG signals, and perform timestamp marking based on the acquired visual scene images and EEG signals to generate an aligned data stream.

[0036] Specifically, a multimodal data acquisition framework is constructed to acquire visual scene images and EEG signals; timestamp marking is performed based on the acquired data; and an aligned data stream is generated according to the time marks.

[0037] In some embodiments, the constructing of the multimodal data acquisition framework for acquiring visual scene images and EEG signals includes: using a preset visual sensor to acquire visual scene images at a preset rate; using multiple dry electrodes to acquire EEG signals at a preset frequency, where the dry electrodes use a nano-coating to reduce contact impedance and are configured with an active shielding ring to suppress electromagnetic interference; and respectively performing visual data enhancement and EEG signal conditioning to acquire the visual scene images and EEG signals.

[0038] Build a multi-modal data acquisition framework to collect visual scene images and electroencephalogram (EEG) signals. The AI glasses use a high-performance CMOS image sensor with a resolution of 3840×2160, a field of view of 120 degrees, and a frame rate adjustable between 30 - 60 fps. The system dynamically adjusts the exposure parameters through an adaptive exposure control algorithm to maintain stable image quality under different environmental conditions. The sensor adopts pixel-level HDR technology, achieving a dynamic range of 120 dB in a single exposure, ensuring clear target details can be obtained in both strong light and low light environments. The image signal processing unit uses a dual-core architecture. One core is dedicated to real-time exposure parameter calculation and image enhancement, while the other core performs scene analysis and data preprocessing. Eight high-performance dry electrodes are integrated into the glasses frame, including sampling sites FP1, FP2, F3, F4, C3, C4, O1, and O2, with a sampling frequency of 1000 Hz and a signal-to-noise ratio better than 45 dB. Each electrode is coated with a nano material, significantly reducing the contact impedance and improving the signal acquisition stability. The actively shielded ring designed around the electrodes can reduce environmental electromagnetic interference, with a shielding effect better than 60 dB. The system uses adaptive impedance matching technology to monitor the electrode status in real time and, in conjunction with a programmable gain amplifier, dynamically adjusts the signal amplification factor. The signal conditioning circuit uses a low-noise amplifier and a high-precision analog-to-digital converter, with an input noise less than 1 μV, meeting the requirements for collecting weak EEG signals. The data acquisition unit uses a pipeline architecture, supporting parallel acquisition and caching of visual and EEG data, and finally outputs two original data streams: a visual data stream with 8.24 million pixels per frame and an 8-channel EEG data stream with 1000 sampling points per second.

[0039] Based on the two original data streams obtained above, the system performs timestamp marking. A temperature-compensated crystal oscillator is used as the system's main clock, with a frequency of 100 MHz and a frequency stability better than 0.1 ppm. The clock management unit generates the clock signals required by each subsystem through phase-locked loop technology and establishes a unified time reference. In the visual data stream, the image acquisition module adds a 64-bit timestamp to each frame of image, accurate to the microsecond level. The timestamp information is directly written into the image file header and transmitted and stored together with the image data. At the same time, parameters such as the start time of exposure and the exposure duration are recorded for accurate reconstruction of the image acquisition timing. The EEG data stream adopts a data packet structure. Each data packet contains 512 sampling points, and the start time of the packet and the sampling interval are recorded at the same time. The data packet header contains a 32-bit timestamp and status flag bits for data synchronization and status monitoring. The system designs a large-capacity circular buffer that can temporarily store 30 seconds of original data. The buffer adopts a dual-port structure, supporting simultaneous writing and reading of data to avoid data loss. The timestamp generation module has a watchdog function and automatically switches to the backup clock source when a clock anomaly is detected. To adapt to data streams with different sampling frequencies, the system implements a configurable timestamp resolution and can adjust the time accuracy according to actual needs, finally obtaining visual data packets and EEG data packets with microsecond-level timestamps.

[0040] Using the visual and EEG data packets with accurate timestamps, the system generates aligned data streams through the dynamic time window method. The data alignment module first calculates the maximum sampling period of the two data streams and uses it as the basic time window unit. The window size can be dynamically adjusted within the range of 50 ms to 200 ms, which not only ensures the real-time nature of the data but also takes into account the integrity of signal processing. Within each time window, the system first aligns the timestamps at the window boundaries to establish the timing correspondence between different data streams. For visual data, the system calculates the exact position of the image frame within the time window and records the offset relative to the start time of the window. For EEG data, the system calculates the theoretical time of each data point according to the sampling frequency and aligns it to the unified time axis through linear interpolation algorithm. Considering the inherent delays brought by sensors and signal processing, the system introduces a delay compensation mechanism in the data stream. The delay in the visual channel mainly comes from the exposure time and data transmission of the image sensor, generally within 20 ms. The delay in the EEG channel mainly comes from the signal conditioning circuit and digital filtering, with a typical value of 5 ms. The system calculates the channel delay through timestamps and compensates it during data alignment. The aligned data is organized in a unified format, including a frame header, time information, visual data block, EEG data block, and status information. The system monitors the data quality in real time, calculates indicators such as the clarity and motion blur degree of visual data, and evaluates the signal-to-noise ratio and spectral characteristics of EEG signals at the same time. When an abnormal data quality is detected, the system dynamically optimizes by adjusting the acquisition parameters and processing parameters, and finally outputs a multi-modal data stream with time synchronization.

[0041] Step S102: Perform scene analysis on the visual part of the aligned data stream to identify target features, perform spatio-temporal localization based on the target features, and generate an attention map; perform time-domain decomposition on the EEG part of the aligned data stream to extract temporal features, and perform spectral conversion based on the temporal features to obtain spectral features for identifying the intention features output by the attention state.

[0042] Specifically, perform scene analysis on the visual frames in the aligned data stream obtained in the above steps to identify target features. The system extracts visual data blocks from the synchronized data stream, and each data block contains a high-definition image frame and corresponding timestamp information. For each image frame, first perform multi-scale decomposition to construct a 5-layer image pyramid structure, and the resolution ratio of adjacent levels is 2:1. At each level, an improved deep convolutional network is used to extract scene features. The network contains three parallel branches: the spatial feature branch uses 3×3 and 5×5 convolutional kernels to extract the shape and texture information of the target; the channel feature branch realizes cross-channel information integration through 1×1 convolution; the edge feature branch uses direction-sensitive convolutional kernels to extract the target contour. The features of these three branches are fused through an adaptive weight mechanism, and the weight coefficients are dynamically adjusted according to the feature response intensity. The scene analysis module also calculates the global semantic descriptor and local saliency features of the image to establish a hierarchical scene representation. For moving targets, the system uses timestamp information to calculate the feature differences between adjacent frames and evaluate the changes in target states. Each layer of the multi-scale feature extraction network is equipped with a feature selection module to screen key features according to the discriminability and stability of the features. After feature extraction and selection, the system obtains the feature representation of each target, including multi-dimensional feature information such as appearance, structure, and time series.

[0043] Based on the obtained target feature representation, the system performs spatio-temporal localization. The localization module first executes a region proposal network in the feature space to generate candidate regions that may contain the target. Each candidate region calculates the target score through a feature matching network, and the regions with scores exceeding the threshold are marked as potential targets. The system uses the non-maximum suppression algorithm to merge overlapping candidate regions to ensure the uniqueness of the detection results. A temporal consistency constraint is introduced during the localization process, and the temporal information in the features is used to establish inter-frame target associations. The system designs a spatio-temporal attention module, which learns the spatial and temporal dependencies in the features to improve the accuracy of target localization. For multi-target scenarios, the system distinguishes between targets by calculating the similarity of target features. To handle occlusion situations, the system maintains a target state cache to record the target positions and feature information in historical frames. When a target is detected to be occluded, historical information is used for trajectory prediction and identity maintenance. Through the above localization process, the system finally obtains the accurate spatial position and motion trajectory of the target.

[0044] In some embodiments, the spatio-temporal positioning based on the target feature to generate an attention map includes: dividing the visual scene image into grid regions, extracting color distribution, texture gradient, and shape contour features within the grid regions to generate a target movement trajectory; calculating the distance weight between the grid region and the scene center, and fusing feature saliency and spatial weight to generate an initial attention distribution; establishing a spatio-temporal correlation model according to the target movement trajectory, dynamically updating the attention value corresponding to the grid region, and generating the attention map through gradient optimization.

[0045] Using the obtained target spatial position and movement trajectory, the system generates an attention map. The attention calculation module adopts a multi-channel architecture, including bottom-up and top-down processing streams. The bottom-up processing stream generates an initial saliency map based on low-level visual features such as color contrast, texture complexity, and edge intensity. The top-down processing stream converts the spatial position information of the target into a two-dimensional Gaussian distribution, with the center of the distribution corresponding to the center position of the target; at the same time, using the movement trajectory information, an attention weight is increased at the predicted position of the trajectory, and the weight size is proportional to the movement speed. The outputs of the two processing streams are fused at multiple scales through a feature pyramid network, and the feature maps at different scales are fused layer by layer through lateral connections and upsampling operations. The system performs spatial normalization and channel normalization on the fused feature map to ensure the numerical stability of the attention distribution. The calculation process of the attention map takes into account the spatial distribution and temporal changes of the target, and assigns a higher attention weight to the region where the moving target is located. The system also implements an attention suppression mechanism. When multiple targets are detected, the attention intensity of each target region is adjusted through competitive inhibition to avoid excessive attention dispersion. For the background region, the system assigns an appropriate attention weight according to its spatial relationship and feature similarity with the target. The finally generated attention map has the same resolution as the input image, and each pixel value represents the importance degree of the corresponding position, with the value range between 0 and 1. This attention map not only identifies the important regions in the current frame but also reflects the dynamic change characteristics of the scene through the integration of temporal information.

[0046] Process the aligned data stream obtained in the above steps. The data stream contains 8-channel EEG signals (FP1, FP2, F3, F4, C3, C4, O1, O2) that have been time-aligned and corresponding timestamp information, with a sampling frequency of 1000 Hz. The system first segments the EEG data into fixed windows according to the timestamps. The window length is set to 1 second (1000 sampling points), and the window overlap rate is 50% to ensure the continuity of data analysis. Perform wavelet decomposition on each data window, using the discrete wavelet transform to decompose the signal into 5 scale levels to obtain the time-domain representation of different frequency bands. These data contain the EEG activities of the user in daily interaction scenarios such as observing the display, operating the mouse, and making menu selections. The Daubechies4 wavelet basis function is used in the decomposition process, which has good time-frequency localization characteristics. The system performs threshold processing on the decomposed coefficients to remove the influence of common artifacts such as blinks and EMG. At the same time, calculate the energy distribution and phase information of each scale level to construct a time-domain feature vector. To maintain the temporal correspondence with the data stream in the above steps, the system retains the timestamp information of the original data during the feature extraction process. To capture the temporal correlation between channels, the system calculates the cross-correlation coefficient and phase synchronization index between channel pairs. For example, when the user performs a spatial attention shift, significant changes in phase synchronization can be observed between the left and right hemispheres of the occipital lobe region. Through time-domain decomposition, the system finally obtains a set of feature sequences reflecting the time-domain characteristics of the EEG signal, including information such as energy, phase, and inter-channel correlation, and maintains the temporal correspondence with the synchronized data stream in the above steps.

[0047] Based on the above-obtained time-domain feature sequence, the system performs spectral conversion. The system uses the energy distribution in the time-domain features to guide the key areas of frequency band analysis, uses phase information to assist in determining the window parameters of time-frequency analysis, and optimizes the spectral coherence analysis based on the inter-channel correlation features. On this basis, the conversion module adopts an improved short-time Fourier transform method to perform time-frequency analysis on the feature sequence of each channel. The Hann window function is used in the transformation process, with a window length of 256 points and an overlap rate of 75%, ensuring the balance of time-frequency resolution. The system calculates the power spectral density of the θ band (4 - 8 Hz), α band (8 - 13 Hz), and β band (13 - 30 Hz). In practical applications, these frequency bands are closely related to different cognitive states. For example, when the user is browsing the web, the α wave will show a characteristic modulation pattern; when performing fine operations (such as aiming at a small target), the energy of the β band will increase significantly; during memory tasks, the θ wave activity will increase. To improve the reliability of spectral estimation, the system adopts a multi-taper spectral analysis method to fuse the spectral estimation results under different window functions. For each frequency band, characteristic parameters such as relative power and central frequency are calculated. The system also analyzes the spectral coherence between different brain regions and constructs a frequency-domain coupling matrix. For example, during a reading comprehension task, significant β-band coherence will be shown between the prefrontal and temporal regions. The significant spectral components are identified through an adaptive threshold method to suppress the influence of background noise. The spectral features of all channels are temporally aligned to form a unified spectral feature representation. The system normalizes the spectral features to eliminate the influence of individual differences and baseline drift. Through spectral conversion, the system obtains a set of feature matrices describing the frequency-domain characteristics of EEG activities, including the energy distribution, coherence, and time-varying characteristics of each frequency band.

[0048] Using the obtained spectral feature matrix, the system conducts the recognition of the attention state. The recognition module first extracts attention-related feature patterns such as alpha wave suppression and beta wave enhancement. The system analyzes the spectral activity differences in the prefrontal lobe (FP1, FP2), motor area (C3, C4), and occipital lobe area (O1, O2), and calculates the hemispheric asymmetry index. Through the adaptive feature selection algorithm, the most discriminative spectral feature combination is determined. The system adopts a hierarchical attention state classifier, including two levels: spatial attention and task attention. The spatial attention classifier, based on the alpha wave modulation feature in the occipital lobe area, identifies the direction and intensity of visual attention. This is particularly important in scenarios where the user needs to switch attention between multiple targets, such as selecting an operation target in a complex graphical interface. The task attention classifier uses the beta wave features in the prefrontal lobe and motor area to evaluate the cognitive load and intention intensity. The outputs of the two classifiers are fused through dynamic weighting to generate a comprehensive attention state assessment result. The system also establishes a temporal model of the attention state, and predicts the change trend of the attention level by analyzing the state transition sequence. For the judgment result of the attention state, the system calculates a confidence index, and triggers the state re-evaluation mechanism when the confidence is lower than the threshold. Through the attention state recognition, the system outputs a state description vector including the attention direction, intensity, and stability.

[0049] In some embodiments, the obtaining of the spectral features for identifying the intention features of the attention state output based on the temporal features includes: after band-pass filtering the electroencephalogram signal, extracting the band energy features through short-time Fourier transform; calculating the band energy asymmetry index of the electrodes in the frontal and occipital regions according to the band energy features, and constructing a multi-dimensional feature vector in combination with the preset motor imagery features; using a support vector machine classifier to identify three types of intention features: target selection, operation type, and execution intensity.

[0050] Based on the obtained attention state description vector, the system generates intention features. The intention feature extraction module adopts a multi-level feature mapping architecture to convert attention state information into specific behavioral intention representations. The system first establishes an attention-intention mapping model, which learns the correspondence between different attention patterns and user intentions. For spatial attention features, the system decodes the intention information of target selection and spatial positioning; for task attention features, they are used to infer the operation type and execution intensity. The system implements an intention state machine, which includes four basic states: idle, ready, executing, and completed, and predicts the user's behavior sequence according to the temporal changes of the attention state. To improve the robustness of intention recognition, the system introduces a context constraint mechanism to correct the current intention using historical intention information and task rules. For multi-target scenarios, the system designs an attention competition mechanism to determine the final operation target by evaluating the attention weights of different targets. The intention feature generation process also considers the continuity of tasks and the user's operation habits, and establishes an intention pattern library to store common operation sequences. The system also establishes an intention credibility evaluation module, which comprehensively considers the stability of the attention state and the consistency of intention inference, and assigns a credibility score to each intention feature. When detecting uncertain or conflicting intentions, the system activates the intention clarification mechanism to determine the final intention feature through cross-verification of multi-modal information. The finally output intention feature is a multi-dimensional vector, which includes information such as target selection, operation type, and execution parameters.

[0051] Step S103, construct a feature matrix according to the attention map and intention features, perform modal alignment on the feature matrix, and perform feature fusion based on the alignment result of the modal alignment to generate a unified representation.

[0052] Specifically, combining the attention map obtained in the above steps and the intention features obtained in the above steps, the system constructs a multi-modal feature matrix. The attention map contains the spatial distribution and importance information of the targets in the scene, and the value range of each pixel is between 0 and 1; the intention features contain high-level semantic information such as target selection, operation type, and execution parameters. The system first divides the attention map into grid regions, and calculates the average attention value and gradient features for each region. For the intention features, the system extracts the spatial pointing information and behavior type encoding therein. The two types of features are organized into an initial feature matrix through tensor concatenation. The rows of the matrix represent time sampling points, and the columns represent different feature dimensions. At the same time, a feature index table is established to record the source and physical meaning of each feature dimension, providing a basis for subsequent feature selection. Through the construction of the feature matrix, the system obtains a unified data structure, which contains the key information of both the visual and intention modalities.

[0053] In some embodiments, performing feature fusion based on the alignment result of the modality alignment to generate a unified representation includes: performing tensor concatenation on the grid region features of the attention map and the semantic encoding of the intent features; calculating the inter-modal feature correlations of the concatenated tensors through a self-attention mechanism, including spatially position-related features and temporally related features; enhancing the local features of the spatially position-related features using a spatial attention mechanism and modeling the state transition of the temporally related features using a recurrent neural network; dynamically adjusting the fusion weights of the spatially position-related features and the temporally related features according to the environmental light intensity and the quality of the electroencephalogram signals, so as to generate a unified representation with hierarchical encoding through principal component analysis dimensionality reduction.

[0054] Based on the constructed feature matrix and feature index table, the system performs a modality alignment operation. The system uses the feature source information recorded in the feature index table to determine the modality type, sampling characteristics, and physical meaning of each feature dimension. For the feature dimensions related to the attention map (such as the attention values and gradient features of the grid regions), the index table records their 30Hz sampling rate and spatial attributes; for the dimensions of the intent features (such as target selection and behavior type encoding), the index table identifies their 10Hz update frequency and semantic attributes. Based on this index information, the system separately performs temporal resampling on the features of different modalities to unify the data of the two modalities to the same time scale. Taking the mouse dragging operation as an example, the system identifies the 33ms sampling period of the visual attention features and the 100ms update period of the intent features according to the index table, and then interpolates the intent features to generate a sampling sequence with the same frequency as the visual features. The piecewise linear interpolation method is used to handle the sampling rate difference, and the sampling density is increased for rapidly changing segments. For example, when quickly swiping the page, the system will increase the sampling rate to capture the rapidly changing attention distribution. At the semantic level, the system establishes a correspondence between the spatial regions in the attention map features and the target selection information in the intent features according to the physical meaning labels in the feature index table. When the user switches windows in a multi-document interface, the system uses the index table to quickly locate the relevant spatial features and intent features to ensure that the spatial migration of the attention focus can accurately correspond to the task switching intent. The alignment process takes into account the signal transmission delay and processing delay, and compensates for the delay differences of different modalities through a time window mechanism. For example, there is usually a physiological delay of 200 - 300ms from when the user gazes at a target to generating a clear operation intention, and the system will compensate for this time difference during feature alignment. The system also implements dynamic clock bias correction to ensure the synchronization between modalities during long-term operation. Through modality alignment, the system obtains a feature representation that is consistent in both time sequence and semantics as the unified representation.

[0055] Using the aligned feature representations, the system performs multimodal feature fusion. The fusion module adopts a multi-level architecture design, including two stages: feature-level fusion and decision-level fusion. In feature-level fusion, the system uses the self-attention mechanism to calculate the intra-modal and inter-modal feature correlations. Taking the intelligent office scenario as an example, when the user gazes at the document on the screen, the visual attention will focus on the text area, and at the same time, the EEG intention feature indicates that the user has a reading intention, and the system will enhance the weights of these features. In the video conferencing scenario, when the user needs to share content, the focus of visual attention on the toolbar area will form a strong correlation with the operation intention feature. For features related to spatial location, a spatial attention mechanism is used to enhance local features. In the multi-window switching task, the system can accurately capture the process of the user's gaze point transferring from one application window to another and correspond it to the task switching signal in the intention feature. For features related to time series, a recurrent neural network is used to model the time series dependencies to accurately identify the user's behavior patterns in continuous tasks such as file browsing and content editing. In the decision-level fusion stage, the system dynamically adjusts the weights of different features according to their reliabilities. For example, when the light condition deteriorates, the weight of the visual feature will be correspondingly reduced; when the user is in a highly concentrated state, the weight of the EEG feature will be increased. The fusion process also considers the task context and adopts specific fusion strategies for different types of interaction tasks such as document processing, image editing, and data analysis. Through feature fusion, the system obtains a feature vector that comprehensively represents the information of the two modalities.

[0056] Based on the fused feature vector, the system generates a unified multimodal representation. The representation generation module first performs dimensionality reduction on the feature vector and uses the principal component analysis method to extract the most discriminative feature combination. The system designs a hierarchical coding structure. The bottom layer encodes the spatial location and attention distribution, the middle layer encodes the behavior type and execution parameters, and the top layer encodes the task objective and interaction intention. At each level, the system establishes a feature dictionary to quantize the continuous feature space into discrete representation symbols. To improve the robustness of the representation, the system implements multi-scale feature aggregation, which can interpret the user's intention at different granularity levels. At the same time, a semantic mapping network is established to associate the low-level perceptual features with the high-level cognitive features. The finally generated unified representation is a hierarchical symbolic structure that not only retains the detailed information of the original features but also provides a clear semantic interpretation.

[0057] In step S104, perform temporal segmentation on the unified representation, extract associated features, construct a state sequence based on the associated features, use the state sequence to label transition nodes, and generate a dynamic pattern according to the transition nodes; perform information analysis using the dynamic pattern to determine the modality weights, perform feature selection based on the modality weights, and perform classification mapping on the feature selection result to generate a control sequence.

[0058] Specifically, perform temporal segmentation on the unified representation obtained in the above steps to extract associated features. The system receives the hierarchical symbolic structure in the unified representation, which contains multi-level information such as spatial position, attention distribution, behavior type, and task objective. The system adopts a multi-scale sliding window strategy, setting three window sizes: 100 ms for capturing instantaneous changes, 500 ms for behavior segment analysis, and 2000 ms for task-level changes. Within each window, the system first extracts the time-varying features in the representation vector, including the change rate of attention distribution, the fluctuation of intention intensity, and the degree of cooperation between modalities. For rapidly changing segments, such as rapid attention transfer or intention mutation, the system increases the sampling density to ensure the accuracy of feature extraction. At the same time, calculate the statistical features of the window sequence, including mean, variance, kurtosis, etc., to characterize the stability and mutation degree of the behavior. The system also constructs an association graph between features, describing the temporal dependence relationship between different features through the graph structure. In practical applications, when the user performs a continuous task, the system can accurately capture the critical moments of the co-variation of attention-intention. For example, in a scenario of alternating reading and editing, the system can extract the close association between the change of attention focus and the enhancement of editing intention from the unified representation. During the segmentation process, the system adaptively adjusts the window parameters and finely segments the detected critical event segments. Through temporal segmentation, a set of feature sequences with temporal associations is finally obtained, and these sequences retain the temporal structure of the behavior segments and the feature association intensity.

[0059] In some embodiments, constructing a state sequence based on the associated features, using the state sequence to label transition nodes, and generating a dynamic pattern according to the transition nodes includes: extracting mean, variance, and transition probability statistics of the associated features after temporal segmentation according to time windows, for constructing a state sequence including spatial distribution, operation type, and execution intensity; detecting the feature jump amplitude of the state sequence, and marking it as a transition node when the feature jump amplitude exceeds a set threshold; generating a dynamic pattern according to the node interval duration and state transition path of the transition node, and predicting the optimal state transition sequence using a hidden Markov model to generate the dynamic pattern.

[0060] Based on the acquired feature sequence and its temporal correlation information, the system constructs a state sequence. The state construction module first analyzes the temporal pattern of the feature sequence and defines a basic state space, including core states such as preparation, execution, completion, and transition. For different interaction patterns, the system establishes corresponding state subsets. To accurately map the feature sequence to the state space, the system adopts a hierarchical coding mechanism to convert continuous feature change patterns into discrete state representations. The state construction process uses a hierarchical approach, first performing a coarse-grained mapping in the basic state space and then refining the state definition according to the specific task context. The system implements a state estimator to determine the current state by evaluating the temporal pattern of the feature sequence. The state transition probability is dynamically updated based on the statistical characteristics of the feature sequence to ensure the accuracy of state recognition. The system also establishes a feature-state mapping dictionary that records the typical state patterns corresponding to different feature combinations. By deeply analyzing the temporal structure of the feature sequence, the system generates a state sequence that reflects the evolution of user behavior.

[0061] Using the constructed state sequence and the behavior evolution information it contains, the system marks the state transition nodes. The transition detection module monitors the mutation points at the feature level and identifies the transition boundaries at the state level. The system defines transition feature metrics, including feature gradient, state probability change rate, and modal consistency measure. When these metrics exceed the adaptive threshold, the moment is marked as a potential transition node. For each transition node, the system calculates its significance score, which is based on factors such as the magnitude, duration, and context relevance of the state change. The system designs a state stability evaluation mechanism, and only when the state change persists for a certain time and the change magnitude is large enough will it be marked as a valid transition node. In the case of rapid multi-state switching, the system merges adjacent transition nodes through the minimum description length criterion to avoid over-segmentation. The system maintains the hierarchical relationship of the transition nodes, distinguishing between task-level and operation-level transitions. Through the marking of transition nodes, the system obtains a structured sequence that describes the dynamic changes of the state.

[0062] Based on the sequence of labeled transformation nodes and their structured information, the system generates a dynamic pattern. The pattern generation module analyzes the temporal patterns in the sequence of transformation nodes and uses the pattern grammar method to parse them into high-level behavior patterns. The system defines a basic pattern dictionary, which contains common interaction pattern templates, such as linear sequence patterns, loop patterns, conditional branch patterns, etc. For different application scenarios, the system expands specific pattern sets. The pattern recognition process takes into account temporal constraints and semantic constraints, and uses the dynamic programming algorithm to find the optimal pattern match. The system implements a pattern combination mechanism that can construct complex combined patterns from basic patterns. The system establishes a pattern evaluation mechanism to evaluate the stability and predictability of patterns by calculating the occurrence frequency, duration, and transition probability of patterns. For newly emerging behavior patterns, the system expands the pattern library through pattern learning algorithms. During the pattern generation process, the system also considers the generality and specificity of patterns and seeks a balance between them. Through dynamic pattern generation, the system outputs a pattern representation that includes the temporal characteristics of user behavior and the semantic information of the interaction process.

[0063] Using the dynamic pattern obtained by the above steps for information analysis to determine the modal weights. The system first evaluates the temporal patterns and state transition rules in the dynamic pattern and calculates the contribution degree of different modal features in the decision-making process. For the visual modality, the system evaluates the stability of the attention distribution and the accuracy of target tracking, uses the normalized information entropy to measure the degree of attention dispersion, and uses the smoothness of the target trajectory to evaluate the tracking quality; for the electroencephalogram modality, the system analyzes the confidence of intention recognition and the accuracy of state prediction, and evaluates the signal reliability by calculating multi-scale signal quality indicators and band energy ratios. Based on the evaluation results, the system uses an adaptive weight adjustment algorithm to dynamically allocate the weight coefficients of visual and electroencephalogram features. The weight calculation uses a sliding time window mechanism, and the window size is adaptively adjusted according to the task complexity to ensure the real-time and stability of weight updates. The system also establishes a multi-level weight modulation mechanism to optimize the weights at the feature level, modal level, and task level respectively. By introducing temporal smoothing constraints, it avoids the drastic fluctuations of modal weights affecting the system stability. Through information analysis, the system obtains a set of weight vectors reflecting the importance of each modality.

[0064] Based on the obtained modal weight vectors, the system performs feature selection. The selection module first constructs a feature importance scoring matrix, and the matrix elements are jointly determined by the weight coefficients, temporal coherence, and task relevance of the features. For each feature dimension, the system not only calculates its discriminative ability in the current task but also evaluates its stability in a continuous task sequence. The feature selection process adopts an iterative strategy, and in each iteration, the feature subset with the highest score is selected. The system implements an adaptive feature screening mechanism that dynamically adjusts the number of features according to task difficulty and execution efficiency. To avoid local optima, the system uses a genetic algorithm-based feature combination optimization method to search for the optimal feature set through multiple generations of evolution. At the same time, the system implements feature redundancy analysis, calculates the conditional mutual information and temporal correlation between features, and ensures that the selected features are both complementary and avoid information duplication. To ensure real-time performance, the system adopts a hierarchical feature screening structure and uses a multi-level screening strategy from coarse-grained to fine-grained to quickly locate the most valuable feature combination. Through feature selection, the system obtains a set of concise and informative feature subsets.

[0065] Classify and map the selected feature subset. The mapping module designs a multi-level classification structure, including three levels: intention level, behavior level, and operation level. At the intention level, the system maps the features to the user's interaction intention space and establishes an intention inference framework based on a probabilistic graphical model; at the behavior level, the system identifies specific operation types and execution methods and uses a hierarchical state machine to describe the behavior conversion rules; at the operation level, the system determines precise control parameters and predicts continuous control quantities through a regression model. For example, in an intelligent office scenario, when the system detects that the user is gazing at a document area and there is a continuous increase in beta waves, the intention level maps it to a "deep reading intention", the behavior level identifies it as a "text selection operation", and the operation level precisely locates the specific text range and selection method. The classification process adopts a hierarchical decision tree structure, and each node is equipped with a scoring function and decision threshold for feature combinations. The system establishes an adaptive mapping rule library that can dynamically update the rule parameters according to user feedback. The mapping process also implements a context awareness mechanism that tracks the operation sequence by maintaining a task state stack to improve the context relevance of classification. Through classification mapping, the system generates a standardized control description.

[0066] Step S105, generate an interaction instruction according to the control sequence to complete the visual-electroencephalogram collaborative control of the multi-modal AI glasses.

[0067] Specifically, according to the control description obtained by mapping, the system generates a control sequence. The control generation module first parses the control description into an ordered combination of basic operation units, and each operation unit includes attributes such as action type, execution parameters, timing constraints, and priority. The system designs a multi-level control template library, including basic action templates, combined operation templates, and task flow templates, to support the generation of control sequences with different granularities. For complex control tasks, the system divides them into multiple subtasks through a task decomposition algorithm and generates corresponding control segments for each subtask. The generation process of the control sequence considers the continuity and smoothness of operations. The system implements a trajectory planning algorithm based on spline interpolation to ensure smooth transitions between adjacent operations. In addition, the system also establishes a control constraint checking mechanism, including the verification of physical constraints, safety constraints, and task constraints, to ensure that the generated control sequence meets various execution requirements. To improve the robustness of control, the system implements an exception handling and failure recovery strategy, which can quickly adjust the control scheme when an execution exception occurs. Through the generation of the control sequence, the system finally outputs an instruction sequence containing complete control information.

[0068] In some embodiments, generating interaction instructions according to the control sequence includes: performing delay analysis on the control sequence, marking processing nodes, dividing execution units based on the processing nodes, constructing a pipeline strategy using the execution units, and generating response instructions; performing behavior analysis based on the control sequence and the response instructions, extracting features from the analysis results of the behavior analysis, using the extracted features to determine control rules, generating an interaction scheme and corresponding interaction instructions, and completing the visual-electroencephalogram collaborative control of the multi-modal AI glasses.

[0069] Perform latency analysis on the control sequence obtained in the above steps (an instruction stream containing target selection, operation type, and execution parameters), and mark the processing nodes. The system first establishes a latency evaluation model for the control sequence, which includes multiple links such as signal acquisition delay, feature processing delay, decision-making delay, and execution delay. Taking the intelligent interaction scenario as an example, when the user performs rapid and continuous document page-turning operations, the system needs to complete the entire process from attention detection to page switching within 100 ms, which requires accurate evaluation of the latency of each processing link. For each control instruction, the system calculates the cumulative latency and critical path on its processing chain. The latency analysis adopts an event-based evaluation method to refine the time overhead of each processing link through task decomposition. The system implements an adaptive time window mechanism to dynamically adjust the processing time limit according to the instruction complexity. For example, for fine-grained image editing operations, the system will expand the time window to ensure processing accuracy; while for rapid page scrolling, the window will be narrowed to improve the response speed. For high-frequency operation sequences, the system adopts a prediction and compensation strategy to plan the processing resources in advance. For each critical time point, the system evaluates the processing load and resource status, and marks the processing nodes that need to be optimized with emphasis. Through latency analysis, the system obtains a sequence of processing nodes with timing constraints, and each node contains the maximum allowable processing time, critical path information, and resource requirement evaluation.

[0070] Based on the sequence of processing nodes with timing constraints, the system performs execution unit partitioning. The partitioning module first uses the maximum allowable processing time of each node as the upper limit of the time budget to ensure that the divided execution units do not exceed the timing requirements. The system constructs a task dependency graph based on the critical path information, and optimizes the task allocation by analyzing the data flow and control flow dependency relationships between the nodes on the critical path. At the same time, using the resource requirement evaluation results of the nodes, it reasonably allocates computing resources such as CPU and memory. The system adopts a task clustering algorithm to group the nodes according to time correlation, resource sharing degree, and processing time limit. For each execution unit, the system defines strict input and output interfaces and processing time limit requirements, and the time limit is allocated according to the cumulative delay in the node sequence. The execution unit partitioning process takes into account the load balancing of the processing modules, and reasonably allocates tasks by evaluating the computational complexity and data traffic to ensure that the processing time of each unit is within the preset timing constraints. The system also establishes a communication mechanism between the units, stipulates the data exchange format and synchronization method, and the communication overhead is included in the total processing latency. In particular, for operations with strict real-time requirements, the system sets up an independent fast channel according to the timing constraints to ensure the timely response of critical instructions. Through execution unit partitioning, the system forms a set of processing modules with independent functions, standardized interfaces, and meeting the timing requirements, and each module is equipped with a clear processing time limit and resource quota.

[0071] Using the partitioned processing modules and their timing information, the system constructs a pipeline strategy. The pipeline design adopts a multi-level structure, including a data preprocessing level, a feature extraction level, a decision control level, and an instruction generation level. In practical applications, for example, when a user inputs through an eye movement controlled virtual keyboard, the preprocessing level is responsible for filtering and enhancing eye movement signals and electroencephalogram signals, the feature extraction level calculates the fixation point position and intention intensity, the decision control level determines the key selection, and the instruction generation level outputs specific keystroke commands. The system allocates appropriate processing time slices for each level of the pipeline according to the timing constraints of each processing module, and designs a buffer size that matches the processing time limit. Each level of the pipeline is equipped with an independent buffer and a status management mechanism, and the buffer size is dynamically adjusted according to the processing delay and data throughput of the module. The system implements a dynamic scheduling algorithm considering timing constraints, and flexibly adjusts the processing order according to task priorities, resource status, and remaining processing time. For different types of interaction tasks, the pipeline can dynamically change the parallelism and processing granularity while ensuring timing constraints. For example, in continuous operations, the system uses fine-grained pipelining to improve throughput; in bursty tasks, it switches to a coarse-grained mode to reduce scheduling overhead. The pipeline also includes an exception handling mechanism that can detect and handle various abnormal situations, such as signal quality mutations, processing timeouts, etc., to ensure a quick recovery to the normal processing flow in case of interference. Through pipeline construction, the system establishes a multi-level pipelined parallel processing framework, and each processing module strictly follows the preset timing constraints.

[0072] According to the constructed multi-level pipelined processing framework, the system generates response instructions. The instruction generation module makes full use of the parallel processing capabilities at all levels of the pipeline, implements multi-level instruction templates, and supports control requirements with different precisions and complexities. First, the system integrates the processing results at all levels of the pipeline according to the time sequence and converts them into a standard instruction format, including an opcode, parameters, and time sequence tags. In complex interaction scenarios, such as when a user controls smart home devices through eye gaze and EEG signals, the system needs to coordinate the processing rhythms at all levels of the pipeline to generate a complete instruction chain that includes device selection, function settings, and execution time sequences. The instruction generation process inherits the real-time processing characteristics of the pipeline and adjusts the execution efficiency of the instruction sequence through an optimization algorithm. For example, for continuous brightness adjustment operations, the system combines multiple fine adjustments into a smooth gradient instruction within the cumulative delay range of the pipeline. At the execution level, the system implements a multi-level cache structure that matches the pipeline processing rhythm, including an immediate response cache, a speculative execution cache, and a historical pattern cache, which are used to handle control requirements at different time scales. At the same time, the system establishes a complete instruction verification mechanism, including syntax checking, parameter range verification, and execution condition evaluation, to ensure that the generated instruction sequence meets the requirements of safety and executability. In particular, for critical control instructions, the system implements a dual verification mechanism, and a secondary confirmation is performed through an independent security monitoring module. The verification process is incorporated into the processing cycle of the pipeline. The instruction generation module also includes an adaptive adjustment function that can dynamically optimize instruction parameters based on the execution feedback at all levels of the pipeline. For example, when a change in the user's reaction time for a specific operation is detected, the system adjusts the execution delay and transition time of the instruction accordingly within the pipeline processing range. Through response instruction generation, the system finally outputs a control instruction stream with a rigorous time sequence and high execution efficiency.

[0073] Combine and analyze the control sequence of the above steps and the response instructions of the above steps. The system uses the target selection and operation parameters in the control sequence to guide the analysis direction, and at the same time uses the processing strategy and timing information in the response instructions to evaluate the execution effect. The analysis module constructs a two-stream analysis architecture. One branch tracks the operation process specified in the control sequence, and the other branch monitors the actual execution situation according to the response instructions. Taking the intelligent office environment as an example, when the user performs a document editing task, the system compares the target location and operation type specified in the control sequence with the actual execution trajectory given by the response instructions to evaluate the matching degree of control and response. The comparison process uses a multi-scale time window to conduct a detailed evaluation from the millisecond-level instant response to the task completion at the second level. The system realizes the consistency check based on timing, and quantifies the accuracy of the interaction by calculating the deviation between the control intention and the actual response. In practical applications, the system can identify the execution characteristics of the user at different operation stages, such as key indicators like instruction delay, response speed, and completion accuracy. For the detected abnormal execution situations, such as the situation where the response instructions deviate from the control sequence expectation, the system will record detailed context information. Through the combined analysis, the system obtains a set of behavior indicators including timing characteristics, execution accuracy, and response delay.

[0074] Feature extraction is performed on this set of behavioral indicators reflecting the control-response relationship. The feature extraction module first processes three types of key indicators separately: calculating the instruction interval distribution and rhythm characteristics using temporal features, extracting operation accuracy and stability indicators from execution accuracy, and analyzing the system response characteristics and user adaptability based on response latency. On this basis, the system designs a hierarchical feature description system, including low-level operation features, middle-level behavior patterns, and high-level interaction strategies. At the operation feature level, the system combines temporal features and response latency to extract the continuity index of instruction execution; analyzes the accuracy of spatial positioning and operation control through execution accuracy data analysis. At the behavior pattern level, based on temporal features, the task switching frequency and pattern are identified, the attention transfer pattern is detected using execution accuracy data, and the change trend of intention intensity is analyzed through response latency. At the interaction strategy level, the operation rhythm reflected by temporal features, the control ability reflected by execution accuracy, and the adaptation degree characterized by response latency are comprehensively analyzed to summarize the user's decision-making tendency and adaptation characteristics. For example, the system can discover the instruction rhythm of the user during graphic interface operations through temporal features, evaluate the operation proficiency in combination with execution accuracy, and analyze the user's adaptation to system feedback through response latency. The feature extraction process adopts an adaptive computing framework, dynamically adjusting the weights of the three types of indicators according to the characteristics of different scenarios. In fast interaction scenarios, the system increases the weight of temporal features to optimize rhythm control; in fine operation tasks, the weight of execution accuracy is increased to ensure operation quality; in the new task adaptation stage, response latency is mainly analyzed to evaluate the learning effect. The system implements a real-time update mechanism for features, continuously monitoring the change trends of the three types of indicators and timely adjusting the feature extraction strategy. For each type of feature, the system establishes a reliability evaluation standard by analyzing the signal-to-noise ratio and stability of the corresponding indicators to ensure the accuracy of the extraction results. Through this multi-level feature extraction process, the system finally obtains a set of feature vectors reflecting the user's operation rules, which comprehensively expresses the user's operation characteristics in three dimensions: temporal control, accuracy control, and response adaptation.

[0075] Based on this set of eigenvectors reflecting the operation laws, the system determines the control rules. The rule generation module first constructs a feature-rule mapping network according to the distribution characteristics of the eigenvectors, and divides the feature space into rule domains corresponding to different control strategies. For each rule domain, the system defines corresponding control parameters and adjustment strategies according to the performance of the eigenvectors. In virtual environment interaction, when the eigenvectors indicate that the user tends to browse quickly, the system enables the large-range navigation rule; when the features show precise operation intentions, the fine control rule is activated. The rule determination process takes into account multiple performance indicators reflected by the eigenvectors, including response speed, operation accuracy, and user experience. The system uses a fuzzy inference method to handle the uncertainty of the feature values, and calculates the activation intensity of different rules through the membership function. For complex interaction scenarios, the system implements a rule combination mechanism, which can dynamically select a set of rules according to the changes of the eigenvectors. For example, in a multi-document processing task, the system can combine the window switching rule and the content editing rule according to the indication of the eigenvectors to achieve a smooth task transition. The rule base has the ability to learn, and can optimize the rule parameters according to the long-term change trend of the eigenvectors. In particular, the system can identify the user's skill improvement process from the feature performance and adjust the intervention degree of the auxiliary rules in a timely manner. Through rule determination, the system establishes a set of control rule systems guided by eigenvectors.

[0076] Exemplarily, the generation of the interaction scheme and the corresponding interaction instructions includes: determining a control rule according to the extracted features to generate the interaction scheme; performing multi-modal detection on the interaction scheme, marking the signal quality, determining a response mode based on the signal quality, performing collaborative processing using the response mode, and outputting the interaction instructions.

[0077] According to the established control rule system, the system generates specific interaction instructions. The scheme generation module adopts a hierarchical design, transforming the parameters and strategies defined in the control rules into executable interaction steps. At the strategy level, the system maps the rules into different interaction templates, including basic gaze selection, intention confirmation, and combined operation templates. For the application scenario of the multimodal AI glasses, the system forms an adapted interaction instruction set based on the rule requirements. For example, in the reading - editing scenario, the system transforms the gaze control rule and the intention confirmation rule into a text selection scheme. The user determines the target position by continuously gazing, and the system performs the selection operation after detecting the enhanced stable beta waves. During the scheme generation process, the system dynamically adjusts the interaction parameters according to the feedback requirements of the rules. For example, it reduces the intention judgment threshold when the attention switches frequently and improves the spatial positioning accuracy during fine operations. For complex interaction tasks, the system generates alternative schemes according to the rules. For example, during text editing, it can choose between "gaze positioning + intention confirmation" or "gaze trajectory circle selection", and selects a more suitable scheme according to the electroencephalogram signal characteristics during the execution process. Each interaction instruction includes the trigger conditions, operation steps, and feedback mechanism defined by the rules, and adaptively adjusts the presentation mode of the auxiliary information according to the user's proficiency. Through this rule - driven scheme generation process, the system finally outputs a complete set of interaction instructions.

[0078] To execute the multimodal AI glasses vision - electroencephalogram collaborative control method corresponding to the above - mentioned method embodiments to achieve the corresponding functions and technical effects. Refer to Figure 2 , Figure 2 FIG. shows a structural block diagram of a multimodal AI glasses vision - electroencephalogram collaborative control device 200 provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to this embodiment are shown. The multimodal AI glasses vision - electroencephalogram collaborative control device 200 provided by the embodiment of the present application includes:

[0079] A data generation module 201, configured to construct a multimodal data acquisition framework for acquiring visual scene images and electroencephalogram signals, perform timestamp marking based on the acquired visual scene images and electroencephalogram signals, and generate an aligned data stream;

[0080] A scene analysis module 202, configured to perform scene analysis on the visual part in the aligned data stream to identify target features, perform spatio - temporal positioning based on the target features to generate an attention map; perform time - domain decomposition on the electroencephalogram part in the aligned data stream to extract time - series features, perform spectral conversion based on the time - series features to obtain spectral features for identifying the intention features output by the attention state;

[0081] The matrix construction module 203 is configured to construct a feature matrix based on the attention map and the intention feature, perform modality alignment on the feature matrix, and execute feature fusion based on the alignment result of the modality alignment to generate a unified representation;

[0082] The time series segmentation module 204 is configured to perform time series segmentation on the unified representation, extract associated features, construct a state sequence based on the associated features, label conversion nodes using the state sequence, and generate a dynamic pattern according to the conversion nodes; perform information analysis using the dynamic pattern to determine modality weights, execute feature selection based on the modality weights, and perform classification mapping on the feature selection result to generate a control sequence;

[0083] The instruction generation module 205 is configured to generate an interaction instruction according to the control sequence to complete the visual - electroencephalogram collaborative control of the multi - modal AI glasses.

[0084] The above - mentioned multi - modal AI glasses visual - electroencephalogram collaborative control device 200 can implement the multi - modal AI glasses visual - electroencephalogram collaborative control method in the above - mentioned method embodiment. The optional items in the above - mentioned method embodiment are also applicable to this embodiment, which will not be elaborated here. The remaining content of the embodiment of the present application can refer to the content of the above - mentioned method embodiment and will not be repeated in this embodiment.

[0085] Figure 3 It is a schematic structural diagram of a computer device provided in an embodiment of the present application. As Figure 3 shown, the computer device 3 in this embodiment includes: at least one processor 30 ( Figure 3 only one is shown here), a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30. When the processor 30 executes the computer program 32, the steps in any of the above - mentioned method embodiments are implemented.

[0086] The computer device 3 can be a computing device such as a smart phone, a tablet computer, a desktop computer, and a cloud server. The computer device may include but is not limited to the processor 30 and the memory 31. Those skilled in the art can understand that Figure 3 this is only an example of the computer device 3 and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input - output devices, network access devices, etc.

[0087] The so-called processor 30 may be a Central Processing Unit (CPU), and the processor 30 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0088] In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In other embodiments, the memory 31 may also be an external storage device of the computer device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 3. Further, the memory 31 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 31 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 31 may also be used to temporarily store data that has been output or is to be output.

[0089] In addition, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0090] An embodiment of the present application provides a computer program product, and when the computer program product runs on a computer device, the computer device is caused to implement the steps in each of the above method embodiments when executed.

[0091] In several embodiments provided by the present application, it can be understood that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved.

[0092] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0093] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A visual - electroencephalogram collaborative control method for a multi - modal AI glasses, characterized in that, Including: Construct a multimodal data acquisition framework for collecting visual scene images and electroencephalogram (EEG) signals, perform timestamp marking based on the collected visual scene images and EEG signals, and generate an aligned data stream; Perform scene analysis on the visual part of the aligned data stream to identify target features, perform spatio-temporal localization based on the target features, and generate an attention map, including: divide the visual scene image into grid regions, extract color distribution, texture gradient, and shape contour features within the grid regions to generate a target movement trajectory; calculate the distance weight between the grid region and the scene center, fuse feature saliency and spatial weight to generate an initial attention distribution; establish a spatio-temporal correlation model according to the target movement trajectory, dynamically update the attention value corresponding to the grid region, and generate the attention map through gradient optimization; adopt a multi-channel architecture, including a bottom-up processing stream and a top-down processing stream; the bottom-up processing stream generates an initial saliency map based on color contrast, texture complexity, and edge intensity as low-level visual features; the top-down processing stream converts the spatial position information of the target into a two-dimensional Gaussian distribution, and the distribution center corresponds to the target center position; utilize the motion trajectory information to increase the attention weight at the predicted position of the trajectory, and the weight size of the attention weight is proportional to the motion speed; the outputs of the bottom-up processing stream and the top-down processing stream are fused at multiple scales through a feature pyramid network, and feature maps at different scales are fused layer by layer through lateral connection and upsampling operations; perform spatial normalization and channel normalization on the fused feature maps, and assign a higher attention weight to the region where the moving target is located; when multiple targets are detected, adjust the attention intensity of each target region through competitive inhibition to avoid over-dispersion of attention; for the background region, assign an appropriate attention weight according to the spatial relationship and feature similarity with the target; the generated attention map has the same resolution as the input image, and each pixel value represents the importance degree of the corresponding position, with the value range between 0 and 1; perform time-domain decomposition on the EEG part of the aligned data stream to extract time-series features, and perform spectral conversion based on the time-series features to obtain spectral features for identifying the intention features output by the attention state; Construct a feature matrix according to the attention map and intention features, perform modal alignment on the feature matrix, and perform feature fusion based on the alignment result of the modal alignment to generate a unified representation; Perform time-series segmentation on the unified representation, extract correlation features, construct a state sequence based on the correlation features, use the state sequence to mark conversion nodes, and generate a dynamic pattern according to the conversion nodes; perform information analysis using the dynamic pattern, determine modal weights, perform feature selection based on the modal weights, and perform classification mapping on the feature selection result to generate a control sequence; Generate an interaction instruction according to the control sequence to complete the visual-EEG collaborative control of the multimodal AI glasses.

2. The method according to claim 1, characterized in that, The constructing of the multimodal data acquisition framework for collecting visual scene images and EEG signals includes: Collect visual scene images at a preset rate using a preset visual sensor; Multiple dry electrodes are used to collect electroencephalogram (EEG) signals at a preset frequency. The dry electrodes adopt a nano - coating to reduce contact impedance and are configured with an active shielding ring to suppress electromagnetic interference; Visual data enhancement and EEG signal conditioning are respectively performed to collect the visual scene images and EEG signals.

3. The method according to claim 1, characterized in that, The obtaining of spectral features based on the timing features for identifying the intention features output by the attention state includes: After band - pass filtering the EEG signals, the band - energy features are extracted through short - time Fourier transform; According to the band - energy features, the band - energy asymmetry index between the frontal and occipital electrodes is calculated, and a multi - dimensional feature vector is constructed by combining the preset motor imagery features; A support vector machine classifier is used to identify three types of intention features: target selection, operation type, and execution intensity.

4. The method according to claim 1, wherein The performing of feature fusion based on the alignment result of the modality alignment to generate a unified representation includes: Performing tensor splicing on the grid - region features of the attention map and the semantic encoding of the intention features; Calculating the inter - modality feature correlations of the spliced tensor through a self - attention mechanism, including spatial - position - related features and timing - related features; Using a spatial - attention mechanism to enhance local features for the spatial - position - related features and using a recurrent neural network to model the state transition for the timing - related features; Dynamically adjusting the fusion weights of the spatial - position - related features and the timing - related features according to the environmental light intensity and the EEG signal quality, and generating a hierarchically - encoded unified representation through principal component analysis for dimensionality reduction.

5. The method according to claim 1, characterized in that, The constructing of a state sequence based on the correlation features, using the state sequence to mark transition nodes, and generating a dynamic pattern according to the transition nodes includes: Extracting the mean, variance, and transition - probability statistics of the correlation features after time - series segmentation according to time windows, which are used to construct a state sequence including spatial distribution, operation type, and execution intensity; Detecting the feature jump amplitude of the state sequence, and marking it as a transition node when the feature jump amplitude exceeds a set threshold; Generating a dynamic pattern according to the node - interval duration and the state - transition path of the transition nodes, and using a hidden Markov model to predict the optimal state - transition sequence to generate the dynamic pattern.

6. The method according to claim 1, characterized in that, The generating of interaction instructions according to the control sequence includes: Performing time - delay analysis on the control sequence, marking processing nodes, dividing execution units based on the processing nodes, constructing a pipeline strategy using the execution units to generate response instructions; performing behavior analysis according to the control sequence and the response instructions, and extracting features from the analysis results of the behavior analysis, which are used to determine control rules according to the extracted features, generate interaction instructions and corresponding interaction instructions, and complete the visual - EEG collaborative control of the multi - modal AI glasses.

7. The method according to claim 6, wherein The generating of the interaction instructions and corresponding interaction instructions includes: Determining control rules according to the extracted features to generate the interaction instructions; Performing multi - modal detection on the interaction instructions, marking the signal quality, determining the response method based on the signal quality, and performing collaborative processing using the response method to output the interaction instructions.

8. A visual-EEG collaborative control device for a multi-modal AI glasses, characterized in that, Including: A data generation module, which is used to construct a multimodal data acquisition framework for collecting visual scene images and electroencephalogram (EEG) signals, perform timestamp marking based on the collected visual scene images and EEG signals, and generate an aligned data stream; A scene analysis module, which is used to perform scene analysis on the visual part of the aligned data stream to identify target features, perform spatio-temporal localization based on the target features, and generate an attention map, including: dividing the visual scene image into grid regions, extracting color distribution, texture gradient, and shape contour features within the grid regions to generate a target movement trajectory; calculating the distance weight between the grid region and the scene center, and fusing feature saliency and spatial weight to generate an initial attention distribution; establishing a spatio-temporal association model according to the target movement trajectory, dynamically updating the attention value corresponding to the grid region, and generating the attention map through gradient optimization; adopting a multi-channel architecture, including a bottom-up processing stream and a top-down processing stream; the bottom-up processing stream generates an initial saliency map based on color contrast, texture complexity, and edge intensity as low-level visual features; the top-down processing stream converts the spatial position information of the target into a two-dimensional Gaussian distribution, and the center of the distribution corresponds to the center position of the target; using the motion trajectory information, increasing the attention weight at the predicted position of the trajectory, and the weight size of the attention weight is proportional to the motion speed; the outputs of the bottom-up processing stream and the top-down processing stream are fused at multiple scales through a feature pyramid network, and feature maps at different scales are fused layer by layer through lateral connection and upsampling operations; performing spatial normalization and channel normalization on the fused feature maps, and assigning a higher attention weight to the region where the moving target is located; when multiple targets are detected, adjusting the attention intensity of each target region through competitive inhibition to avoid over-dispersion of attention; for the background region, assigning appropriate attention weights according to the spatial relationship and feature similarity with the target; the generated attention map has the same resolution as the input image, and each pixel value represents the importance degree of the corresponding position, with the value range between 0 and 1; performing time-domain decomposition on the EEG part of the aligned data stream to extract time-series features, and performing spectrum conversion based on the time-series features to obtain spectrum features for identifying the intention features output by the attention state; A matrix construction module, which is used to construct a feature matrix according to the attention map and intention features, perform modal alignment on the feature matrix, and perform feature fusion based on the alignment result of the modal alignment to generate a unified representation; A time-series segmentation module, which is used to perform time-series segmentation on the unified representation, extract correlation features, construct a state sequence based on the correlation features, mark conversion nodes using the state sequence, and generate a dynamic pattern according to the conversion nodes; performing information analysis using the dynamic pattern, determining modal weights, performing feature selection based on the modal weights, and performing classification mapping on the feature selection result to generate a control sequence; An instruction generation module, which is used to generate an interaction instruction according to the control sequence to complete the visual-EEG collaborative control of the multimodal AI glasses.

9. A computer device, characterized in that, It includes a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program and implementing the method according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Multi-modal interaction method and system based on holographic equipment

    CN116880701A

Cited By

  • AI glasses cooperative control system and method based on offline AI small host

    CN122219218A