Visual identification method and system based on electroencephalogram signal enhancement
The EEG-enhanced visual recognition method, which combines feature alignment and multimodal fusion, solves the problem of fine-grained category differentiation in mixed-granularity recognition tasks by enabling efficient cross-modal feature alignment and fusion, thereby improving recognition accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-19
AI Technical Summary
Existing visual models struggle to distinguish fine-grained categories when handling mixed-granularity recognition tasks, and EEG-enhanced visual recognition technology has limited capabilities in cross-modal alignment and fusion representation.
A visual recognition method based on feature alignment and multimodal fusion is adopted to enhance EEG recognition. This method achieves cross-modal feature alignment and fusion through randomized image sequence design, event-related potential (ERP) paradigm, whole-brain high temporal resolution EEG signal acquisition, image and EEG feature extraction, supervised contrastive learning with hard negative sample weighting mechanism, and a Transformer-based deep fusion model.
It improves the model's ability to distinguish fine-grained categories, increases recognition accuracy, and enhances the model's robustness and generalization ability.
Smart Images

Figure CN122065104A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision, brain-computer interface and multimodal information fusion, and in particular to a brainwave-enhanced visual recognition method and system. Background Technology
[0002] In recent years, artificial intelligence technologies, represented by deep learning, have achieved great success in the field of computer vision, especially in standard object recognition tasks, where their performance has rivaled or even surpassed that of humans. However, existing visual models still face challenges when dealing with more complex mixed-granularity recognition tasks.
[0003] Mixed-granularity recognition requires the model to not only distinguish between different major categories (such as "dog" and "car"), but also to distinguish between similar minor categories (such as "husky" and "border collie"), which places extremely high demands on the model's ability to capture details and its robustness.
[0004] In contrast, the human visual system, benefiting from complex attention mechanisms, can easily and reliably identify objects, unaffected by factors such as scale, lighting, and background clutter. More importantly, the human brain possesses remarkable generalization capabilities, applying knowledge gained from limited visual experience to new cognitive and recognition situations. This indicates that the neurovisual processing mechanisms within the human brain can construct efficient visual representations and effectively utilize these representations to decode brain activity, thereby inferring visual categories.
[0005] Therefore, by integrating human brain responses (such as EEG signals) into image features for machine learning, we can not only gain a deeper understanding of the complexity of the human brain, but also explore complementary information in brain responses to build more advanced and robust computer vision models. Summary of the Invention
[0006] The purpose of this invention is to provide an EEG-enhanced visual recognition algorithm and system based on feature alignment and multimodal fusion, so as to solve the problems of existing EEG-enhanced visual recognition technologies, such as difficulty in fine-grained category differentiation, insufficient cross-modal alignment, and limited fusion representation capabilities.
[0007] To this end, the present invention provides a visual recognition method based on enhanced EEG signals. S1: Group a multi-category image dataset, generating an image sequence containing multiple different image samples for each category, and randomizing the image order within the sequence and the presentation order of different category sequences; S2: Present the image sequence to the subject using the event-related potential (ERP) paradigm, and randomly insert task targets unrelated to the main task into the sequence for attention monitoring; S3: While the subject views the image sequence, simultaneously collect high temporal resolution EEG signals from the whole brain, and preprocess the raw signals to obtain clean EEG data segments corresponding one-to-one with each image; S4: Extract features using an image encoder and an EEG encoder respectively, and jointly train them using a supervised contrastive learning objective including a hard negative sample weighting mechanism to generate aligned feature representations in a unified semantic space; S5: Input the aligned feature representations into a Transformer-based fusion model, perform deep information interaction through a cross-attention mechanism to generate the final fused feature representation, and output the recognition result by a classifier.
[0008] Furthermore, in step S1, the image sequence (also known as the stimulus sequence) is generated using a dual randomization strategy. This strategy effectively reduces visual adaptation effects and class order bias by randomly sorting the presentation order of image samples within each category (intra-class randomization) and pseudo-randomizing the presentation order of sequences from different categories (inter-class pseudo-randomization).
[0009] Furthermore, in the signal preprocessing stage of step S3, the original continuous EEG signal is sequentially filtered to remove environmental noise and physiological artifacts, segmented to extract independent data segments (Epochs) corresponding to each image based on event markers, and artifact removal to automatically remove severely contaminated data segments.
[0010] Furthermore, in step S4, the joint training employs a two-stage training strategy. First, a feature alignment operation is performed to fine-tune the image encoder and train the EEG encoder; then, the parameters of the two encoders are frozen, and a fusion and classification operation is performed, in which only the multimodal fusion module and the classifier are trained.
[0011] Furthermore, in step S4, the supervised contrastive learning objective is a multi-task supervised contrastive learning objective. This objective aims not only to optimize feature alignment between different modalities (EEG-image), but also to simultaneously optimize cluster consistency within each modality (image-image, EEG-EEG), thereby improving the overall quality of the learned feature representations.
[0012] Furthermore, the calculation of intermodal loss and intramodal loss includes a Hard Negative Weighting mechanism, which is configured to dynamically increase the weight of negative samples with high feature similarity but different categories in the loss calculation for a given anchor sample, so as to enhance the model's ability to distinguish fine-grained categories.
[0013] Furthermore, the architecture of the Transformer-based fusion model adopts a two-stream design, which can process the input image feature representation and EEG feature representation in parallel and explicitly model the complex dependencies between the two.
[0014] This invention also provides a visual recognition system based on enhanced EEG signals, comprising: a stimulus presentation and data synchronous acquisition module, used to group multi-category image datasets, generate image sequences containing multiple different image samples for each category, and randomize the image order within the sequence and the presentation order of different category sequences; presenting the image sequences to subjects using the event-related potential (ERP) paradigm, and randomly inserting task targets unrelated to the main task into the sequence for attention monitoring; a signal preprocessing module, used to synchronously acquire high temporal resolution EEG signals of the whole brain while subjects view the image sequences, and preprocess the raw signals to obtain clean EEG data segments corresponding one-to-one with each image; a feature extraction module, used to extract features using an image encoder and an EEG encoder respectively, and jointly train them using a supervised contrastive learning objective including a hard negative sample weighting mechanism to generate aligned feature representations in a unified semantic space; and a fusion and classification module, used to input the aligned feature representations into a Transformer-based fusion model, perform deep information interaction through a cross-attention mechanism to generate the final fused feature representation, and output the recognition result by a classifier.
[0015] Compared with the prior art, the present invention has the following technical advantages / effects:
[0016] First, at the data acquisition level, this invention proposes an event-related potential (ERP) experimental paradigm for multi-granularity visual recognition tasks, ensuring high signal-to-noise ratio and strong semantic relevance of the collected EEG data.
[0017] Secondly, at the core algorithm level, this invention proposes a two-stage framework of "alignment and fusion". The first stage introduces multi-task supervised contrastive learning with a hard negative sample weighting mechanism to learn more discriminative feature representations, solving the problem that traditional methods have difficulty distinguishing fine-grained categories. The second stage designs a Transformer-based deep fusion model and adopts a decoupling strategy of "learning representations first, freezing them, and then learning fusion" to achieve in-depth mining and modeling of the complex relationship between EEG and image data, fully releasing the complementary advantages of multimodal information.
[0018] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the figures. Attached Figure Description
[0019] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0020] Figure 1 This is a schematic diagram of the overall process of the EEG-enhanced visual recognition method based on feature alignment and multimodal fusion of the present invention;
[0021] Figure 2 This is an example segment of the multi-channel raw EEG signal waveform acquired by this invention;
[0022] Figure 3 This is a detailed structural diagram of the alignment and fusion framework of the present invention;
[0023] Figure 4 This is a schematic diagram comparing the effects of the hard negative sample weighting mechanism proposed in this invention with typical contrastive learning. The diagram shows the effect of the weighting mechanism ( Adaptively adjust the repulsion force for difficult samples; concentric circles represent contours equidistant from anchor points, with larger radii indicating lower similarity.
[0024] Figure 5 The chart shows the performance comparison results of the method of this invention on the public dataset EEG-ImageNet, compared with various baseline methods including image-only models, EEG-only models and simple multimodal fusion models.
[0025] Figure 6 The chart shows the performance comparison results of the method of this invention with various baseline methods on the public dataset EEGCVPR;
[0026] Figure 7 The chart shows the ablation experimental results of key modules of the method of this invention on the EEG-ImageNet dataset. Detailed Implementation
[0027] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0028] This invention constructs and implements a complete closed-loop system for joint cognitive decoding of EEG and vision. It presents an experimental paradigm for event-related potentials (ERPs) for multi-granularity visual recognition tasks, realizing a controllable data acquisition process from stimulus design and EEG acquisition to behavioral feedback. At the algorithm level, it integrates multi-task supervised contrastive learning with a hard negative sample weighting mechanism and a Transformer-based deep fusion architecture, realizing the entire process from heterogeneous modality feature extraction, cross-modal alignment, fusion modeling, to classification decision-making.
[0029] Combined with reference Figures 1 to 4 The EEG-enhanced visual recognition algorithm and system based on feature alignment and multimodal fusion of the present invention includes the following steps S1 to S5.
[0030] S1. Divide the multi-class image dataset into groups, generate an image sequence containing multiple different image samples for each class, and randomize the image order within the sequence and the presentation order of different class sequences. Details are as follows.
[0031] In step S1 of this invention, visual stimuli are prepared for subsequent ERP experiments. The core of this step lies in constructing a stimulus sequence that can minimize the effects of sequence and expectation.
[0032] Specifically, several categories (e.g., 40 categories) are first selected from a multi-class image dataset (e.g., a subset of ImageNet), with each category containing multiple distinct image samples (e.g., 50 images). The overall image rendering process is divided into multiple blocks, and a dual randomization strategy is used for sequence generation.
[0033] Intra-class randomization: Image samples within each class are grouped into multiple image sequences (e.g., 50 images per class are divided into 5 sequences, with 10 images per sequence), and the presentation order of the 10 images within each sequence is randomly sorted.
[0034] Inter-class pseudo-randomization: In each block of the experiment, the image sequences from different categories are arranged in a pseudo-randomized manner. That is, while ensuring a balanced presentation between different categories, the order in which the categories appear is randomly distributed, thereby avoiding systematic order bias between categories.
[0035] S2. Present the image sequence to the user using the event-related potential (ERP) paradigm, and randomly insert task targets into the sequence to monitor attention and ensure the user's focus throughout the experiment.
[0036] Step S2 aims to precisely present the randomized image sequence generated in S1 to the user through a rigorous event-related potential (ERP) visual stimulus paradigm with attention monitoring mechanisms.
[0037] In a preferred embodiment, the presentation of each image sequence follows a precise temporally controlled flow designed to elicit clear and temporally locked neural responses. A typical trial structure for the presentation of a single sequence is as follows.
[0038] Preparation phase: First, a fixation point (e.g., a white "+" sign) is presented in the center of the screen for 0.75 seconds. The purpose of this phase is to guide the subject to focus their attention on the center of the screen and prepare them for the upcoming stimulus sequence.
[0039] Stimulus Presentation Phase: Immediately following, the image sequence generated in step S1 (e.g., containing 10 similar images) is played out in a fast sequence visual presentation manner. The presentation duration of each image is 500 milliseconds, and the total sequence duration is 5 seconds.
[0040] Buffering phase: After the image sequence finishes playing, the screen remains blank for 0.75 seconds. This phase separates the stimulus presentation from the subject's response, preventing motion artifacts from interfering with the EEG signals.
[0041] Response and Rest Phase: At the end of each sequence, a 2-second window is provided. During this period, participants can engage in relaxation activities such as blinking and perform the attention monitoring task described below.
[0042] To ensure users remain focused throughout the experiment, an attention monitoring task is also embedded in this paradigm. Specifically, in each block's image sequence, a special target image completely unrelated to the main task category (e.g., an animated image of "Buzz Lightyear") is inserted at random frequencies and positions.
[0043] Task requirements: Participants were asked to use a button press during the response phase after each sequence to indicate whether they had observed the specific target image in that sequence.
[0044] Technical objective: The design of this task can not only effectively maintain the alertness and participation of the participants, but also the accuracy of their behavioral feedback can serve as an objective indicator for screening and eliminating experimental data that may have poor signal quality due to lack of concentration during the data analysis phase.
[0045] S3. While the user is viewing the image sequence, high temporal resolution EEG signals of the whole brain are simultaneously acquired, and the raw signals are subjected to standardized preprocessing such as filtering, segmentation and artifact removal to obtain clean EEG data segments corresponding to each image.
[0046] EEG signal recording: Throughout the experiment, a professional EEG acquisition device (a 64-lead EEG cap with electrodes arranged according to the international 10-20 system) was used to synchronously record the whole brain EEG signals of the subjects. The sampling rate was preferably set to 1000Hz.
[0047] Synchronization Mechanism: In S2, at the exact moment each visual image is presented on the screen, the stimulus presentation computer sends an event tag to the EEG recording system via a parallel or USB port. This tag is precisely embedded into the continuous EEG data stream, providing a millisecond-level time synchronization reference for subsequent signal segmentation.
[0048] Filtering: First, a bandpass filter is applied to the continuous EEG data, for example, with a passband range of 0.5Hz to 95Hz, to filter out DC drift and high-frequency muscle electrical activity noise. Subsequently, a 50Hz notch filter is applied to eliminate power frequency interference.
[0049] Segmentation: Based on the image presentation time points recorded in the experiment, the continuous EEG data was segmented into independent epochs. A fixed-length time window of 40ms to 440ms was extracted after each stimulus presentation time (0ms) to minimize the influence of preceding and following stimuli on the current stimulus.
[0050] Artifact Rejection: Automatically identifies and removes trials containing large-value artifacts such as eye movements, blinks, or muscle activity by setting a voltage threshold (e.g., ±100μV) to ensure the purity of the data used for model training.
[0051] S4. Features are extracted by image encoder and EEG encoder respectively, and jointly trained using a supervised contrastive learning objective that includes a hard negative sample weighting mechanism to generate feature representations that are aligned in a unified semantic space and have high discriminative power for fine-grained categories.
[0052] Image encoder ( This embodiment employs various deep convolutional neural networks (DCNNs), such as the ResNet model pre-trained on the ImageNet-1K dataset, as the image encoder. Given a pre-processed image... The output of the global average pooling layer after the last convolutional block is taken as the initial image feature. .
[0053] EEG encoder ( This embodiment uses the EEGNet model, specifically designed for decoding EEG signals, as the EEG encoder. This model effectively captures the spatiotemporal dynamics of EEG signals through separable convolutions and depthwise convolutions. Given a preprocessed EEG sample... EEGNet outputs the initial EEG features .
[0054] To achieve cross-modal feature alignment and intra-modal discriminative enhancement, this embodiment uses Supervised Contrastive Learning (SCLM) as the core training framework. Its core idea is to bring features from the same class (positive sample pairs) closer together in the feature embedding space, while pushing features from different classes (negative sample pairs) further apart. Unlike traditional cross-entropy loss, which only utilizes label information for independent classification, SCLM learns a more structured and discriminative feature space by constructing relationships between samples.
[0055] To comprehensively optimize the quality of feature representation, this embodiment employs a multi-task learning objective, the objective function of which is... It consists of two parts: intermodal contrast loss ( ) and intramodal contrast loss ( ), Intramodal contrast loss ( )Include and .
[0056]
[0057] in, and It is a hyperparameter used to balance the two parts of the loss.
[0058] Intermodal contrast loss ( This aims to achieve cross-modal feature alignment. For any anchor sample in a batch... (For example, an EEG sample), whose set of positive samples It includes all samples from another modality that belong to the same class as the anchor point (i.e., all image samples of the same class). Its negative samples are all other samples from the other modality. (EEG samples) As the anchor point, its intermodal loss for the image sample can be expressed as:
[0059]
[0060] in, and These represent the feature extraction functions of the EEG and image encoder, respectively. , These are positive and negative image samples, respectively. For temperature hyperparameters, It is a set containing one positive sample and all negative samples. These are hard negative sample weights. Overall. It is the sum of the losses for all samples used as anchors (including EEG anchors and image anchors).
[0061] Intramodal contrast loss ( This aims to enhance the cluster structure within a single modality, making similar samples more compact and dissimilar samples more dispersed. For an anchor sample... Its positive sample set Includes all other samples from the same modality and belonging to the same class as the anchor. (EEG samples) As an anchor point, its intramodal loss can be expressed as:
[0062]
[0063]
[0064] To enhance the model's ability to distinguish fine-grained categories, this embodiment introduces a hard negative sample weighting mechanism in the aforementioned contrastive learning process. This mechanism guides the model to focus on learning to distinguish semantically similar sample pairs by assigning higher weights to difficult-to-distinguish negative samples. In a preferred embodiment, the weights... The calculation method is as follows:
[0065]
[0066] in, and These are anchor points and negative samples The final feature vector after the projection layer. It is a function used to measure the similarity of feature vectors, and cosine similarity is preferred. and These are samples and The respective tags, It is the Sigmoid activation function, which maps similarity values to the interval (0, 1). and These are hyperparameters that control the weighting effect. Control the maximum weighting magnitude. Control the sensitivity of weights to changes in similarity.
[0067] S5. Input the aligned feature representation into a Transformer-based fusion module, perform deep information interaction through a cross-attention mechanism to generate the final fused feature representation, and output the recognition result by the classifier.
[0068] In a preferred embodiment of the invention, this step corresponds to the second stage of the two-stage training strategy. Specifically, before performing this stage of training, the network parameters of the image encoder and EEG encoder, which have already been trained in step S4, are frozen to prevent gradient updates during backpropagation. Subsequently, only the parameters of the multimodal fusion module and classifier in this step are trained and optimized. This decoupled training design ensures that the fusion stage focuses on learning the fusion strategy for cross-modal features, thereby avoiding interference with or damage to the stable and high-quality feature representations obtained in the previous stage.
[0069] In a preferred embodiment of the present invention, the multimodal fusion module is a Transformer-based deep fusion architecture, internally composed of L (preferably L=2) identical fusion blocks stacked together. This architecture employs a two-stream design to process the input image features and EEG features in parallel. Specifically, the internal processing flow of each fusion block includes:
[0070] Co-Attention Layer: This is the core unit for enabling cross-modal information interaction. In one processing stream (e.g., an image stream), this layer uses features from its own stream as a query and features from another processing stream (EEG signal stream) as keys and values. Through an attention mechanism, it computes and aggregates relevant information from the other modality to update the feature representation of its own stream. This interaction process occurs bidirectionally and symmetrically between the two streams, achieving mutual enhancement and complementarity of features.
[0071] Self-Attention Layer: After the information exchange through the cross-attention layer, each stream passes through a standard self-attention layer. This layer models and refines the internal contextual relationships of features that have partially fused with information from other streams, in order to capture higher-order semantic dependencies.
[0072] Feed-Forward Network: After the attention layer, a feed-forward network is used to perform non-linear feature transformation.
[0073] Preferably, residual connections and layer normalization operations are applied after each sublayer (cross attention, self attention, feedforward network) to ensure the stability and convergence speed of deep model training.
[0074] In a preferred embodiment of the present invention, the feature representations of the two streams output from the last fusion block are used to generate a final single fused feature representation through a feature aggregation operation.
[0075] Figure 5 The chart shows the performance comparison results of the method of the present invention on the public dataset EEG-ImageNet, compared with various baseline methods (including image-only models, EEG-only models and simple multimodal fusion models), demonstrating the superior performance of the method of the present invention on this dataset. Figure 6 The chart shows the performance comparison results of the method of the present invention on the public dataset EEGCVPR with various baseline methods, further verifying the effectiveness and robustness of the method of the present invention on different datasets. Figure 7 These are ablation experiment results of the key modules of the method of this invention on the EEG-ImageNet dataset, used to verify the necessity and effectiveness of the contrastive learning (feature alignment) module and the multimodal Transformer (fusion) module proposed in this invention.
[0076] Combination Figures 5 to 7 As can be seen, on the classic EEGCVPR dataset, the multimodal method of this invention significantly improves the recognition accuracy from 92.95% of the visual baseline to 95.82%; on the more challenging EEG-ImageNet dataset, it also improves the accuracy from 89.17% to 92.56%, demonstrating its clear technical advantages.
[0077] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A visual recognition method based on enhanced electroencephalogram (EEG) signals, characterized in that, S1. Group the multi-class image dataset, generate an image sequence containing multiple different image samples for each class, and randomize the image order within the sequence and the presentation order of different class sequences. S2. The image sequence is presented to the subjects using the event-related potential (ERP) paradigm, and task targets unrelated to the main task are randomly inserted into the sequence for attention monitoring. S3. While the subjects were viewing the image sequence, high temporal resolution EEG signals of their whole brain were simultaneously acquired, and the raw signals were preprocessed to obtain clean EEG data segments corresponding to each image. S4. Features are extracted using an image encoder and an EEG encoder respectively, and a supervised contrastive learning objective with a hard negative sample weighting mechanism is used for joint training to generate feature representations aligned in a unified semantic space. S5. Input the aligned feature representation into a Transformer-based fusion model, perform deep information interaction through a cross-attention mechanism to generate the final fused feature representation, and output the recognition result by the classifier.
2. The visual recognition method based on enhanced electroencephalogram (EEG) signals according to claim 1, characterized in that, In step S1, the image sequence is generated using a dual randomization strategy to make the order of images within the sequence and the presentation order of different categories of sequences random.
3. The visual recognition method based on enhanced electroencephalogram (EEG) signals according to claim 1, characterized in that, In step S2, the subject is presented with a configured sequence of images using the event-related potential (ERP) paradigm. This configuration includes setting the stimulus time window, image display duration, and task target insertion frequency.
4. The visual recognition method based on enhanced electroencephalogram (EEG) signals according to claim 1, characterized in that, In step S3, the original signal is preprocessed, including filtering, segmenting, and artifact removal of the EEG signal to obtain clean EEG data segments corresponding to each image.
5. The visual recognition method based on enhanced electroencephalogram (EEG) signals according to claim 1, characterized in that, In step S4, the supervised contrastive learning objective includes intermodal loss and intramodal loss, wherein the intermodal loss is used to optimize the alignment of EEG features and image features, and the intramodal loss is used to optimize the density of clusters within each modality.
6. The visual recognition method based on enhanced electroencephalogram (EEG) signals according to claim 5, characterized in that, The calculation of intermodal and intramodal losses incorporates a hard negative sample weighting mechanism, which is configured to dynamically increase the weight of negative samples with high feature similarity but different categories in the loss calculation for a given anchor sample, thereby enhancing the model's ability to distinguish fine-grained categories.
7. The visual recognition method based on enhanced electroencephalogram (EEG) signals according to claim 5, characterized in that, The Transformer-based fusion model includes at least one cross-attention layer for intermodal information interaction and at least one self-attention layer for intramodal information extraction. In the cross-attention layer, the input aligned image features and EEG features interact through a multi-head attention mechanism to capture intermodal dependencies. In the self-attention layer, the features processed by cross-attention are used to extract deep intramodal features through multi-head attention.
8. The visual recognition method based on enhanced electroencephalogram (EEG) signals according to claim 5, characterized in that, The image encoder is a pre-trained deep convolutional neural network (DCNN), and the EEG encoder is an EEGNet model used to extract EEG features with temporal correlation and spatial dependence.
9. A visual recognition system based on enhanced electroencephalogram (EEG) signals, characterized in that, include: The stimulus presentation and data synchronization acquisition module is used to group multi-category image datasets, generate image sequences containing multiple different image samples for each category, and randomize the image order within the sequence and the presentation order of different category sequences; the image sequence is presented to the subject using the event-related potential (ERP) paradigm, and task targets unrelated to the main task are randomly inserted into the sequence for attention monitoring; The signal preprocessing module is used to simultaneously acquire high temporal resolution EEG signals of the whole brain while the subject views the image sequence, and preprocess the raw signals to obtain clean EEG data segments corresponding to each image. The feature extraction module is used to extract features using the image encoder and the EEG encoder respectively, and to jointly train them using a supervised contrastive learning objective that includes a hard negative sample weighting mechanism, so as to generate feature representations aligned in a unified semantic space. The fusion and classification module is used to input aligned feature representations into a Transformer-based fusion model, perform deep information interaction through a cross-attention mechanism, generate the final fused feature representation, and output the recognition result by the classifier.
10. The visual recognition system based on EEG signal enhancement according to claim 9, characterized in that, The feature alignment module includes a dynamic weight adjustment unit, which is used to calculate the weight coefficients in real time based on the feature similarity between the anchor sample and the negative sample during the training process, so as to achieve adaptive optimization of the hard negative sample weighting mechanism.