Sketch face brain-computer hybrid intelligent recognition method based on RSVP electroencephalogram multi-scale feature extraction and aggregation
By constructing a multi-scale feature extraction and aggregation network using RSVP EEG multi-scale feature extraction, the problem of low accuracy in traditional face recognition sketching was solved, enabling fast and accurate identity recognition in criminal investigation and security scenarios.
Patent Information
- Application Number
- CN202511441380.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-01-13
AI Technical Summary
Traditional computer vision methods struggle to handle the abstract nature and information sparsity of facial sketches. Existing EEG facial recognition technology lacks effective experimental paradigms and recognition methods for sketched faces, resulting in low recognition accuracy and insufficient stability.
The RSVP method for multi-scale EEG feature extraction and aggregation is adopted. By constructing a multi-scale feature extraction and aggregation network, and combining a shallow feature extraction module, a deep feature reconstruction module, a feature information aggregation module, and a spatial processing module, multi-scale features of EEG signals are extracted and aggregated to achieve efficient recognition of sketched faces.
It significantly improves the recognition accuracy and robustness of sketched facial images, enabling fast and accurate identity recognition in scenarios where effective image records are lacking, and is suitable for practical applications such as criminal investigation and security.
Smart Images

Figure CN121330740A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of intelligent face recognition, and particularly relates to a sketch face brain-computer hybrid intelligent recognition method based on RSVP electroencephalogram multi-scale feature extraction and aggregation. BACKGROUND
[0002] Traditional face recognition methods based on computer vision usually rely on clear real photos as input, and for scenes lacking effective image records, only having eyewitness descriptions, etc., which rely on hand-drawn face sketches, such methods are difficult to achieve ideal recognition effect. Therefore, taking hand-drawn face sketches as the target recognition object, developing efficient and accurate recognition technology has become one of the important problems to be solved.
[0003] Face images are complex, multi-dimensional and abstract, and facial features have very abstract semantic features. For computers, based on pixel point calculation, only data-driven features can be extracted. The representation calculated from the original image data has a semantic gap with the high-level features expressed by the image. The human brain can extract rich high-level semantic features of the image, rather than the single visual features extracted by the computer. This makes the human brain have cognitive ability far superior to computer vision in semantic related information extraction.
[0004] With the deep cross-fusion of life science and information technology, using electroencephalogram signals for face recognition has become one of the important trends in the field of brain-computer interface (BCI). The rapid serial visual presentation (RSVP) paradigm is a method of presenting visual stimuli at high speed and continuously and collecting electroencephalogram signals of subjects at the same time, which can help to study the cognitive process of the brain under rapid and high-intensity information stimulation. However, the current electroencephalogram image recognition technology based on the RSVP paradigm mainly focuses on the recognition of photo images or real scene images, and the electroencephalogram recognition research for this special image category of face sketch is still relatively lacking.
[0005] In the field of computer vision, the general idea of face sketch comparison can be divided into two categories: The most traditional method is a generative method, which converts two modal images into a single modal and then matches them. For example, the feature transformation method proposed by Tang and Wang, the local linear embedding method proposed by Liu et al., and the belief propagation method based on Markov random field proposed by Wang and Tang all belong to this method. However, the main disadvantage of this method is that it depends on the quality of the generated image; when the difference between the two image modalities is significant, the quality of the generated result will decrease significantly. In addition, the generation mechanism of the face sketch is difficult to accurately model by rules or grammar, so it is difficult to generate a sketch directly from abstract face features.
[0006] Another method is a discriminative method, which extracts modal-independent features or learns a common subspace to achieve recognition. Representative methods include partial least squares (PLS), coupled information theory projection (CITP), discriminant analysis based on local features (LFDA), canonical correlation analysis (CCA), and self-similarity descriptor dictionary (SSD). However, due to the difficulty in aligning the distributions of the two modalities in the unified feature space, and the limited representation ability of the feature vectors or common space, the recognition effect of this method is still not ideal.
[0007] Traditional computer vision methods cannot cope with the abstractness and sparsity of sketch face images, resulting in low recognition accuracy. Existing EEG face recognition techniques lack experimental paradigms and recognition method designs for sketch faces. EEG signal decoding methods do not fully utilize multi-scale spatiotemporal feature information, limiting the recognition performance of sketch faces. SUMMARY
[0008] To overcome the deficiencies of the prior art, the purpose of the present application is to provide a RSVP EEG multi-scale feature extraction and aggregation sketch face brain-computer hybrid intelligent recognition method, which significantly improves the recognition accuracy and robustness of EEG signals for sketch face images, and solves the problems of low recognition accuracy and insufficient stability in the existing technology for sketch face recognition tasks due to significant modal differences, weak EEG response features, and poor cross-individual generalization ability.
[0009] To achieve the above purpose, the technical solution adopted by the present application is: A RSVP EEG multi-scale feature extraction and aggregation sketch face brain-computer hybrid intelligent recognition method, comprising the following steps: Step 1: Collect sketch face RSVP EEG data, process the RSVP EEG data, and divide it into a training set and a test set; Step 2: Construct a multi-scale feature extraction and aggregation network to decode the processed RSVP EEG data; Step 3: Train the constructed multi-scale feature extraction and aggregation network; Step 4: Acquire real-time EEG signals and preprocess them; Step 5: Decode the preprocessed real-time EEG signal using a trained multi-scale feature extraction and aggregation network and output the classification result; Step 6: When there is a target response, the target map is located and output. If there is no target response, return to step 4 to re-acquire real-time EEG signals.
[0010] In step 1, prepare a sufficient number (at least 300 pairs) of sketch-face paired images. Each pair includes a sketch and a corresponding frontal face photograph. The sketches are drawn in black and white line drawing style, highlighting facial contours and features, and removing background and color interference. The face photographs are standard ID photos with uniform lighting and no obstructions to ensure image quality. All images undergo contrast and brightness standardization before use to ensure consistency and comparability of visual stimuli. Before each trial begins, a crosshair is presented for 2 seconds to help the subject concentrate. Then, the subject views a target sketch (free observation, no time limit). After the observation is completed, the RSVP stage begins. RSVP EEG signals are collected synchronously throughout the presentation. The collected RSVP EEG data is preprocessed and divided into training and test sets.
[0011] The RSVP phase specifically includes: The RSVP sketch face recognition paradigm is designed by learning the features of a sketch face and presenting a sequence of face images in a short period of time with random insertion of target stimuli. This induces event-related potential (ERP) EEG components related to the face recognition process, such as N170, P300, and N250, providing a stable data foundation for subsequent EEG decoding.
[0012] The RSVP sketch face recognition paradigm is as follows: In the RSVP sketch face recognition paradigm, the image materials used include three categories: target images, non-target images, and sketch images; The target image is the real facial ID photo corresponding to the hand-drawn sketch of the face that each subject learned and memorized before the trial, that is, the target identity image that the subject needs to identify in the visual sequence. Non-target images are selected from real facial identification photos of other individuals in the same database, whose appearance and style are similar to the target. Figure One However, their identities differ, and they are primarily used to create interference conditions; The sketch image is a hand-drawn sketch that corresponds one-to-one with the target image, which is provided for the subjects to learn and memorize before the experiment.
[0013] All images are frontal views of the person, clear and unobstructed, and have been standardized in terms of contrast and brightness before use to ensure consistency and comparability of visual stimuli. The subject first learns the target sketch image freely, and then the computer screen presents a visual stimulus sequence composed of 25 face ID photos at a frequency of 5Hz, each image is presented for 200ms without interval; There is only one target photo corresponding to the target sketch in each sequence, and the rest are non-target images. The target image appears randomly, but not in the first five and last five.
[0014] The 25 face ID photos include: 24 non-target face photos (not related to the target identity); 1 target face photo (corresponding to the sketch image), which is inserted randomly to prevent prediction; The subject needs to perform a target identification task during the viewing of the RSVP sequence, and the EEG of the subject is collected simultaneously to obtain the neural response to each image.
[0015] The specific steps of processing the EEG data are: The continuously collected raw EEG signal is time-anchored to each image presentation, and time window slicing is performed; The length of each sample segment is set to the EEG signal within 1 second after the start of the stimulus, i.e. each segment contains 1024 sampling points (sampling rate is 1024 Hz) and covers 64 channels, forming a raw EEG segment with dimensions 64 × 1024; Use a Butterworth band-pass filter (0.1-48Hz) to remove low-frequency drift and high-frequency noise; Downsample the filtered EEG segment from the original sampling rate of 1024 Hz to 256 Hz to retain the main ERP components while reducing data redundancy and computational burden; Then use the Z-score method to normalize each channel to ensure consistency across subjects; Considering that each sequence contains only one target image and the remaining 24 are non-target images, there is a class imbalance problem, the present invention randomly retains one non-target sample and one target sample in each sequence to form a balanced data set, thereby avoiding training bias. Finally, the preprocessed samples are divided into training and test sets in the ratio of 85% and 15%. Through the above design, this step can stably induce key ERP components including N170, N250 and P300 in the sketch and real face cross-modal recognition task, and combined with balanced sampling and block design, it ensures the robustness and reproducibility of the data, providing high-quality input data for subsequent network modeling and decoding. Compared with the prior art which only targets real faces or ignores class balance, the method of the present invention is more suitable for the needs of practical application scenarios such as criminal investigation and security.
[0016] In step 2, a multi-scale feature extraction and aggregation network is used to extract shallow features and deep abstract features in three dimensions, and then the shallow features and deep abstract features are aggregated to maximize the retention of all feature information and achieve efficient decoding of the RSVP electroencephalogram signal. The multi-scale feature extraction and aggregation network specifically includes a shallow feature extraction module, a deep feature reconstruction module, a feature information aggregation module, a spatial processing module, and a classifier module.
[0017] The specific structure of the multi-scale feature extraction and aggregation network is as follows: A shallow feature extraction module is constructed, which is composed of three down-sampling modules. Each down-sampling module includes a convolution layer (extracting local temporal features), a residual connection (retaining original signal information), an average pooling layer (down-sampling and compressing redundancy), and a batch normalization and activation function processing (BN + ReLU); The three down-sampling modules are arranged in hierarchical order and sequentially extract features with different time resolutions, outputting shallow features of three scales respectively, while retaining key information and reducing computational complexity; A deep feature reconstruction module is constructed, which is composed of three up-sampling blocks. Each up-sampling module first performs convolution, then residual fusion with the original input, followed by up-sampling operation to restore the time resolution. The activation function uses ReLU or ELU to improve the non-linear expression ability. Through three-layer structure, it progressively recovers high-order semantic information to obtain deep features of three scales; A feature information aggregation module is constructed, which concatenates the shallow features and the reconstructed deep features of the corresponding scales channel by channel, realizes cross-level temporal feature aggregation, and the spliced features contain both shallow edge detail information and deep abstract semantic information. This structure significantly enhances the model's joint expression ability for fine-grained and global features, avoiding the risk of recognition failure due to relying only on deep abstract representation while ignoring key boundary details; A spatial processing module is constructed. After the feature information aggregation module, the feature is sent to the spatial processing module for three-dimensional spatial dimension information compression and fusion. Spatial convolution kernel is used to integrate the responses of different EEG channels in the channel dimension, and ELU activation function is used to maintain gradient stability. Finally, all spatial dimension features are fused and sent to the classifier. The spatial module retains the regional electrode spatial distribution features while further improving the robustness of classification; Classifier module: the finally fused features are input into a fully connected neural network, and the classification probability result is output after softmax activation, indicating whether the current EEG segment belongs to the target stimulus response.
[0018] Through the above design, the multi-scale feature extraction and aggregation network constructed can fully extract the fine-grained time information of the shallow layer, can reconstruct the high-order semantic features of the deep layer, and can model the time sequence and spatial mode of the EEG through the cross-scale aggregation and spatial convolution at the same time, thereby significantly improving the precision and robustness of the electroencephalogram signal decoding. Compared with the prior art, the present application introduces a shallow-deep symmetric reconstruction and multi-scale aggregation mechanism in the network structure, avoids the problem of loss of boundary details caused by the traditional method of relying only on deep abstract features, and is particularly suitable for processing the characteristics of weak ERP signals and long time span in the RSVP sketch face recognition task.
[0019] The specific training method of step 3 is: (1) Training data preparation: input the training set into the multi-scale feature extraction and aggregation network, and label it as "target stimulus" or "non-target stimulus" binary classification; the training data has been optimized by a balanced sampling strategy to avoid the problem of class imbalance, thereby improving the stability and generalization ability of the model; (2) Training hyperparameter setting: Loss function: use the cross entropy loss function (Cross Entropy Loss) as the objective function to measure the difference between the model prediction probability distribution and the true label; Optimizer: use the Adam optimization algorithm, set the initial learning rate to 0.001, the momentum factor to 0.9, to 0.999 to stabilize the gradient update; Batch size: set to 32 to control the number of samples participating in training at each parameter update; Maximum training rounds (Epochs): set to 100 rounds, and enable the Early Stopping mechanism, which terminates training in advance when the validation set loss does not show a downward trend for multiple Epochs; (3) Regularization strategy: to prevent model overfitting, introduce a Dropout layer or L2 weight decay term at key network layers to reduce the dependence on specific features during training and improve the generalization ability of the model on the test set; (4) Training process execution: the network performs forward propagation (forward pass) in sequence, calculates the cross entropy loss between the model output and the label, and then calculates the gradient through the backward propagation (backward pass) algorithm and updates the parameters based on the Adam optimizer; (5) Performance monitoring and model saving: After each round of training, calculate the accuracy, AUC, F1 value and other indicators on the validation set to evaluate the model performance; Record the network weight corresponding to the moment when the validation set accuracy is optimal; If Early Stopping is triggered, the training is immediately terminated and the current best model is saved as the final network parameters.
[0020] Through the above design, this step ensures that the electroencephalogram features related to target recognition can be fully captured during the training process, while avoiding performance degradation caused by class imbalance or overfitting. Unlike existing technologies that rely solely on offline training, the present application introduces a balanced sampling strategy and dynamic regularization measure in the training process, thereby significantly improving the stability and robustness of the network in cross-individual and online recognition scenarios.
[0021] In step 4, the real-time electroencephalogram signal is preprocessed as follows: In the online recognition phase, before each round of RSVP task begins, a sketch image is first presented to the subject as the target stimulus in that round; the sketch image is a black and white line drawing style face sketch, which the subject can freely observe until they actively press the button to confirm "ready to start", and then enter the formal task phase; In the formal task, 25 ID photo style face images are randomly selected from the face photo library to form a stimulus sequence, and the subject does not know whether the target image is included in the sequence or not, nor does the subject know the specific position of the target image in the sequence, in order to simulate a real recognition task; the image presentation method is to present each frame for 200ms without gaps; During the stimulus presentation process, the subject wears a 64-channel electroencephalogram cap to collect signals, and real-time EEG data stream is acquired at a sampling rate of 1024Hz, and segmented with the image presentation time point as the anchor point, with each segment being 1 second long. The following preprocessing procedures are performed on each segment of data: (1) 0.1-48Hz band-pass filtering to remove low-frequency drift and high-frequency noise; (2) Downsample to 256Hz to reduce data redundancy; (3) Normalize each channel using the Z-score method to ensure data consistency.
[0022] Through the above real-time acquisition and preprocessing design, the present application realizes consistent data stream specification as in the training phase, ensuring that the online signal and offline training signal are completely aligned in terms of frequency range, sampling rate, and normalization scale, thereby effectively avoiding performance degradation caused by distribution shift. This step not only ensures the stability of the online decoding process, but also provides high-quality input for the subsequent multi-scale feature extraction and aggregation network decoding. Compared with existing solutions that only stay in offline experiments, the present application constructs a signal processing procedure that can be directly applied to real-time tasks, which is more suitable for the actual needs of low latency and high reliability in security and criminal investigation scenarios.
[0023] Step 5 is specifically: The preprocessed EEG fragments obtained in step 4 are fed one by one into the multi-scale feature extraction and aggregation network trained in step 3. The network sequentially passes through shallow feature extraction, deep feature reconstruction, cross-scale feature aggregation, and spatial processing modules, and finally outputs the probability distribution in the classifier module. The classification result is given by the Softmax classifier, which outputs the probability value of whether the current EEG fragment belongs to the "target response" or "non-target response".
[0024] Through the aforementioned decoding and output steps, the system can determine whether the current subject has produced a neural response to the target stimulus with a latency of milliseconds. Compared with existing methods that rely solely on offline classification, this invention achieves online reasoning and real-time output during the decoding stage, significantly improving its operability and practical value in criminal investigation and security scenarios.
[0025] Step 6 specifically involves: if the target response probability of the current frame sample is greater than a preset threshold, then it is determined that the target image has been identified; If the target response probability of all frames in the current sequence is below a threshold, the sequence is determined not to contain a valid target response. In this case, the program removes the images already presented in the sequence from the candidate library, automatically generates a new RSVP sequence, and repeats steps 4 to 6 until a target response is detected or the preset maximum number of iterations is reached. This process constitutes the "sequence iteration mechanism" of this invention, which ensures final recognition by progressively reducing the candidate library and continuously regenerating stimulus sequences. This mechanism avoids attentional distraction caused by repeated presentation of target images and improves recognition robustness.
[0026] Once a "target response" is detected, the system records the timestamp and image number of the corresponding image in that frame and feeds it back to the user in real time as a recognition result, which is then displayed on the human-machine interface.
[0027] Through the above design, this step not only achieves accurate localization output when the subject identifies the target, but also ensures the reliability of the recognition result through threshold judgment and sequence iteration mechanism. Compared with existing methods, this invention can achieve rapid recognition of the face corresponding to the sketch in real-world scenarios such as criminal investigation and security, and has significant advantages such as strong real-time performance, low false judgment rate, and high practicality.
[0028] The beneficial effects of this invention are: This invention constructs a sketch-based facial recognition method based on the RSVP (Real-Side Pixel Vertebrate Image) EEG paradigm, breaking through the dependence of traditional facial recognition technology on clear image input. It can assist in identity recognition tasks with only hand-drawn sketches, making it particularly suitable for practical applications such as criminal investigation and security where effective image records are lacking. This invention utilizes the human brain's natural ability to recognize facial images, combined with cognitive response patterns of EEG signals, to construct a highly systematic and operable recognition scheme.
[0029] This invention proposes an RSVP sketch face recognition paradigm that effectively stimulates characteristic EEG activity when users identify targets. The EEG under the paradigm stimulation contains multiple ERP components related to face recognition, providing a high-quality signal foundation for subsequent decoding models.
[0030] This invention proposes a multi-scale feature extraction and aggregation network. This network extracts shallow features at different time scales through a three-layer downsampling module and gradually recovers deep semantic information through a three-layer upsampling module. Downsampling and upsampling are fused across scales through a feature information aggregation module, preserving original temporal details while enhancing abstract semantic expression. Furthermore, the network incorporates a spatial processing module to model the spatial structural relationships between different EEG channels, thereby improving the model's ability to recognize linkage patterns in multi-channel EEG data. The overall structure ensures sensitivity to local features while enhancing the integration ability of global features, significantly improving recognition accuracy and model stability.
[0031] This invention constructs a complete online sketch-based face recognition mechanism, realizing a blind face photo recognition process driven by brainwaves. In practical applications, it locates the face photo identified by the user by decoding brainwave signals in real time. It is simple to operate, highly accurate, and has a short response time, demonstrating good practical deployment capabilities and scalability. Attached Figure Description
[0032] Figure 1 This is a flowchart of the experimental process of the present invention.
[0033] Figure 2 The illustration is for sketching a human face.
[0034] Figure 3 This is for RSVP sketch face recognition paradigm.
[0035] Figure 4 This is a block diagram of the multi-scale feature extraction aggregation network structure. Detailed Implementation
[0036] The present invention will now be described in further detail with reference to the accompanying drawings.
[0037] like Figure 1 As shown, a brain-computer hybrid intelligent recognition method for sketching faces based on RSVP multi-scale feature extraction and aggregation includes the following steps; Step 1: Collect RSVP facial EEG data, process the EEG data, and divide it into training and test sets; Step 2: Construct a multi-scale feature extraction and aggregation network; Step 3: Train the constructed multi-scale feature extraction and aggregation network; Step 4: Acquire real-time EEG signals and preprocess them; Step 5: Decode the preprocessed real-time EEG signal using a multi-scale feature extraction and aggregation network and output the classification result; Step 6: When there is a target response, the target map is located and output; if there is no target response, return to step 4 to start real-time EEG signal acquisition again.
[0038] Step 1 specifically involves: 1.1 Construction of the Paradigm Stimulus Material Library (1) Establish an image library of ≥300 pairs of sketches and real ID photos. Each image pair contains: Black and white line drawing style sketch (background and color removed, retaining facial contours, features and light and shadow structure); A front-facing ID photo that corresponds one-to-one with the sketch (frontal pose, even lighting, no obstructions).
[0039] (2) Standardize image specifications: the shorter side is ≥256 pixels, and the length and width are kept consistent according to the experimental procedure; the contrast and brightness are standardized; the background is uniformly light or solid color to reduce non-semantic variables.
[0040] (3) Non-target library: Select ID photos from the same database that are different from the target identity but have the same style as the target as non-target stimuli.
[0041] 1.2 Subjects and Hardware Conditions (1) Subjects: ≥20 people will be recruited, aged 18–35 years, with normal / corrected vision and no history of neurological diseases; screening will be conducted by the Cambridge Face Memory Test (CFMT) before the experiment; written informed consent will be signed.
[0042] (2) Hardware: 64-channel EEG cap (international 10–20 or its extended positioning), sampling rate 1024 Hz, reference electrodes placed on both earlobes or a common reference; electrode impedance ≤25 kΩ with electrode paste.
[0043] (3) Presentation environment: quiet and interference-free laboratory; subject distance from the monitor is about 60 cm; 27-inch LCD monitor with a refresh rate of ≥60 Hz; constant indoor lighting; good EMI shielding or grounding.
[0044] 1.3 RSVP Paradigm Presentation and Raw Data Acquisition (Same as Paradigm Introduction) (1) Fixation phase: At the beginning of each trial, present a cross fixation for 2 seconds to concentrate attention.
[0045] (2) Target learning: The target sketch of the test is presented, and the subject observes it freely (without time limit). The subject presses a button to confirm that he / she is ready.
[0046] (3) RSVP stage: 25 real ID photos are presented continuously at a frequency of 5 Hz, each frame is 200 ms; there is only one real ID photo corresponding to the target sketch in one of the frames, and its appearance sequence position is random, but it does not appear in the first 5 frames and the last 5 frames.
[0047] (4) Event labeling: Use triggers to write event codes (including frame number, whether it is a target, and image ID) at the beginning of each frame for EEG data alignment.
[0048] (5) Trial organization: Each pair of sketches and photographs constitutes one trial; 20 trials = one block; 20 blocks constitute one complete session; free rest is provided between blocks to avoid fatigue and drift. Each subject accumulates approximately 400 trial samples after completion.
[0049] (6) Raw data acquisition: EEG signals are acquired synchronously throughout the presentation process.
[0050] 1.4 Data Construction and Labeling (1) Segment extraction: Using the trigger of each frame as the anchor point, extract the time window [0, 1.0 s] from the original EEG to obtain a 64×1024 matrix (channels × sampling points).
[0051] (2) EEG signal preprocessing: Bandpass filter: Butterworth 0.1–48 Hz.
[0052] Downsampling: The 1024 Hz Hz is downsampled to 256 Hz, and the time window becomes 64×256.
[0053] Artifact removal: Threshold artifact removal (e.g., ±100 μV), EOG / EMG independent component removal (ICA).
[0054] Baseline correction: The baseline is [-200, 0] ms or the adjacent resting interval.
[0055] Normalization: Perform Z-score by channel (using training set statistics).
[0056] (3) Labeling and Balancing: Label targets / non-targets according to event codes; considering the imbalance of target:non-target = 1:24 in each sequence, randomly retain one non-target in each sequence and pair it with the target to construct a balanced sample set.
[0057] (4) Data partitioning: The samples are divided into training set and test set at a ratio of 85%:15%.
[0058] Technical effects: The RSVP paradigm enables rapid and repetitive stimulation of the target face, which can significantly induce significant event-related potential (ERP) components such as P300. The collected raw data, combined with strict preprocessing and sample balancing measures, ensures the feature stability of the model training in step 3 and the real-time EEG signal acquisition in step 4.
[0059] Through the above steps, the raw EEG signals can be converted into a set of EEG samples with uniform structure, clear annotation, and good balance, providing a data foundation for subsequent deep neural network modeling.
[0060] Step 2 specifically involves: 2.1 Input and Overall Structure: (1) Input tensor: ; (2) Overall module: shallow feature extraction (3-level downsampling) → deep feature reconstruction (3-level upsampling) → inter-scale feature aggregation (3 sets of fused feature maps) → spatial processing (1×C convolution) → classifier (FC+Softmax); (3) Number of channels: Downsampling block output channels: F1=32, F2=64, F3=128; The output channels of the upsampling block are symmetrically set: U1=128, U2=64, U3=32; After spatial processing, the result is unified to Fs=64; before classification, the vector is concatenated to obtain a 3×Fs dimensional vector.
[0061] 2.2 Shallow Feature Extraction Module (3-Level Downsampling) The downsampling block contains: (1) Conv1D (nuclear length K=3, step size 1, padding=same) → BatchNorm1D → ELU; (2) Residual connection: The block input is linearly projected / 1×1 Conv and then added to the main branch; (3) AvgPool1D (kernel 2×1, stride=2) makes the time dimension T successively 256→128→64→32, forming a three-scale feature pyramid of 1 / 2, 1 / 4, 1 / 8, and outputting {D1,D2,D3}.
[0062] The three blocks are arranged hierarchically, and features at different time resolutions are extracted sequentially. Shallow features at three scales are output respectively, preserving key information while reducing computational complexity.
[0063] 2.3 Deep Feature Reconstruction Module (3-Level Upsampling) The temporal features are reconstructed step-by-step using three upsampling blocks symmetrical to the downsampling: (1) UpSampling1D (×2) → Conv1D (K=3) → BatchNorm1D → ELU; (2) Residual connection: Introduce short-circuit summation of the output of the previous stage to alleviate degradation; (3) We obtain {U3,U2,U1}, whose time dimensions are 64, 128, and 256, respectively, which correspond to {D3,D2,D1}.
[0064] By progressively recovering high-order semantic information through a three-layer structure, deep features at three scales are obtained.
[0065] 2.4 Feature Information Aggregation Module (Scale-Aligned Stitching) (1) The components are spliced in the channel dimension according to the one-to-one correspondence of the scale: F1=Concat(D1,U1), F2=Concat(D2,U2), F3=Concat(D3,U3) By concatenating shallow features at the corresponding scale with the reconstructed deep features channel by channel, cross-level temporal feature aggregation is achieved. The concatenated features contain both shallow edge details and deep abstract semantic information. This structure significantly enhances the model's ability to jointly express fine-grained and global features, avoiding the risk of recognition failure that occurs when relying solely on deep abstract representations while ignoring key boundary details.
[0066] 2.5 Spatial Processing Module (Cross-Electrode Spatial Modeling) (1) Apply a 1×C spatial convolution to each Fi to fuse the responses of each electrode in the spatial dimension; (2) Followed by BatchNorm1D + ELU to enhance nonlinearity and stabilize gradient; (3) Compress the spatial output of each scale to a uniform number of channels Fs to obtain {S1,S2,S3}.
[0067] 2.6 Classifier Module (1) Stitching: The three-scale spatial features S1, S2, and S3 are stitched together in the channel dimension; (2) Fully connected (FC): Perform feature integration, followed by Dropout (dropout rate of 0.5); (3) Softmax outputs the probability of "target / non-target".
[0068] Step 3 specifically involves: 3.1 Training Data (1) Data source: The training set data obtained in step 1 was used; (2) Agreement: within the subject (not across subjects); 3.2 Training Setup (1) Loss function: cross-entropy; (2) Optimizer: Adam (initial learning rate 0.001); (3) Batch size: 32; Maximum epoch: 100; (4) Early Stopping: Stop if the validation set loss does not decrease for 10 consecutive rounds; (5) Regularization: L2 weight decay (decay weight is 5e-4); Dropout (dropout rate is 0.5); 3.3 Training Process (1) The network performs forward propagation sequentially, calculating the cross-entropy loss between the model output and the label; (2) Calculate the gradient using the backpropagation algorithm and update the parameters based on the Adam optimizer.
[0069] (3) Validation, evaluation, and model storage: • After each training round, calculate metrics such as accuracy, AUC, and F1 score on the validation set to evaluate model performance; • Record the network weights corresponding to the moment when the validation set accuracy is optimal; • If Early Stopping is triggered, training will be terminated immediately and the current best model will be saved as the final network parameters.
[0070] Technical effect: By systematically setting the network training process and parameter configuration, the convergence speed of the model can be effectively improved, the oscillation or overfitting phenomenon during the training process can be avoided, and the final model can be guaranteed to have stable and reliable classification performance for decoding in step 5.
[0071] In step 4, the preprocessing of the real-time EEG signal specifically involves: Subjects wore 64-channel EEG caps for signal acquisition, and EEG data streams were acquired in real time at a sampling rate of 1024Hz. The data was segmented based on the image presentation time point; each segment was 1 second long, and the following preprocessing procedures were performed on each segment: Low-frequency drift and high-frequency noise are removed by using a 0.1–48Hz bandpass filter; downsampling to 256Hz reduces data redundancy; and the Z-score method is used to normalize each channel to ensure data consistency.
[0072] 4.1 Online RSVP Process and Data Acquisition (1) Repeat the RSVP process in step 1.3: 2 s gaze → target sketch learning → 25 frames of ID photo sequence (200ms / frame, target is random and not in the first or last 5 frames).
[0073] (2) Data acquisition: follow the hardware and reference in step 1.2; use event triggers that are synchronized with the presentation experiment to ensure timescale consistency.
[0074] 4.2 Online preprocessing pipeline (from the same source as training) (1) Trigger Alignment Slice: Extract a window of [0, 1.0 s] from the trigger start point of each frame; (2) Bandpass filter: Butterworth 0.1–48 Hz; (3) Downsampling: 256 Hz; (4) Artifact suppression (lightweighting): Peak-to-peak amplitude threshold (±100 μV) removal or replacement; (5) Normalization reuse: Use the channels saved during training in step 3 to perform Z-score to align the online distribution with the offline training (avoiding domain offset). (6) Output the online sample tensor shape as 64×256, which is consistent with the network input in step 2.
[0075] Technical effect: Obtain online preprocessed samples that are isomorphic to the training samples for real-time decoding in step 5.
[0076] Step 5 specifically includes: 5.1 Forward Inference (1) Input each frame of EEG fragment output in step 4 into the model obtained in step 3; (2) Obtain the frame-by-frame classification confidence P(t), where t is the frame number. 5.2 Prediction Output (1) Give the P(t) value for each frame; (2) Synchronously record the timestamp, frame number, corresponding image ID, and confidence in the log. Technical Effect: This step obtains the frame-by-frame classification result stream through real-time frame-by-frame classification, providing a basis for judgment for target response detection and localization in step 6.
[0077] Step 6 specifically involves: 6.1 Decision Logic Threshold detection: When a frame t satisfies P(t) ≥ a specified threshold, it is determined that a target response has been generated; 6.2 Target Location and Result Output (1) If a target response is generated after the determination, read the image ID flag of that frame; (2) Directly output the corresponding real photo of the target as the recognition result; (3) Display the target face thumbnail, frame number, and confidence level in real time on the display terminal; Output of this step: final recognition result (target image localization), loop closure complete.
[0078] Example 1: RSVP EEG Signal Acquisition Method Stimulus material construction: Prepare sufficient (300 or more pairs) of sketch-face paired images, each pair containing a sketch and a corresponding frontal face photograph; the sketches should be in black and white line drawing style, and the photographs should be in standard ID photo style as attached. Figure 2 As shown.
[0079] Subject preparation and condition control: Multiple subjects were selected for the experiment. All subjects had normal or corrected vision, no history of neurological diseases, and their facial recognition ability was verified through the Cambridge Face Memory Test (CFMT) before the experiment. They were informed of the experimental procedures and precautions and signed written consent forms. Subjects wore 64-channel EEG electrode caps with a sampling rate of 1024Hz, and EEG ointment was applied to keep the impedance of each electrode below 25kΩ to ensure high-quality EEG signals. The experiment was conducted in a quiet and undisturbed environment. Subjects sat approximately 60cm away from the monitor, and stimuli were presented on a 32-inch LCD screen.
[0080] RSVP Task Procedure: Before each trial, a crosshair is presented for 2 seconds to help the subject focus their attention; then, the subject views one target sketch (free observation, no time limit). After observation, the RSVP phase begins, where 25 real face photographs are presented consecutively, each image for 200ms, with no intervals between them; each sequence contains exactly one target photograph corresponding to the target sketch, the rest are non-target images. The target images appear randomly, but not in the first five or last five images, as shown in the attached diagram. Figure 3 As shown.
[0081] Experimental Structure: Each sketched face pair was represented by a stimulus sequence forming a unit. Twenty units were grouped into a block, and a total of 20 blocks comprised the entire process of inducing and collecting data for the sketched face recognition paradigm. Free rest periods were provided between blocks to prevent fatigue from affecting EEG quality. The experiment concluded after data collection from all 20 blocks. This example collected 400 single-trial samples.
[0082] Example 2: EEG sample processing method for constructing training and test sets To illustrate the training set generation method described in this invention, this embodiment, based on Embodiment 1, further performs standardized preprocessing and structured sample generation operations on the collected raw EEG signals for subsequent training and validation of the classification model.
[0083] Data Segmentation: The continuously acquired raw EEG signals are sliced into time windows, with each image presentation serving as a time anchor. The length of each sample segment is set to the EEG signal within 1 second after the stimulus begins, i.e., each segment contains 1024 sampling points (sampling rate of 1024Hz) and covers 64 channels, forming a raw EEG segment with a dimension of 64 × 1024.
[0084] Filtering: The EEG signals were bandpass filtered at 0.1–48 Hz using a Butterworth filter.
[0085] Downsampling: The filtered EEG segments are downsampled from the original sampling rate of 1024 Hz to 256 Hz, preserving the main ERP components while reducing data redundancy and computational burden.
[0086] Normalization: The downsampled data is normalized using the Z-score method.
[0087] Sample selection and class balancing: Since each sequence contains only one target image and the remaining 24 are non-target images, there is a problem of class imbalance. To avoid training bias, the system randomly retains one non-target sample in each sequence to form a balanced dataset together with the target samples.
[0088] Training and test set partitioning: The processed samples are divided into training and test sets according to the following proportions: 85% as the training set and 15% as the test set.
[0089] Through the above steps, the raw EEG signals can be converted into a set of EEG samples with uniform structure, clear annotation, and good balance, providing a data foundation for subsequent deep neural network modeling.
[0090] Example 3: Construction and Training Method of Multi-Scale Feature Extraction Aggregation Network To illustrate the EEG signal classification method described in this invention, this embodiment constructs a multi-scale feature extraction and aggregation network based on the structured sample data generated in Embodiment 2, and trains and optimizes it to decode the target face recognition EEG response under the RSVP sketch face recognition paradigm.
[0091] The network structure includes the following modules (see appendix) Figure 4 The multi-scale feature aggregation network constructed in this example consists of a shallow feature extraction module, a deep feature reconstruction module, a feature information aggregation module, a spatial processing module, and a classifier module, wherein: The shallow feature extraction module comprises three downsampling blocks connected in series, extracting shallow features at different time scales layer by layer. Each downsampling block consists of a one-dimensional convolutional layer, a batch normalization layer, an ELU activation layer, residual connections, and average pooling. First, the one-dimensional convolutional layer performs convolution operations on the EEG signal to extract local temporal features; the batch normalization layer standardizes the convolution results to improve training stability and accelerate convergence. Next, the ELU activation function enhances the nonlinear expressive power of the network and avoids the gradient vanishing problem. Residual connections are used to facilitate information flow and prevent information loss. At the end of each downsampling block, an average pooling layer is used to compress the temporal dimension, progressively compressing the temporal resolution of the signal to 1 / 2, 1 / 4, and 1 / 8 of the original length, forming a multi-scale feature pyramid.
[0092] The deep feature reconstruction module comprises three sequentially connected upsampling blocks, which perform temporal dimension recovery and deep semantic reconstruction on the downsampled feature maps. Each upsampling block includes a one-dimensional convolutional layer, a batch normalization layer, an ELU activation layer, residual connections, and an upsampling operation. The convolutional layer extracts complex temporal features from the downsampled feature map, recovering the deep semantic information of the signal. The batch normalization layer helps stabilize the upsampled features and avoids overfitting. Through the ELU activation function, the network further enhances the nonlinear expressive power of the features. The residual connections help maintain the integrity of the feature information during the upsampling process, avoiding the loss of deep information.
[0093] At each scale, the feature information aggregation module concatenates the shallow features obtained from downsampling with the deep features reconstructed from upsampling along the channel dimension, fusing low-order edge features with high-order abstract information and preserving multi-level temporal features. The upper and lower layer outputs at each scale are matched in a one-to-one correspondence manner to obtain three fused feature maps.
[0094] The space processing module first uses a size of 1× Spatial convolution is performed using convolutional kernels, followed by batch normalization and ELU activation to enhance nonlinear expressive power and alleviate gradient vanishing. Finally, spatial compression is performed to output a unified dimensional fusion feature representation.
[0095] The classifier module concatenates the feature maps output by all spatial modules along the channel dimension and inputs them into a fully connected layer for feature integration. Finally, it connects to a softmax classifier and outputs a binary classification prediction result to determine whether the current EEG sample is the target stimulus.
[0096] The network structure described in this embodiment, through shallow-deep information fusion and multi-scale feature processing mechanisms, can fully extract the spatiotemporal features of EEG related to target face recognition, improve the sensitivity and accuracy of the classification model to the target EEG response, and achieve high recognition accuracy while ensuring computational efficiency.
[0097] Example 4: Training Method for Multi-Scale Feature Extraction Aggregation Network To illustrate the neural network training strategy described in this invention, this embodiment further clarifies the specific process of network training, hyperparameter setting, and model optimization method based on the multi-scale feature extraction and aggregation network constructed in Embodiment 3, aiming to obtain a deep model with good classification performance for RSVP EEG signals.
[0098] The training process in this embodiment includes the following steps: 1. Training data preparation: Input the training set samples constructed in Example 2 into the multi-scale feature extraction and aggregation network, and label them as "target stimulus" or "non-target stimulus" binary classification.
[0099] 2. Training hyperparameter settings: Loss function: Cross-entropy loss is used as the objective function to measure the difference between the model's predicted probability distribution and the true label. Optimizer: Uses the Adam optimization algorithm, with an initial learning rate of 0.001 and a momentum factor. Set it to 0.9. Set it to 0.999 to stabilize gradient updates; Batch Size: Set to 32 to control the number of samples used for training each time parameters are updated; Maximum number of training epochs: Set to 100 epochs, and enable the Early Stopping mechanism to terminate training early when the validation set loss shows no decreasing trend for several consecutive epochs.
[0100] 3. Regularization strategy: To prevent the model from overfitting, Dropout layers or L2 weight decay terms can be introduced into key network layers to reduce the dependence of the training process on specific features and improve the model's generalization ability on the test set.
[0101] 4. Training process execution: The network sequentially performs forward propagation to calculate the cross-entropy loss between the model output and the label; then it calculates the gradient through the backward propagation algorithm and updates the parameters based on the Adam optimizer.
[0102] 5. Performance monitoring and model saving: After each training round, metrics such as accuracy, AUC, and F1 score are calculated on the validation set to evaluate model performance. Record the network weights corresponding to the moment when the validation set accuracy is optimal; If Early Stopping is triggered, training will be terminated immediately and the current best model will be saved as the final network parameters.
[0103] This embodiment, by systematically setting the network training process and parameter configuration, can effectively improve the model convergence speed, avoid oscillations or overfitting during the training process, and ensure that the final model has stable and reliable classification performance, laying the foundation for subsequent online decoding and real-time application scenarios.
[0104] Example 5: Online Target Detection and Recognition Method Based on Electroencephalogram (EEG) Signals To illustrate the decoding and feedback capabilities of the system described in this invention under real-time conditions, this embodiment further constructs a real-time EEG decoding process suitable for online recognition tasks based on the multi-scale feature extraction and aggregation network trained in Embodiment 4, thereby realizing the recognition, judgment, and output of target faces under the RSVP sketch face recognition paradigm.
[0105] The method includes the following steps: Target stimulus preparation Before each round of the RSVP task begins, the system first presents the subject with a sketch image as the target stimulus for that round. The image is a black and white line drawing of a human face, which the subject can observe freely until they actively press a button to confirm "Ready to start," at which point the system proceeds to the formal task phase.
[0106] RSVP Sequence Construction and Presentation The system randomly selects 25 ID-style facial images from a facial image database to form a stimulus sequence. Whether the target image is included in the sequence is unknown to the subject, and the system does not label the specific location of the target image within the sequence, thus simulating a real-world recognition task. Images are presented at 200ms per frame without gaps.
[0107] Real-time EEG signal acquisition and preprocessing Subjects wore a 64-channel EEG cap for signal acquisition. The system acquired EEG data streams in real time at a sampling rate of 1024Hz and segmented the data based on the image presentation time point. Each segment was 1 second long, and the following preprocessing procedures were performed on each segment: 0.1–48Hz bandpass filtering removes low-frequency drift and high-frequency noise; Downsampling to 256Hz reduces data redundancy; The Z-score method is used to normalize each channel to ensure data consistency.
[0108] Classification model inference and response judgment The processed EEG samples are sequentially input into the multi-scale aggregation network trained in Example 4. The model outputs the probability value of each sample belonging to the "target response" or "non-target response".
[0109] The system sets a recognition threshold (e.g., 0.9). If the target response probability of the current frame sample is greater than the threshold, it is determined that the target image has been recognized.
[0110] Rollback mechanism and state update If no high-confidence target response is detected in the current sequence (i.e., the probability of all frames is lower than the set threshold), the system removes the image presented in the sequence from the test image library, generates a new sequence for the next round, and continues to execute the previous steps until the target is successfully identified.
[0111] Results output and feedback presentation Once a "target response" is detected, the system records the timestamp and image number of the corresponding image in that frame and feeds it back to the user in real time as a recognition result, which is then displayed on the human-machine interface.
[0112] This embodiment constructs a complete online recognition closed-loop process. Through low-latency data processing, network decoding, and threshold determination mechanisms, it achieves real-time recognition and localization of target stimuli, and has high accuracy, fast feedback, and dynamic adjustment capabilities, making it suitable for practical neural interface application needs.
Claims
1. A brain-computer hybrid intelligent recognition method for sketching faces based on RSVP multi-scale feature extraction and aggregation, characterized in that, Includes the following steps; Step 1: Collect RSVP EEG data of the sketched face, process the RSVP EEG data, and divide it into training set and test set; Step 2: Construct a multi-scale feature extraction and aggregation network to decode the processed RSVP EEG data; Step 3: Train the constructed multi-scale feature extraction and aggregation network; Step 4: Acquire real-time EEG signals and preprocess them; Step 5: Decode the preprocessed real-time EEG signal using a trained multi-scale feature extraction and aggregation network and output the classification result; Step 6: When there is a target response, the target map is located and output. If there is no target response, return to step 4 to re-acquire real-time EEG signals.
2. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 1, characterized in that, In step 1, sufficient sketch-face pairing images are prepared. Each pair includes a sketch and a corresponding frontal face photograph. The sketches are drawn in black and white line drawing style, highlighting facial contours and features, and removing background and color interference. The face photographs are standard ID photos with uniform lighting and no obstructions to ensure image quality. All images undergo contrast and brightness standardization before use to ensure consistency and comparability of visual stimuli. Before each trial begins, a crosshair is presented for 2 seconds to help the subject concentrate. Then, the subject views a target sketch. After the observation is completed, the RSVP phase begins. RSVP EEG signals are collected synchronously throughout the presentation. The collected RSVP EEG data is preprocessed and divided into training and test sets.
3. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 2, characterized in that, The RSVP phase specifically includes: The RSVP sketch face recognition paradigm was designed by learning the features of sketched faces and presenting a sequence of face images in a short period of time with random insertion of target stimuli to induce event-related potential (ERP) EEG components related to the face recognition process. The RSVP sketch face recognition paradigm is as follows: In the RSVP sketch face recognition paradigm, the image materials used include three categories: target images, non-target images, and sketch images; The target image is the real facial ID photo corresponding to the hand-drawn sketch of the face that each subject learned and memorized before the trial, that is, the target identity image that the subject needs to identify in the visual sequence. Non-target images are selected from real facial ID photos of other individuals in the same database, whose appearance style is consistent with the target image but whose identities are different; The sketch image is a hand-drawn sketch that corresponds one-to-one with the target image, which is provided for the subjects to learn and memorize before the experiment.
4. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 3, characterized in that, All images are frontal photographs of people, clear and unobstructed, and have undergone standardization of contrast and brightness before use to ensure consistency and comparability of visual stimuli. Subjects first freely studied the target sketch image to familiarize themselves with the facial features; then, the computer screen presented a sequence of visual stimuli consisting of facial ID photos at a frequency of 5Hz, with each image presented for 200ms and no interval in between; each sequence contained one and only one target photo corresponding to the target sketch, and the rest were non-target images, and the target images appeared randomly, but not in the first five or the last five; The facial identification photos include: a non-target face photo and one target face photo. The target face photo corresponds to the sketch image, and its insertion position is random to prevent prediction. The subject needs to perform a target recognition task while watching the RSVP sequence. The subject's EEG signals are collected simultaneously to obtain its neural response to each image.
5. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 4, characterized in that, The specific steps for processing the EEG data are as follows: The continuously acquired raw EEG signals were sliced into time windows, with each image presentation serving as a time anchor point. The length of each sample segment was set to the EEG signal within 1 second after the stimulus began, i.e., each segment contained 1024 sampling points and covered 64 channels, forming a raw EEG segment with a dimension of 64 × 1024. Use Butterworth bandpass filters (0.1–48Hz) to remove low-frequency drift and high-frequency noise; The filtered EEG segments were downsampled from the original sampling rate of 1024 Hz to 256 Hz, preserving the ERP components while reducing data redundancy and computational burden. The Z-score method was then used to normalize each channel to ensure consistency of data across subjects; One non-target sample is randomly retained in each sequence, and together with the target sample, they form a balanced dataset.
6. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 5, characterized in that, In step 2, a multi-scale feature extraction and aggregation network is used to extract shallow features and deep abstract features in three dimensions. The shallow features and deep abstract features are then aggregated to achieve efficient decoding of RSVP EEG signals. The constructed multi-scale feature extraction and aggregation network includes a shallow feature extraction module, a deep feature reconstruction module, a feature information aggregation module, a spatial processing module, and a classifier module.
7. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 6, characterized in that, The specific structure of the multi-scale feature extraction aggregation network is as follows: A shallow feature extraction module is constructed, which consists of three downsampling modules. Each downsampling module includes: a convolutional layer, a residual connection, an average pooling layer, and batch normalization and activation function processing (BN + ReLU). The three downsampling modules are arranged hierarchically to extract RSVP EEG data at different time resolutions in sequence, and output shallow features at three scales respectively. A deep feature reconstruction module is constructed, which consists of three upsampling blocks. Each upsampling block is first convolved, then residual fusion is performed with RSVP EEG data, and then upsampling operation is used to restore the temporal resolution. The activation function is ReLU or ELU to enhance the nonlinear expression ability. Through the three-layer structure, high-order semantic information is progressively restored to obtain deep features at three scales. A feature information aggregation module is constructed to connect the shallow features of the corresponding scale with the reconstructed deep features channel by channel, so as to realize cross-level temporal feature aggregation. The spliced features contain both shallow edge detail information and deep abstract semantic information. A spatial processing module is constructed, which contains both shallow edge detail information and deep abstract semantic information. Features are then fed into the spatial processing module for information compression and fusion in three-dimensional space. Spatial convolution kernels are used to integrate the responses of different EEG channels in the channel dimension, and the ELU activation function is used to maintain gradient stability. Finally, features from all spatial dimensions are fused and fed into the classifier. Classifier module: The final fused features are input into a fully connected neural network, and after softmax activation, the classification probability result is output, indicating whether the current EEG fragment belongs to the target stimulus response.
8. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 7, characterized in that, The specific training method for step 3 is as follows: (1) Training data preparation: Input the training set into the multi-scale feature extraction and aggregation network, and label it as "target stimulus" or "non-target stimulus" binary classification; (2) Training hyperparameter settings: Loss function: The cross-entropy loss function is used as the objective function to measure the difference between the model's predicted probability distribution and the true label; Optimizer: Uses the Adam optimization algorithm, with an initial learning rate of 0.001 and a momentum factor. Set it to 0.
9. Set it to 0.999 to stabilize gradient updates; Batch size: Set to 32 to control the number of samples used for training each time parameters are updated; Maximum number of training epochs: set to 100 epochs, and enable the Early Stopping mechanism to terminate training early when the validation set loss shows no decreasing trend for several consecutive epochs. (3) Regularization strategy: To prevent the model from overfitting, Dropout layers or L2 weight decay terms are introduced in key network layers to reduce the dependence of the training process on specific features and improve the model’s generalization ability on the test set. (4) Training process execution: The network performs forward propagation in sequence, calculates the cross-entropy loss between the model output and the label; then calculates the gradient through the backpropagation algorithm, and updates the parameters based on the Adam optimizer; (5) Performance monitoring and model saving: After each training round, metrics such as accuracy, AUC, and F1 score are calculated on the validation set to evaluate model performance. Record the network weights corresponding to the moment when the validation set accuracy is optimal; If Early Stopping is triggered, training will be terminated immediately and the current best model will be saved as the final network parameters.
9. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 8, characterized in that, In step 4, the preprocessing of the real-time EEG signal specifically involves: In the online recognition phase, before each round of the RSVP task begins, a sketch image is presented to the subject as the target stimulus for that round of the task. The sketch image is a black and white line drawing of a human face, which the subject can observe freely until actively pressing the button to confirm "ready to start", and then enters the formal task phase. In the formal task, a set of stimulus sequences is formed by randomly selecting ID photo-style facial images from a facial photo database. Whether the sequence contains the target image is unknown to the subject, and the specific location of the target image in the sequence is not marked, in order to simulate a real recognition task. The images are presented in a 200ms per frame without gaps. During stimulus presentation, subjects wore a 64-channel EEG cap for signal acquisition, obtaining real-time EEG data streams at a sampling rate of 1024Hz. The data was segmented based on the image presentation time point, with each segment lasting 1 second. The following preprocessing procedure was performed on each segment: (1) 0.1–48Hz bandpass filtering to remove low-frequency drift and high-frequency noise; (2) Downsampling to 256Hz reduces data redundancy; (3) Use the Z-score method to normalize each channel to ensure data consistency.
10. The method for RSVP EEG multi-scale feature extraction and aggregation for sketching facial brain-computer hybrid intelligent recognition according to claim 9, characterized in that, Step 5 specifically involves: The preprocessed EEG fragments obtained in step 4 are input one by one into the multi-scale feature extraction and aggregation network trained in step 3. The network goes through shallow feature extraction, deep feature reconstruction, cross-scale feature aggregation and spatial processing modules in sequence, and finally outputs the probability distribution in the classifier module. The classification result is given by the Softmax classifier, which outputs the probability value of the current EEG fragment belonging to "target response" or "non-target response". Step 6 specifically involves: If the target response probability of all frames in the current sequence is lower than the threshold, the sequence is determined not to contain a valid target response. At this time, the program will remove the images already presented in the sequence from the candidate library, automatically generate a new RSVP sequence, and repeat steps 4 to 6 until a target response is detected or the preset maximum number of iterations is reached; this process constitutes a "sequence iteration mechanism", that is, by reducing the candidate library round by round and continuously regenerating stimulus sequences to ensure the final recognition is achieved; Once a "target response" is detected, the system records the timestamp and image number of the corresponding image in that frame and feeds it back to the user in real time as a recognition result, which is then displayed on the human-machine interface.