A brain-computer information fusion classification method and system based on shared subspace learning
By employing a comparative learning method involving shared subspace learning and positive/negative sample sampling, a shared subspace for image-brain response is constructed. This addresses the issues of data scarcity and loop applications in existing technologies, enabling efficient transfer of brain cognitive information and improved image classification performance.
Patent Information
- Application Number
- CN202210257094.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-03-16
AI Technical Summary
Existing technologies are limited by the "brain in loop" application paradigm, making it difficult to achieve high-intensity, real-time, fully automated processing. Furthermore, in situations where brain response data is scarce, it is difficult to learn high-quality image-EEG shared subspaces end-to-end, resulting in insufficient transfer of brain cognitive information and cumbersome model deployment.
A shared subspace learning method is adopted, which utilizes a contrastive learning method based on positive and negative sample sampling. The image-brain response dual-stream network model is optimized through the InfoNCE loss function to construct a shared subspace end-to-end, thereby realizing the transfer of brain cognitive information.
This method efficiently learns the association information between images and brain responses with limited data, improves image classification performance in complex and open scenarios, avoids the limitations of "brain in the loop" applications, and enhances the efficiency and stability of the system.
Smart Images

Figure CN114742092B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of brain-computer interface technology application, and in particular relates to a brain-computer information fusion classification method and system for shared subspace learning. Background Art
[0002] In recent years, artificial intelligence methods, represented by deep learning, have developed rapidly, surpassing human performance in image classification tasks. However, deep learning systems are currently only widely applied in limited, specific, and simple scenarios, such as face recognition, speech recognition, and optical character recognition. They are primarily data-driven, requiring the construction of appropriate models and sufficient computing power to fully exploit the distribution patterns within massive amounts of data, and cannot achieve human-like cognitive capabilities. Therefore, when faced with complex and open scenarios such as autonomous driving and remote sensing image interpretation, where targets and backgrounds are complex and changeable, occlusions, and interference are present, it is difficult to fully simulate the complete distribution of the physical world, and it is difficult to establish a universal and robust representation of the target, resulting in a sharp decline in performance, far from achieving strong human-like generalization capabilities.
[0003] Currently, in complex, open application scenarios with low tolerance for erroneous decisions, such as military applications, medical diagnosis, and autonomous driving, manual interpretation by visual recognition experts remains the mainstream approach to image recognition and decision-making. However, this process involves subjective visual cognition and decision-making, and their behavior can be influenced by external environmental factors, fatigue, injury, and other internal factors, leading to erroneous decisions. Compared to machine intelligence, visual experts struggle to perform long-term, high-intensity, and large-scale real-time interpretation. Furthermore, in most fields, qualified visual experts require extensive and costly training.
[0004] Given that the brain is the material basis and control center for the behavior and cognition of primates, especially humans, the use of brain-computer interface technology to build a brain-computer information fusion system from an engineering perspective can achieve deep information perception, interaction, and integration between biological intelligence and machine intelligence, potentially forming a more advanced intelligent model. By transferring high-level cognitive information from the brain to machine intelligence models, this brain-computer information fusion system provides a new processing paradigm for image classification tasks in complex and open environments.
[0005] Currently, there are two main technologies for building brain-computer information fusion classification systems. Existing technology 1: a fusion method based on complementary information between images and brain responses; existing technology 2: a shared subspace learning method based on correlated information between images and brain responses. The main theoretical basis of existing technology 1 is to treat brain responses and image information as representations of different sources of image targets, and to maximize the complementary information between them through information fusion to obtain a more complete joint representation of the image target. Its technical feature lies in the design of a rational information fusion method to maximize the effective information of different modalities. The main representative methods of the existing technology include: "A brain-computer interface for the detection of mine-like objects in sidescan sonar imagery" (IEEE Journal of Oceanic Engineering, 2016, 41(1): 123-138), which uses a feature cascade method to fuse image Haar-type features and subject EEG features, effectively improving the performance of mine target detection in side-scan sonar images; "An adaptive brain-computer information fusion classification method and system" (application number: CN202111017296.4), which constructs a feature reliability learning model of two modalities, learns the feature reliability of images and brain responses, and adaptively adjusts the fusion weights of different modalities, and uses adaptive fusion features for classification. This method maximizes the use of the complementary information of the two modalities and improves the performance of image classification. However, the existing technology requires real-time participation of the brain in its application paradigm. This "brain-in-the-loop" application paradigm is limited by subjective factors such as fatigue and injury of the subject, making it difficult to achieve real-time, high-intensity fully automated application, and does not fully utilize the respective advantages of the brain and the machine. The main theory of the second existing technology is to construct a shared representation space based on the relevant information between images and brain responses, so as to achieve the goal of migrating high-level cognitive information in the brain's cognitive decision-making process to the machine learning model. Its technical feature lies in the design of an efficient associative information learning model.The main representative methods of the existing technology include: "Bridging the Semantic Gap via Functional Brain Imaging" (IEEE Transactions on Multimedia: 2012, 14 (2): 314-325), which uses the PCA method to establish a correlation prediction model between brain magnetic resonance data features and video low-level image features, and realizes video emotion classification by mapping video features to the brain response representation space. However, with the development of deep learning technology, it has been able to extract high-level semantic information of videos, and its performance is no less than that of human recognition. Therefore, similar applications are gradually decreasing; "Decoding Brain Representations by Multimodal Learning of Neural Activity and Visual Features" (IEEE Transactions on Pattern Analysis and Machine Intelligence: 2020) uses triplet loss to optimize a two-stream network built based on deep learning methods, constraining the high-level semantic space of image features to approximate the feature space of the electroencephalogram (EEG) to achieve the transfer of brain cognitive information. However, training the triplet loss is difficult to converge even with a large amount of training data. Due to the characteristics of brain responses, it is difficult to collect sufficiently high-quality EEG data, and existing public data cannot support the training of such models. "A Brain-Computer Information Fusion Classification Method and System for Brain-Out-of-the-Loop Applications" (Application No.: CN202111017290.7) achieves image-to-brain response prediction by constructing a feature domain reconstruction model. By building a feature reliability prediction module to learn the reliability of features from different sources, it achieves adaptive information fusion classification for "brain-out-of-the-loop" applications. However, the related method uses multiple stages of separate learning and processing, and the process is relatively cumbersome, which is more troublesome when deploying the model. How to construct an image-brain response shared subspace learning model when brain response data is scarce, learn the correlation information between the two end-to-end, and achieve the maximum transfer of brain cognitive information.
[0006] Through the above analysis, the problems and defects of the existing technology are as follows:
[0007] (1) Existing technologies are limited by the “brain-in-the-loop” application paradigm, making it difficult to achieve high-intensity, real-time, fully automated processing and to fully leverage the advantages of machine intelligence and fully automated processing.
[0008] (2) Existing technologies are limited by the difficulty of obtaining brain response data. It is difficult to learn high-quality image-EEG shared subspace end-to-end with limited data, and it is difficult to achieve the complete transfer of brain cognitive information.
[0009] (3) Existing methods for implementing adaptive information fusion classification for “brain out-of-the-loop” applications use multiple stages of separate learning processes, which are cumbersome and troublesome when deploying the model. Summary of the Invention
[0010] In response to the problems existing in the existing technology, the present invention provides a brain-computer information fusion classification method and system based on shared subspace learning, and in particular relates to a brain-computer information fusion classification method and system based on shared subspace learning. Its technical characteristics are to use a contrastive learning method based on positive and negative sample sampling to construct an end-to-end image-brain response shared subspace to achieve the transfer of brain cognitive abilities.
[0011] The present invention is implemented as follows: a brain-computer information fusion classification method based on shared subspace learning, which includes a training stage and an inference stage; wherein, the training stage utilizes paired images and brain response data, optimizes the shared subspace model parameters of images and brain responses through a comparative learning strategy of positive and negative sample sampling, and trains an image classifier; the inference stage extracts image features for classification to achieve the application goal of the entire brain-computer information fusion classification system.
[0012] Furthermore, the brain-computer information fusion classification method of shared subspace learning includes the following steps:
[0013] Step 1, training phase:
[0014] (1) Using the ResNet feature extraction structure and fully connected layers, a dual-stream feature extraction network for images and brain responses was constructed as a feature extraction model in the shared subspace.
[0015] (2) Load paired stimulus images and brain response datasets, and optimize the parameters of the two-stream network model in the shared subspace based on the contrastive learning method of positive and negative sampling until the model converges;
[0016] (3) The converged two-stream network is used to extract the image feature set of the training set stimulus images in the shared subspace, and the image feature set is used to train the SVM classifier.
[0017] Step 2, reasoning stage:
[0018] (1) Load the test image and the image branch model in the two-stream network, and extract the image features of the test image in the shared subspace;
[0019] (2) The image features are fed into the SVM classifier, which outputs the probability category of the image feature classification.
[0020] Furthermore, the step 1 of constructing a shared subspace dual-stream feature extraction model includes:
[0021] 1) Use the PyTorch deep learning framework to build a ResNet34 model structure, remove the fully connected layer, and add a fully connected layer with an input size of 512 and an output size of 168 dimensions. Set the model parameter "pretrained = True" and load the ImageNet pre-trained model parameters as the image feature extraction branch of the two-stream network;
[0022] 2) Using the PyTorch deep learning framework, we constructed a three-layer fully connected network with 168-dimensional input and output dimensions and assigned random initialization parameters to serve as the brain response feature extraction branch of the two-stream network.
[0023] 3) Integrate the image and brain response feature extraction module classes into a common module of the two-stream network.
[0024] Furthermore, the step 1 of loading paired stimulus images and brain response datasets includes:
[0025] 1) Image data loading process:
[0026] ①Use PyTorch's Dataset toolkit to load stimulus images;
[0027] ② Use the torchvision transforms toolkit to transform the image size to 224*224, perform random left-right flipping for data enhancement, and then convert the read image data into tensor format.
[0028] 2) Loading process of brain response data:
[0029] ① Load the brain response dataset and average the brain responses captured when the same stimulus image is presented multiple times;
[0030] ② Select electrodes placed in the inferior temporal lobe area and extract the brain response signals corresponding to the electrodes;
[0031] ③ On the brain response signal of each electrode, average along the time dimension to remove the influence of the time dimension;
[0032] ④ Flip the processed brain response into a 1*168-dimensional feature and convert it into a tensor format as the average brain response feature of the stimulation image on each electrode in the IT area.
[0033] 3) Loading process of paired image-brain response data:
[0034] ① Build the dataset public class, index the stimulus image name information, and load the image data; index the corresponding brain response data information according to the image name and load the brain response data;
[0035] ② Return paired image-brain response data.
[0036] Furthermore, the optimization of the dual-stream network model parameters using the contrastive learning method based on positive and negative sampling in step 1 includes:
[0037] 1) Use the PyTorch deep learning framework to load paired image and brain response data, where the batch size is set to 256, and 256 pairs of data are loaded each time;
[0038] 2) Load the two-stream network model parameters, perform forward reasoning, and obtain the feature set of batch images and brain responses, recorded as<f(v),f(b)> ;
[0039] 3) For any image feature f(v i ), category c, all brain response features of the same category in the batch are all positive sample pairs of the current image features, and the image features f(v i ) is a positive sample pair All brain response features of different categories in the batch Does not belong to category c, recorded as the negative sample pair of the current image feature, image feature f(v i ) is a negative sample pair Then, a set of positive / negative brain response features corresponding to each image feature is obtained;
[0040] 4) Use InfoNCE loss function to calculate each image feature f(v i ) corresponds to the contrast loss L i :
[0041]
[0042] Among them, m and n represent the current image features f(v i ) corresponds to the number of positive and negative brain response samples, S(.) represents the cosine similarity of the two features;
[0043] 5) Backpropagating the contrast loss calculated by the InfoNCE loss function to optimize the model parameters of the two-stream network until the contrast loss converges stably.
[0044] Furthermore, the step 1 of extracting image features using a dual-stream network to train an SVM classifier includes:
[0045] 1) Load the parameters of the dual-stream network image branch model, load the training set image data, perform forward reasoning, and obtain the feature set of the image in the shared subspace;
[0046] 2) Use Python's sklearn toolkit to build a linear SVM classifier, use the extracted image features to train the classifier parameters, and save the model parameters.
[0047] The reasoning stage in step 2 is the application reasoning process of the brain-computer information fusion classification model, including:
[0048] 1) Load the image branch model parameters of the two-stream network. Simply load the test image and perform forward inference on the image branch model to extract image features in the shared subspace.
[0049] 2) Load the model parameters of the SVM classifier, input the extracted image features into the classifier, and obtain the classification results of the image.
[0050] Another object of the present invention is to provide a brain-computer information fusion classification system using the brain-computer information fusion classification method of shared subspace learning, wherein the brain-computer information fusion classification system comprises:
[0051] A data loading device is used to load the test image and perform preliminary size transformation and format conversion functions to make it suitable for the input model;
[0052] A feature extraction device is used to store model parameters successfully trained by the contrastive learning method based on positive and negative sample sampling, load input image data and perform forward reasoning to obtain image features in the shared subspace;
[0053] The classifier device is used to store the successfully trained SVM classifier parameters, load image features for SVM classification, and output the classification results.
[0054] Another object of the present invention is to provide a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:
[0055] During the training phase, a two-stream network is used to map images and brain responses to the same subspace respectively. Paired image and brain response data are used to train the two-stream network model parameters of the shared subspace, and the image and brain response features of the current batch are extracted in the shared subspace. The positive and negative sample sampling method based on category information obtains the positive and negative feature sets of the current sample, and the loss value of the current sample is calculated using the InfoNCE loss function. After optimization, the image features of the shared subspace are extracted to train the SVM classifier. During the inference phase, the test image is loaded, the image features of the shared subspace are extracted, and input into the SVM classifier for classification.
[0056] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:
[0057] During the training phase, a two-stream network is used to map images and brain responses to the same subspace respectively. Paired image and brain response data are used to train the two-stream network model parameters of the shared subspace, and the image and brain response features of the current batch are extracted in the shared subspace. The positive and negative sample sampling method based on category information obtains the positive and negative feature sets of the current sample, and the loss value of the current sample is calculated using the InfoNCE loss function. After optimization, the image features of the shared subspace are extracted to train the SVM classifier. During the inference phase, the test image is loaded, the image features of the shared subspace are extracted, and input into the SVM classifier for classification.
[0058] Another object of the present invention is to provide an information data processing terminal, which is used to implement the brain-computer information fusion classification system.
[0059] In combination with the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solutions to be protected by the present invention from the following aspects:
[0060] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving these problems, this paper closely combines the technical solutions to be protected by the present invention and the results and data during the research and development process, and analyzes in detail and in depth how the technical solutions of the present invention solve the technical problems and some creative technical effects brought about by solving the problems. The specific description is as follows:
[0061] The present invention utilizes a contrastive learning method based on positive and negative sample sampling of category information to optimize a two-stream network model of a shared subspace under the constraints of the InfoNCE loss function; the image classification system includes a data loading device, a feature extraction device, and a classifier device, and by preserving the model parameters of the shared subspace, it can realize a "brain out of the loop" image classification application. Compared with the existing technology, the brain-computer information fusion classification system of shared subspace learning extracted by the present invention can train the shared subspace end-to-end, realize the efficient transfer of brain cognitive information, and greatly improve the performance of image classification tasks in complex open scenes. The brain-computer information fusion image classification system proposed by the present invention has an application paradigm that can naturally avoid the limitations of "brain in the loop" applications. Through "brain out of the loop" applications, it greatly improves the efficiency and stability in real-world applications, and has broad application prospects under the new paradigm of brain-computer information collaboration.
[0062] Second, considering the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by the present invention are described in detail as follows:
[0063] The brain-computer information fusion classification method based on shared subspace learning proposed in the present invention can learn the correlation information between images and brain response data end-to-end under limited brain response data, and compared with the triplet loss function, the InfoNCE contrast loss function of positive and negative sample sampling proposed in the present invention can make the two-stream network converge quickly, and efficiently realize the migration of cognitive information of brain response to image model. In addition, the shared subspace learning method proposed in the present invention can directly realize "brain out of the loop" application, greatly give play to the advantages of machine intelligent automation application, greatly improve the efficiency of deployment and application of brain-computer information fusion classification system, and have extremely high application significance. The present invention proposes a contrast learning method based on positive and negative sample sampling, constructs a shared subspace of image-brain response, effectively realizes the migration of brain cognitive information, and can realize "brain out of the loop" application, which greatly improves the performance of image classification in complex open scenes.
[0064] Third, as auxiliary evidence of the invention's creativity, it is also reflected in the following important aspects:
[0065] (1) The expected benefits and commercial value of the technical solution of the present invention after transformation are as follows: After transformation, the technology of the present invention can be used for image recognition and classification tasks in complex and open application scenarios with low tolerance for error rates, such as autonomous driving, remote sensing image interpretation, synthetic aperture radar image interpretation, and intelligent medical assisted recognition and detection. It can combine the dual advantages of machine intelligence and human intelligence in the above application scenarios and improve the classification accuracy of its application system.
[0066] (2) The technical solution of the present invention fills the technical gap in the industry at home and abroad: The technical solution of the present invention fills the gap in applying the contrastive learning method to the field of brain-computer hybrid intelligent computing, and realizes the purpose of constructing an image-brain response shared subspace using a small amount of data.
[0067] (3) Whether the technical solution of the present invention solves the technical problems that people have always wanted to solve but have never been able to solve: The brain-computer hybrid intelligent computing solution based on shared subspace learning proposed in the present invention constructs a shared subspace through end-to-end learning, breaking through the technical difficulties of "brain-in-the-loop" modeling and "brain-out-of-the-loop" application.
[0068] (4) Whether the technical solution of the present invention overcomes technical bias: The technical solution of the present invention confirms the application prospect of introducing the specific brain responses of visual experts into computer vision methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0070] Figure 1 This is a flow chart of the brain-computer information fusion classification method provided by an embodiment of the present invention.
[0071] Figure 2 It is a process diagram of the training phase and the inference phase provided by an embodiment of the present invention.
[0072] Figure 3 This is a principle framework diagram of shared subspace learning provided by an embodiment of the present invention.
[0073] Figure 4 Schematic diagram of positive and negative sample sampling provided by an embodiment of the present invention.
[0074] Figure 5 This is a computer application system diagram of the brain-computer hybrid intelligent classification system provided by an embodiment of the present invention.
[0075] Figure 6 Schematic diagram of some stimulus images for classification tasks provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0076] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0077] In response to the problems existing in the prior art, the present invention provides a brain-computer information fusion classification method and system for shared subspace learning. The present invention is described in detail below with reference to the accompanying drawings.
[0078] 1. Explanatory Examples In order to enable those skilled in the art to fully understand how to implement the present invention, this section provides an illustrative example that expands upon the technical solutions of the claims.
[0079] In response to the problems existing in the existing technology, the present invention provides a brain-computer information fusion classification method and system based on shared subspace learning, which can realize the efficient migration of the brain's visual cognitive information to the machine learning model under the condition of "brain out of the loop" application, and improve the target recognition performance in complex scenarios.
[0080] like Figure 1As shown, the brain-computer information fusion classification method provided by the embodiment of the present invention includes the following steps:
[0081] S101 uses the ResNet feature extraction structure and fully connected layers to construct a dual-stream feature extraction network for images and brain responses, respectively, as a feature extraction model for the shared subspace;
[0082] S102, loading paired stimulus images and brain response datasets, and optimizing the parameters of the two-stream network model in the shared subspace based on the contrastive learning method of positive and negative sampling until the model converges;
[0083] S103, extracting an image feature set of the training set stimulus image in a shared subspace using the converged two-stream network, and training an SVM classifier using the image feature set;
[0084] S104, loading a test image and an image branch model in the two-stream network, and extracting image features of the test image in a shared subspace;
[0085] S105: Send the image features to the SVM classifier and output the probability category of the image feature classification.
[0086] The brain-computer information fusion classification method based on shared subspace learning provided by an embodiment of the present invention loads paired stimulation images and brain response data; uses the paired brain response data to optimize the parameters of the two-stream network model of the training image-brain response shared subspace based on a contrastive learning method of positive and negative sample sampling; extracts the image feature set of the shared subspace and trains a linear SVM classifier to output the classification results.
[0087] like Figure 2 As shown, the brain-computer information fusion classification method based on shared subspace learning provided by the embodiment of the present invention specifically includes the following steps:
[0088] Step 1: Training phase:
[0089] (1) The ResNet feature extraction structure and fully connected layers are used to construct a dual-stream feature extraction network for images and brain responses, respectively, as a feature extraction model in the shared subspace.
[0090] 1) Use the PyTorch deep learning framework to build a ResNet34 model, remove its fully connected layer, and add a fully connected layer with an input dimension of 512 and an output dimension of 168. Set the model parameter "pretrained = True" to load the parameters of the ImageNet pretrained model. This serves as the image feature extraction branch of the two-stream network.
[0091] 2) Use the PyTorch deep learning framework to build a three-layer fully connected network with input and output dimensions of 168 dimensions and assign random initialization parameters as the brain response feature extraction branch of the two-stream network.
[0092] 3) Integrate the image and brain response feature extraction module classes into a common module of the two-stream network.
[0093] (2) Load paired stimulus images and brain response datasets, and optimize the two-stream network model parameters based on the contrastive learning method of positive and negative sampling until the model converges.
[0094] The data loading process mainly includes loading image data, loading brain response data, and loading paired image-brain response data:
[0095] Image data loading process:
[0096] 1) Use PyTorch's Dataset toolkit to load stimulus images.
[0097] 2) Use the torchvision transforms toolkit to transform the image size to 224*224, randomly flip it left and right for data enhancement, and then convert the read image data into tensor format.
[0098] The process of loading brain response data:
[0099] 1) Load the brain response dataset and average the brain responses captured when the same stimulus image is presented multiple times.
[0100] 2) Select electrodes placed in the inferior temporal lobe (IT) and extract the brain response signals corresponding to the electrodes.
[0101] 3) The brain response signal of each electrode is averaged along the time dimension to remove the influence of the time dimension.
[0102] 4) The processed brain response is flipped into a 1*168-dimensional feature and converted into a tensor format as the average brain response feature of the stimulation image on each electrode in the IT area.
[0103] Paired image-brain response data loading process:
[0104] 1) Build the dataset public class, index the stimulus image name information, load the image data, then index the corresponding brain response data information according to the image name and load the brain response data.
[0105] 2) Return paired image-brain response data.
[0106] The present invention trains a dual-stream network for shared subspace learning in an end-to-end manner, such as Figure 3 As shown in FIG, the embodiment of the present invention provides a principle framework diagram of a two-stream network for shared subspace learning. Paired image and brain response features are extracted respectively, and a positive sample set and a negative sample set of the current sample are constructed within the batch by a class-based positive and negative sample sampling method, such as Figure 4 As shown in the figure, the embodiment of the present invention provides a schematic diagram of the principle of category-based positive and negative sample sampling. After determining the positive and negative sample sets in the current sample batch, the current loss value is calculated through InfoNCE, and gradient backpropagation is performed to optimize the network parameters. The specific steps are as follows:
[0107] 1) Use the PyTorch deep learning framework to load paired image and brain response data, where the batch size is set to 256, and 256 pairs of data are loaded each time.
[0108] 2) Load the two-stream network model parameters, perform forward reasoning, and obtain the feature set of batch images and brain responses, recorded as<f(v),f(b)> .
[0109] 3) For any image feature f(v i ), whose category is c, and all brain response features of the same category in the batch are all positive sample pairs of the current image features, that is, image features f(v i ) is a positive sample pair All brain response features of different categories in the batch does not belong to category c, recorded as the negative sample pair of the current image feature, that is, the image feature f(v i ) is a negative sample pair That is, the positive / negative brain response feature set corresponding to each image feature is obtained accordingly.
[0110] 4) Use InfoNCE loss function to calculate each image feature f(v i ) corresponds to the contrast loss L i :
[0111]
[0112] Among them, m and n represent the current image features f(v i ) corresponds to the number of positive and negative brain response samples, S(.) represents the cosine similarity of the two features;
[0113] 5) Backpropagate the contrastive loss calculated using the InfoNCE loss function to optimize the two-stream network model parameters until the contrastive loss converges stably. The Adam optimizer is used for backpropagation during training of the two-stream network. When the loss function converges, save the model parameters. The batch size is set to 128, the initial learning rate is set to 0.1, and the learning rate is decayed by 0.1 every 30 epochs for a total of 100 epochs.
[0114] (3) The converged two-stream network is used to extract the image feature set of the training set stimulus images in the shared subspace and train the SVM classifier.
[0115] 1) Load the parameters of the dual-stream network image branch model, load the training set image data, perform forward inference, and obtain the feature set of the image in the shared subspace.
[0116] 2) Use Python's sklearn toolkit to build a linear SVM classifier, use the image features extracted in the above steps to train the classifier parameters, and save the model parameters.
[0117] Step 2: Reasoning
[0118] (1) To load the image branch model parameters of the two-stream network, it is only necessary to load the test image and perform forward inference on the image branch model to extract the image features in the shared subspace.
[0119] (2) Load the model parameters of the SVM classifier, input the image features extracted in the above steps into the classifier, and obtain the classification results of the image.
[0120] like Figure 5 As shown, the embodiment of the present invention provides an application example of a computer image classification system for a brain-computer fusion system based on shared subspace learning. The system mainly includes a data loading device, a feature extraction device, and a classifier device. Each device of the system can store the computer program required by the corresponding module and the parameters of the successfully trained model to ensure the correct application of the system. The specific information of each device of the system is as follows:
[0121] (1) Data loading device: loads the test image and performs preliminary size transformation and format conversion functions to make it suitable for the input model.
[0122] (2) Feature extraction device: used to store the model parameters successfully trained by the contrastive learning method based on positive and negative sample sampling, load the input image data, and perform forward reasoning to obtain image features in the shared subspace.
[0123] (3) Classifier device: used to store the successfully trained SVM classifier parameters, load image features for SVM classification, and output the classification results.
[0124] 2. Application Examples: In order to demonstrate the creativity and technical value of the technical solution of the present invention, this section provides application examples of the claimed technical solution on specific products or related technologies.
[0125] The creativity of the technical solution of the present invention lies in proposing a brain-computer information fusion classification method and system based on shared subspace learning, and proposing a contrastive learning method based on positive and negative sample sampling to construct a shared subspace. The application basis of the present invention is to use the contrastive learning strategy based on positive and negative sample sampling proposed by the present invention to train a feature extraction model of the shared subspace on the image-brain response data set, and train a classifier based on the features in the shared subspace. The application implementation of the present invention requires saving the above-mentioned successfully trained shared subspace feature extraction model parameters and classifier parameters to the computer hardware system. Subsequent applications can be achieved through the embodiments Figure 5 The described software system loads the image data to be tested, and then performs feature extraction and classifier inference through the above model parameters to output the classification results.
[0126] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0127] 3. Evidence of the effects of the embodiments: The embodiments of the present invention have achieved some positive effects during the development or use process, and indeed have great advantages over the existing technology. The following content describes them with reference to the data, charts, etc. of the experimental process.
[0128] 1. Experimental conditions:
[0129] The hardware conditions for the experiments in this paper are: an ordinary computer with an Intel i5 CPU, 8GB of memory, and an NVIDIA GeForce GTX 1070 graphics card; the software platform: Ubuntu 18.04, the PyTorch deep learning framework, and Python 3.6 language; the brain response and stimulus image dataset used in this paper comes from the Brain-Score platform of the McGovern Institute for Brain Research at the Massachusetts Institute of Technology.
[0130] 2. Training data and test data:
[0131] The dataset used in this paper consists of two parts: stimulus images and brain response data. The stimulus images are composite images of 8 types of targets and random natural scenes, with a total of 3200 images, 400 images of each type. Each stimulus image contains only one target, and the target image is generated by changing the posture of the target object's 3D model, such as Figure 6 As shown, by varying the target pose and randomizing the natural background, this dataset can effectively simulate complex open scenes with complex transformations of targets and scenes. Brain response data were collected from the ventral stream region of two trained adult rhesus monkeys. Brain responses in the corresponding brain regions were captured using a 168-channel electrode array in the inferotemporal region (IT). During EEG acquisition, stimulus images were presented sequentially in the center of the monitor, each image displayed for 100ms, followed by a 100ms blank interval. The monkeys maintained their gaze fixed on the center of the monitor throughout the process. Each stimulus image was presented multiple times, at least 28 times and an average of 50 times. The brain responses were preprocessed using the publicly available data processing framework on the Brain-Score platform (https: / / brain-score.readthedocs.io / en / latest / index.html) to obtain preprocessed brain response features.
[0132] 3. Experimental content:
[0133] Following the training phase described above, the shared subspace two-stream network training process is accelerated using the computer's GPU until the model converges, and the SVM classifier is trained. After successful model training, the model parameters are saved.
[0134] The inference process loads the parameters of each model, performs forward inference, and obtains the classification results.
[0135] 4. Analysis of experimental results
[0136] The present invention uses classification accuracy to describe classification performance and evaluates the classification results of shared subspace learning under different image feature extraction branches, primarily including four image feature extraction networks: AlexNet, VGG, GoogLeNet, and ResNet. Table 1 compares the performance of IT and unimodal image classification with the brain-computer information fusion classification method based on shared subspace learning. As can be seen from the table, the proposed contrastive learning method based on positive and negative sample sampling, which trains the shared subspace of image-brain responses, can effectively improve image classification performance, achieving an average improvement of 7.43% compared to unimodal SVM classification and a 6.05% improvement compared to direct optimization using the InfoNCE loss. This demonstrates that the proposed contrastive learning method based on positive and negative sample sampling based on category information can efficiently transfer brain cognitive information and improve image recognition performance in complex downstream open scenarios. Furthermore, the application paradigm of the present invention naturally avoids the limitations of "brain-in-the-loop" applications and, through "brain-out-of-the-loop" applications, greatly improves efficiency and stability in real-world applications. Therefore, the present invention has greater practical application value and has broad application prospects within the new paradigm of brain-computer information collaboration.
[0137] Table 1 Simulation results
[0138]
[0139] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A brain-computer information fusion classification method based on shared subspace learning, characterized in that: The brain-computer information fusion classification method for shared subspace learning includes a training phase and an inference phase. The training phase uses paired image and brain response data to optimize the shared subspace model parameters of images and brain responses and train an image classifier through a comparative learning strategy of positive and negative sample sampling. The inference phase extracts image features for classification, achieving the application goal of the entire brain-computer information fusion classification system. The brain-computer information fusion classification method based on shared subspace learning includes the following steps: Step 1, training phase: (1) Using the ResNet feature extraction structure and fully connected layers, a dual-stream feature extraction network for images and brain responses was constructed as a feature extraction model in the shared subspace. (2) Load paired stimulus images and brain response datasets, and optimize the parameters of the two-stream network model in the shared subspace based on the contrastive learning method of positive and negative sampling until the model converges; (3) extracting an image feature set of the training set stimulus images in a shared subspace using a converged two-stream network, and training an SVM classifier using the image feature set; Step 2, reasoning stage: (1) Load the test image and the image branch model in the two-stream network, and extract the image features of the test image in the shared subspace; (2) Send the image features to the SVM classifier and output the probability category of the image feature classification; The step 1 of constructing a shared subspace dual-stream feature extraction model includes: 1) Use the PyTorch deep learning framework to build a ResNet34 model structure, remove the fully connected layer, and add a fully connected layer. Set the input size to 512 and the output size to 168 dimensions. Set the model parameter "pretrained = True" and load the ImageNet pre-trained model parameters as the image feature extraction branch of the two-stream network. 2) Using the PyTorch deep learning framework, we constructed a three-layer fully connected network with 168-dimensional input and output dimensions and assigned random initialization parameters to serve as the brain response feature extraction branch of the two-stream network. 3) Integrate the image and brain response feature extraction module classes into a common module of the two-stream network.
2. The brain-computer information fusion classification method based on shared subspace learning as claimed in claim 1, characterized in that: The step 1 of loading paired stimulus images and brain response datasets includes: 1) Image data loading process: ①Use PyTorch's Dataset toolkit to load stimulus images; ② Use the torchvision transforms toolkit to transform the image size to 224*224, perform random left-right flipping for data enhancement, and then convert the read image data into tensor format; 2) Loading process of brain response data: ① Load the brain response dataset and average the brain responses captured when the same stimulus image is presented multiple times; ② Select electrodes placed in the inferior temporal lobe region and extract the brain response signals corresponding to the electrodes; ③ On the brain response signal of each electrode, average along the time dimension to remove the influence of the time dimension; ④ Flip the processed brain response into a 1*168-dimensional feature and convert it into tensor format as the average brain response feature of the stimulation image on each electrode in the IT area; 3) Paired image-brain response data loading process: ① Build the dataset public class, index the stimulus image name information, and load the image data; index the corresponding brain response data information according to the image name and load the brain response data; ② Return paired image-brain response data.
3. The brain-computer information fusion classification method based on shared subspace learning as claimed in claim 1, characterized in that: The optimization of the dual-stream network model parameters by the contrastive learning method based on positive and negative sampling in step 1 includes: 1) Use the PyTorch deep learning framework to load paired image and brain response data, where the batch size is set to 256, and 256 pairs of data are loaded each time; 2) Load the two-stream network model parameters, perform forward reasoning, and obtain the feature set of batch images and brain responses, recorded as<f(v),f(b)> ; 3) For any image feature f(v i ), category c, all brain response features of the same category in the batch are all positive sample pairs of the current image features, and the image features f(v i ) is a positive sample pair All brain response features of different categories in the batch Does not belong to category c, recorded as the negative sample pair of the current image feature, image feature f(v i ) is a negative sample pair Then, a set of positive / negative brain response features corresponding to each image feature is obtained; 4) Use InfoNCE loss function to calculate each image feature f(v i ) corresponds to the contrast loss L i : Among them, m and n represent the current image features f(v i ) corresponds to the number of positive and negative brain response samples, S(.) represents the cosine similarity of the two features; 5) Backpropagating the contrast loss calculated by the InfoNCE loss function to optimize the model parameters of the two-stream network until the contrast loss converges stably.
4. The brain-computer information fusion classification method based on shared subspace learning as claimed in claim 1, characterized in that: The step 1 of using a two-stream network to extract image features and train an SVM classifier includes: 1) Load the parameters of the dual-stream network image branch model, load the training set image data, perform forward reasoning, and obtain the feature set of the image in the shared subspace; 2) Use Python's sklearn toolkit to build a linear SVM classifier, use the extracted image features to train the classifier parameters, and save the model parameters; The reasoning stage in step 2 is the application reasoning process of the brain-computer information fusion classification model, including: 1) Load the image branch model parameters of the two-stream network. Simply load the test image and perform forward inference on the image branch model to extract image features in the shared subspace. 2) Load the model parameters of the SVM classifier, input the extracted image features into the classifier, and obtain the classification results of the image.
5. A brain-computer information fusion classification system implementing the brain-computer information fusion classification method of shared subspace learning according to any one of claims 1 to 4, characterized in that: The brain-computer information fusion classification system includes: A data loading device is used to load the test image and perform preliminary size transformation and format conversion functions to make it suitable for the input model; A feature extraction device is used to store model parameters successfully trained by the contrastive learning method based on positive and negative sample sampling, load input image data and perform forward reasoning to obtain image features in the shared subspace; The classifier device is used to store the successfully trained SVM classifier parameters, load image features for SVM classification, and output the classification results.
6. A computer device, characterized in that: The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: During the training phase, a two-stream network is used to map images and brain responses to the same subspace. Paired image and brain response data are used to train the two-stream network model parameters in the shared subspace, and the image and brain response features of the current batch are extracted in the shared subspace. The positive and negative sample sampling method based on category information obtains the positive and negative feature sets of the current sample, uses the InfoNCE loss function to calculate the loss value of the current sample, and extracts the image features of the shared subspace after optimization to train the SVM classifier; in the inference stage, the test image is loaded, the image features of the shared subspace are extracted and input into the SVM classifier for classification.
7. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps: During the training phase, a two-stream network is used to map images and brain responses to the same subspace respectively. Paired image and brain response data are used to train the two-stream network model parameters of the shared subspace, and the image and brain response features of the current batch are extracted in the shared subspace. The positive and negative sample sampling method based on category information obtains the positive and negative feature sets of the current sample, and the loss value of the current sample is calculated using the InfoNCE loss function. After optimization, the image features of the shared subspace are extracted to train the SVM classifier. During the inference phase, the test image is loaded, the image features of the shared subspace are extracted, and input into the SVM classifier for classification.
8. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the brain-computer information fusion classification system as described in claim 5.
Citation Information
Patent Citations
Self-adaptive brain-computer information fusion classification method and system
CN113869369A
A brain-computer information fusion classification method and system for brain-out-of-loop applications
CN113887559B