Multitask simultaneous decoding method for natural image-evoked human brain activity

CN118038138BActive Publication Date: 2026-08-11UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0008]本发明针对现有技术中大脑信息解码模型单一任务解码的局限性,为解决模型应用范围受限以及为各种任务单独设计和管理模型难度增大的问题,提出了一种用于自然图像引发的人脑活动的多任务同步解码方法

Benefits of technology

[0070] This invention establishes a multi-task visual information brain decoding model based on an encoding and decoding framework. It performs classification decoding, semantic decoding, and language decoding on color images of complex natural scenes. The decoded category information has high accuracy, and it also shows high accuracy on most semantic labels. Furthermore, all words in the decoded language strongly point to the main elements or events in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118038138B_ABST
    Figure CN118038138B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-task synchronous decoding method for human brain activity induced by natural images, belonging to the field of multi-task decoding technology for biomedical images. Based on functional magnetic resonance imaging (fMRI) signal data from viewing a large number of natural images, this invention establishes a multi-task visual information brain decoding model, including: a visual encoding module that encodes voxel signals from visually relevant regions into a latent feature space; a multi-task encoding module that acquires multi-task feature vectors including visual information feature vectors, category information feature vectors, and feature vectors for semantic decoding tasks; a category decoding module that acquires the probability distribution of predicted categories; a semantic decoding module that predicts the probability distribution of semantic labels; and a language decoding module that captures deep-level structures and semantic relationships in the text, thereby generating more accurate continuous descriptive text. This invention achieves high accuracy in decoding category information and semantic labels, and the decoded image descriptions can point to their main elements or events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multi-task decoding of biomedical images (including visual scene categories, semantics, and language), and specifically relates to a multi-task synchronous decoding method for human brain activity evoked by natural images. Background Technology

[0002] Since 2005, Kamitani, Tong, and others have been conducting in-depth research on methods for decoding visual information. These researchers have explored a wide range of visual information decoding methods, covering information at all levels, from primary, intermediate, and advanced visual features decoded in the brain to higher cognitive activities, and have achieved many important results.

[0003] In the research on decoding primary visual features, Haynes et al. studied the direction decoding of stripes in 2005; Cossell et al. explored spatial frequency decoding in 2015; Schrantee et al. studied contrast decoding in 2023; and Geisler et al. focused on the decoding of motion direction in 2001.

[0004] In the decoding of intermediate visual features, Pitts et al. studied contour decoding in 2012. Decoding of higher visual features includes Horikawa et al.'s research on object category decoding in 2017; Stansbury's exploration of semantic label decoding in 2013; and Huang et al.'s exploration of descriptive language decoding in 2021. Furthermore, in the decoding of higher cognitive activities, Benedek et al. studied attention decoding in 2019; Horikawa et al. explored imagination decoding in 2017; and Horikawa et al. also studied dream decoding in 2013.

[0005] The aforementioned studies primarily focus on decoding brain information in a single task, and can be broadly categorized into main category decoding, multi-semantic label decoding, language decoding, and reconstruction decoding. Although single-task decoding often yields good results, it is limited to a single task and fails to fully utilize the efficiency and potential for widespread application brought about by multi-task joint decoding.

[0006] In the field of artificial intelligence, it has been proven that jointly modeling multiple tasks can significantly improve the efficiency of processing these tasks. For example, in computer vision, in 2022, Wu et al. designed a multi-task network called YOLOP, which can simultaneously handle three key driving perception tasks: object detection, drivable area segmentation, and lane detection. While ensuring efficient parallel processing, it also significantly improved the performance of these three tasks. Furthermore, in natural language processing, cutting-edge research like ChatGPT, developed in 2017 by Vaswani et al., integrated Transformer architecture, GPT, Reinforcement Learning from Human Feedback, and Chain-of-Thought Prompt to achieve multi-task learning. This covers multiple tasks such as machine translation and question-answering systems, bringing breakthroughs to the field of artificial intelligence and driving the intelligent upgrading of society.

[0007] However, in research on brain information decoding, a single decoding model is often only applicable to a specific task. To achieve multi-level decoding, multiple models need to be constructed. This results in insufficient correlation and consistency between models, thus hindering their widespread application in practical applications such as brain-computer interfaces or neuromorphic chips. Summary of the Invention

[0008] This invention addresses the limitations of single-task decoding in existing brain information decoding models. To solve the problems of limited model application scope and increased difficulty in designing and managing models separately for various tasks, this invention proposes a multi-task synchronous decoding method for human brain activity evoked by natural images.

[0009] The technical solution adopted in this invention is as follows:

[0010] A multi-task synchronous decoding method for human brain activity evoked by natural images performs the following steps in a multi-task visual encoding / decoding model (i.e., a multi-task visual information brain decoding model based on an encoding / decoding framework), comprising a visual encoder E1, a multi-task encoder E2, a category decoder D1, a semantic decoder D2, and a language decoder D3:

[0011] Step 1: Using the visual encoder E1, the image information obtained from the fMRI image data of measuring the BOLD response signal based on magnetic resonance imaging when the test subject views natural images is embedded into the latent feature space to obtain visual information feature vectors of several different visual areas.

[0012] Step 2: The multi-task encoder E2 is used to concatenate the visual information feature vector, category information feature vector, and semantic decoding task feature vector before extracting multi-task features; where the category information feature vector refers to the image category keywords of the natural image viewed by the tester, and the semantic decoding task refers to the semantic labels of the natural image.

[0013] Step 3: Use the classification decoder D1 to classify and identify the natural image categories based on the multi-task features, and output the probability of each image category of the natural image.

[0014] Step 4: Use the semantic decoder D2 to predict the semantic labels of natural images from the multi-task visual feature vectors, and output the probability of each semantic label of the natural image.

[0015] Step 5: Use the language decoder D3 to predict the text description words of the natural image from the multi-task visual feature vector to generate a continuous text description of the natural image.

[0016] Furthermore, in step 1, the encoding method of the visual encoder E1 includes:

[0017] Step 1.1: Select Region of Interest (ROI) from the input BOLD response signal and fMRI image data. Each ROI is considered a visual region (i.e., each visual region represents a specific area of ​​brain activity), and extract the activity signal (BOLD response signal and fMRI image data) of each visual region at a given time point, resulting in several visual region signal data. The fMRI image data is generated by mapping BOLD signal measurements into three-dimensional space. These images can be used to visualize the distribution of brain activity under different tasks or conditions. Researchers can use fMRI images to identify active regions, analyze connections between different brain regions, and study brain networks, etc.

[0018] Upsampling or downsampling is performed on signal data from several different visual regions using linear interpolation to unify the data dimension to the embedding space dimension. Then, the signals from different visual regions are sorted according to the visual region number to form a T×M visual region feature sequence (V1, V2, ..., V...). T ), where T represents the number of visual regions selected, and M represents the feature vector dimension of each visual region;

[0019] Step 1.2: Extract the visual region feature sequence (V1, V2, ..., V...). T The data is fed into a bidirectional gated recurrent unit (BiGRU) for processing to obtain the updated M-dimensional visual information feature vector (F1, F2, ..., F) at each time point. T ).

[0020] Furthermore, in step 2, the encoding method of the multi-task encoder E2 includes:

[0021] Step 2.1: Embed the category information and the information from the semantic decoding task into two different vectors E, respectively. [CLS] and E [SMT] In the middle; among them, E [CLS] E represents the feature vector of category information. [SMT] The feature vector representing the semantic decoding task;

[0022] Step 2.2: E [CLS] Vector and E [SMT] Vectors and multidimensional visual information feature vectors (F1, F2, ..., F T The features are concatenated to form a comprehensive feature vector with a dimension of (T+2)×M;

[0023] Step 2.3: Add position embedding (i.e., the position index of the feature vector; the specific embedding position can be defined by yourself, mainly used to determine the position of the visual information feature vector, the category information feature vector, and the feature vector of the semantic decoding task) to the new feature vector obtained in step 2.2. This helps the model learn the positional dependencies in the sequence, thereby obtaining a visual feature vector with positional encoding.

[0024] Step 2.4: Feed the position-encoded visual feature vector obtained in Step 2.3 into the language representation model BERT, and obtain the (T+2)×M dimensional multi-task visual feature vector (Z1, Z2, ..., Z...) based on its output. T+2 ).

[0025] Furthermore, in step 3, the decoding method of the classification decoder D1 includes:

[0026] Step 3.1: Construct two hidden layers using a multi-layer perceptron (MLP);

[0027] Step 3.2: Apply the Leaky-ReLU activation function and Layer Normalization to all internal neurons of the hidden layer constructed in Step 3.1 to help improve the generalization ability of the model and make the network performance more reliable and stable.

[0028] Step 3.3: Construct several neurons in the output layer of the decoder, each neuron corresponding to a category, and use the activation function Softmax to perform a non-linear transformation to obtain the classification prediction of the classification decoder D1; using the Softmax function can help reduce the saturation effect and complete multi-class classification tasks.

[0029] Furthermore, the classification decoder D1 uses cross-entropy loss as its class loss function during training to make the multi-label classification problem more stable and less susceptible to outliers, allowing the model to converge to the optimal solution more quickly.

[0030] Furthermore, in step 4, the decoding method of semantic decoder D2 includes:

[0031] Step 4.1: Construct two hidden layers using MLP;

[0032] Step 4.2: Apply the Leaky-ReLU activation function and Layer Normalization to all neurons inside the hidden layer constructed in Step 4.1;

[0033] Step 4.3: Construct multiple neurons in the decoder output layer, each neuron corresponding to a semantic label, and perform a non-linear transformation using the activation function Softmax.

[0034] Furthermore, the semantic decoder D2 uses cross-entropy loss as its semantic label loss function during training.

[0035] Furthermore, in step 5, the decoding method of the language decoder D3 includes:

[0036] Step 5.1: Convert the multi-task visual feature vector (Z1, Z2, ..., Z...) T+2 ) is mapped to an embedding vector E(Z) i Then embed the vector E(Z) i The first new feature vector is obtained by adding the position embedding Pos(i) to the feature vector.

[0037] In this step, the Embedding model (a machine learning method for mapping high-dimensional data (such as text, images, and audio) to a low-dimensional space) can be used to map multi-task visual feature vectors to embedding vectors E(Z). i ).

[0038] Step 5.2: The first new feature vector is processed sequentially through a masked multi-head attention layer and a multi-head attention layer to obtain the second new feature vector;

[0039] In this process, both the masked multi-head attention mechanism layer and the multi-head attention mechanism layer have residual connections, and the residual results are normalized. That is, the input and output of the masked multi-head attention mechanism layer are superimposed through the Add layer, and the superimposed result is normalized before being fed into the multi-head attention mechanism layer. The input and output of the multi-head attention mechanism layer are superimposed through the Add layer, and the superimposed result is normalized to obtain the second new feature vector.

[0040] Using Multi-Head Attention helps learn rich contextual information, capture deep-level combinations and semantic relationships in the text, and achieve better generalization ability;

[0041] Step 5.3: Pass the second new feature vector through a feedforward neural network (FFN) with residual connections, and normalize the residual results of the feedforward neural network to obtain the third new feature vector;

[0042] For a feedforward neural network, its output can be expressed as FFN(x) = max(0, xW′1+b1)W′2+b2, where x represents the input of the feedforward neural network, W′1 and W′2 represent the two weight matrices of the feedforward neural network, and b1 and b2 represent the two bias terms of the feedforward neural network.

[0043] Step 5.4: Pass the third new feature vector to a linear layer, and then perform a softmax operation to predict the probability distribution of the next word to guide text generation;

[0044] The probability distribution of the next word is expressed as follows:

[0045]

[0046] Wherein, P(word) next =word j |Z) is the next word given the feature Z (output of the linear layer). j The probability, W j It is related to the word "word" j The relevant weights, b j It is related to the word "word" j The relevant biases are LayerOutput, which is the output of the last layer of the Transformer network, namely the output of the linear layer of the language decoder D3, and N is the total number of words in the vocabulary of the GPT model (a natural language processing technique that can model natural language text and generate natural language outputs similar to the input text).

[0047] Furthermore, the training process of the multi-task visual encoding / decoding model includes:

[0048] Step 6.1: Collect the training set for the model;

[0049] The training set includes visual activity information from multiple cerebral cortexes obtained based on fMRI measurements, image categories of stimulus images, and manually annotated semantic labels.

[0050] Step 6.2: Based on the visual encoder E1 and the multi-task encoder E2, transform the visual activity into a multi-task visual feature vector (Z1, Z2, ..., Z). T+2 );

[0051] Based on the outputs of category decoder D1, semantic decoder D2, and language decoder D3, multi-task feature prediction of image category and semantic information, and word-by-word generation of continuous text;

[0052] Step 6.3: Calculate the cross-entropy loss (first cross-entropy loss) of category decoder D1 based on the predicted image category and category label; calculate the cross-entropy loss (second cross-entropy loss) of semantic decoder D2 based on the predicted semantic information and semantic label; calculate the cross-entropy loss (third cross-entropy loss) of language decoder D3 based on the predicted context distribution and label text.

[0053] The total loss function L of the multi-task visual encoding / decoding model is obtained by weighted fusion of the first, second, and third cross losses.

[0054] Step 6.4: Based on the total loss function L, use an optimization algorithm (such as the AdamW algorithm) to iteratively update the model parameters of the multi-task visual encoding and decoding model until the preset training convergence conditions (training times or total loss function value convergence) are met.

[0055] Furthermore, in step 1, the step size of the bidirectional gated recurrent unit is T (the number of selected visual areas of the brain), the number of layers is 1, the input layer size is M-dimensional, and the output layer size is M-dimensional.

[0056] Furthermore, in step 5, the mask multi-head attention mechanism layer and the multi-head attention module used in the multi-head attention mechanism layer have a head count of 8.

[0057] Furthermore, step 6.4 specifically includes:

[0058] (1) Initialize parameters, including learning rate α, two decay coefficients β1 and β2, parameter ∈ to prevent the denominator from being zero, first momentum v and second momentum u, time step t, and convergence condition;

[0059] (2) Iteratively update the model parameters based on the total loss function L:

[0060] Calculate the gradient g based on the total loss function value L(f(x;θ),y): Where L() represents the value of the total loss function L, f(x; θ) represents the output of the multi-task visual encoding and decoding model, i.e. the model prediction result, x represents the input data of the multi-task visual encoding and decoding model, θ represents the model parameters, and y represents the label;

[0061] Update time step: t = t + 1;

[0062] Calculate the first momentum: v = β1 × v + (1 - β1) × g;

[0063] Calculate the second momentum: u = β² × u + (1 - β²) × g 2 ;

[0064] Correcting the first momentum yields the first correction value:

[0065] By correcting the second momentum, we obtain the second correction value:

[0066] Calculate the update amount of the model parameters:

[0067] Update model parameters: θ = θ + Δθ

[0068] (3) Determine whether the convergence condition is met. If yes, stop; otherwise, continue with step (2).

[0069] The technical solution provided by this invention brings at least the following beneficial effects:

[0070] This invention establishes a multi-task visual information brain decoding model based on an encoding and decoding framework. It performs classification decoding, semantic decoding, and language decoding on color images of complex natural scenes. The decoded category information has high accuracy, and it also shows high accuracy on most semantic labels. Furthermore, all words in the decoded language strongly point to the main elements or events in the image. Attached Figure Description

[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0072] Figure 1 This is a schematic diagram of the model structure of the present invention;

[0073] Figure 2This is a schematic diagram of the classification and decoding results of the present invention;

[0074] Figure 3 This is a schematic diagram of the semantic decoding result of the present invention;

[0075] Figure 4 This is a schematic diagram of the language decoding results of this invention. Detailed Implementation

[0076] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0077] This invention discloses a multi-task synchronous decoding method for human brain activity induced by natural images. It integrates deep learning techniques to establish a multi-task visual decoding model based on an encoder-decoder framework, enabling in-depth interpretation of human visual information. The method is based on functional magnetic resonance imaging (fMRI) signal data from viewing a large number of natural images as input, and establishes a multi-task visual information brain decoding model based on an encoder-decoder framework. This model includes five modules: a visual encoding module, a multi-task encoding module, a category decoding module, a semantic decoding module, and a language decoding module. The visual encoding module uses a BIGRU (Bidirectional Recurrent Neural Network) to encode voxel signals from visually relevant regions into a latent feature space. The multi-task encoding module first concatenates the embedding vectors of category and semantic information with the feature vectors of visual information, and then inputs them together into a language representation model BERT (Bidirectional Encoder Representations from Transformers) to obtain multi-task features. The category decoding module inputs the obtained multi-task features into a layered normalization model constructed using an MLP (Multi-Layer Perceptron) with the Leaky-ReLU activation function and the Layer normalization technique. The MLP uses two hidden layers, one for Normalization (the Norm layer) and the other for Softmax activation, to obtain the probability distribution of predicted scene categories (the probability distribution of image categories in the viewed natural images). The semantic decoding module inputs the feature vectors at specified positions in the obtained multi-task feature vectors (the feature vectors corresponding to the semantic positions) into the two hidden layers to obtain the probability distribution of predicted semantic labels. The language decoding module uses the deep self-attention mechanism in GPT (Generative Pre-Trained Transformer) to capture the deep structure and semantic relationships in the text, thereby generating more accurate continuous descriptive text.

[0078] As one possible implementation, such as Figure 1 As shown in the embodiments of the present invention, the multi-task synchronous decoding method for human brain activity evoked by natural images provided by the present invention includes the following steps:

[0079] A. Visual encoding module:

[0080] Step A1: Perform upsampling or downsampling operations on signal data from 10 different visual regions using linear interpolation to unify the data dimension to the embedding space dimension. Then, sort the signals from different visual regions according to the visual region number to form a 10*1024 visual region feature sequence (V1, V2, ..., V...). T The number of visual regions can be set based on the actual application scenario, and the embedding space dimension can be customized based on the actual application scenario. The subscript T represents the number of visual regions. In this embodiment, T = 10.

[0081] Step A2: Input the T*1024 visual features obtained in Step A1 into the BiGRU bidirectional gated recurrent unit. The BiGRU integrates the feature vector sequence into a single feature vector, and takes the T*1024-dimensional hidden features from the final stage of the BiGRU as the visual region feature vector (F1, F2, ..., F...). T );

[0082] Steps A1 and A2 convert the visual region signal into a latent feature space vector.

[0083] B. Multi-task coding module:

[0084] Step B1: Convert the 1024-dimensional visual feature vector (F1, F2, ..., F...) obtained in Step A1 into a multidimensional array. T ) and category information embedding vector E [CLS] Semantic information embedding vector E [SMT] Concatenating these elements yields a (T+2)*1024-dimensional feature vector. Among these, the category information embedding vector E... [CLS] and semantic information embedding vector E [SMT] For a custom identifier, its vector dimension is the same as F. i (i=1,…,T) are consistent.

[0085] Step B2: Based on the 12*1024 dimensional feature vector obtained after concatenation in Step B1, positional embedding is added. This helps the model learn the positional dependencies in the sequence, resulting in a 12*1024 dimensional multi-task visual feature vector (Z1, Z2, ..., Z...). T+2 ).

[0086] C. Category Decoding Module:

[0087] Step C1: Calculate the 12*1024 dimensional multi-task feature vector (Z1, Z2, ..., Z2) obtained in step B2. T+2 The 1024-dimensional feature vector Z1 in the MLP is input into the first hidden layer. Through the activation function Leaky-ReLU and the normalization technique Layer Normalization, all parameters are guaranteed to be continuously updated, avoiding the problem of parameter update stagnation when the activation value is less than 0. Furthermore, through normalization, the scale problem in the parameter update process is reduced, and the parameter update may be made more stable.

[0088] Step C2: Input the 1024-dimensional feature vector obtained in step C1 into the second hidden layer constructed by the MLP. At the output, use the activation function Softmax to perform a non-linear transformation to obtain the probability distribution of the 12 predicted categories.

[0089] D. Semantic Decoding Module:

[0090] Step D1: Calculate the 12*1024 dimensional multi-task feature vector (Z1, Z2, ..., Z2) obtained in step B2. T+2 The 1024-dimensional feature vector Z2 in the model is input into two hidden layers constructed by the MLP to ensure the stability of parameter updates and the uniformity of parameter scale.

[0091] Step D2: Input the 1024-dimensional feature vector obtained in step D1 into the second hidden layer constructed by MLP. At the output, use the activation function Softmax to perform a non-linear transformation to obtain the probability distribution of the predicted 80 semantic labels.

[0092] E. Language decoding module:

[0093] Step E1: Calculate the 12*1024 dimensional multi-task feature vector (Z1, Z2, ..., Z...) obtained in step B. T+2 The 10*1024 dimensional feature vector (Z3, Z4, ..., Z) T+2 First, it is mapped to an embedding vector E(Z). i ).

[0094] Step E2: Input the embedding vector obtained in step E1 into the multi-head attention mechanism of GPT for computation, and then pass it through a feedforward neural network. There is a residual connection between the feedforward neural network and the attention mechanism.

[0095] Step E3: Perform a softmax operation on the output of the last layer to predict the probability distribution of the next word, thereby guiding text generation. The formula is as follows:

[0096]

[0097] Wherein, P(word) next =w j |Z) is the next word given feature Z, where |Z) is w. j (i.e., the word) j The probability of ) W j It is related to the word "word" j The relevant weights, b j It is related to the word "word" j The relevant biases are LayerOutput, which is the output of the last layer of the Transformer network, and N is the total number of words in the GPT vocabulary.

[0098] Step E4: Repeat the above steps to generate continuous text.

[0099] F. Training Phase:

[0100] Step F1: First, obtain visual activity information from multiple cortical regions measured by fMRI, including the main categories of stimulus images and manually labeled semantic tags, divided into training and test sets. The cortical regions include: V1, V2, V3, OFA, PPA, OPA, VWFA, FBA, FFA, and EBA.

[0101] Step F2: Use a visual encoder and a multi-task encoder to transform visual activity into multi-task features, and then predict the main category, semantic information, and word-by-word continuous text of the image based on the multi-task features.

[0102] Step F3: Calculate the cross-entropy loss based on the predicted category and category label; calculate the cross-entropy loss based on the predicted semantic information and semantic label; calculate the cross-entropy loss based on the predicted context distribution and label text.

[0103] Step F4: Based on the loss function L, use the AdamW optimization algorithm to update the corresponding weights of the parameters of the entire model.

[0104] The specific optimization parameters for AdamW are as follows: the learning rate α is initialized to 0.0001; the momentum term decay coefficients β1 and β2 are set to 0.88 and 0.98 respectively; and the minimum value ∈ (to prevent the denominator from being zero) is set to 10. -9 The first and second momentum are initialized as v = 0, u = 0; the time step is initialized as t = 0; each time there are n dataset samples {(x1, y1), (x2, y1)}. 12 ),...,(x n ,y n The following updates will be made:

[0105] Gradient calculation: Where L(f(x;θ),y) represents the loss function. The gradient of the loss function is represented by f(x; θ), which represents the output of the model, i.e., the model's prediction result. x represents the input data of the model, θ represents the model parameters, and y represents the label.

[0106] Time step update: t = t + 1;

[0107] Calculate the first momentum: v = β1 × v + (1 - θ1) × g;

[0108] Calculate the second momentum: u = β² × u + (1 - β²) × g 2 ;

[0109] Correction for the first momentum:

[0110] Correction for the second momentum:

[0111] Update parameters: That is, Δθ represents the amount of model parameter update;

[0112] Application update: θ = θ + Δθ.

[0113] G. Testing Phase:

[0114] Step G1: Collect test data. Each natural image (visual stimulus in the experiment) has multiple labels from the COCO dataset, including a primary category, multiple labels, and five text descriptions. The primary category is a human annotation corresponding to the natural image. In the annotation information of the COCO dataset, the primary category includes 12 "super categories," such as "people," "vehicles," "outdoors," etc. The multiple labels are "names" of multiple human annotations for the natural images. The multiple labels include a total of 80 "names," such as "bicycle," "stop sign," "horse," etc., as well as visual cortex signals evoked by each stimulus image based on the HCP-MMP1 atlas, including 10 cortical regions: V1, V2, V3, EBA, FBA, FFA, OFA, OPA, PPA, and VWFA. This embodiment of the invention yielded a training set containing 24,980 samples and a test set containing 2,770 samples. Each sample includes the following five components:

[0115] (1) Image - A natural image.

[0116] (2) Categories - The main categories of natural images.

[0117] (3) Semantic - Multiple labels for natural images (from 80 labels).

[0118] (4) Sentences - Five descriptive sentences for natural images.

[0119] (5) Visual activity - ten response patterns from ten regions of interest (ROIs).

[0120] Step G2: Based on the training phase, a model is obtained that includes a visual encoding module, a multi-task encoding module, a category decoding module, a semantic decoding module, and a language decoding module. The voxel signal of the visual region corresponding to the test image is input into the visual encoding module to obtain visual features. Then, the visual features are concatenated with category and semantic information and fed into the multi-task encoder to obtain multi-task features. Finally, the multi-task features are input into the category decoder, semantic decoder, and language decoder, respectively. The model will then generate category, semantic labels, and continuous descriptive text corresponding to the visual stimuli of the natural scene color image. This process achieves the synchronous decoding of category, semantic, and language evoked by natural images in human brain activity, and the decoding results are as follows: Figure 2 , Figure 3 and Figure 4 As shown, this embodiment achieves a decoding accuracy of nearly 70% for 12 major categories in natural images, significantly exceeding the chance level of 8.33%. The decoding accuracy for 80 detailed semantic tags reaches 20%, representing a 16-fold improvement over the chance level of 0.0125. In the language decoding task, this embodiment generates continuous text language, achieving scores of 0.1922, 0.7986 and 0.1909, 0.5287, 0.8020 and 0.8612 on six evaluation metrics: BLEU, CIDEr, ROUGE, WCS, GCS and FTCS, respectively, all exceeding the corresponding random levels (0.1421, 0.5843 and 0.1439, 0.4126, 0.7557 and 0.8342).

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0122] The above descriptions are merely some embodiments of the present invention. Those skilled in the art can make various modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the scope of protection of the present invention.

Claims

1. A multi-task synchronous decoding method for human brain activity evoked by natural images, characterized in that, In a multi-task visual encoder-decoder model comprising a visual encoder , a multi-task encoder , a category decoder , a semantic decoder , and a language decoder the following steps are performed: Step 1, using visual encoder Embedding image information acquired in fMRI image data of measured BOLD response signal based on magnetic resonance imaging when the tester watches natural images into a hidden feature space to acquire visual information feature vectors of several different visual areas; Step 2, using a multi-task encoder After feature splicing of the visual information feature vector, the category information feature vector and the feature vector of the semantic decoding task, multi-task features are extracted; wherein the category information feature vector refers to an image category keyword of a natural image viewed by the tester, and the semantic decoding task refers to a semantic label of the natural image. Step 3, using a category decoder The multi-task feature is used to classify and identify the natural image category, and the probability of each image category of the natural image is output. Step 4, using a semantic decoder perform semantic label prediction on the multi-task visual feature vector to output probabilities of each semantic label of the natural image; Step 5, employing a language decoder performing text description word prediction on the multi-task visual feature vector for the natural image to generate a continuous text description of the natural image; in, In step 1, the visual encoder The encoding method of the present application comprises: Step 1.1: Select the region of interest from the input BOLD response signal and fMRI image data. Each region of interest is regarded as a visual region. At a given time point, acquire the BOLD response signal and corresponding fMRI image data of each visual region to obtain several visual region signal data. The data dimension is unified to the embedding space dimension based on a linear interpolation method, and the obtained signals of different visual areas are sorted according to the visual area number to form a visual area feature sequence of T*M (T*M visual area feature sequence) , ), wherein T represents the number of selected visual areas, and M represents the feature vector dimension of each visual area. Step 1.2: Sequence of visual region features ( , The data is fed into a bidirectional gated recurrent unit for processing to obtain the updated M-dimensional visual information feature vector at each time point. ); In step 2, the multi-task encoder The encoding methods include: Step 2.1: Embed the category information and the information from the semantic decoding task into two different vectors respectively. and In; among them, Represents category information feature vectors, The feature vector representing the semantic decoding task; Step 2.2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] Vector sum Vectors and multidimensional visual information feature vectors The features are concatenated to form a comprehensive feature vector with a dimension of (T+2)×M; Step 2.3: Add position embeddings to the new feature vector obtained in step 2.2 to obtain a visual feature vector with position encoding; Step 2.4: Feed the position-encoded visual feature vector obtained in Step 2.3 into the language representation model BERT, and obtain a (T+2)×M dimensional multi-task visual feature vector based on its output. ; Where T represents the number of visual regions selected, and M represents the feature vector dimension of each visual region.

2. The method as described in claim 1, characterized in that, The bidirectional gated loop unit has a step size of T, a layer size of 1, an input layer size of M dimensions, and an output layer size of M dimensions.

3. The method as described in claim 1, characterized in that, In step 3, the classification decoder Decoding methods include: Step 3.1: Construct two hidden layers using a multilayer perceptron; Step 3.2: Apply activation functions and normalization techniques to all internal neurons of the hidden layer constructed in Step 3.1; Step 3.3: Construct the output layer of the decoder with several neurons, each neuron corresponding to a class, and use an activation function. The probabilities of each category after nonlinear transformation are used to obtain the classification decoder. Classification prediction.

4. The method as described in claim 1, characterized in that, In step 4, the semantic decoder Decoding methods include: Step 4.1: Construct two hidden layers using MLP; Step 4.2: Apply activation functions and normalization techniques to all neurons within the hidden layer constructed in Step 4.1; Step 4.3: Construct multiple neurons in the decoder output layer, each neuron corresponding to a semantic label, and use an activation function. Perform nonlinear transformation.

5. The method as described in claim 1, characterized in that, In step 5, the language decoder Decoding methods include: Step 5.1: Map the multi-task visual feature vectors to an embedding vector Then embed the vector With position embedding The summation yields the first new eigenvector; Step 5.2: The first new feature vector is processed sequentially through the masked multi-head attention mechanism layer and the multi-head attention mechanism layer to obtain the second new feature vector; wherein, both the masked multi-head attention mechanism layer and the multi-head attention mechanism layer are configured with residual connections, and the residual results are normalized. Step 5.3: Pass the second new feature vector through a feedforward neural network with residual connections, and normalize the residual results of the feedforward neural network to obtain the third new feature vector. Step 5.4: Pass the third new feature vector to a linear layer, and then through one... The operation is used to predict the probability distribution of the next word.

6. The method as described in claim 4, characterized in that, The masked multi-head attention mechanism layer in step 5 uses a multi-head attention module with 8 heads.

7. The method as described in claim 1, characterized in that, The training process of the multi-task visual encoding and decoding model includes: Step 6.1: Collect the training set for the model; The training set includes visual activity information from multiple cerebral cortexes obtained based on fMRI measurements, image categories of stimulus images, and manually annotated semantic labels; Step 6.2: Based on the visual encoder and multi-task encoder Transform visual activities into multi-task visual feature vectors; Based on category decoder Semantic decoder and language decoder The output obtains multi-task feature prediction of image category, semantic information, and word-by-word generation of continuous text; Step 6.3: Calculate the first cross-entropy loss based on the predicted image category and category label; calculate the second cross-entropy loss based on the predicted semantic information and semantic label; calculate the third cross-entropy loss based on the predicted context distribution and label text; The total loss function L of the multi-task visual encoding and decoding model is obtained by weighted fusion of the first, second and third cross losses; Step 6.4: Based on the total loss function L, use an optimization algorithm to iteratively update the model parameters of the multi-task visual encoding and decoding model until the preset training convergence condition is met.

8. The method as described in claim 7, characterized in that, Step 6.4 specifically includes: (1) Initialize parameters, including learning rate Two attenuation coefficients , Parameters to prevent the denominator from being zero First Momentum Second Momentum The time step t and the convergence condition; (2) Iteratively update the model parameters based on the total loss function L: Based on the total loss function value Calculate gradient : ,in, This represents the function value of the total loss function L. This represents the output of the multi-task visual encoding / decoding model, i.e., the model's prediction result. This represents the input data for a multi-task visual encoding / decoding model. Indicates model parameters, Indicates a label; Update time step: ; Calculate the first momentum: ; Calculate the second momentum: ; Correcting the first momentum yields the first correction value: ; By correcting the second momentum, we obtain the second correction value: ; Calculate the update amount of the model parameters: ; Update model parameters: (3) Determine whether the convergence condition is met. If yes, stop; otherwise, continue with step (2).