A learner concentration detection method and device based on MRI images and a medium
By employing an MRI-based learner attention detection method, utilizing the Local Global Feature Aggregator (LGA) and the Cross-Attention Mechanism (MAB), combined with a brain-based captioning generation and visual stimulus reconstruction model, the accuracy problem of learner attention detection in existing technologies has been solved, achieving a more reliable and accurate detection result.
Patent Information
- Application Number
- CN202510410941.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-04-02
AI Technical Summary
In existing technologies, learner attention detection methods rely on behavioral observation or subjective questionnaires, which suffer from response delays and measurement biases. Furthermore, functional magnetic resonance imaging (fMRI) technology has failed to effectively combine neural decoding and cognitive assessment, making it difficult to achieve accurate detection.
This study employs an MRI-based method for detecting learner attention. By acquiring learners' fMRI brain signals, it utilizes a Local Global Feature Aggregator (LGA) and a Cross-Attention Mechanism (MAB), combined with brain caption generation and visual stimulus reconstruction models, to achieve multimodal feature fusion and weight optimization, ultimately detecting learners' attention levels.
It significantly improves the objectivity and accuracy of learner attention detection. Through multimodal feature fusion and weight allocation, it enhances the reliability and accuracy of detection, helping researchers understand learner status and improve teacher-student communication.
Smart Images

Figure CN120298842B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method, device and medium for detecting learner attention based on MRI images. Background Technology
[0002] Magnetic resonance imaging (MRI), as an important non-invasive diagnostic tool, is widely used in the diagnosis and monitoring of diseases of the nervous system, tumors, and cardiovascular diseases due to its excellent soft tissue contrast and multiplanar imaging capabilities. Especially in the field of neuroimaging, MRI can clearly display brain structure and function, helping doctors to identify lesions and formulate treatment plans in a timely manner. However, despite the significant advantages of MRI in medical imaging, its application in the field of education is not yet widespread. There are few methods to judge the learner's state based on MRI, which limits researchers from judging the learner's state based on direct brain signals.
[0003] Traditional methods for detecting learners' attention rely primarily on behavioral observation or subjective questionnaires, which suffer from response delays and measurement biases. In recent years, while functional magnetic resonance imaging (fMRI) can capture the spatiotemporal characteristics of brain activity, existing technologies often separate neural decoding from cognitive assessment, leading to a disconnect between feature extraction and behavioral interpretation. For example, fMRI-based visual reconstruction models lack a direct mapping to cognitive states, while attention detection models neglect the multimodal representation of neural activity. Therefore, this separation in existing technologies limits the interpretable correlation between brain signals and cognitive states, making it difficult to achieve accurate detection of learners' attention. Summary of the Invention
[0004] This invention provides a method, device, and medium for detecting learner attention based on MRI images, which can solve the problem of difficulty in accurately detecting learner attention in the prior art.
[0005] This invention provides a method for detecting learner attention based on MRI images, comprising the following steps: Acquire fMRI signals of the learner's brain while viewing original visual stimulus images; The brain fMRI signal is input into the attention detection model, which includes a feature extraction module and a stimulus reconstruction module; The feature extraction module is used to extract features from the brain fMRI signal to obtain fMRI features associated with the learner's learning process and fMRI features associated with the learner's cognitive activities. The fMRI features associated with the learner's cognitive activities are then converted into brain subtitles representing cognitive activities. The stimulus reconstruction module employs the cross-attention mechanism MAB, assigning different weights to brain captions representing cognitive activities, fMRI features associated with the learner's learning process, and the original visual stimulus image. These different weights guide the utilization of brain captions, fMRI features, and the original visual stimulus image during visual stimulus image reconstruction, resulting in the reconstructed visual stimulus image. Learners’ attention levels are detected based on the similarity between the reconstructed visual stimulus image and the original visual stimulus image.
[0006] Preferably, the acquisition of fMRI features associated with the learner's learning process specifically involves feature extraction using a Local Global Feature Aggregator (LGA), including: The Local Global Feature Aggregator (LGA) consists of the Efficient Relation Aggregator (ERA) and a feedforward neural network. The Effective Relation Aggregator (ERA) employs multi-head fusion learning, as shown below: ; ; in: Indicates the first Effective relation aggregator for individual heads; matrix Indicates integration The learnable parameter matrix of the size; Local relation aggregator and global relationship aggregator The system consists of shallow convolutional networks with local relation aggregators that focus on local neighborhood features to generate a learnable parameter matrix. Its formula is: ; in: Represents a token. Represents local neighborhood Any token; Represents the learnable parameter matrix; A global relationship aggregator is established by calculating the similarity of global context features. Global Relation Aggregator Represented as: ; in: and All are linear transformation functions; the aggregated features are input into a feedforward neural network for further processing to obtain... for: ; in: FFN Indicates feedforward layer; Indicates a linear layer; Finally, the three features are aggregated through residual connections to form fMRI features associated with the learner's learning process. , is represented as: .
[0007] Preferably, the generation of the brain-based subtitles includes: The Local Global Feature Aggregator (LGA) was used to extract features from brain fMRI signals to obtain fMRI features associated with learners' cognitive activities. The fMRI features associated with learners' cognitive activities are represented using a text feature extractor, and the representation equation is as follows: ; Using a text feature encoder This is transformed into a series of trainable fMRI text query features corresponding to relevant visual and semantic information. The transformation equation is: ; Where: FFN represents a feedforward neural network; MAB represents a cross-attention mechanism; SA represents a self-attention mechanism; Brain captions were generated from fMRI text query features using a text decoder.
[0008] Preferably, the process by which the stimulus reconstruction module reconstructs the visual stimulus image is a reverse generation process within the reverse diffusion process of the Conditional Diffusion Model (ELDM), including: In the reverse diffusion process, it is assumed As the initial input, we can obtain the following from Bayes' theorem: ; In order to rebuild , use reduction and Training the generative network using Euclidean distance between them ;but From this, it can be inferred that ,but: ; in: express ; The generation process can then be rewritten using a formula as follows: ; Its loss function is expressed as: .
[0009] Preferably, the reconstruction of the visual stimulus image includes: After adding the conditional mechanism, the distribution is represented as follows: ; in: This represents the fMRI features extracted by the encoder. ; and These represent the conditions after the addition of the conditional mechanism. The time step and the first Reconstructed image generated after one time step; when and When they are very close, the value will be significantly greater than 0; the Taylor expansion of the diffusion model after adding the conditional mechanism is expressed as: ; in: Represents gradient change; assuming ,but: ; in: Indicates direct proportion; express Noise at any given moment; that is Approximately After adding the conditional mechanism, it is represented as: ; By continuously advancing the time step, the reconstructed image with the addition of a conditional mechanism is finally obtained. .
[0010] This invention also provides an electronic device, including a memory and a processor; The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the steps of the learner attention detection method based on MRI images as described above.
[0011] This invention also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of a learner attention detection method based on MRI images as described above.
[0012] This invention provides a method, device, and medium for detecting learner attention based on MRI images. Compared with existing technologies, its advantages are as follows: This invention generates brain captions related to cognitive activity based on fMRI signals. In the visual stimulus reconstruction part, it introduces a three-condition integrated guidance diffusion model for the reverse generation process. In addition to the stimulus signals and fMRI latent features in the traditional visual stimulus image, it adds guidance conditions for brain-generated captions. The brain captions representing the learner's cognitive activity, the original stimulus image, and the fMRI latent features representing the learner's learning process are used as conditions to guide image reconstruction. These three different conditions are weighted by cross-attention calculation to optimize the final image generation result. That is, the weights indicate which information should be strong and which should be weak under different conditions, and the optimal information fusion is achieved according to the degree of strength to obtain the reconstructed visual stimulus image. Learning attention is detected based on the similarity between the reconstructed visual stimulus image and the original visual stimulus image. The similarity indicates the degree of change in the learner's cognitive activity and learning process during the learning process. Thus, through multimodal feature fusion and weight allocation, the objectivity and accuracy of learner attention detection are significantly improved. Attached Figure Description
[0013] Figure 1 A schematic diagram of the overall process of a learner attention detection method based on MRI images provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall structure of a model built for a learner attention detection method based on MRI images, as provided in an embodiment of the present invention. Figure 3 A schematic diagram of the Intellect-vis two-stage neurovisual decoding framework structure constructed for a learner attention detection method based on MRI images provided in an embodiment of the present invention; Figure 4 This is a schematic diagram comparing the signal features of LGA before and after reconstruction in a learner attention detection method based on MRI images provided in an embodiment of the present invention; where the left side is the initial input signal, the middle side is the signal representation after masking, and the right side is the reconstructed signal; Figure 5 A schematic diagram of the structure of the Local Global Aggregator constructed by a learner attention detection method based on MRI images provided in an embodiment of the present invention; Figure 6 A schematic diagram of the MAB module architecture constructed for a learner attention detection method based on MRI images provided in an embodiment of the present invention; Figure 7 This is a temporal mapping diagram of a learner attention detection method based on MRI images provided in an embodiment of the present invention. Detailed Implementation
[0014] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0015] See Figure 1 This invention provides a learner attention detection method based on MRI images. Specifically, it constructs an end-to-end deep learning model specifically designed to interpret the brain activity of learners when receiving external visual stimuli. In the visual stimulus reconstruction part of the model, a three-condition integrated guidance diffusion model reverse generation process is introduced for the first time. In addition to traditional visual stimuli and fMRI feature extraction, this invention adds guidance conditions for brain-generated subtitles. This innovation significantly improves the quality of stimulus reconstruction, making the reconstructed visual stimuli more reliable. Furthermore, by ensuring the consistency of meaning between the subtitle text and the reconstructed image, this invention assigns different weights to the conditions through cross-attention to achieve comprehensive consideration, effectively avoiding potential contradictions between the two, thereby enhancing the effectiveness and reliability of learner attention detection. This comprehensive method not only promotes the development of image reconstruction technology but also provides a clearer basis for understanding brain activity, laying an important foundation for subsequent research.
[0016] I. Precursor conditions for the brain interpretation model.
[0017] When building a brain decoding model, the first step is to collect data on learners' viewing of visual stimulus images. fMRI signal Next, brain subtitles will be generated and visual stimulus reconstruction will be completed based on paired visual stimuli and their brain imaging images.
[0018] To address the difficulty existing models face in extracting latent features from fMRI, this model first trains a Local Global Feature Aggregator (LGA) using a masking modeling task. Then, it uses L1 loss to calculate the difference between the acquired learner fMRI signal and the reconstructed fMRI signal, improving the model's ability to extract latent features. Next, LGA is used to extract text-generation-related features from the learner fMRI signal, and these features are then fed into a text feature encoder via a latent feature mapper to generate fMRI-text query features. Next, it is fed into a text feature encoder to generate brain-related subtitles. Finally, in the visual stimulus generation part, a conditionally stable diffusion model is chosen as the tool for generating images. Specifically, this invention uses brain captions, the original stimulus image, and latent features from fMRI as conditions to guide image reconstruction. These three different conditions are weighted through cross-attention calculation to optimize the final image generation result, ultimately generating reconstructed stimuli. .
[0019] Ultimately, by combining the generated brain captions and reconstructed visual stimuli, this model aims to determine learners' concentration levels using MRI. This method not only helps researchers better understand learners' psychological states but also enables learners to more intuitively understand their own cognitive states, improving communication effectiveness and satisfaction between teachers and students, and has broad application prospects in the field of education.
[0020] II. Construction of the Brain Interpretation Model
[0021] 1. The model architecture is as follows: Figure 2 As shown.
[0022] 2. The model's workflow.
[0023] ① fMRI signals from learners are the foundation of model research; in this invention, it is first necessary to select appropriate visual stimulus I to present to learners, and scan their brains while they are viewing the visual stimulus to obtain functional magnetic resonance imaging, which will facilitate the following work.
[0024] ② Pre-trained local and global feature aggregators for feature extraction.
[0025] Before performing brain captioning and visual stimulus reconstruction, a latent feature extractor based on fMRI signals needs to be pre-trained, which is the foundation for the following work. Studies have shown that hemodynamic response and spatial smoothness in the BOLD signal of fMRI can lead to spatial ambiguity, resulting in spatial redundancy in fMRI data. Even if most of the signal is masked, the fMRI data may still be recovered. Therefore, in order to improve the performance of the feature extractor, this invention uses a masking modeling task to train a local-to-global feature aggregator that can efficiently extract features for the upstream task, laying the foundation for subsequent captioning generation and stimulus reconstruction.
[0026] The first step is to receive the input fMRI signal. Next, the vectorized voxels need to be converted into patches, and then encoded using convolutions with kernels and strides equal to the patch size to obtain the patched signal. Next, a random masking strategy was used to mask the signal with a masking ratio of 0.75, resulting in the masked signal. .
[0027] The second step involves feeding the unmasked portion of the signal into a local-to-global feature aggregator. The encoder learns an effective fMRI representation, while the decoder predicts the masked portion, thus achieving masked reconstruction and obtaining the reconstructed signal. .
[0028] The encoder consists of an efficient relation aggregator and a feedforward neural network; the efficient relation aggregator can solve local redundancy and complex global dependencies, achieving efficient and effective spatiotemporal representation learning; the efficient relation aggregator adopts multi-head fusion learning, which can be represented as: .
[0029] .
[0030] in: Indicates the first Effective relation aggregator for individual heads; matrix Indicates integration The learnable parameter matrix of the size; Local relation aggregator and global relationship aggregator The system consists of shallow convolutional networks with local relation aggregators that focus on local neighborhood features to generate a learnable parameter matrix. Its formula is: .
[0031] in: Represents a token. Represents local neighborhood Any token; The learnable parameter matrix represents the relative offset positions of the two. Local relation aggregation only establishes neighborhood feature associations. Therefore, additional positional encoding and feedforward layers are introduced. This special combination effectively enhances the feature representation.
[0032] Deep networks establish long-term relationships in the global context feature space, consistent with the idea of self-attention mechanisms. They design a token relevance matrix from a global perspective and establish a global relationship aggregator by calculating the similarity of global context features. Its formula is: .
[0033] in: and This represents a linear transformation function. The aggregated features are then input into a feedforward neural network for further processing to obtain... Its formula is:
[0034] .
[0035] in: This indicates a linear layer.
[0036] Finally, the three features are aggregated into latent features through residual connections. : .
[0037] The decoder also employed the same strategy to achieve masking reconstruction, laying the foundation for subsequent research.
[0038] ③ Train the brain's subtitle generator.
[0039] This invention proposes a general and efficient brain language model for extracting brain captions from human fMRI activity. The brain caption generator consists of a text feature extractor, a text feature encoder, and a text decoder, which are used for brain representation learning, multimodal alignment of the brain and language, and text generation, respectively.
[0040] The text feature extractor originates from the encoder portion of a pre-trained feature relation aggregator, enabling efficient extraction of features suitable for downstream tasks from fMRI signals. The latent features used for caption generation are expressed as follows: .
[0041] To connect fMRI with the semantic space, the extracted features are then fed into a feature mapper consisting of fully connected layers and 1D convolutional layers for dimensionality adjustment, so that the feature dimensions meet the requirements of the text feature encoder.
[0042] The text feature encoder transforms fMRI features into a series of trainable fMRI-text feature queries corresponding to relevant visual and semantic information. The process is as follows: .
[0043] Wherein: FFN represents a feedforward neural network, which autonomously learns higher-level features through multiple hidden layers to improve the model's subsequent performance; MAB represents a cross-attention module, which helps to consider global information, thereby more accurately realizing the connections between contexts, and can dynamically adjust according to different input features, making the model more flexible and enhancing its adaptability; SA represents a self-attention mechanism, which can capture long-distance dependencies between elements in a sequence, helping the model understand contextual information; generated under the action of the text feature encoder. .
[0044] The text decoder is built upon a weight-invariant large language model; here, this invention conducts experiments on BERT. To obtain BERT's generative language capabilities, this invention designs a text latent feature mapper to connect the text feature encoder and BERT, which will output... The embedding is linearly projected to the same dimension as the text embedding in LLM, and then the query embedding of the projection is pre-applied to the input text embedding; since the initialized text feature encoder can extract brain representations of linguistic information, it effectively acts as an information bottleneck, feeding the most useful information to BERT while removing irrelevant brain information; then, BERT is able to generate text based on brain representations learned from visually evoked functional magnetic resonance imaging; furthermore, this invention constructs a text latent feature mapper like the fully connected module in BLIP-2 to share its pre-trained weights and keep them frozen during training; therefore, this invention trains the entire brain caption generator, with language modeling of the loss between the generated brain captions and the ground-based live COCO title, while actually only training the encoder and fMRI feature mapper as follows: .
[0045] in: and They represent the first A set of real-stimulus captions based on fMRI and captions predicted by the model; This indicates the number of COCO captions corresponding to each stimulus image; this invention uses five captions to train the model to enhance the semantic richness and structural flexibility of the text description. This indicates the data size of the training fMRI samples for the subjects.
[0046] ④ Visual stimulus reconstruction model under training conditions.
[0047] In visual reconstruction, this invention chooses to use a conditional diffusion model to generate stimulus images, consisting of a forward process and a backward process, both of which are parameterizable Markov chains. The forward process is used to decompose the image into noise, while the backward process is used to recover the generated image from the noise. The initial image representing the input to the stable diffusion model is used to progressively add noise during the forward propagation process to obtain the noise. The reverse process, on the other hand, predicts the addition of latent features by training a denoised EU-Net. The image is used to generate the denoised image.
[0048] The forward diffusion process requires gradually adding noise to the image to make it standard Gaussian noise, satisfying the following: .
[0049] By choosing the appropriate , making Then the random variable The mean compared to the sample It will approach 0, and the variance will also gradually approach 0 as t increases. From the perspective of probability distribution, the initial sample After infinite diffusion, The original signal distribution has almost disappeared, and it is almost entirely noise. Therefore, a mapping from the initial sample distribution to the standard Gaussian distribution is achieved. .
[0050] Acquiring latent features in fMRI signals It is necessary to use the previously trained encoder to process the fMRI signal. The specific process of extracting latent features from the visual stimulus reconstruction model is as follows: .
[0051] in: Indicates a linear layer; Indicates the activation function; The temporal encoder is responsible for encoding temporal information into the model to help understand and utilize time-related features. By modeling the relationships between time steps, the temporal encoder provides conditional information for the generation process, ensuring that the model can generate consistent outputs at different time points. At the same time, it helps handle Markov properties, supports multimodal applications, and significantly improves the quality of generated results. In summary, by introducing the time dimension, the temporal encoder makes the diffusion model more flexible and coherent in generation tasks.
[0052] Next, captions generated by a brain caption generator and latent features obtained from fMRI are used as conditions to guide stimulus reconstruction, and these conditions are integrated through cross-attention. In the cross-attention model, different input conditions (e.g., image features and text descriptions) are mapped to a common representation space. The model generates a weight matrix by calculating the similarity between these inputs, and then dynamically adjusts the influence of each input during information fusion. This approach not only allows the model to focus on more relevant information, but also improves the quality of the output by capturing and fusing cross-modal information. Furthermore, the cross-attention mechanism significantly improves the model's ability to model complex and long-range dependencies by directly weighting and associating different inputs, effectively handling long-range dependencies and generating stimulus reconstruction results under conditional guidance. .
[0053] The loss function for the visual stimulus reconstruction model under the given conditions is as follows: .
[0054] Ultimately, by integrating the brain captions and visual stimulus reconstruction results generated by two independent models, we are able to present this key information to researchers and learners simultaneously. In this process, the present invention introduces a multimodal information combination approach, enabling researchers to more comprehensively understand learners' feelings and needs during the diagnostic process, while helping learners better understand their condition and treatment plan, thereby reducing communication barriers, improving the quality of medical care and learner satisfaction, and pioneering a completely new method for attention detection.
[0055] III. Detailed Explanation of Each Structure of the Model.
[0056] 1. Overview.
[0057] This invention proposes a two-stage agnostic neural visual decoding framework, Intellect-vis, which reconstructs images from fMRI signals induced by visual stimuli.
[0058] like Figure 3 As shown, in stage A, the present invention trains an autoencoder of Local Global Aggregator (LGA) on the large public dataset HCP to learn an effective feature representation of functional magnetic resonance imaging (fMRI) signals, with sparse masking signals as the upstream task, and the learned feature representation is used as a condition to guide the image reconstruction in the next stage; in stage B, the pre-trained fMRI encoder is integrated with the Effective Latent Diffusion Model (ELDM) through the cross attention module (MAB), the temporal mapping module (TMB) to perform conditional synthesis to obtain the final result of visual decoding.
[0059] After the fMRI signal is input into Intellect-vis, it is first sparsely masked and encoded. The vectorized voxels are divided into patches, and then the patches are transformed into embeddings using convolution with a stride of the patch size. During this transformation, 75% of the voxel signal is masked, and the unmasked embeddings are sent to the LGA to perform the sparse signal modeling task in Stage A. In this task, the encoder learns an effective fMRI voxel signal embedding representation, and then the decoder reconstructs the masked signal. At the same time, the encoder trained by the sparse masking signal modeling task learns the fMRI signal features. In order not to increase the number of additional parameters of the decoder, it will be reused in Stage B. Since the decoder is no longer used in Stage B, the computational burden is reduced as much as possible in the design.
[0060] To address the modality conversion problem between fMRI signals and visual stimuli, this invention constructs an Effective Latent Diffusion Model (ELDM) for efficient image reconstruction. During the forward diffusion process, noise is gradually added to blur the signal until it conforms to a standard Gaussian distribution. In the backward reconstruction process, denoising is performed at each time step to achieve image reconstruction. EU-Net is constructed within the ELDM to enhance its denoising capabilities. Furthermore, this invention employs dual conditionally encoded voxel signals and feeds them into the ELDM, using MAB to deeply integrate local features into global dependencies, further improving image reconstruction. After joint training with EU-Net and MAB, modality conversion between fMRI signals and visual stimuli is achieved, addressing the problem of the human brain's invisibility.
[0061] 2. Phase A Local Global Aggregator (LGA).
[0062] The activity in the human brain involves nonlinear interactions between hundreds of millions of neurons, making it highly complex. Measuring the BOLD signal in fMRI is an indirect and comprehensive measure of neuronal activity, which can be used to analyze the implicit correlations between voxels when functional networks respond to external stimuli. Therefore, in Stage A, this invention learns these implicit correlations by recovering masked voxels, thereby enabling a pre-trained model to gain a deeper understanding of the context of fMRI signals.
[0063] The hemodynamic response and spatial smoothing in the BOLD signal of fMRI can lead to spatial ambiguity, resulting in spatial redundancy in the fMRI data. Even if most of the signal is masked, the fMRI data may still be recovered. Therefore, in Intellect-vis Stage A, in order to mask most of the fMRI patches and improve computational performance without affecting the learning ability of LGA masking modeling, this invention uses a large embedding patch scale, which significantly preserves the feature capacity while expanding the representation space of fMRI.
[0064] In the sparse masked signal reconstruction task of Stage A, an automatic encoding and decoding method is used to reconstruct the masked portion of the signal; where LGA consists of an encoder and a decoder, the former mapping the voxel signal to a latent feature embedding. The latter will Reconstructed signal representation ;like Figure 4 As shown, this is a comparison of signal features before and after LGA reconstruction; the leftmost part is the initial input signal, the middle part is the signal representation after masking, and the rightmost part is the reconstructed signal. It can be seen that the LGA coding model retains the extracted salient features.
[0065] ① Patch encoding.
[0066] Upon receiving the input voxel signal Next, the vectorized voxels need to be converted into patches, and then encoded using convolutions with kernels and strides equal to the patch size to obtain the patched signal. ,in express , Indicates the sequence length. This represents the feature dimension at each time step.
[0067] ② Masking signal pre-training for effective and robust EEG representation.
[0068] Since the masking strategy and masking range both affect the final image reconstruction result, in order to better learn the feature representation of fMRI signals, it is necessary to select an appropriate masking strategy and masking ratio; the fMRI signal patch is encoded as follows: Then, a strategy of random sampling without replacement is used to mask the patch; random sampling with a higher masking ratio can largely eliminate redundancy, so this invention chooses a high masking ratio of 0.75 to solve the task of inferring the remaining regions based on the visible adjacent voxel regions, which creates opportunities for the design of highly sparse data encoders.
[0069] This invention employs a random masking strategy to mask patched data. Using masking ratio To control the size of the masking range; first, determine the length of the sequence to be retained. Then, a noise tensor is generated based on this, where each element is a random number in the range [0,1]. Next, the distribution of the mask is controlled by a weight tensor, using these weights to select the positions to be masked, and the corresponding elements in the noise tensor are set to 1.1, so these positions will be removed during subsequent sorting. The function then sorts the noise tensor to determine which positions need to be retained; the sorted indices... This indicates the positions that need to be retained. Then, by sorting the sorted indices again, we obtain... This is used to restore the original sequence order; finally, the function will adjust the sorted indexes. From the input sequence Extract the parts that need to be retained and generate a binary mask tensor. Where 0 indicates to keep and 1 indicates to remove; finally, for The binary mask is then reordered to restore its original order, yielding the masked signal. .
[0070] ③ Encoder and decoder.
[0071] Inspired by the Uniformer, this invention designs the Local Global Aggregator (LGA) as the main component of masking reconstruction. Current mainstream solutions using convolution and ViT struggle to balance local redundancy and global dependency. Convolution's advantage lies in aggregating context within a small local domain, but this limits the receptive field and makes it difficult to model global dependencies. In ViT, the self-attention module establishes long-distance contextual relationships by calculating global token similarity, which generates a lot of unnecessary computation and fails to address local redundancy. In comparison, convolution is more efficient in extracting shallow features. By adopting a transformer-like design to address the differences between shallow and deep features, convolution and self-attention can be combined to fully leverage the iteration of local and global features at different depths, achieving efficient feature learning with LGA.
[0072] The structure of the designed LGA is as follows Figure 5 As shown, LGA consists of two parts: Effective Relationship Aggregator (ERA) and Feed Forward Network (FFN).
[0073] ERA can address local redundancy and complex global dependencies, achieving efficient and effective spatiotemporal representation learning. Specifically, ERA employs multi-head fusion learning, expressed by the following formula: .
[0074] .
[0075] in: Indicates the first The effective relation aggregator (ERA) is a single-headed, efficient relation aggregator. It is integration The learnable parameter matrix of the size; All are local relation aggregators and global relationship aggregator The system consists of shallow convolutional networks with local relation aggregators that focus on local neighborhood features to generate a learnable parameter matrix. Its formula is: .
[0076] in: express ; Represents local neighborhood Any token; The learnable parameter matrix represents the relative offset positions between the two, indicating that local relation aggregation only establishes neighborhood feature associations. Therefore, additional positional encoding and a feedforward layer are introduced. This special combination effectively enhances the features. Express.
[0077] Deep networks establish long-term relationships in the global context feature space, consistent with the idea of self-attention. They design a token relevance matrix from a global perspective and establish a global relationship aggregator by calculating the similarity of global context features. Global Relation Aggregator Represented as: .
[0078] in: and All are linear transformation functions; the aggregated features are input into a feedforward neural network for further processing to obtain... for: .
[0079] in: FFN Indicates feedforward layer; This indicates a linear layer.
[0080] Finally, the three features are aggregated through residual connections to form latent features. for: .
[0081] The decoder also uses LGA for masking reconstruction. Since the goal of the entire task is visual decoding, and Stage A is for training the encoder to extract features more effectively in order to achieve modality conversion between fMRI signals and visual stimulus images, this invention uses a lightweight decoder to minimize the number of parameters and reduce the computational burden without affecting the overall task performance.
[0082] 3. Feature relation mapping.
[0083] In order to effectively denoise the potentially salient features extracted in Stage A, this invention constructs a convolutional module with attention mechanism and temporal mapping encoding to achieve effective mapping of feature relationships.
[0084] ① Memory Attention Block (MAB) mechanism.
[0085] In the task of reconstructing visual stimuli from fMRI signals, the reconstruction of inverse spatiotemporal features of brain-imaging visual stimuli is more closely related to the reconstruction task of natural images. On the one hand, the cut-off points between image patches affect the reconstruction effect of visual stimuli, and on the other hand, the contribution weight of each voxel to the reconstruction, the distribution of the visual encoding sequence will affect the final reconstruction effect. To address this, this invention constructs a new and more effective attention module, MAB (Memory Attention Block), whose structure is as follows: Figure 6 As shown, MAB can calculate the contribution weight of each image patch to the reconstruction task and establish a correlation with the weight-related sequence, thereby improving the final visual decoding effect.
[0086] MAB will identify the potential features of Stage A. The latent features and the context tensor (CE) embedding of the image patch are used as input, and the scaling factor N controls the scaling size. Specifically, this invention first maps the latent features and the context tensor to the same dimensional space through linear transformations, obtaining... , with Q and Multiply to get This achieves the goal of reducing the amount of computation, and then... A series of operations, including matrix multiplication with K, are performed to obtain the multi-head attention tensor. It enables cross-block attention; its multi-head attention tensor Represented as: .
[0087] in: This represents matrix multiplication.
[0088] Finally, V and Attn are multiplied by matrix to obtain the features after cross-attention. Its formula is: .
[0089] ②Time Mapping Block (TMB).
[0090] In the diffusion model of Stage B, image reconstruction is performed step-by-step. The reconstruction effect of each time step directly or indirectly affects the reconstruction effect of subsequent time steps. Therefore, ensuring effective feature generation at each time step is crucial. Since the fMRI signal features extracted in Stage A do not include temporal features, this invention optimizes the feature generation to better integrate them into the diffusion model. Add time-mapping features, such as Figure 7 As shown.
[0091] It is a one-dimensional tensor of length *n*, representing the time step of the corresponding processed element; the time step embedding feature representation is a sine function embedding of the time step created after time mapping. , is represented as: .
[0092] in: Indicates time; This indicates the embedding dimension of the input.
[0093] Then After linear activation and spliced together This allows the fMRI features extracted from Stage A to be effectively infused with temporal information and mapped into the potential diffusion model of Stage B. The formula is as follows: .
[0094] 4. An effective conditional diffusion model based on dual-conditional mapping for stage B.
[0095] In Intellect-vis stage B, this invention uses ELDM to generate stimulus images, consisting of a forward process and a reverse process, both of which are parameterizable Markov chains. The forward process is used to decompose the image into noise, while the reverse process is used to recover the generated image from the noise. The initial image input to the ELDM is used to represent the noise. During the forward pass, noise is gradually added to obtain the noise. The reverse process, on the other hand, predicts the addition of latent features by training a denoised EU-Net. The image is used to generate the denoised image.
[0096] ① Forward diffusion - add noise.
[0097] To avoid confusion, this invention will be referred to as This represents an image without a conditional mechanism, in order to This indicates that different time steps added The generated image, This indicates that standard Gaussian noise has been added after forward diffusion.
[0098] During forward diffusion, by using any initial sample By continuously adding T-order Gaussian noise, a series of images can be obtained. And as T approaches infinity, the original image Its characteristics completely disappear, becoming standard Gaussian noise.
[0099] This invention models the process of adding noise as follows: .
[0100] Right now .
[0101] in, And satisfy ;noise Indicates the signal Interference; This indicates the degree of noise added to the original signal at each time step. After multiple time steps, we obtain:
[0102] .
[0103] in: This represents the sum of multiple independent normal noises, with a mean of 0 and variances of [missing information]. .because By gradually adding the squares of each coefficient and summing them up, the noisy signal at time step t can be obtained. for:
[0104] .
[0105] Based on the superposition property of the normal distribution, we can conclude that: .
[0106] By choosing the appropriate , making Then the random variable The mean compared to the sample It will approach 0, and the variance will also gradually approach 0 as t increases. From the perspective of probability distribution, the initial sample After infinite diffusion, The original signal distribution has almost disappeared, and it is almost entirely noise. Therefore, a mapping from the initial sample distribution to the standard Gaussian distribution is achieved. .
[0107] ② Reverse generation - noise removal.
[0108] A: Reverse generation.
[0109] During the forward diffusion process, this invention adds noise to the original features to denoise them; the reverse generation is a process of removing noise in reverse. In the reverse generation, standard Gaussian noise is injected. ,in accordance with Sampling, inferring from, and reconstructing the image—it's important to note that if… Small enough The sampling results also conform to a Gaussian distribution, but their values are difficult to evaluate. Therefore, it is difficult to gradually deduce the original true distribution by solving a formula. So this invention utilizes the noise-adding features in the forward process to... By training a denoising model The reverse generation of the reconstructed image will and The mapping relationship between them is defined as a parameter-learnable neural network, using... This represents the reverse inference process at each time step.
[0110] Since ELDM is constructed by adding a conditional mechanism to the unconditional diffusion model, this invention explains the reverse process as a diffusion process and a conditional diffusion process.
[0111] At this point, assume As the initial input, we can obtain the following from Bayes' theorem: .
[0112] In order to rebuild This invention uses reduction and Training the generative network using Euclidean distance between them Specifically, From this, it can be inferred that ,but: .
[0113] in: express The generation process can then be rewritten using a formula as follows:
[0114] .
[0115] Its loss function can be expressed as: .
[0116] B: Conditional diffusion.
[0117] After adding the conditional mechanism, the distribution can be represented as: .
[0118] in: This represents the fMRI features extracted by the encoder. ; and These represent the conditions after the addition of the conditional mechanism. The time step and the first The reconstructed image is generated after [number] time steps. When and It will only be significantly greater than 0 when they are very close; the Taylor expansion of the diffusion model after adding the conditional mechanism is:
[0119] .
[0120] in: Represents gradient change; assuming ,but: .
[0121] in: Indicates direct proportion; express Noise at any given moment; that is Approximately In other words, after adding the conditional mechanism, it can be represented as: .
[0122] By continuously advancing the time step, the reconstructed image after incorporating the conditional mechanism can eventually be obtained. .
[0123] Therefore, the loss function of ELDM can be expressed as: .
[0124] This invention introduces for the first time a three-condition integrated guidance process for the reverse generation of a diffusion model in the visual stimulus reconstruction part of the model. In addition to traditional visual stimuli and fMRI feature extraction, it adds guidance conditions for brain-generated subtitles. This innovation significantly improves the quality of stimulus reconstruction, making the reconstructed visual stimuli more reliable. Furthermore, by ensuring the consistency of meaning between the subtitle text and the reconstructed image, this invention assigns different weights to the conditions through cross-attention to achieve comprehensive consideration, effectively avoiding potential contradictions between the two, thereby enhancing the effectiveness and reliability of learner attention detection. This comprehensive method not only promotes the development of image reconstruction technology but also provides a clearer basis for understanding brain activity, laying an important foundation for subsequent research.
[0125] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A method for detecting learner attention based on MRI images, characterized in that, Includes the following steps: Acquire fMRI signals of the learner's brain while viewing original visual stimulus images; The brain fMRI signal is input into the attention detection model, which includes a feature extraction module and a stimulus reconstruction module; The feature extraction module is used to extract features from the brain fMRI signal to obtain fMRI features associated with the learner's learning process and fMRI features associated with the learner's cognitive activities. The fMRI features associated with the learner's cognitive activities are then converted into brain subtitles representing cognitive activities. The stimulus reconstruction module employs the cross-attention mechanism MAB, assigning different weights to brain captions representing cognitive activities, fMRI features associated with the learner's learning process, and the original visual stimulus image. These different weights guide the utilization of brain captions, fMRI features associated with the learner's learning process, and the original visual stimulus image during visual stimulus image reconstruction, resulting in the reconstructed visual stimulus image. Learners’ attention levels are detected based on the similarity between the reconstructed visual stimulus image and the original visual stimulus image.
2. The learner attention detection method based on MRI images according to claim 1, characterized in that, The acquisition of fMRI features associated with the learner's learning process specifically involves feature extraction using a Local Global Feature Aggregator (LGA), including: The Local Global Feature Aggregator (LGA) consists of the Efficient Relation Aggregator (ERA) and a feedforward neural network. The Effective Relation Aggregator (ERA) employs multi-head fusion learning, as shown below: ; ; in: Indicates the first Effective relation aggregator for individual heads; matrix Indicates integration The learnable parameter matrix of the size; Local relation aggregator and global relationship aggregator The system consists of shallow convolutional networks with local relation aggregators that focus on local neighborhood features to generate a learnable parameter matrix. Its formula is: ; in: Represents a token. Represents local neighborhood Any token; Represents the learnable parameter matrix; A global relationship aggregator is established by calculating the similarity of global context features. Global Relation Aggregator Represented as: ; in: and All are linear transformation functions; the aggregated features are input into a feedforward neural network for further processing to obtain... for: ; in: FFN Indicates feedforward layer; Indicates a linear layer; Finally, the three features are aggregated through residual connections to form fMRI features associated with the learner's learning process. , is represented as: 。 3. The learner attention detection method based on MRI images according to claim 2, characterized in that, The generation of the brain-based subtitles includes: The Local Global Feature Aggregator (LGA) was used to extract features from brain fMRI signals to obtain fMRI features associated with learners' cognitive activities. The fMRI features associated with learners' cognitive activities are represented using a text feature extractor, and the representation equation is as follows: ; Using a text feature encoder This is transformed into a series of trainable fMRI text query features corresponding to relevant visual and semantic information. The transformation equation is: ; Where: FFN represents a feedforward neural network; MAB represents a cross-attention mechanism; SA represents a self-attention mechanism; Brain captions were generated from fMRI text query features using a text decoder.
4. The learner attention detection method based on MRI images according to claim 1, characterized in that, The process by which the stimulus reconstruction module reconstructs the visual stimulus image is a reverse generation process within the reverse diffusion process of the Conditional Diffusion Model (ELDM), including: In the reverse diffusion process, it is assumed As the initial input, we can obtain the following from Bayes' theorem: ; Generative network trained To minimize and To rebuild the Euclidean distance between them ;but From this, it can be inferred that ,but: ; in: express ; The generation process can then be rewritten using a formula as follows: ; Its loss function is expressed as: 。 5. The learner attention detection method based on MRI images according to claim 4, characterized in that, The reconstruction of the visual stimulus image includes: After adding the conditional mechanism, the distribution is represented as follows: ; in: This represents the fMRI features extracted by the encoder. ; and These represent the conditions after the addition of the conditional mechanism. The time step and the first Reconstructed image generated after one time step; when and When they are very close, the value will be significantly greater than 0; the Taylor expansion of the diffusion model after adding the conditional mechanism is expressed as: ; in: Represents gradient change; assuming ,but: ; in: Indicates direct proportion; express Noise at any given moment; that is Approximately After adding the conditional mechanism, it is represented as: ; By continuously advancing the time step, the reconstructed image with the addition of a conditional mechanism is finally obtained. .
6. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the steps of the learner attention detection method based on MRI images as described in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the steps of a learner attention detection method based on MRI images as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method for fMRI data embedding and pattern recognition based on attention mechanism
CN119207821A
Apparatus for analysing focus and nonfocus states and method therof
KR1020120124772A