MRI (Magnetic Resonance Imaging) image analysis method based on multi-time-point and multi-modal Mama model

By fusing MRI and PET data using a multi-temporal, multimodal Mamba model, and combining a time-aware Mamba state-space model with a pixel-level dual-cross attention mechanism, this approach solves the problems of high computational complexity and low information utilization in existing multi-temporal, multimodal data processing technologies, enabling accurate detection and interpretable assessment of early lesions in brain diseases.

CN121860928APending Publication Date: 2026-04-14NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies for assessing brain diseases using MRI and PET images are insufficient to fully depict the dynamic trajectory of the brain from compensation to decompensation. Furthermore, traditional models suffer from high computational complexity and low data utilization in processing multi-time-point and multi-modal data.

Method used

Employing a multi-temporal, multimodal Mamba model, this study integrates MRI structural, PET metabolic, and neuropsychological scale data. By utilizing a time-aware Mamba state-space model and a pixel-level dual-cross attention mechanism, it achieves cross-modal alignment and information fusion, capturing the dynamic evolution of brain structures.

Benefits of technology

It improves the accuracy and interpretability of detecting early lesions in brain diseases, overcomes the limitations of insufficient sensitivity and high computational complexity of traditional models, and enables early and accurate clinical decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860928A_ABST
    Figure CN121860928A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent analysis of medical images, in particular to an MRI (Magnetic Resonance Imaging) image analysis method based on a multi-time-point and multi-modal Mama model, which comprises the following steps: acquiring T1 weighted structure MRI, PET (Positron Emission Tomography) images and neuropsychological assessment scale data of the same subject at each time point through a database; bias field correction, tissue segmentation and spatial registration are carried out on the MRI and PET images, and corresponding potential features are synthesized by a 3D GAN-ViT generation network when the PET images are missing; mRI potential features, PET potential features, scale category and numerical value embedding and baseline time difference coding at each time point are spliced into a multi-modal token, and an individual specificity longitudinal sequence is constructed; modeling by using a time-aware Mama state space model, introducing a time decay factor to describe the brain structure evolution dynamic state, fusing original image space information, and outputting the hidden state of each time point; and sending the accurate lesion image features to a terminal for providing an image basis for risk detection and evaluation for a doctor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent medical image analysis technology, specifically to an MRI image analysis method based on a multi-time-point, multi-modal Mamba model. Background Technology

[0002] With the continuous improvement of my country's scientific and technological level, imaging technology has been widely used in the field of medical testing. In clinical practice, doctors usually rely on single-timepoint neuropsychological scales, cerebrospinal fluid biomarkers, or single-segment structural magnetic resonance imaging (MRI) for assessment. However, scale results are easily affected by education level, emotional state, and assessment environment; biomarker collection is invasive; and a single imaging modality can only provide local information on structure, metabolism, or function, making it difficult to comprehensively depict the dynamic trajectory of the brain from compensation to decompensation.

[0003] In recent years, researchers have attempted to combine machine learning with neuroimaging, using support vector machines, random forests, or 3D convolutional neural networks to extract features and classify MRI and PET images. However, most of these approaches are limited to cross-sectional single-modal data, neglecting the intrinsic temporal evolution of individuals. Furthermore, traditional convolutional networks have a large number of parameters and fixed receptive fields in 3D medical images, making it difficult to capture long-range dependencies, and often discarding samples lacking the PET modality, resulting in decreased data utilization. In addition, while Transformer-type models can model long sequences, their quadratic computational complexity presents a dual bottleneck of memory and training time for multi-timepoint and multi-modal longitudinal expansion. Therefore, how to fully utilize the three heterogeneous information types—"MRI-structure," "PET-metabolism," and "scale-cognition"—while maintaining linear complexity, and explicitly embedding the passage of time into network memory, has become a crucial scientific problem that urgently needs to be solved. Summary of the Invention

[0004] The purpose of this invention is to provide an MRI image analysis method based on a multi-time-point, multimodal Mamba model. By fusing individual longitudinal MRI structure, PET metabolic and neuropsychological scale data, and utilizing the time-aware Mamba state-space model to capture brain evolution dynamics, combined with a pixel-level double-cross attention mechanism, this method provides doctors with more accurate images of lesion evolution, enabling early, precise, and interpretable clinical decision support.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0006] A method for MRI image analysis based on a multi-timepoint, multimodal Mamba model, characterized in that: the method includes:

[0007] S100. Obtain T1-weighted structural MRI, PET images, and neuropsychological assessment scale data of the same subject at at least 3 time points within 18 months from the Alzheimer's Disease Neuroimaging Project (ADNI) database to form a longitudinal multimodal raw data sequence.

[0008] S200 performs bias field correction, tissue segmentation and spatial registration on MRI and PET images. For time points with missing PET images, a 3D GAN-ViT generative network is used to synthesize PET images and their potential spatial features with the corresponding MRI as input, so as to achieve cross-modal alignment between MRI and PET.

[0009] S300: The latent features of MRI and PET at each time point, the embedded categories of the scale, the embedded values ​​of the scale, and the sinusoidal encoding of the time difference from the baseline are spliced ​​together and linearly projected into a multimodal token of a unified dimension, and stacked according to the scanning time sequence to construct an individual longitudinal sequence.

[0010] S400. Input the longitudinal sequence into the time-aware Mamba state-space model, introduce a time decay factor into the memory unit to model the sequence, and output the hidden state features at each time point.

[0011] S500: Using the latent state features as the query vector, perform pixel-level double cross-attention calculation with the key and value vectors after flattening the original MRI and PET images, and fuse global temporal features with local image spatial information to obtain weighted fusion features;

[0012] The S600 accurately locates lesion images based on weighted fusion features and sends the image features to the terminal to provide doctors with image-based evidence for risk detection and assessment.

[0013] Preferably, S100 includes:

[0014] Participants with a baseline diagnosis of mild cognitive impairment were selected from the Alzheimer's Disease Neuroimaging Project (ADNI) database. Each participant was required to undergo at least three T1-weighted three-dimensional structural MRI scans and corresponding PET scans within 18 months of their initial scan, with at least two months between any two scans, to create a longitudinal observation covering the individual's early disease trajectory. Simultaneously, participants were required to maintain a mild cognitive impairment diagnosis for 18 months after their initial scan to exclude those who had already converted to mild cognitive impairment in a short period, thereby improving the model's ability to identify those with slow progression. A conversion event was defined as a Mini-Mental State Examination (MMSE) score below 24 or a Clinical Dementia Rating Scale (CDR) score of 1 or higher at any follow-up 18 months later. For each scan, the actual number of years since the initial scan was recorded. (Unit: Year) As a time difference, MRI, PET, and neuropsychological scale data obtained within the same visit were grouped together to construct an individual-level multi-time-point multimodal sequence. If PET images are missing at certain time points, MRI and scale data are retained and supplemented using a generative network in subsequent steps. Finally, the dataset is randomly divided into training, validation, and test sets in a 7:1:2 ratio according to the subject level to ensure that the proportion of transformation events in each set is consistent, so as to obtain a longitudinal multimodal raw data sequence with no information leakage and repeatability, providing a foundation for subsequent modeling.

[0015] Preferably, S200 includes:

[0016] S201. Obtain the original longitudinal multimodal data sequence. First, perform secondary correction of head movement and gradient distortion in a unified coordinate space. Then, use the SPM12 unified segmentation algorithm to perform bias field correction, gray matter, white matter, and cerebrospinal fluid probability segmentation, and MNI152 template spatial registration on each T1-weighted MRI to obtain the standardized brain volume after skull dissection.

[0017] For PET images, rigid registration with the corresponding MRI is first performed, followed by SUV normalization and Gaussian filtering with a smooth kernel half-width of 6 mm to make the image resolution consistent with the MRI voxel size and reduce noise. After spatial normalization, all images are resampled to 1.5 mm isotropic voxels, and the intensity values ​​are z-score normalized to eliminate the distribution shift caused by the difference between the scanner and the field strength.

[0018] S202. When PET data is missing at a certain time point, modality completion is performed using a 3DGAN-ViT generative network pre-trained on MRI-PET paired data: the corrected MRI is used as the generator input G to synthesize PET images and their latent spatial features. The generation process is as follows:

[0019] ;

[0020] in, This indicates that the generator weights have been pre-trained and frozen on the MRI-PET pairing set and will not be updated during the training phase to ensure that the synthesized features are in the same latent space distribution as the real PET.

[0021] Thus, aligned MRI and PET latent representations can be obtained at each time point, providing a complete and consistent source of features for subsequent multimodal token construction. After preprocessing, a set of image data packets corresponding one-to-one with the original scans, with offset correction, spatial registration, intensity normalization, and missing data completion, is output for direct use by the S300.

[0022] Preferably, S300 includes:

[0023] S301. Obtain the latent feature vectors of MRI and PET aligned at each time point, as well as the neuropsychological scale data collected during the same visit, and then divide the scale fields into two categories: categorical and numerical.

[0024] Categorical variables (such as CDR grade and APOE genotype) are mapped to dense embeddings using a lookup table, while numerical variables (such as MMSE score and FAQ total score) are first Z-score standardized and then linearly transformed to obtain continuous embeddings; the two types of embeddings are compared with the scan time difference. After concatenating the sinusoidal position codes, the scale-time joint embedding vector is obtained;

[0025] S302. The latent features of MRI and PET are embedded together with the above scale-time joint and spliced ​​according to the channel dimension. The overall dimension is compressed to a preset uniform length D through a linear projection layer, thereby forming a multimodal token at that time point: ;

[0026] In this context, ";" indicates concatenation of channel dimensions. This represents the total number of input feature channels. This indicates the preset unified token dimension. This indicates the time difference between the first scan and the current scan in the user's scan record. and Represents latent features from MRI and PET, extracted from S200's generative network or real images; and These represent scale embedding and time coding, respectively.

[0027] Obtain the time difference consisting of three time points for each target user. , and Then, by arranging the tokens in the order of scanning, a vertical sequence of individuals with a length of 3 is obtained:

[0028] ;

[0029] This sequence retains cross-modal information and explicitly includes time-varying information, allowing it to be directly input into subsequent Mamba networks for dynamic modeling. Through the above process, this step transforms heterogeneous, multi-source, and multi-time-point raw data into a unified, aligned, and serializable multimodal token representation, laying the data foundation for the vertical state-space modeling of the S400.

[0030] Preferably, S400 includes:

[0031] Individual longitudinal sequences are input into a time-aware Mamba state-space model to capture the dynamic patterns of brain structure evolution over time.

[0032] The Mamba state-space model consists of six stacked Mamba Blocks. Each Block sequentially performs RMS normalization, selective scanning, and residual connections. The selective scanning SSM model achieves long-range dependency modeling with linear complexity O(L) through a trainable state matrix A and input-related gating mechanisms B and C. Simultaneously, a time decay factor is introduced during the memory update phase, causing the contribution of observations with distances greater than a threshold to the current state to decrease exponentially, and a time decay coefficient is introduced. Enhance time difference The impact, among which >1, thus explicitly reflecting the clinical hypothesis that "the longer the time, the weaker the impact";

[0033] Specifically, first, the time difference corresponding to each token in the sequence is... Convert to attenuation coefficient:

[0034] ;

[0035] in, , These represent the time points corresponding to the current memory and the memory to be updated, respectively. Determined through validation set optimization. ∈(0,1) controls the degree of forgetting;

[0036] Before the SSM state derivation, the memory cell C is weighted and corrected:

[0037] ;

[0038] The corrected state is used in subsequent gating calculations to ensure that the smaller the difference between adjacent time points, the more memory is retained, and the greater the difference, the faster the forgetting.

[0039] After being refined layer by layer through six blocks, the model output is a hidden state sequence of equal length to the input. ,in and This one-to-one correspondence integrates cross-modal structural information as well as individual-specific temporal evolution characteristics. This hidden state sequence will serve as a unified representation for subsequent pixel-level cross-attention and dual-task prediction, realizing the transformation from a "static snapshot" to a "dynamic trajectory".

[0040] Preferably, S500 includes:

[0041] S501, Complete time-aware Mamba modeling and obtain the hidden state sequence. Subsequently, to further integrate fine-grained pathological information from the original image space, a pixel-level dual-cross attention mechanism is introduced. This mechanism uses the final hidden state of the Mamba output (or the global representation after self-attention weighting) as the query vector Q: The feature matrices flattened from the original MRI and PET images, respectively. , Cross-attention calculation is performed to map global temporal features to local pixel space;

[0042] Specifically, the MRI and PET three-dimensional images at corresponding time points are first flattened along the spatial dimension into a two-dimensional feature matrix, resulting in... , and , ;

[0043] MRI branch: , ;

[0044] PET branch: , ;

[0045] Subsequently, the query vector Q is mapped to the same channel dimension as the image features using linear projection, and the attention weights for the MRI and PET branches are calculated separately:

[0046] ;

[0047] ;

[0048] in, Indicates the channel dimension, used for scaling to prevent gradient saturation;

[0049] The weighted features of the MRI and PET branches are calculated and compared with the original data using a residual method. The fusion process creates a unified expression that incorporates both temporal evolutionary information and preserves spatial pathological details. The fusion output is:

[0050] .

[0051] The fusion features simultaneously carry individualized longitudinal evolution information and spatial lesion weights, which is more conducive to providing doctors with accurate images for lesion detection and assessment.

[0052] This design allows the model to "look back" at key regions in the image (such as the hippocampus and temporal cortex) when making decisions, improving its sensitivity to early atrophy or metabolic abnormalities, while providing clear pixel-level weighting for subsequent interpretability analysis (such as Grad-CAM visualization).

[0053] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0054] This invention integrates individual longitudinal MRI structural images, PET metabolic images, and neuropsychological scale data. It uses a 3D GAN-ViT network to complete the modalities missing from PET, and employs a time-aware Mamba state-space model to capture the dynamic evolution of brain structure with linear complexity. It also combines a pixel-level double-cross attention mechanism to fuse global temporal and local spatial features. This invention overcomes the limitations of traditional single-time-point and single-modal models, which suffer from insufficient sensitivity and poor interpretability. It solves the key problems of "incomplete information, lack of dynamics, and computational complexity" in the early prediction of the risk of conversion in individuals with mild cognitive impairment. Attached Figure Description

[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0056] Figure 1 This is a flowchart of an MRI image analysis method based on a multi-time-point multimodal Mamba model according to the present invention;

[0057] Figure 2 This is a flowchart of the MRI-to-PET generation network in an embodiment of the present invention;

[0058] Figure 3 This is a flowchart of the multimodal feature fusion module in an embodiment of the present invention;

[0059] Figure 4 This is a flowchart of the pixel-level dual-cross attention module in an embodiment of the present invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Please see Figures 1-4 The present invention provides the following technical solution:

[0062] Example 1: As Figure 1As shown, the purpose of this invention is to provide an MRI image analysis method based on a multi-time-point, multi-modal Mamba model. By fusing individual longitudinal MRI structure, PET metabolic and neuropsychological scale data, the method uses a time-aware Mamba state-space model to capture brain evolution dynamics and combines a pixel-level double-cross attention mechanism to provide clear lesion images, enabling early, accurate, and interpretable individualized prediction and clinical decision support.

[0063] Step 1: Subjects diagnosed with mild cognitive impairment at baseline were selected from the Alzheimer's Disease Neuroimaging Project (ADNI) database. Each subject was required to undergo at least three T1-weighted 3D structural MRI scans and corresponding PET scans within 18 months of the initial scan, with at least two months between any two scans, to create a longitudinal observation covering the individual's early disease trajectory. Simultaneously, subjects must maintain a mild cognitive impairment diagnosis for 18 months after the initial scan to exclude those who have already converted to mild cognitive impairment in a short period, thereby improving the model's ability to identify those with slow progression. A conversion event was defined as a Mini-Mental State Examination (MMSE) score below 24 or a Clinical Dementia Rating Scale (CDR) score of 1 or higher at any follow-up 18 months later. For each scan, the actual number of years since the initial scan was recorded. (Unit: Year) As a time difference, MRI, PET, and neuropsychological scale data obtained within the same visit were grouped together to construct an individual-level multi-time-point multimodal sequence. If PET images are missing at certain time points, MRI and scale data are retained and supplemented using a generative network in subsequent steps. Finally, the dataset is randomly divided into training, validation, and test sets in a 7:1:2 ratio according to the subject level to ensure that the proportion of transformation events in each set is consistent, so as to obtain a longitudinal multimodal raw data sequence with no information leakage and repeatability, providing a foundation for subsequent modeling.

[0064] Step 2: After receiving the original DICOM sequence obtained in Step 1, the head motion and gradient distortion are first corrected in a unified coordinate space. Then, the SPM12 unified segmentation algorithm is used to perform bias field correction, gray matter, white matter, and cerebrospinal fluid probability segmentation, and MNI152 template spatial registration on each T1-weighted MRI to obtain the standardized brain volume after skull dissection. For PET images, rigid registration with the corresponding MRI is first performed, and then SUV normalization and Gaussian filtering with a smooth kernel half-width of 6 mm are used to make the image resolution consistent with the MRI voxel size and reduce noise. After spatial normalization, all images are resampled to 1.5 mm isotropic voxels, and the intensity values ​​are z-score normalized to eliminate the distribution shift caused by the difference between the scanner and the field strength.

[0065] When PET data is missing at a certain time point, modal completion is performed using a 3D GAN-ViT generative network pre-trained on MRI-PET paired data: the corrected MRI is used as the generator input G to synthesize PET images and their latent spatial features, where the generation process can be represented as:

[0066] ;

[0067] in, This indicates that the generator weights have been pre-trained and frozen on the MRI-PET pairing set and will not be updated during the training phase to ensure that the synthesized features and real PET images are distributed in the same latent space. Thus, aligned MRI and PET latent representations can be obtained at each time point, providing a complete and consistent feature source for subsequent multimodal token construction. After preprocessing, a set of image data packets corresponding one-to-one with the original scans, with bias correction, spatial registration, intensity normalization, and missing data completion, is output and directly used in step 3.

[0068] Step 3: After processing in Step 2, each time point has aligned latent feature vectors from MRI and PET scans, along with neuropsychological scale data collected during the same visit. To form a unified representation that can be read by the sequence model, the scale fields are first divided into categorical and numerical categories: categorical variables (such as CDR level, APOE genotype) are mapped to dense embeddings using a lookup table, while numerical variables (such as MMSE score, FAQ total score) are first Z-score standardized and then linearly transformed to obtain continuous embeddings; the two types of embeddings and the scan time difference are then analyzed. After concatenating the sinusoidal position codes, the scale-time joint embedding vector is obtained.

[0069] Subsequently, the latent features of MRI and PET are embedded together with the aforementioned scale-time joint embedding and spliced ​​along the channel dimension. Then, the overall dimension is compressed to a preset uniform length D through a single linear projection layer, thereby constructing the multimodal token at that time point: ;

[0070] In this context, ";" indicates concatenation of channel dimensions. The total number of input feature channels, This is a pre-defined unified token dimension. f_MRI and f_PET have been extracted from the generative network or real images in step 2, and f_scale and f_time are the scale embedding and time encoding, respectively. The above operation is repeated for each subject's three time points, arranging the tokens sequentially according to the scan order to obtain an individual longitudinal sequence of length 3.

[0071] ;

[0072] This sequence retains cross-modal information and explicitly includes temporal evolution information, making it directly input into subsequent Mamba networks for dynamic modeling. Through the above process, this step transforms heterogeneous, multi-source, and multi-time-point raw data into a unified, aligned, and serializable multimodal token representation, laying the data foundation for the longitudinal state-space modeling in step 4.

[0073] Step 4: Input the individual longitudinal sequences obtained in Step 3 into the time-aware Mamba state-space model to capture the dynamic laws of brain structure evolution over time. The Mamba module consists of six stacked Mamba Blocks. Within each Block, RMS normalization, selective scan (SSM), and residual connections are performed sequentially. SSM achieves long-range dependency modeling with linear complexity (O(L)) through a trainable state matrix A and input-related gating mechanisms B and C. Simultaneously, a time decay factor is introduced during the memory update stage, causing the contribution of distant observations to the current state to decrease exponentially, and a time decay coefficient is introduced. ( >1), enhancing the effect of time difference on γ, thus explicitly reflecting the clinical hypothesis that "the longer the time, the weaker the effect".

[0074] In terms of specific implementation, firstly, the token corresponding to each token in the sequence... Convert to attenuation coefficient:

[0075] ;

[0076] in , These represent the time points corresponding to the current memory and the memory to be updated, respectively. Determined through validation set optimization. ∈(0,1) controls the degree of forgetting.

[0077] Before the SSM state derivation, the memory cell C is weighted and corrected:

[0078] ;

[0079] The corrected state participates in subsequent gating calculations, ensuring that smaller differences between adjacent time points result in greater memory retention, while larger differences lead to faster forgetting. After six blocks of progressive refinement, the model output is a hidden state sequence of equal length to the input. ,in and This one-to-one correspondence integrates cross-modal structural information as well as individual-specific temporal evolution characteristics. This hidden state sequence will serve as a unified representation for subsequent pixel-level cross-attention and dual-task prediction, realizing the transformation from a "static snapshot" to a "dynamic trajectory".

[0080] Step 5: Complete time-aware Mamba modeling and obtain the hidden state sequence. Subsequently, to further integrate fine-grained pathological information from the original image space, this step introduces a pixel-level dual-cross attention mechanism. This mechanism uses the final hidden state output by Mamba (or the global representation after self-attention weighting) as the query vector Q: ;

[0081] Feature matrices flattened with the original MRI and PET images respectively , Cross-attention calculation is performed to map global temporal features to local pixel space.

[0082] Specifically, the MRI and PET three-dimensional images at corresponding time points are first flattened along the spatial dimension into a two-dimensional feature matrix, resulting in... , and , ;

[0083] MRI branch:

[0084] ; ;

[0085] PET branch:

[0086] ; ;

[0087] Subsequently, the query vector Q is mapped to the same channel dimension as the image features using linear projection, and the attention weights for the MRI and PET branches are calculated separately:

[0088] ;

[0089] ;

[0090] in The channel dimension is used for scaling to prevent gradient saturation. The weighted features of the two branches are compared with the original features using a residual method. The fusion process creates a fused expression that includes both temporal evolutionary information and preserves spatial pathological details. The fusion output is: ;

[0091] This design allows the model to "look back" at key regions in the image (such as the hippocampus and temporal cortex) during decision-making, improving its sensitivity to early atrophy or metabolic abnormalities. It also provides clear pixel-level weighting for subsequent interpretability analyses (such as Grad-CAM visualization). Finally, the fused feature vector is fed into a dual-task prediction head to complete classification and risk score output.

[0092] Step 6: The fusion features obtained in Step 5 simultaneously carry individualized longitudinal evolution information and spatial lesion weights. This step accurately locates the lesion image based on the weighted fusion features and sends the image features to the terminal to provide doctors with image-based evidence for risk detection and assessment.

[0093] Preferably, this step can also set up a dual-task prediction head to initially detect lesion images and achieve integrated output of "whether it will transform" and "risk ranking". The upstream features are first compressed to 256 dimensions through a shared dimensionality reduction layer, and then branch into two independent branches: the binary classification head uses a single hidden layer with Sigmoid activation to output the probability that the current subject will develop Alzheimer's disease within a certain period of time. The loss function used is binary cross-entropy; the survival risk head uses linear units to directly output continuous risk scores. Optimization is performed using Cox partial likelihood loss, making Maintain a monotonous consistency with the actual conversion time; that is, the higher the risk, the earlier the conversion.

[0094] Binary classification head probability output: ;

[0095] Survival risk head: ;

[0096] in With binary head weights Dimensions are the same but parameters are independent. A dedicated bias is applied to the survival risk head, ensuring that the two tasks can be optimized independently according to their own objectives, thereby improving the consistency and accuracy of the output of the two tasks.

[0097] Losses and joint losses:

[0098] ;

[0099] ;

[0100] ;

[0101] Where i represents the individual that underwent the transformation event, and j represents the individual that had not yet transformed at the transformation time point i. Let be the risk set when individual i undergoes transformation, i.e., all individuals whose transformation time is greater than or equal to the transformation time of individual i. The gradient magnitudes for balancing the two tasks are determined through a validation set grid search. During training, both heads backpropagate simultaneously, sharing the backbone parameters; during inference, a single forward pass is sufficient to obtain the parameters simultaneously. and This meets the dual clinical needs of "both diagnosis and prognosis." Ultimately, the system is based on... A conversion warning is given if the value is >0.5, and based on... The population is divided into three risk levels: low, medium, and high. Kaplan-Meier curves and risk factor contributions are generated accordingly, enabling prediction of the entire process from imaging to decision-making, and providing a reference basis for doctors.

[0102] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for MRI image analysis based on a multi-time-point, multi-modal Mamba model, characterized in that: The method includes: S100. Based on the image database, obtain T1-weighted structural MRI, PET images and neuropsychological assessment scale data of at least 3 time points for the same target user to form a longitudinal multimodal raw data sequence; S200 performs bias field correction, tissue segmentation and spatial registration on MRI and PET images. For time points with missing PET images, a 3D GAN-ViT generative network is used to synthesize PET images and their potential spatial features with the corresponding MRI as input, so as to achieve cross-modal alignment between MRI and PET. S300: The latent features of MRI and PET at each time point, the embedded categories of the scale, the embedded values ​​of the scale, and the sinusoidal encoding of the time difference from the baseline are spliced ​​together and linearly projected into a multimodal token of a unified dimension, and stacked according to the scanning time sequence to construct an individual longitudinal sequence. S400. Input the longitudinal sequence into the time-aware Mamba state-space model, introduce a time decay factor into the memory unit to model the sequence, and output the hidden state features at each time point. S500: Using the latent state features as the query vector, perform pixel-level double cross-attention calculation with the key and value vectors after flattening the original MRI and PET images, and fuse global temporal features with local image spatial information to obtain weighted fusion features; The S600 accurately locates lesion images based on weighted fusion features and sends the image features to the terminal to provide doctors with image-based evidence for risk detection and assessment.

2. The MRI image analysis method based on a multi-time-point, multi-modal Mamba model as described in claim 1, characterized in that, The S200 includes: S201. Obtain the longitudinal multimodal raw data sequence. First, complete the secondary correction of head movement and gradient distortion in a unified coordinate space. Then, use the SPM12 unified segmentation algorithm to perform bias field correction, gray matter, white matter, and cerebrospinal fluid probability segmentation and MNI152 template spatial registration on each T1-weighted MRI to obtain the standardized brain volume after skull stripping. For PET images, rigid registration with the corresponding MRI is first performed, followed by SUV normalization and Gaussian filtering with a smooth kernel half-width of 6 mm to make the image resolution consistent with the MRI voxel size and reduce noise. After spatial normalization, all images are resampled to 1.5 mm isotropic voxels, and the intensity values ​​are z-score normalized to eliminate the distribution shift caused by the difference between the scanner and the field strength. S202. When PET data is missing at a certain time point, modality completion is performed using a 3D GAN-ViT generative network pre-trained on MRI-PET paired data: the corrected MRI is used as the generator input G to synthesize PET images and their latent spatial features. The generation process is as follows: ; in, This indicates that the generator weights have been pre-trained and frozen on the MRI-PET pairing set and will not be updated during the training phase to ensure that the synthesized features are in the same latent space distribution as the real PET. Thus, aligned MRI and PET latent representations are obtained at each time point; after preprocessing, a set of image data packets corresponding one-to-one with the original scans, with offset correction, spatial registration, intensity normalization and missing data completion, is output.

3. The MRI image analysis method based on a multi-time-point, multi-modal Mamba model as described in claim 1, characterized in that, The S300 includes: S301. Obtain the latent feature vectors of MRI and PET aligned at each time point, as well as the neuropsychological scale data collected during the same visit, and then divide the scale fields into two categories: categorical and numerical. Categorical variables are mapped to dense embeddings via a lookup table, while numerical variables are standardized using Z-scores and then linearly transformed to obtain continuous embeddings. The two types of embeddings are then compared with the scan time difference. After concatenating the sinusoidal position codes, the scale-time joint embedding vector is obtained; S302. The latent features of MRI and PET are embedded together with the above scale-time joint and spliced ​​according to the channel dimension. The overall dimension is compressed to a preset uniform length D through a linear projection layer, thereby forming a multimodal token at that time point: ; In this context, ";" indicates concatenation of channel dimensions. This represents the total number of input feature channels. This indicates the preset unified token dimension. This indicates the time difference between the first scan and the current scan in the user's scan record. and Indicates potential features of MRI and PET. and These represent scale embedding and time coding, respectively. Obtain the time difference consisting of three time points for each target user. , and Then, by arranging the tokens in the order of scanning, a vertical sequence of individuals with a length of 3 is obtained: 。 4. The MRI image analysis method based on a multi-time-point, multi-modal Mamba model as described in claim 1, characterized in that, The S400 includes: Individual longitudinal sequences are input into a time-aware Mamba state-space model to capture the dynamic patterns of brain structure evolution over time. The Mamba state-space model consists of six stacked Mamba Blocks. Each Block sequentially performs RMS normalization, selective scanning, and residual connections. The selective scanning SSM model achieves long-range dependency modeling with linear complexity O(L) through a trainable state matrix A and input-related gating mechanisms B and C. Simultaneously, a time decay factor is introduced during the memory update phase, causing the contribution of observations with distances greater than a threshold to the current state to decrease exponentially, and a time decay coefficient is introduced. Enhance time difference The impact, among which >1, thus explicitly reflecting the clinical hypothesis that "the longer the time, the weaker the impact"; Specifically, first, the time difference corresponding to each token in the sequence is... Convert to attenuation coefficient: ; in, , These represent the time points corresponding to the current memory and the memory to be updated, respectively. Determined through validation set optimization. ∈(0,1) controls the degree of forgetting; Before the SSM state derivation, the memory cell C is weighted and corrected: ; The corrected state participates in subsequent gating calculations; then, after being refined layer by layer through six blocks, the model output is a hidden state sequence of equal length to the input. ,in, and The one-to-one correspondence integrates cross-modal structural information and carries individual-specific temporal evolutionary characteristics.

5. The MRI image analysis method based on a multi-time-point multimodal Mamba model as described in claim 1, characterized in that, The S500 includes: S501, Sequence of hidden states As the query vector Q: The feature matrices flattened from the original MRI and PET images, respectively. , Cross-attention calculation is performed to map global temporal features to local pixel space; Specifically, the MRI and PET three-dimensional images at corresponding time points are first flattened along the spatial dimension into a two-dimensional feature matrix, resulting in... , and , ; MRI branch: , ; PET branch: , ; Subsequently, the query vector Q is mapped to the same channel dimension as the image features using linear projection, and the attention weights for the MRI and PET branches are calculated separately: ; ; in, Indicates the channel dimension, used for scaling to prevent gradient saturation; The weighted features of the MRI and PET branches are calculated and compared with the original data using a residual method. The fusion process creates a unified expression that incorporates both temporal evolutionary information and preserves spatial pathological details. The fusion output is: 。