Transformer-model-based modeling and feature clustering method for longitudinal neuroimaging data

WO2026200147A1PCT designated stage Publication Date: 2026-10-01NANJING MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/146919
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-08-25
Filing Date
2025-12-30
Publication Date
2026-10-01

Smart Images

  • Figure CN2025146919_01102026_PF_FP_ABST
    Figure CN2025146919_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a Transformer-model-based modeling and feature clustering method for longitudinal neuroimaging data. The method comprises: acquiring brain functional connectivity matrices, and generating second-order-difference brain networks; partitioning each second-order-difference brain network into a plurality of interaction sub-matrices; encoding the interaction sub-matrices to obtain an interaction embedding vector of a subject; matching the interaction embedding vector of the subject with memory vectors, in order to obtain a new representation vector of the subject and a corresponding score list; using the new representation vector to perform graph reconstruction; and on the basis of the score list, determining a cluster center by means of a K-means method. In the present invention, deep learning is combined with longitudinal neuroimaging analysis to construct an efficient and automated method for temporal feature modeling and clustering analysis. The method can serve as an intelligent auxiliary analysis tool for psychiatric disorder research, and has a significant scientific research value and broad application prospects in terms of supporting research on individual differences and connectivity pattern exploration.
Need to check novelty before this filing date? Find Prior Art

Description

Transformer-based methods for longitudinal neuroimaging data modeling and feature clustering Technical Field

[0001] This invention belongs to the interdisciplinary field of psychiatry, medical image analysis and machine learning. Specifically, it involves a method for longitudinal neuroimaging data modeling and feature clustering based on the Transformer model, which is applicable to the structured representation and pattern mining of brain imaging data at multiple time points. Background Technology

[0002] Mental illnesses exhibit high clinical manifestations and biological heterogeneity, posing numerous challenges to their diagnosis and treatment. Existing research largely focuses on grouping based on symptomatological characteristics; however, this subjective information-based grouping approach lacks a dynamic characterization of underlying neurobiological processes and fails to delve into the intrinsic mechanisms and progression of the diseases. Therefore, while symptom-based classification aids in disease analysis and treatment, it often becomes a major bottleneck in basic research and mechanistic exploration of mental illnesses, and also limits the implementation of precision medicine goals.

[0003] On the other hand, previous cross-sectional studies have shown that different patients with mental illnesses may possess unique underlying biological mechanisms, which are not static but dynamically change during disease treatment. More importantly, the developmental trajectory of mental illness may also differ among individuals. For example, structural and functional changes in the brain during disease progression are closely related to alterations in clinical symptoms, indicating that a deeper understanding of these changes is crucial for improving treatment outcomes and predicting disease progression. Therefore, extracting characteristic patterns reflecting disease evolution from a neuroimaging perspective can not only help understand individual differences but also provide critical information for precision diagnosis and treatment.

[0004] In recent years, with the development of neuroimaging technology, especially the widespread application of longitudinal neuroimaging data, researchers have been able to track the evolution of individual brain functional networks over time by analyzing imaging data at different time points. Compared with traditional cross-sectional studies, longitudinal imaging studies provide a more dynamic perspective, revealing the evolutionary trajectory of brain structure and function in patients with different mental illnesses during the course of their diseases. However, despite the richer information provided by longitudinal neuroimaging data, traditional statistical analysis or shallow machine learning methods have certain limitations in processing high-dimensional, temporally complex neuroimaging data, making it difficult to effectively extract deep-seated time-dependent structures and latent feature representations. To address this issue, deep learning technology, especially temporal modeling based on Transformer models, has shown superior performance in the field of medical image analysis. Transformer models, due to their powerful sequence modeling capabilities and self-attention mechanism, have been widely applied in neuroimaging research. Compared with traditional methods, their advantage lies in their ability to effectively handle information interaction across time points and capture the dynamic evolution of hidden patterns.

[0005] Therefore, how to leverage the powerful modeling capabilities of Transformer to develop a method with temporal modeling and feature induction capabilities, capable of identifying potential evolutionary patterns of individuals in the course of disease from a dynamic perspective, and thus providing strong technical support for individual differences research, scientific classification, and precision medicine of mental illnesses, is a pressing problem that needs to be solved. Summary of the Invention

[0006] This invention aims to overcome the shortcomings of existing technologies in multi-time point neuroimaging data modeling and individualized feature extraction, and provides a longitudinal neuroimaging data modeling and feature clustering method based on the Transformer model. This method can model and cluster change patterns across time points in neuroimaging data, thereby extracting discriminative connectivity feature evolution trajectories to support subsequent individual difference analysis and data-driven grouping studies.

[0007] To achieve the above objectives, this invention provides a method for longitudinal neural imaging data modeling and feature clustering based on the Transformer model, comprising the following steps:

[0008] S1. Obtain the brain functional connectivity matrix of the subject at three time points, and calculate the second-order difference relationship between the connectivity matrices of adjacent time points to generate a second-order difference brain network.

[0009] S2. Divide the brain functional connectivity matrix into multiple resting state sub-networks, introduce a Transformer-based multi-head attention mechanism, capture the interaction relationship between each sub-network, and divide the subject's second-order difference brain network into multiple interaction sub-matrices between sub-networks.

[0010] S3. Encode the interaction submatrix using an encoder that includes graph convolution operations and interaction attention operations to obtain the subject's interaction embedding vector.

[0011] S4. Construct a memory bank containing multiple memory vectors, match the subject's interaction embedding vector with the memory vectors to obtain the subject's new representation vector, and a score list of the subject's interaction embedding vector matching each memory vector;

[0012] S5. Introduce a graph decoder to reconstruct the graph using a new representation vector; simultaneously, based on the obtained score list, determine the cluster centers using the K-means method.

[0013] A further preferred embodiment of the present invention is that, in step S1, the brain functional connectivity matrices of the subject at three time points are obtained, and the second-order difference relationship between the connectivity matrices at adjacent time points is calculated to generate a second-order difference brain network; specifically:

[0014] Obtain the brain functional connectivity matrix of the subjects at three time points, and perform standardized preprocessing on the connectivity matrix;

[0015] Based on adjacent time points, the second-order difference relation of the connectivity matrix is ​​calculated using the following formula:

[0016] ;

[0017] in, This represents the brain functional connectivity matrix at time point T.

[0018] The generated second-order difference brain network data is represented as follows , where n represents the sample size. The integrated data representing the two second-order difference relationships of the s-th subject. It is the second-order difference relationship data between the first and second time points. It is the second-order difference relationship data between the second and third time points.

[0019] As a preferred option, step S2 uses the Yeo-7 network map to divide the brain functional connectivity matrix into 7 resting state sub-networks, namely the visual network (VIS), the somatic motor network (SMN), the dorsal attention network (DAN), the ventral attention network (VAN), the limbic network (LIM), the frontoparietal network (FPN), and the default mode network (DMN).

[0020] The introduction of a Transformer-based multi-head attention mechanism to capture the interaction relationships between sub-networks specifically involves:

[0021] The fMRI time series of each sub-network is divided into segments of fixed length. Each segment is mapped to a feature vector of fixed dimension through linear transformation, and learnable positional encoding is added to preserve spatiotemporal information.

[0022] The attention head is computed independently for each sub-network, the correlation weights between sub-networks are calculated, the features are weighted and fused after normalization by softmax, and then the outputs of the attention heads are concatenated and integrated through a linear layer.

[0023] Spatial self-attention is applied to the attention-weighted features, and the dimensionality is compressed through a fully connected layer to output a representation that includes the interaction features between subnetworks.

[0024] Preferably, the encoder in step S3 includes a graph convolution operation. An interactive attention operation and aggregate functions The second-order difference brain network of the p-th subnetwork of the s-th subject was generated using an encoder. Mapping to higher-dimensional space , represented as Specifically, this includes:

[0025] Convolution operation on graphs Defined as ,in These represent the learnable weights of the horizontal and vertical filters, respectively. The weights of the last layer of information aggregation operation in graph convolution are used to calculate the network interaction embedding of each second-order difference brain network of the s-th subject, expressed as:

[0026] ;

[0027] Multi-head attention is used to focus on the relationships between different network nodes. The attention embedding is obtained through the QKV attention mechanism and is represented as follows: ;

[0028] Using aggregate functions The network interaction of the embedded attention mechanism of each second-order differential brain network is embedded. By combining these data, we obtain the participants' interaction embedding vectors. , represented as:

[0029] .

[0030] Preferably, in step S4, a memory bank containing multiple memory vectors is constructed, and the subject's interaction embedding vector is matched with the memory vectors to obtain a new representation vector for the subject, as well as a score list of the matching between the subject's interaction embedding vector and each memory vector; specifically:

[0031] initialization memory vectors As a brain biomarker library;

[0032] The subject's interaction embedding vector is matched with the memory vector.

[0033] Each participant's interaction embedding vector is defined as a weighted combination of memory vectors from a biomarker library. A matching mechanism is introduced to interact the participant's interaction embedding vector with the memory vectors:

[0034] ;

[0035] ;

[0036] ;

[0037] in, These are learnable weights; It is the first The memory vector is the first one. The weight assigned to the nth subject represents the weight of the nth subject. The first subject Each participant was assigned a score list based on the pattern. ; It is a new representation vector generated by weighted combination of memory vectors.

[0038] Preferably, a graphics decoder is introduced in step S5, through... Output the reconstructed graphics.

[0039] Preferably, after image reconstruction, the encoder, memory vector, and reconstructed image are optimized using a comprehensive loss function, expressed as:

[0040] ;

[0041] in, To compare the loss, it is used to facilitate the learning of interactive embedding vectors;

[0042] The triple loss function, used to update the memory vector, is expressed as:

[0043] ;

[0044] in, It is based on weight The first two memory vectors of the s-th subject in the sorting will Treat it as an anchor to ensure Compare Store more information from the anchor;

[0045] For the reconstructed image A reconstruction loss was used to approximate the input map. The reconstruction loss is expressed as:

[0046] ;

[0047] The orthogonal regularization loss of the memory vector is expressed as:

[0048] .

[0049] In another aspect, the present invention provides a non-transitory computer-readable storage medium having computer instructions stored thereon, which cause a computer to execute the above-described method for longitudinal neural imaging data modeling and feature clustering based on the Transformer model.

[0050] In another aspect, the present invention provides an electronic device, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus, and the processor calls logical instructions in the memory to execute the above-mentioned method for longitudinal neural imaging data modeling and feature clustering based on the Transformer model.

[0051] In another aspect, the present invention provides a computer program product comprising a computer program stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer performs the aforementioned method for longitudinal neural imaging data modeling and feature clustering based on the Transformer model.

[0052] Compared with existing technologies, the longitudinal neural imaging data modeling and feature clustering method based on the Transformer model of this invention has the following advantages:

[0053] (1) Deep modeling of longitudinal dynamic features: This invention introduces the Transformer model to perform deep modeling of neuroimaging data of individuals with mental illness at multiple time points, effectively capturing the temporal evolution characteristics of brain functional connectivity. Compared with traditional static analysis methods, this invention overcomes the limitation of ignoring temporal dimension information, providing more comprehensive and refined trends and evolutionary trajectories of brain functional changes. This dynamic modeling approach helps to reveal the potential neurobiological mechanisms in the disease process and promotes a deeper understanding of the relationship between disease development and clinical symptoms.

[0054] (2) Enhanced sensitivity and expressive power of connectivity pattern grouping: This invention utilizes temporal feature embedding extracted by Transformer and combines it with clustering algorithms for individual difference analysis, which can accurately identify significant structural differences in dynamic connectivity features among different patients. Compared with traditional methods, this invention exhibits higher sensitivity and stronger expressive power in connectivity pattern recognition, effectively distinguishing subgroups that exhibit different evolutionary trajectories during the course of the disease, and improving the recognition accuracy and interpretability of potential connectivity evolution patterns.

[0055] (3) Good adaptability and scalability: This method is compatible with multiple longitudinal neuroimaging modalities (such as DTI), and can be widely applied to different research scenarios related to mental illnesses. It has strong cross-task transferability and practical application potential. In addition to mental illness research, this invention is also applicable to other tasks that require long-term tracking and dynamic modeling of individual characteristics of neurological diseases.

[0056] (4) Advanced technology and high degree of automation: This invention adopts the Transformer deep learning framework, which can automatically extract key connectivity features from high-dimensional, complex, and multi-time-point neuroimaging data. This method significantly reduces the reliance on manual feature design, improves the efficiency of data processing and the intelligence level of the system, and avoids the manual intervention and subjective bias in traditional methods. In addition, the high degree of automation of this method makes it have good prospects for system implementation and engineering transformation.

[0057] In summary, this invention combines deep learning with longitudinal neuroimaging analysis to construct an efficient and automated method for temporal feature modeling and clustering analysis. This method can serve as an intelligent auxiliary analysis tool in mental illness research and has significant scientific application value and broad prospects for promotion in supporting research on individual differences and exploration of connection patterns. Attached Figure Description

[0058] Figure 1 is a flowchart of the longitudinal neuroimaging data modeling and feature clustering method based on the Transformer model of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0060] The following description, with reference to Figure 1, illustrates the longitudinal neuroimaging data modeling and feature clustering method based on the Transformer model provided by this invention.

[0061] Example 1: This example provides a method for longitudinal neuroimaging data modeling and feature clustering based on the Transformer model. The method consists of two stages, each composed of multiple ordered steps, with clear data flow and logical dependencies between the steps.

[0062] As shown in Figure 1, the first stage is the encoding stage based on sub-network connections, and the second stage is the learning network for building a sub-network interaction pattern library.

[0063] The first stage aims to construct subnetwork connectivity representations that express the dynamic features of brain region interactions from longitudinal neuroimaging data, providing structured input for subsequent clustering analysis. Specifically, it includes the following steps:

[0064] S1, Construction of second-order differential brain network.

[0065] Obtain the brain functional connectivity matrix of the subjects at three time points, and perform standardized preprocessing on the connectivity matrix;

[0066] Based on adjacent time points, the second-order difference relation of the connectivity matrix is ​​calculated using the following formula:

[0067] ;

[0068] in, This represents the brain functional connectivity matrix at time point T.

[0069] The generated second-order difference brain network data is represented as follows , where n represents the sample size. The integrated data representing the two second-order difference relationships of the s-th subject. It is the second-order difference relationship data between the first and second time points. It is second-order difference relational data between the second and third time points. This network is designed to highlight regions in the entire brain network that have undergone significant changes between two consecutive time points.

[0070] S2, Brain Sub-network Interactive Learning.

[0071] Based on the brain region segmentation template, the Yeo-7 network atlas was used to divide the brain functional connectivity matrix into 7 resting state sub-networks, namely the visual network (VIS), the somatic motor network (SMN), the dorsal attention network (DAN), the ventral attention network (VAN), the limbic network (LIM), the frontoparietal network (FPN), and the default mode network (DMN).

[0072] This embodiment studies the interaction features between sub-networks, rather than relying on manually designed features, which can more effectively capture salient features among subjects, thus achieving more accurate clustering. Compared with traditional methods, this method can automatically learn potential interaction relationships from raw data, enhancing the characterization of the dynamic evolution of complex functional brain networks. Learning the connections between sub-networks is crucial when modeling this interaction process using graph theory, and Transformer-based methods have proven to be robust for modeling these connections. Therefore, this embodiment introduces a Transformer-based multi-head attention brain network interaction module to capture interaction features helpful for subsequent analysis.

[0073] This embodiment introduces a Transformer-based multi-head attention mechanism. The specific method for capturing the interaction relationships between sub-networks is as follows:

[0074] The fMRI time series of each sub-network is divided into segments of fixed length. Each segment is mapped to a feature vector of fixed dimension through linear transformation, and learnable positional encoding is added to preserve spatiotemporal information.

[0075] The attention head is computed independently for each sub-network, the correlation weights between sub-networks are calculated, the features are weighted and fused after normalization by softmax, and then the outputs of the attention heads are concatenated and integrated through a linear layer.

[0076] Spatial self-attention is applied to the attention-weighted features, and the dimensionality is compressed through a fully connected layer to output a representation that includes the interaction features between subnetworks.

[0077] S3, Interactive Embedded Vector Generation.

[0078] This embodiment constructs an encoder that includes a graph convolution operation. The first step is to learn the representation of subnetwork connectivity information; the second step is interactive attention computation. To effectively simulate the interactions of brain subnetworks and aggregation functions Further extract relevant information. The encoder is used to extract the second-order difference brain network of the p-th subnetwork of the s-th subject. Mapping to higher-dimensional space , represented as .

[0079] Specifically, graph convolution operations Defined as ,in These represent the learnable weights of the horizontal and vertical filters, respectively. The weights of the last layer of information aggregation operation in graph convolution are used to calculate the network interaction embedding of each second-order difference brain network of the s-th subject, expressed as:

[0080] ;

[0081] Then, the encoder utilizes a multi-head attention mechanism to focus on the relationships between different network nodes, effectively simulating their interactions. The attention embedding is obtained through the QKV attention mechanism, represented as... ;

[0082] Next, using aggregate functions The network interaction of the embedded attention mechanism of each second-order differential brain network is embedded. By combining these data, we obtain the participants' interaction embedding vectors. , represented as:

[0083] .

[0084] The second stage, based on the aforementioned embedding representation, introduces representative brain connectivity prototypes and clustering mechanisms to achieve dynamic modeling of individual differences and discovery of potential patterns.

[0085] S4, Brain biological model library learning.

[0086] Inspired by memory networks, this embodiment proposes the construction of a brain biomarker library. This library can store typical evolutionary patterns, facilitating personalized matching with individual subjects.

[0087] initialization memory vectors As a library of brain biomarkers, these can be updated during the optimization process. These memory vectors are orthogonal to each other and can remember different information. Therefore, orthogonal regularization is introduced. .

[0088] The interaction between the biomarker library and different participants is achieved through a search and matching process. It is assumed that the high-dimensional embedding of each participant can be represented by a weighted combination of memory vectors from the biomarker library, and the specificity among participants is reflected by assigning different weights to these memory vectors. Specifically, the following matching mechanism is introduced:

[0089] ;

[0090] ;

[0091] ;

[0092] in, These are learnable weights; It is the first The memory vector is the first one. The weights assigned to each participant, more specifically... Also represents the first The first subject Each pattern score, therefore, the pattern score of the subject represented by the memory vector can be expressed as a list of combined values. These individual scores sensitively reflect individual characteristics; It is a new representation vector generated by weighted combination of memory vectors.

[0093] S5, Structural Reconstruction.

[0094] To enable the model to focus on more useful meta-path-based interaction information, the input graph is reconstructed using bank-memory vectors. Specifically, a graph decoder is introduced, through... Output the reconstructed graphics.

[0095] S6. Optimize target design.

[0096] After image reconstruction, the encoder, memory vector, and reconstructed image are optimized using a comprehensive loss function, as follows:

[0097] ;

[0098] in, To compare the loss, it is used to facilitate the learning of interactive embedding vectors;

[0099] The triple loss function, used to update the memory vector, is expressed as:

[0100] ;

[0101] in, It is based on weight The first two memory vectors of the s-th subject in the sorting will Treat it as an anchor to ensure Compare Store more information from the anchor;

[0102] For the reconstructed image A reconstruction loss was used to approximate the input map. The reconstruction loss is expressed as:

[0103] ;

[0104] The orthogonal regularization loss of the memory vector is expressed as:

[0105] .

[0106] S7. Score-based interaction pattern clustering.

[0107] By matching the interaction embeddings with memory vectors, each participant was assigned a score list. Analysis of these score lists revealed that the participants' scores were similar, but they belonged to different groups. The K-means method was used to determine the cluster centers.

[0108] Example 2: This example provides a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute a longitudinal neural imaging data modeling and feature clustering method based on the Transformer model. The method includes the following steps:

[0109] S1. Obtain the brain functional connectivity matrix at at least three time points for the same subject, and calculate the second-order difference relationship between the connectivity matrices at adjacent time points to generate a second-order difference brain network.

[0110] S2. Divide the brain functional connectivity matrix into multiple resting state sub-networks, introduce a Transformer-based multi-head attention mechanism, capture the interaction relationship between each sub-network, and divide the second-order difference brain network into multiple interaction sub-matrices between sub-networks.

[0111] S3. Encode the interaction submatrix using an encoder that includes graph convolution operations and interaction attention operations to obtain the subject's interaction embedding vector.

[0112] S4. Construct a memory bank containing multiple memory vectors, match the subject's interaction embedding vector with the memory vectors to obtain the subject's new representation vector, and a score list of the subject's interaction embedding vector matching each memory vector;

[0113] S5. Introduce a graph decoder to reconstruct the graph using a new representation vector; simultaneously, based on the obtained score list, determine the cluster centers using the K-means method.

[0114] Example 3: This example provides an electronic device that may include a processor, a communication interface, memory, and a communication bus. The processor, communication interface, and memory communicate with each other via the communication bus. The processor can call logical instructions from the memory to execute a Transformer-based longitudinal neural imaging data modeling and feature clustering method. This method includes the following steps:

[0115] S1. Obtain the brain functional connectivity matrix at at least three time points for the same subject, and calculate the second-order difference relationship between the connectivity matrices at adjacent time points to generate a second-order difference brain network.

[0116] S2. Divide the brain functional connectivity matrix into multiple resting state sub-networks, introduce a Transformer-based multi-head attention mechanism, capture the interaction relationship between each sub-network, and divide the second-order difference brain network into multiple interaction sub-matrices between sub-networks.

[0117] S3. Encode the interaction submatrix using an encoder that includes graph convolution operations and interaction attention operations to obtain the subject's interaction embedding vector.

[0118] S4. Construct a memory bank containing multiple memory vectors, match the subject's interaction embedding vector with the memory vectors to obtain the subject's new representation vector, and a score list of the subject's interaction embedding vector matching each memory vector;

[0119] S5. Introduce a graph decoder to reconstruct the graph using a new representation vector; simultaneously, based on the obtained score list, determine the cluster centers using the K-means method.

[0120] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0121] Example 4: This example provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform a longitudinal neural imaging data modeling and feature clustering method based on the Transformer model. The method includes the following steps:

[0122] S1. Obtain the brain functional connectivity matrix at at least three time points for the same subject, and calculate the second-order difference relationship between the connectivity matrices at adjacent time points to generate a second-order difference brain network.

[0123] S2. Divide the brain functional connectivity matrix into multiple resting state sub-networks, introduce a Transformer-based multi-head attention mechanism, capture the interaction relationship between each sub-network, and divide the second-order difference brain network into multiple interaction sub-matrices between sub-networks.

[0124] S3. Encode the interaction submatrix using an encoder that includes graph convolution operations and interaction attention operations to obtain the subject's interaction embedding vector.

[0125] S4. Construct a memory bank containing multiple memory vectors, match the subject's interaction embedding vector with the memory vectors to obtain the subject's new representation vector, and a score list of the subject's interaction embedding vector matching each memory vector;

[0126] S5. Introduce a graph decoder to reconstruct the graph using a new representation vector; simultaneously, based on the obtained score list, determine the cluster centers using the K-means method.

[0127] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

A longitudinal neural image data modeling and feature clustering method based on a Transformer model, characterized in that, Includes the following steps: S1. Obtain the brain functional connectivity matrix of the subject at three time points, and calculate the second-order difference relationship between the connectivity matrices of adjacent time points to generate a second-order difference brain network. S2. Divide the brain functional connectivity matrix into multiple resting state sub-networks, introduce a Transformer-based multi-head attention mechanism, capture the interaction relationship between each sub-network, and divide the subject's second-order difference brain network into multiple interaction sub-matrices between sub-networks. S3. Encode the interaction submatrix using an encoder that includes graph convolution operations and interaction attention operations to obtain the subject's interaction embedding vector. S4. Construct a memory bank containing multiple memory vectors, match the subject's interaction embedding vector with the memory vectors to obtain the subject's new representation vector, and a score list of the subject's interaction embedding vector matching each memory vector; S5. Introduce a graph decoder to reconstruct the graph using a new representation vector; simultaneously, determine the cluster centers using the K-means method based on the obtained score list. The Transformer model-based longitudinal neural image data modeling and feature clustering method according to claim 1, characterized in that, In step S1, the brain functional connectivity matrices of the subject at three time points are obtained, and the second-order difference relationship between the connectivity matrices at adjacent time points is calculated to generate a second-order difference brain network; specifically: Obtain the brain functional connectivity matrix of the subjects at three time points, and perform standardized preprocessing on the connectivity matrix; Based on adjacent time points, the second-order difference relation of the connectivity matrix is ​​calculated using the following formula: ; wherein This represents the brain functional connectivity matrix at time point T; The generated second-order difference brain network data is represented as n represents the sample size, integrated data representing the second-order difference relationship for the s-th subject, is second-order difference relationship data of the first time point and the second time point, It is the second-order difference relationship data between the second and third time points. The Transformer model-based longitudinal neural image data modeling and feature clustering method according to claim 1, characterized in that, Step S2 uses the Yeo-7 network map to divide the brain functional connectivity matrix into 7 resting state sub-networks, namely the visual network (VIS), somatic motor network (SMN), dorsal attention network (DAN), ventral attention network (VAN), limbic network (LIM), frontoparietal network (FPN), and default mode network (DMN). The introduction of a Transformer-based multi-head attention mechanism to capture the interaction relationships between sub-networks specifically involves: The fMRI time series of each sub-network is divided into segments of fixed length. Each segment is mapped to a feature vector of fixed dimension through linear transformation, and learnable positional encoding is added to preserve spatiotemporal information. The attention head is computed independently for each sub-network, the correlation weights between sub-networks are calculated, the features are weighted and fused after normalization by softmax, and then the outputs of the attention heads are concatenated and integrated through a linear layer. Spatial self-attention is applied to the attention-weighted features, and the dimensionality is compressed through a fully connected layer to output a representation that includes the interaction features between subnetworks. The Transformer model-based longitudinal neuroimaging data modeling and feature clustering method according to claim 1, characterized in that, The encoder in step S3 comprises a graph convolution operation , an interaction attention operation and aggregation functions ; using the encoder to encode the second-order differential brain network of the pth subnetwork of the s th subject Mapping to high dimensional space , is represented as Specifically, it includes: Applying graph convolution operations defined as wherein learnable weights representing a horizontal filter and a vertical filter, respectively, The weights of the last layer of information aggregation operation in graph convolution are used to calculate the network interaction embedding of each second-order difference brain network of the s-th subject, expressed as: ; The multi-head attention mechanism is used to focus on the relationship between different network nodes, and the attention embedding is obtained through the QKV attention mechanism, represented as ; Utilizing aggregation functions embedding attention mechanism of each second-order differential brain network aggregating to obtain an interactive embedding vector for the subject , is represented as: 。 The Transformer model-based longitudinal neural image data modeling and feature clustering method according to claim 4, characterized in that, Step S4 involves constructing a memory bank containing multiple memory vectors, matching the subject's interaction embedding vector with the memory vectors to obtain the subject's new representation vector, and a score list of the matching between the subject's interaction embedding vector and each memory vector; specifically: initialization a memory vector As a brain biomarker library; The subject's interaction embedding vector is matched with the memory vector. Each participant's interaction embedding vector is defined as a weighted combination of memory vectors from a biomarker library. A matching mechanism is introduced to interact the participant's interaction embedding vector with the memory vectors: ; ; ; wherein, are learnable weights; is the first a memory vector is the first The weight assigned to each subject represents the number of times the subject was in the top 10% of the group the first a mode score, each subject being assigned a list of scores ; It is a new representation vector generated by weighted combination of memory vectors. The Transformer model-based longitudinal neural image data modeling and feature clustering method according to claim 5, characterized in that, The introduction of the graphics decoder in step S5 is done by Output the reconstructed graphics. The Transformer model-based longitudinal neural image data modeling and feature clustering method according to claim 6, characterized in that, After image reconstruction, the encoder, memory vector, and reconstructed image are optimized using a comprehensive loss function, as follows: ; wherein, To compare the loss, it is used to facilitate the learning of interactive embedding vectors; The triple loss function, used to update the memory vector, is expressed as: ; wherein, is according to the weight The first two memory vectors of the s-th subject in the ranking are combined to form a new vector, which is then used to determine the score of the s-th subject in the ranking. considered as anchor, make sure Ratio Store more information from the anchor; For the reconstructed picture A reconstruction loss was used to approximate the input map. The reconstruction loss is expressed as: ; The orthogonal regularization loss of the memory vector is expressed as: 。 A non-transitory computer-readable storage medium, characterized by It stores computer instructions that cause the computer to execute the longitudinal neuroimaging data modeling and feature clustering method based on the Transformer model as described in any one of claims 1-7. An electronic device, characterized by include: The system includes a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other via the communication bus. The processor calls logical instructions from the memory to execute the longitudinal neural imaging data modeling and feature clustering method based on the Transformer model as described in any one of claims 1-7. A computer program product, characterized by The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer performs the longitudinal neural imaging data modeling and feature clustering method based on the Transformer model as described in any one of claims 1-7.