A method for predicting autism based on nuclear magnetic resonance images and finding biomarkers
By using multi-headed attention encoder and graph generator in NMR image processing, the nonlinear relationship between the regions of interest in the brain is learned, and combined with graph convolution and top-k pooling operations, the data consistency, computational complexity and interpretability problems of nuclear magnetic resonance images predicting autism in the prior art are solved, achieving higher prediction accuracy and biomarker discovery.
Patent Information
- Application Number
- CN202211000237.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-08-19
AI Technical Summary
The prior art has problems such as inconsistent data attributes, explosive computational dimensions, limited training samples and poor interpretability in predicting autism using NMR images.
By obtaining brain MRI images, a preset brain region segmentation template was used to segment the brain region of interest, and a time series recording the changes in blood oxygen level in each brain region of interest was obtained. Using multi-headed attention encoder and graph generator, nonlinear relationships between ROIs are learned, and the interpretability and computational efficiency of the model are improved through graph convolution and top-k pooling operations.
Improved the accuracy of MRI images in autism prediction, and found important brain regions related to autism, providing reliable biomarkers, providing strong support for clinical diagnosis and treatment.
Smart Images

Figure CN115331809B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of medical image processing, and specifically relates to a method for predicting autism based on magnetic resonance images and finding biomarkers. Background Art
[0002] Autism spectrum disorder (ASD) is a lifelong developmental disability, clinically characterized by impaired social communication, repetitive behaviors, and restricted interests. Recent research has proposed that the combination of genetics and living environment may be the main cause of ASD. However, neuroscientists are still looking for reliable and objective methods to improve the current diagnostic methods that tend to be subjective and inaccurate, such as questionnaires and clinical observations. According to the Autism and Developmental Disabilities Monitoring (ADDM) Network, approximately 1 in every 59 children has ASD, which reveals the urgent need to find objective biomarkers to facilitate clinical diagnosis and treatment.
[0003] Using magnetic resonance image preprocessing tools, brain functional connectivity data can be obtained, and with this as input, many deep learning methods have emerged. However, most related studies have the following problems:
[0004] (1) Inconsistencies in all data attributes due to different scanners or research groups at multiple locations.
[0005] (2) Computational dimensionality explosion due to the millions of voxels generated by each fMRI scan.
[0006] (3) A limited number (tens or hundreds) of training samples.
[0007] (4) Poor interpretability.
[0008] Therefore, it is particularly necessary to conduct in-depth research on magnetic resonance images to find objective biomarkers to facilitate clinical diagnosis and treatment. Summary of the Invention
[0009] This application proposes a method for predicting autism based on magnetic resonance images and finding biomarkers, promoting the research process of ASD diagnosis based on magnetic resonance images, and providing reliable tools and biomarkers for clinical medicine to diagnose and treat ASD.
[0010] To achieve the above objective, the technical solution of this application is as follows:
[0011] A method for predicting autism based on magnetic resonance images and finding biomarkers, comprising:
[0012] Obtain the magnetic resonance imaging (MRI) of the brain of the detection object, segment the regions of interest (ROIs) in the brain using a preset brain segmentation template, calculate the average of all voxel sequence values contained in each ROI in the brain to obtain the time series recording the blood oxygenation level change in each ROI in the brain;
[0013] Perform a correlation operation on the obtained time series using the Pearson correlation coefficient calculation method to obtain the feature vectors corresponding to each ROI in the brain;
[0014] Use the one-hot encoding method to create a position vector for each ROI in the brain, and through the learning of a multi-layer perceptron for the position vector, obtain the final position vector;
[0015] Take each ROI in the brain as a node, and add the position vector and the feature vector of each ROI in the brain correspondingly to obtain the node feature vector;
[0016] Input the time series into a multi-head attention encoder, input the obtained encoded features into a graph generator, take each ROI in the brain as a node, and obtain the relational adjacency matrix of the graph formed by each node;
[0017] Perform a graph convolution operation on the node feature vector matrix composed of all node feature vectors and the relational adjacency matrix to obtain the final node feature vector matrix containing the relational adjacency node relationship;
[0018] Perform a top-k pooling operation on the final node feature vector matrix to obtain the importance scores of all nodes, select the top k nodes with the highest importance scores, and flatten the feature vectors of the selected nodes to obtain the feature vector for prediction;
[0019] Input the feature vector for prediction into a multi-layer perceptron to obtain the prediction result.
[0020] Furthermore, the formula representing the graph convolution operation is as follows:
[0021] h S = BatchNorm1D(ReLU(Ah S-1 W S ))
[0022] where BatchNorm1D represents a one-dimensional batch normalization operation, ReLU is an activation function, A is the relational adjacency matrix, h S represents the node feature vector matrix obtained after S times of graph convolution, and h 0 is the initial node feature vector matrix.
[0023] Furthermore, the step of performing a top-k pooling operation on the final node feature vector matrix to obtain the importance scores of all nodes includes:
[0024] Multiply the final node feature vector matrix by the pooling weight vector to obtain the importance scores of all nodes.
[0025] Furthermore, the method for predicting autism and finding biomarkers based on nuclear magnetic resonance images further includes:
[0026] Calculate the joint loss for network training through the joint loss function, and the joint loss function is as follows:
[0027] L = L ce + αL intra + βL inter + γL TPK
[0028] where α, β, and γ are the weight coefficients of each loss;
[0029]
[0030] where y is the label value for autism, M represents the number of samples in each batch during batch training, z i is the score for the i-th example being autism, and Sigmoid is the activation function;
[0031]
[0032] where c represents a label, C is the set of all labels, S c = {i|Y i,c = 1} is the set of examples with label c, and i represents an example; μ c is the mean of the weights of the adjacency matrices of all examples in S c ; A i represents the weight value of the adjacency matrix of example i; represents S c the variance of the weights in the adjacency matrices of all examples in;
[0033]
[0034] where a and b are two labels in C, and μ a and μ b are the mean of the adjacency matrix of the set of examples with label a and the mean of the adjacency matrix of the set of examples with label b, respectively;
[0035]
[0036] where N represents the number of nodes in each input example; s m,i represents the importance score of the i-th node in the m-th input example.
[0037] A method for predicting autism and finding biomarkers based on nuclear magnetic resonance images proposed in this application encodes and constructs a graph through the time series of ROIs using a multi-head self-attention model, learning the non-linear relationships between ROIs instead of simply using the Pearson correlation coefficient method to obtain linear relationships; using the top-k pooling method to improve the interpretability of the model and reduce the computational complexity; adding a position feature vector to distinguish nodes at different positions and introducing the position information of nodes for model prediction; using three loss functions in addition to the standard binary cross-entropy loss function to regularize the training effect of the model. Finally, the obtained network model is used to detect the brain magnetic resonance images of the detection object, improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the network model structure of this application;
[0039] Figure 2 It is a functional flowchart of the top-k pooling of this application;
[0040] Figure 3 It is a list of important brain regions related to autism in the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0042] Based on the nuclear magnetic resonance images of the detection object, through preprocessing the images and using a brain segmentation method for segmentation, a series of time series of blood oxygen level changes in each brain region are obtained. According to this time series, a graph is constructed, and with graph convolution as the core, autism prediction is performed on the detection object, and based on the prediction result data of all detection objects, the brain regions significantly related to ASD are found. Provide reliable tools and biomarkers for clinical medicine.
[0043] As Figure 1 shown, a method for predicting autism and finding biomarkers based on nuclear magnetic resonance images includes:
[0044] S1. Obtain the brain magnetic resonance images of the detection object, perform segmentation on the regions of interest in the brain using a preset brain segmentation template, and calculate the average of all voxel sequence values included in each region of interest in the brain to obtain the time series of blood oxygen level changes recorded in each region of interest in the brain.
[0045] In this embodiment, first, obtain the magnetic resonance imaging (MRI) of the brain of the detection object, and perform preprocessing using a domain-related method (the Configurable Pipeline for the Analysis of Connectomes, CPAC). The preprocessing process includes: normalizing the voxel intensity values measured by magnetic resonance; finding the one-to-one correspondence between voxels before and after movement and correcting the movement; and slice timing correction.
[0046] In this embodiment, select appropriate brain segmentation templates (Atlases) such as AAL, CC200, etc., and perform segmentation of regions of interest (ROIs) in the brain on the preprocessed image, that is, segment the preprocessed image according to the segmentation regions defined on the brain segmentation template. After segmentation, calculate the average of all voxel sequence values included in each segmentation region to obtain the time series t recording the change in blood oxygen level of each ROI. The matrix formed by the time series of all ROIs has a size of N×d, where N is the number of ROIs and d is the length of the time series.
[0047] S2. Perform a correlation operation on the obtained time series using the Pearson correlation coefficient calculation method to obtain the eigenvectors corresponding to each region of interest in the brain.
[0048] The specific formula is:
[0049]
[0050] where ρ X,Y represents the correlation coefficient value, cov(X,Y) represents the covariance of X and Y, σ X represents the standard deviation of X, σ Y represents the standard deviation of Y, E[(X - μ X )(Y - μ Y )] represents the calculation formula of covariance, μ X represents the mean of X, μ Y represents the mean of Y. X and Y represent the time series corresponding to different regions of interest in the brain.
[0051] Then, obtain a matrix Z containing all ROI eigenvectors, with a size of N×f, where f is the length of the eigenvector, generally equal to N (the number of ROIs). One row in the matrix represents the eigenvector of a node. Each eigenvalue in the eigenvector Z[i] of the i-th ROI is the correlation coefficient value between the i-th ROI and other corresponding ROIs. For example, Z[i][j] represents the correlation coefficient value between the i-th ROI and the j-th ROI. This makes the eigenvectors of the generated nodes (ROIs) contain the relationship information with other nodes.
[0052] S3. Use the one-hot encoding method to create a position vector for each brain region of interest. After the position vector is learned by a multi-layer perceptron, the final position vector is obtained.
[0053] In this embodiment, one-hot encoding is used for position encoding, enabling the model to dynamically learn the position vectors of each ROI. In the position vector of the i-th ROI, the i-th bit is 1 and the other bits are 0. Then, the position vector is learned by a multi-layer perceptron (Multilayer Perceptron (MLP), also known as an artificial neural network, which can have multiple hidden layers in addition to the input and output layers) to obtain the final position vector p.
[0054] S4. Treat each brain region of interest as a node, and add the position vector and the feature vector of each brain region of interest correspondingly to obtain the node feature vector.
[0055] In this application, each brain region of interest is treated as a node to generate a graph. The edges between the nodes in the graph represent the relationships between two nodes. In this step, the position vector and the feature vector of each brain region of interest are added correspondingly to obtain the node feature vector, which contains its position information. All node feature vectors form the node feature vector matrix.
[0056] S5. Input the time series into the multi-head attention encoder, and input the obtained encoded features into the graph generator. Treat each brain region of interest as a node to obtain the relational adjacency matrix of the graph formed by each node.
[0057] The function implemented by the multi-head attention encoder in this embodiment can be expressed as:
[0058]
[0059] where Q, K, and V are all obtained by feature mapping of the time series t of each ROI through a single-layer perceptron. H a has a size of N×l, where l is the length of the extracted feature vector. Let H a pass through the Softmax activation function to obtain the encoded feature H e , with the same size. Input H e into the graph generation component to obtain the relational adjacency matrix A of the graph. In matrix A, the connection strength values between each ROI are stored, that is, the weights of the edges.
[0060] The principle of the graph generation component is with a size of N×N. This step enables the model to construct and learn the adjacency matrix of the non-linear relationships between the nodes on the graph, better fitting the functional connections of the human brain.
[0061] S7. Perform a graph convolution operation on the node feature vector matrix composed of all node feature vectors and the relational adjacency matrix to obtain the final node feature vector matrix containing the adjacency node relationships.
[0062] This application takes the node feature vector matrix containing all node feature vectors and the relational adjacency matrix A, and performs S graph convolution operations and normalization operations. The core principle of graph convolution is to perform matrix multiplication between A and so that each node updates its own feature vector through its relationship with adjacent nodes, obtaining the node feature vector matrix after S times of information passing and updating (graph convolution). At this time, the feature vectors of all nodes in the matrix have saturated and absorbed the feature information of adjacent nodes.
[0063] For each graph convolution, the node feature vectors are input into a multi-layer perceptron to further extract node features (the principle of the multi-layer perceptron is equivalent to multiplying the object input to the perceptron by the parameter matrix of the multi-layer perceptron. The parameter matrix of the multi-layer perceptron is learned through training, and finally, after passing through the activation function ReLU), and then a normalization method is used to normalize the extracted node features, and the result is used as the final node feature vector.
[0064] The overall formula is h S = BatchNorm1D(ReLU(Ah S-1 W S ))), where W S is the parameter that the multi-layer perceptron needs to learn during the S-th graph convolution, that is, the node feature vector matrix initially input to the first graph convolution. Among them, h S represents the node feature vector matrix after S times of graph convolution BatchNorm1D represents a one-dimensional batch normalization operation.
[0065] S8. Perform a top-k pooling operation on the final node feature vector matrix to obtain the importance scores of all nodes, select the top k nodes with the highest importance scores, and flatten the feature vectors of the selected nodes to obtain the feature vectors for prediction.
[0066] This application uses the top-k pooling method, such as Figure 2As shown. The main function of top-k pooling is to assign an importance score to all nodes on the graph according to a certain calculation method, retain the top-k nodes with the highest importance scores, and discard the other nodes. The pooling process is also a process of reducing the computational amount. By constructing a learnable pooling weight vector W, which is the parameter that the model needs to learn in the top-k pooling layer and has a size of f×1. The final node feature vector obtained after multiple graph convolution operations is multiplied by the pooling weight vector W in matrix form to obtain the importance scores of all nodes, with a size of N×1. According to the importance scores, the top-k nodes with the highest scores are retained.
[0067] Then, the feature vectors of the retained nodes are flattened to obtain a vector with a size of 1×(N×f).
[0068] S9. Input the feature vector used for prediction into a multi-layer perceptron to obtain the prediction result.
[0069] In this application, the feature vector used for prediction is put into a multi-layer perceptron. After multi-layer feature mapping extraction, a vector with a size of 2×1 is finally obtained, corresponding to the two categories of prediction. In this method, the categories are Autism or Control. If the prediction result is Autism, it means that the subject of the magnetic resonance image input to the model has autism, and Control means that the subject is normal.
[0070] In this application, as Figure 1 shown in the network model, during training, the parameters of each module of the network are updated by calculating the joint loss. The joint loss function is expressed as follows:
[0071] L = L ce + αL intra + βL inter + γL TPK
[0072] Among them, α, β, and γ are the weight coefficients of each loss.
[0073] 1. Cross-entropy loss function L ce :
[0074]
[0075] Among them, y is the label value for the result of autism; M represents the number of examples in a batch during batch training, and z i is the score obtained by the model for the i-th example being autism, and Sigmoid is the activation function that controls the score value between [0, 1].
[0076] 2. Intra-group loss L intra :
[0077]
[0078] where L intra represents the calculated loss value; c represents a label; C is the set of all labels, and S c ={i|Y i,c =1} is a set of examples with label c, where i represents a certain example; μ c is the mean of the weights of the adjacency matrices of all examples in S c ; A i represents the weight value of the adjacency matrix of a certain example; represents c the variance of the weights in the adjacency matrices of all examples in S.
[0079] This loss function regularizes the model so that for examples of the same class, the relational adjacency matrices they learn through the model are as similar as possible.
[0080] 3. Inter-group loss L inter :
[0081]
[0082] where L inter represents the calculated loss value; a and b are two labels in C, and μ a and μ b are the mean of the adjacency matrix of the set of examples with label a and the mean of the adjacency matrix of the set of examples with label b, respectively.
[0083] This loss function regularizes the model so that for examples of different classes, the relational adjacency matrices they learn through the model are as different as possible.
[0084] 4. Top-k loss L TPK :
[0085]
[0086] where L TPK represents the calculated loss value; M represents the number of examples included in each batch during batch training; N represents the number of nodes in each input example; s m,i represents the importance score of the i-th node in the m-th input example.
[0087] Regularize the model so that after the examples pass through the model, for the top k selected important nodes selected by the model, their importance scores (ranging from 0 to 1) tend to be 1, and for the unselected nodes, the importance scores tend to be 0.
[0088] In the technical solution of this application, the relationship adjacency matrix A is compared with the adjacency matrix obtained by using Pearson to calculate the correlation coefficient in the traditional method. It is found that the heat in the precuneus part of the brain is significantly higher than that in other regions. Based on this, it is speculated that the precuneus is related to ASD. The precuneus belongs to the DMN (default mode network), and medical research shows that DMN is closely related to ASD.
[0089] All ASD subjects are predicted by the method of this application. The frequency of the top 10 influence rankings of each ROI in each prediction result is counted, and the ASD-related biomarkers are found according to the frequency. The ranking results are as Figure 3 shown. Finally, it is speculated that the middle temporal gyrus and the related regions of the cerebellum are closely related to ASD, and this result is consistent with the existing medical research conclusions. Figure 3 In it. The first column is the brain region ID, the second column is the brain region label, and the third column is the frequency of occurrence. Among them, Temporal_Pole_Mid_L represents the middle temporal gyrus of the temporal pole (left), Temporal_Pole_Mid_R represents the middle temporal gyrus of the temporal pole (right), Temporal_Mid_R represents the middle temporal gyrus (right), Temporal_Pole_Sup_L represents the superior temporal gyrus of the temporal pole (left), Temporal_Mid_L represents the middle temporal gyrus (left), Temporal_Pole_Sup_R represents the superior temporal gyrus of the temporal pole (right), Cerebelum_Crus1_L represents the cerebellum (left), Cerebelum_Crus1_R represents the cerebellum (right), Temporal_Sup_R represents the superior temporal gyrus (right), and Temporal_Sup_L represents the superior temporal gyrus (left).
[0090] The test effect of the technical solution of this application on the ABIDE dataset is finally shown in Table 1:
[0091] Brain segmentation method Sensitivity % Specificity % Accuracy % AAL 68.6 77.2 72.9 CC200 73.3 78.6 76.4
[0092] Table 1
[0093] Among them, the brain region segmentation adopts AAL (Anatomical Automatic Labeling) and CC200 (Craddock 200) respectively. Compared with the existing technical solutions in this field, the above performance indicators have been greatly improved.
[0094] In this application, through the time series of ROI, a multi-head self-attention model is used for encoding and graph construction to learn the non-linear relationship between ROIs, rather than simply using the method of calculating the Pearson correlation coefficient to obtain a linear relationship. The top-k pooling method is used to improve the interpretability of the model and reduce the computational complexity. A position feature vector is added to distinguish nodes at different positions, introducing the position information of nodes for model prediction. Three other loss functions in addition to the standard binary cross-entropy loss function are used to regularize the training effect of the model.
[0095] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application patent shall be subject to the appended claims.
Claims
1. A method for predicting autism based on nuclear magnetic resonance images and finding biomarkers, characterized in that, The method for predicting autism and finding biomarkers based on magnetic resonance images includes: Obtain the magnetic resonance images of the brain of the detection object, segment the regions of interest in the brain using a preset brain segmentation template, and average all the voxel sequence values contained in each region of interest in the brain to obtain the time series recording the change in blood oxygen level for each region of interest in the brain; Perform correlation operations on the obtained time series using the Pearson correlation coefficient calculation method to obtain the feature vectors corresponding to each region of interest in the brain; Use the one-hot encoding method to create a position vector for each region of interest in the brain, and pass the position vector through the learning of a multi-layer perceptron to obtain the final position vector; Take each region of interest in the brain as a node, and add the position vector and the feature vector of each region of interest in the brain correspondingly to obtain the node feature vector; Input the time series into a multi-head attention encoder, input the obtained encoded features into a graph generator, and take each region of interest in the brain as a node to obtain the relational adjacency matrix of the graph composed of each node; Perform a graph convolution operation on the node feature vector matrix composed of all node feature vectors and the relational adjacency matrix to obtain the final node feature vector matrix containing the relational adjacency node relationship; Perform a top-k pooling operation on the final node feature vector matrix to obtain the importance scores of all nodes, select the top k nodes with the highest importance scores, and flatten the feature vectors of the selected nodes to obtain the feature vectors for prediction; Input the feature vectors for prediction into a multi-layer perceptron to obtain the prediction result.
2. The method for predicting autism based on nuclear magnetic resonance images and finding biomarkers according to claim 1, characterized in that, The formula for the graph convolution operation is as follows: h S = BatchNorm1D(ReLU(Ah S-1 W S )) Among them, BatchNorm1D represents a one-dimensional batch normalization operation, ReLU is an activation function, A is a relational adjacency matrix, and h S represents the node feature vector matrix obtained after S graph convolutions, and h 0 is the initial node feature vector matrix.
3. The method for predicting autism based on nuclear magnetic resonance images and finding biomarkers according to claim 1, characterized in that, The step of performing a top-k pooling operation on the final node feature vector matrix to obtain the importance scores of all nodes includes: Multiply the final node feature vector matrix by the pooling weight vector to obtain the importance scores of all nodes.
4. The method for predicting autism based on nuclear magnetic resonance images and finding biomarkers according to claim 1, characterized in that, The method for predicting autism and finding biomarkers based on magnetic resonance images further includes: Calculate the joint loss for network training through a joint loss function, and the joint loss function is as follows: L = L ce + αL intra + βL inter + γL TPK where α, β, and γ are the weight coefficients of each loss; where y is the label value for autism, M represents the number of samples in each batch during batch training, and z i is the score for the i-th example being autistic, and Sigmoid is the activation function; Among them, c represents a label, C is the set of all labels, and S c = {i | Y i,c = 1} is a set of examples with the label c, and i represents an example; μ c is the mean of the weights of the adjacency matrices of all examples in S c ; A i represents the weight value of the adjacency matrix of example i; represents the variance of the weights in the adjacency matrices of all examples in S c ; where a and b are two labels in C, and μ a and μ b are the mean of the adjacency matrix of the sample set with label a and the mean of the adjacency matrix of the sample set with label b, respectively; Among them, N represents the number of nodes in each input example; s m,i represents the importance score of the i-th node in the m-th input example.