Magnetic resonance imaging classification method based on multi-stage fusion
Through the multi-stage fusion strategy and adaptive graph learning module, the problems of feature inconsistency and information loss in the multimodal fusion process in the prior art are solved, and high accuracy and reliability of magnetic resonance imaging classification are achieved.
Patent Information
- Application Number
- CN202510174497.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-30
AI Technical Summary
The existing magnetic resonance imaging classification methods have problems of inconsistent features and high sample complexity in the multimodal fusion process. Early fusion may lose key information, while late fusion may ignore the complex associations between modes.
The magnetic resonance imaging classification method based on multi-stage fusion is adopted, and the information of different modes is gradually integrated and optimized through multi-stage strategies of early fusion and late fusion, and the complex relationship between modes is automatically captured in combination with the adaptive graph learning module.
The deep integration of information in different modes is achieved, the accuracy and reliability of magnetic resonance imaging classification is improved, information loss and sample complexity problems are reduced, and the model's ability to pay attention to information in different modes is enhanced.
Smart Images

Figure CN120070997A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and specifically relates to a magnetic resonance imaging classification method based on multi-stage fusion. Background Art
[0002] In the technical field of medical image processing, deep learning models can utilize large-scale multi-modal medical data, such as medical records, physiological indicators, gene data, etc., to assist in the classification of magnetic resonance imaging.
[0003] In the existing multi-modal fusion methods for magnetic resonance imaging classification, generally, features are simply concatenated and fused at an early stage and then processed, or each modality is processed independently and then fused at a later stage. However, fusing information from different modalities at an early stage will have problems of inconsistent features and high sample complexity. Similarly, fusing information from different modalities at a later stage may lose some key information.
[0004] In the magnetic resonance image classification task, existing graph-based methods often require manually defining graphs for specific modalities. For example, node connections are constructed based on demographic information, and information from other modalities is integrated through graph representation learning to obtain the representation of the person to be tested. However, constructing an appropriate graph is not an easy task, and these methods ignore the complex associations between modalities and the differential information existing among the persons to be tested. These factors inevitably lead to the model being unable to fully focus on information from different modalities of the person to be tested, thus affecting reliable magnetic resonance image classification. Summary of the Invention
[0005] The present invention is proposed to solve the above-mentioned deficiencies existing in the prior art, and provides a magnetic resonance imaging classification method based on multi-stage fusion, aiming to more effectively integrate medical data information from different modalities, improve the accuracy and reliability of magnetic resonance imaging classification recognition, and thus better assist doctors in making decisions.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A magnetic resonance imaging classification method based on multi-stage fusion according to the present invention is characterized by including the following steps:
[0008] Step 1: After obtaining the brain magnetic resonance imaging pictures of N persons to be tested and preprocessing them, a magnetic resonance imaging modality matrix is obtained , where represents the magnetic resonance imaging modality vector of the i-th person to be tested after preprocessing, and h represents the dimension of the magnetic resonance imaging modality vector;
[0009] After obtaining and preprocessing the M-1 non-imaging modality data of N persons to be tested, M-1 non-imaging modality data matrices are obtained , , where represents the M-1 non-imaging modality vector, and , represents the M-1 non-imaging modality vector of the i-th person to be tested;
[0010] Define the true class labels of N persons to be tested ; represents the true class of the i-th person to be tested. If = 1, it means that the brain magnetic resonance imaging picture of the i-th person to be tested is of the normal class, = 0 means that the brain magnetic resonance imaging picture of the i-th person to be tested is of the abnormal class;
[0011] Step 2, Concatenate with , to obtain the early fusion feature matrix of the i-th person to be tested = , , thereby concatenating with , to obtain the early fusion feature matrix V of N persons to be tested E = ; where represents the m-th modality vector of the i-th person to be tested;
[0012] Input V i E into a multi-layer perceptron for processing, output the early prediction class probability of the i-th person to be tested, and select the maximum probability as the early prediction class Y of the magnetic resonance imaging image of the i-th person to be tested E i ;
[0013] Step 3, Construct a multi-modal late fusion network based on cross-modal attention and process to obtain the late fusion feature matrix V of the i-th person to be tested i L ;
[0014] Step 3.1, Use the attention mechanism to calculate the query vector Q of i m , the key vector K i m , and the value vector V i m;
[0015] Step 3.2. Calculate the cross-modal attention matrix G of i m :
[0016] (1)
[0017] In formula (1), K i m is the key vector of the m-th modality of the i-th person to be tested, m≠n; d is the dimension of K i m ; T represents transpose; represents the activation function; is the key vector of the n-th modality of the i-th person to be tested;
[0018] Step 3.3. Calculate the aggregated feature matrix F of i m :
[0019] (2)
[0020] In formula (2), I is the identity matrix, α n is the hyperparameter that controls the self-preservation strength of the n-th modality, W i n is the linear transformation; is the value vector of the n-th modality of the i-th person to be tested;
[0021] Step 3.4. Calculate the cross-modal attention feature fusion matrix F of the i-th person to be tested i :
[0022] (3)
[0023] Step 3.5. Calculate the query matrix Q i of F i , the key matrix K i and the value matrix V i , and then calculate the attention matrix G of F i using the Softmax function; i ;
[0024] Step 3.6. Calculate the late fusion feature matrix V of the i-th person to be tested i L :
[0025] (4)
[0026] Step 4: Based on the early fusion feature matrix and the late fusion feature matrix, use the relational graph to obtain the relational graph among the persons to be tested;
[0027] Step 4.1: Use Equations (5) and (6) to calculate the early fusion embedding space of the i-th person to be tested and the late fusion embedding space :
[0028] (5)
[0029] (6)
[0030] In Equations (5) and (6), Vec represents the matrix vectorization operation of concatenating row by row;
[0031] Step 4.2: Use Equations (7) and (8) to calculate the similarity A between the i-th person to be tested and the j-th person to be tested in the early fusion embedding space ij and the similarity B in the late fusion embedding space ij :
[0032] (7)
[0033] (8)
[0034] In Equations (7) and (8), represents the early fusion embedding space of the j-th person to be tested, represents the late fusion embedding space of the j-th person to be tested, W A , W B are two weight matrices to be learned, cos(·,·) represents the weighted cosine similarity, and there is:
[0035] (9)
[0036] Step 4.3: Use Equation (10) to calculate the similarity C between the i-th person to be tested and the j-th person to be tested in the joint embedding space ij :
[0037] (10)
[0038] In Equation (10), γ is the weight hyperparameter used to balance the two similarities;
[0039] Step 4.4: Take the fusion feature of the i-th person to be tested as node i. When C ij is less than the threshold th, it means there is no edge connection between node i and node j; otherwise, it means there is an edge connection between node i and node j, thus generating the relational graph C of the persons to be tested;
[0040] Step 5: The prediction module uses the Node2Vec method to map the node i in the relationship graph of the person to be tested into a low-dimensional vector V i C and input V i C into the fully connected layer for processing, output the predicted class probability of the i-th person to be tested, and select the maximum probability as the predicted class of the magnetic resonance imaging image of the i-th person to be tested;
[0041] Step 6: Construct the total loss function L and use it to train the network composed of the multi-modal late fusion network, the relationship graph construction module and the prediction module to update the network parameters until the total loss function L converges, so as to obtain a trained magnetic resonance functional imaging classification and recognition model for classifying the input nuclear magnetic resonance imaging image.
[0042] The characteristics of the magnetic resonance imaging classification method based on multi-stage fusion according to the present invention also lie in that the total loss function L in Step 6 includes:
[0043] Step 6.1: Based on Y E i and construct the cross-entropy loss L of early feature fusion E ;
[0044] Step 6.2: Calculate the similarity loss function L using Equation (11) C :
[0045] (11)
[0046] In Equation (11), represents the m-th modal vector of the j-th person to be tested;
[0047] Step 6.3: Calculate L = L E + L C .
[0048] The characteristics of an electronic device according to the present invention, including a memory and a processor, are that the memory is used to store a program supporting the processor to execute the magnetic resonance functional imaging classification and recognition method, and the processor is configured to execute the program stored in the memory.
[0049] The characteristics of a computer-readable storage medium according to the present invention, on which a computer program is stored, are that the computer program executes the steps of the magnetic resonance functional imaging classification and recognition method when run by a processor.
[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0051] 1. The present invention proposes a multi-modal feature extraction module based on multi-stage fusion. This module effectively solves the respective deficiencies of early fusion and late fusion in the prior art through a multi-stage fusion strategy. Through step-by-step fusion and optimization, it achieves the deep integration of different modal information, providing richer and more accurate feature extraction for the final magnetic resonance image classification decision.
[0052] 2. The present invention proposes a step-by-step graph learning module based on adaptive graph learning. This module automatically generates representations between nodes through adaptive graph learning, and by generating adjacency matrices in two embedding spaces respectively, it introduces the differences between the persons to be measured while reducing information loss, overcoming the problem in traditional magnetic resonance image classification methods that the graph constructed manually is significantly separated from the prediction module, resulting in a cumbersome model tuning process and relatively poor generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a flowchart of the magnetic resonance imaging classification method based on multi-stage fusion of the present invention;
[0054] Figure 2 is a model structure diagram of the magnetic resonance imaging classification method based on multi-stage fusion of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0055] In this embodiment, a magnetic resonance imaging classification method based on multi-stage fusion introduces a multi-modal feature extraction module based on multi-stage fusion. This module effectively solves the respective deficiencies of early fusion and late fusion in the prior art through a multi-stage fusion strategy. Through step-by-step fusion and optimization, it achieves the deep integration of different modal information, providing richer and more accurate feature extraction for the final classification decision; specifically, as Figure 1 shown, the method includes the following steps:
[0056] Step 1: After obtaining the brain magnetic resonance imaging pictures of N persons to be measured and performing preprocessing, a magnetic resonance imaging modal matrix is obtained, where represents the magnetic resonance imaging modal vector of the i-th person to be measured after preprocessing, and h represents the dimension of the magnetic resonance imaging modal vector;
[0057] In the specific implementation, a series of preprocessing is performed on the brain magnetic resonance imaging data, including deboning, brain tissue segmentation, global mean intensity normalization, time slice correction, head motion correction, and interference signal regression, to obtain the brain magnetic resonance imaging data of N persons to be measured.
[0058] After obtaining and preprocessing the M-1 non-imaging modality data of N persons to be tested, M-1 non-imaging modality data matrices are obtained , , where represents the M-1 non-imaging modality vector, and , represents the M-1 non-imaging modality vector of the i-th person to be tested; in a specific embodiment, the value of M is 5, and the model structure of this classification method is as shown in Figure 2 as follows:
[0059] After obtaining and preprocessing the demographic information of N persons to be tested, a demographic information modality matrix is obtained, where represents the demographic information modality vector of the i-th person to be tested after preprocessing;
[0060] After obtaining and preprocessing the physiological index data of N persons to be tested, a physiological index modality matrix is obtained, where represents the physiological index modality vector of the i-th person to be tested after preprocessing;
[0061] After obtaining and preprocessing the past medical history data of N persons to be tested, a past medical history modality matrix is obtained, where represents the past medical history modality vector of the i-th person to be tested after preprocessing;
[0062] After obtaining and preprocessing the living habit data of N persons to be tested, a living habit modality matrix is obtained, where represents the living habit modality vector of the i-th person to be tested after preprocessing.
[0063] Define the true class labels of N persons to be tested ; represents the true class of the i-th person to be tested. If = 1, it means that the brain magnetic resonance imaging picture of the i-th person to be tested is of the normal class, = 0 means that the brain magnetic resonance imaging picture of the i-th person to be tested is of the abnormal class;
[0064] In a specific embodiment, professional clinicians identify and label the resting-state functional magnetic resonance imaging data to obtain a label set.
[0065] Step 2, Concatenate with , to obtain the early fusion feature matrix = , , so as to combine with , and obtain the early fusion feature matrix V of N to-be-tested persons after splicing E = ; where represents the m-th modal vector of the i-th to-be-tested person;
[0066] Input V i E into a multi-layer perceptron for processing, output the early prediction class probability of the i-th to-be-tested person, and select the maximum probability as the early prediction class Y of the magnetic resonance imaging image of the i-th to-be-tested person E i .
[0067] Step 3, as Figure 2 shown, construct a cross-modal attention-based multi-modal late fusion network and process to obtain the late fusion feature matrix V of the i-th to-be-tested person i L ;
[0068] Step 3.1, use the attention mechanism to calculate the query vector Q of i m , key vector K i m , value vector V i m ;
[0069] Step 3.2, use Equation (1) to calculate the cross-modal attention matrix G of i m :
[0070] (1)
[0071] In Equation (1), K i m is the key vector of the m-th modality of the i-th to-be-tested person, m≠n; d is the dimension of K i m ; T represents transpose; represents the activation function, and through the calculation of the Softmax function, G i n is normalized; is the key vector of the n-th modality of the i-th to-be-tested person.
[0072] Step 3.3, use Equation (2) to calculate Aggregation feature matrix F i m :
[0073] (2)
[0074] In formula (2), I is the identity matrix, and α n is the hyperparameter that controls the self-preservation intensity of the nth modality. In the specific implementation, α n is 0.5, and W i n is a linear transformation; is the value vector of the nth modality of the ith person to be tested.
[0075] Step 3.4. Calculate the cross-modal attention feature fusion matrix F of the ith person to be tested using formula (3) i :
[0076] (3)
[0077] Step 3.5. Calculate the query matrix Q i of F i , key matrix K i and value matrix V i of F, and then calculate the attention matrix G i of F using the Softmax function i .
[0078] Step 3.6. Calculate the late fusion feature matrix V of the ith person to be tested using formula (4) i L :
[0079] (4).
[0080] Step 4. Based on the early fusion feature matrix and the late fusion feature matrix, use the relationship graph construction module to obtain the relationship graph among the persons to be tested;
[0081] Step 4.1. Calculate the early fusion embedding space and the late fusion embedding space of the ith person to be tested using formulas (5) and (6):
[0082] (5)
[0083] (6)
[0084] In formulas (5) and (6), Vec represents the matrix vectorization operation of concatenating rows.
[0085] Step 4.2: Calculate the similarity A between the i-th person to be tested and the j-th person to be tested in the early fusion embedding space using Equations (7) and (8). ij And the similarity B in the late fusion embedding space ij :
[0086] (7)
[0087] (8)
[0088] In Equations (7) and (8), represents the early fusion embedding space of the j-th person to be tested, represents the late fusion embedding space of the j-th person to be tested, and W A , W B are two weight matrices to be learned. cos(·,·) represents the weighted cosine similarity. Here, methods such as Pearson correlation and Euclidean distance can be used to replace the cosine similarity to measure the similarity between the persons to be tested. It is found in the specific implementation that using the weighted cosine similarity has better measurement effects than Pearson correlation and Euclidean distance. And there is:
[0089] (9)
[0090] Step 4.3: Calculate the similarity C between the i-th person to be tested and the j-th person to be tested in the joint embedding space using Equation (10). ij :
[0091] (10)
[0092] In Equation (10), γ is a weight hyperparameter used to balance the two similarities. In the specific implementation, γ is 0.3.
[0093] Step 4.4: Take the fusion feature of the i-th person to be tested as node i, and set a threshold th. In the specific implementation, th is 0.8. When C ij is less than th, it means there is no edge connection between node i and node j; otherwise, it means there is an edge connection, thereby generating the relationship graph C of the persons to be tested.
[0094] Step 5: The prediction module uses the Node2Vec method to map node i in the relationship graph of the persons to be tested to a low-dimensional vector V i C , and input V i C into the fully connected layer for processing, output the predicted class probability of the i-th person to be tested, and select the maximum probability as the predicted class of the magnetic resonance imaging image of the i-th person to be tested;
[0095] In the specific implementation, the execution process of Node2Vec is as follows: 1. Initialize the Node2Vec model: First, set the parameters of Node2Vec, including the dimension of the embedding vector as 128, the length of the random walk as 10, the number of walks for each node as 80, and the window size as 5; 2. Generate node sequences through random walks to train Node2Vec, thereby obtaining the low-dimensional vector representation V of each node i C .
[0096] Step 6, construct the total loss function L and use it to train the network composed of the multi-modal late fusion network, the relationship graph construction module, and the prediction module to update the network parameters until the total loss function L converges, thereby obtaining a trained magnetic resonance functional imaging classification and recognition model for classifying the input nuclear magnetic resonance imaging images.
[0097] Step 6.1, based on Y E i and construct the cross-entropy loss L of early feature fusion E ;
[0098] Step 6.2, use Equation (11) to calculate the similarity loss function L C :
[0099] (11)
[0100] In Equation (11), represents the m-th modal vector of the j-th person to be tested;
[0101] Step 6.3, calculate L = L E + L C .
[0102] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above magnetic resonance functional imaging classification and recognition method, and the processor is configured to execute the program stored in the memory.
[0103] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium. When the computer program is run by the processor, it executes the steps of the above magnetic resonance functional imaging classification and recognition method.
[0104] In summary, the present invention improves the accuracy of magnetic resonance functional imaging classification, reduces the problems of information loss and sample complexity caused by improper feature fusion. And through adaptive graph learning, it automatically captures and expresses the complex associations between modalities, enhancing the model's attention ability to different modality information.
Claims
1. A magnetic resonance imaging classification method based on multi-stage fusion, characterized in that: The following steps are involved: Step 1: Obtain brain MRI images of N test subjects and preprocess them to obtain the MRI modality matrix. ,in, represents the MRI modality vector of the i-th person to be tested after preprocessing, h represents the dimension of the MRI modality vector; After obtaining M-1 non-imaging modality data of N test subjects and preprocessing them, we get M-1 non-imaging modality data matrices [ , ],in, represents the M-1th non-imaging modality vector, and , represents the M-1th non-imaging modality vector of the i-th person to be tested; Define the true category labels of N people to be tested ; represents the true category of the i-th person to be tested, if =1, indicating that the brain MRI image of the i-th person under test is of the normal category, =0 means that the brain MRI image of the i-th person under test is of the abnormal category; Step 2: and[ , ] After splicing, the early fusion feature matrix of the i-th person to be tested is obtained =[ , ], thus and[ , ] After splicing, we get the early fusion feature matrix V of N people to be tested. E = ;in, represents the mth modal vector of the i-th person to be tested; V i E The input is processed into a multi-layer perceptron, and the early prediction category probability of the i-th person to be tested is output, and the maximum probability is selected as the early prediction category Y of the magnetic resonance imaging image of the i-th person to be tested E i ; Step 3: Construct a multimodal late fusion network based on cross-modal attention and Processing is performed to obtain the late fusion feature matrix V of the i-th person to be tested i L ; Step 3.1: Calculate using attention mechanism The query vector Q i m , key vector K i m , value vector V i m ; Step 3.2: Calculate using formula (1) The cross-modal attention matrix G i m : (1) In formula (1), K i m is the key vector of the mth mode of the i-th person to be tested, m≠n; d is K i m The dimension of , T represents transpose; represents the activation function; is the key vector of the nth mode of the i-th person to be tested; Step 3.3: Calculate using formula (2) The aggregate feature matrix F i m : (2) In formula (2), I is the unit matrix, α n W is a hyperparameter that controls the self-preservation strength of the nth mode. i n is a linear transformation; is the value vector of the nth mode of the i-th person to be tested; Step 3.4: Use formula (3) to calculate the cross-modal attention feature fusion matrix F of the i-th person to be tested: i : (3) Step 3.5: Calculate F using the attention mechanism i The query matrix Q i , key matrix K i and the value matrix V i , so as to use the Softmax function to calculate F i The attention matrix G i ; Step 3.6: Use formula (4) to calculate the late fusion feature matrix V of the i-th person to be tested: i L : (4) Step 4: Based on the early fusion feature matrix and the late fusion feature matrix, a relationship diagram between the persons to be tested is obtained using the relationship diagram; Step 4.1: Use equations (5) and (6) to calculate the early fusion embedding space of the i-th person to be tested: and late fusion embedding space : (5) (6) In equations (5) and (6), Vec represents the matrix vectorization operation of row-by-row concatenation; Step 4.2: Use equations (7) and (8) to calculate the similarity A between the i-th person to be tested and the j-th person to be tested in the early fusion embedding space: ij and the similarity B in the late fusion embedding space ij : (7) (8) In formula (7) and formula (8), represents the early fusion embedding space of the jth person to be tested, represents the late fusion embedding space of the jth person to be tested, W A , W B are two weight matrices to be learned, cos(·,·) represents the weighted cosine similarity, and: (9) Step 4.3: Use formula (10) to calculate the similarity C between the i-th person to be tested and the j-th person to be tested in the joint embedding space ij : (10) In formula (10), γ is a weight hyperparameter used to balance the two similarities; Step 4.4: Take the fusion feature of the i-th person to be tested as node i. ij When it is less than the threshold th, it means that there is no edge connection between node i and node j, otherwise, it means that there is an edge connection between node i and node j, thereby generating a relationship graph C of the persons to be tested; Step 5: The prediction module uses the Node2Vec method to map the node i in the relationship graph of the person to be tested into a low-dimensional vector V i C , and V i C The input is processed in the fully connected layer, and the predicted category probability of the i-th person to be tested is output, and the maximum probability is selected as the predicted category of the magnetic resonance imaging image of the i-th person to be tested; Step 6: construct a total loss function L and use it to train the network composed of the multimodal late fusion network, the relationship graph construction module and the prediction module to update the network parameters until the total loss function L converges, thereby obtaining a trained magnetic resonance functional imaging classification and recognition model for classifying the input magnetic resonance imaging images.
2. The magnetic resonance imaging classification method based on multi-stage fusion according to claim 1, characterized in that: The total loss function L in step 6 includes: Step 6.1: Based on Y E i and Constructing the cross entropy loss for early feature fusion E ; Step 6.2: Calculate the similarity loss function L using formula (11) C : (11) In formula (11), represents the mth modal vector of the jth person to be tested; Step 6.3, calculate L=L E +L C .
3. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the magnetic resonance functional imaging classification and recognition method according to claim 1 or 2, and the processor is configured to execute the program stored in the memory.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the magnetic resonance functional imaging classification and recognition method according to claim 1 or 2 are executed.
Citation Information
Cited By
Patient classification method and device based on magnetic resonance parameters, equipment and medium
CN120804919A