A globally optimized supervised multi-modal image fusion method
By employing a globally optimized supervised multimodal image fusion method, the independence and correlation of independent components are maximized. By combining gradient descent algorithm and working memory scoring, the method addresses the shortcomings in accuracy and robustness of existing multimodal image fusion techniques, achieving more accurate multimodal image analysis and supporting personalized treatment plans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2024-04-03
- Publication Date
- 2026-04-28
AI Technical Summary
Existing multimodal image fusion methods cannot fully utilize the complementary information between modalities, and existing supervised fusion methods are insufficient in terms of accuracy and robustness. In particular, they are highly dependent on accuracy when utilizing genetic information, and phased optimization may lead to the loss of correlation.
A globally optimized supervised multimodal image fusion method is adopted. By maximizing the independence of each modal independent component and its correlation with reference information, gradient descent algorithm is used for iterative optimization. Combining the features of functional magnetic resonance imaging and structural magnetic resonance imaging, working memory score is introduced as reference information to identify multimodal covariant components.
It improves the accuracy and robustness of multimodal image fusion, enabling targeted identification of covariant components related to clinical indicators, deepening the understanding of the pathological mechanisms of mental illnesses, and providing support for personalized treatment plans.
Smart Images

Figure CN118351406B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to medical image analysis, specifically to a globally optimized supervised multimodal image fusion method. Background Technology
[0002] In recent years, the rapid development of neuroimaging technology has enabled researchers to acquire neuroimaging data of the same subject in different modalities using various imaging methods, thereby providing multi-dimensional information about brain anatomy and function. For example, common structural imaging techniques include structural magnetic resonance imaging (sMRI), which can be used to assess changes in the local concentration or volume of gray and white matter on each voxel, reflecting corresponding anatomical changes. In the field of functional imaging, functional magnetic resonance imaging (fMRI) is based on blood oxygen level-dependent signals, reflecting the activity of neurons in the brain under task or resting conditions. All of these imaging methods have their own advantages but also limitations. Using only one imaging modality cannot simultaneously obtain complete information about brain function and structure. Although multimodal imaging data of the scanned subject coexist, in practice they are often analyzed separately, or only a simple comparison or correlation of the analysis results of two magnetic resonance imaging data is performed, and the complementary information between modalities cannot be fully utilized. Multimodal data fusion analysis methods can integrate complementary information between different modalities to gain a deeper understanding of the pathological mechanisms associated with neuropsychiatric diseases, thereby improving the accuracy of diagnosis and prediction.
[0003] Supervised fusion methods are more goal-oriented because they utilize reference information to guide fusion analysis, enabling precise localization of specific components of interest from large, complex datasets and effectively improving the targeting of feature extraction. pICA-R (pICA with references) is a method that uses genes as reference information to guide multimodal fusion. However, the reference information in this method is limited to spatial component distribution, such as specific brain regions or sites, and cannot cover individual subject indicators, such as various behavioral scores of the subject samples. Furthermore, the fusion effect heavily depends on the accuracy of the genetic information. MCCAR+jICA (multi-set CCA with reference + joint ICA), while maintaining the basic performance of joint source separation, can further detect covariant multimodal features that are significantly correlated with the reference information. However, this method optimizes the objective function in stages, which may lead to the loss of correlations extracted in the first stage in the second stage analysis. Summary of the Invention
[0004] Purpose of the invention: To address the above-mentioned shortcomings, this invention provides a globally optimized supervised multimodal image fusion method with improved accuracy.
[0005] Technical Solution: To solve the above problems, this invention employs a globally optimized supervised multimodal image fusion method, comprising the following steps:
[0006] (1) Preprocess brain imaging data of each modality and extract features to obtain features of each modality;
[0007] (2) Preprocess the obtained modal features to obtain the preprocessed feature matrix of each modality;
[0008] (3) The independence of each modality independent component is calculated based on the feature matrix, the mixture matrix and the weight matrix of each modality;
[0009] (4) Find the optimal solution of the mixing matrix when maximizing the independence of each modal independent component, the correlation between the mixing matrices of each modality, and the correlation between the mixing matrix of each modality and the reference information;
[0010] (5) Based on the mixing matrix obtained in step (4), calculate the covariance component and mixing matrix of the multimodal related to the reference information.
[0011] Furthermore, the independence of each modal independent component in step (3) is described by the information entropy function H(·), and the formula for maximizing the independence of each modal independent component is:
[0012]
[0013]
[0014] U k =W k ·X k +W k0 ;
[0015]
[0016] Where E(.) is the expected value; Y is the probability density function; k U k W is an intermediate variable in the calculation. k Let A be the mixture matrix. k The unmixing matrix; A k W is the mixing matrix of mode k; k0 X is the bias weight matrix for mode k; k This is the feature matrix of mode k after preprocessing.
[0017] Furthermore, the formulas for solving step (4) to maximize the independence of each modal independent component, the correlation between the mixing matrices of each modality, and the correlation between the mixing matrices of each modality and the reference information are as follows:
[0018]
[0019] Where, m k (k = 1, 2) represents the number of independent components of mode k; A 1i A is the i-th column of the mixing matrix of mode 1; 2j The j-th column of the mixing matrix of mode 2; ref h Represents the h-th column of the reference information matrix; Corr(A 1i A 2j ) represents A 1i With A 2j Correlation; Corr(A) 1i ,ref k ) represents A 1i With ref h Correlation; Corr(A) 2j ,ref h ) represents A2j and ref h The correlation; α1, α2, α3 represent regularization parameters.
[0020] Furthermore, the gradient descent algorithm is used to iteratively solve the formula until convergence, yielding the optimal solution for the mixture matrix:
[0021] The iteration rules for the unmixing matrices of each mode satisfy:
[0022]
[0023] Where, λ k The learning rate for independent component analysis of mode k; I is the identity matrix; T is the transpose of the matrix.
[0024] Furthermore, step (4) specifically includes the steps of maximizing the correlation between the mixing matrices of each mode and the correlation between the mixing matrices of each mode and the reference information:
[0025] (41) Calculate the correlation between the mixing matrix of each mode and the reference information, and select the triplet A with the highest correlation. 1i A 2j and ref k As a constraint;
[0026] (42) Relating the sum of squares of correlation to A 1i A 2j Take the partial derivatives separately and set them equal to 0 to obtain the product containing A.1i A 2j Partial differential equations;
[0027] (43) Set an initial point, and place the A 1i A 2j Perform iterative updates, ensuring the value of the correlation sum of squares increases until the correlation sum of squares converges, and then set A at this point. 1i A 2j This is the optimal solution for this iteration process;
[0028] (44) Update A1 and A2 according to the optimal solution described in step (43), and repeat steps (41), (42) and (43) to obtain the optimal solution A. 1i A 2j .
[0029] Furthermore, in step (43), A 1i A 2j The iterative update rules for performing iterative updates satisfy:
[0030]
[0031]
[0032] Where Std(·) is the standard deviation; Cov(·) is the covariance; and var(·) is the variance. They represent A respectively 1i A 2j ref k The average value of η; 11 η 12 η 21 η 22 The descent step size is obtained using the gradient descent algorithm; λ c1 and λ c2 This is the learning rate.
[0033] Furthermore, in step (1), the modal images include functional magnetic resonance imaging (fMRI) and structural magnetic resonance imaging (sMRI). The modal features extracted from the fMRI include fractional low-frequency fluctuation amplitude; the modal features extracted from the sMRI include gray matter volume. In step (2), the obtained modal features are preprocessed, including matrix transformation, centering, whitening, and dimensionality reduction using principal component analysis.
[0034] Beneficial Effects: Compared to existing technologies, the significant advantage of this invention is that it achieves a globally optimal solution by maximizing the independence of independent components and simultaneously maximizing the sum of squared correlations between independent components across modalities and between independent components and reference information. By introducing working memory scores as reference information, it specifically identifies covariant components that are spatially independent and correlated with clinical indicators. This allows for a deeper exploration of the interaction between cognitive abilities and multimodal neuroimaging in mental illnesses, providing strong support for understanding the pathological mechanisms of mental illnesses and developing personalized treatment plans. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating the multimodal brain image fusion method of the present invention;
[0036] Figure 2 (a) is a schematic diagram comparing the absolute values of the correlation between each modal target component and the reference information estimated by MCCA, jICA, MCCA+jICA, separate ICA, PICA, MCCAR, MCCAR+jICA and the fusion method of the present invention with the true absolute values of the correlation between each modal target component and the reference information.
[0037] Figure 2 (b) A schematic diagram comparing the accuracy of the mixing matrix and the accuracy of the independent components of MCCA, jICA, MCCA+jICA, separate ICA, PICA, MCCAR, MCCAR+jICA and the fusion method of the present invention relative to the absolute value of the true correlation between the target components of each modality and the reference information.
[0038] Figure 3 (a) is a schematic diagram of the accuracy of MCCA, jICA, MCCA+jICA, separate ICA, PICA, MCCAR, MCCAR+jICA and the supervised PICA method of the present invention relative to the independent components averaged with respect to 14 noise levels;
[0039] Figure 3 (b) is a schematic diagram of the accuracy of the MCCA, jICA, MCCA+jICA, separate ICA, PICA, MCCAR, MCCAR+jICA and supervised PICA methods relative to the averaged mixture matrix of 14 noise levels;
[0040] Figure 4 (a) is a spatial view of the multimodal joint covariant components related to working memory scores detected by the fusion method of the present invention;
[0041] Figure 4(b) A graph showing the differences between the schizophrenia patient group and the healthy control group for each modality;
[0042] Figure 4 (c) Correlation fitting plot of each modality and working memory score;
[0043] Figure 4 (d) is the correlation fitting diagram between each mode. Detailed Implementation
[0044] like Figure 1 As shown, a globally optimized supervised multimodal image fusion method in this embodiment includes the following steps:
[0045] S1: Preprocess the brain imaging data of each modality and extract its features.
[0046] The various modalities can be brain imaging techniques such as resting-state functional MRI (fMRI), structural magnetic resonance imaging (sMRI), and diffusion tensor imaging (dMRI). In this embodiment, two modalities, fMRI and sMRI, are used, but in practical applications, multiple modalities can be selected as needed.
[0047] Although this embodiment uses multimodal data fusion of schizophrenia as an example for illustration, the scope of protection of this invention is not limited thereto. The data of each modality can cover multimodal data of all mental disorders, including but not limited to schizophrenia, depression and other mental illnesses, as well as multimodal data of non-mental illnesses such as bipolar disorder.
[0048] In this embodiment, the fractional amplitude of low-frequency fluctuations (fALFF) of resting-state functional brain images (fMRI) is extracted as its feature. The preprocessing and feature extraction of fMRI are as follows: In the MATLAB 2019 environment, based on the standard preprocessing of statistical parametric mapping, the data is processed as follows: (1) Head motion correction, the image at each time point is adjusted to ensure that the brain structure is accurately corresponded during the analysis; (2) Inter-slice time correction, the acquisition time difference between different layers (slices) in the fMRI image is adjusted; (3) Normalization to the Montreal MNI standard space and resampling; preferably, resampling to 3×3×3mm; (4) Regression of 6 head motion parameters and white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF); (5) Using 8 mm full width at half maximum (FWHM) ... (6) Spatial smoothing is performed using a Gaussian kernel (max, FWHM); the sum of amplitude values in the low-frequency power range of 0.01 to 0.08 Hz is divided by the sum of amplitude values in the entire detectable power spectrum to calculate the fractional amplitude of low-frequency fluctuations (fALFF).
[0049] In this embodiment, gray matter volume (GMV) is extracted from structural magnetic resonance imaging (sMRI) as its feature. The preprocessing and feature extraction of sMRI are as follows: (1) Using the unified segmentation method in SPM12, the sMRI data is standardized to the Montreal MNI standard space and resampled; preferably, it is resampled to 3×3×3mm; (2) it is segmented into gray matter, white matter and cerebrospinal fluid, and the result is output as gray matter volume; (3) the gray matter is smoothed using an 8 mm half maximum full width Gaussian filter kernel; (4) outliers of the subjects are detected to ensure correct segmentation.
[0050] After performing the above preprocessing on the fMRI and sMRI data of each subject, the features of each modality were converted into 3D matrices. For fMRI, available features included fractional low-frequency fluctuation amplitude features. For sMRI, available modal features included gray matter volume features.
[0051] S2: Preprocess the features of each modality, including matrixing, centering, whitening, and dimensionality reduction using principal component analysis.
[0052] The matrix transformation process can be as follows: The 3D feature matrices of each modality obtained after step S1 are stretched into row vectors (1×L), where L is the number of voxels. This transformation is performed on all the 3D matrices of the subjects to obtain an N×L matrix X. k (k = 1, 2), where N represents the number of subjects and k represents the number of modalities. Performing the above transformation on both fMRI and sMRI modalities yields the feature matrices for both modalities.
[0053] Next, the aforementioned feature matrix is centered and whitened to eliminate the effects of data translation, ensuring that it has the same variance and is uncorrelated with each other, thus guaranteeing that the feature values of the two modes are within the same range.
[0054] Finally, principal component analysis (PCA) is used to reduce the dimensionality of the feature matrix, removing redundant information and irrelevant features, thereby reducing computational complexity and noise interference. This helps preserve the main direction of data variation and improves the model's generalization performance.
[0055] S3: Based on the dimensionality-reduced modal features obtained in step S2, add the reference information matrix, and maximize the independence of each modal component, the correlation between modes, and the correlation between components and reference information. Use the gradient descent algorithm to perform iterative loops until convergence.
[0056] Specifically, maximization is achieved according to the following formula:
[0057]
[0058] In formula (1), H(·) is the information entropy function that guarantees the independence of each modal independent component; m k (k = 1, 2) represents the number of mutually independent components in mode k; A 1i A represents the i-th column of the mixing matrix for mode 1; 2j Represents the j-th column of the mixing matrix of mode 2; ref k Represents the k-th column of the reference information matrix; Corr(A 1i A 2j ) represents A 1i With A 2j Correlation; Corr(A) 1i ,ref k ) represents A 1i With ref k Correlation; Corr(A) 2j ,ref k ) represents A 2j With ref k The correlation; α1, α2, α3 represent the control and The regularization parameter between the weights ensures the convergence of the algorithm optimization process.
[0059] Specifically, parallel independent component analysis is performed on each modal feature according to the following formula:
[0060]
[0061]
[0062]
[0063] in, and Let Y1 and Y2 represent the probability density functions respectively; E(.) is the expected value; H(.) is the entropy function; W k (k = 1, 2) is the mixture matrix A k The inverse of is A k The unmixing matrix; W k0 (k = 1, 2) is the bias weight matrix for mode k; X k (k = 1, 2) represents the feature matrix of mode k after preprocessing.
[0064] For solving the ICA problem, the widely used Infomax algorithm is employed here.
[0065] Solving the maximization problem of formula (1) may specifically include: calculating the correlation of columns A1, A2, and ref, and selecting the triplet A with the highest correlation. 1iA 2j and ref k As a constraint, the sum of squares of correlations is relative to A. 1i A 2j Take the partial derivatives respectively and set them equal to 0 to obtain the product containing the aforementioned A. 1i A 2j The partial differential equation. Since the sum of squared correlations in the objective function is A... 1i A 2j The quadratic equation, therefore the sum of squares of the correlations with respect to A 1i A 2j The partial differential is A 1i A 2j Since it is a linear function, an approximate solution can be obtained. Let an initial point be set, and let A... 1i A 2j The iteration is performed using the formula (1), ensuring that the value of the sum of squares of correlation increases. When the convergence criterion of the objective function is met, the iteration stops, and A is then set to... 1i A 2j This is the optimal solution for this iteration process. A1 and A2 are updated based on the optimal solution obtained in each iteration process. The above steps are repeated until a stable solution is obtained, which is the optimal solution of formula (1).
[0066] The reference information is used to guide multimodal fusion and includes, but is not limited to, working memory scores, cognitive scores, symptom scores, and genes at a specific locus. The reference information can be selected based on the actual research objectives.
[0067] S4: Step S3 yields multimodal covariance components and a mixture matrix that are significantly related to the reference information, thereby achieving globally optimized supervised multimodal brain image fusion.
[0068] Figure 2 (a) The absolute values of the correlation between each modal target component and the reference information estimated by each method were compared with the true absolute values of the correlation between each modal target component and the reference information.
[0069] In this embodiment, MCCA refers to multi-group canonical correlation analysis; jICA refers to joint independent component analysis; separate ICA refers to independent component analysis of individual modes; PICA refers to parallel independent component analysis; MCCAR refers to supervised multivariate canonical correlation analysis; MCCAR+jICA refers to supervised multivariate canonical correlation analysis + joint independent component analysis; and supervised PICA refers to supervised parallel independent component analysis. Additionally, true correlation refers to the actual correlation.
[0070] pass Figure 2(a) As can be seen, the method proposed in this embodiment of the invention can still extract independent components related to the reference information even when the true correlation is low. As the true correlation increases, the method proposed in this embodiment of the invention exhibits strong robustness in extracting independent components.
[0071] Figure 2 (b) An illustrative example is shown, which is a mixture matrix A averaging the absolute values of the correlations between each method and the reference information for each modal target component. k and independent component S k The accuracy.
[0072] Accuracy refers to the correlation between the mixture matrix and independent components obtained by each method and the true mixture matrix and components. The higher the correlation between the two, the higher the accuracy.
[0073] pass Figure 2 (b) It can be seen that the method proposed in the embodiments of the present invention is effective in the hybrid matrix A. k and independent component S k In terms of accuracy, it is optimal compared to other supervised multimodal fusion methods, which illustrates the importance and necessity of using global optimization solution methods.
[0074] Figure 3 (a) Exemplarily shown is the independent component S, averaged relative to 14 noise levels (peak signal-to-noise ratio PSNR = [-1110]), in each method. k The accuracy. Figure 3 (b) An illustrative example is shown of the mixing matrix A, which is averaged over 14 noise levels (peak signal-to-noise ratio PSNR = [-11 10]). k The accuracy. Through Figure 3 (a) and Figure 3 (b) It can be seen that the method proposed in this embodiment of the invention has a relatively stable ability to extract independent components under different noise influences, and in the mixing matrix A k and independent component S k In terms of accuracy, it is optimal.
[0075] Table 1 shows the comparison results of the ability of each method to extract independent components significantly correlated with the reference when the peak signal-to-noise ratio (PSNR) is 7. Here, real corr represents the true correlation; A~A * The accuracy of the mixture matrix extracted by the method proposed in the example of this invention; S~S * The accuracy of the independent components extracted by the method proposed in the example of this invention.
[0076] Table 1 compares the ability of each method to extract independent components significantly related to the reference information.
[0077]
[0078] The following is a detailed description of the tests performed on the methods proposed in the examples of this invention.
[0079] In this example, the simulation data was generated by using the simTB toolbox to simulate two datasets with the same dimensions as fMRI and sMRI data. Each dataset contained six brain networks with different distributions and two random mixing matrices.
[0080] The real-world data used in this example came from the Function Biomedical Informatics Research Network (FBIRN) Phase III study dataset, which included 294 participants (147 patients with schizophrenia and 147 demographically matched healthy controls). All participants' working memory scores were assessed using the CMINDS cognitive assessment system. This data was used to test the method proposed in this embodiment of the invention. Working memory scores were used as reference information.
[0081] Figure 4 (a) Figure 4 (b) Figure 4 (c) and Figure 4 (d) An exemplary result of an embodiment of the present invention is shown on the FBIRN dataset, where IC is an independent component. Figure 4 (a) is a spatial view of the multimodal covariance components related to working memory scores detected by the method proposed in this embodiment of the invention. It can be seen that the location of the brain regions and the degree of activation contained therein are basically consistent with existing studies. Figure 4 (b) is a plot of differences between groups for each modality; Figure 4 (c) is a correlation fitting plot between each modality and the Working Memory score; Figure 4 (d) is a fitted plot of the correlation between modalities. In summary, it can be seen that the method proposed in this invention can separate independent components that have inter-group differences, inter-modal correlations, and are significantly correlated with working memory scores.
Claims
1. A globally optimized supervised multimodal image fusion method, characterized in that, Includes the following steps: (1) Preprocess brain imaging data of each modality and extract features to obtain features of each modality; (2) Preprocess the obtained modal features to obtain the preprocessed feature matrix of each modality; (3) The independence of each modality independent component is calculated based on the feature matrix, the mixture matrix and the weight matrix of each modality; (4) Find the optimal solution of the mixing matrix when maximizing the independence of each modal independent component, the correlation between the mixing matrices of each modality, and the correlation between the mixing matrix of each modality and the reference information; (5) Based on the mixing matrix obtained in step (4), calculate the covariance component and mixing matrix of the multimodal related to the reference information; The independence of each modal independent component in step (3) is described by the information entropy function H(·), and the formula for maximizing the independence of each modal independent component is: U k =W k ·X k +W k0 ; Where E(.) is the expected value; Y is the probability density function; k U k W is an intermediate variable in the calculation. k Let A be the mixture matrix. k The unmixing matrix; A k W is the mixing matrix of mode k; k0 X is the bias weight matrix for mode k; k This is the feature matrix of mode k after preprocessing; The formulas for solving step (4) to maximize the independence of each modal independent component, the correlation between the mixing matrices of each modality, and the correlation between the mixing matrices of each modality and the reference information are as follows: Where, m k (k = 1, 2) represents the number of independent components of mode k; A 1i A is the i-th column of the mixing matrix of mode 1; 2j The j-th column of the mixing matrix of mode 2; ref h Represents the h-th column of the reference information matrix; Corr(A 1i A 2j ) represents A 1i With A 2j Correlation; Corr(A) 1i ,ref k ) represents A 1i With ref h Correlation; Corr(A) 2j ,ref h ) represents A 2j With ref h The correlation; α1, α2, α3 represent regularization parameters.
2. The supervised multimodal image fusion method according to claim 1, characterized in that, The gradient descent algorithm is used to iteratively solve the formula until convergence, yielding the optimal solution for the mixture matrix. The iteration rules for the unmixing matrices of each mode satisfy: Where, λ k The learning rate for independent component analysis of mode k; I is the identity matrix; T is the transpose of the matrix.
3. The supervised multimodal image fusion method according to claim 2, characterized in that, The steps in step (4) that maximize the correlation between the mixing matrices of each mode and the correlation between the mixing matrices of each mode and the reference information specifically include: (41) Calculate the correlation between the mixing matrix of each mode and the reference information, and select the triplet A with the highest correlation. 1i A 2j and ref k As a constraint; (42) Relating the sum of squares of correlation to A 1i A 2j Take the partial derivatives separately and set them equal to 0 to obtain the product containing A. 1i A 2j Partial differential equations; (43) Set an initial point, and place the A 1i A 2j Perform iterative updates, ensuring the value of the correlation sum of squares increases until the correlation sum of squares converges, and then set A at this point. 1i A 2j This is the optimal solution for this iteration process; (44) Update A1 and A2 according to the optimal solution described in step (43), and repeat steps (41), (42) and (43) to obtain the optimal solution A. 1i A 2j .
4. The supervised multimodal image fusion method according to claim 3, characterized in that, In step (43), A 1i A 2j The iterative update rules for performing iterative updates satisfy: Where Std(·) is the standard deviation; Cov(·) is the covariance; and var(·) is the variance. They represent A respectively 1i A 2j ref k The average value of η; 11 η 12 η 21 η 22 The descent step size is obtained using the gradient descent algorithm; λ c1 and λ c2 This is the learning rate.
5. The supervised multimodal image fusion method according to claim 1, characterized in that, The modal images in step (1) include functional magnetic resonance imaging (fMRI) and structural magnetic resonance imaging (sMRI).
6. The supervised multimodal image fusion method according to claim 5, characterized in that, The modal features extracted from the functional magnetic resonance imaging (fMRI) include fractional low-frequency fluctuation amplitude; the modal features extracted from the structural magnetic resonance imaging (sMRI) include gray matter volume.
7. The supervised multimodal image fusion method according to claim 1, characterized in that, The preprocessing of the modal features obtained in step (2) includes matrixing, centering, whitening, and dimensionality reduction using principal component analysis.
Citation Information
Patent Citations
Supervised multimodal brain image fusion method
CN105957047A
Multi-modal image fusion method of regression interference covariable
CN116402731A