Brain Functional Connectivity Clustering With Multi-Site MRI Harmonization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in using functional brain imaging for neurological/mental disorders is the difficulty in generating precise biomarkers due to small sample sizes and site-to-site differences in brain activity measurements, leading to overfitting and poor generalization of classifiers across different facilities.
Innovation Solution
A clustering device and method that utilizes supervised machine learning with under-sampling and sub-sampling, ensemble learning, and harmonization to correct measurement biases, enabling the generation of a classifier model that can accurately discriminate and subtype psychiatric disorders across multiple facilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If supervised machine learning is used to generate classifier models for psychiatric disorders, then diagnostic accuracy can be improved, but the models suffer from overfitting and poor generalization due to small sample sizes and site-to-site differences
Solution Approach 1:
The patent segments the training process into multiple iterations with progressive inclusion of sites. In each iteration, a subset of sites is used for training while others are reserved for testing. This segmentation allows the model to learn from limited data without overfitting to site-specific characteristics, thereby improving generalization while maintaining diagnostic accuracy.
Solution Approach 2:
The patent implements a dynamic training framework where the composition of training and testing sites changes across multiple iterations. The model is retrained in each iteration with different site combinations, allowing it to adapt to various data distributions. This dynamic approach enhances the model's robustness and generalization capability across different facilities while preserving diagnostic precision.
2Quantity of substance
If brain functional connectivity data is collected from multiple facilities to increase sample size, then statistical power is improved, but site-to-site differences introduce measurement bias and reduce data homogeneity
Solution Approach 1:
The patent applies local quality by allowing each site to maintain its own data characteristics while contributing to the overall model training. Instead of forcing uniformity across sites, the method accepts site-specific variations and uses cross-validation to ensure the model performs well across different local conditions. This approach preserves data homogeneity in terms of model performance while accommodating local differences in measurement protocols and populations.
Solution Approach 2:
The patent creates a universal classifier model that functions effectively across multiple sites with different characteristics. The model is designed to be site-agnostic, learning patterns that generalize across diverse facilities. This multi-functionality allows the same model to accurately diagnose psychiatric disorders regardless of which facility's data is used, thereby increasing sample size without compromising data quality.
3Productivity
If conventional machine learning methods are used with small sample sizes, then training time and computational resources are reduced, but the resulting classifiers lack robustness and fail to generalize to new facilities
Solution Approach 1:
The patent employs periodic action through iterative cross-validation where the model is trained and tested in repeated cycles with different site combinations. Each iteration provides a periodic check on model performance, ensuring robustness is maintained. This periodic evaluation approach builds reliable classifiers that generalize well, while the iterative nature allows efficient use of computational resources by reusing training data in different configurations.
Data Source
Figure 1
Figure 2A~2B
Figure 3A~3B
AI summary
A brain functional connectivity correlation value clustering device for clustering subjects having a prescribed attribute on the basis of brain measurement data obtained from a plurality of facilities, wherein a plurality of MRI devices capture resting state fMRI image data of a healthy cohort and a patient cohort; a computing system 300 performs generation of an identifier as ensemble learning of "supervised learning" between harmonized component values of correlation matrixes and disease labels of each of the subjects, selects, during the ensemble learning, features for clustering in accordance with importance from the features specified for generating an identifier for a disease label, and performs multiple co-clustering by "unsupervised learning."