Method and apparatus for generating information based on brain image data
By constructing cross-disease brain morphological similarity features and utilizing longitudinal brain imaging analysis and machine learning models, the problem of lacking integration of longitudinal imaging and clinical evolution information in existing technologies has been solved, enabling high-precision diagnosis and treatment prediction of schizophrenia.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-12
AI Technical Summary
Current technologies in neuroimaging research on schizophrenia lack the ability to integrate longitudinal imaging and clinical evolution information, making it difficult to fully analyze the neurobiological continuum between diseases. Furthermore, existing classification models mostly rely on single-time-point data, which limits the practicality of the models in efficacy assessment and prognosis prediction.
By constructing cross-disease brain morphology similarity features, using longitudinal brain imaging analysis and machine learning models, combined with pre-trained cortical partition maps and brain morphology norms, an individual brain morphology deviation vector is generated. The correlation is calculated with statistical data of various mental illnesses, and the classification information of whether an individual is a schizophrenia patient and the prediction information of clinical indicators are output.
It breaks through the limitations of traditional single-disease research, generates more accurate classification and prediction information, and can assess the commonalities and differences in brain structural characteristics between schizophrenia and other mental illnesses, thus achieving high-precision diagnosis and treatment prediction for schizophrenia.
Smart Images

Figure CN121549772B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and in particular to a method and apparatus for generating information based on brain image data. Background Technology
[0002] In neuroimaging studies of schizophrenia (SCZ), quantitative analysis of brain morphology based on brain imaging data such as magnetic resonance imaging (MRI) and nuclear magnetic resonance imaging (NMRI) has become an important approach to revealing disease-related brain structural abnormalities. Numerous studies have shown that schizophrenia patients exhibit significant volume reduction in subcortical regions such as the hippocampus, thalamus, and amygdala, while also showing significant reductions in gray matter volume and cortical thickness in the frontal, temporal, and parietal lobes. Large-scale, multi-site neuroimaging studies have confirmed these morphological changes, which are closely related to patients' cognitive impairment, symptom severity, and disease progression. Therefore, mining information related to schizophrenia from brain imaging data to aid in the diagnosis and treatment of schizophrenia is of great significance. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method and apparatus for generating information based on brain imaging data, in order to eliminate or improve one or more defects existing in the prior art.
[0004] According to the first aspect, a method for generating information based on brain imaging data is provided, comprising: performing longitudinal brain imaging analysis on brain imaging data of the target individual's brain collected at at least two time points; dividing the target individual's brain into regions and extracting morphological measurements of each region based on the longitudinal brain imaging analysis results and a pre-defined cortical partition map; determining the degree of deviation of the morphological measurements of each region of the target individual's brain at each time point from the pre-established brain morphology norms, and obtaining an individual brain morphology deviation vector corresponding to the target individual at each time point; performing correlation calculation between the individual brain morphology deviation vector corresponding to the target individual at each time point and pre-acquired statistical data of multiple mental illnesses in the whole brain, and obtaining cross-disease brain morphology similarity features of the target individual at each time point; inputting the cross-disease brain morphology similarity features of the target individual into a pre-trained machine learning model, and outputting classification information on whether the target individual is a patient with schizophrenia, and / or predictive information of clinical indicators.
[0005] According to a second aspect, an apparatus for generating information based on brain imaging data is provided, comprising: an extraction unit configured to perform longitudinal brain imaging analysis on brain imaging data of a target individual collected at at least two time points, and to divide the brain of the target individual at each time point into regions and extract morphological measurement values of each region based on the longitudinal brain imaging analysis results and a preset cortical partition map; a determination unit configured to determine the degree of deviation of the morphological measurement values of each region of the target individual's brain at each time point from the brain morphological norms based on a pre-established brain morphological norm, thereby obtaining an individual brain morphological deviation vector corresponding to the target individual at each time point; a calculation unit configured to perform correlation calculation between the individual brain morphological deviation vector corresponding to the target individual at each time point and pre-acquired statistical data of multiple mental illnesses in the whole brain, thereby obtaining cross-disease brain morphological similarity features of the target individual at each time point; and a generation unit configured to input the cross-disease brain morphological similarity features of the target individual into a pre-trained machine learning model, and output classification information on whether the target individual is a patient with schizophrenia, and / or predictive information of clinical indicators.
[0006] According to a third aspect, an apparatus for generating information based on brain imaging data is provided, including a processor, a memory, and a computer program / instructions stored in the memory, wherein the processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, the apparatus performs the steps of any of the methods described in the first aspect.
[0007] According to a fourth aspect, a computer-readable storage medium is provided that stores a computer program / instructions thereon, which, when executed by a processor, implement the steps of the methods described in any of the first aspects.
[0008] The method and apparatus for generating information based on brain imaging data of the present invention first perform longitudinal brain imaging analysis on brain imaging data of the target individual collected at at least two time points, and then divide the brain of the target individual at each time point into regions and extract morphological measurements of each region based on the longitudinal brain imaging analysis results and a preset cortical partition map. Next, based on brain morphological norms, the degree of deviation of the morphological measurements of each region of the target individual's brain at each time point from the brain morphological norms is determined, resulting in an individual brain morphological deviation vector for each time point. Then, the individual brain morphological deviation vectors for each time point are correlated with statistical data of various mental illnesses across the whole brain to obtain cross-disease brain morphological similarity features of the target individual at each time point. Finally, the cross-disease brain morphological similarity features of the target individual are input into a pre-trained machine learning model, which outputs classification information on whether the target individual is a patient with schizophrenia and / or predictive information of clinical indicators. Therefore, by constructing cross-disease brain morphology similarity features, it is possible to assess the commonalities and differences in brain structural features between schizophrenia and other mental illnesses, breaking through the limitations of traditional single-disease studies, thereby making the generated classification and prediction information more accurate.
[0009] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.
[0010] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description
[0011] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to limit the scope of the invention.
[0012] Figure 1 A flowchart is shown showing a method for generating information based on brain imaging data according to one embodiment;
[0013] Figure 2 A schematic diagram illustrating an example of a method for generating information based on brain imaging data according to one embodiment is shown.
[0014] Figure 3 This diagram illustrates an example of a comparison of intergroup differences between schizophrenic patients and healthy individuals.
[0015] Figures 4A-4FThe diagram shows the ROC plot of the SVM classification results (repeated 100 times) for SCZ and hc, along with a schematic diagram illustrating the contributions of features and brain regions. Figure 4A The baseline period results are shown; Figure 4B The follow-up results are shown; Figure 4C The figure shows the percentage difference between the AUC values of the SVM classification model using all features at baseline and the AUC values after removing one feature at a time. Figure 4D The figure shows the percentage difference between the AUC values of the SVM classification model using all features and the AUC values after removing one feature at a time during the follow-up period. Figure 4E The percentage difference between the AUC values of the SVM classification model at baseline using all brain regions and the AUC values after sequentially removing one brain region at a time is shown. Figure 4F The percentage difference in AUC values of the SVM classification model using all brain regions versus AUC values after sequentially removing one brain region at the follow-up period is shown.
[0016] Figure 5 A schematic diagram illustrating an example of a multivariate correlation analysis between MSP scores and clinical symptoms is shown.
[0017] Figure 6 A schematic diagram showing an example of PANSS difference prediction results with 100 repeated nested cross-validations;
[0018] Figure 7 A schematic block diagram of an apparatus for generating information based on brain imaging data according to one embodiment is shown. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0020] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0021] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0022] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.
[0023] It is understood that the ordinal numbers such as "first" and "second" mentioned in this specification are only used to distinguish multiple objects of the same or different categories (such as components, steps, parameters, etc.), and do not indicate the priority, importance or order relationship between objects, nor do they constitute a limitation on the technical features.
[0024] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.
[0025] As mentioned earlier, mining information related to schizophrenia through brain imaging data is of great significance in assisting the diagnosis and treatment of schizophrenia.
[0026] In recent years, domestic and international research has further revealed that brain structural changes in schizophrenia patients can also be observed in other mental illnesses. For example, similar patterns of brain structural abnormalities have been observed in mental illnesses such as bipolar disorder (BD), major depressive disorder (MDD), and autism spectrum disorder (ASD), suggesting that there may be common pathobiological mechanisms among mental illnesses.
[0027] Existing methods have made some progress in identifying disease heterogeneity, such as various machine learning methods applied to subtype identification and classification prediction. However, these methods still have several technical limitations. For example, existing studies are mostly based on cross-sectional analysis and lack the ability to integrate longitudinal imaging and clinical evolution information. Schizophrenia, as a progressive disease, has brain structural changes closely related to symptom evolution, but existing classification models mostly rely on single-time-point data and do not fully consider the dynamic changes in brain morphology at multiple time points and their correlation with clinical indicators (such as PANSS scores and treatment response), limiting the practicality of the models in efficacy assessment and prognostic prediction. PANSS stands for Positive and Negative Syndrome Scale. Furthermore, current machine learning methods mostly focus on subtype classification within a single disease and lack effective computational models for systematically identifying common brain abnormality patterns and genetic associations across mental illnesses, making it difficult to comprehensively analyze the neurobiological continuum between diseases.
[0028] Therefore, this specification provides a method for generating information based on brain imaging data. By constructing cross-disease brain morphological similarity features, it is possible to assess the commonalities and differences in brain structural features between schizophrenia and other mental illnesses, breaking through the limitations of traditional single-disease studies, thereby making the generated classification and prediction information more accurate.
[0029] Please see Figure 1 , Figure 1 A flowchart illustrating a method for generating information based on brain imaging data according to one embodiment is shown. It will be understood that this method can be executed by any device, apparatus, platform, or cluster of devices with computing and processing capabilities. Figure 1 As shown, the method for generating information based on brain imaging data may include the following steps 101 to 104, specifically:
[0030] Step 101 involves performing longitudinal brain imaging analysis on brain imaging data of the target individuals collected at at least two time points, and dividing the brain of the target individuals at each time point into regions and extracting morphological measurements of each region based on the results of the longitudinal brain imaging analysis and a preset cortical partition map.
[0031] In this embodiment, brain imaging data of the target individual's brain can be acquired at at least two time points, such as magnetic resonance imaging (MRI) and T1-weighted magnetic resonance imaging (T1-weighted MRI). The target individual may include a patient with schizophrenia or a suspected patient with schizophrenia. As an example, the at least two time points may include a baseline and a follow-up time point, where the baseline may represent the time of the first acquisition of brain imaging data, and the follow-up time point may represent the time of subsequent acquisitions. Subsequently, longitudinal brain imaging analysis can be performed on the brain imaging data of the target individual acquired at the at least two time points. Based on the longitudinal brain imaging analysis results and a pre-defined cortical region atlas, the brain of the target individual at each time point is divided into regions, and morphological measurements of each region are extracted. As an example, the pre-defined cortical region atlas may include various brain atlases, such as the Harvard-Oxford atlas and the Desikan-Killiany atlas.
[0032] In some examples, the aforementioned brain imaging data may include T1-weighted structural magnetic resonance imaging. Based on this, step 101 may include steps 1011-1013, specifically:
[0033] Step 1011: Longitudinal brain imaging analysis is performed on brain imaging data of the target individual acquired at at least two time points using a longitudinal processing workflow of a software suite for processing and analyzing magnetic resonance imaging data of the human brain.
[0034] In this example, for T1-weighted structural MRI scans of the target individual's brain acquired at at least two time points, longitudinal preprocessing can be performed using software suites for processing and analyzing human brain MRI data, such as FreeSurfer (v6.0), ANTs (Advanced Normalization Tools), FastSurfer, and sMRIPrep (Structural Magnetic Resonance Imaging PREProcessing pipeline). This preprocessing may include cranial dissection, spatial normalization, registration, and brain region segmentation. For instance, using FreeSurfer (v6.0)'s longitudinal processing workflow ensures consistency across multiple time points, resulting in more accurate longitudinal brain imaging analysis. As a control, brain imaging data from healthy individuals can also be acquired; for example, single-time-point T1-weighted structural MRI scans of healthy individuals can be acquired and processed using FreeSurfer's default processing workflow.
[0035] Step 1012: Based on the longitudinal brain imaging analysis results and the preset cortical partition map, the brain of the target individual at each time point is divided into regions to obtain multiple cortical thickness regions and multiple subcortical volume regions.
[0036] In this example, the brain regions of the target individual at each time point can be divided based on a preset cortical region map, thereby obtaining multiple cortical thickness regions and multiple subcortical volume regions. Taking the Desikan-Killiany cortical region map as an example, 68 cortical thickness regions and 14 subcortical volume regions can be obtained based on the Desikan-Killiany map, totaling 82 brain regions. Specifically, the 68 cortical thickness regions can include the left superior temporal sulcus, left caudal anterior cingulate cortex, left middle frontal gyrus tail, left cuneus, left entorhinal cortex, left fusiform gyrus, left inferior parietal lobe, left inferior temporal lobe, left cingulate isthmus, left lateral occipital lobe, left lateral orbitofrontal lobe, left lingual gyrus, left medial orbitofrontal lobe, left middle temporal lobe, left parahippocampal gyrus, left paracentral lobule, left operculum, left orbitomegaly, left pericalcaneal fissure, left postcentral gyrus, left posterior cingulate cortex, left precentral gyrus, left precuneus, left rostral anterior cingulate cortex, left rostral middle frontal lobe, left superior frontal lobe, left superior parietal lobe, left supramarginal gyrus, left frontal pole, left temporal pole, left transverse temporal gyrus, and left insula, as well as the corresponding right-sided portion. The 14 subcortical volume regions can include the left thalamus, left caudate nucleus, left putamen, left globus pallidus, left hippocampus, left amygdala, and left nucleus accumbens, as well as the corresponding right-sided portion. It is understandable that dividing the brain into 68 cortical thickness regions and 14 subcortical volume regions is a very classic and widely used partitioning scheme in modern brain imaging analysis. It is mainly based on the Desikan-Killiany (DK) atlas and its extensions, which will not be elaborated here.
[0037] Step 1013: Extract the measured values of the cortical thickness index of each cortical thickness region and the measured values of the subcortical volume index of each subcortical volume region as morphological measurement values for each region.
[0038] In this example, for regions of cortical thickness, the measured values of cortical thickness indices for each region can be extracted as morphological measurements. For regions of subcortical volume, the measured values of subcortical volume indices for each region can be extracted as morphological measurements. It is understood that after processing magnetic resonance imaging (MRI) data using software suites for processing and analyzing human brain MRI data, the processing results can include morphological measurements of each segmented brain region, and therefore, these measurements can be directly extracted.
[0039] Step 102: Based on the pre-established brain morphology norms, determine the degree of deviation of the morphological measurement values of each region of the target individual's brain at each time point from the brain morphology norms, and obtain the individual brain morphology deviation vector corresponding to the target individual at each time point.
[0040] In this embodiment, brain morphology norms can be pre-established. These norms refer to a set of standardized reference data on the normal structural characteristics of various brain regions, established by collecting and analyzing brain structural imaging data from a large number of healthy individuals. Simply put, it's a "standard brain structure atlas" or a "reference standard for a healthy brain," describing the normal ranges of structural indicators such as volume, thickness, and surface area that a healthy brain should possess under specific age, gender, and population conditions. For example, the CentileBrain framework can be used to construct gender- and age-specific brain morphology norms, and the MFPR (Multivariate Fractional Polynomial Regression) algorithm can be used to build an optimized model. Thus, using a large-scale sample and the CentileBrain framework, high-precision, individualized brain imaging norms for schizophrenia can be constructed. This provides a unified and reliable benchmark for cross-disease and cross-center brain imaging comparisons, improving the accuracy and reproducibility of subsequent analyses.
[0041] Based on a pre-established brain morphology norm that matches the gender and age of the target individual (e.g., same gender, similar age or the same), the degree of deviation of the morphological measurements of each region of the target individual's brain at each time point from the brain morphology norm can be determined, thereby obtaining the individual brain morphology deviation vector corresponding to the target individual at each time point. Each dimension of the individual brain morphology deviation vector corresponds to each region of the region division.
[0042] In some examples, step 102 above may specifically include the following: based on a pre-established brain morphology norm, calculate the z-score of the morphological measurement values of each region of the target individual's brain at each time point relative to the measurement values of the brain morphology norm, and obtain the z-score vector corresponding to the target individual at each time point as the individual brain morphology deviation vector.
[0043] In this example, for the target individual at each time point, the z-score of their actual measurement relative to the norm can be calculated for each brain region to quantify the degree of deviation from the normal level. For example, the formula for calculating the z-score is as follows:
[0044] ,
[0045] in, This can represent the morphological measurement value of the i-th brain region. and These can represent the mean and standard deviation of morphological measurements of the brain region in the age- and sex-matched brain morphological norms, respectively.
[0046] The z-scores of each brain region can be used to form a z-score vector, which can be used as an individual brain morphological deviation vector.
[0047] Step 103: The individual brain morphology deviation vector corresponding to the target individual at each time point is correlated with the statistical data of various mental diseases in the whole brain obtained in advance to obtain the cross-disease brain morphology similarity characteristics of the target individual at each time point.
[0048] In this embodiment, statistical data on multiple mental illnesses across the whole brain can be obtained in advance. For example, comprehensive statistics on multiple mental disorders across the whole brain can be extracted from the ENIGMA Consortium's public database. These statistics may include Cohen's d effect size and p-values corrected for FDR (False Discovery Rate) for morphological measurements of each brain region. Then, the correlation between the individual brain morphological deviation vectors corresponding to the target individuals at each time point and the statistical data on multiple mental illnesses across the whole brain can be calculated. For example, the Pearson correlation coefficient can be calculated to obtain the cross-disease brain morphological similarity characteristics of the target individuals at each time point.
[0049] In some examples, the aforementioned mental illnesses may include at least two of the following: 22q11.2 deletion syndrome, attention deficit hyperactivity disorder (ADHD), autism spectrum disorder, bipolar disorder, epilepsy, major depressive disorder, obsessive-compulsive disorder (OCD), and schizophrenia.
[0050] In some examples, statistical data for each mental illness across the whole brain may include an effect size vector of morphological measurements. Based on this, step 103 above may include steps 1) and 2), specifically:
[0051] Step 1) Calculate the Pearson correlation between the individual brain morphology deviation vector corresponding to the target individual at each time point and the effect size vector of the morphological measurement values of various mental illnesses, and obtain the correlation coefficient between the individual brain morphology deviation vector and the effect size vector of the morphological measurement values of various mental illnesses.
[0052] As an example, suppose we consider eight mental illnesses, including 22q11.2 deletion syndrome, attention deficit hyperactivity disorder (ADHD), autism spectrum disorder, bipolar disorder, epilepsy, major depressive disorder, obsessive-compulsive disorder (OCD), and schizophrenia. For the target individual, we can calculate the Pearson correlation coefficient between the individual's brain morphological deviation vector at each time point and the effect size vector of the whole-brain morphological measurements for each disease. This is calculated separately for cortical thickness and subcortical volume indices. The formula for the Pearson correlation coefficient is:
[0053]
[0054] in, This can represent the z-score of the i-th brain region of the target individual. The Cohen's d value can represent the morphological measurement of the disease in the i-th brain region. and Let z and d represent the mean of the z-score and Cohen's d-value, respectively.
[0055] Ultimately, each target individual at each time point can obtain 16 cross-disease MSP (morphological similarity profile) scores. These 16 scores include scores from 8 diseases × 2 modalities (including cortical thickness and subcortical volume indicators). These 16 cross-disease MSP scores constitute cross-disease brain morphological similarity features, which can form the core input features for subsequent machine learning models. Therefore, whole-brain summary statistical features can be extracted using the Enigma Toolbox, and similarity network analysis can be performed with 8 types of mental illnesses to construct cross-disease MSP scores. This systematically assesses the commonalities and differences in brain structural features between schizophrenia and other mental disorders, thereby identifying specific brain features of schizophrenia and shared pathological bases across diseases, providing new evidence for the classification and mechanistic analysis of mental illnesses.
[0056] Step 104: Input the cross-disease brain morphology similarity features of the target individual into the pre-trained machine learning model, and output classification information on whether the target individual is a schizophrenia patient and / or prediction information of clinical indicators.
[0057] In this embodiment, various machine learning models can be pre-built and trained for schizophrenia. These models can take cross-disease brain morphological similarity features as input and output information related to schizophrenia. For example, these machine learning models may include classification models for outputting whether an individual is a schizophrenia patient; for instance, the classification model could be a decision tree model, a convolutional neural network model, etc. These machine learning models may also include predictive models for outputting predictive information for clinical indicators; for instance, the predictive model could be a logistic regression model, a convolutional neural network model, etc. In practice, the cross-disease brain morphological similarity features of the target individual at each time point can be input into the machine learning models to obtain classification information and / or predictive information. Thus, classification information and / or predictive information for the target individual at each time point can be obtained for analysis.
[0058] In some examples, the machine learning model may include a Support Vector Machine (SVM) classification model based on Radial Basis Function (RBF). This RBF-based SVM classification model can be used to output classification information based on an individual's cross-disease brain morphological similarity features, indicating whether the individual has schizophrenia. For example, the input to an RBF-based SVM classification model can be an individual's cross-disease brain morphological similarity features, and the output can include 0 and 1, where 0 represents a healthy individual and 1 represents an individual with schizophrenia.
[0059] In this example, a supervised training method can be used to train an RBF-based SVM classification model. For instance, the model can be trained using the following steps: First, obtain a training sample set. Each training sample may include cross-disease brain morphological similarity features corresponding to the individual sample, as well as label information indicating whether the individual sample is a schizophrenic patient. Then, input the cross-disease brain morphological similarity features from the training samples into the classification model, which then outputs a classification result. Finally, determine the difference loss between the classification result output by the model and the label information, and adjust the model's parameters with the goal of minimizing this difference loss.
[0060] During the classification model training phase, five-fold cross-validation can be repeated multiple times (e.g., 100 times) to evaluate model performance. For example, evaluation metrics could include AUC, accuracy, sensitivity, and specificity. To further identify key features, ablation experiments can be used to systematically remove features or all features contributing to a specific brain region to assess their impact on model performance. In this example, using the aforementioned MSP features, a nonlinear support vector machine (SVM) model with a radial basis function (RBF) is constructed, effectively capturing the nonlinear discriminant boundary between patients and healthy individuals. Through five-fold cross-validation and repeated experiments, the model's performance in AUC, accuracy, sensitivity, and specificity is systematically evaluated. Ablation experiments are used to clarify key features and brain region contributions, improving the model's interpretability and generalization ability.
[0061] In some examples, the above-described method for generating information based on brain imaging data may also include: using partial least squares regression to calculate the correlation between cross-disease brain morphological similarity features and measurements of clinical indicators among multiple schizophrenia patients.
[0062] In this example, partial least squares (PLS) regression can be used to explore the multivariate correlation between cross-disease brain morphological similarity features and clinical indicators among multiple schizophrenia patients. In this example, the cross-disease brain morphological similarity features of multiple schizophrenia patients can form an X matrix, and the clinical indicators of multiple schizophrenia patients can form a Y matrix. The clinical indicators may include the PANSS (Positive and Negative Syndrome Scale) total score, sub-scores, and treatment response rate, etc. The process of calculating the multivariate correlation of the X and Y matrices using PLS may include the following steps (1)-(4), specifically:
[0063] Step (1), calculate the cross covariance matrix:
[0064] .
[0065] matrix It captures the sum of the interrelationships between all brain features and all clinical indicators, laying the foundation for extracting common components.
[0066] Step (2), Singular value decomposition (SVD):
[0067] Singular value decomposition is performed on the cross-covariance matrix R to generate three low-dimensional matrices U, S, and V, where U and V are singular vectors (called behavioral and imaging saliency, reflecting the contribution of the original variables to the latent components), and S is a diagonal matrix containing the singular values.
[0068] .
[0069] Step (3), calculate subject-specific score:
[0070] For each latent component, subject-specific imaging (Lx) and behavioral scores (Ly) are calculated by projecting X and Y onto their respective saliences (V and U).
[0071] Step (4), calculate the load or significance:
[0072] Imaging and behavioral structure coefficients (or “loadings”) are obtained by calculating the Pearson correlation between the raw data (X and Y) and subject-specific scores (Lx and Ly). It is understood that using partial least squares regression to explore multivariate correlations between two matrices is a mature and widely used existing technique, which will not be elaborated upon here. In this example, cross-disease brain morphology similarity indicators are used as predictive features. A multivariate association model between these indicators and clinical indicators such as the PANSS scale is established using partial least squares (PLS). Subsequently, LASSO regression is combined for feature selection and predictive modeling. Nested cross-validation is used to optimize predictive accuracy, ultimately achieving quantitative prediction of treatment response and symptom changes, providing data-driven support for clinical intervention strategy development. This reveals the multidimensional association between brain morphology similarity indicators and clinical phenotypes. Traditional analyses often struggle to integrate multimodal data to identify significant associations.
[0073] In some examples, the machine learning model may include a LASSO (Least Absolute Shrinkage and Selection Operator) regression prediction model, which can be used to predict an individual's clinical indicators based on cross-disease brain morphological similarity features. Furthermore, nested cross-validation can be used to train the LASSO regression prediction model and optimize its regularization parameters.
[0074] In this example, a LASSO regression prediction model can be established using cross-disease brain morphological similarity features as independent variables and clinical indicators as dependent variables. Taking the brain imaging data acquisition time points as t1 and t2 as an example, the cross-disease brain morphological similarity features of the target individual at the two time points t1 and t2 can be input into the LASSO regression prediction model to obtain predictive information on clinical indicators and treatment response. Clinical indicators can include the treatment response rate, the PANSS (Positive and Negative Syndrome Scale) clinical rating scale at t1 and t2, and the change in PANSS during the time interval from t1 to t2, etc. Here, t1 and t2 can represent the acquisition time of the brain imaging data, with t1 earlier than t2. The treatment response rate can represent the percentage of patients who achieve a predefined "effective" or "responsive" standard after receiving a certain treatment.
[0075] In building a LASSO regression prediction model, λ is a hyperparameter controlling the model complexity; it's the coefficient of the L1 penalty term (the sum of the absolute values of all regression coefficients) in the loss function. The value of λ directly determines the strength of the penalty. In this example, the regularization parameter λ can be optimized using 5-fold nested cross-validation. In each run of nested cross-validation, the inner loop is used to fine-tune the penalty parameter from a pool of candidate λ values using a grid search. The candidate λ value that minimizes the mean square error (MSE) is selected, and this λ is used to train the regression model on the complete inner loop dataset. The trained model is then applied to the test data of the outer loop. After 100 iterations of nested cross-validation, predicted values are generated based on the outer loop test data. Finally, the Pearson correlation coefficient between the predicted and true values is calculated to evaluate the prediction performance.
[0076] For example, the training process of nested cross-validation can be as follows: Step 1) Divide all data into 5 folds, taking 4 folds as the training set and 1 fold as the test set each time (outer loop); Step 2) Divide the training set into 5 folds again for grid search to select λ (inner loop); Step 3) In the inner loop, for each λ value, use 5-fold cross-validation to calculate the mean squared error and select the λ that minimizes the mean squared error; Step 4) Train the LASSO regression prediction model on the entire training set using the optimal λ; Step 5) Make predictions on the test set and store the actual and predicted values; Step 6) Repeat the entire nested cross-validation process 100 times, randomly shuffling the data each time. Finally, average the 100 prediction results for each sample and calculate the correlation between the average predicted value and the actual value.
[0077] The method for generating information based on brain imaging data in the embodiments of this specification can be implemented by first setting up a corresponding operating environment and then executing the method. For example, it can be achieved through the following steps (1)-(4):
[0078] Step (1), Environment configuration and data preparation.
[0079] In practice, users can configure MATLAB (R2021a or later) and Python (3.8 or later) environments and install necessary toolkits (e.g., tools for providing machine learning models, plotting tools, data processing tools, etc.). T1-weighted MRI data and demographic information of the subjects are used to calculate brain morphological norms using the CentileBrain framework.
[0080] Step (2), calculate cross-disease brain morphological similarity features.
[0081] Taking eight mental illnesses as an example, the multi-disease statistics provided by the ENIGMA Alliance public database are loaded, and the Pearson correlation coefficient between each subject and the eight mental illnesses in two modalities, cortical thickness and subcortical volume, is calculated. Finally, an N×16 feature matrix (N is the number of subjects) is output.
[0082] Step (3): Train and evaluate the classification and prediction model.
[0083] Taking the RBF-based SVM classification model as an example, with cross-disease brain morphological similarity features as input, the RBF-based SVM classification model is used to distinguish between patients and healthy controls. The default is to use 5-fold cross-validation repeated 100 times, and the output is the average AUC, accuracy, sensitivity, specificity and individual classification probability.
[0084] Key features and brain region contributions are assessed, for example, by conducting feature ablation and brain region ablation experiments to evaluate the impact of each feature and brain region on model performance.
[0085] Taking the LASSO regression prediction model as an example, the prediction of clinical symptoms includes: using cross-disease brain morphological similarity features as independent variables and clinical indicators (such as PANSS scores) as dependent variables, performing LASSO regression with nested cross-validation, and outputting prediction scores, feature weights, and prediction correlations. The prediction score can represent the value of the predicted clinical indicator, the feature weights can represent the trained model parameters, and the prediction correlation can refer to the correlation between the predicted value and the true value.
[0086] In addition, multivariate association analysis can be performed to analyze the multivariate correlation between cross-disease brain morphological similarity features and clinical behavioral data, extract latent variables, and visualize the covariance structure.
[0087] Step (4), result visualization and report generation.
[0088] During the results visualization and report generation process, box plots for group comparisons and longitudinal comparisons can be generated for easy viewing. Additionally, changes in model performance after feature and brain region ablation can be visualized for easy viewing.
[0089] As can be seen from the above process, the method for generating information based on brain imaging data in this application embodiment can establish norms through CentileBrain, introduce ENIGA multi-disease summary statistical data, and construct cross-disease brain morphological similarity features. This allows for a systematic assessment of the commonalities and differences in brain structural features between schizophrenia and other mental illnesses, breaking through the limitations of traditional single-disease studies. This reveals the neurobiological basis of cross-diagnosis and provides new evidence for the reclassification and mechanism analysis of mental illnesses. Furthermore, high-precision patient identification is achieved using SVM with RBF nuclei, and key features and brain region contributions are clarified through ablation experiments. Further, PLS and LASSO regressions are combined to establish a multivariate predictive model of brain features versus clinical symptoms. This model assists in diagnosis while enabling individualized prediction of treatment response and symptom evolution, balancing classification identification and predictive interpretability. The application scenarios of this invention are wide-ranging, applicable to the diagnosis and efficacy prediction of schizophrenia patients, and suitable for multi-center, large-sample mental illness research and clinical practice.
[0090] See also Figure 2 , Figure 2 A schematic diagram illustrating an example of a method for generating information based on brain imaging data according to one embodiment is shown. Figure 2 The example shown may include subgraphs A, B, C, D, E, and F. Specifically:
[0091] Subfigure A illustrates the data preprocessing stage, where FreeSurfer's Longitudinal Processing workflow was used to process T1-weighted MRI images of 100 (N=100) patients with schizophrenia (SCZ) during the baseline and follow-up periods. Simultaneously, data from 97 (N=97) healthy controls (HC) were processed using FreeSurfer's default workflow to obtain morphological features.
[0092] Subfigure B illustrates the process of calculating the MSP score. Individualized region bias (z-score) is obtained by using sex-specific norms for cortical thickness and subcortical volume through the CentileBrain framework. Then, Pearson correlations between brain region z-scores and multiple comprehensive statistics of mental illnesses (Cohens' d values) extracted from the ENIGMA toolkit are calculated to obtain similarity features.
[0093] The C subplot shows the difference in MSP scores between patients with schizophrenia and healthy controls at baseline and follow-up.
[0094] Subgraph D illustrates the SVM classification models trained on schizophrenia patients and healthy controls.
[0095] The E subplot illustrates the use of partial least squares (PLS) analysis to explore the correlation between cross-disease pattern scores and clinical indicators.
[0096] The F subplot illustrates the use of Lasso regression to predict clinical indicators. In this case, the regularization parameter λ can be optimized through nested cross-validation.
[0097] Please see Figure 3 , Figure 3 This diagram illustrates an example of a comparison of differences between groups of patients with schizophrenia and healthy individuals. Figure 3 The example shown may include subgraph A and subgraph B, specifically:
[0098] Subplot A shows individualized cross-disease MSP scores for subcortical volume in healthy controls (HC), schizophrenia patients at baseline (SCZ_T1), and during follow-up (SCZ_T2). Similarity is compared with bipolar disorder (BD), major depressive disorder (MDD), 22q11.2 syndrome, and schizophrenia (SCZ). The brain map also displays the Cohens'd value for subcortical volume in the ENIGA statistics.
[0099] Subplot B shows individualized cross-disease MSP scores for cortical thickness in healthy controls, schizophrenia patients at baseline (SCZ_T1), and during follow-up (SCZ_T2). Similarity comparisons are made with bipolar disorder, major depressive disorder, obsessive-compulsive disorder (OCD), and schizophrenia. The brain map also displays Cohens'd values for cortical thickness in the ENIGA statistics. p<0.05; p < 0.005.
[0100] Please continue reading Figures 4A-4F , Figures 4A-4F The diagram shows the ROC plot of the SVM classification results (repeated 100 times) for SCZ and HC, along with a schematic representation of the contributions of features and brain regions. In this example, it may include... Figure 4A Subgraph Figure 4B Subgraph Figure 4C Subgraph Figure 4D Subgraph Figure 4E Subgraph and Figure 4F Subgraph.
[0101] Figure 4A The subfigures show the baseline results. The left side shows the results using subcortical volume and cortical thickness as features (AUC = 0.83 ± 0.01); the upper right image shows the results using only subcortical volume (AUC = 0.68 ± 0.02); and the lower right image shows the results using only cortical thickness (AUC = 0.77 ± 0.01).
[0102] Figure 4B The subfigures show the follow-up results. The left side shows the results when both subcortical volume and cortical thickness were used (AUC = 0.87 ± 0.01); the upper right image shows the results when only subcortical volume was used (AUC = 0.68 ± 0.02); and the lower right image shows the results when only cortical thickness was used (AUC = 0.82 ± 0.01).
[0103] Figure 4C The subplot shows the percentage difference between the AUC values of the SVM classification model using all features at baseline and the AUC values after removing one feature at a time. In this example, we use 16 scores from cross-disease brain morphological similarity features, comprising 8 disease × 2 modalities (including cortical thickness and subcortical volume indices).
[0104] Figure 4DThe subplot shows the percentage difference between the AUC values of the SVM classification model using all features and the AUC values after removing one feature at a time during the follow-up period. In this example, the cross-disease brain morphological similarity features include 16 scores across 8 disease × 2 modalities (including cortical thickness and subcortical volume indices).
[0105] Figure 4E The subplot shows the percentage difference between the AUC values of the SVM classification model using all brain regions at baseline and the AUC values after sequentially removing one brain region at a time.
[0106] Figure 4F The subplot shows the percentage difference between the AUC values of the SVM classification model using all brain regions and the AUC values after sequentially removing one brain region at a time during the follow-up period.
[0107] Figure 5 This diagram illustrates an example of a multivariate correlation analysis between MSP scores and clinical symptoms. Figure 5 The example shown may include subgraphs A, B, and C.
[0108] Subplot A shows the correlation between the MSP composite score and the subject's symptom composite score.
[0109] Subplot B shows the z-score histograms of the imaging and behavioral structural coefficients for each ENIGMA disease, with blue bars representing statistically significant values.
[0110] Subplot C shows the z-score values of the imaging and behavioral structure coefficients for each clinical symptom, with blue indicating statistical significance. As an example, subplot C includes the total score of the PANSS scale, positive symptom score, negative symptom score, and general psychopathological symptom score.
[0111] Figure 6 A schematic diagram shows an example of the PANSS difference prediction results with 100 repeated nested cross-validations.
[0112] According to another embodiment, an apparatus for generating information based on brain imaging data is provided. This apparatus for generating information based on brain imaging data can be deployed in any device, platform, or cluster of devices with computing and processing capabilities.
[0113] Figure 7 A schematic block diagram of an apparatus for generating information based on brain imaging data according to one embodiment is shown. Figure 7As shown, the device 700 for generating information based on brain imaging data includes: an extraction unit 701, configured to perform longitudinal brain imaging analysis on brain imaging data of the target individual's brain collected at at least two time points, and to divide the brain of the target individual at each time point into regions and extract morphological measurement values of each region based on the longitudinal brain imaging analysis results and a preset cortical partition map; a determination unit 702, configured to determine the degree of deviation of the morphological measurement values of each region of the target individual's brain at each time point from the brain morphological norm based on a pre-established brain morphological norm, and obtain an individual brain morphological deviation vector corresponding to the target individual at each time point; a calculation unit 703, configured to perform correlation calculation between the individual brain morphological deviation vector corresponding to the target individual at each time point and pre-acquired statistical data of multiple mental illnesses in the whole brain, and obtain cross-disease brain morphological similarity features of the target individual at each time point; and a generation unit 704, configured to input the cross-disease brain morphological similarity features of the target individual into a pre-trained machine learning model, and output classification information and / or prediction information of clinical indicators of whether the target individual is a patient with schizophrenia.
[0114] According to another embodiment, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform actions such as... Figure 1 The method described herein. The computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0115] According to another embodiment, a computing device is also provided, including a processor, a memory, and a computer program / instructions stored in the memory, characterized in that the processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, the device performs the following functions: Figure 1 The method described.
[0116] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.
[0117] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0118] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A device for generating information based on brain imaging data, comprising: The extraction unit is configured to perform longitudinal brain imaging analysis on brain imaging data of the target individual collected at at least two time points, and to perform regional division of the brain of the target individual at each time point and extract morphological measurement values of each region based on the longitudinal brain imaging analysis results and a preset cortical partition map; wherein, the at least two time points include a baseline and a follow-up time point, the baseline representing the time of the first acquisition of brain imaging data, and the follow-up time point representing the time of subsequent acquisition of brain imaging data; the longitudinal brain imaging analysis on the brain imaging data of the target individual collected at at least two time points includes: performing longitudinal brain imaging analysis on the brain imaging data of the target individual collected at at least two time points using a longitudinal processing flow of a software suite for processing and analyzing magnetic resonance imaging data of the human brain; The determination unit is configured to, based on a pre-established brain morphology norm, determine the degree of deviation of the morphological measurement values of each region of the target individual's brain at each time point from the brain morphology norm, and obtain the individual brain morphology deviation vector corresponding to the target individual at each time point. The calculation unit is configured to perform correlation calculation between the individual brain morphology deviation vector corresponding to the target individual at each time point and the statistical data of multiple mental diseases in the whole brain obtained in advance, so as to obtain the cross-disease brain morphology similarity features of the target individual at each time point. The generation unit is configured to input the cross-disease brain morphological similarity features of the target individual into a pre-trained machine learning model, and output classification information and / or prediction information of clinical indicators of whether the target individual is a schizophrenia patient.
2. The apparatus according to claim 1, characterized in that, The machine learning model includes a support vector machine classification model based on radial basis functions. This model is used to output classification information on whether an individual is a patient with schizophrenia based on cross-disease brain morphological similarity features.
3. The apparatus according to claim 1 or 2, wherein, The machine learning model includes a LASSO regression prediction model, which is used to output predictive information of an individual's clinical indicators based on the individual's cross-disease brain morphological similarity features; and the LASSO regression prediction model is trained and its regularization parameters are optimized using a nested cross-validation method.
4. The apparatus according to claim 1, characterized in that, Based on a pre-established brain morphology norm, the deviation of morphological measurements of various brain regions of the target individual at each time point from the brain morphology norm is determined, resulting in an individual brain morphology deviation vector for the target individual at each time point, including: Based on the pre-established brain morphology norms, the z-scores of the morphological measurements of each region of the target individual's brain at each time point are calculated relative to the measurements of the brain morphology norms. The z-score vector corresponding to the target individual at each time point is then used as the individual brain morphology deviation vector.
5. The apparatus according to claim 1, characterized in that, The statistical data for each mental illness across the whole brain includes the effect size vector of morphological measurements; and the correlation calculation of the individual brain morphological deviation vector corresponding to the target individual at each time point with the pre-acquired statistical data for multiple mental illnesses across the whole brain, to obtain the cross-disease brain morphological similarity features of the target individual at each time point, including: The Pearson correlation between the individual brain morphology deviation vector corresponding to the target individual at each time point and the effect size vector of the morphological measurement values of various mental illnesses is calculated to obtain the correlation coefficient between the individual brain morphology deviation vector and the effect size vector of the morphological measurement values of various mental illnesses. Based on the obtained multiple correlation coefficients, the cross-disease brain morphological similarity characteristics of the target individual are determined.
6. The apparatus according to claim 1, characterized in that, Brain imaging data includes T1-weighted magnetic resonance imaging data; and the process of dividing the brain of the target individual at each time point into regions and extracting morphological measurements of each region based on longitudinal brain imaging analysis results and a pre-defined cortical region map includes: Based on the results of longitudinal brain imaging analysis and the pre-set cortical partition map, the brains of the target individuals at each time point were divided into regions, resulting in multiple cortical thickness regions and multiple subcortical volume regions. The measured values of cortical thickness index for each cortical thickness region and the measured values of subcortical volume index for each subcortical volume region were extracted as morphological measurements for each region.
7. The apparatus according to claim 1, characterized in that, The device further includes a unit that performs the following steps: Partial least squares regression was used to calculate the correlation between cross-disease brain morphological similarity features and clinical indicators among multiple schizophrenia patients.
8. The apparatus according to claim 1, characterized in that, The aforementioned mental illnesses include at least two of the following: 22q11.2 deficiency syndrome, attention deficit hyperactivity disorder, autism spectrum disorder, bipolar disorder, epilepsy, major depressive disorder, obsessive-compulsive disorder, and schizophrenia.
9. A computing device, comprising a processor, a memory, and computer programs / instructions stored in the memory, characterized in that, The processor is configured to execute the computer program / instructions. When the computer program / instructions are executed, the device implements a method for generating information based on brain imaging data, the method comprising: The method involves performing longitudinal brain imaging analysis on brain imaging data of the target individual collected at at least two time points, and dividing the brain of the target individual at each time point into regions and extracting morphological measurements of each region based on the longitudinal brain imaging analysis results and a pre-defined cortical region map. The at least two time points include a baseline and a follow-up time point, where the baseline represents the time of the first acquisition of brain imaging data, and the follow-up time point represents the time of subsequent acquisitions of brain imaging data. The longitudinal brain imaging analysis of the target individual's brain imaging data collected at at least two time points includes: performing longitudinal brain imaging analysis on the brain imaging data of the target individual's brain collected at at least two time points using a longitudinal processing workflow of a software suite for processing and analyzing magnetic resonance imaging data of the human brain. Based on the pre-established brain morphology norms, the degree of deviation of the morphological measurement values of each region of the target individual's brain at each time point from the brain morphology norms is determined, and the individual brain morphology deviation vector corresponding to the target individual at each time point is obtained. The correlation between the individual brain morphology deviation vectors of the target individuals at each time point and the statistical data of various mental diseases in the whole brain obtained in advance is calculated to obtain the cross-disease brain morphology similarity features of the target individuals at each time point. The cross-disease brain morphology similarity features of the target individual are input into a pre-trained machine learning model, which outputs classification information on whether the target individual is a patient with schizophrenia and / or prediction information on clinical indicators.
10. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When executed by a processor, the computer program / instructions implement a method for generating information based on brain imaging data, the method comprising: The method involves performing longitudinal brain imaging analysis on brain imaging data of the target individual collected at at least two time points, and dividing the brain of the target individual at each time point into regions and extracting morphological measurements of each region based on the longitudinal brain imaging analysis results and a pre-defined cortical region map. The at least two time points include a baseline and a follow-up time point, where the baseline represents the time of the first acquisition of brain imaging data, and the follow-up time point represents the time of subsequent acquisitions of brain imaging data. The longitudinal brain imaging analysis of the target individual's brain imaging data collected at at least two time points includes: performing longitudinal brain imaging analysis on the brain imaging data of the target individual's brain collected at at least two time points using a longitudinal processing workflow of a software suite for processing and analyzing magnetic resonance imaging data of the human brain. Based on the pre-established brain morphology norms, the degree of deviation of the morphological measurement values of each region of the target individual's brain at each time point from the brain morphology norms is determined, and the individual brain morphology deviation vector corresponding to the target individual at each time point is obtained. The correlation between the individual brain morphology deviation vectors of the target individuals at each time point and the statistical data of various mental diseases in the whole brain obtained in advance is calculated to obtain the cross-disease brain morphology similarity features of the target individuals at each time point. The cross-disease brain morphology similarity features of the target individual are input into a pre-trained machine learning model, which outputs classification information on whether the target individual is a patient with schizophrenia and / or prediction information on clinical indicators.