Brain function connection correlation value clustering device, brain function connection correlation value clustering system, brain function connection correlation value clustering method, brain function connection correlation value classifier program, and brain activity marker classification system
The system addresses the challenge of developing practical biomarkers for neurological disorders by using machine learning to generate classifiers for brain activity data, correcting measurement biases, and stratifying subjects for effective treatment selection and screening.
Patent Information
- Application Number
- JP2022536448
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-17
- Filing Date
- 2021-07-15
- Publication Date
- 2025-09-10
- Estimated Expiration
- 2041-07-15
AI Technical Summary
Current methods for diagnosing neurological and psychiatric disorders using brain imaging data face challenges in developing practical genetic biomarkers, leading to difficulties in assessing drug effectiveness and treatment development, particularly for depression, due to issues like overfitting and variability in measurement data across sites.
A treatment selection support system using machine learning to generate discriminators and classifiers based on brain functional connectivity correlation values, employing unsupervised learning and multiple co-clustering to stratify subjects into clusters, correcting for measurement bias, and integrating classifiers for improved treatment selection and screening support.
The system assists in selecting appropriate treatments for depression by providing information on therapeutic responsiveness, supporting clinical trial screening, and enhancing the accuracy of treatment recommendations based on brain activity measurements.
Smart Images

Figure 0007737099000021 
Figure 0007737099000022 
Figure 0007737099000023
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technology for clustering patterns of brain function connectivity correlation values measured by functional brain imaging in multiple devices, and more specifically to a brain function connectivity correlation value clustering device, a brain function connectivity correlation value clustering system, a brain function connectivity correlation value clustering method, a brain function connectivity correlation value classifier program, and a brain activity marker classification system. [Background technology]
[0002] (Data-driven clustering method) Recent advances in artificial intelligence technology, particularly data-driven artificial intelligence technology, have led to the realization of applications in fields such as speech recognition, translation, and image recognition that rival human capabilities in some areas, or even surpass them in some areas (for example, Patent Document 1).
[0003] In the field of medical technology, machine learning such as deep learning has been increasingly used in image diagnosis, etc. Deep learning is machine learning that uses a multi-layer neural network, and in the field of image recognition, it is known that a learning method that uses a convolutional neural network (hereinafter referred to as CNN), which is one type of deep learning, exhibits significantly higher performance than conventional methods (for example, Patent Document 2). For example, in the field of endoscopic imaging diagnosis of colon cancer, diagnostic equipment has been put into practical use that can provide diagnostic accuracy that exceeds that of humans (Non-Patent Document 1).
[0004] However, in terms of machine learning classification, most of these artificial intelligence technologies fall into the category of so-called "supervised learning," in which a large number of pairs of correct answer data and input data (e.g., image data) are prepared and used as input for the artificial intelligence to perform learning processes.
[0005] On the other hand, one application of data-driven AI is to perform the task of classifying given data into several clusters based on its features. In this case, methods such as "unsupervised learning," in which no correct answer data exists, and "semi-supervised learning," which combines learning using a small amount of "training data with correct answer labels" with learning using a large amount of "training data without correct answer labels" are known (for example, Patent Document 3).
[0006] For example, Patent Document 3 states that "semi-supervised learning is a learning method that performs learning based on a relatively small amount of labeled data and unlabeled data, and includes, for example, a bootstrap method that generates a learning model that performs classification using labeled data (teaching data T including state data S and judgment data L) and performs additional learning on the learning model using the learning model and unlabeled data (state data S) to improve the accuracy of learning, and a graph-based algorithm that generates a learning model as a classifier by grouping labeled data and unlabeled data based on their data distribution." However, as shown in this example, in "semi-supervised learning," the teaching data exists as a small amount of training data, and it is assumed that a classifier is first generated using this, and then the classifier itself is retrained using a large amount of "correct answer unlabeled training data."
[0007] (biomarker) Below, we will use the medical field as an example of a field in which discrimination and clustering using artificial intelligence technology can be applied. Indicators that quantify and digitize biological information in order to quantitatively understand biological changes within the body are called "biomarkers."
[0008] The U.S. Food and Drug Administration (FDA) defines biomarkers as "items that can be objectively measured and evaluated as indicators of normal or pathological processes or pharmacological responses to treatment." Biomarkers that characterize disease state, changes, or the degree of cure are also used as surrogate markers to confirm the effectiveness of new drugs in clinical trials. Blood glucose and cholesterol levels are typical biomarkers used to indicate lifestyle-related diseases. These include not only biological substances contained in urine and blood, but also electrocardiograms, blood pressure, PET images, bone density, and pulmonary function. Advances in genome and proteome analysis have led to the discovery of various biomarkers related to DNA, RNA, and biological proteins.
[0009] Biomarkers are expected to be used not only to measure the effectiveness of treatment after a disease has been contracted, but also as everyday indicators for preventing disease before it occurs, and also for personalized medicine, which selects effective treatments that avoid side effects.
[0010] For example, Patent Document 4 discloses a biomarker for determining the likelihood of lung disease using genetic information. Patent Document 4 defines a "biomarker" or "marker" as "a biological molecule that can be objectively measured as a characteristic of the physiological state of a biological system." Patent Document 4 also states, "Typically, biomarker measurements are information relating to the quantitative measurement of expression products, typically proteins or polypeptides. The present invention contemplates determining biomarker measurements at the RNA (pre-translational) level or the protein level (which may include post-translational modifications)." Patent Document 4 also lists examples of classifiers used as "classification systems" for such biomarker measurements, such as decision trees, Bayesian classifiers, Bayesian belief networks, k-nearest neighbor methods, case-based reasoning, and support vector machines.
[0011] On the other hand, in the case of neurological and psychiatric disorders, current diagnosis is based on symptoms, as is the case with the DSM-5 (Diagnostic and Statistical Manual of Mental Disorders, 5th Edition), and although research is being conducted into molecular markers that can be used as objective indicators from a biochemical or molecular genetic perspective, these are still in the investigation stage.
[0012] However, there have been reports of a disease diagnosis system that uses NIRS (Near-infrared Spectroscopy) technology to classify mental illnesses such as schizophrenia and depression based on features from hemoglobin signals measured by biophotometric measurement (Patent Document 5).
[0013] (Brain activity-based biomarkers) On the other hand, in the field of diagnostic imaging, there are also so-called "imaging biomarkers," which differ from the concept of biomarkers as "biological molecules" mentioned above. For example, there are attempts to use PET (positron emission tomography) for molecular imaging in the neurological field to analyze neurotransmitter and receptor functions.
[0014] Furthermore, magnetic resonance imaging (MRI) can visualize brain activity in response to external stimuli, utilizing the fact that changes in detected signals occur in response to changes in blood flow. This type of magnetic resonance imaging is specifically called functional MRI (fMRI). For fMRI, the equipment used is a regular MRI device equipped with the necessary hardware and software for fMRI measurements.
[0015] The reason why changes in blood flow cause changes in NMR signal intensity is because oxygenated and deoxygenated hemoglobin in blood have different magnetic properties. Oxygenated hemoglobin is diamagnetic and does not affect the relaxation time of the hydrogen atoms in the surrounding water, whereas deoxygenated hemoglobin is paramagnetic and alters the surrounding magnetic field. Therefore, when the brain is stimulated, local blood flow increases, and deoxygenated hemoglobin changes, the change can be detected as an MRI signal. Such stimuli are typically provided to the subject through visual or auditory stimulation, or by performing a specific task.
[0016] In brain function research, brain activity is measured by measuring the increase in nuclear magnetic resonance signals (MRI signals) of hydrogen atoms, which corresponds to the phenomenon of a decrease in the concentration of deoxygenated hemoglobin in red blood cells in microvenules and capillaries (the BOLD effect).
[0017] In this way, the blood oxygen level dependent signal that reflects brain activity measured by an fMRI device is called the BOLD signal (Blood Oxygen Level Dependent Signal). In particular, in research into human motor function, subjects are asked to perform some kind of exercise while brain activity is measured using the above-mentioned fMRI measurement.
[0018] In humans, non-invasive measurement of brain activity is necessary, and in this case, decoding technology that can extract more detailed information from fMRI data has been developed. In particular, fMRI analyzes brain activity at the voxel level (volumetric pixel) in the brain, making it possible to estimate stimulus input and cognitive state from the spatial pattern of brain activity.
[0019] Furthermore, as an advanced version of this decoding technology, Patent Literature 6 discloses a brain activity analysis method for realizing "diagnostic biomarkers" for neurological and psychiatric disorders using functional brain imaging. In this method, a correlation matrix (brain functional connectivity parameters) of activity between specified brain regions is derived for each subject from resting-state functional connectivity MRI data measured in a healthy control group and a patient group. Feature extraction is performed using regularized canonical correlation analysis on the correlation matrix and subject attributes, including the subject's disease / health label. Based on the results of the regularized canonical correlation analysis, a classifier functioning as a biomarker is generated using discriminant analysis using sparse logistic regression (SLR). It has been shown that this machine learning technology can predict the diagnosis of neurological disorders based on the connectivity between brain regions derived from resting-state fMRI data. Furthermore, verification of the prediction performance has demonstrated that it is possible to generalize to a certain extent not only to brain activity measured at a single facility, but also to brain activity measured at other facilities. Furthermore, technological improvements have been made to further improve the generalization performance of such "diagnostic biomarkers" (Patent Document 7).
[0020] Recently, it has been recognized that obtaining and sharing large-scale brain imaging data, such as the Human Connectome Project in the United States, is important for bridging the gap between basic neuroscience research and clinical applications such as the diagnosis and treatment of psychiatric disorders (Non-Patent Document 2).
[0021] In 2013, the Japan Agency for Medical Research and Development (AMED), a national research and development agency in Japan, organized the Decoded Neurofeedback (DecNef) project, in which eight research institutes collected multisite resting-state functional magnetic resonance (resting-state fMRI) data covering 2,239 samples and five disorders and publicly shared them through the Strategic Research Program for Brain Sciences (SRPBS) multisite, multidisease database (https: / / bicr-resource.atr.jp / decnefpro / ). This project has identified resting-state functional connectivity (resting-state fMRI)-based biomarkers for several psychiatric disorders that can be generalized to completely independent cohorts.
[0022] In this way, certain progress is being made in diagnosing healthy and diseased groups. It is known that within diseased groups, for example, patients generally diagnosed with "depression" are actually divided into multiple subtypes. For example, while there are groups of patients who achieve remission with the administration of standard "antidepressants," there are also "treatment-resistant" groups of patients who do not achieve remission easily.
[0023] There have been attempts to classify such "depression" patients by applying clustering using data-driven artificial intelligence to the "brain function connectivity parameters" mentioned above, and some literature has shown that certain trends exist (Non-Patent Documents 3 and 4).
[0024] However, to put such a method of classifying subtypes of a group of diseases into practical use, large-scale data on the group of diseases is required, but collecting large-scale brain imaging data is not easy, even for healthy individuals, and especially for patients.
[0025] Therefore, when measurements are carried out at multiple sites to collect large-scale data, differences in the measurement data at each measurement site become a problem. Non-Patent Document 4 also mentions that the "generalization" of clustering for large amounts of measurement data from multiple facilities is a future challenge.
[0026] For example, in the aforementioned Non-Patent Document 3, it was pointed out that depression patients were stratified into four subtypes and that there were differences in the therapeutic response to TMS (transcranial magnetic stimulation). However, another document pointed out that in the process of discovering the brain function connectivity index, depression symptom data was used twice, which resulted in overfitting, and therefore the statistical significance of the association with depression symptoms could not be confirmed, and the stability of the stratification was poor (Non-Patent Document 5). Therefore, for example, with regard to depression, the accuracy of stratification using independent validation data has not yet been confirmed.
[0027] On the other hand, for example, in order to evaluate differences in measurement data between sites when MRI measurements are performed at multiple measurement sites, attempts have been made to investigate the effect of measurement bias on functional connectivity at rest by employing so-called "traveling subjects," in which a large number of participants travel to multiple sites and undergo measurements (Non-Patent Documents 6 and 7).
[0028] In any case, when classifying subject attributes from fMRI data, in machine learning, to avoid the problem of overfitting, classifiers are often evaluated using leave-one-subject-out cross validation, in which one subject is excluded for validation, or 10-fold cross validation, in which the data is divided into 10 parts, 9 parts are trained, and the remaining 1 part is validated. However, in recent years, the risk of prediction inflation when applying machine learning to a small number of samples obtained from a single institution has become recognized, even in the field of psychiatry.
[0029] When machine learning is performed on a small amount of data, there is a high possibility of overfitting to specific trends or noise present in the training data, such as the fMRI equipment, measurement methods, experimenters, and participant groups at a particular facility.
[0030] For example, a classifier that identifies autism spectrum disorder from anatomical brain images has been reported to show high performance with a sensitivity and specificity of over 90% for the British training data used in its development, but only achieves 50% for Japanese data. This suggests that classifiers that have not been validated in an independent validation cohort consisting of subjects from facilities completely different from those used for the training data are of little scientific or practical value. The applicant of the present application has also reported on a "harmonization method" for compensating for the inter-site differences between measurement sites as described above (Non-Patent Document 8). [Prior art documents] [Patent documents]
[0031] [Patent Document 1] Republished Patent Publication No. 2018 / 147193 (International Publication WO2018 / 147193) [Patent Document 2] Japanese Patent Application Publication No. 2019-198376 [Patent Document 3] Japanese Patent Publication No. 2020-024139 [Patent Document 4] Patent Publication No. 2019-516950 (International Publication WO2017 / 162773) [Patent Document 5] Republished Patent Publication No. 2005 / 025421 (International Publication WO2005 / 025421) [Patent Document 6] Japanese Patent Application Laid-Open No. 2015-62817 [Patent Document 7] Japanese Patent Application Publication No. 2017-196523 [Non-patent literature]
[0032] [Non-Patent Document 1] Japan Agency for Medical Research and Development (AMED) press release dated December 10, 2018: "AI-enabled endoscopic diagnostic support program approved - To be used to assist doctors in diagnosis" https: / / www.amed.go.jp / news / release_20181210.html [Non-patent document 2] Glasser MF, et al. The Human Connectome Project's neuroimaging approach. Nat Neurosci 19, 1175-1187 (2016). [Non-patent document 3] Andrew T Drysdale, Logan Grosenick, Jonathan Downar, Katharine Dunlop, Farrokh Mansouri, Yue Meng1, Robert N Fetcho, Benjamin Zebley, Desmond J Oathes, Amit Etkin, Alan F Schatzberg, Keith Sudheimer, Jennifer Keller, Helen S Mayberg, Faith M Gunning, George S Alexopoulos, Michael D Fox, Alvaro Pascual-Leone, Henning U Voss, BJ Casey, Marc J Dubin & Conor Liston, “Resting-state connectivity biomarkers define neurophysiological subtypes of depression”, natural medicine, VOLUME 23, NUMBER 1, JANUARY 2017 [Non-patent document 4] Tomoki Tokuda, Junichiro Yoshimoto,, Yu Shimizu, Go Okada, Masahiro Takamura, Yasumasa Okamoto, Shigeto Yamawaki, Kenji Doya, “Identification of depression subtypes and relevant brain regions using a data-driven approach”, SCIENTIFIC REPORTS | (2018) 8:14082 | DOI:10.1038 / s41598-018-32521-z
Direct Environment 5
Outdoor Configuration6
Direct Environment 7
Outdoor Track 8
[0033] As described above, when considering the application of analysis of brain activity using functional brain imaging methods such as functional magnetic resonance imaging to the treatment of neurological and psychiatric disorders, for example, analysis of brain activity using functional brain imaging is expected to be applied as a non-invasive functional marker as a biomarker as mentioned above, in the development of diagnostic methods, and in the search and identification of target molecules for drug discovery to realize fundamental treatment.
[0034] For example, to date, practical genetic biomarkers for psychiatric disorders have not been developed, making it difficult to assess the effectiveness of drugs and therefore difficult to develop therapeutic drugs.
[0035] The present invention has been made to solve the above-mentioned problems, and its purpose is to provide a treatment selection support system, a treatment selection support device, a treatment selection support method, and a treatment selection support program that use machine learning to generate discriminators (identifiers) as diagnostic markers and classifiers as stratification markers based on measurement data of brain activity, and use these as biomarkers to provide information on selecting a treatment for a subject exhibiting depressive symptoms based on the measurement results of the subject's brain activity.
[0036] Another object of the present invention is to provide a screening support system, a screening support device, a screening support method, and a screening support program for supporting screening of subjects based on the results of measuring the subjects' brain activity in clinical trials of candidate therapeutic substances for depressive symptoms. [Means for solving the problem]
[0037] According to one aspect of the present invention, an embodiment of the present invention relates to a treatment selection support system for providing information regarding the selection of a treatment for a first subject exhibiting depressive symptoms based on measurement results of the brain activity of the first subject. The treatment selection support system includes a clustering device for performing stratification into a plurality of clusters by clustering measurement results of brain functional connectivity correlation values obtained from a plurality of second subjects, the plurality of second subjects including a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression. The clustering device includes a calculation device and a storage device for performing the clustering process on the plurality of second subjects. The computing device, in the process of generating a clustering classifier, i) stores in the storage device, feature quantities based on multiple brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined pair of brain areas, for each of the plurality of second subjects, ii) performs machine learning under supervision to generate a classifier model for discriminating the presence or absence of the diagnostic label based on the feature quantities stored in the storage device, iii) in the machine learning to generate the classifier model, selects feature quantities for clustering according to the importance of the feature quantities used in generating the classifier by machine learning, and iv) generates a clustering classifier by clustering the first group using a multiple co-clustering method under unsupervised learning based on the selected feature quantities for clustering. The treatment selection support system further includes a database device for storing clusters resulting from stratification by the clustering classifier in association with corresponding predetermined treatment information, and a support information providing device that receives as input brain activity measurement results of the first subjects and outputs corresponding treatment information according to the classification result of the clustering classifier for the measurement results.
[0038] Preferably, in the machine learning for generating the classifier model, the computing device performs undersampling and subsampling from the first group and the second group to generate a plurality of training subsamples, selects, for each of the training subsamples, features for clustering from a union of features used in generating the classifier by machine learning according to the importance of the features belonging to the union, and generates the clustering classifier by the multiple co-clustering method based on the selected features for clustering.
[0039] Preferably, the support information providing device comprises a clustering calculation device and an interface device, the clustering calculation device calculates the probability that the first subject belongs to each of the clusters using the clustering classifier, reads out at least two pieces of treatment information selected according to the probability from the database device, and the interface device outputs data for displaying the selected clusters in association with the corresponding treatment information. Preferably, the treatment information is information indicating responsiveness to a specific therapeutic agent. Preferably, the treatment information is information indicating responsiveness to a particular physical treatment.
[0040] Preferably, the process of generating a classifier by machine learning is ensemble learning in which a plurality of classifier sub-models are generated for each of the plurality of training sub-samples, and the plurality of classifier sub-models are integrated to generate the classifier model.
[0041] Preferably, the clustering device receives information representing the temporal correlation of brain activity between a predetermined plurality of pairs of brain areas of each of the plurality of second subjects from a plurality of brain activity measuring devices respectively installed at a plurality of measurement sites.
[0042] Preferably, the arithmetic unit includes a harmonization calculation means for correcting the plurality of brain function connectivity correlation values for each of the plurality of second subjects so as to remove measurement bias at the measurement site, and storing the corrected adjusted value as the feature in the storage device. Preferably, the predetermined therapeutic information is information relating to therapeutic response to a selective serotonin reuptake inhibitor.
[0043] According to one aspect of the present invention, an embodiment of the present invention relates to a treatment selection support device for providing information regarding the selection of a treatment for a first subject exhibiting depressive symptoms based on brain activity measurement results of the first subject. The treatment selection support device includes a database device for storing clusters of stratification results for subjects with a depression diagnostic label among multiple second subjects in association with corresponding predetermined treatment information, and a support information providing device for receiving the brain activity measurement results of the first subjects as input and outputting corresponding treatment information according to the stratification results based on the measurement results. The multiple second subjects include a first group with a depression diagnostic label and a second group without the depression diagnostic label. The clusters of the stratification results are obtained by a clustering classifier obtained by clustering the measurement results of brain function connectivity correlation values using a clustering device. The clustering device includes an arithmetic unit and a storage device for executing the clustering process for the first group. In the process of generating the clustering classifier, the computing device i) stores in the storage device, for each of the plurality of second subjects, features based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) performs machine learning using supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the features stored in the storage device; iii) in the machine learning to generate the classifier model, selects features for clustering according to the importance of the features used in generating the classifier by machine learning; and iv) generates the clustering classifier by clustering the first group using a multiple co-clustering method using unsupervised learning based on the selected features for clustering.
[0044] Preferably, in the machine learning to generate the classifier model, the computing device generates a plurality of training subsamples by undersampling and subsampling from the first control group and the second control group, selects features for clustering for each of the training subsamples from a union of features used in generating the classifier by machine learning according to the importance of the features belonging to the union, and generates the clustering classifier by the multiple co-clustering method based on the selected features for clustering.
[0045] Preferably, the support information providing device includes a clustering calculation device and an interface device. The clustering calculation device calculates the probability that the first subject belongs to each of the clusters using the clustering classifier, and reads at least two pieces of treatment information selected according to the probabilities from the database device. The interface device outputs data for displaying the selected clusters in association with the corresponding treatment information.
[0046] Preferably, the treatment information is information indicating responsiveness to a specific therapeutic agent. Preferably, the treatment information is information indicating responsiveness to a particular physical treatment. Preferably, the process of generating a classifier by machine learning is ensemble learning in which a plurality of classifier sub-models are generated for each of the plurality of training sub-samples, and the plurality of classifier sub-models are integrated to generate the classifier model.
[0047] Preferably, the clustering device receives information representing the temporal correlation of brain activity between a predetermined pair of brain areas of each of the subjects from a plurality of brain activity measuring devices provided at a plurality of measurement sites, and the arithmetic device includes harmonization calculation means for correcting the plurality of brain function connectivity correlation values for each of the subjects to remove measurement bias of the measurement sites, and storing the corrected adjusted values in the storage device as the feature quantities. Preferably, the predetermined therapeutic information is information relating to therapeutic response to a selective serotonin reuptake inhibitor.
[0048] According to one aspect of the present invention, an embodiment of the present invention relates to a treatment selection support method for providing information regarding selection of a treatment for a first subject exhibiting depressive symptoms based on measurement results of the brain activity of the first subject. The treatment selection support method includes a step of generating and preparing a clustering classifier for performing stratification into a plurality of clusters by clustering measurement results of brain functional connectivity correlation values obtained from a plurality of second subjects, the plurality of second subjects including a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression. The preparing step includes a calculation step of performing the clustering process for the plurality of second subjects. The calculation step includes: (i) acquiring, for each of the plurality of second subjects, features based on a plurality of brain function connectivity correlation values each representing a temporal correlation of brain activity between a plurality of predetermined pairs of brain areas; (ii) performing machine learning using supervised learning to generate a classifier model for discriminating the presence or absence of the diagnostic label based on the acquired features; (iii) selecting features for clustering in the machine learning to generate the classifier model according to the importance of the features used in generating the classifier by machine learning; and (iv) clustering the first group by a multiple co-clustering method using unsupervised learning based on the selected features for clustering to generate the clustering classifier. The treatment selection support method further includes a support information provision step of acquiring and outputting corresponding treatment information from a database that stores, in association with a cluster of the stratification result obtained by the clustering classifier for the brain activity measurement results of the first subject, predetermined treatment information corresponding to the clusters obtained by the clustering classifier. Preferably, the predetermined therapeutic information is information relating to therapeutic response to a selective serotonin reuptake inhibitor.
[0049] According to one aspect of the present invention, an embodiment of the present invention relates to a treatment selection support method for providing information regarding the selection of a treatment for a first subject exhibiting depressive symptoms based on brain activity measurement results of the first subject. The treatment selection support method includes a support information providing step of acquiring and outputting corresponding treatment information from a database that associates and stores stratification results for subjects with a depression diagnostic label among multiple second subjects, in accordance with clusters of stratification results based on the brain activity measurement results of the first subject. The multiple second subjects include a first group with a depression diagnostic label and a second group without the depression diagnostic label. The clusters of the stratification results are obtained by a clustering classifier obtained by a clustering process on the measurement results of brain function connectivity correlation values. The clustering classifier is generated by a calculation step for executing the clustering process for the multiple second subjects. The calculation step includes: i) acquiring features for each of the second subjects based on multiple brain function connectivity correlation values, each representing the temporal correlation of brain activity between a predetermined number of pairs of brain areas; ii) performing supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the acquired features; iii) selecting features for clustering in the machine learning to generate the classifier model based on the importance of the features used in generating the classifier by machine learning; and iv) clustering the first group using an unsupervised multiple co-clustering method based on the selected features for clustering to generate the clustering classifier. Preferably, the predetermined therapeutic information is information relating to therapeutic response to a selective serotonin reuptake inhibitor.
[0050] According to one aspect of the present invention, an embodiment of the present invention relates to a treatment selection support program for providing information on selecting a treatment for a first subject exhibiting depressive symptoms based on brain activity measurement results of the first subject. When executed by a computer, the treatment selection support program causes the computer to execute the following steps: generating a clustering classifier for performing stratification into multiple clusters using a clustering process on measurement results of brain functional connectivity correlation values obtained from multiple second subjects; and receiving the brain activity measurement results of the first subjects as input and, in accordance with the classification results of the clustering classifier on the measurement results, acquiring and outputting corresponding treatment information from a database device that associates and stores clusters resulting from the stratification by the clustering classifier with corresponding predetermined treatment information. The multiple second subjects include a first group with a diagnostic label of depression and a second group without the diagnostic label of depression. The clustering process includes a calculation step for performing the clustering process on the multiple second subjects. The calculation step includes: i) storing, in a storage device, feature quantities based on multiple brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined pair of brain areas, for each of the multiple second subjects; ii) performing supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the feature quantities stored in the storage device; iii) selecting, in the machine learning to generate the classifier model, feature quantities for clustering according to the importance of the feature quantities used in generating the classifier by machine learning; and iv) clustering the first group using a multiple co-clustering method in unsupervised learning based on the selected feature quantities for clustering to generate a clustering classifier. Preferably, the predetermined therapeutic information is information relating to therapeutic response to a selective serotonin reuptake inhibitor.
[0051] According to one aspect of the present invention, an embodiment of the present invention relates to a treatment selection support program for providing information regarding the selection of a treatment for a first subject exhibiting depressive symptoms based on brain activity measurement results of the first subject. When executed by a computer, the treatment selection support program causes the computer to execute a support information providing step of acquiring and outputting corresponding treatment information from a database that associates and stores stratification results for subjects with a depression diagnostic label among multiple second subjects, in accordance with clusters of the stratification results based on the brain activity measurement results of the first subject. The clusters of the stratification results are obtained by a clustering classifier obtained by a clustering process on the measurement results of brain function connectivity correlation values. The multiple second subjects include a first group with a depression diagnostic label and a second group without the depression diagnostic label. The clustering classifier is generated by a calculation step for executing the clustering process on the multiple second subjects. The calculation step includes: i) a step of acquiring features for each of the plurality of second subjects based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) a step of performing machine learning using supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the acquired features; iii) a step of selecting features for clustering in the machine learning to generate the classifier model according to the importance of the features used in generating the classifier by machine learning; and iv) a step of clustering the first group using a multiple co-clustering method using unsupervised learning based on the selected features for clustering to generate the clustering classifier. Preferably, the predetermined therapeutic information is information relating to therapeutic response to a selective serotonin reuptake inhibitor.
[0052] According to one aspect of the present invention, an embodiment of the present invention relates to a screening support system for supporting screening of first subjects in a clinical trial of candidate treatments for depression based on measurement results of the first subjects' brain activity. The screening support system includes a clustering device for performing stratification into multiple clusters using a clustering process on measurement results of brain functional connectivity correlation values obtained from multiple second subjects. The multiple second subjects include a first group with a diagnostic label of depression and a second group without the diagnostic label of depression. The clustering device includes a calculation device and a storage device for performing the clustering process on the multiple second subjects. The computing device i) stores in the storage device, for each of the plurality of second subjects, feature quantities based on a plurality of brain function connectivity correlation values each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas, ii) performs machine learning under supervision to generate a classifier model for discriminating the presence or absence of the diagnostic label based on the feature quantities stored in the storage device, iii) in the machine learning to generate the classifier model, selects feature quantities for clustering according to the importance of the feature quantities used in generating the classifier by machine learning, and iv) generates a clustering classifier by clustering the first group using a multiple co-clustering method under unsupervised learning based on the selected feature quantities for clustering. The screening support system further includes a support information providing device that receives as input measurement results of the brain activity of the first subject, records a classification result by the clustering classifier for the measurement results in association with the first subject, and outputs information to support screening of the first subject based on the classification result.
[0053] Preferably, in the machine learning for generating the classifier model, the computing device performs undersampling and subsampling from the first group and the second group to generate a plurality of training subsamples, selects, for each of the training subsamples, features for clustering from a union of features used in generating the classifier by machine learning according to the importance of the features belonging to the union, and generates the clustering classifier by the multiple co-clustering method based on the selected features for clustering. Preferably, the potential treatment is a treatment using a selective serotonin reuptake inhibitor.
[0054] According to one aspect of the present invention, an embodiment of the present invention relates to a screening support device for supporting screening of a first subject based on brain activity measurement results of the first subject in a clinical trial of candidate treatments for depression. The screening support device includes a support information providing device having a storage device for storing information identifying a clustering classifier, receiving as input the brain activity measurement results of the first subject, recording classification results based on the clustering classifier in association with the first subject, and outputting information to support screening of the first subject based on the classification results. The clusters resulting from the stratification are obtained by the clustering classifier obtained by clustering processing of the measurement results of brain function connectivity correlation values measured by the clustering device. The clustering device includes an arithmetic unit and a storage device for executing the clustering processing for a plurality of second subjects, including a first group with a diagnostic label of depression and a second group without the diagnostic label of depression. In the process of generating the clustering classifier, the computing device i) stores in the storage device, for each of the plurality of second subjects, features based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) performs supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the features stored in the storage device; iii) in the machine learning to generate the classifier model, selects features for clustering according to the importance of the features used in generating the classifier by machine learning; and iv) generates the clustering classifier by clustering the first group using a multiple co-clustering method of unsupervised learning based on the selected features for clustering. Preferably, the potential treatment is a treatment using a selective serotonin reuptake inhibitor.
[0055] According to one aspect of the present invention, an embodiment of the present invention relates to a screening support method for supporting screening of a first subject in a clinical trial of a candidate treatment for depression based on a measurement result of the first subject's brain activity. The screening support method includes the steps of: using a computing device to classify the first subject using the measurement result of the brain activity based on a clustering classifier identified by information stored in a storage device; and recording the classification result in association with the first subject and outputting information to support screening of the first subject based on the classification result. The process for generating the clustering classifier includes a computing step of performing the clustering process on a plurality of second subjects, including a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression. The calculation step includes: i) a step of acquiring features for each of the plurality of second subjects based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) a step of performing machine learning using supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the acquired features; iii) a step of selecting features for clustering in the machine learning to generate the classifier model according to the importance of the features used in generating the classifier by machine learning; and iv) a step of clustering the first group using a multiple co-clustering method using unsupervised learning based on the selected features for clustering to generate a clustering classifier. Preferably, the potential treatment is a treatment using a selective serotonin reuptake inhibitor.
[0056] According to one aspect of the present invention, an embodiment of the present invention relates to a screening support program for supporting screening of a first subject in a clinical trial of candidate treatments for depression based on measurement results of the first subject's brain activity. When executed by a computer, the screening support program causes the computer to execute the following steps: classifying the first subject using a clustering classifier specified by information stored in a storage device, and recording the classification results in association with the first subject, and outputting information to support screening of the first subject based on the classification results. The process for generating the clustering classifier includes a calculation step for executing the clustering process for a plurality of second subjects, including a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression. The calculation step includes the steps of: i) storing, in a storage device, feature quantities based on a plurality of brain function connectivity correlation values representing temporal correlations of brain activity between a plurality of predetermined pairs of brain areas for each of the plurality of second subjects; ii) performing supervised machine learning to generate a classifier model for discriminating between the presence or absence of the diagnostic label based on the feature quantities stored in the storage device; iii) in the machine learning to generate the classifier model, selecting feature quantities for clustering according to the importance of the feature quantities used in generating the classifier by machine learning; and iv) clustering the first group using a multiple co-clustering method of unsupervised learning based on the selected feature quantities for clustering to generate a clustering classifier. Preferably, the potential treatment is a treatment using a selective serotonin reuptake inhibitor. [Effects of the Invention]
[0057] Using discriminators as diagnostic markers and classifiers as stratification markers generated by machine learning, it is possible to assist in the selection of treatment methods for subjects with depression who require treatment.
[0058] Furthermore, it is possible to support the screening of subjects participating in clinical trials of candidate therapeutic agents for depressive symptoms by using a discriminator as a diagnostic marker and a classifier as a stratification marker. [Brief explanation of the drawings]
[0059] [Figure 1] FIG. 1 is a conceptual diagram for explaining harmonization processing of data measured by MRI measurement systems installed at multiple measurement sites. [Figure 2] FIG. 1 is a conceptual diagram showing the procedure for extracting a correlation matrix showing the correlation of functional connectivity at rest for a region of interest (ROI) in a subject's brain. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the contents of "measurement parameters" and "subject attribute data." [Figure 4] FIG. 1 is a schematic diagram showing the overall configuration of an MRI apparatus 100.i (1≦i≦Ns) installed at each measurement site. [Figure 5] FIG. 2 is a hardware block diagram of a data processing unit 32. [Figure 6] FIG. 10 is a conceptual diagram illustrating a process of generating a classifier serving as a diagnostic marker from a correlation matrix and a clustering process. [Figure 7] FIG. 3 is a functional block diagram for explaining the configuration of a computer processing system 300. [Figure 8] FIG. 3 is a functional block diagram for explaining the configuration of a computer processing system 300. [Figure 9] 1 is a flowchart illustrating a machine learning procedure for generating a disease classifier through ensemble learning. [Figure 10] FIG. 1 is a diagram showing demographic characteristics of a training dataset (dataset 1). [Figure 11] FIG. 1 shows the demographic characteristics of the independent validation dataset (dataset 2). [Figure 12]FIG. 10 is a diagram showing the prediction performance (output probability distribution) of MDD for the training dataset for all imaging sites. [Figure 13] FIG. 10 is a diagram showing the prediction performance (probability distribution of the classifier output) of the MDD for the training dataset for each imaging site. [Figure 14] FIG. 10 shows the probability distribution of the output of the classifier for MDD on an independent validation dataset. [Figure 15] FIG. 10 shows the probability distribution of the output of the MDD classifier for the independent validation dataset for each imaging site. [Figure 16] 10 is a flowchart illustrating a process of selecting feature amounts and performing clustering through unsupervised learning. [Figure 17] FIG. 10 is a diagram showing the concept of selecting features when there are a plurality of features (for example, Nch features) by the "learning process involving feature selection." [Figure 18] FIG. 1 is a conceptual diagram showing feature amounts that are finally selected when generating one classifier through a learning process involving feature amount selection. [Figure 19] FIG. 10 is a conceptual diagram showing how feature values are selected when a classifier is generated by performing undersampling and subsampling processes multiple times. [Figure 20] FIG. 10 is a conceptual diagram for explaining a case where there are multiple ways of dividing into clusters depending on feature amounts. [Figure 21] FIG. 1 is a conceptual diagram for explaining the concept of clustering when multiple objects are characterized by multiple feature amounts. [Figure 22] FIG. 1 is a conceptual diagram illustrating multiple clustering and multiple co-clustering. [Figure 23] This is a conceptual diagram showing the case where probabilistic models of different types of probability distributions are assumed within one view in "multiple co-clustering." [Figure 24] 1 is a flowchart for explaining an overview of a multiple co-clustering learning method. [Figure 25]FIG. 1 illustrates a graphical representation of Bayesian inference in the multiple co-clustering learning method. [Figure 26] FIG. 1 is a diagram showing two divided datasets, dataset 1 and dataset 2. [Figure 27] FIG. 1 is a conceptual diagram illustrating the concept of performing clustering on each data set. [Figure 28] FIG. 1 is a conceptual diagram showing an example of multiple co-clustering of subject data. [Figure 29] FIG. 10 is a diagram showing the results of an actual multiple co-clustering process performed on Data Set 1 and Data Set 2. [Figure 30] 1 is a table showing the number of brain functional connections (FC) assigned to each view in Dataset 1 and Dataset 2. [Figure 31] FIG. 1 is a conceptual diagram for explaining a method for evaluating the similarity of clustering (generalization performance of stratification). [Figure 32] FIG. 1 is a conceptual diagram for explaining ARI. [Figure 33] FIG. 10 is a diagram showing the evaluation results of the similarity between clustering 1 and clustering 1′, and between clustering 2 and clustering 2′. [Figure 34] 10 is a table showing the distribution of the number of subjects assigned to each cluster in view 1 in clustering 1 and clustering 1'. [Figure 35] FIG. 1 is a conceptual diagram for explaining a method for evaluating inter-site differences using a mobile subject who undergoes measurement while moving between sites. [Figure 36] This is a conceptual diagram to explain the expression of the bth functional connectivity of subject a. [Figure 37] 10 is a flowchart illustrating a process for calculating a measurement bias for harmonization. [Figure 38] FIG. 10 is a functional block diagram showing an example of a case where data collection, estimation processing, and measurement of a subject's brain activity are processed in a distributed manner. [Figure 39]FIG. 1 is a diagram showing the functional configuration of a treatment method selection support system 1000. [Figure 40] FIG. 5 is a diagram showing an example of a treatment information database 5100. [Figure 41] 10 is a flowchart illustrating the flow of a process for supporting selection of a treatment method for a subject. [Figure 42] FIG. 1 is a diagram illustrating a general drug discovery process. [Figure 43] 10 is a flowchart illustrating the flow of a process for supporting screening of a subject. [Figure 44] FIG. 1 is a diagram illustrating the configuration of a screening support device 1000′. [Figure 45] FIG. 1 shows a view of the clustering results with multiple co-clustering for dataset 1. [Figure 46] FIG. 10 shows a view of the clustering results with multiple co-clustering for dataset 2. [Figure 47] FIG. 1 shows the clustering stability between Dataset 1 and Dataset 2. [Figure 48] FIG. 10 shows a view of the clustering results for all depression patient data (dataset 1+2). [Figure 49] This figure shows the number of all depression patients and the number of depression patients for which clinical data exists, by subtype, for view 3 generated by clustering datasets 1+2. [Figure 50] FIG. 10 shows the brain functional connectivity (view 3) used for clustering. [Figure 51] FIG. 10 is a diagram showing the relationship between the severity of depression and the rate of improvement in severity for each subtype in View 3. DETAILED DESCRIPTION OF THE INVENTION
[0060] I. Learning Phase In the following, in order to explain the "device for clustering brain function connectivity correlation values" and "method for clustering brain function connectivity correlation values" of the present invention, we will use as an example "clustering" using artificial intelligence technology on brain function connectivity image data of subjects (including patients with psychiatric disorders) measured using a measurement system consisting of multiple brain activity measurement devices.
[0061] Therefore, the configuration of a measurement system according to an embodiment of the present invention, more specifically, an MRI measurement system, will be described below with reference to the drawings. In the following embodiments, components and processing steps denoted by the same reference numerals are identical or equivalent, and their description will not be repeated unless necessary.
[0062] Furthermore, in this embodiment, the present invention will be described as a system in which brain activity between multiple brain areas is measured in time series using "brain activity measuring devices," more specifically, "MRI devices," installed in multiple facilities, and subjects with a specific disease are further classified into multiple groups (subgroups) based on the pattern of temporal correlation between these areas (referred to as "brain functional connectivity") in a manner that can be generalized to multiple facilities.
[0063] Although this is also not particularly limited, the "specific disease" will be described using "major depression" as an example. However, as will be shown in the following explanation, the present invention relates to a technology for data-driven classification of a subject's "brain function connectivity correlation value," and the subject's disease is not limited to "major depression" and may be other diseases. Furthermore, as long as the subject's attribute can be classified based on the pattern of the subject's "brain function connectivity correlation value," it does not necessarily have to be a disease and may be other attributes.
[0064] In this "MRI measurement system," multiple "MRI devices" are installed in multiple different facilities. As described below, inter-site differences in measurements between these measurement facilities (measurement sites) are independently evaluated for measurement bias caused by the measurement equipment and differences due to the subject population at the measurement site (sample bias). Then, a process is performed to correct inter-site differences by eliminating the effects of measurement bias for the measurement values at each measurement site, thereby achieving harmonization of the measurement results between measurement sites. After harmonization, the brain function connectivity values are subjected to "feature selection" using ensemble learning with diagnostic labels of specific diseases as training data, followed by unsupervised clustering to classify subject attributes (e.g., psychiatric disorder subtypes).
[0065] [Embodiment 1] FIG. 1 is a conceptual diagram for explaining clustering (stratification) processing of data measured by MRI measurement systems installed at multiple measurement sites. Referring to FIG. 1, it is assumed that MRI apparatuses 100.1 to 100.Ns are installed at measurement sites MS.1 to MS.Ns (Ns: number of sites), respectively.
[0066] Furthermore, measurement sites MS.1 to MS.Ns each measure subjects in subject groups PA.1 to PA.Ns. Subject groups PA.1 to PA.N are also referred to herein as the second subject group. Each person belonging to the second subject group is also referred to as the second subject. That is, the second subject group includes multiple second subjects. Each of subject groups PA.1 to PA.Ns includes at least two groups, for example, a patient group and a healthy subject group. In this specification, the patient group is referred to as the first group when it is necessary to distinguish it from a non-patient group of healthy subjects. Subjects belonging to the first group are subjects who have a diagnostic label for depression based on a diagnosis using a diagnostic method such as DSM-5. In this specification, the healthy subject group is referred to as the second group when it is necessary to distinguish it from the patient group. Subjects belonging to the second group are subjects who do not have a diagnostic label for depression. Furthermore, the patient group is not particularly limited, but may be, for example, a group of patients with mental illness, more specifically, a group of "patients with major depression." Note that the method for determining a "diagnostic label" is not limited to the conventional "symptom-based diagnostic method such as DSM-5" as described above, but may be, for example, a method in which the results of an analysis of brain activity measurement data are used as auxiliary information, as exemplified below. In principle, at each measurement site, measurements of each subject will be carried out using a unified measurement protocol to the extent possible given the specifications of the MRI equipment. Although not particularly limited, the measurement protocol is assumed to specify, for example, the following:
[0067] 1) Direction of head scan: For example, it is necessary to specify whether the scan will be performed in the direction from the posterior (hereinafter abbreviated as "P") to the anterior (hereinafter abbreviated as "A") of the head (hereinafter referred to as "P→A direction"), or in the opposite direction, i.e., from the anterior to the posterior (hereinafter referred to as "A→P direction"). Depending on the situation, it may be possible to specify that scans will be performed in both cases. Depending on the MRI device, the default direction may differ, or it may not be possible to set both directions arbitrarily. The direction of the scan may, for example, determine the "distortion" of the image, and conditions are set as a protocol. 2) Imaging conditions for brain structure images Using the so-called spin echo method, conditions are set to capture either a "T1 weighted image" or a "T2 weighted image," or both. 3) Imaging conditions for functional brain images Using the fMRI (functional Magnetic Resonance Imaging) method, conditions are set to capture functional brain images of subjects at rest. 4) Diffusion-weighted imaging conditions Set whether or not to take a diffusion weighted image (DWI) and the conditions for doing so. Diffusion-weighted imaging is a type of MRI imaging sequence that visualizes the diffusive movement of water molecules. While signal attenuation due to diffusion can be ignored in the pulse sequence of the commonly used spin-echo method, when a large gradient magnetic field is applied for a long period of time, the phase shift caused by the movement of each magnetization vector during that time cannot be ignored, and areas with active diffusion appear as low signals. 5) Imaging to correct EPI distortion by image processing For example, one method for correcting EPI distortion by image processing is known as the "field map method," which sets imaging conditions for correcting spatial distortion. The field map method acquires EPI images at multiple echo times and calculates the amount of EPI distortion based on these EPI images. The field map method can be applied to correct the EPI distortion in the new images. Given a set of images of the same anatomical structure at different echo times, the EPI distortion can be calculated and the image distortion can be corrected. For example, the "field map method" is disclosed in the following publicly known document. Publicly known document 1: JP 2015-112474 A In addition, the measurement protocol may include extracting the necessary sequence portions from the above conditions as appropriate, and may include other sequences and their conditions as needed.
[0068] Referring again to Figure 1, the act of recruiting subjects to be measured at each measurement site MS.1 to MS.Ns is called "sampling subjects," and the cause of inter-site differences in measurement values arising from sampling bias at each measurement site is called "sample bias." For example, in the above example, it is known that patients diagnosed with "major depression" according to conventional diagnostic criteria actually include several subtypes.
[0069] Typical subtypes include "melancholic," "atypical," "seasonal," and "postpartum" depression. Furthermore, when "moderate to severe symptoms persist despite adequate treatment with two or more antidepressants with different mechanisms of action," this is called "treatment-resistant depression," and it is estimated to account for 10-20% of depression cases. In other words, it is generally known that patients diagnosed with "major depression" are by no means homogeneous. However, methods for classifying these subtypes based on objective measurement data have not yet been put to practical use.
[0070] At each measurement site, due to various factors such as biases in the tendencies of patients visiting the hospital at that measurement site due to regional differences and diagnostic trends at that hospital, it cannot necessarily be said that the distribution of subtypes included in the patient group is uniform, even if they are simply referred to as "major depression patients." As a result, it is rather common for the distribution of subtypes to be biased in the patient group at each measurement site, which is thought to result in the "sampling bias" mentioned above.
[0071] Furthermore, even in a group of subjects called a "healthy subject group," it is common for multiple subtypes to exist within them, and in that sense, "sample bias" exists even in a "healthy subject group." Furthermore, it cannot be said that the MRI devices 100.1 to 100.Ns used at each measurement site have the same measurement characteristics.
[0072] For example, differences in measurement data between sites can occur depending on the measurement conditions of the MRI device, such as the manufacturer of the MRI device, the model number of the MRI device, the static magnetic field strength of the MRI device, the number of coils (channels) in the (transmitting) and receiving coils in the MRI device, etc. The differences between sites that arise due to such measurement conditions are called "measurement bias."
[0073] Even if MRI devices with the same model number are manufactured by the same manufacturer, it is not necessarily true that they will achieve completely identical measurement characteristics due to the unique characteristics of the devices.
[0074] Here, "multi-array coils" are generally used for the (transmitting) and receiving coils in order to improve the signal-to-noise ratio of the measured signal. The "number of receiving coils" refers to the number of "element coils" that make up the multi-array coil. The aim is to improve receiving sensitivity by increasing the sensitivity of each element coil and bundling their outputs.
[0075] Although not particularly limited, in this embodiment, a harmonization method as will be described later makes it possible to evaluate "sample bias" and "measurement bias" independently.
[0076] Referring again to FIG. 1, the measurement-related data DA100.1 to DA100.Ns from each of the measurement sites MS.1 to MS.Ns are accumulated and stored in a storage device 210 in a data center 200.
[0077] Here, the "measurement-related data" includes the "measurement parameters" at each measurement site, and the "patient group data" and "healthy subject group data" measured at each measurement site.
[0078] Furthermore, the "patient group data" and the "healthy subject group data" each include "MRI measurement data of patients" and "MRI measurement data of healthy subjects" corresponding to each subject. Such "measurement-related data" will be explained below.
[0079] Figure 2 is a conceptual diagram showing the procedure for extracting a correlation matrix that indicates the correlation of functional connectivity at rest for a region of interest (ROI) in the subject's brain.
[0080] In FIG. 1, the "MRI measurement data of patients" and the "MRI measurement data of healthy subjects" in the "patient group data" and "healthy subject group data" include at least the following data.
[0081] i) Time-series "brain function image data" for calculating correlation matrix data, and / or the correlation matrix data itself That is, these data are used as the basis for calculation of brain activity biomarkers, as will be described later, by the computer processing system 300 in FIG. 1 based on the data stored in the storage device 210.
[0082] Here, the correlation matrix data can be configured to be calculated at each measurement site based on time-series "brain function image data" and then stored in the storage device 210, and the calculation processing system 300 calculates brain activity biomarkers based on the correlation matrix data in the storage device 210.
[0083] Alternatively, it is also possible to configure the system so that time-series "brain function image data" is stored in the storage device 210, and the computing system 300 calculates correlation matrix data based on the "brain function image data" in the storage device 210, and further calculates brain activity biomarkers.
[0084] Therefore, each of the "MRI measurement data of a patient" and the "MRI measurement data of a healthy individual" will contain at least either the time-series "brain function image data" for calculating the correlation matrix data or the correlation matrix data itself.
[0085] ii) Structural imaging data and diffusion-weighted imaging data of the subject Although not particularly limited, the process of correcting EPI distortion by image processing can be configured such that the data is stored in the storage device 210 after calculation processing is performed at each measurement site.
[0086] Furthermore, although not limited to this, from the viewpoint of protecting personal information, it is possible to configure the system so that anonymization processing is performed at each measurement site before the data is stored in storage device 210. However, in cases where the entity operating computer processing system 300 is legally permitted to handle personal information, the anonymization processing may be performed in computer processing system 300.
[0087] Returning to Figure 2, as shown in Figure 2(a), the average "activity" of each region of interest is calculated from the fMRI data of the resting state measured in real time for n time periods (n: natural number), and as shown in Figure 2(b), a correlation matrix of the functional connectivity ("activity correlation value") between brain regions (regions of interest) is calculated.
[0088] (Parcellation of brain regions) Functional connectivity was calculated for each participant as the temporal correlation of resting-state functional MRI blood oxygen level-dependent (BOLD) signal between two brain regions. Here, as mentioned above, the Nr region is considered as the region of interest, so the independent off-diagonal elements in the correlation matrix are, taking symmetry into consideration, Nr×(Nr-1) / 2(pcs) This means that... The following methods are conceivable for setting the region of interest.
[0089] Method 1) “Define the region of interest based on anatomical brain areas.” Here, for example, 140 regions are adopted as regions of interest for brain activity biomarkers.
[0090] In this method, we use 137 ROIs from the Brain Sulci Atlas (BAL) and ROIs from the cerebellum (left and right) and vermis from the Automated Anatomical Labeling Atlas. The functional connectivity (FC) between these 140 ROIs is used as a feature.
[0091] The Brain Sulci Atlas (BAL) and Automated Anatomical Labeling Atlas are disclosed below. Publication 2: Perrot et al., Med Image Anal, 15(4), 2011 Publication 3: Tzourio-Mazoyer et al., Neuroimage, 15(1), 2002 Such regions of interest include, for example, the following regions: Dorsomedial prefrontal cortex (DMPFC) Ventromedial prefrontal cortex (VMPFC) Anterior cingulate cortex (ACC) Cerebellar vermis, left thalamus, right inferior parietal lobe, right caudate nucleus, Right middle occipital lobe, Right midcingulate cortex However, the brain areas used are not limited to these regions. For example, the region selected may be changed depending on the neurological or psychiatric disorder of interest.
[0092] Method 2) “Functional connectivity is defined based on brain regions in a functional brain map covering the whole brain.” Here, the brain regions of such functional brain maps are also disclosed in the following documents, and although not particularly limited, they can be configured, for example, to consist of 268 nodes (brain regions).
[0093] Publication 4: Noble S, et al. Multisite reliability of MR-based functional connectivity. Neuroimage 146,959-970 (2017).
[0094] Publication 5: Finn ES, et al. Functional connectome fingerprinting: identifying individuals using patterns of brain connectivity. Nat Neurosci 18, 1664-1671 (2015).
[0095] Published literature 6: Rosenberg MD, et al. A neuromarker of sustained attention from whole-brain functional connectivity. Nat Neurosci 19, 165-171 (2016).
[0096] Publication 7: Shen X, Tokoglu F, Papademetris X, Constable RT. Groupwise whole-brain parcellation from resting-state fMRI data for network node identification. Neuroimage 82, 403-415 (2013).
[0097] Method 3) Surface-based methods Regarding brain region parcellation, it is also possible to use Human Connectome Project (HCP) style multimodality imaging (myelin, task, and functional) to analyze data using a "surface-based method" based on a brain atlas created by converting the brain into sheets along the sulci.
[0098] For such parcellation methods, a toolbox such as that disclosed at the following site can be used (ciftify toolbox version 2.0.2). https: / / edickie.github.io / ciftify / # /
[0099] This toolbox allows for the use and analysis of data in HCP-like surface-based pipelines (e.g., even when lacking T2-weighted images required for the HCP pipeline).
[0100] In the analysis of Method 3, 379 surface-based compartments (360 cortical compartments + 19 subcortical compartments) disclosed in the following known literature are used as regions of interest (ROIs).
[0101] Publication 8: Glasser, MF, Coalson, TS, Robinson, EC, Hacker, CD, Harwell, J., Yacoub, E., et al. (2016). A multi-modal parcellation of human cerebral cortex. Nature 536(7615), 171-178. doi: 10.1038 / nature18933. Therefore, the temporal changes of the BOLD signal are extracted from these 379 regions of interest (ROIs).
[0102] Furthermore, the anatomical names of important ROIs and the names of intrinsic brain networks containing ROIs can be identified using automatic anatomical labeling (AAL) and Neurosynth (http: / / neurosynth.org / locations / ), as disclosed in the following references:
[0103] Publication 9: Tzourio-Mazoyer, N., Landeau, B., Papathanassiou, D., Crivello, F., Etard, O., Delcroix, N., et al. (2002). Automated anatomical labeling of activations in SPM using a macroscopic anatomical parcellation of the MNI MRI single-subject brain. Neuroimage 15(1), 273-289. doi: 10.1006 / nimg.2001.0978.
[0104] Method 4) Data-driven method for determining brain regions As disclosed in the literature below, this is a method for identifying networks anew from phase-aligned voxels without prior information (brain maps), and is known as "Canonical ICA" or "dictionary learning." Publication: Kamalaker Dadi, Mehdi Rahim, Alexandre Abraham, Darya Chyzhyk, Michael Milham, Bertrand Thirion, Gael Varoquaux, “Benchmarking functional connectome-based predictive models for resting-state fMRI”, Preprint submitted to NeuroImage, October 31, 2018.
[0105] In the following, we will basically use "Method 3," which defines functional connectivity based on brain regions of a surface-based brain map.
[0106] There are also several candidates for calculating correlation values to measure functional connectivity, such as the tangent method and the partial correlation method. However, in the following description, although not limited to this, it is assumed that the Pearson correlation coefficient is used. The Fisher's z-transformed Pearson correlation coefficient between each possible node pair time course of the preprocessed BOLD signal was calculated and used to construct a 379 × 379 contrastive brain functional connectivity matrix, where each element represents the strength of connectivity between two nodes. FIG. 3 is a conceptual diagram showing an example of the contents of the "measurement parameters" and "subject attribute data."
[0107] The "subject attribute data" is stored in association with the "MRI measurement data of patients" and the "MRI measurement data of healthy subjects" in the "patient group data" and "healthy subject group data" in FIG. 1, respectively.
[0108] As shown in FIG. 3(a), the information includes a site ID for identifying a measurement site, a site name, a condition ID for identifying a measurement parameter, information about the measurement device, and information about the measurement conditions. "Measurement parameters" include "information about the measurement device" and "information about the measurement conditions." "Information about the measurement device" includes the manufacturer name, model number, and number of (transmitting) and receiving coils of the MRI device used to measure the brain activity of the subject at each measurement site.
[0109] It should be noted that the "information about the measuring device" is not limited to these, and may also include, for example, indicators that represent the performance of other measuring devices, such as static magnetic field strength, magnetic field uniformity after shim adjustment, and the like.
[0110] "Information about measurement conditions" includes information such as the phase encoding direction during image reconstruction (P→A or A→P), image type (T1 weighted, T2 weighted, diffusion weighted, etc.), imaging sequence (spin echo, etc.), and whether the subject's eyes are open or closed during imaging. The "information regarding measurement conditions" is not limited to these.
[0111] As shown in Figure 3(b), the "subject attribute data" includes a pseudonymized subject ID to prevent the subject from being identified, a condition ID representing the measurement conditions under which the subject was measured, and the subject's attribute information.
[0112] The "subject attribute information" includes information such as the subject's gender, age, a label indicating whether the subject is healthy or ill, the name of the subject's illness diagnosed by a doctor, the subject's medication history, and the diagnosis history. It should be noted that the "subject attribute information" is assumed to be anonymized, for example, at the measurement site, as necessary.
[0113] For example, for age and gender, by converting the data so that there are k or more pieces of quasi-identifier (same attribute) data, it is possible to reduce the probability of an individual being identified to one-kth or less, thereby maintaining "k-anonymity," which makes identification difficult. Here, "quasi-identifiers" refer to attributes such as "age," "gender," and "place of residence" that cannot be used alone, but can be used in combination to identify an individual. In addition, medication history and diagnosis history will be processed for anonymity as necessary, such as by randomizing or shifting (relativizing) the dates.
[0114] In the following, the functional connectivity calculated as the correlation of activity over time between each brain area of each subject using the above-mentioned method in the "MRI measurement data of patients" and "MRI measurement data of healthy subjects" will be collectively referred to as "functional connectivity" (abbreviated as "FC") for each area. When it is necessary to distinguish between functional connectivity for each brain area, a subscript will be added to distinguish them, as described below.
[0115] (MRI equipment configuration) FIG. 4 is a schematic diagram showing the overall configuration of an MRI apparatus 100.i (1≦i≦Ns) installed at each measurement site. 4, the MRI apparatus 100.1 at the first measurement site is illustrated in detail as an example. The other MRI apparatuses 100.2 to 100.Ns have the same basic configuration.
[0116] As shown in Figure 4, the MRI device 100.1 includes a magnetic field application mechanism 11 that applies a controlled magnetic field to a region of interest of a subject 2 (which may be a first subject or a second subject) and irradiates RF waves, a receiving coil 20 that receives a response wave (NMR signal) from the subject 2 and outputs an analog signal, a drive unit 21 that controls the magnetic field applied to the subject 2 and controls the transmission and reception of RF waves, and a data processing unit 32 that sets the control sequence of the drive unit 21 and processes various data signals to generate an image. Here, the central axis of the cylindrical bore on which the subject 2 is placed is taken as the Z axis, and the X axis is defined as the horizontal direction perpendicular to the Z axis, and the Y axis is defined as the vertical direction.
[0117] Since the MRI device 100.1 is configured as described above, the static magnetic field applied by the magnetic field application mechanism 11 causes the nuclear spins of the nuclei constituting the subject 2 to orient in the magnetic field direction (Z axis) and to precess around this magnetic field direction at the Larmor frequency specific to the nuclei.
[0118] When an RF pulse with the same frequency as this Larmor frequency is applied, the atoms resonate, absorb energy, and become excited, resulting in the nuclear magnetic resonance (NMR) phenomenon. After this resonance, when the RF pulse irradiation is stopped, the atoms release energy and return to their original steady state during the relaxation process, emitting electromagnetic waves (NMR signals) with the same frequency as the Larmor frequency. The output NMR signal is received by the receiving coil 20 as a response wave from the subject 2, and the region of interest of the subject 2 is imaged in the data processing unit 32. The magnetic field application mechanism 11 includes a static magnetic field generating coil 12, a gradient magnetic field generating coil 14, an RF irradiation unit 16, and a bed 18 on which the subject 2 is placed in the bore.
[0119] The subject 2 lies, for example, supine on a bed 18. The subject 2 can view a screen displayed on a display 6 that is installed perpendicular to the Z axis, for example, through prism glasses 4, although this is not a particular limitation. The image on this display 6 can also provide a visual stimulus to the subject 2 as needed. The visual stimulus to the subject 2 may be configured to project an image in front of the subject 2 using a projector. When neurofeedback is performed on a subject, such visual stimulation corresponds to the presentation of feedback information.
[0120] The driving unit 21 includes a static magnetic field power supply 22, a gradient magnetic field power supply 24, a signal transmitting unit 26, a signal receiving unit 28, and a bed driving unit 30 that moves the bed 18 to an arbitrary position in the Z-axis direction.
[0121] The data processing unit 32 includes an input unit 40 that accepts various operations and information inputs from an operator (not shown), a display unit 38 that displays various images and information related to the region of interest of the subject 2 on a screen, a memory unit 36 that stores programs, control parameters, image data (structural images, etc.) and other electronic data for executing various processes, a control unit 42 that controls the operation of each functional unit, such as generating a control sequence for driving the drive unit 21, an interface unit 44 that transmits and receives various signals to and from the drive unit 21, a data collection unit 46 that collects data consisting of a group of NMR signals derived from the region of interest, an image processing unit 48 that forms an image based on this NMR signal data, and a network interface 50 for communicating with the network.
[0122] The data processing unit 32 may be a dedicated computer or may be a general-purpose computer that executes functions to operate each functional unit, and performs designated calculations, data processing, and control sequence generation based on programs installed in the storage unit 36. In the following description, the data processing unit 32 will be described as a general-purpose computer.
[0123] The static magnetic field generating coil 12 generates an induced magnetic field by passing a current supplied from a static magnetic field power supply 22 through a helical coil wound around the Z axis, thereby generating a static magnetic field in the bore in the Z-axis direction. The region of interest of the subject 2 is set in the highly uniform region of the static magnetic field formed in this bore. More specifically, the static magnetic field generating coil 12 is composed of, for example, four air-core coils, which combine to create a uniform magnetic field inside and impart orientation to the spins of predetermined atomic nuclei, more specifically hydrogen atomic nuclei, within the body of the subject 2. The gradient magnetic field generating coil 14 is made up of an X coil, a Y coil, and a Z coil (not shown), and is provided on the inner circumferential surface of the static magnetic field generating coil 12, which has a cylindrical shape. In order to improve the uniformity of the gradient magnetic field, shim coils (not shown) are provided and "shim adjustment" is performed.
[0124] These X, Y, and Z coils superimpose a gradient magnetic field on the uniform magnetic field inside the bore while switching between the X-axis, Y-axis, and Z-axis directions, respectively, to impart a strength gradient to the static magnetic field. The Z coil gradients the magnetic field strength in the Z direction during excitation to limit the resonance surface, the Y coil applies a short-term gradient immediately after applying the magnetic field in the Z direction to impart phase modulation proportional to the Y coordinate to the detection signal (phase encoding), and the X coil subsequently applies a gradient during data acquisition to impart frequency modulation proportional to the X coordinate to the detection signal (frequency encoding).
[0125] The switching of the superimposed gradient magnetic fields is realized by outputting different pulse signals to the X coil, Y coil, and Z coil from the transmitter 24 in accordance with a control sequence. This makes it possible to identify the position of the subject 2 where the NMR phenomenon occurs, and provides the position information in three-dimensional coordinates required to form an image of the subject 2.
[0126] As described above, three sets of orthogonal gradient magnetic fields are used, each of which is assigned a slice direction, a phase encoding direction, and a frequency encoding direction, and imaging can be performed from various angles by combining these. For example, in addition to transverse slices in the same direction as those generally imaged by X-ray CT scanners, it is possible to image sagittal slices and coronal slices that are orthogonal to the transverse slices, as well as oblique slices in which the direction perpendicular to the plane is not parallel to the axes of the three sets of orthogonal gradient magnetic fields.
[0127] The RF irradiator 16 irradiates a region of interest of the subject 2 with an RF (Radio Frequency) pulse based on a high frequency signal transmitted from the signal transmitter 26 in accordance with a control sequence.
[0128] In FIG. 1, the RF irradiation unit 16 is built into the magnetic field application mechanism 11, but it may be provided in the bed 18 or integrated with the receiving coil 20 to form the transmitting and receiving coil 20.
[0129] The receiving coil 20 detects a response wave (NMR signal) from the subject 2, and is placed close to the subject 2 in order to detect this NMR signal with high sensitivity.
[0130] When the electromagnetic waves of the NMR signal cut through the coil wires, a weak current is generated due to electromagnetic induction in the receiving coil 20. This weak current is amplified in the signal receiving unit 28, and further converted from an analog signal to a digital signal and sent to the data processing unit 32. As described above, a multi-array coil is used for the (transmitting) and receiving coil 20 in order to improve the signal-to-noise ratio.
[0131] That is, when a high-frequency electromagnetic field of a resonant frequency is applied to the subject 2 in a state in which a Z-axis gradient magnetic field is added to a static magnetic field via the RF irradiation unit 16, predetermined atomic nuclei (e.g., hydrogen atomic nuclei) in the area where the magnetic field strength meets the resonance condition are selectively excited and begin to resonate. The predetermined atomic nuclei in the area that meets the resonance condition (e.g., a slice of the subject 2 of a predetermined thickness) are excited, and (in a classical image) their spins rotate simultaneously. When the excitation pulse is stopped, the electromagnetic waves emitted by the rotating spins induce a signal in the receiving coil 20, and this signal is detected for a while. This signal is used to observe the tissue containing the predetermined atoms inside the body of the subject 2. Then, to determine the signal's origin, X and Y gradient magnetic fields are applied and the signal is detected.
[0132] Based on the data stored in the memory unit 36, the image processing unit 48 measures the detection signal while repeatedly applying an excitation signal, performs a first Fourier transform calculation to reduce the resonance frequency to an X coordinate, and performs a second Fourier transform to restore the Y coordinate to obtain an image, and displays the corresponding image on the display unit 38.
[0133] For example, such an MRI system can capture the above-mentioned BOLD signals in real time, and the control unit 42 can perform analytical processing, as will be described later, on the images captured in time series, thereby enabling resting-state functional connectivity MRI (rs-fcMRI).
[0134] 4, measurement data, measurement parameters, and subject attribute data from MRI apparatus 100.1 and MRI apparatuses 100.2 to 100.Ns at other measurement sites are accumulated and stored in storage device 210 via communication interface 202 in data center 200. Furthermore, computer processing system 300 is configured to access the data in storage device 210 via communication interface 204. FIG. 5 is a hardware block diagram of the data processing unit 32. As shown in FIG. As mentioned above, the hardware of the data processing unit 32 is not particularly limited, but a general-purpose computer can be used.
[0135] In FIG. 5, the computer main body 2010 of the data processing unit 32 includes, in addition to a memory drive 2020 and a disk drive 2030, a processing unit (CPU) 2040, a bus 2050 connected to the disk drive 2030 and the memory drive 2020, a ROM 2060 for storing programs such as a boot-up program, a RAM 2070 connected to the RAM 2060 for temporarily storing application program instructions and providing temporary storage space, a nonvolatile storage device 2080 for storing application programs, system programs, and data, and a communication interface 2090. The communication interface 2090 corresponds to the interface unit 44 for exchanging signals with the drive unit 21 and the like and the network interface 50 for communicating with other computers via a network (not shown). The nonvolatile storage device 2080 can be a hard disk drive (HDD), a solid state drive (SSD), or the like. The nonvolatile storage device 2080 corresponds to the storage unit 36.
[0136] The CPU 2040 executes arithmetic processing based on a program to realize each function of the data processing unit 32, for example, each function of the control unit 42, the data collection unit 46, and the image processing unit 48.
[0137] A program that causes data processing unit 32 to execute the functions of the above-described embodiment may be stored on DVD-ROM 2200 or memory medium 2210, inserted into disk drive 2030 or memory drive 2020, and further transferred to non-volatile storage device 2080. The program is loaded into RAM 2070 when executed.
[0138] The data processing unit 32 further includes a keyboard 2100 and a mouse 2110 as input devices, and a display 2120 as an output device. The keyboard 2100 and the mouse 2110 correspond to the input unit 40, and the display 2120 corresponds to the display unit .
[0139] The program for functioning as the data processing unit 32 described above does not necessarily have to include an operating system (OS) that causes the computer main body 2010 to execute functions of an information processing device, etc. The program need only include instructions for calling appropriate functions (modules) in a controlled manner and achieving desired results. How the data processing unit 32 operates is well known, and a detailed description thereof will be omitted.
[0140] The computer that executes the program may be a single computer or a plurality of computers, that is, it may perform centralized processing or distributed processing.
[0141] Furthermore, although there may be differences in the configuration of the hardware within the computing system 300, such as the use of parallelized processing units or GPGPUs (general-purpose computing on graphics processing units), the basic configuration is the same as that shown in FIG. (Generation of disease / health classifiers and clustering based on brain function connectivity) FIG. 6 is a conceptual diagram illustrating the process of generating a classifier serving as a diagnostic marker from the correlation matrix as described in FIG. 2, and the clustering process.
[0142] In machine learning processing, the generation of a classifier is performed by so-called "supervised learning" processing, and the clustering processing is performed by so-called "unsupervised learning" processing.
[0143] Furthermore, the clustering process itself is "unsupervised learning" and does not use information such as doctor's diagnoses. Therefore, the individual clusters obtained as a result are data-driven groups of patients, and when patients are divided into subtypes, they form the basis for "patient stratification" using brain functional connectivity as a feature.
[0144] As shown in Figure 6, first, resting-state fMRI image data is acquired for a group of healthy subjects and a group of patients using multiple MRI devices, and the computer processing system 300 performs "preprocessing" on the fMRI image data, as described below. Next, the computer processing system 300 performs brain area parcellation processing for each subject from the measured resting-state functional connectivity MRI data to derive a correlation matrix of activity between brain regions (regions of interest).
[0145] Next, for the off-diagonal elements of the correlation matrix, corresponding measurement biases are derived as described below, and the computing system 300 performs harmonization processing by subtracting the measurement biases from the values of the elements of the correlation matrix.
[0146] Furthermore, between the element values of the correlation matrix that has undergone the harmonization process and the disease label (label indicating disease or health) for each subject, the computational processing system 300 suppresses overlearning and generates a classifier involving feature selection as a "classifier generation process using ensemble learning" as described below, to generate a disease classifier (diagnostic marker) that can predict whether the subject is diseased or healthy.
[0147] On the other hand, during ensemble learning, the computational processing system 300 performs a feature selection process for clustering, as described below, from among the features (brain function connections) identified during the process of generating a classifier for a disease label, and then performs multiple co-clustering processing using ``unsupervised learning.'' Each process in FIG. 6 will be described in more detail below. [Overview of preprocessing, generation of disease classifiers, and clustering processing] Preprocessing and calculation of resting-state functional connectivity FC matrix For example, the first 10 seconds of the measured fMRI data are discarded to allow for T1 equilibrium.
[0148] In the pre-processing step, the computer processing system 300 performs processes such as slice timing calibration, realignment to correct for head motion artifacts, co-registration of functional brain images (EPI images) and morphological images, distortion correction, segmentation of T1-weighted structural images, normalization to Montreal Neurological Institute (MNI) space, and spatial smoothing using, for example, an isotropic Gaussian kernel with a half-width of 6 mm. Such pre-processing pipeline processing is disclosed, for example, at the following site: http: / / fmriprep.readthedocs.io / en / latest / workflows.html
[0149] (Parcellation of brain regions) Brain region parcellation can be performed using a "surface-based method" according to "Method 3" described above, but is not limited to this.
[0150] (Physiological noise regression) Physiological noise regression is performed using CompCor, which is disclosed in the following reference: Publication: Behzadi, Y., Restom, K., Liau, J., and Liu, TT (2007). A component based noise correction method (CompCor) for BOLD and perfusion based fMRI. Neuroimage 37(1), 90-101. doi: 10.1016 / j.neuroimage.2007.04.042. To remove some spurious sources, linear regression with regression parameters such as the six movement parameters, the whole brain, etc. is used.
[0151] (Temporal filtering) The computing system 300 applies a temporal bandpass filter, for example a Butterworth filter with a passband between 0.01 Hz and 0.08 Hz, to the time series data to limit the analysis to low-frequency fluctuations characteristic of BOLD activity.
[0152] (head movement) Frame-wise displacement (FD) is calculated for each functional session, and volumes with FD > 0.5 mm are removed to reduce spurious changes in functional connectivity FC due to head movement. FD represents head motion between two temporally consecutive volumes as a scalar quantity (i.e., the sum of absolute translational and rotational displacements).
[0153] For example, in the specific dataset described below, if the percentage of volume removed after scrubbing exceeded (mean ± 3 × standard deviation), the participant's data was excluded from analysis. As a result, 35 participants were excluded from the entire dataset. Therefore, data from 683 participants (545 HCs, 138 MDD patients) were used in the training dataset, and data from 444 participants (263 HCs, 181 MDD patients) were used in the independent validation dataset for the following analysis.
[0154] (Calculation of functional connectivity (FC) matrix) In this specific example, functional connectivity FC is calculated as the temporal correlation of the BOLD signal across 379 regions of interest (ROIs) for each participant after the regions have been segmented using the parcellation method described above. In calculating the functional binding, although not particularly limited, Pearson's correlation coefficient is used here.
[0155] Fisher's z -transformed Pearson's correlation coefficient between the time courses of the preprocessed BOLD signals of each possible pair of ROIs was calculated, and a symmetric connectivity matrix of 379 rows × 379 columns was constructed, with each element representing the strength of connectivity between two ROIs. Furthermore, for the analysis, 71,631 (=(379 × 378) / 2) functional connectivity FC values of the lower triangular matrix of the connectivity matrix are used. (Harmonization processing for brain activity biomarkers)
[0156] When collecting big data related to psychiatric disorders, as mentioned above, it is almost impossible for a single site to collect large-scale brain imaging data (connectomes related to human diseases), so it is necessary to collect imaging data from multiple sites.
[0157] Because it is difficult to fully control the type of MRI scanner, protocol, and patient demographics, brain imaging data acquired under heterogeneous conditions is used to analyze the collected data.
[0158] In particular, because disease factors tend to be confounded with site factors, differences between sites are the biggest obstacle when applying machine learning techniques to extract disease factors from data under such heterogeneous conditions.
[0159] Confounding occurs because a site (or hospital) is likely to sample only a few types of psychiatric disorder (e.g., primarily schizophrenia from site A, primarily autism from site B, primarily major depression from site C, etc.). To properly manage data under such heterogeneous conditions, it is necessary to harmonize data between sites. Site differences essentially involve two types of bias. These are technical bias (i.e., measurement bias) and biological bias (i.e., sample bias).
[0160] Measurement bias includes differences in MRI scanner characteristics such as imaging parameters, field strength, MRI equipment manufacturer and scanner type, while sample bias is related to differences in subject groups between sites.
[0161] Therefore, a "harmonization process" is required to compensate for such differences between sites. Details of this harmonization process are described in the above-mentioned Non-Patent Document 8 (Ayumu Yamashita et al.), and the contents of that document will be discussed later. (Disease classifier using ensemble learning)
[0162] In this specification, the term "ensemble learning" refers to a process of creating K sets of training data by extracting and restoring original training data, independently generating K classifiers for each set of training data through machine learning processing, and integrating these K classifiers to generate a discriminator.
[0163] In particular, the objective here is to determine whether a subject has a specific disease or is healthy based on the subject's brain functional connectivity pattern, so each classifier is a classifier for a two-class classification problem.
[0164] When creating K sets of training data by sampling with replacement from the original training data, "undersampling" and "subsampling" are performed, as will be described later.
[0165] Here, it is also possible to use a regularization learning method that uses a "classifier using a learning process involving feature selection" such as a classifier using L1 regularization (LASSO (Least Absolute Shrinkage and Selection Operator) method) or a ridge regularization method (L2 regularization).
[0166] Here, "regularized learning" refers to a learning method that aims to improve generalization performance by using all features of the original training data as the learning target, but imposes a penalty for increased complexity when learning a model in the learning algorithm, and seeks to find a learning model that minimizes the amount of this penalty added to the training error. L1 regularization uses the sum of the absolute values of the parameters (corresponding to the features) of the learning model as the penalty, while L2 regularization uses the sum of the squares of the parameters of the learning model as the penalty. Note that L0 regularization, which uses the number of features used in the model itself as the penalty, is also possible.
[0167] In addition, the LASSO method (L1 regularization) is a method that enables so-called sparse estimation, and its derivatives include the Elastic Net method, the Group Lasso method, the Fused Lasso method, the Adaptive Lasso method, and the Graphical Lasso method.
[0168] On the other hand, as a "classifier using a learning process involving feature selection," it is also possible to use a method such as the "random forest method," which, in generating a classifier, obtains the importance of features together with the selection of features.
[0169] Although the following description focuses on the LASSO method as a "method for generating a classifier using ensemble learning," the method is not limited to the above. For example, the parcellation method may be an ensemble learning method in which a dictionary learning method is used to set a brain region to be analyzed in a data-dependent manner, the tangent method (tangent-space covariance) is used as the value of functional brain connectivity, the brain functional connectivity FC in the dataset is corrected between facilities using the ComBat method described below, and a classifier is generated using ridge regularization. The parcellation method, the brain functional connectivity calculation method, the harmonization method, and the classifier generation method may be combined in other ways.
[0170] As will be described later, in this embodiment, in such ensemble learning, when a classifier is trained, the "importance" for achieving the function of classifying each feature is specified. (Feature selection for clustering and clustering)
[0171] In the process of generating K classifiers as "ensemble learning" for K sets of training data, a process is executed to identify a set of second features to be used for clustering by "unsupervised learning" from the union of the "first features" used in generating each classifier. In particular, but not limited to, the "importance" is determined as follows, for example.
[0172] i) In ensemble learning, if the learning method for generating K classifiers is "learning processing involving feature selection," then the features are ranked in the union of the "first features" selected in generating the classifiers according to the frequency with which they are used in generating the K classifiers.
[0173] ii) In ensemble learning, if the learning method for generating K classifiers is a method that obtains the importance of features in generating classifiers, such as the "random forest method," the ranking list of features can be configured to be generated according to such importance.
[0174] iii) In ensemble learning, if the learning method for generating K classifiers is "ridge regularization (L2 regularization)" (which does not necessarily involve feature selection) and is a classifier generation method that uses a weighted sum of features as an argument, a feature ranking list is generated using the median value, which is obtained by aggregating the absolute values of the weight coefficients of each feature in each classifier across the K classifiers, as the importance. Note that the importance is not necessarily limited to this "median" and other statistical representative values, such as the "integrated value across the K classifiers," may also be used.
[0175] It is possible to configure the system so that a predetermined number of top feature quantities in the ranking list generated by the methods i) to iii) are set as "second feature quantities."
[0176] The condition for identifying the "second feature" is not limited to a predetermined number at the top of the ranking list, but may be, for example, a predetermined frequency or more in the ranking list (using the frequency itself of being selected at a certain rate or more in generating K classifiers).
[0177] Based on the feature amounts selected in this way, clustering processing (patient stratification) is performed by the "multiple co-clustering method," which is unsupervised learning, as will be described later.
[0178] [Classifier generation process for two-class classification] Among the processes described in FIG. 6, generation of a classifier by ensemble learning will be described in more detail below. That is, we will explain as an example the process of generating a classifier for two-class classification, more specifically, the process of constructing a biomarker for MDD using a learning dataset as training data for a disease classifier (a two-class classifier for "healthy" or "disease").
[0179] Here, the process of generating a classifier will be described using major depressive disorder, a type of mental illness, as an example, that is, a group of patients diagnosed with major depressive disorder by a doctor using a conventional symptom-based diagnostic method. An example of the process executed by the disease classifier generation unit 3008 shown in Fig. 8 to generate a classifier that outputs auxiliary information for diagnosis to distinguish between a patient group and a healthy control group will be described. Therefore, in the following, we will explain the procedure for constructing an MDD classifier that distinguishes between healthy subjects (HC) and MDD patients based on functional connectivity FC.
[0180] In the following, as an example of a "classifier using a learning process involving feature selection" for creating a disease classifier (MDD classifier), a classifier learning method using L1 regularization (LASSO method) will be described.
[0181] Then, as described below, in order to identify functional connectivity FCs related to MDD diagnosis, features to be used for clustering are selected according to the "importance" of each functional connectivity FC to the construction of a disease classifier.
[0182] Figures 7 and 8 are functional block diagrams for explaining the configuration of a computing system 300 that performs harmonization processing, disease classifier generation processing, clustering classifier generation processing, and discrimination processing based on data stored in the storage device 210 of the data center 200.
[0183] Here, the "discrimination process" includes disease discrimination (discrimination between disease and health) and classification process to determine which "cluster" (subtype) the subject to be tested belongs to.
[0184] 7, computer system 300 includes a storage device 2080 for storing data from storage device 210 and data generated during calculation, and an arithmetic unit 2040 for executing arithmetic processing on the data in storage device 2080. An example of arithmetic unit 2040 is a CPU.
[0185] The arithmetic device 2040 includes a correlation matrix calculation unit 3002 that calculates correlation matrix elements for MRI measurement data 3102 of a patient group and a healthy subject group by executing a program and stores the correlation matrix data 3106 in the storage device 2080, a harmonization calculation unit 3020 that performs harmonization processing, and a learning and discrimination processing unit 3000 that performs disease classifier generation processing, clustering classifier generation processing, and discrimination processing using the generated disease classifier or clustering classifier based on the results of the harmonization processing. FIG. 8 is a functional block diagram illustrating the configuration of FIG. 7 in more detail. FIG. 9 is a flowchart illustrating a machine learning procedure for generating a disease classifier by ensemble learning.
[0186] First, the process from harmonization processing to generation of a classifier (disease classifier) by ensemble learning as shown in FIG. 6 will be described with reference to FIGS. 8 and 9. FIG.
[0187] First, it is assumed that the storage device 210 of the data center 200 has collected fMRI measurement data of subjects (healthy subjects and patients) from each measurement site, attribute data of the subjects, and measurement parameters as a “learning dataset.”
[0188] 8 and 9, we use this learning dataset as training data for a disease classifier (a two-class classifier for "healthy" or "disease") to construct a biomarker for MDD discrimination, which distinguishes between a healthy control group (a group diagnosed as healthy (HC)) and an MDD patient group (a group diagnosed as major depressive disorder) based on the functional connectivity FC values of 71,631.
[0189] As described below, in the training process for generating a classifier for MDD (hereinafter referred to as the "MDD classifier"), a logistic regression analysis (a type of sparse modeling method) with L1 regularization (LASSO method) is used to select an optimal subset of functional connection FCs from 71,631 functional connection FCs.
[0190] In general, when using L1 regularization, some parameters (weight elements in the following explanation) can be set to 0. This means that feature selection is performed, resulting in a sparse model.
[0191] However, the sparse modeling method is not limited to the LASSO method, and other methods can also be used, such as sparse logistic regression (SLR), which applies variational Bayesian methods to logistic regression, as will be described later.
[0192] Referring to FIG. 9, when the learning process for the MDD classifier is started (S100), the correlation matrix calculation unit 3002 calculates the components of the coupling matrix using a learning data set that has been prepared in advance (stored in the storage device 2080) (S102). Next, the harmonization calculation unit 3020 performs harmonization processing using the calculated measurement bias (S104). As will be described later, the harmonization process is preferably a method using a traveling subject, but other methods may also be used.
[0193] For example, it is possible to configure the discovery dataset and the independent validation dataset to perform harmonization between the datasets using a combat method as described below.
[0194] Next, the disease classifier generation unit 3008 generates an MDD classifier for the training data using a so-called "ensemble learning method," which is a modified version of the "nested cross validation" method.
[0195] First, the disease classifier generation unit 3008 divides the training data into 10 parts, for example, K=10, to perform a training process on the training data using "K-fold cross-validation" (K: natural number) (outer cross-validation).
[0196] That is, the disease classifier generation unit 3008 sets one of the K (10) divided partial datasets as a "test dataset" for verification, and sets the remaining (K-1) (9) divided data as a training dataset (S108, S110).
[0197] Next, the disease classifier generation unit 3008 performs "under-sampling processing" and "sub-sampling processing" on the training dataset (S112).
[0198] Here, "undersampling processing" refers to processing that is carried out when the number of data corresponding to each of the specific attribute data (two or more types) to be classified is not uniform in the training dataset, in order to make these numbers uniform by excluding the data of the attribute with the larger number.
[0199] Here, since the number of subjects in the MDD patient group and the number of subjects in the healthy control group are not equal in the training data set, this corresponds to carrying out processing to make them equal. Furthermore, "subsampling" refers to the process of randomly extracting a predetermined number of samples from a training data set.
[0200] That is, in the cross-validation repeated K times via steps S108 to S118 and S122, since the training data set is imbalanced in terms of the number of MDD patients and healthy control subjects in each cross-validation, an undersampling method is used to construct a classifier, and as a subsampling process, a predetermined number, for example, 130 MDD patients and the same number of 130 healthy control subjects are randomly sampled from the training data set.
[0201] The number 130 is not limited to this value, but is determined appropriately to enable the above-mentioned undersampling depending on the number of data in the learning dataset (683 people in "Dataset 1" described below), the division number K (here, for example, K=10), and the degree of imbalance in the number of data included in the specific attribute to be classified.
[0202] The reason for performing such subsampling is that undersampling has a disadvantage in that the classifier cannot learn using the excluded data, and in order to eliminate this disadvantage, the random sampling procedure (i.e., subsampling) is repeated M times (M: natural number, for example, M=10).
[0203] As will be described later, there is technical significance in performing such undersampling and subsampling processes for "feature selection" when generating a "classifier" for "stratification," and this will be discussed later.
[0204] Next, the disease classifier generation unit 3008 performs hyperparameter adjustment processing for each of the subsampled subsamples 1 to 10 (S114.1 to S114.10).
[0205] Here, for each subsample, a classifier submodel is generated using a logistic function such as: This logistic function is used to define the likelihood of a participant belonging to the MDD class within the subsample as follows:
[0206]
number
[0207] where y sub represents the participant's class label (MDD, y=1; HC, y=0), and c sub Let σ represent the FC vector for a given participant, and w represent the weight vector. The weight vector w is determined so as to minimize the following evaluation function (cost function) (LASSO calculation).
[0208]
number
[0209] For each subsample, the disease classifier generation unit 3008 uses, although not limited to, a predetermined number of data as hyperparameter adjustment data, and determines the weight vector w using the remaining data (for example, data from n=250 or 248 people). At this time, although not limited to, for example, assuming that the hyperparameter λ is 0<λ≦1.0, the disease classifier generation unit 3008 divides this interval into P equal parts (P: natural number), for example, 25 equal parts, and determines the weight vector w by the LASSO calculation as described above using λ for each value.
[0210] In this case, as described above, "nested cross-validation" is performed, and hyperparameter tuning is performed as "inner cross-validation." In inner cross-validation, the "test dataset" of outer cross-validation is not used.
[0211] Then, the disease classifier generation unit 3008 compares the discrimination performance (e.g., accuracy) of the hyperparameter adjustment data using the logistic functions corresponding to each generated λ value, and determines the logistic function corresponding to λ with the highest discrimination performance (hyperparameter adjustment process).
[0212] Next, the disease classifier generation unit 3008 sets a "classifier submodel" that outputs the average of the output values of the logistic function corresponding to each subsample generated in the current cross-validation loop (S116). This can also be considered a type of "ensemble learning" in that the classification performance is determined by the average of the output values of the classifier calculated for each subsample.
[0213] The disease classifier generation unit 3008 uses the test dataset prepared in step S110 as input and executes validation of the classifier sub-model generated in the current cross-validation loop (S118).
[0214] Note that as a method of generating subsamples by undersampling and subsampling, and performing feature selection on each subsample to generate a classifier submodel, other sparse modeling techniques may be used in addition to the above-described method of performing the LASSO method and adjusting hyperparameters.
[0215] If the disease classifier generation unit 3008 determines that the cross-validation loop has not been completed K times (here, 10 times) (N in S122), it sets another partial dataset from the K divided data that is different from the ones used in the previous loops as the test dataset, sets the remaining partial datasets as the training dataset (S108, S110), and repeats the process.
[0216] On the other hand, if the disease classifier generation unit 3008 has completed K (10) cross-validation loops (Y in S122), it generates a classifier model for MDD (MDD classifier) by outputting the average of the outputs of K×M (in this case, 10×10=100) logistic functions (classifiers) for the input data (S120).
[0217] Ultimately, the MDD classifier can be said to be a "classifier" obtained as a result of "ensemble learning" in the sense that its classification output is the average of the outputs of K × M classifiers. When the output of the MDD classifier (probability value of diagnosis) exceeds 0.5, it can be considered as an indicator of an MDD patient.
[0218] Furthermore, in this embodiment, the Matthews correlation coefficients (MCC), the area under the receiver operating characteristic curve (ROC curve) (AUC), accuracy, sensitivity, and specificity are used as evaluation indices for the performance of the MDD classifier generated in this manner.
[0219] The method of generating a classifier for a target disease (e.g., MDD) using the feature-selected values in each subsample (in this case, the elements of the correlation matrix after harmonization processing for measurement bias) is not limited to averaging the outputs of multiple classifier submodels, but may be a majority vote process, or may be configured to generate a classifier using other modeling methods, particularly other sparse modeling methods, for the feature-selected values.
[0220] (Examples of data used for MDD classifier and performance) As mentioned previously, building reliable classifiers and regression models using machine learning algorithms requires the use of large sample sizes of data collected from multiple imaging sites.
[0221] Therefore, in the following, we use a training resting-state fMRI dataset of approximately 700 participants, including MDD patients, recruited from four different imaging sites. FIG. 10 is a diagram showing the demographic characteristics of such a learning dataset (dataset 1). Dataset 1 is the data in the SRPBS mentioned above. FIG. 11 shows the demographic characteristics of the independent validation dataset (dataset 2). Dataset 2 is also basically the data in the SRPBS mentioned above. Namely, the following analysis uses two resting-state functional MRI (rs-fMRI) datasets.
[0222] (1) As shown in Figure 10, Dataset 1 contains data for 713 participants (564 healthy controls from four sites and 149 MDD patients from three sites).
[0223] (2) As shown in Figure 11, Dataset 2 contains data for 449 participants (264 healthy controls (HC) from four sites and 185 MDD patients (MDD) from four sites).
[0224] Additionally, "depressive symptoms" were assessed using the Beck Depression Inventory II (BDI) from the majority of participants in each dataset. Dataset 1 is the "training dataset" and is used to build a discriminator for MDD and a classifier for clustering. Participants underwent a single 10-minute resting-state functional MRI (rs-fMRI) session.
[0225] Here too, resting-state functional MRI (rs-fMRI) data were acquired under a unified imaging protocol (http: / / www.cns.atr.jp / rs-fmri-protocol-2 / ).
[0226] However, it is difficult to ensure that imaging was performed using the same parameters at all sites, as two phase modulation directions (P→A and A→P), two MRI equipment manufacturers (Siemens and GE), three different numbers of coils (12, 24, and 32), and three scanner model numbers were used for the measurements. During resting-state functional MRI (rs-fMRI) scanning, participants were instructed as follows: "Relax. Don't fall asleep. Keep your eyes on the central crosshair and don't think about anything in particular."
[0227] "Demographic characteristics" in a dataset are characteristics used in so-called "demographics," and include attributes in a table such as age, gender, and diagnosis. In Figures 10 and 11, the number in parentheses indicates the number of participants who have BDI score data. Demographic distributions were not statistically significantly different between MDD and HC populations across all training datasets (p>0.05). Dataset 2 is an "independent validation dataset" and is used to test the MDD classifier and the clustering classifier. The sites imaged for Dataset 2 were not included in Dataset 1.
[0228] Although the demographic distribution of age was consistent between MDD and HC populations in the independent validation dataset (p>0.05), the demographic distribution of sex ratio was not consistent between MDD and HC populations in the independent validation dataset (p<0.05).
[0229] (site effect control) In addition, the following description assumes that a traveling subject harmonization method for the training dataset, as described below, is used to control for site effects on functional connectivity FC. However, the harmonization method is not limited to this method, and other methods such as the ComBat method may also be used. The ComBat method is disclosed in, for example, the following document:
[0230] Publication: Johnson WE, Li C, Rabinovic A. “Adjusting batch effects in microarray expression data using empirical Bayes methods.” Biostatistics 8, 118-127 (2007). Using a travelling subject harmonization method makes it possible to remove genuine inter-site differences (measurement bias).
[0231] Note that, since no travelling subject datasets existed for the sites included in the independent validation dataset, the ComBat harmonization method was used to control for site effects in the independent validation dataset. FIG. 12 is a diagram showing the prediction performance (output probability distribution) of the MDD for the training dataset for all imaging sites.
[0232] For the training dataset, a threshold of 0.5 clearly separates the two diagnostic probability distributions corresponding to the MDD patient and healthy control populations into right (MDD) and left (HC) in the output from the classifier model. The classifier model separates MDD patients from the HC population with 66% accuracy. The corresponding AUC was 0.77, indicating high discriminatory power. Also, the MCC is approximately 0.33. FIG. 13 is a diagram showing the prediction performance (probability distribution of the output of the classifier) of the MDD for the training dataset for each imaging site.
[0233] From Figure 13, it can be seen that almost equally high classification accuracy is achieved for the entire dataset as well as for the individual datasets from the three imaging sites (site 1, site 2, and site 4). Note that the dataset from Site 3 (SWA) simply contains a group of healthy subjects, but its probability distribution corresponds to that of the healthy subjects from the other sites.
[0234] (Classifier generalization performance) FIG. 14 shows the probability distribution of the output of the MDD classifier on an independent validation data set. That is, an independent validation data set is used to test the generalization performance of the classifier model.
[0235] For MDD, in the process of Figure 12, 100 (10 divisions × 10 subsampling) logistic function classifiers are generated by machine learning, and an independent validation data set is input to all of the generated 100 classifiers (classifier model as a set of classifiers).
[0236] Then, the outputs of 100 classifiers for each participant were averaged (diagnostic probability), and if the averaged diagnostic probability value was >0.5, the diagnostic label for that participant was determined to be major depressive disorder. On an independent validation dataset, the generated classifier model separates the MDD population from the HC population with approximately 70% accuracy. The corresponding AUC was 0.75, indicating high discriminatory ability (permutation test p<0.01).
[0237] For the independent validation dataset, a threshold of 0.5 clearly separates the two diagnostic probability distributions corresponding to the MDD patient and healthy control populations into right (MDD) and left (HC) at the output from the classifier model. The sensitivity was 68% and the specificity was 71%, which corresponds to a high MCC value of 0.38 (permutation test p<0.01). FIG. 15 shows the probability distribution of the output of the MDD classifier for the independent validation data set for each imaging site. It can be seen that high classification accuracy is achieved for the entire dataset as well as for each individual dataset from the four imaging sites.
[0238] [Clustering process for subject data] Among the processes described with reference to FIG. 6, the selection of features for clustering and the clustering process using the selected features will be described in more detail below.
[0239] That is, the processes described as "feature selection" and "clustering process" in FIG. 6 will be described as processes executed by the disease classifier generation unit 3008 and the clustering classifier generation unit 3010 in FIG. FIG. 16 is a flowchart illustrating a process of selecting feature amounts and performing clustering through unsupervised learning.
[0240] In the following, we will explain the process of ranking the features (brain function connections) used in generating each classifier sub-model in the learning process of the "two-class classifier" as described in Figure 9, and performing clustering by unsupervised learning using a predetermined number of the top features.
[0241] As mentioned above, brain functional connectivity is highly dimensional, exceeding 70,000 dimensions depending on the brain parcellation method, and it is generally difficult to perform clustering processing using unsupervised learning using conventional methods. The method of this embodiment addresses this clustering problem by ranking features according to their importance in "classifier generation using supervised learning," and combining this with "clustering using unsupervised learning" using features selected based on this ranking, thereby enabling such clustering processing.
[0242] For convenience, the following describes the clustering process as separate from the disease classifier generation process. However, in Figure 16, steps S200 to S210 are equivalent to steps S100 to S120 in Figure 9, and the disease classifier generation process and the clustering process can be executed as a series of processes.
[0243] Referring to Figure 16, when the clustering learning process is started, the disease classifier generation unit 3008 prepares subject data of Nh healthy subjects and Nm depressed subjects (S202), and the disease classifier generation unit 3008 performs brain area parcellation processing, calculation of brain function connectivity values, and harmonization processing on the subjects' functional brain activity data (S204).
[0244] Next, the disease classifier generation unit 3008 performs data division for Ncv-division cross-validation (Ncv: natural number, and Ncv≧2) to prepare a training data set and a test data set for each divided data, and performs undersampling and subsampling for each training data set to generate Ns test data subsets (S206). Furthermore, the disease classifier generation unit 3008 performs the following for each subsample: A classifier is generated by a learning method involving feature selection (S208). It is assumed here that feature selection is performed using L1 regularization (LASSO) in the same manner as in FIG.
[0245] The processing of steps S206 to S208 is repeated for the learning dataset divided into Ncv pieces, by sequentially recombining the training dataset ((Ncv-1) of the divided datasets) with the test dataset (one of the divided datasets) until cross-validation is performed Ncv times. An integrated classifier that outputs the average of the (Ns×Ncv) classifiers thus generated is generated as a disease classifier (diagnostic marker) (S210). As described above, the processing up to this point is equivalent to steps S100 to S120 in FIG.
[0246] On the other hand, in steps S206 to S208, which are repeated Ncv times, the clustering classifier generation unit 3010 performs ranking of the features in the union of the features (brain function connections) selected when generating a classifier by a learning method involving feature selection, although this is not particularly limited, the selected number of times (S220). Here, the "number of times selected as a feature" is referred to as the importance of that feature in this ranking.
[0247] In other words, in the example shown in Figure 9, the LASSO method generates 100 (=10 x 10) classifiers, and the number of selections is counted so that brain function connections with a weight other than zero in each classifier are weighted +1. Connections are ranked as important in descending order of the number of counts.
[0248] Next, the clustering classifier generation unit 3010 selects, for example, a predetermined number of feature quantities from the union according to the importance in order to perform clustering on the depression patient group by unsupervised learning (S222).
[0249] Furthermore, the clustering classifier generation unit 3010 performs clustering processing using a multiple co-clustering method, which will be described later, as a method of unsupervised learning (S224). Through the above processing, the clustering classifier generation unit 3010 generates a clustering classifier for the depression patient group (S226).
[0250] II. Model Use Phase That is, through the above processing, the clustering classifier generation unit 3010 specifies, for each cluster, a probability distribution model that generates such observation data from the observation data, and information on each model is stored in the disease classifier data in the storage device 2080. Then, the discriminant value calculation unit 3012, as a clustering classifier, calculates, for input data other than the training data, the posterior probability that the input data belongs to each cluster based on each probability distribution model, and outputs a classification result that the input data belongs to the cluster with the maximum posterior probability (MAP estimation method (Maximum A Posteriori Probability Estimation method)).
[0251] [Additional explanation of the training method and trained model] In the above description, the clustering process has been described as being based on the classifier generation process performed on subject data of a healthy subject group of Nh people and a depression group of Nm people by the disease classifier generation unit 3008. However, the clustering method of this embodiment is not limited to this case, and can also be used to cluster disease groups other than the "depression patient group," for example, groups of patients with other mental disorders, such as a "schizophrenia patient group," a "autism patient group," or a "obsessive-compulsive disorder patient group."
[0252] Alternatively, more generally, for attributes that have been found to have a certain relationship between attribute labels that humans have empirically classified (for example, a person's personality, a person's areas of expertise, etc.) and the pattern of correlation between areas of temporal changes in brain activity, it is possible to use this information to perform data-driven "clustering of subjects belonging to that attribute" (classification into subtypes) for subjects classified by that attribute label. (undersampling and subsampling processes) In the processing described above, "undersampling and subsampling processing" is performed, so the technical significance of this will be briefly explained. First, the effect of the "undersampling process" is that the classification boundary in the classifier is appropriately set.
[0253] For example, when considering a two-class classification task, the less bias there is in the number of data belonging to each class in the training data, the more accurate the evaluation of the classifier's performance (for example, accuracy) will be in the processing flow.
[0254] In steps S114.1 to S114.10 in FIG. 9, the hyperparameter setting involves the process of "determining the logistic function corresponding to λ with the highest discrimination performance," so it is necessary to accurately evaluate the "discrimination performance."
[0255] As an extreme example, when training a classifier using training data in which the number of data items belonging to class 1 is 100 and the number of data items belonging to class 2 is 1, the accuracy may not be significantly affected even if the classifier determines that all data items are class 1. In this respect, it is significant to make the number of data items belonging to each of the two classes uniform by random sampling. Furthermore, subsampling is premised on being performed multiple times in conjunction with undersampling for the following reasons.
[0256] First, even if undersampling and subsampling processes are performed using random sampling, a single process may result in biased data.
[0257] Secondly, as will be explained below, the generation of the "classifiers" repeated in steps S108 to S122 in FIG. 9 is performed by "ensemble learning" as described above. At this time, the importance of the feature amount for classification is determined in generating each classifier.
[0258] The importance is determined according to the contribution of each feature to the classification, such as the selection of the feature in the "learning process with feature selection" or the weight of the feature in the calculated classification in the "learning process without feature selection."
[0259] Below, we will explain the significance of "undersampling and subsampling processes" in determining such importance using "learning processes with feature selection" as an example. Note that even in "learning processes without feature selection" such as L2 regularization, the phenomenon of "the weight of a feature increasing" can be considered to arise for essentially the same technical reasons as the phenomenon of "selection as a feature." Here, examples of the "learning process involving feature selection" include so-called "sparse modeling" techniques such as the above-mentioned LASSO method.
[0260] In sparse modeling, features are selected sparsely, i.e., certain features are selected with non-zero weights, while other features are selected with zero weights. One of the reasons for this sparse feature selection is that when there is a "group of features that make a similar contribution" to the "discrimination (identification) process," there is a "penalty term corresponding to the number of features" that allows the learning process to proceed so that one feature from the group is selected and the weights of the other features in the group are set to zero. This tendency is particularly noticeable in the LASSO method.
[0261] In other words, when feature A and feature B are equally involved in the "discrimination process," for example, when feature A and feature B are highly correlated, the discrimination process can be performed without degrading the discrimination performance even if only feature A is selected as the feature.
[0262] However, in clustering, for example, there may be cases where it is necessary to consider both feature A and feature B. However, if such "sparsification" is performed and features are selected based solely on their contribution to the discrimination process in a single classifier generation process, there is a possibility that the "feature selection" for clustering will be insufficient.
[0263] FIG. 17 is a diagram showing the concept of feature selection when there are a plurality of (for example, Nch) feature amounts by such a "learning process involving feature amount selection." Referring to FIG. 17, it is assumed that the subject group to be subjected to the learning process includes a group of healthy subjects and a group of patients. The subjects in the healthy subject group are assigned a label H, and the healthy subject group includes subtypes h1 and h2. It is assumed that a label M is associated with the subjects in the patient group, and the patient group includes subtype m1, subtype m2, and subtype m3.
[0264] Here, in terms of observables, it is unknown how many subtypes the healthy subject group and patient group will be divided into, and the identification labels of the subtypes are assumed to be latent labels that are not explicitly associated with the subjects. The goal of "clustering" is then to perform a data-driven clustering of the observables into these subtypes.
[0265] Since "undersampling" and "subsampling" are performed randomly from the healthy subject group and patient group as described above, subjects are selected from each of the healthy subject group and patient group, for example, from the area surrounded by the dotted line, as shown in "Subject Group" in Figure 17.
[0266] Now, the features (correlation values of brain function connections) that can be used to distinguish between label M and label H are assumed to be the features in the range indicated by the dashed dotted line in Figure 17 (the "union of brain function connections" in step S220 of Figure 16) among the brain function connections of the features (Nch in total) that characterize each subject.
[0267] Then, as a result of learning a classifier for distinguishing between label M and label H through a learning process involving feature selection (here, processing using the LASSO method), the features indicated by the black circles within the dashed dotted line in Figure 17 are selected.
[0268] FIG. 18 is a conceptual diagram showing features that are finally selected when generating one classifier through a learning process involving feature selection after undersampling and subsampling processes.
[0269] As shown in FIG. 18, the features that can be used to distinguish between label M and label H within the dashed line are further divided into groups of features that are highly correlated with each other, as shown in the dotted line frame. In the LASSO method, sparsification is achieved by selecting one feature for each group within the dotted line frame. FIG. 19 is a conceptual diagram showing how feature amounts are selected when a classifier is generated by performing undersampling and subsampling processes multiple times.
[0270] As shown in FIG. 19, if subsampling is performed, for example, Ns times, different subjects are subsampled from the healthy subject group and the patient group in each subsampling.
[0271] Then, for each subsample, a learning process involving feature selection is performed to learn a classifier for distinguishing between labels M and H. As a result, for each subsample, different features are selected by each classifier from each group with a high degree of correlation within the dot-dash line, which is the union described above, as indicated by the black circles.
[0272] As a result, by performing multiple undersampling and subsampling processes, a union of features that can be used to distinguish between labels M and H is selected.
[0273] In this embodiment, for each subsample, the features selected by the classifier learning process involving feature selection using the LASSO method are ranked according to the frequency with which they are selected.
[0274] Then, using a predetermined number of top ranked features, for example, 100 features, in step S224 of FIG. 16, clustering is performed by unsupervised learning using "multiple co-clustering" as described below.
[0275] In the above explanation, the LASSO method is used as an example of the "classifier learning process involving feature selection," and the feature selection for clustering is performed based on the ranking of the selected frequency.
[0276] However, as described above, in the clustering process of this embodiment, the "classifier learning process involving feature selection" is not limited to such a method, and may be, for example, a method such as the random forest method, and it is also possible to select features for clustering according to a predetermined level of importance.
[0277] For example, as described above, in the random forest method, the importance of features is calculated based on Gini impurity during the learning process of the classifier. Therefore, it is possible to rank the features based on this importance, and then use a predetermined number of top-ranked features to perform clustering through unsupervised learning by "multiple co-clustering" in step S224 of FIG. 16 .
[0278] Furthermore, as the "process of learning a classifier as ensemble learning," it is also possible to use a ridge regularization method or the like, rank features according to the importance corresponding to the median value obtained by aggregating the absolute values of the weighting coefficients, as described above, and select features for clustering using a predetermined number of features from the top of the ranking. [Multiple co-clustering processing] Below, the concept of "multiple co-clustering" in step S224 of FIG. 16 will be explained, and the term "multiple co-clustering" will be defined.
[0279] (Normal clustering method) As a premise, "clustering" refers to a data classification method using unsupervised learning performed by a computer, and more specifically, a method for automatically classifying given data without external criteria. In contrast, "classification" generally refers to a classification method using "supervised learning." Furthermore, a "cluster" is defined as a subset of data that has the properties of internal cohesion and external segregation. Here, external segregation refers to the property that objects in different clusters are dissimilar, and internal cohesion refers to the property that objects in the same cluster are similar to each other. Furthermore, the distance between elements of a set is defined as a measure of "similarity." Generally, the distance is defined to satisfy the so-called "distance axioms," and methods such as Euclidean distance, Mahalanobis distance, city block distance, and Minkowski distance are sometimes used.
[0280] In addition, commonly known methods for clustering using unsupervised learning include "partitional optimization clustering," which is a method for searching for a division that optimizes an objective function that defines the quality of a cluster, such as the non-hierarchical "k-means method," and hierarchical clustering methods such as "agglomerative hierarchical clustering" and "partitional hierarchical clustering."
[0281] However, such conventional clustering methods have the characteristic that they use all feature quantities to divide the objects into clusters (groups), and the resulting clusters are only divided in a single way.
[0282] Therefore, when there are multiple cluster divisions depending on the feature, there is a problem that it is difficult to deal with. Generally, it is thought that the more the number of features, the more likely it is that such multiple cluster structures exist.
[0283] In addition to the above-mentioned clustering methods, there are also algorithms that perform clustering by assuming that the multiple objects to be clustered in each cluster have arisen according to a certain probability distribution, and then estimating this "probability distribution." Known examples of such clustering methods include the "Gaussian mixture clustering method," which is known to enable more flexible clustering.
[0284] (Different clustering depending on feature selection) In the following, we will first explain the case where there are multiple ways to divide clusters depending on features, and assume that for a group of objects containing multiple objects to be clustered, each object is characterized by multiple features. FIG. 20 is a conceptual diagram for explaining a case where there are multiple ways of dividing into clusters depending on the feature amount.
[0285] As shown in FIG. 20, it is assumed that the data to be clustered (hereinafter simply referred to as "target") are six letters "A", "B", "C", "D", "E", and "F". These characters have different background patterns and different fonts (character styles).
[0286] Therefore, the features that characterize these characters can be considered to be "background pattern," "character style," and "number of holes contained in the character (number of areas completely surrounded by lines)." Therefore, even if a set of the same characters is considered, they will be divided into different clusters depending on which feature value is used for clustering.
[0287] In Figure 20, for example, when based on "background pattern," the images are divided into three clusters: {A, D} {B, E} {C, F}; when based on "character style," the images are divided into two clusters: {A, B, C} {D, E, F}; and when based on "number of holes," the images are divided into three clusters: {C, E, F} {A, D} {B}, corresponding to 0, 1, and 2 holes, respectively.
[0288] FIG. 20 shows an example in which one cluster is characterized by one feature amount, but in general, one cluster is characterized by a plurality of feature amounts. FIG. 21 is a conceptual diagram for explaining the concept of clustering when multiple objects are characterized by multiple feature amounts.
[0289] First, as shown in FIG. 21(i), consider a "data matrix" in which the objects to be clustered are arranged in the row direction and the feature quantities that characterize these objects are arranged in the column direction.
[0290] As shown in Figure 21(a), a method of clustering objects (dividing the objects into multiple object clusters) and simultaneously clustering features so that they are associated with each object cluster is called "co-clustering," and this method is disclosed, for example, in the following publicly known document.
[0291] Publications: Madeira SC, Oliveira AL. Biclustering algorithms for biological data analysis: a survey. IEEE / ACM Transactions on Computational Biology and Bioinformatics (TCBB). 2004; 1(1):24±45. https: / / doi.org / 10.1109 / TCBB.2004.2
[0292] In "co-clustering," as shown in Figure 21(a), the rows or columns of the data matrix are swapped, i.e., the objects and features are rearranged according to their similarity, to divide the data into cluster blocks represented by, for example, (i,j) (i=1,2:j=1,2,3).
[0293] At this time, a generative model (probabilistic model) for the objects contained in each cluster is assumed, and the parameters of each probabilistic model are determined so that the likelihood is high for the observed data.
[0294] In this way, once a probabilistic model is estimated for each cluster, it becomes possible to determine (classify) to which cluster specific observed data (test data) belongs. FIG. 22 is a conceptual diagram for explaining multiple clustering and multiple co-clustering.
[0295] In the "co-clustering" shown in Figure 21(a), block-structured clusters are generated by swapping the rows and columns of the "data matrix." Therefore, when the features are divided into multiple feature clusters, each object is clustered as if it were aligned in common to these multiple feature clusters.
[0296] However, if we assume that the features are divided into multiple feature clusters and that the targets are also divided into target clusters for each feature cluster, it is expected that a more likely probability model can be estimated if the arrangement of the targets in the target cluster (the arrangement of the targets contained in a target cluster) is also different for each feature cluster.
[0297] In such a case, the feature clusters are particularly called "views" in response to the fact that the method of dividing the object (clustering of the object) differs for each feature cluster. Executing clustering of different objects for each viewpoint of feature amounts in this manner is called "multiple clustering" as shown in FIG. 22(b).
[0298] Furthermore, if a probability model with a high likelihood for the observed data can be estimated by clustering the feature columns and the target rows at each viewpoint, this is called "multiple co-clustering," as shown in Figure 22(c).
[0299] Here, we refer to this as "multiple co-clustering," including cases where there is only one view, and cases where there is only one type of feature cluster in at least one view when there are multiple views, and we consider "co-clustering" and "multiple clustering" to be sub-concepts of "multiple co-clustering."
[0300] In this embodiment, the term "clustering" simply refers to generating a set of clusters in one view. For example, when dividing features into views and clustering the objects as shown in FIG. 22(b), it is called "multiple clustering," while when dividing features into views and co-clustering are performed simultaneously as shown in FIG. 22(c), it is called "multiple co-clustering." FIG. 23 is a conceptual diagram showing a case where probability models of different types of probability distributions are assumed within one view in "multiple co-clustering." FIG. 23(d) shows a case where the white blocks and the shaded blocks follow different types of probability distribution probability models. For example, the white blocks represent the probability distribution of continuous random variables, while the shaded areas represent the case where the probability distribution of discrete random variables is assumed.
[0301] As will be explained later, in the "multiple co-clustering learning method" of this embodiment, it is possible to perform clustering processing on a family of distributions that includes different distributions in this way. FIG. 24 is a flowchart for explaining an outline of the multiple co-clustering learning method.
[0302] When the processing of the multiple co-clustering learning method starts (S300), the clustering classifier generation unit 3010 randomly divides features into subgroups for the data matrix, and generates feature views and feature clusters within the views (S302: corresponds to the generation of Y (initialization of Y) described below).
[0303] Next, the clustering classifier generation unit 3010 generates and optimizes the division of the target cluster in accordance with the feature views and feature clusters generated in step S302 (step S304: corresponds to generating Z, which will be described later).
[0304] Furthermore, the clustering classifier generation unit 3010 optimizes the division of the feature quantities for the obtained target clusters (S306: corresponds to the generation process of Y described later, and optimizes Y using the generated Z).
[0305] Next, the clustering classifier generation unit 3010 determines whether the objective function satisfies predetermined conditions and has converged (S308). Note that this objective function corresponds to the function L(q(φ)) described below. The function L(q(φ)) has the property of monotonically increasing as Y and Z, also described below, are updated, and it is determined that convergence has occurred when it is determined that the rate of increase has become sufficiently small. If the clustering classifier generation unit 3010 has not converged (N in S308), it returns the process to step S304, and if the objective function has converged (Y in S308), it proceeds to the next step. Then, the clustering classifier generation unit 3010 stores the magnitude of the objective function in the storage device 2080 (S310).
[0306] Next, the clustering classifier generation unit 3010 determines whether the processes from steps S302 to S310 have been performed a predetermined number of times. If the processes have not been performed the predetermined number of times (N in S312), the clustering classifier generation unit 3010 returns the process to step S302, and if the processes have been performed the predetermined number of times (Y in S312), the clustering classifier generation unit 3010 proceeds to the next step.
[0307] The clustering classifier generation unit 3010 ends the learning process for multiple co-clustering with the feature division and cluster division that maximizes the objective function as the final result (S314), and generates a clustering classifier. (More details on the multiple co-clustering process) The learning method for multiple co-clustering explained in FIG. 24 will be explained in more detail below. Details of the multiple co-clustering process are disclosed in the following document, and an outline of the process will be described below.
[0308] Published literature: Tomoki Tokuda, Junichiro Yoshimoto, Yu Shimizu, Go Okada, Masahiro Takamura, Yasumasa Okamoto, Shigeto Yamawaki, Kenji Doya, “Multiple co-clustering based on nonparametric mixture models with heterogeneous marginal distributions”, PLOS ONE | https: / / doi.org / 10.1371 / journal.pone.0186566 October 19, 2017 FIG. 25 is a diagram showing a graphical representation of Bayesian inference in the multiple co-clustering learning method of FIG.
[0309] The multiple co-clustering model is summarized in the graphical model in Figure 25, which clarifies the causal links between the relevant parameters and data matrices. (Multiple Co-clustering Model) The feature values (brain function connectivity values) and subjects (here, subjects in the patient group) are represented as a data matrix as shown in FIG. 21(i). It is assumed that the data matrix X is composed of a family of M distributions known in advance. The probability distributions belonging to the distribution family may include Gaussian distribution, Poisson distribution, categorical distribution / multinomial distribution, etc. The clustering classifier generation unit 3010 generates X (m) For each data size is n× d (m) So, divide it as follows: X = {X (1) ,…,X (m) ,…,X (M)}
[0310] Here, m is an index indicating the distribution family (m = 1, ..., M). Furthermore, the number of views (points of view) is V (common to all distribution families), and the number of feature clusters for view v and distribution family m is G. ν (m) Let the number of object clusters in view v be K v (common to all distribution families)
[0311] Furthermore, for the sake of notational simplicity, we allow the existence of empty clusters to indicate the number of features and the number of clusters, G (m) =max v G v (m) and K=max v K v It is written as follows.
[0312] In this notation, the independent and identically distributed (iid) d (m) dimensional random vector X1 (m) , …, X n (m) About d (m) × V × G (m) (3rd order) feature split tensor Y (m) When feature j in distribution family m belongs to feature cluster g of view v, Y j,v,g (m) =1 (otherwise 0). Combining this for different families of distributions, we can get Y = {Y (m)} m Let's say.
[0313] Similarly, consider an n×V×K object partitioning (3rd order) tensor Z. If object i belongs to object cluster k in view v, then Z i、v、k = 1.
[0314] Feature j is one of the views (Σ v,g Y j, v, g (m) =1), and object i belongs to each view (i.e., Σ k Z i,v,k (m) = 1). Furthermore, Z is common to all distribution families, which means that the estimated probability model uses information from all distribution families to estimate the subject clustering solution.
[0315] First, as shown in Fig. 25, for the pre-generated model of Y, a hierarchical structure of views and feature clusters is considered, where views are generated first and then feature clusters are generated. Therefore, features are divided by the membership of the view-feature cluster pair, and the allocation of feature divisions is determined jointly by the view and feature cluster.
[0316] On the other hand, as shown in Figure 25, since the object is divided into object clusters for each view, we only consider one structure of object clusters for Z. We assume that these generative models are all based on the Stick Breaking Process (SBP), as explained below.
[0317] (Generative model of feature cluster Y) Y j.. (m) Let denote the view / feature cluster membership vector of feature j in distribution family m, generated by the hierarchical fold process, then the following equation holds:
[0318]
number
[0319] Publication: Blei DM, Jordan MI, et al. Variational inference for Dirichlet process mixtures. Bayesian analysis. 2006; 1(1) : 121-143, https: / / doi.org / 10.1214 / 06-BA104
[0320] Y j,v,g (m) = 1, feature j belongs to feature cluster g of view v. By default, the concentration parameters α1 and α2, which are hyperparameters, are set to 1. (Generative model of object cluster Z) is the subject cluster membership vector for object i in view v, Z i, v. A vector denoted as is generated by the following formula:
[0321]
number
[0322] Here Z i, v. is Z i,v = (Z i, v, 1 ,…, Z i, v, K ) T The concentration parameter β is set to 1.
[0323] (Likelihood and prior distribution) Each instance X i,j (m) is assumed to follow a specific distribution independently, conditional on Y and Z. The parameters of the distribution family m for the cluster block of view v, feature cluster g, and object cluster k are defined as θ v,g,k (m) It is expressed as: Furthermore, Θ={θ v,g,k (m)} v,g,k,m Denoting it as , the logarithm of the likelihood of X follows the formula:
[0324]
number
[0325]
number
[0326] (Variational Estimation) The variational Bayes EM algorithm is used for MAP (maximum a posteriori) estimation of Y and Z. Such a variational Bayes EM algorithm is disclosed in the following document:
[0327] Publication: Guan Y, Dy JG, Niu D, Ghahramani Z. Variational inference for nonparametric multiple clustering. In: MultiClust Workshop, KDD-2010; 2010. The log marginal likelihood p(X) is approximated using Jensen's inequality as follows:
[0328]
number
[0329] Publication: Jensen V. Acta Mathematica. 1906; 30(1):175-193. https: / / doi.org / 10.1007 / BF02418571
[0330] where q(φ) is an arbitrary distribution of parameter φ. We prove that the Kullback-Leibler divergence between q(φ) and p(φ), i.e., the difference between the left and right sides, is given by KL(q(φ), p(φ|X)). Thus, an approach to choosing q(φ) is to minimize KL(q(φ), p(φ|X)), which is generally difficult to evaluate. Here we choose q(φ) to be factorized for different parameters (mean field approximation).
[0331]
number
[0332] where each q(·) is a subset of parameters w v , w´ g,v (m) ,Y j.. (m) , u k, v , Z i,v. and θ v,g,k (m) is further factorized for In general, KL(Π l=1 L q l (φ l ), p (φ|x)) i (φ i ) is given by the following formula:
[0333]
number
[0334] Published literature: Murphy K. Machine Learning: A Probabilistic Perspective. Cambridge, Massachusetts: MIT Press; 2012. Applying this property to the model currently under consideration, we can show the following.
[0335]
number
[0336]
number
[0337]
number
[0338]
number
[0339] τ j,g,v (m) is normalized over pairs (g,v) for each pair (j,m), while η i, v, k is normalized with respect to k for each (i, v) pair. The observation model and the prior distribution of the parameters Θ will be described later.
[0340] (observation model) The observation model considers Gaussian, Poisson, and categorical / multinomial distributions. For each cluster block, we fit a univariate distribution from these families, assuming that the features within the cluster block are independent. The parameters of these distribution families assume conjugate prior distributions.
[0341] (Optimization algorithm) The hyperparameter update equation is calculated using the variational Bayes EM algorithm as follows:
[0342] First, randomly select {τ (m)} m and {η v} v We initialize and update the hyperparameters until the lower bound L(q(φ)) in equation (1) converges. This produces a locally optimal distribution q(φ) with respect to L(q(φ)). We repeat this procedure many times and select the best solution with the largest lower bound as the approximate posterior distribution q*(φ). The MAP estimates of Y and Z are argmax Y q * Y (Y) and argmax Z q * ZIt is evaluated as (Z). The lower bound L(q(φ)) is given by the following formula:
[0343]
number
[0344] Both terms on the right-hand side can be derived in closed form. It can be shown that as q(φ) is optimized, this value increases monotonically. That is, as described above, the function L(q(φ)) has the property of increasing monotonically as Y and Z are updated, and convergence is determined when it is determined that the rate of increase has become sufficiently small (for example, when the condition that the increment is equal to or less than a predetermined value is met, although this is not particularly limited).
[0345] First, the distribution family of each feature is identified and a data matrix of the corresponding distribution family is generated. Then, for the set of data matrices, MAP estimates of Y and Z are further generated, and the object / feature clusters in each view are analyzed using the Y and Z estimates.
[0346] (Model representation) The multiple co-cluster model is flexible enough to represent various clustering models because the number of views and the number of feature / object clusters are derived using a data-driven approach. For example, when the number of views is 1, the model matches the co-cluster model. When the number of feature clusters is one for all views, the model matches the multi-clustering model. Furthermore, when the number of views is 1 and the number of feature clusters is the same as the number of features, the model matches the traditional mixture model with independent features. Furthermore, the model can detect uninformative features that do not distinguish object clusters. In such cases, the model generates a view with a single object cluster. The advantage of the model is that it automatically detects such underlying data structures. The above-described "multiple co-clustering method" makes the following possible:
[0347] 1) It is possible to identify the division of multiple clusters behind the data (including not only the division of the objects but also the division of the features) and the corresponding set of features in a data-driven manner. 2) This method makes it possible to identify clusters that could not be found by other methods. 3) Furthermore, the division of each cluster can be given meaning by features, making it easier to interpret each cluster.
[0348] [Evaluating clustering results for a dataset] In the following, we will examine the generalization performance of clustering by dividing the multi-center, large-scale fMRI data collected from a large number of subjects, which has been published as the SRPBS described above, into two parts and using the multiple co-clustering method described above on each part. FIG. 26 is a diagram showing the data set 1 and data set 2 thus divided into two.
[0349] Dataset 1 consists of data on 545 healthy subjects and 138 depressed patients obtained at facilities 1 to 4, and data set 2 consists of data on 263 healthy subjects and 181 depressed patients obtained at facilities 5 to 8. These basically correspond to the data sets shown in Figures 10 and 11. FIG. 27 is a conceptual diagram illustrating the concept of performing clustering on each data set.
[0350] As shown in FIG. 27, clustering is performed independently on each of the datasets 1 by the multiple co-clustering method according to the flow shown in FIG.
[0351] Here, the question of interest is how similar (what is the degree of agreement) the clusters obtained by the data-driven clustering method performed separately on Data Set 1 and Data Set 2 are to each other.
[0352] If the clustering in Dataset 1 and Dataset 2 can be classified (grouped) into clusters (subject groups) with identical or similar characteristics, then this type of data-driven clustering is performed with high generalizability, independent of facilities, measurement equipment, etc. The question then becomes how to quantitatively evaluate "classification into clusters with identical or similar characteristics." FIG. 28 is a conceptual diagram showing an example of multiple co-clustering on subject data. As shown in FIG. 28(a), the input data matrix has subjects arranged in the row direction and feature amounts arranged in the column direction.
[0353] When multiple co-clustering is performed on this input data matrix, the features are divided into two views, and the subjects are clustered in each view, as shown in FIG. 28(b), for example. FIG. 29 shows the results of actually performing multiple co-clustering processing on Data Set 1 and Data Set 2. In FIG. 29, 99 features are selected for each of Data Set 1 and Data Set 2 as features for clustering.
[0354] Then, multiple co-clustering was performed on 138 depressed patients for Dataset 1, and on 181 depressed patients for Dataset 2.
[0355] For dataset 1, features are divided into two views: view 1 and view 2. For view 1, features are further co-clustered into two feature clusters, and subjects are divided into five subject clusters. For view 2, subjects are also divided into five clusters.
[0356] For Dataset 2, the features are also divided into two views: View 1 and View 2. For View 1, the features are further co-clustered into two feature clusters, and the subjects are divided into four subject clusters, and for View 2, the subjects are divided into five clusters. FIG. 30 is a table showing the number of brain functional connections (FC) assigned to each view in Dataset 1 and Dataset 2. In dataset 1, 92 FCs are assigned as features to view 1 and 7 FCs are assigned to view 2. In dataset 2, 93 FCs are assigned as features to view 1 and 6 FCs are assigned to view 2.
[0357] In addition, this table lists the number of matching brain function connections assigned in Dataset 1 and Dataset 2 on the diagonal of the table. It can be seen that the brain function connections assigned to View 1 and View 2 in Dataset 1 and Dataset 2 are almost identical. (Method for verifying generalizability (similarity between data sets) of clustering (stratification))
[0358] Below, we quantitatively evaluate the degree of similarity (degree of agreement) between the clusters obtained by data-driven clustering performed separately on Data Set 1 and Data Set 2. FIG. 31 is a conceptual diagram for explaining a method for evaluating the similarity of such clustering (generalization performance of stratification).
[0359] First, as shown in Figure 31(a), when Data Set 1 and Data Set 2 are independently divided into clusters using the multiple co-clustering method described above, it is difficult to compare the similarity of the clusterings because the subjects in the clusters of each dataset are mutually independent.
[0360] Here, the result of classifying the subjects of Dataset 1 using Classifier 1 created with Dataset 1 is called Clustering 1. On the other hand, the result of classifying the subjects of Dataset 2 using Classifier 2 created with Dataset 2 is called Clustering 2.
[0361] In contrast, as shown in Figure 31(b), the result of classifying the subjects of Dataset 2 using Classifier 1 created with Dataset 1 is called Clustering 1'. On the other hand, the result of classifying the subjects of Dataset 1 using Classifier 2 created with Dataset 2 is called Clustering 2'.
[0362] In this case, the classification is performed for common subjects between clustering 1 and clustering 1', and between clustering 2 and clustering 2', so that the degree of similarity between them can be evaluated. (An evaluation measure of similarity (recall) between clusters)
[0363] Here, the clustering process is performed in a data-driven manner, and the cluster index values themselves in Figure 29 (the order of the index values) are not significant. Therefore, when different clustering processes are performed on the same data set, the question arises as to how to evaluate the similarity.
[0364] For example, when there are two clustering results π and ρ for the same data set X, the Rand index is known as a measure for evaluating the similarity (external validity measure) between these two clustering results.
[0365] For all data pairs {x1, x2}∈X (M=N(N-1) / 2) in the dataset, there are the following types of pairs, and the number of pairs belonging to each type is defined as follows:
[0366]
number
[0367]
number
[0368] However, it is known that, for example, if there is a bias in the number of elements in each cluster of a data set, "even if the clustering is random, the Rand Index may end up being a high value." For this reason, more strictly, the Adjusted Rand Index (ARI) as shown below is used. FIG. 32 is a conceptual diagram for explaining ARI. ARI is disclosed in the following documents, for example:
[0369] Publication: Jorge M. Santos and Mark Embrechts, “On the Use of the Adjusted Rand Index as a Metric for Evaluating Supervised Classification”, ICANN 2009, Part II, LNCS 5769, pp. 175-184, 2009.
[0370] As mentioned above, when two clustering results are applied to the same data set, there are cases where the data are classified into the same cluster both times, as shown in Figure 32(a), and cases where the data are classified into different clusters both times, while there are also cases where the data are classified into the same cluster once and into a different cluster the other time, as shown in Figure 32(b).
[0371] The ARI is calculated by calculating the expected value when a pair of data is classified into the same cluster or into different clusters in two clusterings, assuming that the two clusterings are independent of each other, and subtracting the expected value from the numerator and denominator of the Rand index. Therefore, the ARI is adjusted so that its value is 0 when there is no correlation between the clusterings.
[0372]
number
[0373] Here, A represents the number of subject pairs that were classified into the same cluster both times + (classified into different clusters both times) for the two clusterings, max(A) represents the total number of pairs, and E represents the number of subject pairs whose assignment results match despite the two clusterings being independent. FIG. 33 is a diagram showing the evaluation results of the similarity between clustering 1 and clustering 1', and between clustering 2 and clustering 2'. FIG. 33(a) is a table showing the calculated ARI for each view of Dataset 1 and Dataset 2.
[0374] For Dataset 1 and Dataset 2, the ARI for View 1 is 0.47 and the ARI for View 2 is 0.51, which indicates a significant similarity.
[0375] Figure 33(b) shows the results of the permutation test corresponding to Figure 33(a). It can be seen that the ARI values (shown by the solid line) for View 1 and View 2 are statistically significantly higher than when the elements are swapped (shown by the histogram).
[0376] The "permutation test" is the result of calculating the ARI value when the cluster attribute labels of subjects are randomly swapped between subjects, and the figure shows the distribution of the results when such swaps are performed a predetermined number of times as a histogram. If the similarity between clusters is statistically significant, the ARI value between the clusters being compared will be significantly higher than when elements are swapped randomly.
[0377] From the above, it can be concluded that significant similarity was observed in the clustering (stratification) between the data sets, that is, generalized clustering was achieved.
[0378] As explained above, both the multiple co-clustering for Dataset 1 and the multiple co-clustering for Dataset 2 were achieved through data-driven methods, and can therefore be said to form the basis for "patient stratification" using brain functional connectivity as a feature. FIG. 34 is a table showing the distribution of the number of subjects assigned to each cluster in view 1 in clustering 1 and clustering 1′.
[0379] By appropriately sorting the subject cluster indexes, most subjects can be distributed near the diagonal of the table, and it can be visually confirmed that the two clusterings are similar to each other.
[0380] [Harmonization Processing] The following describes the details of the process called harmonization process in FIG. 6, which is disclosed in the following document. [Harmonization of Travel Subject Laws]
[0381] Below, we will explain a method used to generate the ``disease classifier'' described above and for the ``clustering process'' for stratification, which is a method for evaluating measurement bias and harmonizing measurement data independently of sample bias.
[0382] Figure 35 is a conceptual diagram for explaining a method for evaluating inter-site differences in the rs-fcMRI method of this embodiment using a moving subject (hereinafter referred to as a "traveling subject") who undergoes measurement while moving between sites.
[0383] As will be explained below, the present embodiment describes a harmonization method that uses a dataset of traveling subjects and can remove only measurement bias.
[0384] Referring to FIG. 35, a data set of traveling subject TS1 (number of people: Nts) is acquired to evaluate measurement bias across measurement sites MS.1 to Ms.Ns.
[0385] The resting-state brain activity of Nts healthy participants is assumed to be imaged at each of Ns sites, where Ns sites include all sites that imaged patient data. The acquired dataset of the traveling subject is stored as moving subject data in a storage device 210 in the data center 200 . Then, as will be described later, the processing for the "brain activity biomarker harmonization method" is executed in the computer processing system 300.
[0386] The dataset for traveling subjects includes only healthy controls, and participants are the same across all sites. Therefore, for traveling subjects, differences between sites consist only of "measurement bias."
[0387] In the harmonization method of this embodiment described below, a "brain activity biomarker harmonization method" performs a process of correcting the measurement data at each measurement site to remove the influence of "measurement bias."
[0388] In other words, in the following, we will evaluate "measurement bias" and "sample bias" using the "generalized linear mixed model (GLMM)" among the "statistical modeling" methods.
[0389] Typically, a GLM (Generalized Linear Model) is a model that incorporates "explanatory variables" that explain the probability distribution of a "response variable." A GLM has three main components: a "probability distribution," a "link function," and a "linear predictor." By specifying how these components are combined, various types of data can be represented.
[0390] Furthermore, GLMM (Generalized Linear Mixed Model) is a statistical model that can incorporate "individual differences that cannot be measured by humans or that were not measured by humans" that cannot be explained by GLM. For example, GLMM can incorporate location differences into the model even when the subject consists of several subsets (for example, subsets measured at different locations). In other words, it is a model that uses (mixes) multiple probability distributions as components. For example, the GLMM is disclosed in the following document: Publicly known literature 8: Takuya Kubo, "Introduction to Statistical Modeling for Data Analysis," Iwanami Shoten, 1st printing, 2012, 14th printing, 2017
[0391] However, in the statistical model of this embodiment described below, the terms "bias" and "factor" are used for what are usually called "effects," with "bias" being used for "measurement bias" and "sample bias," and "factor" being used for other factors (subject factors and disease factors).
[0392] The following analysis differs from the simple GLMM approach in that it analyzes factors without distinguishing between "fixed effects" and "random effects." This is because, typically, when using GLMM, only the variance of random effects is estimated, making it impossible to determine the magnitude of the effect of each factor. Therefore, in the following, in order to evaluate the magnitude of the effect of each factor, we transform the variables as follows so that each factor becomes a fixed effect with a mean of 0, and then perform estimation. i) Define the measurement bias for each site as the deviation of the correlation value for each functional connectivity from the mean across all sites.
[0393] ii) We assume that the sampling biases of healthy individuals and patients with psychiatric disorders are different from each other. Therefore, we calculate the sampling bias of each site separately for the healthy individuals and the patients with each disorder. iii) Disease factors are defined as deviations from the values of healthy subjects.
[0394] That is, in the following, a generalized linear mixed-effects model is applied to a dataset containing patients and a dataset containing traveling subjects as follows:
[0395] The number of traveling subjects is Nts people, and of the Ns measurement sites, the number of sites where measurements were taken of healthy individuals is Nsh, and the number of sites where measurements were taken of patients with a certain disease (here represented by the subscript "dis") is Nsd.
[0396] Participant factors (p), measurement bias (m), sample bias (Shc, Sdis) and psychiatric factors (d) are assessed by fitting regression models to the correlation values of functional connectivity for all participants from the patient outcome dataset and the traveling subject dataset. In what follows, vectors are denoted by lowercase letters (e.g., m) and all vectors are assumed to be column vectors. The elements of the vector are m k It is indicated by a subscript, such as: The regression model of a functional connectivity vector (a column vector) consisting of n correlation values between brain areas is expressed as follows:
[0397]
number
[0398] To represent participant characteristics, we use a 1-of-K binary coding scheme, where the target vector (e.g., xm) for measurement bias m belonging to site k has all elements equal to zero except for element k, which is equal to one. If the participant does not belong to any class (healthy subject, patient, traveling subject), the target vector is a vector with all elements equal to 0. The superscript T denotes a matrix or vector permutation, x T represents a row vector.
[0399] where m represents the measurement bias (Ns × 1 column vector), shc represents the sample bias of the healthy control group (Nsh × 1 column vector), sdis represents the sample bias of the patient group (Nsd × 1 column vector), d represents the disease factor (2 × 1 column vector with healthy and disease as elements), p represents the participant factor (Nts × 1 column vector), const represents the average functional connectivity across all participants (including healthy controls, patients, and traveling subjects) from all measurement sites, and e~N(0, γ -1 ) represents noise.
[0400] For the sake of simplicity, the description here is given assuming that there is one type of disease, but cases where there are multiple types of disease will be described later. For each functional connectivity correlation, we used L2 regularized least squares regression to estimate each parameter because the regression model design matrix was rank-deficient. Note that other evaluation methods, such as Bayesian inference, can also be used. After the regression calculations above, the bth connectivity of subject a can be written as:
[0401]
number
[0402] First, fMRI measurement data of subjects (healthy subjects and patients), attribute data of the subjects, and measurement parameters are collected from each measurement site into the storage device 210 of the data center 200 (FIG. 37, S402).
[0403] Next, the brain activity of traveling subject TS1 is measured at each measurement site at a predetermined period (for example, once a year), although this is not limited to this, and the fMRI measurement data of the traveling subject, attribute data of the subject, and measurement parameters are collected from each measurement site in the storage device 210 of the data center 200 (Figure 37 S404).
[0404] The harmonization calculation unit 3020 evaluates the measurement bias of each measurement site for functional connectivity by using the above-mentioned GLMM (generalized linear mixed model) (FIG. 37, S406).
[0405] The harmonization calculation unit 3020 stores the measurement bias of each measurement site calculated in this way in the storage device 2080 as measurement bias data 3108 (FIG. 37, S408).
[0406] (Harmonization in classifier generation process) A brief description will be given of harmonization of brain function connectivity values in the process in which the discrimination processing unit 3000 generates a disease classifier for the disease or healthy label of a subject. Such a disease classifier provides auxiliary information (support information) for diagnosing a subject.
[0407] The correlation value correction processing unit 3004 reads out the measurement bias data 3108 for each measurement site stored in the storage device 2080, and performs harmonization processing on the off-diagonal elements of the correlation matrix of each subject who is the training target for machine learning to generate a disease classifier, using the following equation.
[0408]
number
[0409] Here, "Connectivity" represents the functional connectivity vector before harmonization, and "Csub" represents the functional connectivity vector after harmonization. Furthermore, "m" (hereinafter, the letter "x" with a ^ at the beginning will be referred to as "x") represents the measurement bias at the measurement site evaluated by least-squares regression with L2 regularization as described above. Thus, the measurement bias corresponding to the measurement site where the functional connectivity was measured is subtracted from the functional connectivity, and the result is subjected to harmonization processing. The data after the correction process is stored in the storage device 2080 as corrected correlation value data 3110.
[0410] The method for eliminating bias between measurement sites is not limited to the traveling subject method described above. For example, other methods such as the ComBat method may also be used.
[0411] [Embodiment 2] In embodiment 1, an example configuration using distributed processing has been described in which a brain activity measuring device (fMRI device) measures brain activity data measured at multiple measurement locations, and based on this brain activity data, biomarkers are generated and diagnostic labels are estimated (predicted) using the biomarkers.
[0412] However, it is also possible to configure the following processes to be performed in different facilities: i) measurement of brain activity data (data collection) for training biomarkers through machine learning; ii) generation of biomarkers through machine learning and estimation (prediction) of diagnostic labels using biomarkers for a specific subject (the subject of estimation, also referred to as the "first subject" hereinafter); and iii) measurement of brain activity data for the specific subject (brain activity measurement of the first subject). FIG. 38 is a functional block diagram showing an example of a case where data collection, estimation processing, and measurement of subject's brain activity are processed in a distributed manner.
[0413] Referring to FIG. 38, sites 100.1 to 100.N are facilities that measure data of a patient group and a healthy subject group (i.e., a second subject group) using a brain activity measuring device, and management server 200 manages the measurement data from sites 100.1 to 100.Ns. The computer processing system 300 generates a classifier from the data stored in the server 200 .
[0414] Furthermore, the harmonization calculation unit 3020 of the computer processing system 300 executes harmonization processing for the sites 100.1 to 100.Ns and the site of the MRI measurement device 410, as well. The MRI device 410 is provided at a separate site that uses the results of the classifier on the computer processing system 300, and measures brain activity data for a specific subject.
[0415] The computer 400 is installed at a separate site where the MRI device 410 is installed, calculates correlation data of functional connectivity of the brain of a specific subject from measurement data of the MRI device 410, transmits the correlation data of functional connectivity to the computing system 300, and uses the results of the classifier returned.
[0416] The server 200 stores MRI measurement data 3102 of the patient group and the healthy subject group transmitted from the sites 100.1 to 100.Ns, and personal attribute information 3104 of the subjects associated with the MRI measurement data 3102, and transmits these data to the computing system 300 in response to access from the computing system 300.
[0417] The computer processing system 300 receives the MRI measurement data 3102 and the subject's personal attribute information 3104 from the server 200 via the communication interface 2090 .
[0418] The hardware configurations of server 200, computing system 300, and computer 400 are basically the same as the configuration of "data processing unit 32" described in FIG. 5, and therefore the description thereof will not be repeated.
[0419] Returning to Figure 38, the correlation matrix calculation unit 3002, correlation value correction processing unit 3004, disease classifier generation unit 3008, clustering classifier generation unit 3010, and discriminant value calculation unit 3012, as well as functional connection correlation matrix data 3106, measurement bias data 3108, corrected correlation value data 3110, and classifier data 3112 are the same as those described in embodiment 1, and therefore their description will not be repeated.
[0420] The MRI device 410 measures brain activity data of a subject for whom a diagnostic label is to be estimated, and the processing device 4040 of the computer 400 stores the measured MRI measurement data 4102 in the nonvolatile storage device 4100 .
[0421] Furthermore, the processing unit 4040 of the computer 400 calculates functional connection correlation matrix data 4106 based on the MRI measurement data 4102 in the same manner as the correlation matrix calculation unit 3002 , and stores the data 4106 in the nonvolatile storage device 4100 .
[0422] A disease to be diagnosed is designated by a user of the computer 400, and in accordance with a transmission instruction from the user, the computer 400 transmits functional connectivity correlation matrix data 4106 to the computer processing system 300. In response to this, the computer processing system 300 executes harmonization processing corresponding to the site where the MRI apparatus 410 is installed, and the discriminant value calculation unit 3010 calculates a discrimination result for the designated diagnostic label and an evaluation result for the subtype, which the computer processing system 300 transmits to the computer 400 via the communication interface 2090. The computer 400 notifies the user of the determination result via a display device (not shown). With this configuration, it becomes possible to provide the results of diagnostic label estimation by the classifier based on data collected from a larger number of subjects.
[0423] It is also possible for server 200 and computer processing system 300 to be managed by separate administrators, in which case the security of subject information stored on server 200 can be improved by restricting the computers that can access server 200.
[0424] Furthermore, from the perspective of the operator of the computing system 300, it is possible to provide a "service of providing discrimination results" without providing any information about the discriminator or information about the "measurement bias" to the "party (computer 400) receiving the service of discrimination using a discriminator."
[0425] In the above descriptions of the first and second embodiments, a real-time fMRI has been used as a brain activity detection device for measuring brain activity over time using functional brain imaging. However, the above-mentioned fMRI, magnetoencephalography, near-infrared spectroscopy (NIRS), electroencephalography, or a combination thereof can also be used as the brain activity detection device. For example, when a combination of these is used, fMRI and NIRS detect signals related to changes in blood flow in the brain and have high spatial resolution. On the other hand, magnetoencephalography and electroencephalography are characterized by high temporal resolution for detecting changes in electromagnetic fields associated with brain activity. Therefore, for example, by combining fMRI and magnetoencephalography, brain activity can be measured with high spatial and temporal resolution. Alternatively, by combining NIRS and electroencephalography, a system for measuring brain activity with similar high spatial and temporal resolution can be configured in a compact and portable size.
[0426] With the above configuration, it is possible to realize a brain activity analysis device and a brain activity analysis method that function as biomarkers for neurological and psychiatric disorders using functional brain imaging.
[0427] Furthermore, in the above explanation, an example has been described in which a classifier is generated by machine learning and functions as a biomarker when a "diagnostic label" is included as an attribute of a subject. However, the present invention is not necessarily limited to such a case, and the present invention may be used for other discrimination purposes as long as a group of subjects from whom measurement results that are the subject of machine learning are obtained are divided into multiple classes in advance using an objective method, the correlation (connection) of activity between brain regions (regions of interest) of the subjects is measured, and classifiers for the classes can be generated by machine learning on the measurement results. Also, as mentioned above, such a determination may be made by displaying the possibility of belonging to a certain attribute as a probability.
[0428] Therefore, for example, it is possible to objectively evaluate whether a certain "training" or "behavioral pattern" is useful for improving the health of a subject. In fact, even if a subject is not yet in a disease state (pre-disease state), it is possible to objectively evaluate whether certain foods, drinks, and other intakes, or certain activities, are effective in bringing the subject closer to a healthier state.
[0429] Furthermore, even in a pre-disease state, as described above, if a display such as "The probability of being healthy is XX%" is output, the health state can be displayed to the user as an objective numerical value. In this case, what is output does not necessarily have to be a probability, but a "continuous value of the health degree, for example, the probability of being healthy" converted into a score can also be displayed. By displaying such a value, the device of this embodiment can be used not only for diagnostic support but also as a device for managing the user's health.
[0430] [Embodiment 3] [Treatment selection support system] (Configuration of treatment selection support system) An embodiment of the present invention relates to a treatment method selection support system 1000 for providing information regarding the selection of a treatment method for a subject exhibiting depressive symptoms, based on the results of measuring the brain activity of the subject. FIG. 39 is a diagram showing the functional configuration of the treatment method selection support system 1000.
[0431] The treatment selection support system 1000 includes a computer 400 connected to a medical institution 4 so as to be able to receive data from an MRI device 410 for measuring brain activity data of a subject for whom a diagnostic label is to be estimated, a support information providing device 300a, a classifier generating device 300b that performs processing to generate a clustering classifier, a data server 200', and a treatment information providing server 500.
[0432] Referring to Figure 39, the healthy subject / patient database in the storage device 210 in the data server 200' manages MRI measurement data 3102 and subject attribute information 3104 collected from sites where MRI devices 100.1 to 100.N (not shown) are installed.
[0433] The support information providing device 300a corresponds to the configuration of the computing system 300 shown in FIG. 38 except that i) the processing unit 3200 receives as input the measurement results of the subject's brain activity transmitted from the computer 400 of the medical institution 4, and the computing device 2040a executes the processing of the treatment information generation unit 3200, which outputs corresponding treatment information according to the classification results of the measurement results by the clustering classifier executed as the discriminant value calculation unit 3012, and ii) the processing of the disease classifier generation unit 3008 and the clustering classifier generation unit 3010 is executed by the computing device 2040b in the classifier generation device 300b, which is a computing system different from the support information providing device 300a. As in FIG. 38, the support information providing device 300a and the classifier generating device 300b may be implemented as functions on the same computer system.
[0434] The classifier generation device 300b generates a disease classifier and a clustering classifier using the data stored in the server 200′ as “learning data” (or “discovery cohort data”). The data (other than the learning data) stored in the server 200′ can also be used as validation data (or “validation cohort data”) for the disease classifier and the clustering classifier. Although not particularly limited, in the configuration of FIG. 39, the healthy subject / patient database in the storage device 210 of the server 200′ not only stores MRI measurement data 3102 and the subject's personal attribute information / measurement parameters 3104, but also pre-calculated and stored data on the correlation matrix of functional connectivity calculated based on the MRI measurement data 3102 and data on correlation values corrected by harmonization processing. The disease classifier generation unit 3008 and the clustering classifier generation unit 3010 will be described as executing the learning process and validation process based on the corrected correlation values. The classifier generating device 300b itself may calculate the data of the correlation matrix of functional connections and perform the process of correcting the correlation values through harmonization processing. In addition, basically, parts that are the same as those in the configuration of FIG. 38 are given the same reference numerals.
[0435] Although not particularly limited, the data acquired by the MRI device 410 may be added to the healthy subject / patient database in the storage device 210, and the disease classifier and clustering classifier may be retrained.
[0436] Furthermore, the harmonization calculation unit 3020 of the support information providing device 300a executes harmonization processing for the MRI measurement device 410 for the MRI devices 100.1 to 100.Ns (not shown).
[0437] The MRI device 410 is installed in a medical institution 4 that uses the results of the clustering classifier of the support information providing device 300a, and measures brain activity data for a specific subject. The measured imaging data (including brain structural image data and brain function image data) of the specific subject is anonymized in the anonymization processing unit 4042 of the computer 400, and then transmitted to the support information providing device 300a.
[0438] Although not particularly limited, a so-called cloud computer may be used for the support information providing device 300a. Also, although not particularly limited, the computer 400 may be configured to calculate correlation data of functional connectivity of a specific subject's brain from measurement data of the MRI device 410 and transmit the correlation data of functional connectivity to the support information providing device 300a. In this case, the conversion of the imaging data into correlation data of functional connectivity itself corresponds to anonymization.
[0439] The treatment information generation unit 3200 receives data in the treatment information DB 5100 in the treatment information provision server 500 from the treatment information provision system 5200, and returns corresponding treatment method selection support data according to the classification result of the clustering classifier to the information presenting unit 4044. The treatment information generation unit 3200 may be configured to receive and store in advance information on the correspondence between the classification result and the treatment method selection support data from the treatment information provision system 5200, or may be configured such that the treatment information generation unit 3200 transmits information on the classification result as inquiry information to the treatment information provision system 5200 each time, and receives treatment method selection support data returned from the treatment information provision system 5200 as a response to the inquiry.
[0440] Furthermore, although not particularly limited, for example, the treatment information providing server 500 may be managed by a pharmaceutical manufacturer that has developed a therapeutic drug for a disease or a medical device manufacturer that has developed a treatment device.
[0441] The hardware configurations of the server 200′, the support information providing device 300a, the classifier generating device 300b, and the computer 400 are basically the same as the configuration of the “data processing unit 32” described in FIG. 5, and therefore the description thereof will not be repeated.
[0442] Returning to Figure 39, the correlation matrix calculation unit 3002, correlation value correction processing unit 3004, disease classifier generation unit 3008, clustering classifier generation unit 3010, and discriminant value calculation unit 3012, as well as functional connection correlation matrix data 3106, measurement bias data 3108, corrected correlation value data 3110, and classifier data 3112 are the same as those described in embodiment 1, and therefore their description will not be repeated.
[0443] Here, the disease classifier generation unit 3008, the clustering classifier generation unit 3010, and the storage device 210 constitute a clustering device. The disease classifier generation unit 3008 and the clustering classifier generation unit 3010 are functions executed by the calculation device 2040b of the clustering device.
[0444] The MRI device 410 measures brain activity data of a subject to be stratified by the clustering classifier, and the computer 400 stores the measured MRI measurement data in the non-volatile storage device 4100.
[0445] The computer 400, in accordance with a transmission instruction from a user (for example, a doctor), transmits the measurement data of the subject anonymized by the anonymization processing unit 4042 to the support information providing device 300a as the measurement results of brain activity. At this time, the measurement data to be transmitted is assigned a temporary ID for identifying the subject. Preferably, a correspondence table between temporary IDs and personal information such as patient names can be managed in a state inaccessible to the computer 400. Furthermore, the support information providing device 300a can be configured so that the contents of the correspondence table cannot be accessed at all, at least.
[0446] In response to this, the support information providing device 300a preferably executes a harmonization process corresponding to the site where the MRI device 410 is installed, as necessary. Furthermore, the discriminant value calculation unit 3012 outputs a classification result, which is the result of stratification by the clustering classifier, for the input measurement results of the subject's brain activity. In accordance with the classification result, the treatment information output unit 3200 transmits the classification result and corresponding information for supporting treatment selection to the computer 400 via the communication interface.
[0447] The information presenting unit 4044 of the computer 400 presents the information for supporting the selection of a treatment method for a specific subject and the classification results returned from the treatment method information generating unit 3200 of the support information providing device 300a to the doctor on a display device such as a display (not shown). FIG. 40 is a diagram showing an example of the treatment information database 5100.
[0448] The discriminant value calculation unit 3012 is assumed to be able to classify subjects into, for example, clusters 1 to 5 using a clustering classifier. The treatment method information database 5100 stores predetermined treatment method information related to the treatment method according to each cluster.
[0449] The treatment information may store information on the past treatment history of the subjects classified into each cluster (particularly, subjects belonging to the patient group whose data was used to generate the clustering classifier), and treatment information to be referenced (reference treatment information) based on effects and / or side effects reported in literature, etc. Preferably, information indicating the responsiveness of each cluster to a specific therapeutic drug, information indicating the responsiveness to a specific physical treatment, etc. may be stored.
[0450] In Figure 40, recommended treatment methods in the information on treatment history are stored for each cluster in order of effectiveness, from first to third candidate. In addition, non-recommended treatment methods that may cause side effects, etc., are also stored for each cluster as reference treatment information.
[0451] Drugs for treating depression include, but are not limited to, tricyclic antidepressants such as amitriptyline hydrochloride, amoxapine, and imipramine hydrochloride; tetracyclic antidepressants such as setiptiline maleate, maprotiline hydrochloride, and mianserin hydrochloride; selective serotonin reuptake inhibitors such as escitalopram oxalate, sertraline hydrochloride, paroxetine hydrochloride hydrate, and fluvoxamine maleate; serotonin / noradrenaline reuptake inhibitors such as duloxetine hydrochloride, venlafaxine hydrochloride, and milnacipran hydrochloride; and noradrenergic / specific serotonergic agents such as mirtazapine. Physical treatments for depression include transcranial magnetic stimulation, neurofeedback, electroconvulsive therapy, and cognitive behavioral therapy.
[0452] (Treatment selection support processing) Next, a description will be given of a treatment method selection support process performed by the treatment method information generating unit 3200 of the support information providing device 300a. The treatment method selection support process is achieved by executing a treatment method selection support program as processing by the treatment method information generating unit 3200 on a computer. FIG. 41 is a flowchart illustrating the flow of the support process for selecting a treatment method for a subject.
[0453] 41, the support information providing device 300a receives the measurement results of the subject's brain activity from the computer 400. The measurement results of the subject's brain activity are imaging data (brain structure image data and brain function image data) that have been anonymized in the computer 400.
[0454] In step S504 shown in FIG. 41, the support information providing device 300a calculates the data of the correlation matrix of functional connectivity and performs a harmonization process to correct the correlation values for the measurement results of the subject's brain activity, and then inputs the results into the clustering classifier generated by the classifier generation device 300b to obtain the cluster belonging probability, which is the result of stratification of the subject. The cluster belonging probability is output as the probability that the subject belongs to each cluster for all clusters that can be stratified by the clustering classifier. The support information providing device 300a can determine the cluster with the highest probability as the cluster to which the subject belongs. Alternatively, it may determine the top two clusters with the highest probability as the clusters to which the subject may belong. In this case, the support information providing device 300a functions as a clustering calculation device.
[0455] 41, the support information providing device 300a acquires treatment method information corresponding to the subject's cluster from the treatment method information database 5100 based on the cluster resulting from the stratification of the subject acquired in step S504. At this time, the support information providing device 300a may acquire at least two treatment methods, i.e., a first candidate and a second candidate, for the treatment method information corresponding to the subject's cluster. In this way, it is possible to widen the treatment options.
[0456] In step S508 shown in FIG. 41, the support information providing device 300a outputs the treatment method information acquired in step S506 to the computer 400 via the communication interface.
[0457] With the above configuration, the support information providing device 300a can provide a doctor with information to support the selection of a treatment method for the subject, based on the measurement data of the subject's brain activity.
[0458] [Screening support system, screening support device] The following describes an aspect in which the support information providing device 300a described in FIG. 39 is used as a support device for screening subjects in drug discovery.
[0459] In this specification, "therapy" means "prescribing a specific drug and administering it to a subject at a prescribed dose and method" by a physician, or "selecting a specific treatment process and implementing treatment with that therapy," and "candidate therapy" means a "candidate therapy" that has not yet received approval or certification from a regulatory agency based on a specific clinical trial, etc. If the target disease is a mental illness, "therapy process" means, for example, cognitive behavioral therapy. FIG. 42 is a diagram showing a general drug discovery process.
[0460] As shown in Figure 42, in general, targets are searched for according to the purpose of drug discovery, and candidate substances are screened, optimized, and their efficacy, safety, and pharmacokinetics are examined through animal experiments and in vitro tests in which candidate drug substances are administered to cells cultured in test tubes and the reaction is measured (non-clinical trials), and industrialization is also considered. After going through the above process, candidate substances that pass non-clinical trials are then subjected to so-called "clinical trials."
[0461] "Clinical trials" refer to clinical trials conducted to obtain approval under the Pharmaceutical and Medical Device Act for the manufacture and sale of pharmaceuticals. In the case of pharmaceuticals, clinical trials are often conducted in three phases, from Phase I to Phase III. Phase I trials target healthy adult volunteers. Phase II trials, based on the results of Phase I, target a small number of relatively mild patients to examine efficacy, safety, pharmacokinetics, etc. Phase III trials are conducted on a larger scale, targeting patients who will actually use the compound after it is released to the market, with the primary purpose of verifying efficacy and examining safety.
[0462] The problem with drug discovery for the neuropsychiatric system is that, at present, it is difficult to predict the efficacy of candidate therapeutic substances in humans using animal models, and the probability of ultimately bringing a substance to market after Phase I trials is lower than for candidate therapeutic substances for other diseases.
[0463] One solution to this problem is to identify an appropriate subject population after Phase I (or Phase II) trials in order to efficiently advance potentially effective candidate compounds to Phase III trials. In the following, a case will be described in which the support information providing device 300a is used as a support device for identifying such a subject group (subject screening). That is, the support information providing device 300a is configured to return information about the classified clusters to the computer 400 as screening support data.
[0464] As shown in Figure 40, physical therapy, such as rTMS, is also known as a treatment for neuropsychiatric disorders rather than a therapeutic drug. Medical devices and medical device programs used in such physical therapy must undergo clinical trials, just like drug discovery candidate substances, before being approved as medical devices by regulatory authorities. Even in such clinical trials, subject screening, as described below, may be effective.
[0465] Therefore, the screening support process described in this embodiment can be applied not only to clinical trials of "candidate therapies" as described above, but also to clinical trials of "candidate treatment methods." Here, in this specification, "treatment method" refers to "devices used by doctors to implement a specific treatment" or "programs used to implement a specific treatment (or media on which the programs are recorded, or devices on which the programs are installed)," and "candidate treatment method" refers to "treatment devices" or "programs (or media on which the programs are recorded or devices on which the programs are installed)" that have not yet been approved or certified by a regulatory agency based on a specific clinical trial, etc. For example, if the target disease is a psychiatric disorder, the "treatment method (or treatment device)" refers to a TMS device for implementing transcranial magnetic stimulation therapy, a pulse wave therapy device used for electroconvulsive therapy, a smartphone application program for supporting cognitive behavioral therapy, or a recording medium on which such application programs are recorded or a smartphone on which such application programs are installed, all of which are used in a predetermined manner.
[0466] (Operation as a screening support system) FIG. 44 is a diagram illustrating the configuration of a screening support device 1000'. Similar to the treatment method selection support system 1000, the screening support device 1000' includes a computer 400 connected to a medical institution 4 so as to be able to receive data from an MRI device 410 for measuring brain activity data of a subject for whom a diagnostic label is to be estimated, a support information providing device 300a, a classifier generation device 300b that performs the process of generating a clustering classifier, and a data server 200'.
[0467] In the configuration of the screening support device 1000', parts that are the same as those in the configuration of the treatment selection support system 1000 shown in Figure 39 are given the same symbols, and their descriptions will not be repeated. Below, the main operations of the screening support device 1000' will be described.
[0468] That is, the MRI device 410 measures brain activity data of a subject to be stratified by the clustering classifier, and the computer 400 stores the measured MRI measurement data in the nonvolatile storage device 4100.
[0469] The computer 400, in accordance with a transmission instruction from a user (for example, a doctor), transmits the measurement data of the subject anonymized by the anonymization processing unit 4042 to the support information providing device 300a as the measurement results of brain activity. At this time, the measurement data to be transmitted is assigned a temporary ID for identifying the subject. Preferably, a correspondence table between temporary IDs and personal information such as patient names can be managed in a state inaccessible to the computer 400. Furthermore, the support information providing device 300a can be configured so that the contents of the correspondence table cannot be accessed at all, at least.
[0470] In response to this, the support information providing device 300a preferably executes a harmonization process corresponding to the site where the MRI device 410 is installed, as necessary. Furthermore, the discriminant value calculation unit 3012 outputs a classification result, which is the result of stratification by the clustering classifier, for the input measurement results of the subject's brain activity. This classification result is associated with a temporary ID that identifies the subject and stored in, for example, the storage device 2080'-1. The screening information output unit 3202 generates information to support screening based on the classification result in accordance with the classification result, and transmits the information to the computer 400 via the communication interface.
[0471] The information presenting unit 4044 of the computer 400 presents to the doctor on a display device such as a display (not shown) information for supporting screening for a specific subject returned from the screening information output unit 3202 of the support information providing device 300a.
[0472] Here, the information returned from the screening information output unit 3202 to the computer 400 is referred to as "information to support screening" because, for example, when using the results of screening to conduct a clinical trial or a clinical trial, such as a so-called "double-blind trial" or "randomized controlled trial," the stratification results themselves will not be returned to the medical institution, but will be displayed on the computer as information to support how to conduct screening depending on the characteristics of these trials. When evaluating the test results, such as after the test is completed, the test results can be compared with the classification results for the subjects based on the temporary ID.
[0473] (Screening Support Processing in Support Information Providing Device 300a) Next, a description will be given of the screening support process performed by the arithmetic unit 2040a of the support information providing device 300a. The screening support process is achieved by executing a screening support program on a computer. FIG. 43 is a flowchart illustrating the flow of a process for supporting screening of subjects in drug discovery.
[0474] 43, the support information providing device 300a receives the measurement results of the subject's brain activity from the computer 400. The measurement results of the subject's brain activity are imaging data (brain structure image data and brain function image data) that have been anonymized in the computer 400.
[0475] In step S604 shown in Fig. 43, the support information providing device 300a calculates data on the correlation matrix of functional connectivity and performs processing to correct correlation values by harmonization processing on the measurement results of the subject's brain activity, and then inputs the results into the clustering classifier generated by the classifier generating device 300b to obtain cluster belonging probabilities, which are the results of stratification of the subject. The cluster belonging probabilities are output as the probability that the subject belongs to each cluster for all clusters that can be stratified by the clustering classifier. The support information providing device 300a can determine the cluster with the highest probability as the cluster to which the subject belongs.
[0476] In step S606 shown in FIG. 43, the support information providing apparatus 300a generates information indicating clusters that are the results of stratification of the subjects acquired in step S604.
[0477] In step S608 shown in FIG. 43, the support information providing device 300a outputs the cluster information or screening support information generated in step S606 to the computer 400 via the communication interface.
[0478] With the above configuration, the support information providing device 300a can provide a doctor with information on the cluster to which a subject belongs during the clinical trial process, based on the measurement data of the subject's brain activity. A physician can then screen the subjects in the corresponding cluster according to a predetermined clinical trial protocol and conduct the clinical trial.
[0479] [Verification data for stratification of patients with depression] The treatment method selection support system 1000 has been described above with reference to FIG. 39, and the screening support device 1000' has been described with reference to FIGS.
[0480] In the following, an example of data for verifying the clinical significance of the stratification results calculated by the discriminant value calculation unit 3012 using a clustering classifier for the input measurement results of the subject's brain activity in the embodiment described above will be described.
[0481] (Clustering classifier used for this validation data) The configuration of the MDD classifier generated by the disease classifier generation unit 3008 when generating the clustering classifier used for this verification data is as follows. First, the brain regions were divided according to the BSA method (Brainvisa Sulci Atlas) as a parcellation method. For example, the BSA method is disclosed in the following document:
[0482] Publication: Matthieu Perro, Denis Riviere, and Jean-Francois Mangin a, Cortical sulci recognition and spatial normalization. Medical Image Analysis Volume 15, Issue 4, August 2011, Pages 529-550
[0483] Next, we calculated the Pearson correlation coefficient of fMRI signals as an index of connectivity between brain regions, and used the random forest method based on whole-brain functional connectivity to train and generate an MDD classifier using the discovery cohort data. In this case, we did not specifically apply inter-center correction using the traveling subject method. The clustering classifier generation process in the clustering classifier generation unit 3010 is as follows.
[0484] First, in the process of training the MDD classifiers, nested cross validation is performed, resulting in the generation of 100 MDD classifiers.
[0485] For each classifier, the importance of each brain function connection in classification is determined using the random forest method, and the overall importance is calculated by summing the scores of the 100 classifiers.
[0486] Based on this overall importance, the brain function connections are ranked, and multiple co-clustering is performed using the top connections. Here, although not limited to, the number of top connections to be used can be determined by performing clustering on multiple numbers of connections in a predetermined pattern and adopting the number with the highest stability between data sets. Here, the ARI value explained in FIG. 33 can be used as an evaluation index to evaluate the "stability between data sets."
[0487] FIG. 45 shows a view of the clustering results by multiple co-clustering using the top 30 brain function connections of importance in the MDD classifier for dataset 1. FIG. 46 shows a view of the clustering results with multiple co-clustering for dataset 2. FIG. 47 shows the clustering stability between Dataset 1 and Dataset 2. FIG. 47 is a table in which the ARI for each view is calculated for Dataset 1 and Dataset 2, and corresponds to the table shown in FIG.
[0488] Referring to FIG. 47, view 1 of dataset 1 and view 1 of dataset 2, and view 2 of dataset 1 and view 3 of dataset 2 have ARI values of 0.67 and 0.68, respectively, and can be said to be similar to each other. FIG. 48 shows a view of the clustering results for all depression patient data (dataset 1+2).
[0489] FIG. 49 is a diagram showing the number of all depression patients and the number of depression patients for which clinical data exists, for each subtype, for view 3 generated by clustering dataset 1+2. Here, "clinical data" refers to "medication history information" for each patient, "diagnosis information on the degree of depression" for each patient, and so on.
[0490] Referring to Figure 49, the patient data for which clinical data exists is only a portion of the whole (part of both Data Set 1 and Data Set 2), and even if only a portion of these patients is extracted, there is no significant difference in the distribution compared to when the whole data is clustered. Therefore, it can be seen that there is no problem in examining the clinical significance of clustering only for "patient data for which clinical data exists."
[0491] Below, we will further explain view 3 shown in Figures 48 and 49 as a view in which distinctive clinical significance, specifically, differences in therapeutic response between subtypes, were observed.
[0492] Furthermore, it was also confirmed that the overall mean and standard deviation of brain functional connectivity characteristics of each subtype (mean and standard deviation of connectivity) did not change significantly even when selecting only a subset of patients with clinical data. FIG. 50 shows the brain function connections (view 3) used for clustering. Figure 50(a) shows the locations in the brain of the brain functional connections used for clustering, and Figure 50(b) shows the regions of interest for the functional connections. In view 3, only the three connections shown in Figures 50(a) and (b) are used for clustering.
[0493] The top 30 combinations are selected based on their importance in the MDD classifier used to generate the clustering classifier, and multiple co-clustering is performed using these combinations.
[0494] In other words, clustering was performed on the 30 selected connections without weighting them. As a result, three views were generated. In other words, the 30 connections were divided into three groups (View 1: 22 connections, View 2: 5 connections, View 3: 3 connections), and the subjects were clustered based on each connection group. In multiple co-clustering, the algorithm automatically finds the optimal solution for how to divide connections and how to cluster patients (subjects). The names of the regions of interest in FIG. 50(b) indicate the following: Thalamus_LorR: Left or right thalamus Precentral_R: Right precentral gyrus Postcentral_L: Left postcentral gyrus FIG. 51 is a diagram showing the relationship between the severity of depression and the rate of improvement in severity for each subtype of view 3 shown in FIGS. FIG. 51(a) is a graph comparing HAMD scores at week 0 and week 6 after the start of treatment.
[0495] Here, HAMD stands for the Hamilton Rating Scale for Depression, a scale used to assess the severity of depression. The main 17-item version, which consists of 17 items that indicate the severity of depression, and the 21-item version, which adds four additional items to this, are commonly used. 51(a), subtypes 1 to 5 in FIG. 49 are the top five clusters in the "Results of subject clustering (view 3)" in FIG.
[0496] The subtype numbers are arranged in descending order of the number of subjects for whom clinical data are available. Furthermore, because the analysis targeted subtypes with data for 10 or more subjects, the graph shows subtypes 1, 2, 4, and 5 out of subtypes 1 to 5.
[0497] Week 0 after treatment initiation refers to the time of study entry and the start of SSRI treatment, and week 6 after treatment initiation refers to 6 weeks after study entry and the start of treatment, excluding subjects who started SSRI treatment before study entry.
[0498] Here, SSRI stands for "selective serotonin reuptake inhibitor" and includes, for example, escitalopram. The treatment protocol in this study is as follows:
[0499] Patients with major depression who have not received treatment or who have received insufficient amounts of treatment for an insufficient period will be enrolled in the study, and treatment will begin with SSRIs such as escitalopram from the time of study entry. The dose can be increased at clinical discretion, there are no restrictions on concomitant medications, and if the patient does not improve with SSRIs, another drug therapy will be implemented. FIG. 51(b) is a diagram showing the improvement rate of HAMD when comparing week 0 and week 6. Here, the improvement rate is expressed by the following formula. (HAMD(0) - HAMD(6W)) / HAMD(0) x 100 Figure 51(a) shows that in the early stages, there is no difference in HAMD between subtypes, meaning that the severity of depression is about the same.
[0500] Figure 51(a) shows that differences in HAMD were observed between subtypes 6 weeks after the start of SSRI treatment. Figure 51(b) shows that, as can be seen from the improvement rate from the initial value, subtype 1 is less effective with SSRI treatment than subtypes 2 and 5, while subtypes 2 and 5 are more effective with SSRI treatment than subtype 1.
[0501] From the above explanation, it can be seen that by using the clustering classifier described in the embodiment, it is possible to predict the therapeutic response of patients in the early stages of treatment to a drug called SSRI for each cluster.
[0502] The embodiments disclosed herein are merely examples of configurations for specifically implementing the present invention, and do not limit the technical scope of the present invention. The technical scope of the present invention is defined by the claims, not by the description of the embodiments, and is intended to include modifications within the literal scope of the claims and within the scope of equivalent meanings. [Explanation of symbols]
[0503] 2 Subject, 6 Display, 10 MRI device, 11 Magnetic field application mechanism, 12 Static magnetic field generating coil, 14 Gradient magnetic field generating coil, 16 RF irradiation unit, 18 Bed, 20 Receiving coil, 21 Drive unit, 22 Static magnetic field power supply, 24 Gradient magnetic field power supply, 26 Signal transmission unit, 28 Signal reception unit, 30 Bed drive unit, 32 Data processing unit, 36 Memory unit, 38 Display unit, 40 Input unit, 42 Control unit, 44 Interface unit, 46 Data collection unit, 48 Image processing unit, 50 Network interface, 300a Support information providing device, 300b Classifier generation device, 500 Treatment method information providing server.
Claims
1. 1. A treatment selection support system for providing information regarding selection of a treatment for a first subject exhibiting a depressive symptom, based on a measurement result of brain activity of the first subject, the system comprising: a clustering device for performing stratification into a plurality of clusters by clustering processing on measurement results of brain functional connectivity correlation values obtained from a plurality of second subjects, wherein the plurality of second subjects include a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; The clustering device includes: a computing device and a storage device for executing the clustering process on the plurality of second subjects, The arithmetic device, in the process of generating the clustering classifier, i) storing, in the storage device, feature quantities based on a plurality of brain function connectivity correlation values representing temporal correlations of brain activities between a plurality of predetermined pairs of brain areas for each of the plurality of second subjects; ii) performing supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the feature amounts stored in the storage device; iii) in the machine learning for generating the classifier model, selecting features for clustering according to the importance of the features used in generating the classifier by machine learning; iv) clustering the first group by an unsupervised learning multiple co-clustering method based on the selected features for clustering to generate a clustering classifier; The treatment method selection support system further includes: a database device for storing clusters resulting from stratification by the clustering classifier in association with corresponding predetermined treatment information; a support information providing device that receives as input the measurement results of the brain activity of the first subject and outputs corresponding treatment information according to the classification results of the measurement results by the clustering classifier.
2. The computing device, in the machine learning for generating the classifier model, performing undersampling and subsampling from the first set and the second set to generate a plurality of training subsamples; selecting, for each of the training subsamples, features for clustering from a union of features used in generating a classifier by machine learning, according to the importance of the features belonging to the union; The treatment method selection support system according to claim 1 , wherein the clustering classifier is generated by the multiple co-clustering method based on the selected features for clustering.
3. the support information providing device includes a clustering calculation device and an interface device; the clustering calculation device calculates the probability that the first subject belongs to each of the clusters using the clustering classifier, and reads out at least two pieces of treatment information selected according to the probability from the database device; 3. The treatment selection support system according to claim 1, wherein the interface device outputs data for displaying the selected clusters in association with the corresponding treatment information.
4. 4. The treatment method selection support system according to claim 1, wherein the treatment method information is information indicating responsiveness to a specific therapeutic agent.
5. 4. The treatment method selection support system according to claim 1, wherein the treatment method information is information indicating responsiveness to a specific physical treatment method.
6. 3. The treatment method selection support system according to claim 2, wherein the process of generating the classifier by machine learning is ensemble learning that generates a plurality of classifier sub-models for each of the plurality of training sub-samples and integrates the plurality of classifier sub-models to generate the classifier model.
7. the clustering device receives information representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas of each of the plurality of second subjects from a plurality of brain activity measuring devices respectively provided at a plurality of measurement sites; 3. The treatment method selection support system according to claim 1, wherein the arithmetic unit includes a harmonization calculation unit that corrects the plurality of brain function connectivity correlation values for each of the plurality of second subjects to remove measurement bias at the measurement site, and stores the corrected adjusted values as the feature quantities in the storage device.
8. 1. A treatment selection support device for providing information regarding selection of a treatment for a first subject exhibiting a depressive symptom, based on a measurement result of brain activity of the first subject, the device comprising: a database device for storing clusters resulting from stratification of subjects with a diagnostic label of depression among the plurality of second subjects in association with corresponding predetermined treatment information; a support information providing device that receives as input a measurement result of the brain activity of the first subject, and outputs corresponding treatment information according to a result of stratification based on the measurement result; the plurality of second subjects includes a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; the clusters resulting from the stratification are obtained by a clustering classifier obtained by clustering processing of the measurement results of the brain function connectivity correlation values by a clustering device, The clustering device includes: a computing device and a storage device for executing the clustering process on the first group; In the process of generating the clustering classifier, the arithmetic device i) storing, in the storage device, feature quantities based on a plurality of brain function connectivity correlation values representing temporal correlations of brain activities between a plurality of predetermined pairs of brain areas for each of the plurality of second subjects; ii) performing supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the feature amounts stored in the storage device; iii) in the machine learning for generating the classifier model, selecting features for clustering according to the importance of the features used in generating the classifier by machine learning; iv) A treatment method selection support device that generates the clustering classifier by clustering the first group using an unsupervised learning multiple co-clustering method based on the selected features for clustering.
9. the support information providing device includes a clustering calculation device and an interface device; the clustering calculation device calculates the probability that the first subject belongs to each of the clusters using the clustering classifier, and reads out at least two pieces of treatment information selected according to the probability from the database device; 9. The treatment selection support device according to claim 8, wherein said interface device outputs data for displaying said selected clusters in association with said corresponding treatment information.
10. 10. The treatment method selection support device according to claim 8, wherein the treatment method information is information indicating responsiveness to a specific treatment method.
11. 11. The treatment method selection support device according to claim 8, wherein the treatment method information is information indicating responsiveness to a specific physical treatment method.
12. 1. A treatment selection support method for providing information regarding selection of a treatment for a first subject exhibiting a depressive symptom, based on a measurement result of brain activity of the first subject, the method comprising: generating and preparing a clustering classifier for performing stratification into a plurality of clusters by clustering processing on measurement results of brain functional connectivity correlation values obtained from a plurality of second subjects, wherein the plurality of second subjects include a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; The preparing step includes: a calculation step for performing the clustering process on the plurality of second subjects, The calculation step includes: i) acquiring, for each of the second subjects, feature quantities based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) performing machine learning using supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the acquired feature amount; iii) in machine learning for generating the classifier model, selecting features for clustering according to the importance of features used in generating the classifier by machine learning; iv) clustering the first group by a multiple co-clustering method of unsupervised learning based on the selected features for clustering to generate the clustering classifier; The treatment method further includes: a support information providing step of acquiring and outputting corresponding treatment information from a database that stores clusters resulting from stratification by the clustering classifier in association with corresponding predetermined treatment information, in accordance with the classification result by the clustering classifier of the measurement results of the brain activity of the first subject.
13. 1. A treatment selection support method for providing information regarding selection of a treatment for a first subject exhibiting a depressive symptom, based on a measurement result of brain activity of the first subject, the method comprising: a support information providing step of acquiring corresponding treatment method information from a database that stores, in association with each other, results of stratification for subjects with a diagnostic label of depression among a plurality of second subjects, in accordance with a cluster of the results of stratification based on the measurement results of brain activity of the first subject, and outputting the corresponding treatment method information; the plurality of second subjects includes a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; the clusters resulting from the stratification are obtained by a clustering classifier obtained by a clustering process on the measurement results of the brain function connectivity correlation values, The clustering classifier The clustering process is performed on the plurality of second subjects by a calculation step, the calculation step including: i) acquiring, for each of the second subjects, feature quantities based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) performing machine learning through supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the acquired feature amount; iii) in machine learning for generating the classifier model, selecting features for clustering according to the importance of features used in generating the classifier by machine learning; iv) clustering the first group by an unsupervised learning multiple co-clustering method based on the selected features for clustering to generate the clustering classifier.
14. 1. A treatment selection support program for providing information regarding selection of a treatment for a first subject exhibiting a depressive symptom, based on a measurement result of brain activity of the first subject, the program comprising: When the treatment method selection support program is executed by a computer, the program causes the computer to: generating a clustering classifier for performing stratification into a plurality of clusters by clustering processing on the measurement results of brain function connectivity correlation values obtained from a plurality of second subjects; receiving as input the measurement results of the brain activity of the first subject, and, in accordance with the classification results of the measurement results by the clustering classifier, acquiring and outputting corresponding treatment method information from a database device that stores clusters resulting from stratification by the clustering classifier in association with corresponding predetermined treatment method information; the plurality of second subjects includes a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; The clustering process includes: The method includes a calculation step of performing a clustering process on the plurality of second subjects, the calculation step including: i) storing, for each of the plurality of second subjects, feature quantities based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a plurality of predetermined pairs of brain areas, in a storage device; ii) performing supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the feature amounts stored in the storage device; and iii) in machine learning for generating the classifier model, selecting features for clustering according to the importance of features used in generating the classifier by machine learning; iv) clustering the first group by an unsupervised learning multiple co-clustering method based on the selected features for clustering to generate a clustering classifier.
15. 1. A treatment selection support program for providing information regarding selection of a treatment for a first subject exhibiting a depressive symptom based on a measurement result of brain activity of the first subject, the program comprising: When the treatment method selection support program is executed by a computer, causing the computer to execute a support information providing step of acquiring corresponding treatment method information from a database that stores, in association with each other, results of stratification for subjects with a diagnostic label of depression among a plurality of second subjects, in accordance with a cluster of the results of stratification based on the measurement results of the brain activity of the first subject, and outputting the corresponding treatment method information; the clusters of the stratification results are obtained by a clustering classifier obtained by a clustering process on the measurement results of the brain function connectivity correlation values, the plurality of second subjects includes a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; The clustering classifier The clustering process is performed on the plurality of second subjects by a calculation step, the calculation step including: i) acquiring, for each of the second subjects, feature quantities based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) performing machine learning using supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the acquired feature amount; iii) in machine learning for generating the classifier model, selecting features for clustering according to the importance of features used in generating the classifier by machine learning; iv) clustering the first group by an unsupervised learning multiple co-clustering method based on the selected features for clustering to generate the clustering classifier.
16. 1. A screening support system for supporting screening of a first subject based on a measurement result of brain activity of the first subject in a clinical trial of a candidate treatment for a depressive symptom, comprising: a clustering device for performing stratification into a plurality of clusters by clustering processing on measurement results of brain functional connectivity correlation values obtained from a plurality of second subjects, wherein the plurality of second subjects include a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; The clustering device includes: a computing device and a storage device for executing the clustering process on the plurality of second subjects, wherein the computing device i) storing, in the storage device, feature quantities based on a plurality of brain function connectivity correlation values representing temporal correlations of brain activities between a plurality of predetermined pairs of brain areas for each of the plurality of second subjects; ii) performing supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the feature amounts stored in the storage device; iii) in the machine learning for generating the classifier model, selecting features for clustering according to the importance of the features used in generating the classifier by machine learning; iv) clustering the first group by an unsupervised learning multiple co-clustering method based on the selected features for clustering to generate a clustering classifier; The screening support system further comprises: a support information providing device that receives as input measurement results of the brain activity of the first subject, records classification results of the measurement results by the clustering classifier in association with the first subject, and outputs information to support screening of the first subject based on the classification results.
17. The computing device, in the machine learning for generating the classifier model, performing undersampling and subsampling from the first set and the second set to generate a plurality of training subsamples; selecting, for each of the training subsamples, features for clustering from a union of features used in generating a classifier by machine learning, according to the importance of the features belonging to the union; The screening support system according to claim 16, wherein the clustering classifier is generated by the multiple co-clustering method based on the selected features for clustering.
18. 1. A screening support device for supporting screening of a first subject based on a measurement result of brain activity of the first subject in a clinical trial of a candidate treatment for a depressive symptom, comprising: a support information providing device having a storage device for storing information for specifying a clustering classifier, the support information providing device receiving as input a measurement result of brain activity of the first subject, recording a classification result based on the clustering classifier for the measurement result in association with the first subject, and outputting information for supporting screening of the first subject based on the classification result; the clusters resulting from the stratification are obtained by the clustering classifier obtained by clustering processing of measurement results of brain functional connectivity correlation values obtained from a plurality of second subjects by a clustering device, the plurality of second subjects including a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression; The clustering device includes: a computing device and a storage device for executing the clustering process on the plurality of second subjects, In the process of generating the clustering classifier, the arithmetic device i) storing, in the storage device, feature quantities based on a plurality of brain function connectivity correlation values representing temporal correlations of brain activities between a plurality of predetermined pairs of brain areas for each of the plurality of second subjects; ii) performing supervised machine learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the feature amount stored in the storage device; iii) in the machine learning for generating the classifier model, selecting features for clustering according to the importance of the features used in generating the classifier by machine learning; iv) A screening support device that generates the clustering classifier by clustering the first group using an unsupervised learning multiple co-clustering method based on the selected features for clustering.
19. 1. A screening support method for supporting screening of a first subject based on a measurement result of brain activity of the first subject in a clinical trial of a candidate treatment for a depressive symptom, comprising: a step of causing a computing device to classify the first subject based on the measurement results of the brain activity, based on a clustering classifier specified by information stored in a storage device; and recording the classification result in association with the first subject, and outputting information to support screening of the first subject based on the classification result; The process for generating the clustering classifier includes: The method includes a calculation step of performing a clustering process on a plurality of second subjects including a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression, the calculation step including: i) acquiring, for each of the second subjects, feature quantities based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a predetermined plurality of pairs of brain areas; ii) performing machine learning through supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the acquired feature amount; iii) in machine learning for generating the classifier model, selecting features for clustering according to the importance of features used in generating the classifier by machine learning; iv) clustering the first group by an unsupervised learning multiple co-clustering method based on the selected features for clustering to generate a clustering classifier.
20. 1. A screening support program for supporting screening of a first subject based on a measurement result of brain activity of the first subject in a clinical trial of a candidate treatment for depressive symptoms, comprising: When the screening support program is executed by a computer, the computer a step of causing a computing device to classify the first subject based on the measurement results of the brain activity, based on a clustering classifier specified by information stored in a storage device; recording the classification result in association with the first subject, and outputting information to support screening of the first subject based on the classification result; Execute The process for generating the clustering classifier includes: The method includes a calculation step of performing a clustering process on a plurality of second subjects including a first group having a diagnostic label of depression and a second group not having the diagnostic label of depression, wherein the calculation step includes: i) storing, for each of the plurality of second subjects, feature quantities based on a plurality of brain function connectivity correlation values, each representing a temporal correlation of brain activity between a plurality of predetermined pairs of brain areas, in a storage device; ii) performing machine learning using supervised learning to generate a classifier model for determining the presence or absence of the diagnostic label based on the feature amounts stored in the storage device; iii) in machine learning for generating the classifier model, selecting features for clustering according to the importance of features used in generating the classifier by machine learning; iv) a step of generating a clustering classifier by clustering the first group using an unsupervised learning multiple co-clustering method based on the selected features for clustering.
21. 2. The treatment selection support system according to claim 1, wherein the predetermined treatment information is information regarding treatment response to a selective serotonin reuptake inhibitor.
22. 9. The treatment selection support device according to claim 8, wherein the predetermined treatment information is information relating to treatment response to a selective serotonin reuptake inhibitor.
23. 14. The method for supporting selection of a treatment method according to claim 12, wherein the predetermined treatment information is information relating to therapeutic response to a selective serotonin reuptake inhibitor.
24. 16. The treatment selection support program according to claim 14, wherein the predetermined treatment information is information regarding therapeutic response to a selective serotonin reuptake inhibitor.
25. The screening support system according to claim 16, wherein the candidate treatment is a treatment using a selective serotonin reuptake inhibitor.
26. 19. The screening support device according to claim 18, wherein the candidate treatment is a treatment using a selective serotonin reuptake inhibitor.
27. 20. The screening support method according to claim 19, wherein the candidate treatment is a treatment using a selective serotonin reuptake inhibitor.
28. The screening support program according to claim 20 , wherein the candidate treatment is a treatment using a selective serotonin reuptake inhibitor.
Citation Information
Patent Citations
Brain activity analysis device, brain activity analysis method, and biomarker device
JP2015062817A
Brain activity analyzer, brain activity analysis method, program, and biomarker device
JP2017196523A
Discrimination device, discrimination method for depressive symptom, determination method for depressive symptom level, stratification method for patient with depression, determination method for treatment effect on depressive symptom, and brain activity training device
JP2019063478A
Medical image processor, medical image processing method, and medical image processing system
JP2019198376A
Analysis method and array used therefor
JP2019516950A