Cluster device for brain function connection related values, clustering system, clustering method, brain activity marker classification system, and program product
By performing feature quantity selection and multiple coclustering processing in the computing processing system, a clustering model that can be used to connect the relevant values of brain function common among multiple facilities is generated, which solves the problem of insufficient diagnostic accuracy and universality in the prior art, and realizes efficient classification of subtypes of mental illness.
Patent Information
- Application Number
- CN202180027512.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-06
- Filing Date
- 2021-04-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-04-02
AI Technical Summary
The prior art is difficult to apply the clustering of brain function connection correlation values in general between multiple facilities, resulting in insufficient diagnostic accuracy and universality, especially in subtype classification of mental illnesses.
By performing feature quantity selection and multiple coclustering processing based on brain activity measurement data, the computing processing system generates a clustering model that can connect the relevant values of brain functions common among multiple facilities. The model includes a coordinated computing unit to correct measurement bias and generate a recognizer model through integrated learning, selecting important feature quantities for clustering.
The clustering of brain function connection correlation values common among multiple facilities is realized, which improves the diagnostic accuracy and universality of mental illness subtype classification, and solves the problems of small sample size and overfitting.
Smart Images

Figure CN115484864B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for clustering patterns of brain function connection related values measured by a brain function imaging method in multiple devices. More specifically, the present invention relates to a clustering device for brain function connection related values, a clustering system for brain function connection related values, a clustering method for brain function connection related values, a classifier program for brain function connection related values, a brain activity marker classification system, and a clustering classifier model for brain function connection related values. This application claims the priority of Japanese Patent Application No. 2020-068669 filed on April 6, 2020, and incorporates all the descriptions recorded in the Japanese application by reference into this application. Background Art
[0002] (Data-driven clustering method)
[0003] With the development of artificial intelligence technologies in recent years, especially data-driven artificial intelligence technologies, applications comparable to human capabilities have been partially realized in fields such as speech recognition, translation, and image recognition, or applications that exceed human capabilities have been realized in some fields (for example, Patent Document 1).
[0004] In the field of medical technology, the use of machine learning such as deep learning in image diagnosis and the like has also increased. Deep learning is machine learning using a multi-layer neural network. In the field of image recognition, it is known that a learning method using a convolutional neural network (Convolutional Neural Network, hereinafter referred to as "CNN"), which is one of deep learning, exhibits very high performance compared to conventional methods (for example, Patent Document 2).
[0005] For example, in image diagnosis using an endoscope for colorectal cancer, a diagnostic device with a diagnostic accuracy exceeding that of human diagnosis has been put into practical use (Non-Patent Document 1).
[0006] However, almost all of these artificial intelligence technologies fall into the category of so-called "supervised learning" in machine learning classification, that is, a large number of sets of correct answer data and input data (such as image data) are prepared, and the artificial intelligence is made to perform learning processing with this set as the input.
[0007] On the other hand, as an application of data-driven artificial intelligence, there is also a task of classifying the provided data into several clusters based on their feature quantities. In this case, there are known so-called "unsupervised learning" where there is no correct answer data, "semi-supervised learning" that combines learning based on a small amount of "learning data with correct answer labels" and learning based on a large amount of "learning data without correct answer labels", etc. (for example, Patent Document 3).
[0008] For example, Patent Document 3 describes that "semi-supervised learning is a learning method based on relatively little labeled data and unlabeled data. For example, it includes the bootstrap method, graph-based algorithms, etc. The bootstrap method uses labeled data (training data T including state data S and decision data L) to generate a learning model for classification, and uses this learning model and unlabeled data (state data S) to perform additional learning on this learning model, thereby improving the accuracy of learning. The graph-based algorithm groups based on the data distribution of labeled data and unlabeled data to generate a learning model as a classifier". However, as shown in this example, in "semi-supervised learning", the training data available is a small amount of learning data. Thus, the premise is to first generate a classifier and then use a large amount of "learning data without correct labels" to make the classifier itself perform re-learning, etc.
[0009] (Biomarker)
[0010] Next, as an application area of discrimination and clustering based on artificial intelligence technology, the medical field is taken as an example.
[0011] An index obtained by numericalizing / quantifying biological information in order to quantitatively grasp biological changes in an organism is called a "biomarker".
[0012] The FDA (U.S. Food and Drug Administration) defines the positioning of a biomarker as "an item objectively measured / evaluated as an indicator of normal and pathological processes, or of the pharmacological response to treatment". In addition, biomarkers used to characterize the state, changes, and degree of cure of a disease are used as surrogate markers for confirming the effectiveness of a new drug in a clinical trial. Blood glucose level, cholesterol level, etc. are representative biomarkers as indicators of lifestyle diseases. It includes not only substances derived from organisms contained in urine and blood, but also electrocardiograms, blood pressure, PET (positron emission tomography) images, bone density, lung function, etc. In addition, with the development of genome analysis and proteome analysis, various biomarkers associated with DNA, RNA, biological proteins, etc. have been discovered.
[0013] Biomarkers are not only applied to the measurement of the treatment effect after a disease is contracted, but also applied to the prevention of diseases as daily indicators for disease prevention, and are also expected to be applied to personalized medicine for selecting effective treatment methods to avoid side effects.
[0014] For example, for lung diseases, biomarkers for determining the likelihood of developing a disease using genetic information are publicly available (Patent Document 4). In Patent Document 4, a "biomarker" or "marker" refers to "a biological molecule that can be objectively measured as a substance that represents the physiological state of the biological system." Further, in Patent Document 4, it is described that "usually, biomarker measurement values are typically information related to the quantitative measurement of proteins or polypeptides, i.e., expression products. The present invention contemplates determining biomarker measurement values at the RNA (pre-translation) level or the protein level (which may also include post-translational modifications)." Further, in Patent Document 4, examples of classifiers used as a "classification system" for such biomarker measurement values include decision trees, Bayesian classifiers, Bayesian belief networks, k-nearest neighbor methods, case-based reasoning, and support vector machines, etc.
[0015] On the other hand, in the case of neuro / psychiatric diseases, current diagnoses are sometimes based on so-called symptom-based diagnoses such as DSM-5 (Diagnostic and Statistical Manual of Mental Disorders, 5th Edition), etc. Although molecular markers that can be used as objective indicators have been studied from the viewpoints of biochemistry or molecular genetics, the situation is still in what should be called the research stage.
[0016] However, it has also been reported that a disease determination system, etc. that uses NIRS (Near-infraRed Spectroscopy) technology to classify mental diseases such as schizophrenia and depression based on characteristic quantities of hemoglobin signals measured by in vivo optical measurement (Patent Document 5).
[0017] (Biomarkers based on brain activity)
[0018] On the other hand, in the field of so-called image diagnosis, different from the concept of biomarkers such as the above-mentioned "biological molecules," there are also biomarkers called "image biomarkers." For example, attempts have also been made to use PET in molecular imaging of the cranial nerve region to analyze neurotransmission function and receptor function.
[0019] Moreover, in magnetic resonance imaging (MRI: Magnetic Resonance Imaging), it is also possible to use the situation where the detected signal changes in response to changes in blood flow to visualize the active parts of the brain for external stimuli, etc. This magnetic resonance imaging method is particularly called fMRI (functional MRI: functional magnetic resonance imaging).
[0020] In fMRI, as a device, a device is used that is formed by further equipping a general MRI device with the hardware and software required for fMRI measurement.
[0021] Here, the change in blood flow causing a change in the NMR signal intensity makes use of the different magnetic properties of oxyhemoglobin and deoxyhemoglobin in the blood. Oxyhemoglobin has the property of a diamagnetic substance and has no effect on the relaxation time of the hydrogen atoms of the surrounding water. In contrast, deoxyhemoglobin is a paramagnetic substance and causes a change in the surrounding magnetic field. Therefore, when the brain is stimulated and local blood flow increases, resulting in a change in deoxyhemoglobin, the amount of this change can be detected as an MRI signal. Such stimulation of the subject generally uses visual stimulation, auditory stimulation, or the execution of a specified task, etc.
[0022] Moreover, in brain function research, the measurement of brain activity is carried out by measuring the increase in the magnetic resonance signal (MRI signal) of hydrogen atoms corresponding to the phenomenon (BOLD effect) of the decrease in the concentration of deoxyhemoglobin in red blood cells in venules and capillaries.
[0023] In this way, the blood oxygen concentration-dependent signal that reflects the brain activity measured by the fMRI device is called the BOLD signal (Blood Oxygen Level Dependent Signal).
[0024] In particular, in research related to human motor function, the subject is made to perform some movements, and the brain activity is measured by the above fMRI measurement.
[0025] In addition, in the case of humans, non-invasive measurement of brain activity is required. In this case, decoding techniques that can extract more detailed information from fMRI data have gradually developed. In particular, by analyzing brain activity in units of voxels (volumetric pixels) in the brain using fMRI, it is possible to estimate the stimulus input and identify the state based on the spatial pattern of brain activity.
[0026] Also, as a technology that has developed such decoding techniques, a method for analyzing brain activity of "diagnostic biomarkers" for neuro / psychiatric diseases by means of brain functional imaging methods is disclosed in Patent Document 6. In this method, based on the MRI data of resting-state functional connectivity measured for a healthy group and a patient group, a correlation matrix (brain functional connectivity parameter) of the activity levels between specified brain regions is derived for each subject. For the attributes of the subjects including the disease / health labels of the subjects and the correlation matrix, feature extraction is performed by regularized canonical correlation analysis. Based on the results of the regularized canonical correlation analysis, a discriminator that functions as a biomarker is generated by discriminant analysis based on sparse logistic regression (SLR). By such machine learning techniques, it is shown that the diagnostic results of neurological diseases can be predicted based on the connections between brain regions derived from fMRI data of the resting state. Moreover, the verification of the prediction performance shows that it is applicable not only to brain activities measured in one facility, but also to some extent to brain activities measured in other facilities.
[0027] Also, for such "diagnostic biomarkers", technical improvements for further improving the generalization performance have been made (Patent Document 7).
[0028] In addition, recently, obtaining and sharing large-scale brain image data such as the Human Connectome Project in the United States has been considered to be of great significance for filling the gap between basic neuroscience research and clinical applications such as the diagnosis and treatment of mental diseases (Non-Patent Document 2).
[0029] In 2013, the National Institute of Biomedical Innovation, Health and Nutrition of Japan organized the following Decoding Neurofeedback (DecNef) project: 8 research institutes collected resting-state functional magnetic resonance (resting-state fMRI) data at multiple sites including 2,239 samples and 5 diseases, and made these data publicly available and shared through a database of multiple diseases at multiple sites of SRPBS (Strategic Research Program for Brain Sciences (https: / / www.amed.go.jp / program / list / 01 / 04 / 001_nopro.html)) (https: / / bicr-resource.atr.jp / decnefpro / ). This project identified biomarkers based on resting-state functional connectivity (resting-state functional connectivity MRI) for several mental diseases that are applicable to completely independent cohorts.
[0030] In this way, certain achievements have been made in the diagnosis of healthy groups and disease groups. Additionally, it is known that, for example, the patient group generally diagnosed as "depression" in the disease group is actually further divided into multiple subtypes. For example, there is a patient group that can be alleviated by taking ordinary "antidepressants", but there is also a "refractory" patient group that is difficult to alleviate, etc.
[0031] For such "depression" patients, there are also literatures (Non-Patent Literatures 3 and 4) that attempt to classify and show certain tendencies by applying data-driven artificial intelligence clustering to the above-mentioned "brain functional connectivity parameters".
[0032] However, in order to put this method of classifying subtypes of disease groups into practical use, large-scale data of the disease group is required. However, it is not easy to collect brain image data on a large scale, especially for patients.
[0033] Therefore, when measurements are carried out at multiple locations for large-scale data collection, the inter-site differences in measurement data at each measurement location become a problem. In the above Non-Patent Literature 4, it is also mentioned that the "generalization" of clustering for a large amount of measurement data from multiple facilities is a future issue.
[0034] For example, in the above Non-Patent Literature 3, it is pointed out that depressive patients are stratified into 4 subtypes and there are differences in the treatment responsiveness to TMS (transcranial magnetic stimulation). However, in other literatures, it is pointed out that during the process of discovering brain functional connectivity indicators, depressive symptom data was used twice, and due to overlearning, the statistical significance of the correlation with depressive symptoms could not be confirmed, and the stability of the stratification was also poor (Non-Patent Literature 5).
[0035] Therefore, for example, regarding depression, the current situation is that the accuracy confirmation of stratification in independent validation data has not been carried out.
[0036] On the other hand, for example, in order to evaluate the inter-site differences in measurement data when MRI measurements are carried out at multiple measurement locations, attempts have also been made to investigate the effect of measurement bias on resting-state functional connectivity by adopting the so-called "traveling subject" where a large number of participants go to multiple locations for measurement (Non-Patent Literatures 6 and 7).
[0037] In general, in the case of classifying the attributes of a subject based on fMRI data, in machine learning, in order to avoid the problem of overfitting, in most cases, a cross-validation method that removes one subject and is used for validation is employed: leave-one-subject-out cross-validation, or 10-fold cross-validation in which the data is divided into 10 parts, nine-tenths is used for learning, and the remaining one-tenth is used for validation, to evaluate the classifier. However, in the field of psychiatry, in recent years, it has also been recognized that there is a risk of prediction inflation when applying machine learning to a small number of samples obtained from a single facility.
[0038] In machine learning for a small amount of data, there is a high possibility of overfitting to specific tendencies or noises existing in the fMRI device, measurement method, experimenter, participant group, etc. of a specific facility in the learning data.
[0039] For example, there is also an example where a classifier for discriminating autism spectrum disorder based on anatomical images of the brain showed high performance with a sensitivity and specificity of over 90% for the learning data from the UK used for development, but dropped to 50% for Japanese data. From this, it can be said that a classifier that has not been validated using an independent validation cohort composed of facilities and participant groups completely different from the learning data has little scientific or practical significance.
[0040] Regarding the "harmonization method" for compensating for the inter-site differences between measurement sites as described above, the applicant of this case has also reported it (Non-Patent Document 8).
[0041] The descriptions of Patent Documents 1 to 7 and Non-Patent Documents 1 to 8 are all incorporated herein by reference.
[0042] Prior Art Documents
[0043] Patent Documents
[0044] Patent Document 1: Japanese Re-Examination Publication No. 2018 / 147193 (International Publication WO2018 / 147193)
[0045] Patent Document 2: Japanese Patent Application Laid-Open No. 2019-198376
[0046] Patent Document 3: Japanese Patent Application Laid-Open No. 2020-024139
[0047] Patent Document 4: Japanese Patent Application Publication No. 2019-516950 (International Publication WO2017 / 162773)
[0048] Patent Document 5: Japanese Re-Published Gazette No. 2005 / 025421 (International Publication WO2005 / 025421)
[0049] Patent Document 6: Japanese Unexamined Patent Application Publication No. 2015-62817
[0050] Patent Document 7: Japanese Unexamined Patent Application Publication No. 2017-196523
[0051] Non-Patent Document
[0052] Non-Patent Document 1: Press Release of National Center for Global Health and Medicine on December 10, 2018, "Endoscopic Diagnosis Support Program with AI Approved - to be Utilized for Physician Diagnosis Assistance -" https: / / www.amed.go.jp / news / release_20181210.html
[0053] Non-Patent Document 2: Glasser MF, et al. The Human Connectome Project's neuroimaging approach. Nat Neurosci 19, 1175-1187 (2016).
[0054] Non-Patent Document 3: Andrew T Drysdale, Logan Grosenick, Jonathan Downar, Katharine Dunlop, Farrokh Mansouri, Yue Meng1, Robert N Fetcho, Benjamin Zebley, Desmond J Oathes, Amit Etkin, Alan F Schatzberg, Keith Sudheimer, Jennifer Keller, Helen S Mayberg, Faith M Gunning, George S Alexopoulos, Michael D Fox, Alvaro Pascual-Leone, Henning U Voss, BJ Casey, Marc J Dubin & Conor Liston, "Resting-state connectivity biomarkers define neurophysiological subtypes of depression", nature medicine, VOLUME 23, NUMBER 1, JANUARY 2017
[0055] Non-Patent Document 4: Tomoki Tokuda, Junichiro Yoshimoto,, Yu Shimizu, Go Okada, Masahiro Takamura, Yasumasa Okamoto, Shigeto Yamawaki, Kenji Doya, "Identification of depression subtypes and relevant brain regions using a data-driven approach", SCIENTIFIC REPORTS│(2018)8:14082│DOI:10.1038 / s41598-018-32521-z
[0056] Non-Patent Document 5: Richard Dinga, Lianne Schmaal, Brenda W.J.H. Penninx, Marie Jose van Tol, Dick J. Veltman, Laura van Velzen, Maarten Mennes, Nic J.A. van der Wee, Andre F. Marquand, "Evaluating the evidence for biotypes of depression: Methodological replication and extension of Drysdale et al. (2017)", NeuroImage: Clinical 22 (2019) 101796
[0057] Non-Patent Document 6: Noble S, et al. Multisite reliability of MR-based functional connectivity. Neuroimage 146, 959-970 (2017).
[0058] Non-Patent Document 7: Pearlson G. Multisite collaborations and large databases in psychiatric neuroimaging advantages, problems, and challenges. Schizophr Bull 35, 1-2 (2009).
[0059] Non-patent document 8: Ayumu Yamashita, Noriaki Yahata, Takashi Itahashi, Giuseppe Lisi, Takashi Yamada, Naho Ichikawa, Masahiro Takamura, Yujiro Yoshihara, Akira Kunimatsu, Naohiro Okada, Hirotaka Yamagata, Koji Matsuo, Ryuichiro Hashimoto, GoOkada, Yuki Sakai, Jun Morimoto, Jin Narumoto, Yasuhiro Shimada, Kiyoto Kasai, Nobumasa Kato, Hidehiko Takahashi, Yasumasa Okamoto, Saori C Tanaka, Mitsuo Kawato, Okito Yamashita, and Hiroshi Imamizu, “Harmonization of resting-statefunctional MRI data across multiple imaging sites via the separation of sitedifferences into sampling bias and measurement "Reconciling resting-state functional MRI data across multiple imaging sites by separating site differences into sampling bias and measurement bias." PLOS Biology. DOI: 10.1371 / journal.pbio.3000042, http: / / journals.plos.org / plosbiology / article?id=10.1371 / journal.pbio.3000042 Summary of the invention
[0060] Problem that the invention aims to solve
[0061] As described above, when considering applying the analysis of brain activity based on functional brain imaging methods such as functional magnetic resonance imaging to the treatment of neurological / psychiatric diseases, for example, the analysis of brain activity based on functional brain imaging methods as biomarkers as described above is also expected to be used as a non-invasive functional marker for the development of diagnostic methods, the search and identification of target molecules for drug development to achieve fundamental treatment, etc.
[0062] For example, so far, for mental illnesses, no practical biomarker using genetic information has been completed yet. Therefore, it is difficult to determine the efficacy of drugs, etc., and thus it is also difficult to develop therapeutic drugs.
[0063] In order to generate a discriminator (recognizer) as a diagnostic biomarker and a classifier as a stratification biomarker based on measurement data of brain activity by machine learning and actually use them as biomarkers, it is necessary to improve the prediction accuracy of the biomarker generated by machine learning for the brain activity measured in one facility. In addition, it is necessary to be able to also apply the biomarker generated in this way to the brain activity measured in other facilities.
[0064] That is, when constructing a discriminator for identifying a disease and a classifier for classifying the disease into subtypes based on measurement data of brain activity by machine learning, there are mainly two problems.
[0065] The first problem is the small sample size.
[0066] The data volume N of the number of subjects is much smaller than the dimension M of the measured brain activity measurement data. Therefore, the parameters of the discriminator are likely to overfit the training data.
[0067] Due to this overfitting, the constructed discriminator shows very poor performance for newly sampled test data. This is because these test data were not used for training the discriminator.
[0068] Therefore, for the desired generalization of the discriminator, it is necessary to appropriately introduce feature selection and dimension reduction to only identify and utilize essential features.
[0069] The second problem is that the discriminator is clinically useful and scientifically reliable only when the constructed discriminator maintains high performance for MRI data scanned at a different imaging location from the location where the training data was collected.
[0070] This is the so-called generalization ability involving all imaging locations.
[0071] When collecting large-scale brain image data related to mental illnesses, since there are limitations in the amount of data that can be obtained at one location, it is necessary to obtain it from multiple locations.
[0072] However, when obtaining brain image data at multiple locations, the differences between locations become the biggest obstacle.
[0073] That is, in clinical applications, it is often observed that a discriminator trained using data obtained at a specific location cannot be generalized to data imaged at a different location.
[0074] Therefore, in the above-mentioned Human Connectome Project, it has been assumed so far that measurements are made using a single scanner at a single location.
[0075] The present invention has been completed to solve the above-mentioned problems, and an object thereof is to provide a clustering device for brain functional connection-related values that clusters subjects having a prescribed attribute based on brain measurement data acquired from multiple facilities, a clustering system for brain functional connection-related values, a clustering method for brain functional connection-related values, and a clustering classifier model for brain functional connection-related values.
[0076] Another object of the present invention is to provide a classifier program for brain functional connection-related values for realizing a classifier marker based on brain activity measurement and a brain activity marker classification system.
[0077] Solution for solving the problem
[0078] According to one aspect of the present invention, a clustering device for brain functional connection-related values that clusters a subject having at least one prescribed attribute among subjects based on a measurement result of brain activity of the subject. The clustering device for brain functional connection-related values includes a calculation processing system that performs a clustering process for a plurality of subjects including a first group of subjects having a prescribed attribute and a second group of subjects not having the prescribed attribute based on measurement values of brain activity. The calculation processing system includes a storage device and an arithmetic device. The arithmetic device is configured to: i) for each of the plurality of subjects, store a feature quantity based on a plurality of brain functional connection-related values representing temporal correlations of brain activity between a prescribed plurality of pairs of brain regions in the storage device; ii) perform machine learning for generating a recognition model for discriminating the presence or absence of an attribute by supervised learning based on the feature quantity stored in the storage device. In the machine learning for generating the recognition model, the arithmetic device performs the following processes: perform undersampling and downsampling according to the first group of subjects and the second group of subjects to generate a plurality of learning sub-samples; for each of the learning sub-samples of the learning sub-samples, select feature quantities for clustering from the union of the feature quantities used when generating the recognition model by machine learning according to the importance of the feature quantities belonging to the union. The arithmetic device further performs the following process: perform clustering on the first group of subjects by the multi-co-clustering method of unsupervised learning based on the selected feature quantities for clustering to generate a cluster classifier.
[0079] Preferably, the clustering device for brain functional connectivity related values receives information on the temporal correlation of brain activities between specified pairs of brain regions of each subject among multiple subjects from multiple brain activity measurement devices respectively provided at multiple measurement sites. The calculation processing system includes a coordination calculation unit. The coordination calculation unit corrects the multiple brain functional connectivity related values of each subject among multiple subjects to remove the measurement bias of the measurement site, and stores the adjusted values obtained by the correction as feature quantities in the storage device.
[0080] Preferably, the process of generating an identifier by machine learning is the following ensemble learning: generating multiple identifier sub-models for multiple learning sub-samples respectively, and integrating the multiple identifier sub-models to generate an identifier model.
[0081] Preferably, the attribute is represented by a label of a diagnosis result of a specified mental illness, and the clustering is a process of classifying the first subject group into at least one subtype cluster by data-driven machine learning.
[0082] Preferably, when generating an identifier by machine learning, the arithmetic device performs the following processes: i) dividing the adjusted value into a training data set for machine learning and a test data set for verification; ii) performing a specified number of undersampling and downsampling on the training data set to generate a specified number of learning sub-samples; iii) generating an identifier sub-model for each learning sub-sample; iv) integrating the outputs of the identifier sub-models to generate an identifier model for the presence or absence of an attribute.
[0083] Preferably, the process of generating an identifier by machine learning is a cross-validation with a nested structure having outer cross-validation and inner cross-validation. In the process of nested structure cross-validation, the arithmetic device performs the following processes: i) setting the outer cross-validation as K-fold cross-validation to divide the adjusted value into a training data set for machine learning and a test data set for verification; ii) performing a specified number of undersampling and downsampling on the training data set to generate a specified number of learning sub-samples; iii) in each cycle of K-fold cross-validation, adjusting hyperparameters by inner cross-validation to generate an identifier sub-model for each learning sub-sample; iv) generating an identifier model for the presence or absence of an attribute based on the identifier sub-model.
[0084] Preferably, the process of generating an identifier by machine learning is a machine learning method accompanied by feature quantity selection. In the selection of feature quantities for clustering, the importance of feature quantities is determined according to the ranking of the frequencies of the feature quantities belonging to the union when generating the identifier sub-model.
[0085] Preferably, the process of generating the recognizer through machine learning is the random forest method. In the selection of the feature quantities for clustering, the importance of the feature quantities belonging to the union is the importance calculated for each feature quantity based on the Gini impurity in the random forest method.
[0086] Preferably, the process of generating the recognizer through machine learning is a machine learning method based on L2 regularization. In the selection of the feature quantities for clustering, the importance of the feature quantities belonging to the union is determined according to the ranking based on the weights of the feature quantities in the recognizer sub-model calculated through L2 regularization.
[0087] Preferably, the storage device pre-saves the results obtained by measuring the brain activities of a plurality of predetermined brain regions for each of a plurality of traveling subjects who are commonly measured at a plurality of measurement locations. The arithmetic device performs the following processes: calculating a predetermined element of the brain functional connectivity array for each traveling subject, where the brain functional connectivity array represents the temporal correlation of the brain activities of a plurality of pairs of brain regions; calculating the measurement bias for each predetermined element of the brain functional connectivity array by using the general linear mixed model method as the fixed effect for each measurement location with respect to the average of this element involving a plurality of measurement locations and a plurality of traveling subjects.
[0088] Preferably, the arithmetic device performs the classification process into subtypes based on the measurement data measured for the subject at measurement locations other than the plurality of measurement locations.
[0089] According to another aspect of the present invention, a clustering system for brain function connection-related values is configured to perform clustering of subjects having at least one specified attribute in a subject based on measurement results of the brain activity of the subject. The clustering system for brain function connection-related values includes: a plurality of brain activity measurement devices respectively arranged at a plurality of measurement locations to measure the brain activities of a plurality of subjects in time series, the plurality of subjects including a first group of subjects having a specified attribute and a second group of subjects not having the specified attribute; and a calculation processing system configured to perform clustering processing on the plurality of subjects based on the measured values of the brain activities. The calculation processing system includes a storage device and an arithmetic device. The arithmetic device is configured to: i) for each of the plurality of subjects, save a feature quantity based on a plurality of brain function connection-related values representing the temporal correlation of the brain activities between a specified plurality of pairs of brain regions into the storage device; ii) perform machine learning for generating a discriminator model for discriminating the presence or absence of an attribute by supervised learning based on the feature quantities saved in the storage device. In the machine learning for generating the discriminator model, the arithmetic device performs the following processing: perform undersampling and downsampling according to the first group of subjects and the second group of subjects to generate a plurality of learning sub-samples; for each of the learning sub-samples of the learning sub-samples, select feature quantities for clustering from the union of the feature quantities used when generating the discriminator by machine learning according to the importance of the feature quantities belonging to the union. The arithmetic device further performs the following processing: perform clustering on the first group of subjects by the multi-co-clustering method of unsupervised learning based on the selected feature quantities for clustering to generate a cluster classifier.
[0090] Preferably, the calculation processing system receives information representing the temporal correlation of the brain activities between a specified plurality of pairs of brain regions of each of the plurality of subjects from the plurality of brain activity measurement devices respectively arranged at the plurality of measurement locations. The calculation processing system includes a coordination calculation unit that corrects the plurality of brain function connection-related values of each of the plurality of subjects to remove the measurement bias of the measurement location, and saves the adjusted value obtained by the correction into the storage device as a feature quantity.
[0091] Preferably, the attribute is represented by a label of a diagnosis result of a specified mental illness, and the clustering is a process of classifying the first group of subjects into at least one subtype cluster by data-driven machine learning.
[0092] According to another aspect of the present invention, a clustering method for brain functional connectivity-related values is used to perform clustering processing on a subject having at least one specified attribute in a subject by a computing processing system based on measurement results of the brain activity of the subject. The computing processing system includes a storage device and an arithmetic device. The clustering method for brain functional connectivity-related values includes the following steps: the arithmetic device stores, in the storage device, a feature quantity of a brain functional connectivity-related value based on the temporal correlation of brain activities between a plurality of specified pairs of brain regions for each of a plurality of subjects. The plurality of subjects includes a first group of subjects having a specified attribute and a second group of subjects not having the specified attribute; and the arithmetic device performs machine learning for generating a recognition model for discriminating the presence or absence of an attribute by supervised learning based on the feature quantities stored in the storage device. The step of performing machine learning for generating a recognition model includes the following steps: undersampling and downsampling are performed according to the first group of subjects and the second group of subjects to generate a plurality of learning sub-samples; and for each of the learning sub-samples of the learning sub-samples, feature quantities for clustering are selected from the union of the feature quantities used when generating a recognition model by machine learning according to the importance of the feature quantities belonging to the union. The clustering method for brain functional connectivity-related values further includes the following steps: the arithmetic device clusters the first group of subjects by the multi-co-clustering method of unsupervised learning based on the selected feature quantities for clustering to generate a classifier for clusters.
[0093] According to another aspect of the present invention, a classifier program for brain functional connectivity related values is generated by a computing processing system performing clustering processing on subjects having at least one specified attribute based on measurement results of the brain activities of the subjects. The classifier program for brain functional connectivity related values is used for a computer to classify input data into each cluster corresponding to the result of the clustering processing. The classifier program has the following classification function: the computer classifies the input data into the cluster with the maximum posterior probability based on the model of the probability distribution of each cluster. The computing processing system includes a storage device and an arithmetic device. In the generation processing of generating the classifier program based on the clustering processing, the computing processing system performs the following steps: for each of a plurality of subjects, the arithmetic device stores a feature quantity based on the brain functional connectivity related value representing the temporal correlation of the brain activities between a specified plurality of pairs of brain regions in the storage device. The plurality of subjects includes a first group of subjects having a specified attribute and a second group of subjects not having the specified attribute; and the arithmetic device performs machine learning for generating a recognition model for discriminating the presence or absence of the attribute based on the feature quantities stored in the storage device by supervised learning. The step of performing machine learning for generating the recognition model includes the following steps: undersampling and downsampling are performed according to the first group of subjects and the second group of subjects to generate a plurality of learning sub-samples; and for each of the learning sub-samples of the learning sub-samples, feature quantities for clustering are selected from the union of the feature quantities used when generating the recognition model by machine learning according to the importance of the feature quantities belonging to the union. The arithmetic device also performs the following steps: based on the selected feature quantities for clustering, the first group of subjects is clustered by the multi-co-clustering method of unsupervised learning to generate a cluster classifier.
[0094] Preferably, the computing processing system performs the following steps: receiving information representing the temporal correlation of the brain activities between a specified plurality of pairs of brain regions of each of a plurality of subjects from a plurality of brain activity measurement devices respectively provided at a plurality of measurement locations; and performing coordination for correcting the plurality of brain functional connectivity related values of each of the plurality of subjects to remove the measurement bias of the measurement location, and storing the adjusted value obtained by the correction as a feature quantity in the storage device.
[0095] Preferably, the attribute is represented by a label of a diagnosis result of a specified mental illness, and the clustering is a process of classifying the first group of subjects into at least one subtype of clusters by data-driven machine learning.
[0096] According to another aspect of the present invention, a brain activity marker classification system is generated by a computing processing system performing clustering processing on a subject having at least one specified attribute based on measurement results of the subject's brain activity. The brain activity marker classification system is used for a computer to classify input data into each cluster corresponding to the result of the clustering processing. The brain activity marker classification system has the following classification function: the computer classifies the input data into the cluster with the maximum posterior probability based on the model of the probability distribution of each cluster. The computing processing system includes a storage device and an arithmetic device. In the generation processing of generating the brain activity marker classification system based on the clustering processing, the computing processing system performs the following steps: for each of a plurality of subjects, the arithmetic device stores a feature quantity based on a brain functional connection correlation value representing the temporal correlation of brain activity between a specified plurality of pairs of brain regions in the storage device. The plurality of subjects includes a first group of subjects having a specified attribute and a second group of subjects not having the specified attribute; and the arithmetic device performs machine learning for generating a discriminator model for discriminating the presence or absence of an attribute based on the feature quantities stored in the storage device by supervised learning. The step of performing machine learning for generating the discriminator model includes the following steps: performing undersampling and downsampling according to the first group of subjects and the second group of subjects to generate a plurality of learning sub-samples; and for each of the learning sub-samples of the learning sub-samples, selecting feature quantities for clustering from the union of the feature quantities used when generating the discriminator by machine learning according to the importance of the feature quantities belonging to the union. The arithmetic device further performs the following steps: based on the selected feature quantities for clustering, clustering the first group of subjects by the co-clustering method of unsupervised learning to generate a cluster classifier.
[0097] Preferably, the computing processing system performs the following steps: receiving information representing the temporal correlation of brain activity between a specified plurality of pairs of brain regions of each of a plurality of subjects from a plurality of brain activity measurement devices respectively provided at a plurality of measurement locations; and performing harmonization, where the harmonization is used to correct a plurality of brain functional connection correlation values representing the temporal correlation of brain activity of each of the plurality of subjects to remove the measurement bias of the measurement location, and thus storing the adjusted values obtained by the correction as feature quantities in the storage device.
[0098] Preferably, the attribute is represented by a label of a diagnosis result of a specified mental illness, and the clustering is a process of classifying the first group of subjects into at least one subtype cluster by data-driven machine learning.
[0099] According to another aspect of the present invention, a clustering classifier model for brain functional connectivity-related values is generated by a computing processing system performing clustering processing on subjects having at least one prescribed attribute based on measurement results of the brain activities of the subjects. The clustering classifier model for brain functional connectivity-related values is used for a computer to classify input data into respective clusters corresponding to the results of the clustering processing. The clustering classifier model has the following functions: for each view obtained by dividing a group of feature quantities representing a subject included in learning data, calculating the value of the probability density function for the input data based on the information of the feature quantities included in each view and the information of the probability density functions of the respective clusters for determining the subjects of each view, and classifying the input data into the cluster having the maximum posterior probability. The computing processing system includes a storage device and an arithmetic device. In the generation processing of generating the clustering classifier model based on the clustering processing, the computing processing system performs the following steps: the arithmetic device stores, for each of a plurality of subjects, the feature quantities based on the brain functional connectivity-related values representing the temporal correlation of the brain activities between a prescribed plurality of pairs of brain regions in the storage device. The plurality of subjects includes a first group of subjects having a prescribed attribute and a second group of subjects not having the prescribed attribute; and the arithmetic device performs machine learning for generating a discriminator model for discriminating the presence or absence of the attribute by supervised learning based on the feature quantities stored in the storage device. The step of performing machine learning for generating the discriminator model includes the following steps: performing undersampling and downsampling based on the first group of subjects and the second group of subjects to generate a plurality of learning sub-samples; and for each of the learning sub-samples of the learning sub-samples, selecting the feature quantities for clustering from the union of the feature quantities used when generating the discriminator by machine learning according to the importance of the feature quantities belonging to the union. The arithmetic device further performs the following steps: clustering the first group of subjects by the co-clustering method of unsupervised learning based on the selected feature quantities for clustering, dividing the feature quantities into views, and generating the probability density functions of the respective clusters of the subjects in each view.
[0100] Effects of the Invention
[0101] According to the present invention, it is possible to adjust and correct the measurement biases of each facility for the measurement data of the brain activities measured at a plurality of facilities. Thereby, it is possible to adjust the brain functional connectivity-related values based on the measurement data at the plurality of facilities and perform clustering.
[0102] In addition, according to the present invention, it is possible to implement a clustering device for brain functional connectivity-related values, a clustering system for brain functional connectivity-related values, a clustering method for brain functional connectivity-related values, a classifier program for brain functional connectivity-related values, and a brain activity marker classification system that can coordinate the measurement data of the brain activities measured at a plurality of facilities to objectively judge the subtypes of disease groups. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 It is a conceptual diagram for explaining the concept of coordinated processing for data measured by an MRI measurement system provided at multiple measurement locations.
[0104] Figure 2A It is a conceptual diagram showing the process of extracting a correlation matrix representing the correlation of functional connections in a resting state for a region of interest (ROI: Region of Interest) of a subject's brain, and is a diagram showing the time series data of rsfMRI.
[0105] Figure 2B It is a conceptual diagram showing the process of extracting a correlation matrix representing the correlation of functional connections in a resting state for a region of interest of a subject's brain, and is a diagram showing the correlation matrix.
[0106] Figure 3A It is a conceptual diagram showing an example of the content of "measurement parameters".
[0107] Figure 3B It is a conceptual diagram showing an example of the content of "subject attribute data".
[0108] Figure 4 It is a schematic diagram showing the overall structure of the MRI device 100.i (1 ≤ i ≤ Ns) provided at each measurement location.
[0109] Figure 5 It is a hardware block diagram of the data processing unit 32.
[0110] Figure 6 It is a conceptual diagram explaining the process of generating a discriminator that becomes a diagnostic marker based on the correlation matrix and the clustering process.
[0111] Figure 7 It is a functional block diagram for explaining the structure of the calculation processing system 300.
[0112] Figure 8 It is a functional block diagram for explaining the structure of the calculation processing system 300.
[0113] Fig. 9 It is a flowchart for explaining the process of machine learning for generating a disease identifier through ensemble learning.
[0114] Fig.10 It is a diagram showing the population characteristics of the learning dataset (dataset 1).
[0115] Fig.11 It is a diagram showing the population characteristics of the independent verification dataset (dataset 2).
[0116] Fig.12It is a graph showing the prediction performance (probability distribution of the output) of MDD for the learning dataset regarding all shooting locations.
[0117] Fig.13 It is a graph showing the prediction performance (probability distribution of the output of the recognizer) of MDD for the learning dataset regarding each shooting location.
[0118] Fig.14 It is a graph showing the probability distribution of the output of the recognizer of MDD in an independent validation dataset.
[0119] Fig.15 It is a graph showing the probability distribution of the output of the recognizer of MDD for the independent validation dataset regarding each shooting location.
[0120] Fig.16 It is a flowchart for explaining the process of selecting feature quantities and performing clustering through unsupervised learning.
[0121] Fig.17 It is a graph showing the concept of implementing the selection of feature quantities through "learning process with feature quantity selection" in the case where there are multiple (e.g., Nch) feature quantities.
[0122] Fig.18 It is a conceptual graph showing the feature quantities finally selected when generating one recognizer through the learning process with feature quantity selection.
[0123] Fig.19 It is a conceptual graph showing the situation of selecting feature quantities when generating a recognizer by repeatedly performing the processes of undersampling and downsampling.
[0124] Fig. 20 It is a conceptual graph for explaining the situation where there are multiple ways of dividing clusters according to feature quantities.
[0125] Fig.21A It is a conceptual graph for explaining the concept of clustering in the case of characterizing multiple objects with multiple feature quantities.
[0126] Fig. 21B It is a conceptual graph for explaining the concept of clustering in the case of characterizing multiple objects with multiple feature quantities.
[0127] Fig.22A It is a conceptual graph for explaining the concept of multi-clustering.
[0128] Fig. 22B It is a conceptual graph for explaining the concept of multi-co-clustering.
[0129] Fig.23 It is a conceptual graph showing the situation of a probability model in which different types of probability distributions are envisioned in one view in "multi-co-clustering".
[0130] Fig.24 It is a flowchart for explaining the outline of the learning method for multi - co - clustering.
[0131] Fig.25 It is a graph showing the graphical representation of Bayesian estimation in the learning method of multi - co - clustering.
[0132] Fig.26A It is a graph showing Dataset 1 in the dataset divided into two.
[0133] Fig.26B It is a graph showing Dataset 2 in the dataset divided into two.
[0134] Fig. 27 It is a concept map explaining the concept of performing clustering in each dataset.
[0135] Fig.28 It is a concept map showing an example of multi - co - clustering for subject data.
[0136] Fig.29 It is a graph showing the result obtained by actually performing multi - co - clustering processing on Dataset 1 and Dataset 2.
[0137] Fig.30 It is a table showing the number of brain functional connections (FC) assigned to each view in Dataset 1 and Dataset 2.
[0138] Fig.31 It is a concept map for explaining the evaluation method of the similarity (hierarchical generalization performance) of clustering.
[0139] Fig.32A It is a concept map for explaining ARI.
[0140] Fig.32B It is a concept map for explaining ARI.
[0141] Fig.33A It is a graph showing the result obtained by calculating ARI for each view of Dataset 1 and Dataset 2 in tabular form.
[0142] Fig.33B It is a graph showing the evaluation results of the similarity between Cluster 1 and Cluster 1′, and between Cluster 2 and Cluster 2′, and is a graph showing the result of the Fig.33A corresponding permutation test.
[0143] Fig.34 It is a graph showing the Figure 1 distribution of the number of subjects assigned to each cluster of each view of Cluster 1 and Cluster 1′.
[0144] Fig.35 It is a conceptual diagram for explaining an evaluation method for the difference between locations of a mobile subject who moves between locations to receive measurements.
[0145] Fig.36 It is a conceptual diagram for explaining the expression of the b-th functional connectivity of subject a.
[0146] Fig.37 It is a flowchart for explaining the process of calculating measurement bias for harmonization.
[0147] Fig.38 It is a conceptual diagram for explaining the calculation process of measurement bias for harmonization processing in the case where a new measurement location is added.
[0148] Fig.39A It is a conceptual diagram showing the data set of the multi-disease database used in the harmonization process.
[0149] Fig.39B It is a conceptual diagram showing the data set of multi-facility examinees used in the harmonization process.
[0150] Fig.40 It is a diagram showing the content of the multi-disease data set of SRPBS.
[0151] Fig.40A It is a diagram showing the content of the multi-disease data set of SRPBS.
[0152] Fig.40B It is a diagram showing the content of the multi-disease data set of SRPBS.
[0153] Fig.40C It is a diagram showing the content of the multi-disease data set of SRPBS.
[0154] Fig.40D It is a diagram showing the content of the multi-disease data set of SRPBS.
[0155] Fig.41 It is a diagram showing the shooting protocol at each measurement location.
[0156] Fig.41A It is a diagram showing the shooting protocol at each measurement location.
[0157] Fig.41B It is a diagram showing the shooting protocol at each measurement location.
[0158] Fig.41C It is a diagram showing the shooting protocol at each measurement location.
[0159] Fig.41D It is a diagram showing the shooting protocol at each measurement location.
[0160] Fig.42 It is a graph showing the visualization of the differences between locations and disease effects obtained by principal component analysis.
[0161] Fig.43 It is a dendrogram based on hierarchical cluster analysis.
[0162] Fig.44 It is a graph showing the contribution size of each factor.
[0163] Fig.45 It is a graph visualizing the influence of the harmonization process and is a graph for comparison with Fig.42 ...
[0164] Fig.46 It is a functional block diagram showing an example in the case of collecting data, performing estimation processing, and measuring the brain activity of subjects in a decentralized manner.
[0165] Fig.47 It is a graph showing the tour mode of multi-facility examinees.
[0166] Fig.48 It is a conceptual diagram showing the structure of a clustering classifier. Detailed Implementation Manner
[0167] Hereinafter, in order to explain the "clustering device for brain function connection related values", "clustering method for brain function connection related values", etc. of the present invention, an example of "clustering" of brain function connection image data of examinees (including patients with mental diseases) measured by a measurement system composed of multiple brain activity measurement devices using artificial intelligence technology will be described.
[0168] Therefore, the structure of the measurement system of the embodiment of the present invention, more specifically, the MRI measurement system, will be described with reference to the drawings. In addition, in the following embodiments, the components and processing steps with the same reference numerals are the same or equivalent components and processing steps, and in cases where it is not necessary, they will not be repeatedly described.
[0169] In addition, in the present embodiment, the present invention will be described as follows: The "brain activity measurement device", more specifically, the "MRI device", provided at multiple facilities measures the brain activities between multiple regions of the brain in time series, and based on the patterns of the temporal correlation (referred to as "brain function connection") between these regions, the examinees with a specific disease are further classified into multiple groups (subgroups) in a manner that can be applied to multiple facilities.
[0170] In addition, although not particularly limited, for the "specific disease", "major depressive disorder" is taken as an example for illustration. However, as shown in the following description, the present invention relates to a technique for classifying the "brain functional connectivity related value" of a subject in a data-driven manner, and the disease of the subject is not limited to "major depressive disorder", but may also be other diseases. In addition, as long as it is an attribute of the subject classified according to the pattern of the "brain functional connectivity related value" of the subject, it does not necessarily have to be a disease, but may also be other attributes.
[0171] Moreover, for such an "MRI measurement system", a plurality of "MRI devices" are arranged at a plurality of different facilities. As will be described later, for the between-site differences in measurement between these measurement facilities (measurement locations), the measurement bias caused by the measurement equipment and the differences caused by the number of subjects at the measurement location (sampling bias) are independently evaluated. On this basis, by performing a process of removing the effect of measurement bias on the measurement values at each measurement location to correct the between-site differences, the reconciliation process (coordination) of the measurement results between the measurement locations is thus achieved. Moreover, the following is set for illustration: for the brain functional connectivity values after coordination, "feature quantity selection" is performed using ensemble learning with the diagnostic labels of specific diseases as training data, and then clustering is performed through unsupervised learning, thereby performing the classification of the subject attributes (for example, subtypes of mental diseases).
[0172] [Embodiment 1]
[0173] Figure 1 It is a conceptual diagram for explaining the clustering (stratification) process for the data measured by an MRI measurement system arranged at a plurality of measurement locations.
[0174] Refer to Figure 1 , and it is assumed that MRI devices 100.1 to 100.Ns are respectively arranged at measurement locations MS.1 to MS.Ns (Ns: number of locations).
[0175] In addition, at measurement locations MS.1 to MS.Ns, subject groups PA.1 to PA.Ns are respectively measured. It is assumed that each of subject groups PA.1 to PA.Ns contains at least two or more classified groups, such as a patient group and a healthy group. In addition, as the patient group, although not particularly limited, for example, it is a group equivalent to patients with mental diseases, and more specifically, a group of "patients with major depressive disorder".
[0176] Moreover, at each measurement location, in principle, the subjects are measured with a unified measurement protocol within the possible range according to the specifications of the MRI device.
[0177] Here, although not particularly limited, as the measurement protocol, for example, it is assumed to stipulate the following content.
[0178] 1) Direction of performing head scanning
[0179] For example, it is necessary to specify whether to perform scanning in the direction from the posterior side of the head (posterior: hereinafter abbreviated as "P") to the anterior side (anterior: hereinafter abbreviated as "A") (hereinafter referred to as the "P→A direction"), or the opposite direction, i.e., from the anterior side to the posterior side (hereinafter referred to as the "A→P direction"). Depending on the situation, it may also be specified to perform scanning in both directions.
[0180] Depending on the MRI device, there may be different default directions, or it may not be possible to arbitrarily set the two directions.
[0181] The scanning direction may be specified, for example, as the "deformation method" of the image, so as to set the conditions as a protocol.
[0182] 2) Imaging conditions for brain structure images
[0183] Set the conditions for imaging either or both of the "T1-weighted image" and the "T2-weighted image" by the so-called spin echo method.
[0184] 3) Imaging conditions for brain functional images
[0185] Set the conditions for imaging the brain functional images of a subject in the "resting state" by the fMRI (functional Magnetic Resonance Imaging) method.
[0186] 4) Imaging conditions for diffusion-weighted images
[0187] Set whether to image diffusion-weighted images (DWI: diffusion (weighted) image), and also set the conditions for this imaging.
[0188] Diffusion-weighted images are a type of MRI imaging sequence and are obtained by visualizing the diffusion movement of water molecules. Diffusion-weighted images utilize the following situation: in the pulse sequence of the commonly used spin echo method, the signal attenuation due to diffusion can be ignored, but when a large tilt magnetic field is continuously applied for a long time, the phase shift generated due to the movement of each magnetization vector during this period can no longer be ignored, and the more active the diffusion area, the lower the signal it shows.
[0189] 5) Imaging for correcting EPI distortion through image processing
[0190] For example, as a method for correcting EPI distortion through image processing, the "field map method" is known, and shooting conditions are set for correcting spatial distortion.
[0191] The field map method collects EPI images according to multiple echo times, and calculates the amount of EPI distortion based on these EPI images. The field map method can be applied to correct the EPI distortion included in a new image. It is possible to calculate the EPI distortion on the premise of a set of images obtained for the same anatomical structure at different echo times, and correct the distortion of the image.
[0192] For example, the "field map method" is disclosed in the following known document 1. All the descriptions in known document 1 are incorporated herein by reference.
[0193] Known Document 1: Japanese Patent Application Laid-Open No. 2015-112474
[0194] In addition, for the measurement protocol, the required sequence part can be appropriately extracted from the above conditions, and other sequences and the conditions of these other sequences can be added as needed.
[0195] Refer again Figure 1 , the subject used for measurement at each measurement site MS.1 to MS.Ns is called "sampling the subject", and the reason for the inter-site difference in measurement values caused by the unevenness of sampling at each measurement site is called "sampling bias".
[0196] For example, in the above example, it is known that among the patients diagnosed as "major depressive disorder" according to the previous diagnostic criteria, some subtype patients are actually included.
[0197] Typically, there are "melancholic", "atypical", "seasonal", "postpartum" depression, etc. In addition, it has also been reported that "the situation where symptoms of moderate or more severity persist without sufficient improvement after sufficient treatment with two or more antidepressants with different mechanisms of action" is called "treatment-resistant depression", and it is estimated to account for 10% to 20% of depression. That is, it is known that the patient group usually diagnosed as "major depressive disorder" is by no means a homogeneous group of patients. However, a method for classifying such subtypes based on objective measurement data has not been put into practical use so far.
[0198] At each measurement site, due to various factors such as the bias in the nature of patients visiting the hospital at that measurement site due to regionality and the tendency of diagnosis in that hospital, even if the patients are generally "major depressive disorder" patients, the distribution of subtypes included in this patient group is not necessarily uniform. As a result, the subtype distribution of the patient groups at each measurement site usually shows unevenness, and the above-mentioned "sampling bias" is considered to be generated as a result.
[0199] In addition, even in the group of subjects called the "healthy group", multiple subtypes usually exist within this group. In this regard, the "healthy group" also has "sampling bias".
[0200] Moreover, it cannot be said that the MRI apparatuses 100.1 to 100.Ns use MRI apparatuses with absolutely identical measurement characteristics at each measurement location.
[0201] For example, location - to - location differences in measurement data can be generated corresponding to conditions of the MRI apparatus such as the manufacturer of the MRI apparatus, the model of the MRI apparatus, the static magnetic field strength of the MRI apparatus, the number of coils (channels) of the (transmitting) receiving coil in the MRI apparatus, and measurement conditions such as the measurement conditions of the MRI apparatus. The location - to - location differences caused by such measurement conditions are referred to as "measurement bias".
[0202] Even for MRI apparatuses of the same model from the same manufacturer of MRI apparatuses, due to the inherent uniqueness of the apparatuses, it cannot be said that they necessarily achieve exactly the same measurement characteristics.
[0203] Here, the (transmitting) receiving coil generally uses a "multi - array coil" in order to improve the SN ratio of the measured signal. The "number of coils of the receiving coil" refers to the number of "element coils" that make up the multi - array coil. By improving the sensitivity of each element coil and bundling their outputs, an improvement in receiving sensitivity is thereby achieved.
[0204] Moreover, although not particularly limited, in the present embodiment, "sampling bias" and "measurement bias" can be independently evaluated by a harmonization method as described later.
[0205] Refer again to Figure 1 , the measurement - related data DA100.1 to DA100.Ns from each of the measurement locations MS.1 to MS.Ns are accumulated and stored in the storage device 210 within the data center 200.
[0206] Here, the "measurement - related data" includes "measurement parameters" at each measurement location, as well as "patient group data" and "healthy group data" measured at each measurement location.
[0207] And, in the "patient group data" and "healthy group data", "MRI measurement data of patients" and "MRI measurement data of healthy subjects" are respectively included corresponding to each subject.
[0208] Next, an explanation will be given of such "measurement - related data".
[0209] Figure 2A and Figure 2BIt is a conceptual diagram showing a process of extracting a correlation matrix representing the functional connectivity in a resting state for a region of interest in a subject's brain. Figure 2A It is a diagram showing the time series data of rsfMRI, Figure 2B and is a diagram showing its correlation matrix.
[0210] Here, in Figure 1 at least the following data are included in the "MRI measurement data of patients" and the "MRI measurement data of healthy subjects" in the "patient group data" and the "healthy group data".
[0211] i) Time-series "brain functional image data" for calculating the data of the correlation matrix, and / or the data of the correlation matrix itself
[0212] That is, Figure 1 the calculation processing system 300 in uses these data as the basis data for calculating a brain activity biomarker as described later based on the data stored in the storage device 210.
[0213] Here, it is possible to adopt the following structure: After calculating the data of the correlation matrix based on the time-series "brain functional image data" at each measurement location, the data of the correlation matrix is stored in the storage device 210, and the calculation processing system 300 calculates the brain activity biomarker based on the data of the correlation matrix in the storage device 210.
[0214] Alternatively, it is also possible to adopt the following structure: The time-series "brain functional image data" is stored in the storage device 210, and the calculation processing system 300 calculates the data of the correlation matrix based on the "brain functional image data" in the storage device 210, and then calculates the brain activity biomarker.
[0215] Therefore, for each of the "MRI measurement data of patients" and the "MRI measurement data of healthy subjects", at least either the time-series "brain functional image data" for calculating the data of the correlation matrix or the data of the correlation matrix itself is included.
[0216] ii) Subject's structural image data and diffusion-weighted image data
[0217] In addition, although not particularly limited, for the process of correcting EPI distortion by image processing, it is possible to adopt a structure in which data is stored in the storage device 210 after arithmetic processing is performed at each measurement location.
[0218] In addition, although not particularly limited, from the perspective of personal information protection, a structure can be adopted in which anonymization processing is performed at each measurement location before saving the data in the storage device 210. Regarding the anonymization processing, a structure can be adopted in which the anonymization processing is performed in the computing processing system 300 when the entity operating the computing processing system 300 is legally authorized to process personal information, etc.
[0219] Return Figure 2A and Figure 2B , such as Figure 2A shown, calculate the average "activity" of each region of interest based on the fMRI data at n (n: natural number) moments of the resting-state fMRI measured in real time. As Figure 2B shown, calculate the correlation matrix of the functional connectivity ("correlation value of activity") between brain regions (regions of interest).
[0220] (Partitioning (Segmentation) of Brain Regions)
[0221] Calculate the functional connectivity as the temporal correlation of the blood oxygenation level-dependent (BOLD) signals of the resting-state functional MRI between two brain regions that depends on each participant.
[0222] Here, as the regions of interest, since the Nr region is considered as described above, when considering symmetry, the independent non-diagonal components in the correlation matrix are
[0223] Nr×(Nr - 1) / 2 (pieces).
[0224] As a method for setting the regions of interest, assume the following method.
[0225] Method 1) "Define the regions of interest based on anatomically defined brain regions."
[0226] Here, for brain activity biomarkers, for example, 140 regions are used as the regions of interest.
[0227] That is, in this method, regarding the ROI, in addition to using the 137 ROIs included in the Brain Sulci Atlas (BAL), the ROIs of the cerebellum (left and right) and vermis of the Automated Anatomical Labeling Atlas are also used. The functional connectivity FC between these 140 ROIs is used as a feature quantity.
[0228] Here, the Brain Sulci Atlas (BAL) and the Automated Anatomical Labeling Atlas are publicly available below. The descriptions of each of the following publicly known documents 2 and 3 are hereby incorporated by reference in their entirety.
[0229] Publicly known document 2: Perrot et al., Med Image Anal, 15(4), 2011
[0230] Publicly known document 3: Tzourio-Mazoyer et al., Neuroimage, 15(1), 2002
[0231] As such regions of interest, for example, are the following regions.
[0232] Dorsomedial prefrontal cortex (DMPFC),
[0233] Ventral medial prefrontal cortex (VMPFC),
[0234] Anterior cingulate cortex (ACC),
[0235] Cerebellar vermis,
[0236] Left thalamus,
[0237] Right inferior parietal lobe,
[0238] Right caudate nucleus,
[0239] Right middle occipital gyrus,
[0240] Right middle cingulate cortex.
[0241] Among them, the regions of the brain used are not limited to such regions.
[0242] For example, the regions to be selected can also be changed according to the neural / psychiatric disease being targeted.
[0243] Method 2) "Define functional connectivity based on brain regions of a functional brain map covering the entire brain."
[0244] Here, regarding the brain regions of such a functional brain map, they are also publicly available in the following documents. Although not particularly limited, for example, such a structure can be configured to be composed of 268 nodes (brain regions). The descriptions of each of the following publicly known documents 4 to 7 are hereby incorporated by reference in their entirety.
[0245] Prior art document 4: Noble S, et al. Multisite reliability of MR-based functional connectivity. Neuroimage 146, 959-970 (2017).
[0246] Prior art document 5: Finn ES, et al. Functional connectome fingerprinting: identifying individuals using patterns of brain connectivity. Nat Neurosci 18, 1664-1671 (2015).
[0247] Prior art document 6: Rosenberg MD, et al. A neuromarker of sustained attention from whole-brain functional connectivity. Nat Neurosci 19, 165-171 (2016).
[0248] Prior art document 7: Shen X, Tokoglu F, Papademetris X, Constable RT. Groupwise whole-brain parcellation from resting-state fMRI data for network node identification. Neuroimage 82, 403-415 (2013).
[0249] Method 3) Surface-based method
[0250] Regarding the segmentation of brain regions, by using multimodal bioimaging of the Human Connectome Project (HCP) type (myelin task functional), it is also possible to analyze data using the "surface-based method" based on a brain map created by transforming the brain into thin sheets along the sulci.
[0251] For such a segmentation method, a toolbox (ciftify toolbox version 2.0.2) publicly available on the following website can be used. https: / / edickie.github.io / ciftify / # /
[0252] This toolbox enables the analysis of the data used in an HCP-like surface-based pipeline (e.g., even in the case where the T2-weighted image required by the HCP pipeline is missing).
[0253] Moreover, in the analysis of Method 3, as regions of interest (ROIs), 379 surface-based parcellations (360 cortical parcellations + 19 subcortical parcellations) publicly available in the following known reference 8 are used. All the descriptions in the following known reference 8 are hereby incorporated by reference.
[0254] Known reference 8: Glasser, M.F., Coalson, T.S., Robinson, E.C., Hacker, C.D., Harwell, J., Yacoub, E., et al. (2016). A multi-modal parcellation of human cerebral cortex. Nature 536(7615), 171 - 178. doi: 10.1038 / nature18933.
[0255] Therefore, the temporal changes in the BOLD signal are extracted from these 379 regions of interest (ROIs).
[0256] Furthermore, by using the anatomical automatic labeling (AAL) and Neurosynth (http: / / neurosynth.org / locations / ) publicly available in the following known reference 9, the anatomical names of important ROIs and the names of the intrinsic brain networks including the ROIs can be determined. All the descriptions in the following known reference 9 are hereby incorporated by reference.
[0257] Prior Art Document 9: Tzourio-Mazoyer, N., Landeau, B., Papathanassiou, D., Crivello, F., Etard, O., Delcroix, N., et al. (2002). Automated anatomical labeling of activations in SPM using a macroscopic anatomical parcellation of the MNI MRI single-subject brain. Neuroimage 15(1), 273-289. doi:10.1006 / nimg.2001.0978.
[0258] Method 4) A method for determining regions of the brain in a data-driven manner
[0259] As disclosed in Prior Art Document 10 below, it is a method for newly identifying a network based on phase-consistent voxels without prior information (brain map), and is a method such as "Canonical ICA", "dictionary learning", etc. All the descriptions of Prior Art Document 10 below are hereby incorporated by reference.
[0260] Prior Art Document 10: Kamalaker Dadi, Mehdi Rahim, Alexandre Abraham, Darya Chyzhyk, Michael Milham, Bertrand Thirion, Gael Varoquaux, "Benchmarking functional connectome-based predictive models for resting-state fMRI", Preprint submitted to NeuroImage, October 31, 2018.
[0261] Hereinafter, a method of defining functional connectivity of brain regions using a surface-based brain map based on "Method 3" will be described for the sake of illustration.
[0262] In addition, as for the calculation of the correlation value, there are several candidates for measuring functional connectivity, such as the tangent method and the local correlation method.
[0263] However, hereinafter, although not particularly limited, it will be described using the Pearson correlation coefficient for the sake of illustration.
[0264] For each possible node group of the preprocessed BOLD signal, the Pearson correlation coefficient after Fisher-z transformation is calculated over the period of time elapsed, and is used to construct a 379×379 symmetric brain functional connectivity matrix, the elements of which respectively represent the strength of the connectivity between two nodes.
[0265] Figure 3A and Figure 3B are conceptual diagrams respectively showing examples of the contents of "measurement parameters" and "subject attribute data".
[0266] "Subject attribute data" is set to be stored in association with the "patient MRI measurement data" and the "healthy subject MRI measurement data" respectively in the "patient group data" or "healthy group data" of Figure 1 .
[0267] As Figure 3A shown, it includes a location ID for identifying the measurement location, a location name, a condition ID for identifying the measurement parameters, information related to the measurement device, and information related to the measurement conditions.
[0268] "Measurement parameters" includes "information related to the measurement device" and "information related to the measurement conditions".
[0269] In the "information related to the measurement device", it includes the manufacturer name, model number, and the number of (transmitting) receiving coils of the MRI device used to measure the brain activity of the subject at each measurement location.
[0270] In addition, the "information related to the measurement device" is not limited to these, and for example, it may also include other performance indicators of the measurement device such as the static magnetic field strength and the uniformity of the magnetic field after shimming adjustment.
[0271] In the "information related to the measurement conditions", it includes the direction of phase encoding during image reconstruction (P→A or A→P), the type of image (T1-weighted, T2-weighted, diffusion-weighted, etc.), the imaging sequence (spin echo, etc.), and information such as whether the subject's eyes are open / closed during imaging.
[0272] The "information related to the measurement conditions" is not limited to these either.
[0273] As Figure 3B shown, the "subject attribute data" includes a subject temporary ID that has been anonymized in a way that the subject cannot be identified, a condition ID indicating the measurement conditions when measuring the subject, and the subject's attribute information.
[0274] Moreover, as the "subject's attribute information", it includes information such as the subject's gender, age, a label indicating either health or disease, the name of the diagnosed disease diagnosed by a doctor for the subject, the medication history of administering medications to the subject, the diagnosis history, etc.
[0275] In addition, the "subject's attribute information" is set to be anonymized as needed, for example, at the measurement location.
[0276] For example, regarding age, gender, etc., it can be processed in a way that maintains "k-anonymity". "K-anonymity" is to reduce the probability of an individual being identified to less than 1 / k by processing such as converting data so that there are k or more pieces of data of quasi-identifiers (the same attributes). Here, a "quasi-identifier" refers to an attribute of an individual that cannot be determined by a single piece of information such as "age", "gender", "residence", etc., but can be determined by combining them.
[0277] In addition, the medication history and the diagnosis history are processed for anonymization as needed, such as randomly generating dates, adjusting (relatively changing) dates, etc.
[0278] Moreover, hereinafter, for the "MRI measurement data of patients" and the "MRI measurement data of healthy subjects", the functional connectivity calculated by the above method as the correlation of activities between brain regions of each subject over time is collectively referred to as the "Connectivity" between regions (abbreviated as "FC" when omitted). When it is necessary to distinguish the functional connectivity of each brain region, it is set to be distinguished by adding a suffix as described later.
[0279] [Structure of MRI Device]
[0280] Figure 4 It is a schematic diagram showing the overall structure of the MRI device 100.i (1 ≤ i ≤ Ns) provided at each measurement location.
[0281] In Figure 4 it, the MRI device 100.1 at the first measurement location is illustratively and detailedly described. Regarding the other MRI devices 100.2 to 100.Ns, the basic structure is the same.
[0282] As Figure 4As shown, the MRI apparatus 100.1 includes: a magnetic field applying mechanism 11 that applies a controlled magnetic field to a region of interest of a subject 2 and irradiates an RF wave; a receiving coil 20 that receives a response wave (NMR signal) from the subject 2 and outputs an analog signal; a driving unit 21 that controls the magnetic field applied to the subject 2 and controls the transmission and reception of the RF wave; and a data processing unit 32 that sets a control sequence for the driving unit 21 and processes various data signals to generate an image.
[0283] In addition, herein, the central axis of a cylindrical bore for placing the subject 2 is taken as the Z-axis, the X-axis is defined as the horizontal direction orthogonal to the Z-axis, and the Y-axis is defined as the vertical direction orthogonal to the Z-axis.
[0284] Since the MRI apparatus 100.1 has such a structure, the nuclear spins of the nuclei constituting the subject 2 are oriented in the magnetic field direction (Z-axis) by the static magnetic field applied by the magnetic field applying mechanism 11, and precess about the magnetic field direction at the Larmor frequency inherent to the nuclei.
[0285] Moreover, when an RF pulse having the same frequency as the Larmor frequency is irradiated, the atoms resonate, absorb energy and are excited, generating a magnetic resonance phenomenon (NMR phenomenon; Nuclear Magnetic Resonance). After this resonance, when the irradiation of the RF pulse is stopped, the atoms output an electromagnetic wave (NMR signal) having the same frequency as the Larmor frequency during the relaxation process of releasing energy and returning to the original stable state.
[0286] The NMR signal output is received by the receiving coil 20 as a response wave from the subject 2, and the region of interest of the subject 2 is imaged in the data processing unit 32.
[0287] The magnetic field applying mechanism 11 includes a static magnetic field generating coil 12, an inclined magnetic field generating coil 14, an RF irradiation unit 16, and a couch 18 for placing the subject 2 in the bore.
[0288] The subject 2 lies on the couch 18, for example, in a supine position. Although not particularly limited, the subject 2 can view, for example, a screen displayed on a display 6 provided perpendicular to the Z-axis using prism glasses 4. Visual stimulation can also be applied to the subject 2 according to need through the image on the display 6. In addition, the visual stimulation given to the subject 2 can also be a structure in which an image is projected in front of the eyes of the subject 2 by a projector.
[0289] In the case of performing neurofeedback on the subject, such visual stimulation corresponds to the presentation of feedback information.
[0290] The drive unit 21 includes a static magnetic field power supply 22, an inclined magnetic field power supply 24, a signal transmission unit 26, a signal reception unit 28, and a bedding drive unit 30 that moves the bedding 18 to an arbitrary position in the Z-axis direction.
[0291] The data processing unit 32 includes: an input unit 40 that receives various operations and information inputs performed by an operator (not shown); a display unit 38 that displays various images and various information related to the region of interest of the subject 2 on a screen; a display control unit 34 that controls the display of the display unit 38; a storage unit 36 that stores programs / control parameters / image data (such as structural images) for performing various processes and other electronic data; a control unit 42 that controls the operations of each functional unit, such as generating a control sequence for driving the drive unit 21; an interface unit 44 that transmits and receives various signals to and from the drive unit 21; a data collection unit 46 that collects data composed of a set of NMR signals from the region of interest; an image processing unit 48 that forms an image based on the data of the NMR signals; and a network interface 50 that is used to communicate with a network.
[0292] In addition, except for the case where the data processing unit 32 is a dedicated computer, it also includes the case where it is a general-purpose computer that executes functions for operating each functional unit and performs specified operations, data processing, and generation of control sequences based on programs installed in the storage unit 36. Hereinafter, the data processing unit 32 will be described as a general-purpose computer.
[0293] The static magnetic field generating coil 12 generates an induced magnetic field when a current supplied by the static magnetic field power supply 22 flows through a helical coil wound around the Z axis, and generates a static magnetic field in the Z-axis direction in the cavity. The region of interest of the subject 2 is set in a region with high uniformity of the static magnetic field formed in this cavity. Here, more specifically, the static magnetic field generating coil 12 is composed of, for example, four air-core coils, and a uniform magnetic field is generated inside through this combination, giving orientation to the spins of specified atomic nuclei, more specifically hydrogen nuclei, in the body of the subject 2.
[0294] The inclined magnetic field generating coil 14 is composed of an X coil, a Y coil, and a Z coil (not shown), and is provided on the inner peripheral surface of the cylindrical static magnetic field generating coil 12.
[0295] In addition, a shim coil (not shown) is provided to improve the uniformity of the inclined magnetic field, and "shim adjustment" is performed.
[0296] These X-coils, Y-coils, and Z-coils respectively superimpose an inclined magnetic field on the uniform magnetic field in the cavity while sequentially switching the X-axis direction, Y-axis direction, and Z-axis direction one by one, so as to impart an intensity gradient to the static magnetic field. When the Z-coil is excited, the magnetic field intensity is inclined in the Z-direction to define the resonance plane. Immediately after the magnetic field in the Z-direction is applied, the Y-coil gives a short-time inclination to apply a phase modulation (phase encoding) proportional to the Y coordinate to the detection signal. Next, when data is extracted, the X-coil gives an inclination to apply a frequency modulation (frequency encoding) proportional to the X coordinate to the detection signal.
[0297] The switching of the superimposed inclined magnetic field is achieved by respectively outputting different pulse signals from the inclined magnetic field power supply to the X-coil, Y-coil, and Z-coil according to the control sequence. Thereby, the position of the subject 2 where the NMR phenomenon is detected can be determined, and the position information on the three-dimensional coordinates required to form the image of the subject 2 can be provided.
[0298] Here, as described above, three sets of orthogonal inclined magnetic fields can be used, and a slice selection direction, a phase encoding direction, and a frequency encoding direction are assigned to each inclined magnetic field. Through their combination, photography can be performed from various angles. For example, in addition to the transverse slice in the same direction as the image generally taken by an X-ray CT device, it is also possible to take pictures of the sagittal slice, coronal slice, and inclined slice perpendicular to the plane whose direction perpendicular to the plane is not parallel to the axes of the three sets of orthogonal inclined magnetic fields, etc.
[0299] The RF irradiation unit 16 irradiates an RF (Radio Frequency) pulse to the region of interest of the subject 2 according to the control sequence based on the high-frequency signal transmitted from the signal transmission unit 26.
[0300] In addition, the RF irradiation unit 16 is Figure 1 built into the magnetic field application mechanism 11, but it can also be provided on the bedding 18 or integrated with the receiving coil 20 to form a transmit-receive coil.
[0301] The receiving coil 20 is used to detect the response wave (NMR signal) from the subject 2 and is arranged close to the subject 2 in order to detect the NMR signal with high sensitivity.
[0302] Here, in the receiving coil 20, when the electromagnetic wave of the NMR signal cuts the coil wire of the receiving coil 20, a weak current is generated based on electromagnetic induction. This weak current is amplified in the signal receiving unit 28 and then sent to the data processing unit 32 after being converted from an analog signal to a digital signal.
[0303] Regarding the (transmit) receive coil 20, a multi-array coil is used as described above to improve the SN ratio.
[0304] That is, when a high-frequency electromagnetic field of a resonance frequency is applied to the subject 2 in a state where a Z-axis inclined magnetic field is applied in a static magnetic field by the RF irradiation unit 16, the specified atomic nuclei in the part where the magnetic field strength is the resonance condition, such as hydrogen atomic nuclei, are selectively excited and start to resonate. The specified atomic nuclei in the part that satisfies the resonance condition (for example, a tomogram of a specified thickness of the subject 2) are excited, and (in conventional image drawing) the spin axes rotate simultaneously. When the excitation pulse is stopped, in the receiving coil 20, an electromagnetic wave radiated from the spin axis that is rotating at this time induces a signal, and this signal is detected for a certain period of time. Through this signal, tissues containing the specified atoms in the body of the subject 2 are observed. Moreover, in order to know the position where the signal is emitted, it is configured to apply inclined magnetic fields in the X and Y directions to detect the signal.
[0305] The image processing unit 48 repeats applying an excitation signal and measures the detection signal based on the data constructed in the storage unit 36, obtains an image by restoring the resonance frequency to the X coordinate through the first Fourier transform calculation and restoring the Y coordinate through the second Fourier transform, and displays the corresponding image on the display unit 38.
[0306] For example, through such an MRI system, the above-mentioned BOLD signal is captured in real time, and the image captured in time series is subjected to the analysis processing as described later by the control unit 42, whereby MRI of resting-state functional connectivity (rs-fcMRI) can be performed.
[0307] In Figure 4 the measurement data, measurement parameters, and subject attribute data from the MRI device 100.1 and the MRI devices 100.2 to 100.Ns at other measurement locations are accumulated and stored in the storage device 210 via the communication interface 202 in the data center 200. And the calculation processing system 300 is configured to access the data in the storage device 210 via the communication interface 204.
[0308] Figure 5 is a hardware block diagram of the data processing unit 32.
[0309] As the hardware of the data processing unit 32, as described above, although not particularly limited, a general-purpose computer can be used.
[0310] In Figure 5In [the figure], in addition to a memory drive 2020 and a disk drive 2030, the computer main body 2010 of the data processing unit 32 further includes an arithmetic unit 2040, a bus 2050 connected to the disk drive 2030 and the memory drive 2020, a ROM 2060 for storing programs such as a startup program, a RAM 2070 for temporarily storing commands of application programs and providing a temporary storage space, a non-volatile storage device 2080 for storing application programs, system programs, and data, and a communication interface 2090. The communication interface 2090 corresponds to an interface unit 44 for transmitting and receiving signals to and from the drive unit 21 etc. and a network interface 50 for communicating with other computers via a network (not shown). In addition, as the non-volatile storage device 2080, a hard disk (HDD), a solid state drive (SSD), etc. can be used. The non-volatile storage device 2080 corresponds to the storage unit 36.
[0311] The arithmetic unit 2040 implements the respective functions of the data processing unit 32, such as the respective functions of the control unit 42, the data collection unit 46, and the image processing unit 48, through arithmetic processing based on program execution.
[0312] The program for causing the data processing unit 32 to execute the functions of the above-described embodiment may also be stored in a DVD-ROM 2200 or a memory medium 2210, and is further transmitted to the non-volatile storage device 2080 by being inserted into the disk drive 2030 or the memory drive 2020. The program is loaded into the RAM 2070 when it is executed.
[0313] The data processing unit 32 further includes a keyboard 2100 and a mouse 2110 as input devices, and a display 2120 as an output device. The keyboard 2100 and the mouse 2110 correspond to the input unit 40, and the display 2120 corresponds to the display unit 38.
[0314] The program for functioning as the data processing unit 32 as described above does not necessarily have to include an operating system (OS) for causing the computer main body 2010 to execute functions such as an information processing device. The program only needs to include a part of commands for calling appropriate functions (modules) in a controlled manner to obtain a desired result. How the data processing unit 32 operates is well-known, and detailed description is omitted.
[0315] In addition, the computer that executes the above program may be single or multiple. That is, centralized processing can be performed, or distributed processing can also be performed.
[0316] In addition, although there may be structural differences in the hardware within the computing processing system 300, such as parallelizing the arithmetic processing unit or using a GPGPU (General-Purpose Computing on Graphics Processing Units), the basic structure is the same as that of Figure 5 the structure shown.
[0317] (Generation process and clustering process of discriminator for disease / health based on brain functional connectivity)
[0318] Figure 6 This is a conceptual diagram that explains the process of generating a discriminator that becomes a diagnostic marker and the clustering process based on the correlation array as described in Figure 2B The discriminator is generated through a process of generating a discriminator that becomes a diagnostic marker based on the correlation array described in
[0319] As a machine learning process, for the generation of the discriminator, a process of so-called "supervised learning" is performed, and for the clustering process, a process of "unsupervised learning" is performed.
[0320] Moreover, the clustering process itself is "unsupervised learning" and does not use information such as a doctor's diagnosis. Therefore, each cluster obtained as a result of the clustering process is a group of patients obtained in a data-driven manner, and when patients are classified into subtypes, it becomes the basis for "stratification of patients" with brain functional connectivity as a feature quantity.
[0321] As shown in Figure 6 , first, in multiple MRI devices, resting-state fMRI image data is taken for the healthy group and the patient group, and the computing processing system 300 performs "preprocessing" on such fMRI image data as described later. Then, based on the measured resting-state functional connectivity MRI data, the computing processing system 300 performs brain region segmentation processing for each subject and derives a correlation array of the activity levels between brain regions (between regions of interest).
[0322] Next, for the off-diagonal components of the correlation array, the corresponding measurement bias is derived in advance as described later, and the computing processing system 300 performs normalization processing by subtracting the measurement bias from the element values of the correlation array.
[0323] Furthermore, between the element values of the correlation array after normalization processing and the disease labels (labels indicating disease or health) for each subject, the computing processing system 300 suppresses overlearning and generates an identifier with feature selection as a "discriminator generation process based on ensemble learning" as described later, thereby generating a disease identifier (diagnostic marker) that can predict the disease or health of the subject.
[0324] On the other hand, the computational processing system 300 performs a feature quantity selection process of selecting a feature quantity for clustering as described below from the feature quantities (brain functional connections) determined in the generation process of generating an identifier for disease labels in integrated learning, and then performs a multi-co-clustering process by "unsupervised learning".
[0325] Next, each process in Figure 6 will be described in more detail.
[0326] [Outline from preprocessing to generation and clustering processes for disease identification]
[0327] (Preprocessing and calculation of the resting-state functional connectivity FC matrix)
[0328] For example, the first 10 seconds of the measured fMRI data are discarded due to consideration of T1 equilibrium.
[0329] In the preprocessing step, the computational processing system 300 performs processes such as temporal layer correction, recalibration processing for correcting body motion artifacts observable in the head, co-registration of brain functional images (EPI images) and morphological images, distortion correction, segmentation of T1-enhanced structural images, normalization to the Montreal Neurological Institute (MNI) space, and spatial smoothing using an isotropic Gaussian kernel with a full width at half maximum of, for example, 6 mm.
[0330] Regarding such a preprocessing pipeline process, it is publicly available, for example, at the following site.
[0331] http: / / fmriprep.readthedocs.io / en / latest / workflows.html
[0332] (Partitioning of brain regions (Segmentation: Parcellation))
[0333] Regarding the segmentation of brain regions, although not particularly limited, it can be performed by a "surface-based method" according to the above "Method 3".
[0334] (Physiological noise regression)
[0335] Physiological noise regression is performed by applying CompCor disclosed in the following Literature 11. All the descriptions in the publicly known Literature 11 are incorporated herein by reference.
[0336] Known Document 11: Behzadi, Y., Restom, K., Liau, J., and Liu, T.T. (2007). A component-based noise correction method (CompCor) for BOLD and perfusion based fMRI. Neuroimage 37(1), 90 - 101. doi:10.1016 / j.neuroimage.2007.04.042.
[0337] To remove some spurious sources (extra signal sources), linear regression with six motion parameters and whole-brain equivalent regression global parameters was used.
[0338] (Temporal filtering)
[0339] The processing system 300 uses a Butterworth filter with a passband between, for example, 0.01 Hz and 0.08 Hz as a temporal bandpass filter and applies this Butterworth filter to the time series data to limit the analysis to low-frequency fluctuations that are characteristic of BOLD activity.
[0340] (Head movement)
[0341] Frame-wise displacement (FD) is calculated in each functional session. To reduce spurious variations in functional connectivity (FC) caused by head movement, for example, volumes with FD > 0.5 mm are removed.
[0342] FD represents head movement between two temporally consecutive volumes as a scalar (i.e., the sum of absolute displacements in translational and rotational movements).
[0343] For example, in the specific example described later, in the above-mentioned specific dataset, when the ratio of volumes removed after cleaning exceeds (mean ± 3 × standard deviation), the data of that participant is excluded from the analysis. As a result, 35 participants were removed from the entire dataset. Thus, 683 participants (545 HC, 138 MDD) were used in the training dataset, and the data of 444 participants (263 HC, 181 MDD patients) were used in the independent validation dataset for the following analysis.
[0344] (Calculation of functional connectivity (FC) matrix)
[0345] In a specific example of the present embodiment, after dividing the region by the above-described division method, for each participant, functional connectivity FC is calculated for 379 regions of interest (ROIs) as the temporal correlation of the BOLD signal.
[0346] In the calculation of functional connectivity, although not particularly limited, the Pearson correlation coefficient is used here.
[0347] During the time course of the preprocessed BOLD signal of the ROIs in each possible group, the Pearson correlation coefficient after Fisher-z transformation is calculated, and a 379-row × 379-column symmetric connectivity matrix is constructed, where the elements represent the connection strength between two ROIs.
[0348] And, for analysis, the values of 71,631 (= (379 × 378) / 2) functional connectivity FCs in the lower triangular array of the connectivity matrix are used.
[0349] (For the harmonization process of brain activity biomarkers)
[0350] When collecting big data related to mental diseases, as described above, it is almost impossible to collect large-scale brain image data (connectomes related to human diseases) at one location. Therefore, it is necessary to collect image data from multiple locations.
[0351] It is very difficult to completely control the type of MRI device (scanner), protocol, and patient population. Thus, to analyze the collected data, brain image data taken under different conditions is used.
[0352] In particular, disease factors tend to be intertwined with location factors. Therefore, when extracting disease factors by applying machine learning techniques to data under such different conditions, the differences between locations become the biggest obstacle.
[0353] One location (or hospital) often samples only a few types of mental diseases (for example, mainly samples schizophrenia from location A, mainly samples autism from location B, mainly samples major depressive disorder from location C, etc.). Thus, intertwining occurs.
[0354] To appropriately manage data under such different conditions, it is necessary to harmonize the data between locations.
[0355] Differences between locations essentially include two types of biases.
[0356] They are technical biases (that is, measurement biases) and biological biases (that is, sampling biases).
[0357] Measurement bias includes differences in the characteristics of the MRI scanner such as imaging parameters, the intensity of the electric field, the MRI device manufacturer, and the type of scanner, and sampling bias is associated with differences in the groups of subjects between locations.
[0358] Therefore, "harmonization processing" for compensating for such differences between locations is required. The details of this harmonization processing are described in Non-Patent Document 8 (Ayumu Yamashita et al.) mentioned above, and will be described later.
[0359] (Disease identifier based on ensemble learning)
[0360] In this specification, the term "ensemble learning" is assumed to refer to the following processing: resampling from the original learning data to create K sets of learning data, and for each learning data, independently generating K identifiers through machine learning processing, and integrating these K identifiers to generate a discriminator.
[0361] In particular, here it is for the purpose of discriminating whether a subject has a disease or is healthy based on the brain functional connectivity pattern of a certain subject for a specific disease, so each identifier is an identifier for two types of recognition problems.
[0362] Moreover, when creating K sets of learning data by resampling from the original learning data, "undersampling" and "downsampling" are performed as described later.
[0363] Here, it is also possible to use an "identifier based on learning processing with feature quantity selection" such as an identifier based on L1 regularization (LASSO (Least Absolute Shrinkage and Selection Operator) method), a regularization learning method such as the ridge regularization method (L2 regularization).
[0364] Here, the "regularization learning method" refers to the following learning method: although the entire feature quantity of the original training data is used as the object of learning, in the learning algorithm, when making the model learn, a penalty for increasing complexity is set, and the learning model that minimizes the sum of this penalty and the training error is obtained, thereby aiming to improve the generalization performance. Moreover, L1 regularization uses the sum of the absolute values of the parameters (corresponding to the feature quantities) of the learning model as the penalty, and L2 regularization uses the sum of the squares of the parameters of the learning model as the penalty. In addition, there may also be L0 regularization that uses the number of feature quantities used in the model itself as the penalty.
[0365] In addition, the LASSO method (L1 regularization) is a technique capable of performing so-called sparse estimation. As its derivative forms, there are the Elastic Net method, Group Lasso method, Fused Lasso method, Adaptive Lasso method, Graphical Lasso method, etc.
[0366] On the other hand, as an "identifier based on learning processing accompanied by feature quantity selection", it is also possible to use a technique such as the "random forest method" that selects feature quantities during the generation of the identifier and simultaneously obtains the importance of the feature quantities.
[0367] Furthermore, as such a "generation technique for generating an identifier through ensemble learning", although the explanation will be centered on the LASSO method hereafter, this technique is not limited to the techniques described above. For example, it can also be the following ensemble learning methods: Use the technique of dictionary learning as a segmentation method to set the analysis target brain region in a data-dependent manner, use the tangent method (tangent-space covariance) as the value of brain functional connectivity, and use the ComBat method described later to perform between-facility calibration on the brain functional connectivity FC within the dataset, and generate an identifier through ridge regularization. The segmentation method, the calculation method of brain functional connectivity, the harmonization method, the generation method of the identifier, etc. can also be other combinations. For example, as the calculation method of brain functional connectivity, it is also possible to use distance correlation. Here, "distance correlation" is a method that calculates the similarity of activity patterns without averaging the activity patterns within the brain region like the Pearson correlation method, and is disclosed in the following publicly known documents 12 and 13. The entire descriptions of publicly known documents 12 and 13 are hereby incorporated by reference herein.
[0368] Publicly known document 12: G.J. Szekely, M.L. Rizzo, N.K. Bakirov, Measuring and testing dependence by correlation of distances, Ann. Statist., 35(6)(2007), pp. 2769-2794
[0369] Publicly known document 13: https: / / en.wikipedia.org / wiki / Distance_correlation
[0370] As will be described later, in the present embodiment, in such ensemble learning, when making the identifier learn, the "importance" for realizing the function of identifying each feature quantity is determined.
[0371] (Feature quantity selection and clustering for clustering)
[0372] As "ensemble learning", a process is performed for K sets of learning data to further determine a set of second feature quantities for performing clustering by "unsupervised learning" from the union of "first feature quantities" used in generating each of the K recognizers during the process of generating the K recognizers.
[0373] Although not particularly limited, the "importance" is determined as follows, for example.
[0374] i) When the learning method for generating K recognizers by ensemble learning is "learning process with feature quantity selection", in the union of the "first feature quantities" selected in generating the recognizers, the feature quantities are sorted according to the frequency of use in generating the K recognizers.
[0375] ii) When the learning method for generating K recognizers by ensemble learning is a method such as the "random forest method" that can obtain the importance of feature quantities in generating the recognizers, a structure can be set to generate a sorted list of feature quantities according to such importance.
[0376] iii) When the learning method for generating K recognizers by ensemble learning is the "ridge regularization method (L2 regularization)" (not necessarily with feature quantity selection) and is a method for generating recognizers with the weighted sum of feature quantities as the independent variable, a sorted list of feature quantities is generated with the median value obtained by summing the absolute values of the weight coefficients of each feature quantity in each of the K recognizers as the importance. In addition, as the importance, it is not limited to such a "median value", and other statistical representative values such as "cumulative value obtained by cumulatively summing over K recognizers" can also be used, for example.
[0377] A structure can be set such that a specified number of feature quantities from the top of the sorted list generated as in i) to iii) are used as the "second feature quantities".
[0378] In addition, as a condition for determining the "second feature quantities", it is not limited to a specified number from the top of the sorted list. For example, it can also be set as a condition that the frequency of use of the feature quantities in the sorted list is above a specified frequency (using the frequency of being selected at a certain ratio or more in generating the K recognizers).
[0379] As described above, based on the selected feature quantities, clustering processing (patient stratification) is performed by the "multi-co-clustering method" which is an unsupervised learning as described later.
[0380] [Recognizer generation processing for classification into two classes]
[0381] Next, a more detailed description will be given Figure 6 of the generation of an integrative learning-based recognizer in the described processing.
[0382] That is, as an example of the classifier generation process for classifying into two categories, more specifically, the process of using a learning dataset as training data for a disease recognizer (a two-category classifier for "healthy" or "disease") to construct a biomarker for MDD will be described.
[0383] Here, taking major depressive disorder among mental illnesses as an example, that is, taking the group of patients diagnosed with major depressive disorder by a doctor through a conventional symptom-based diagnosis method as an example, the process of generating a classifier will be described. Moreover, the Figure 8 example of the process executed by the disease recognizer generation unit 3008 shown for generating a classifier that outputs auxiliary information for discriminating the diagnosis between the patient group and the healthy group will be described.
[0384] Therefore, next, the process of constructing an MDD recognizer for identifying the healthy group (HC) and MDD patients based on functional connectivity FC will be described.
[0385] Next, as an "recognizer based on a learning process with feature quantity selection" for creating a disease recognizer (MDD recognizer), an example will be given using a learning method (LASSO method) of a recognizer based on L1 regularization for description.
[0386] Moreover, as will be described later, in order to determine the functional connectivity FC related to MDD diagnosis, feature quantities used in clustering are selected according to the "importance" of each functional connectivity FC for constructing a disease recognizer.
[0387] Figure 7 and Figure 8 are functional block diagrams showing the structure of the computing processing system 300 that performs reconciliation processing, disease recognizer generation processing, clustering classifier generation processing, and discrimination processing based on the data stored in the storage device 210 of the data center 200.
[0388] In addition, here, it is assumed that the "discrimination processing" includes discrimination of diseases (discrimination between disease and health) and classification processing for determining which "cluster" (subtype) the subject being targeted belongs to.
[0389] Referring to Figure 7 , the computing processing system 300 includes: a storage device 2080 for storing data from the storage device 210 and data generated during the computation; and an arithmetic device 2040 for performing arithmetic processing on the data in the storage device 2080. As the arithmetic device 2040, for example, a CPU is suitable.
[0390] The arithmetic unit 2040 includes: a correlation array calculation unit 3002 that calculates elements of a correlation array for MRI measurement data 3102 of a patient group and a healthy group by executing a program, and stores the elements of the correlation array as correlation array data 3106 in a storage device 2080; a coordination calculation unit 3020 that executes a coordination process; and a learning and discrimination processing unit 3000 that, based on the result of the coordination process, executes a generation process of a disease recognizer, a generation process of a clustering classifier, and a discrimination process using the generated disease recognizer or clustering classifier.
[0391] Figure 8 is to explain in more detail Figure 7 the functional block diagram of the structure.
[0392] In addition, Fig. 9 is a flowchart for explaining the process of machine learning for generating a disease recognizer based on ensemble learning.
[0393] Therefore, first, as Figure 6 shown, with reference to Figure 8 and Fig. 9 the processes up to the coordination process and the generation process of generating a recognizer (disease recognizer) through ensemble learning will be explained.
[0394] First, it is assumed that fMRI measurement data of subjects (healthy and patients), attribute data of subjects, and measurement parameters have been collected in the storage device 210 of the data center 200 as a "learning dataset".
[0395] With reference to Figure 8 and Fig. 9 , such a learning dataset is used as training data for a disease recognizer (a two-class classifier for "healthy" or "disease") to construct a biomarker for MDD discrimination. That is, this biomarker is used to identify a healthy group (a group of individuals with a diagnostic label of being healthy (HC)) and an MDD patient group (a group of individuals with a diagnostic label of major depressive disorder) based on the values of 71,631 functional connections FC.
[0396] As will be explained below, in the learning process for generating a recognizer for MDD (hereinafter referred to as the "MDD recognizer"), a logistic regression analysis (one of the sparse modeling methods) based on L1 regularization (LASSO method) is used to select the best subset of functional connections FC from 71,631 functional connections FC.
[0397] Generally, when using L1 regularization, several parameters (weight elements in the following explanation) can be set to 0. That is, feature selection is performed to obtain a sparse model.
[0398] Among them, as a sparse modeling method, it is not limited to the LASSO method. As will be described later, other methods can also be used, such as using sparse logistic regression (SLR: Sparse Logistic Regression) that applies the variational Bayesian method to logistic regression, etc.
[0399] Refer to Fig. 9 , when starting the learning process of the MDD recognizer (S100), using the learning dataset prepared in advance (stored in the storage device 2080), the correlation array calculation unit 3002 calculates the components of the connection matrix.
[0400] Next, the harmonization calculation unit 3020 performs a harmonization process (S104) using the calculated measurement bias.
[0401] As will be described later, the desired harmonization process is a method using multi-facility examinees, but other methods can also be used.
[0402] For example, it can also be configured to perform dataset inter-harmonization between the discovery dataset and the independent validation dataset using the combat method described later.
[0403] Next, the disease recognizer generation unit 3008 generates an MDD recognizer for the learning data by a method called the "ensemble learning method", that is, a method that modifies the "nested cross-validation" method.
[0404] First, the disease recognizer generation unit 3008 divides the learning data into 10 parts (S106) by setting K = 10, for example, in order to perform a learning process on the learning data using "K-fold cross-validation" (K: natural number) (outer cross-validation).
[0405] That is, the disease recognizer generation unit 3008 uses one part of the K parts (10 parts) of the dataset as the "test dataset" for verification, and sets the remaining (K - 1) parts (9 parts) of the data as the training dataset (S108, S110).
[0406] Next, the disease recognizer generation unit 3008 performs "undersampling processing" and "downsampling processing" on the training dataset (S112).
[0407] Here, the "undersampling processing" refers to the following process: in the training dataset, when the number of data corresponding to specific attribute data (two or more) to be classified is inconsistent, in order to make the number the same, the data with the larger number of attributes is removed to become the same number.
[0408] Here, it is equivalent to the following situation: in the training dataset, since the number of subjects in the MDD patient group is not equal to the number of subjects in the healthy group, a process for making them consistent is performed.
[0409] Moreover, the "downsampling process" refers to a process of randomly extracting a specified number of samples from the training dataset.
[0410] That is, in the K-fold cross-validation repeated through steps S108 to S118 and S122, in each cross-validation, the training dataset is unbalanced in terms of the number of MDD patients and healthy controls (HC). Therefore, an undersampling method for constructing a classifier is adopted. As the downsampling process, a specified number, for example, 130 MDD patients and the same number of 130 healthy controls are randomly sampled from the training dataset.
[0411] In addition, the value of 130 is not limited to such a value and is determined in the following manner: it can be appropriately undersampled as described above according to the number of data in the learning dataset (683 people in "Dataset 1" described later), the number of folds K (here, for example, K = 10), and the degree of imbalance in the number of data included in the specific attribute to be classified.
[0412] Performing such a downsampling process is because, in undersampling, it is disadvantageous in that the recognizer can no longer use the excluded data for learning. To eliminate this disadvantage, a process of randomly extracting M times (M: natural number, for example, M = 10) (that is, downsampling) is repeated.
[0413] In addition, as will be described later, performing such an undersampling process and downsampling process also has technical significance for "feature quantity selection" when generating a "classifier" for "stratification". Therefore, this will be described later.
[0414] Next, the disease recognizer generation unit 3008 performs hyperparameter adjustment processing (S114.1 to S114.10) on each of the downsampled subsamples 1 to 10.
[0415] Here, in each subsample, a recognizer submodel is generated by using the following logical function. This logical function is used to define the possibility of participants belonging to the MDD class within the subsample.
[0416] [Equation 1]
[0417]
[0418] Among them, y sub represents the class label of the participant (MDD, y = 1; HC, y = 0), and c subThe FC vector represents a given participant, and w represents the weight vector.
[0419] The weight vector w is determined (LASSO calculation) in such a way as to minimize the following evaluation function (cost function).
[0420] [Equation 2]
[0421] When set to t j = P j (y j = 1|c j ; ω)
[0422]
[0423] In the LASSO calculation, in the cost function, there is a sum (L1 norm) of the absolute values (first order) of the elements of the weight vector as the second term.
[0424] Here, λ represents a hyperparameter that controls the amount of shrinkage applied to the evaluation.
[0425] In each subsample, although not particularly limited, the disease identifier generation unit 3008 uses a specified number of data as hyperparameter adjustment data and determines the weight vector w using the remaining data (e.g., data of n = 250 people or 248 people). At this time, although not particularly limited, the disease identifier generation unit 3008 sets the hyperparameter λ to 0 < λ ≤ 1.0, for example, and uses each value of λ obtained by dividing this interval into P equal parts (P: natural number), for example, 25 equal parts, and determines the weight vector w through the LASSO calculation as described above.
[0426] At this time, as described above, as "nested structure cross-validation", the adjustment of the hyperparameter is performed in the form of "inner cross-validation". In the inner cross-validation, the "test data set" of the outer cross-validation is not used.
[0427] On this basis, the disease identifier generation unit 3008 compares and discriminates the performance (e.g., accuracy) of the hyperparameter adjustment data through a logical function corresponding to each generated value of λ, and determines the logical function (hyperparameter adjustment process) corresponding to the λ with the highest discrimination performance.
[0428] Next, the disease identifier generation unit 3008 sets the "identifier sub-model" to output the average (S116) of the output values of the logical functions corresponding to each subsample generated in the current cross-validation loop. In terms of determining the recognition performance based on the average of the output values of the identifier calculated in each subsample, it can also be regarded as a kind of "ensemble learning".
[0429] The disease recognizer generation unit 3008 validates the recognizer sub-model generated in the current cross-validation loop by using the test data set prepared in step S110 as input (S118).
[0430] In addition, as a method of generating sub-samples by undersampling and downsampling and performing feature selection in each sub-sample to generate a recognizer sub-model, other sparse modeling techniques can be used in addition to the method of performing the LASSO method and hyperparameter tuning as described above.
[0431] When the disease recognizer generation unit 3008 determines that the loop of K-fold (here, 10-fold) cross-validation has not ended (S122: "No"), it sets another partial data set different from the data used in the loop so far in the data divided into K parts as the test data set, and sets the remaining partial data set as the training data set (S108, S110), and repeats the process.
[0432] On the other hand, when the loop of K-fold (10-fold) cross-validation ends (S122: "Yes"), the disease recognizer generation unit 3008 outputs the average of the outputs of K×M (in this case, 10×10 = 100) logical functions (recognizers) for the input data, and generates a recognizer model (MDD recognizer) for MDD (S120).
[0433] As a result, the meaning that the MDD recognizer takes the average of the outputs of K×M recognizers as its recognition output can be said to be a "recognizer" obtained as a result of "ensemble learning".
[0434] When the output (diagnostic probability value) of the MDD recognizer exceeds 0.5, it can be regarded as an index indicating an MDD patient.
[0435] Moreover, in the present embodiment, as evaluation indexes for the performance of the MDD recognizer generated in this way, the Matthews correlation coefficient (MCC), the area under the ROC curve (AUC: area under the curve) for the ROC curve (Receiver Operating Characteristic curve), accuracy, sensitivity, and specificity are used.
[0436] In addition, the method of generating an identifier for an object disease (e.g., MDD) using the feature quantities selected by performing feature selection in each subsample (in this case, the elements of the correlation array after harmonizing measurement bias) is not limited to the method based on the averaging process of the outputs of such multiple identifier submodels, and can also be set to a process based on majority decision, or other modeling methods, particularly other sparse modeling methods, can be used for the feature quantities selected by performing feature selection to generate the structure of the identifier.
[0437] (Examples and Performance of Data Used in MDD Identifier)
[0438] As already described, for the construction of reliable classifiers and regression models using machine learning algorithms, it is necessary to use data with a large sample size collected from a large number of shooting locations.
[0439] Therefore, below, a study is conducted using a learning resting-state fMRI dataset of approximately 700 participants including MDD patients collected from four different shooting locations.
[0440] Fig.10 It is a graph showing the population characteristics of such a learning dataset (Dataset 1).
[0441] Dataset 1 is the data in the above-mentioned SRPBS.
[0442] Fig.11 It is a graph showing the population characteristics of an independent verification dataset (Dataset 2).
[0443] Dataset 2 is basically also the data in the above-mentioned SRPBS.
[0444] That is, in the following analysis, two resting-state functional MRI (rs-fMRI) datasets are used.
[0445] (1) As Fig.10 shown, Dataset 1 contains data of 713 participants (a healthy group HC of 564 people from 4 locations, an MDD patient group of 149 people from 3 locations).
[0446] (2) As Fig.11 shown, Dataset 2 contains data of 449 participants (a healthy group HC of 264 people from 4 locations, an MDD patient group of 185 people from 4 locations).
[0447] In addition, the Beck Depression Inventory (BDI) II obtained from most of the participants in each dataset is used simultaneously to evaluate "depressive symptoms".
[0448] Dataset 1 is the "learning dataset" and is used to construct the recognizer of MDD and the classifier of clustering.
[0449] The measurements of the participants were performed in a single 10-minute resting-state functional MRI (rs-fMRI) session.
[0450] Here, the resting-state functional MRI data were also acquired under a unified imaging protocol (http: / / www.cns.atr.jp / rs-fmri-protocol-2 / ).
[0451] However, it was actually difficult to ensure that image acquisition was performed with the same parameters at all sites. In the measurement, two phase modulation directions (P→A and A→P), two MRI device manufacturers (Siemens and GE), three different coil numbers (12, 24, 32), and three scanner models were used.
[0452] In the scanning of resting-state functional MRI, the participants were instructed as follows in principle.
[0453] "Please relax. Please stay awake. Please fix your eyes on the crosshair in the center and don't think about specific things."
[0454] The "population characteristics" in the dataset are the characteristics used in so-called "demographics" and include attributes in the table such as diagnosis names in addition to age, gender, etc.
[0455] In addition, in Fig.10 and Fig.11 the number in parentheses indicates the number of participants with BDI score data.
[0456] The population distribution in all learning datasets shows no statistically significant difference between the MDD and HC subgroups (p > 0.05).
[0457] Dataset 2 is the "independent validation dataset" and is used to test the classifier of MDD and the classifier of clustering.
[0458] The sites where Dataset 2 was acquired are not included in Dataset 1.
[0459] The population distribution of age is consistent between the MDD and HC subgroups in the independent validation dataset (p > 0.05), but the population distribution of gender is inconsistent between the MDD and HC subgroups in the independent validation dataset (p < 0.05).
[0460] (Control of site effect)
[0461] In addition, hereinafter, a method for harmonizing multi-site subjects in a learning dataset as described below is used to control the site effect on the functional connection FC for explanation.
[0462] Among them, as the harmonization method, it is not limited to this method. For example, other methods such as the ComBat method can also be used.
[0463] In addition, the ComBat method is disclosed, for example, in the following well-known document 14. All the descriptions in well-known document 14 are incorporated herein by reference.
[0464] Well-known document 14: Johnson WE, Li C, Rabinovic A. "Adjusting batch effects in microarray expression data using empirical Bayes methods." Biostatistics 8, 118 - 127 (2007).
[0465] By using the method for harmonizing multi-site subjects, simple inter-site differences (measurement biases) can be removed.
[0466] In addition, for the sites included in the independent validation dataset, there is no dataset of multi-site subjects. Therefore, in order to control the site effect in the independent validation dataset, a harmonization method based on the ComBat method is used.
[0467] Fig.12 It is a graph showing the prediction performance (output probability distribution) of MDD for the learning dataset for all shooting sites.
[0468] For the learning dataset, in the output from the recognizer model, the probability distributions of the two diagnoses corresponding to the populations of MDD patients and healthy subjects are clearly separated into right (MDD) and left (HC) by a threshold of 0.5.
[0469] The recognizer model separates MDD patients from the HC population with an accuracy of 66%.
[0470] The corresponding AUC is 0.77, showing high discriminative power.
[0471] In addition, the MCC is approximately 0.33.
[0472] Fig.13 It is a graph showing the prediction performance (probability distribution of the output of the recognizer) of MDD for the learning dataset for each shooting site.
[0473] From Fig.13 It can be seen that not only for the entire dataset, but also for each dataset at the three shooting locations (Location 1, Location 2, and Location 4), a nearly identical high classification accuracy was achieved.
[0474] In addition, although only the healthy group was included in the dataset at Location 3 (SWA), its probability distribution was equivalent to that of the healthy group at other locations.
[0475] (Generalization performance of the recognizer)
[0476] Fig.14 It is a graph showing the probability distribution of the output of the recognizer for MDD in an independent validation dataset.
[0477] That is, an independent validation dataset was used to test the generalization performance of the recognizer model.
[0478] For MDD, in Fig.12 the process, 100 recognizers of logical functions (10-fold × 10 downsamplings) were generated by machine learning, and the independent validation dataset was input into all 100 generated recognizers (the recognizer model as a set of recognizers).
[0479] Then, for each participant, the average of the outputs of the 100 recognizers (diagnostic probability) was obtained. When the average diagnostic probability value > 0.5, it was set as the diagnostic label for that participant, indicating compliance with major depressive disorder.
[0480] In the independent validation dataset, the generated recognizer model separated the MDD individual group from the HC individual group with an accuracy of approximately 70%.
[0481] The corresponding AUC was 0.75, showing a high recognition ability (permutation test p < 0.01).
[0482] For the independent validation dataset, in the output from the recognizer model, the probability distributions of the two diagnoses corresponding to the MDD patients and the healthy individual groups were clearly separated into right (MDD) and left (HC) by the threshold of 0.5.
[0483] The sensitivity was 68% and the specificity was 71%. This resulted in a relatively high MCC value of 0.38 (permutation test p < 0.01).
[0484] Fig.15 It is a graph showing the probability distribution of the output of the MDD recognizer for the independent validation dataset for each shooting location.
[0485] It can be learned that not only for the entire datasets at the four shooting locations, but also for each dataset, a high classification accuracy can be achieved.
[0486] [Clustering Process for Subject Data]
[0487] Next, a more detailed description will be given Figure 6 of the selection of feature quantities for clustering and the clustering process based on the selected feature quantities in the described process.
[0488] That is, Figure 6 the processes described as "feature quantity selection" and "clustering process" in Figure 8 will be described as the processes executed by the disease recognizer generation unit 3008 and the clustering classifier generation unit 3010 in
[0489] Fig.16 is a flowchart for explaining the process of selecting feature quantities and performing clustering by unsupervised learning.
[0490] Next, the following process will be described: in the learning process of the "two-class recognizer" as described in Fig. 9 the feature quantities (brain functional connections) used in the generation of each recognizer sub-model are sorted, and a specified number of feature quantities starting from the highest order are used to perform clustering by unsupervised learning.
[0491] As described above, the brain functional connection has a high dimension of more than 70,000 dimensions corresponding to the brain segmentation method. If it is a normal method, it is generally difficult to perform clustering processing by unsupervised learning. In the method of this embodiment, for such a clustering problem, in the "recognizer generation using supervised learning", sorting corresponding to the importance of the feature quantities is performed, and "clustering based on unsupervised learning" based on the feature quantities selected according to this sorting is combined, so that such clustering processing can be performed.
[0492] In addition, next, for convenience, the clustering process will be described as a process different from the generation process of the disease recognizer. However, in Fig.16 steps S200 to S210 are equivalent processes to steps S100 to S120 in Fig. 9 and the generation process of the disease recognizer and the clustering process can be executed as a series of processes.
[0493] Referring to Fig.16 , when starting the clustering learning process, the disease recognizer generation unit 3008 prepares the subject data of Nh healthy people and Nm people with depression (S202), and the disease recognizer generation unit 3008 performs brain region parcellation processing, calculation of brain functional connection values, and normalization processing on the functional brain activity data of the subjects (S204).
[0494] Next, the disease recognizer generation unit 3008 performs data division for Ncv-fold cross-validation (Ncv: natural number, and Ncv ≥ 2). For each divided data, a training data set and a test data set are prepared, and undersampling and downsampling are performed on each training data set to generate Ns subsets of test data (S206).
[0495] Moreover, for each subsample obtained by downsampling, the disease recognizer generation unit 3008 generates a recognizer by a learning method accompanied by feature quantity selection (S208).
[0496] In addition, here, it is assumed that Fig. 9 feature quantity selection is performed by L1 regularization (LASSO) in the same way.
[0497] For the learning data sets divided into Ncv, the training data set ((Ncv - 1) of the divided data sets) and the test data set (1 of the divided data sets) are reorganized in sequence, and the processing of steps S206 to S208 is repeated until Ncv times of cross-validation are performed.
[0498] An integrated recognizer with the average of (Ns × Ncv) recognizers generated in this way as the output is generated as a disease recognizer (diagnostic marker) (S210).
[0499] As described above, the processing up to this point is the same as the processing of steps S100 to S120 in Fig. 9 .
[0500] On the other hand, in the process of repeating steps S206 to S208 Ncv times, the clustering classifier generation unit 3010 sorts the feature quantities (brain functional connections) selected when generating a recognizer by a learning method accompanied by feature quantity selection in the union set. Although there is no particular limitation, the feature quantities in this union set are sorted according to the number of times they are selected (S220).
[0501] Here, the "number of times of being selected as a feature quantity" is called the importance of this feature quantity in this sorting.
[0502] In other words, for example, according to the Fig. 9 example shown, 100 (= 10 × 10) recognizers are generated by the LASSO method. For the brain functional connections with non-zero weights in each recognizer, the selection times are counted in the way of +1. They are sorted in descending order of the counted times as important connections.
[0503] Next, the clustering classifier generation unit 3010 selects, for example, a specified number of feature quantities from the above union set according to the importance in order to perform clustering on the group of depression patients by unsupervised learning (S222).
[0504] Further, the clustering classifier generation unit 3010 performs clustering processing (S224) using the multi co-clustering method as described below as a method of unsupervised learning.
[0505] Through the above processing, the clustering classifier generation unit 3010 generates a clustering classifier for the group of depression patients (S226).
[0506] Specifically, through the above processing, the clustering classifier generation unit 3010 determines, for each cluster, a model for generating the probability distribution of such observed data based on the observed data, and stores the information of each model in the storage device 2080. Further, the discrimination value calculation unit 3012, as a clustering classifier, calculates, for input data other than the learning data, the posterior probability that the input data belongs to each cluster based on the models of the respective probability distributions, and outputs a classification result that the input data belongs to the cluster with the maximum posterior probability (MAP estimation method (Maximum A posteriori Probability estimation method)).
[0507] In addition, in the above description, it is assumed that the clustering process is performed based on the generation process of the identifier implemented on the subject data of Nh healthy people and Nm depression patients by the disease identifier generation unit 3008. However, the clustering method of the present embodiment is not limited to such a case, and can also be used for clustering of patient groups of diseases other than the "group of depression patients", such as "group of schizophrenia patients", "group of autism patients", "group of obsessive-compulsive disorder patients", etc.
[0508] Or, more generally, for an attribute that makes the relationship between the attribute labels (such as the personality of the person, the field of expertise of the person, etc.) classified by a person according to experience and the pattern of the correlation between the regions of the temporal change of brain activity clear, it can also be used to perform "clustering of the group of subjects classified according to the attribute label" (classification into subtypes) on the subjects classified according to the attribute label in a data-driven manner.
[0509] (Undersampling and downsampling processing)
[0510] In the processing described above, "undersampling and downsampling processing" is performed, so its technical significance will be briefly described.
[0511] First, regarding the effect of the "undersampling process", it is cited that the boundary of recognition in the identifier is appropriately set.
[0512] For example, in the case of considering a two-class classification task, in the learning data, the less the deviation in the number of data belonging to each class, the higher the accuracy of the evaluation of the performance (e.g., accuracy) of the recognizer in the processing flow.
[0513] In Fig. 9 In steps S114.1 to 114.10 in , in the setting of hyperparameters, the process of "determining the logical function corresponding to the λ with the highest discrimination performance" is performed, so it is necessary to accurately evaluate the "discrimination performance".
[0514] As an extreme example, in the case of learning a recognizer for learning data where the number of data belonging to class 1 is 100 and the number of data belonging to class 2 is 1, even if the recognizer determines that all data is class 1, there will be a situation where it does not have a great impact on accuracy, etc. In this regard, it is meaningful to make the number of data belonging to the two classes consistent through random sampling.
[0515] In addition, regarding downsampling, for the following reasons, its premise is originally to be implemented multiple times together with undersampling.
[0516] First, even if it is assumed that undersampling and downsampling are performed through random sampling, if it is only a single process, there may still be a deviation in the data.
[0517] Second, as will be described below, in Fig. 9 the generation of the "recognizer" repeatedly performed in steps S108 to S122 in is implemented through "ensemble learning" as described above.
[0518] At this time, in the generation of each recognizer, the importance of the feature quantity for recognition is determined.
[0519] In such a way that in the "learning process with feature quantity selection", this feature quantity is selected, or in the "learning process without feature quantity selection", the weight of this feature quantity for recognition is calculated, and the importance is determined according to the contribution degree of each feature quantity to recognition.
[0520] Next, taking the "learning process with feature quantity selection" as an example, the significance of the "undersampling and downsampling process" in the determination of such importance will be described. In addition, even in the "learning process without feature quantity selection" such as L2 regularization, it can be considered that the event of "the weight of this feature quantity increases" occurs for basically the same technical reasons as the event of "being selected as a feature quantity".
[0521] Here, as the "learning process with feature quantity selection", for example, a so-called "sparse modeling" method such as the above-mentioned LASSO method is cited.
[0522] In sparse modeling, feature quantities are sparsely selected, that is, by making the weights of specific feature quantities non-zero, and in contrast, making the weights of other feature quantities zero, thereby selecting feature quantities. As one of the reasons for implementing such sparse selection of feature quantities, the following situation is cited: there is a "penalty term corresponding to the number of feature quantities" for performing learning processing, such that in the case where there is a "group of feature quantities that contribute to similarity" for "discrimination (recognition) processing", one feature in the group is selected and the weights of other feature quantities in the group are made zero. Regarding the LASSO method, this tendency is particularly significant.
[0523] That is, in "discrimination processing", if feature quantity A participates in the same way as feature quantity B, for example, when the correlation between feature quantity A and feature quantity B is high, even if only feature quantity A is selected as the feature quantity, discrimination processing can be performed without degrading the discrimination performance.
[0524] However, for example, in clustering, it is assumed that both feature quantity A and feature quantity B need to be considered. However, if such "sparsification" is performed and feature quantities are selected only based on the contribution degree to the discrimination processing in the generation process of a single recognizer, it may be insufficient for "feature quantity selection" for clustering.
[0525] Fig.17 It is a diagram showing the concept of implementing feature quantity selection through such "learning processing with feature quantity selection" in the case where there are multiple (for example, Nch) feature quantities.
[0526] Refer to Fig.17 and assume that the subject group includes a healthy group and a patient group.
[0527] Assume that the subjects in the healthy group correspond to the label H, and the healthy group includes subtypes h1 and h2.
[0528] Assume that the subjects in the patient group correspond to the label M, and the patient group includes subtypes m1, m2, and m3.
[0529] Here, for the observables, it is unknown how many subtypes the healthy group and the subject group are divided into, and the recognition labels of the subtypes are not clearly corresponding to the subjects, which are potential labels.
[0530] Moreover, the purpose of "clustering" is to perform clustering from observables to these subtypes in a data-driven manner.
[0531] Since "undersampling" and "downsampling" are randomly performed on the subjects in the healthy group and the patient group as described above, as shown in the "subject group" in Fig.17 subjects in the dotted part are respectively selected from each of the healthy group and the patient group.
[0532] In addition, the feature quantity (the related value of brain functional connectivity) that can be used to identify label M and label H is set to be the feature quantity within the range indicated by the one-dot dash line in the brain functional connectivity that characterizes the feature quantities (a total of Nch) of each subject, Fig.17 i.e., the feature quantity within the range indicated by the one-dot dash line shown in Fig.16 ("union of brain functional connectivities") in step S220.
[0533] Moreover, as a result of learning an identifier for identifying label M and label H through a learning process accompanied by feature quantity selection (here, a process based on the LASSO method), it is set to select Fig.17 the feature quantity further indicated by a black dot within the one-dot dash line shown in
[0534] Fig.18 is a conceptual diagram showing the feature quantity finally selected when generating one identifier through a learning process accompanied by feature quantity selection after undersampling and downsampling processes.
[0535] As Fig.18 shown, the feature quantity that can be used to identify label M and label H within the one-dot dash line is further divided into groups of feature quantities with strong mutual correlation as indicated by the dashed box.
[0536] In the LASSO method, sparsification is achieved by selecting one feature quantity for each group within such a dashed box.
[0537] Fig.19 is a conceptual diagram showing the situation of selecting feature quantities when generating an identifier by repeatedly performing undersampling and downsampling processes.
[0538] As Fig.19 shown, when performing downsampling, for example, Ns times, in each time, different subjects are downsampled from the healthy group and the patient group respectively.
[0539] Then, as a result of learning an identifier for identifying label M and label H through a learning process accompanied by feature quantity selection for each subsample, corresponding to each subsample, different feature quantities are selected as indicated by black dots from the groups with strong correlation within the one-dot dash line that is the above-mentioned union.
[0540] As a result, by repeatedly performing undersampling and downsampling processes, the union of the feature quantities that can be used to identify label M and label H is selected.
[0541] In this embodiment, for each subsample, the feature quantities selected through the learning process of the identifier accompanied by feature quantity selection are sorted according to the selected frequency by the LASSO method.
[0542] Then, a predetermined number of features, for example 100, are used from the highest position in the sorting. Fig.16 In step S224, clustering based on unsupervised learning is implemented through "multiple co-clustering" as described later.
[0543] In the above description, the LASSO method is used as an example of “a learning process of a classifier with feature quantity selection”, and the selection of feature quantities for clustering is performed in a ranking of the frequencies of selection.
[0544] However, as described above, in the clustering process of this embodiment, "learning process of the identifier accompanied by feature quantity selection" is not limited to such a method, for example, it can also be a method such as the random forest method, and the selection of feature quantities for clustering can also be implemented according to a specified importance.
[0545] For example, as described above, in the random forest method, in the learning process of the recognizer, the importance of the feature quantity is calculated based on Gini impurity and permutation importance, so it can also be configured to sort the feature quantities based on the importance, and use a specified number of feature quantities from the high position in the sorting, through Fig.16 The “multiple co-clustering” in step S224 is used to implement clustering based on unsupervised learning.
[0546] In addition, here, permutation importance refers to "the difference between the error of the model created with the randomized feature quantity and the error of the original model", and is used to calculate "which feature quantity contributes the most to the accuracy of the model, or does not contribute to the accuracy". For example, the permutation importance is disclosed in the following public document 15. The records of public document 15 are all cited here by reference.
[0547] Publicly known document 15: Breiman, Leo. "Random forests." Machine learning 45.1 (2001): 5-32. https: / / link.springer.com / article / 10.1023 / A:1010933404324
[0548] In addition, regarding "the process of learning an identifier as an ensemble learning", the ridge regularization method or the like can also be used. As described above, the feature quantities are sorted according to the importance corresponding to the median value obtained by summing the absolute values of the weight coefficients, and the specified number of feature quantities starting from the higher positions in the sorting are used to perform the selection of the feature quantities for clustering.
[0549] [Multi-co-clustering process]
[0550] Next, regarding Fig.16 "multi-co-clustering" in step S224 of
[0551] As a premise, "clustering" refers to a data classification method based on unsupervised learning performed by a computer. More specifically, it refers to a method of automatically classifying the provided data without an external reference. In contrast, "class classification" generally refers to a classification method based on "supervised learning". In addition, a "cluster" is defined as a subset of data having the properties of internal connection and external separation. Here, external separation refers to the property that objects in different clusters are not similar, and internal connection refers to the property that objects within the same cluster are similar to each other. And the distance between the elements of the set is defined as the criterion of "similarity". Usually, the distance is defined in a way that satisfies the so-called "axioms of distance". As the distance, the Euclidean distance, the Mahalanobis distance, the city block distance, the Minkowski distance, etc. are sometimes used.
[0552] In addition, generally, as methods for clustering through unsupervised learning, there are known methods such as "k-means method" as a non-hierarchical method, which is called "partition optimization clustering" as a method of searching for the optimal partition of the objective function that defines the goodness of the cluster, and "agglomerative hierarchical clustering", "divisive hierarchical clustering", etc. as hierarchical clustering methods.
[0553] However, such conventional clustering methods have the following characteristics: all feature quantities are used to divide the objects into clusters (groups), and there is one way of dividing the obtained clusters.
[0554] Therefore, in the case where there are multiple ways of dividing clusters according to feature quantities, there will be problems that cannot be well handled. Generally, the larger the number of feature quantities, the higher the possibility of such multiple cluster structures.
[0555] In addition, as a clustering method, not only the methods described above, but also there are algorithms that assume that multiple objects to be clustered are generated according to a certain probability distribution in each cluster and perform clustering in the direction of estimating such a "probability distribution". As such a clustering method, for example, the "clustering method based on Gaussian mixture distribution" is known, and it is known that more flexible clustering can be performed.
[0556] Next, first, in order to illustrate the case where there are multiple ways of dividing clusters according to characteristic quantities, each object in a group of objects including a plurality of objects to be clustered is characterized by a plurality of characteristic quantities.
[0557] Fig. 20 It is a conceptual diagram for explaining the case where there are multiple ways of dividing clusters according to characteristic quantities.
[0558] As Fig. 20 shown, assume that the data to be clustered (hereinafter simply referred to as "objects") are 6 characters "A", "B", "C", "D", "E", "F".
[0559] Moreover, these characters have different background patterns and different fonts (character styles).
[0560] Therefore, as characteristic quantities for characterizing these characters, "background pattern", "character style", and "number of holes contained in the character (number of regions completely surrounded by lines)" can be considered.
[0561] Thus, even when considering the same set of characters, different clusters are divided depending on which characteristic quantity is used for clustering.
[0562] In Fig. 20 , for example, in the case based on "background pattern", it is divided into 3 clusters {A, D}, {B, E}, {C, F}; in the case based on "character style", it is divided into 2 clusters {A, B, C}, {D, E, F}; in the case based on "number of holes", it is divided into 3 clusters {C, E, F}, {A, D}, {B} corresponding to 0, 1, and 2 respectively.
[0563] In Fig. 20 , the case where one characteristic quantity characterizes one cluster is taken as an example, but generally, one cluster is characterized by a plurality of characteristic quantities.
[0564] Fig.21A And Fig. 21B is a conceptual diagram for explaining the concept of clustering in the case where a plurality of objects are characterized by a plurality of characteristic quantities.
[0565] First, as Fig.21A shown, consider a "data array" in which the objects to be clustered are arranged in the row direction and the characteristic quantities characterizing these objects are arranged in the column direction.
[0566] As Fig. 21BAs shown, a method of clustering features in a manner associated with each object cluster simultaneously with clustering an object (dividing the object into multiple object clusters) is called "co-clustering". For example, this method is disclosed in the following known document 16. All the descriptions in known document 16 are incorporated herein by reference.
[0567] Known document 16: Madeira SC, Oliveira AL. Biclustering algorithms for biological data analysis: a survey (Biclustering algorithms for biological data analysis: a survey). IEEE / ACM Transactions on Computational Biology and Bioinformatics (TCBB). 2004;1(1):24±45. https: / / doi.org / 10.1109 / TCBB.2004.2
[0568] In "co-clustering", as Fig. 21B shown, by changing the rows or columns of the data array, that is, by rearranging the objects and features according to the similarity, it is divided into, for example, cluster blocks represented by (i, j) (i = 1, 2; j = 1, 2, 3).
[0569] At this time, assuming a generative model (probability model) for the objects included in each cluster, the parameters of each probability model are determined in such a way as to increase the likelihood for the observed data.
[0570] In this way, when estimating the probability model for each cluster, it is possible to discriminate (classify) which cluster the specific observed data (test data) belongs to.
[0571] Fig.22A And Fig. 22B are concept diagrams for explaining the concepts of multi-clustering and multi-co-clustering.
[0572] In Fig. 21B the "co-clustering" shown, by changing the rows and columns of the "data array", a block-structured cluster is generated. Therefore, when the features are divided into multiple feature clusters, each object is clustered into an object that is commonly arranged in the multiple feature clusters.
[0573] However, in the case of assuming that the features are divided into multiple feature clusters and the objects are also divided into object clusters for each feature cluster, if it is assumed that the arrangement of the objects in the object cluster (the arrangement of the objects included in a certain object cluster) is also different in each feature cluster, a probability model with a higher likelihood can be estimated.
[0574] In such a case, for each cluster of feature quantities, the method of dividing the object (clustering of the object) is different. Correspondingly, the cluster of feature quantities is specifically referred to as a "viewpoint".
[0575] As Fig.22A shown, performing different object clusterings according to each viewpoint of feature quantities as described above is called "multi-clustering".
[0576] And, as Fig. 22B shown, the case where the likelihood of the probability model for the observed data can be further estimated by changing the columns of the feature quantities and the rows of the objects and performing clustering under each viewpoint is called "multi-co-clustering".
[0577] Here, in the case of only one view and the case of multiple views, including the case where there is only one type of feature quantity cluster in at least one view, it is called "multi-co-clustering". "Co-clustering" and "multi-clustering" are subordinate concepts of "multi-co-clustering".
[0578] In addition, in the present embodiment, only the term "clustering" refers to generating a group of clusters in one view. For example, the case of dividing the feature quantities into viewpoints and performing object clustering as Fig.22A shown is called "multi-clustering", and the case of performing co-clustering while dividing into viewpoints as Fig. 22B shown is called "multi-co-clustering", and thus they are distinguished.
[0579] Fig.23 is a conceptual diagram showing the case of a probability model in which different types of probability distributions are assumed in one view in "multi-co-clustering".
[0580] In Fig.23 it shows the case of a probability model in which different types of probability distributions are followed in white blocks and shaded blocks.
[0581] For example, it shows cases such as: the white blocks are the probability distribution of continuous probability variables, while the shaded part is assumed to be the probability distribution of discrete probability variables.
[0582] As will be described later, in the "learning method of multi-co-clustering" of the present embodiment, clustering processing can be performed on a distribution family containing different distributions in this way.
[0583] Fig.24 is a flowchart for explaining the outline of the learning method of multi-co-clustering.
[0584] When starting the processing of the learning method for multi-co-clustering (S300), the clustering classifier generation unit 3010 randomly divides the feature quantities into subgroups for the data array, generates views of the feature quantities, and feature quantity clusters within the views (S302: corresponding to the generation of Y (initialization of Y) described later).
[0585] Next, the clustering classifier generation unit 3010 generates and optimizes the division of the object clusters corresponding to the views of the feature quantities and the feature quantity clusters generated in step S302 (step S304: corresponding to the generation of Z described later).
[0586] Moreover, the clustering classifier generation unit 3010 optimizes the division of the feature quantities for the obtained object clusters (S306: corresponding to the generation process of Y described later, using the generated Z to optimize Y).
[0587] Next, the clustering classifier generation unit 3010 determines whether the objective function satisfies a specified condition and has converged (S308). In addition, this objective function is a function L(q(φ)) as described later. The function L(q(φ)) also has the property of monotonically increasing as Y and Z described later are updated, and when it is determined that the increase has become very small, it is determined to have converged. If it has not converged (S308: "No"), the clustering classifier generation unit 3010 returns the process to step S304, and if it has converged (S308: "Yes"), the process proceeds to the next step.
[0588] Then, the clustering classifier generation unit 3010 saves the magnitude of the objective function to the storage device 2080 (S310).
[0589] Next, the clustering classifier generation unit 3010 determines whether the processing of steps S302 to S310 has been performed a specified number of times. If it has not been performed a specified number of times (S312: "No"), the clustering classifier generation unit 3010 returns the process to step S302, and if it has been performed a specified number of times (S312: "Yes"), the process proceeds to the next step.
[0590] The clustering classifier generation unit 3010 takes the feature quantity division and the cluster partitioning method that maximize the objective function as the final result (S314), and ends the processing of learning for multi-co-clustering, and generates a clustering classifier.
[0591] Fig.48 It is a conceptual diagram showing the structure of the clustering classifier.
[0592] In Fig.48 Three views are subjected to division, and Figure 1The object is segmented corresponding to the feature quantity group 1, and the object is clustered into clusters 1 to 3. The view 2 is segmented corresponding to the feature quantity group 2, and the object is clustered into clusters 1 to 4. The view 3 is segmented corresponding to the feature quantity group 3, and the object is clustered into clusters 1 to 2.
[0593] Therefore, the clustering classifier generation unit 3010 stores information on the feature quantities corresponding to each view and information for determining the probability density function for each view (for example, when the probability density function is a normal distribution, it is the center coordinate μ and variance σ of the distribution 2 ) in the disease recognizer data 3112 of the storage device 2080.
[0594] As Fig.48 shown, when "new data" with feature quantity groups 1 to 3 is input to the clustering classifier, the discriminant value calculation unit 3012 calculates the posterior probability of belonging to each cluster for the feature quantity group 1 corresponding to the view Figure 1 based on the probability density function within the view Figure 1 , and this "new data" is "new data (data of a new examinee)" not included in the learning data. In Fig.48 , the posterior probability of belonging to cluster 2 of the view Figure 1 is the highest, and the classification result that the new data belongs to cluster 2 of the view Figure 1 is output. Similarly, the following classification results are output from the clustering classifier: In view 2, the new data belongs to cluster 3, and in view 3, the new data belongs to cluster 1.
[0595] (Details of the processing of multi - co - clustering)
[0596] Next, the learning method of multi - co - clustering described in Fig.24 will be described in more detail.
[0597] In addition, the details of the processing of multi - co - clustering are disclosed in the following well - known document 17, so an outline thereof will be described below. The entire description of well - known document 17 is incorporated herein by reference.
[0598] Prior Art Document 17: Tomoki Tokuda, Junichiro Yoshimoto, Yu Shimizu, Go Okada, Masahiro Takamura, Yasumasa Okamoto, Shigeto Yamawaki, Kenji Doya, "Multiple co-clustering based on nonparametric mixture models with heterogeneous marginal distributions", PLOS ONE https: / / doi.org / 10.1371 / journal.pone.0186566 October 19, 2017
[0599] Fig.25 is a diagram showing Fig.24 the graphical representation of the Bayesian estimation in the learning method of multiple co-clustering.
[0600] The multiple co-clustering model is summarized as Fig.25 a graphical model that clarifies the links of the causal relationships between the relevant parameters and the data array.
[0601] (Multiple co-clustering model)
[0602] The feature quantity (brain functional connectivity value) and the subjects (here, the subjects in the patient group) are shown as Fig.21A the data array shown.
[0603] Moreover, it is assumed that the data array X is composed of a distribution family including M pre-known distributions.
[0604] As the probability distributions belonging to the distribution family, it is assumed that Gaussian distribution, Poisson distribution, and categorical distribution / multinomial distribution, etc. can be included.
[0605] The clustering classifier generation unit 3010 divides X (m) in the manner of each data size being n×d (m) as follows.
[0606] X = {X (1) , …, X (m) , …, X (M)}
[0607] Here, m is an index indicating the distribution family (m = 1, ···, M). And, the number of views (viewpoints) is set to V (common to all distribution families), and the number of feature quantity clusters of view v and distribution family m is set to Gν (m) The number of object clusters in view v is denoted by K v (common to all distribution families).
[0608] Also, to simplify the presentation and to show the number of features and the number of clusters, empty clusters are allowed, and thus it is expressed as G (m) = max v G v (m) and K = max v K v .
[0609] In this expression, for the independent and identically distributed (i.i.d.) d (m) -dimensional random vectors X1 (m) , ···, X n (m) Consider the (third-order) feature partition tensor Y of d (m) ×V×G (m) . When the feature j of distribution family m belongs to the feature cluster g of view v, it is set that Y (m) j,v,g (m) = 1 (0 otherwise).
[0610] Combining this feature partition tensor with different distribution families gives Y = {Y (m)} m .
[0611] Similarly, consider the object partition (third-order) tensor Z of n×V×K. When the object i belongs to the object cluster k of view v, it is set that Z i、v、k = 1.
[0612] The feature j belongs to one in the view (Σ v,g Y j,v,g (m) = 1), and the object i belongs to each view (i.e., Σ k Z i,v,k (m) = 1). Also, Z is common to all distribution families, which means that the estimated probability model uses the information of all distribution families to estimate the subject clustering solution.
[0613] First, as Fig.25 shown, for the pre-generated model of Y, consider the hierarchical structure of views and feature clusters. First, generate the views, and then generate the feature clusters. Thus, the features are partitioned by the members of the pairs of views and feature clusters, and the assignment of the feature partition is jointly determined by the views and feature clusters.
[0614] On the other hand, as Fig.25 As shown, the object is segmented into object clusters for each view. Therefore, for Z, only one configuration of the object cluster is considered. Assume that these generative models are all based on the "Stick Breaking Process" (SBP) as described below.
[0615] (Generative model of feature quantity cluster Y)
[0616] Assume Y j.. (m) represents the view / feature quantity cluster member vector of feature quantity j of distribution family m generated by the hierarchical stick breaking process, then the following equation holds.
[0617] [Equation 3]
[0618] w v ~Beta(·1,α1),v = 1,2,...
[0619]
[0620] Here, τ (m) represents a 1×GV vector (τ 1,1 (m) , ···, τ G,V (m) ) T .
[0621] Mul(·|π) is a multinomial distribution with a sample size of one and probability parameter π.
[0622] β(·|a, b) is a beta distribution with presample size (a, b).
[0623] Y j.. (m) represents a 1×GV vector (Y j,1,1 (m) , ···, Y j,V,G (m) ) T .
[0624] Here, according to the specified conditions, the number of views of very large V and the number of feature quantity clusters of G are discarded.
[0625] Known literature 18: Blei DM, Jordan MI, et al. Variational inference for Dirichlet process mixtures. Bayesian analysis. 2006; 1(1): 121 - 143,
[0626] https: / / doi.org / 10.1214 / 06-BA104
[0627] When Y j,v,g (m) = 1, the feature quantity j belongs to the feature quantity cluster g of the view v. By default, the concentration parameters α1 and α2 as hyperparameters are set to 1.
[0628] [Generative Model of Object Cluster Z]
[0629] Generate the vector expressed as Z i,v. by the following formula, and this vector is the subject cluster member vector of the target object i of the view v.
[0630] [Equation 4]
[0631] u k,v ~Beta(·1,β), v = 1, 2,..., k = 1, 2,...
[0632]
[0633] Z i,v ~Mul(·|η v )
[0634] Here, Z i,v. is a 1×K (K takes a sufficiently large value) vector given by Z i,v =(Z i,v,1 , ···, Z i,v,K ). T The concentration parameter β is set to 1.
[0635] (Likelihood and Prior Distribution)
[0636] Assume that each instance X i,j (m) independently follows a specific distribution under the condition of conditioning on Y and Z. Denote the parameter of the distribution family m in the cluster block of the view v, the feature quantity cluster g, and the target object cluster k as θ v,g,k (m) .
[0637] And, expressed as Θ = {θ v,g,k (m)} v,g,k,m , the logarithm of the likelihood of X follows the following formula:
[0638] [Equation 5]
[0639]
[0640] Here, I(x) is the indicator function, which returns 1 when x is true and returns 0 otherwise. The likelihood does not depend on w = {wv} v , w′ = {w′ g,v (m)} g,v and u = {u k,v} k,v are directly related.
[0641] Give the joint prior distribution φ = {Y, Z, w, w′, u, Θ} (i.e., the member variables of the class and the model parameters) of the unknown variables as follows.
[0642] [Equation 6]
[0643] p(w)p(w′)p(Y|w, w′)p(u)p(Z|u)p(Θ).
[0644] (Variational estimation)
[0645] Use the variational Bayesian EM algorithm in the MAP (maximum a posteriori) estimation of Y and Z.
[0646] Regarding such a variational Bayesian EM algorithm, it is disclosed in the following well-known document 19. Incorporate all the descriptions of well-known document 19 by reference herein.
[0647] Well-known document 19: Guan Y, Dy JG, Niu D, Ghahramani Z. Variational inference for nonparametric multiple clustering. In: MultiClust Workshop, KDD - 2010; 2010.
[0648] Approximate the log marginal likelihood p(X) using Jensen's inequality as follows.
[0649] [Equation 7]
[0650]
[0651] In addition, regarding Jensen's inequality, it is disclosed in the following well-known document 20. Incorporate all the descriptions of well-known document 20 by reference herein.
[0652] Well-known document 20: Jensen V. Sur les fonctions convexes et les inegalites entre les valeurs moyennes (On convex functions and mean inequalities). Acta Mathematica. 1906; 30(1): 175 - 193.
[0653] https: / / doi.org / 10.1007 / BF02418571
[0654] Here, \(q(\varphi)\) is an arbitrary distribution of the parameter \(\varphi\). It is shown that the difference between the left - hand side and the right - hand side is given by the Kullback - Leibler divergence between \(q(\varphi)\) and \(p(\varphi|X)\), i.e., \(KL(q(\varphi), p(\varphi|X))\). Thus, the method of choosing \(q(\varphi)\) is to minimize \(KL(q(\varphi), p(\varphi|X))\), which is usually difficult to evaluate.
[0655] Here, for different parameter (mean - field approximation) choices, \(q(\varphi)\) which is factorized is chosen.
[0656] [Equation 8]
[0657] \(q(\varphi)=q\) w (w)q w′ (w′)q Y (Y)q u (u)q Z (Z)q Θ (Θ)
[0658] Here, each \(q(\cdot)\) is further factorized with respect to subsets of the parameters \(w\) v , \(w′\) g,v (m) , \(Y\) j.. (m) , \(u\) k,v , \(Z\) i,v. and \(\theta\) v,g,k (m) is factorized.
[0659] Generally, the distribution \(q\) l=1 L q l (\varphi l ) that minimizes \(KL(\Pi\) i (\varphi i ), \(p(\varphi|x))\) is given by the following formula.
[0660] [Equation 9]
[0661]
[0662] Here, denotes the average with respect to \(\Pi\) l≠i q i (\varphi i ).
[0663] Regarding this property, it is disclosed in the well - known document 21 below. The entire description of the well - known document 21 is incorporated herein by reference.
[0664] Known Document 21: Murphy K. Machine Learning: A Probabilistic Perspective. Cambridge, Massachusetts: MIT Press; 2012.
[0665] When this property is applied to the model currently under study, the following can be shown.
[0666] [Number 10]
[0667]
[0668] Here, consider a function represented in the following form.
[0669] [Number 11]
[0670] q Θ (Θ)
[0671] Hyperparameters other than the above formula are represented by the following formula.
[0672] [Number 12]
[0673]
[0674] [Number 13]
[0675]
[0676] Here, E q(θ) represents the average of q(θ) corresponding to θ v,g,k (m) The digamma function ψ(·) is defined as the first derivative of the logarithm of the gamma function.
[0677] The digamma function ψ(·) represents the first derivative of the logarithm of the gamma function.
[0678] τ j,g,v (m) is normalized over multiple pairs (g, v) for each pair (j, m). On the other hand, η i,v,k is normalized with respect to k in each pair (i, v).
[0679] The observation model and the prior distribution of the parameter Θ will be described later.
[0680] (Observation model)
[0681] For the observation model, Gaussian distribution, Poisson distribution, and categorical / multinomial distribution are considered. For each cluster, assuming that the features within the cluster are independent, the univariate distributions of these families are fitted. The parameters of these distribution families are assumed to be conjugate prior distributions.
[0682] (Optimization Algorithm)
[0683] In the update equation of the hyperparameters, the variational Bayesian EM algorithm is used to perform calculations as follows.
[0684] First, randomly initialize {τ (m)} m and {η v}, v and update the hyperparameters until the lower bound L(q(φ)) of Equation (1) converges. This generates a locally optimal distribution q(φ) from the perspective of L(q(φ)). Repeat this process multiple times, and select the optimal solution with the maximum lower bound as the approximate posterior distribution q*(φ).
[0685] The MAP estimates of Y and Z are evaluated as argmax Y q * Y (Y) and argmax Z q * Z (Z), respectively.
[0686] The lower bound L(q(φ)) is given by the following equation.
[0687] [Equation 14]
[0688]
[0689] The two terms on the right-hand side can be derived in a closed form. As q(φ) is optimized, this value shows a monotonically increasing situation. That is, as described above, the function L(q(φ)) has the property of monotonically increasing as Y and Z are updated. When it is determined that the way of its increase becomes very small (although not particularly limited, for example, when conditions such as the increment being below a specified value are satisfied), it is determined that convergence has occurred.
[0690] First, determine the distribution family of each feature and generate a data array corresponding to the distribution family. Then, for the set of data arrays, further generate the MAP estimates of Y and Z, and use the estimated values of Y and Z to analyze the object / feature quantity clusters in each view.
[0691] (Model Performance)
[0692] The multi - co - clustering model has sufficient flexibility to represent various clustering models because it derives the number of views and the number of feature / target clusters through a data - driven method. For example, when the number of views is 1, the model is consistent with the co - clustering model. When the number of feature clusters is 1 for all views, it is consistent with the multi - clustering model. And when the number of views is 1 and the number of feature clusters is the same as the number of features, it is consistent with the conventional mixture model with independent features. Moreover, the model can detect uninformative features that do not distinguish target clusters. In this case, the model generates a view with the number of target clusters being one. The advantage of the model is that it can automatically detect such underlying data structures.
[0693] The following situations can be achieved by the "multi - co - clustering method" as described above.
[0694] 1) It is possible to identify, in a data - driven manner, the partitioning methods of multiple clusters underlying the data (including not only the partitioning methods of objects but also those of feature quantities) and the corresponding groups of feature quantities.
[0695] 2) It is possible to identify clusters that cannot be discovered by other methods using this approach.
[0696] 3) Moreover, it is possible to assign meanings to the partitioning methods of each cluster through feature quantities, and it is easy to interpret each cluster.
[0697] [Evaluation of the results of clustering the dataset]
[0698] Next, the multi - facility large - scale fMRI data collected from a large number of subjects publicly available as SRPBS as described above is segmented into two parts, and each part is used in the multi - co - clustering method as described above to verify the generalization performance of clustering.
[0699] Fig.26A and Fig.26B are diagrams showing the dataset 1 and dataset 2 segmented into two parts like this.
[0700] Fig.26A The dataset 1 shown is composed of data from 545 healthy subjects and 138 depressed patients obtained at facilities 1 - 4, Fig.26B The dataset 2 shown is composed of data from 263 healthy subjects and 181 depressed patients obtained at facilities 5 - 8. It basically corresponds to Fig.10 and Fig.11 the datasets shown.
[0701] Fig. 27 is a conceptual diagram explaining the concept of performing clustering on each dataset.
[0702] As Fig. 27 As shown, for dataset 1, independently and respectively, according to Fig.24 the process shown, clustering is performed by the multi - co - clustering method.
[0703] Here, the problem of concern is to what extent the clusters obtained by such data - driven clustering methods independently performed on dataset 1 and dataset 2 are similar to each other (what is the degree of consistency).
[0704] If the clusters in dataset 1 and dataset 2 can be classified (grouped) into clusters (subject groups) with the same or similar characteristics, then this data - driven clustering is performed in a highly generalizable state without depending on facilities, measuring devices, etc. Therefore, the problem is how to quantitatively evaluate being “classified into clusters with the same or similar characteristics”.
[0705] Fig.28 is a conceptual diagram showing an example of multi - co - clustering for subject data.
[0706] As Fig.28 shown in (a) of, as the data array that becomes the input, it is assumed that subjects are arranged in the row direction and feature quantities are arranged in the column direction.
[0707] When performing multi - co - clustering on this input data array, for example, as Fig.28 shown in (b) of, the feature quantities are divided into 2 views, and in each view, the subjects are clustered.
[0708] Fig.29 is a diagram showing the results obtained by actually performing multi - co - clustering processing on dataset 1 and dataset 2.
[0709] In Fig.29 , as the feature quantities for clustering, 99 are respectively selected for both dataset 1 and dataset 2.
[0710] On this basis, for dataset 1, multi - co - clustering processing was performed on 138 depressive patients, and for dataset 2, multi - co - clustering processing was performed on 181 depressive patients.
[0711] For dataset 1, the feature quantities are divided into Figure 1 view and view 2. For view Figure 1 , it is further co - clustered into 2 feature quantity clusters, and the subjects are divided into 5 subject clusters. For view 2, the subjects are also divided into 5 clusters.
[0712] For dataset 2, the feature quantities are also divided into Figure 1 view and view 2. For view Figure 1, further co-clustered into 2 clusters of feature quantities, the subjects were segmented into 4 subject clusters, and for View 2, the subjects were segmented into 5 clusters.
[0713] Fig.30 is a table showing the number of brain functional connections (FC) assigned to each view in Dataset 1 and Dataset 2.
[0714] In Dataset 1, for view Figure 1 92 FCs were assigned as feature quantities, and 7 FCs were assigned as feature quantities for View 2.
[0715] In Dataset 2, for view Figure 1 93 FCs were assigned as feature quantities, and 6 FCs were assigned as feature quantities for View 2.
[0716] In addition, in this table, for Dataset 1 and Dataset 2, the number of agreements in the assigned brain functional connections is recorded on the diagonal of the table. It can be seen that the brain functional connections assigned to view Figure 1 and View 2 in Dataset 1 and Dataset 2 are roughly the same.
[0717] (Verification method for the generality (similarity between datasets) of clustering (hierarchical))
[0718] Next, quantitatively evaluate to what extent the clusters obtained by data-driven clustering independently performed on Dataset 1 and Dataset 2 are similar to each other (what is the degree of agreement).
[0719] Fig.31 is a conceptual diagram for explaining the evaluation method of the similarity (generalization performance of hierarchical) of such clustering.
[0720] First, as shown in Fig.31 (A) of, when Dataset 1 and Dataset 2 are independently segmented into clusters by the above-mentioned multi-co-clustering method, the subjects are independent of each other in the clusters of each dataset, so it is difficult to compare the similarity of clustering.
[0721] Here, the result of classifying the subjects in Dataset 1 using classifier 1 generated from Dataset 1 is set as clustering 1. On the other hand, the result of classifying the subjects in Dataset 2 using classifier 2 generated from Dataset 2 is set as clustering 2.
[0722] In contrast, as shown in Fig.31 (B) of, the result of classifying the subjects in Dataset 2 using classifier 1 generated from Dataset 1 is set as clustering 1'. On the other hand, the result of classifying the subjects in Dataset 1 using classifier 2 generated from Dataset 2 is set as clustering 2'.
[0723] In this case, common subjects are classified between cluster 1 and cluster 1', and between cluster 2 and cluster 2', respectively, so that the similarity of each can be evaluated.
[0724] (Evaluation criteria for measuring the similarity (reproducibility) between clusters)
[0725] Here, the clustering process is performed in a data-driven manner. Fig.29 The value of the cluster index itself (the order of the index values) in [ ] has no meaning. Therefore, when different clusterings are performed on the same data set, how to evaluate its similarity becomes a problem.
[0726] For example, when there are two clustering results π and ρ for the same data set X, as a criterion for evaluating the similarity (external validity criterion) of these two clustering results, the Rand index is known.
[0727] Regarding all pairs {x1, x2} ∈ X (M = N(N - 1) / 2) of data in the data set, as the types of pairs, there are the following types, and the number of pairs belonging to each type is defined as follows.
[0728] [Equation 15]
[0729] a 11 : The number of pairs where both π and ρ are in the same cluster
[0730] a 01 : The number of pairs where π is in a different cluster but ρ is in the same cluster
[0731] a 10 : The number of pairs where ρ is in a different cluster but π is in the same cluster
[0732] a 00 : The number of pairs where both π and ρ are in different clusters
[0733] At this time, as the correct rate of determining whether the clusters classified by the two clusterings are the same cluster, the Rand index is defined by the following formula.
[0734] [Equation 16]
[0735]
[0736] However, for example, it is known that in the case where there is a deviation in the number of elements in each cluster of the data set, etc., there is a situation where "even when clustering is performed randomly, the Rand index becomes a high value". Therefore, more strictly speaking, the following adjusted Rand index (ARI: Adjusted Rand Index) is used.
[0737] Fig.32A and Fig.32B are conceptual diagrams for explaining ARI.
[0738] In addition, regarding ARI, for example, it is disclosed in the following well-known document 22. All the descriptions in well-known document 22 are incorporated herein by reference.
[0739] Well-known document 22: Jorge M. Santos and Mark Embrechts, “On the Use of the Adjusted Rand Index as a Metric for Evaluating Supervised Classification (Regarding the Use of the Adjusted Rand Index as an Index for Evaluating Supervised Classification)”, ICANN 2009, Part II, LNCS 5769, pp. 175 - 184, 2009.
[0740] As described above, in the case of applying two clustering results to the same data set respectively, Fig.32A as shown, there are cases where the same cluster is classified twice and cases where different clusters are classified twice. On the other hand, Fig.32B as shown, there are also cases where the same cluster is classified once but a different cluster is classified the other time.
[0741] In the case where the two clusterings are independent of each other, ARI is calculated by calculating the expected value in the case where the data pairs are classified into the same cluster or different clusters in both clusterings, and subtracting this expected value from the numerator and denominator of the Rand index respectively. Thus, in ARI, when there is no correlation between the clusterings, its value is adjusted to 0.
[0742] [Equation 17]
[0743]
[0744] Here, A represents the number of subject pairs for the two clusterings [(classified into the same cluster twice) + (classified into different clusters twice)], max(A) represents the number of all pairs, and E represents the number of subject pairs whose assignment results are consistent despite the two clusterings being independent.
[0745] Fig.33A and Fig.33B are respectively graphs showing the evaluation results of the similarity between clustering 1 and clustering 1' and between clustering 2 and clustering 2'.
[0746] Fig.33A is a table obtained by calculating ARI for each view of data set 1 and data set 2.
[0747] Regarding datasets 1 and 2, for view Figure 1 , ARI = 0.47, and for view 2, ARI = 0.51, indicating significant similarity.
[0748] Fig.33B Shows the results of the permutation test corresponding to Fig.33A . Compared with the case where element exchanges were performed for view Figure 1 and view 2 (represented by a histogram), it can be seen that in view Figure 1 and view 2, the ARI values (represented by solid lines) become statistically significantly higher values.
[0749] In addition, "permutation test" refers to the result obtained by calculating the ARI value when the cluster attribute labels of the subjects are randomly exchanged among the subjects. In the figure, the distribution when such exchanges are performed a specified number of times is shown in the form of a histogram. If the similarity between the clusters is statistically significant, the ARI value between the clusters being compared becomes significantly higher than in the case of randomly exchanging elements.
[0750] Based on the above content, regarding the clustering (stratification) between datasets, it can be judged that significant similarity has been confirmed, that is, a generalized clustering has been achieved.
[0751] As described above, the multi-co-clustering for dataset 1 and the multi-co-clustering for dataset 2 are both achieved through data-driven means, and thus can be said to be the basis of "stratification of patients" characterized by brain functional connectivity.
[0752] Fig.34 Is a table showing the distribution of the number of subjects assigned to each cluster for each view of cluster 1 and cluster 1'. Figure 1
[0753] By appropriately rearranging the subject clustering indices, most subjects can be distributed near the diagonal of the table, and the visual similarity between the two clusterings can also be confirmed.
[0754] [Harmonization processing]
[0755] Next, the content of the processing called harmonization processing in the following literature and disclosed in Figure 6 will be described.
[0756] [Harmonization of the multi-facility examinee method]
[0757] Next, the method for generating the "disease identifier" described above, used in the "clustering process" for stratification, to harmonize measurement data for evaluating measurement bias independently of sampling bias will be described.
[0758] Fig.35 This is a conceptual diagram of an evaluation method for the between-site differences of mobile subjects (hereinafter referred to as "multi-facility subjects: traveling subjects") who move between locations and undergo measurements in the rs-fcMRI method of this embodiment.
[0759] As described below, in this embodiment, a harmonization method that can use the dataset of multi-facility subjects to eliminate only measurement bias is described.
[0760] Refer to Fig.35 , in order to evaluate measurement bias across measurement sites MS.1 to Ms.Ns, a dataset of multi-facility subjects TS1 (number of people: Nts people) is obtained.
[0761] The resting-state brain activities of Nts healthy participants are set to be photographed at each of the Ns sites, and the Ns sites include all the sites where patient data is photographed.
[0762] The obtained dataset of multi-facility subjects is saved as mobile subject data in the storage device 210 of the data center 200.
[0763] Moreover, as described later, processing for the "harmonization method of brain activity biomarkers" is executed in the calculation processing system 300.
[0764] The dataset of multi-facility subjects only includes healthy groups. In addition, it is assumed that the participants are the same at all sites. Therefore, for multi-facility subjects, the between-site differences only include "measurement bias".
[0765] In the harmonization method of this embodiment described below, as the "harmonization method of brain activity biomarkers", the measurement data at each measurement site is processed as follows: The influence of "measurement bias" is removed for correction.
[0766] That is, below, the "Generalized Linear Mixed Model (GLMM)" in the "statistical modeling" technique is used to evaluate "measurement bias" and "sampling bias".
[0767] Here, the Generalized Linear Model (GLM) is usually a model that embeds "explanatory variables" for explaining the probability distribution of the "response variable". In GLM, there are mainly three components: "probability distribution", "link function", and "linear predictor". By specifying the combination method of these components, various types of data can be represented.
[0768] Moreover, GLMM (Generalized Linear Mixed Model) is a statistical model that can incorporate "individual differences that cannot be measured or have not been measured by humans," etc., which cannot be explained by GLM. Regarding GLMM, for example, when the object consists of several subsets (e.g., subsets with different measurement locations), the location differences can also be incorporated into the model. In other words, it is a model (formed by mixing) that uses multiple probability distributions as components.
[0769] For example, regarding GLMM, it is disclosed in the following well-known document 23. All the descriptions in well-known document 23 are incorporated herein by reference.
[0770] Well-known document 23: Written by Takuya Kubo, "Introduction to Statistical Modeling for Data Analysis," Iwanami Shoten, 1st edition in 2012, 14th edition in 2017
[0771] Among them, in the statistical model of the present embodiment described below, generally, for terms such as "bias" and "factor" referred to as "effects," "bias" is used as "measurement bias" and "sampling bias," and "factor" is used as other factors (subject factors, disease factors).
[0772] Moreover, the following analysis is different from the process of simple GLMM, and the factors are analyzed without distinguishing between "fixed effects" and "random effects." This is because, usually when using GLMM, only the variance is estimated for random effects, and the magnitudes of the effects of each factor are unknown. Therefore, below, in order to evaluate the magnitudes of the effects of each factor, for each factor, the variables are transformed as follows for estimation to become fixed effects with an average of 0.
[0773] i) Define the measurement bias at each location as the deviation from the average of the correlation values of each functional connectivity at all locations.
[0774] ii) Assume that the sampling biases of healthy individuals and patients with mental illnesses are different. Therefore, for the healthy group and the patient groups with various diseases, the sampling biases at each location are calculated independently.
[0775] iii) The disease factor is defined as the deviation from the value of the healthy group.
[0776] That is, below, for the dataset including patients and the dataset of multi-facility examinees, the generalized linear mixed effects model is applied as follows.
[0777] The multi-facility examinees are Nts people. Let the number of locations among the Ns measurement locations where healthy individuals were measured be Nsh, and the number of locations where patients with a certain disease (represented by the suffix "dis" here) were measured be Nsd.
[0778] The participant factor (p), measurement bias (m), sampling bias (Shc, Sdis), and mental illness factor (d) are evaluated by fitting a regression model to the dataset of the measurement results of the patients and the relevant values of the functional connectivity of all participants in the multi-facility examinee dataset.
[0779] Below, vectors are represented by lowercase letters (e.g., m), and it is assumed that all vectors are column vectors.
[0780] Vector elements such as m k are represented using suffixes in this way.
[0781] The regression model of the functional connectivity vector (assumed to be a column vector) composed of n correlation values between brain regions is expressed by the following formula.
[0782] [Equation 18]
[0783] Connectivity = x m T m + x Shc T s hc + x Sdis T s dis + x d T d + x p T p + const + e
[0784] d1(HC) = 0
[0785] To represent the characteristics of the participants, a 1-of-K (dummy coding) binary code system is used. For the target vector (e.g., xm) corresponding to the measurement bias m belonging to location k, all elements except the element k that is equal to 1 are equal to zero.
[0786] If the participant does not belong to any category (healthy, patient, multi-facility examinee), then the target vector is a vector with all elements equal to 0.
[0787] The superscript T represents the transposition of a matrix or vector, and x T represents a row vector.
[0788] Here, m represents the measurement bias (a column vector of Ns×1), shc represents the sampling bias of the healthy group (a column vector of Nsh×1), sdis represents the sampling bias of the patients (a column vector of Nsd×1), d represents the disease factor (a column vector of 2×1 with elements of healthy and disease), p represents the participant factor (a column vector of Nts×1), const represents the average of the functional connectivity involving all participants from all measurement sites (including healthy individuals, patients, and multi-facility examinees), and e~N(0,γ -1 ) represents noise.
[0789] In addition, here, for the sake of simplicity of explanation, it is assumed that there is only one type of disease. The case where there are multiple types of diseases will be described later.
[0790] For the correlation values of each functional connectivity, since the design matrix of the regression model is rank-deficient, the least-squares regression based on L2 normalization is used to evaluate each parameter. In addition, other evaluation methods such as the Bayesian estimation method can also be used in addition to the least-squares regression method based on L2 normalization.
[0791] After the regression calculation as described above, the b-th connectivity of the examinee a can be described as follows:
[0792] [Equation 19]
[0793]
[0794] Fig.36 is a conceptual diagram for explaining the expression of the b-th functional connectivity of the examinee a.
[0795] In Fig.36 , the meanings of the target vectors of the first and second terms and the measurement bias vector and the sampling bias vector of the healthy individuals are shown.
[0796] The same applies to the terms after the third term.
[0797] (Flow of the harmonization process)
[0798] Fig.37 is a flowchart for explaining the process of calculating the measurement bias for harmonization.
[0799] First, in the storage device 210 of the data center 200, the fMRI measurement data of the examinees (healthy individuals and patients), the attribute data of the examinees, and the measurement parameters ( Fig.37 of S402) are collected from each measurement site.
[0800] Next, although not particularly limited, for example, the brain activity of the multi-facility subject TS1 is measured by making rounds of each measurement site at a prescribed cycle (for example, at an annual cycle), and in the storage device 210 of the data center 200, fMRI measurement data of the multi-facility subject, subject attribute data, and measurement parameters ( Fig.37 of S404).
[0801] The harmonization calculation unit 3020 evaluates the measurement bias at each measurement site for functional connectivity by using the GLMM (general linear mixed model) as described above ( Fig.37 of S406).
[0802] The harmonization calculation unit 3020 saves the measurement bias at each measurement site calculated in this way as measurement bias data 3108 to the storage device 2080 ( Fig.37 of S408).
[0803] (Harmonization in discriminator generation processing)
[0804] Briefly describe the harmonization for the brain functional connectivity values in the process of the discrimination processing unit 3000 generating a disease identifier for the disease or health label of the subject.
[0805] Such a disease identifier provides auxiliary information (assistance information) for diagnosing the subject.
[0806] The correlation value correction processing unit 3004 reads out the measurement bias data 3108 at each measurement site saved in the storage device 2080, and performs a harmonization process on the non-diagonal components of the correlation array of each subject that is the training object for the machine learning for disease identifier generation as shown in the following formula.
[0807] [Equation 20]
[0808]
[0809] Here, Connectivity represents the functional connectivity vector before harmonization, and Csub represents the functional connectivity vector after harmonization. In addition, m (hat) (hereinafter, the expression with a ^ on the top of the letter x will be expressed as "x (hat)") represents the measurement bias at the measurement site evaluated by the least squares regression based on L2 normalization as described above. Thus, the functional connectivity Connectivity is subtracted by the measurement bias corresponding to the measurement site where the functional connectivity Connectivity is measured to undergo the harmonization process.
[0810] The data after the correction process is saved as corrected correlation value data 3110 to the storage device 2080.
[0811] In addition, as described in the following well-known document 24, it is appropriate to assume that the disease factor is not related to the connectivity of the entire brain but to a specific subset of the connectivity. The entire description of the well-known document 24 is incorporated herein by reference.
[0812] Well-known document 24: Yahata N, et al. A small number of abnormal brain connections predicts adult autism spectrum disorder. Nat Commun 7, 11254 (2016).
[0813] Therefore, next, for the subject's disease label including the disease / health label of the subject and the corrected functional connectivity, the disease identifier generation unit 3008 generates a disease identifier through learning processing with feature selection as described above.
[0814] In addition, as a method for performing feature selection and modeling to suppress overlearning, it is not limited to the regularized logistic regression using LASSO. For example, other methods such as the sparse logistic regression and other sparse Bayesian estimation methods disclosed in the following well-known document 25 can also be used. The entire description of the following well-known document 25 is incorporated herein by reference.
[0815] Well-known document 25: Okito Yamashita, Masaaki Sato, Taku Yoshioka, Frank Tong, and Yukiyasu Kamitani. “Sparse Estimation automatically selects voxels relevant for the decoding of fMRI activity patterns.” NeuroImage, Vol. 42, No. 4, pp. 1414 - 1429, 2008.
[0816] Based on the above, for the disease identifier (which functions as a discriminator for diagnosis using brain activity biomarkers), the disease identifier generation unit 3008 saves the information for determining the discriminator as disease identifier data 3112 in the storage device 2080.
[0817] The above processing is not particularly limited as described above, but can be set to be performed at regular intervals (e.g., once a year).
[0818] Moreover, when a measurement of the functional connectivity of a subject is newly performed at a certain measurement location, it can be assumed that the measurement bias at the measurement location is fixed within a specified period. Therefore, the discrimination value calculation unit 3012 subtracts the "measurement bias" corresponding to the measurement location where the input data of the subject is measured, which has been calculated in the above process, from the value of the element of the correlation matrix of the input data of the subject, thereby performing the harmonization process. Moreover, through the above process, the "disease identifier" that has been created outputs a discrimination label for the subject as the discrimination result.
[0819] The discrimination result can be a value indicating one of "disease" and "health", or can also be a value of the probability indicating at least one of "disease" and "health".
[0820] In addition, when performing the discrimination process implemented by the discrimination value calculation unit 3012, as the "input data" input, it can be the MRI measurement data itself indicating the brain function activity of the subject, or can also be the data of the correlation value itself, which is the non-diagonal element of the correlation matrix after calculating the value of the correlation matrix at each measurement location based on the MRI measurement data indicating the brain function activity of the subject.
[0821] (Harmonization calculation process in the case of adding a new measurement location)
[0822] Fig.38 is a conceptual diagram for explaining the calculation process of the measurement bias for harmonization processing in the case of adding a new measurement location after performing the calculation of the measurement bias for harmonization processing through the process described in Fig.37 The concept map of the calculation process of the measurement bias for harmonization processing in the case of adding a new measurement location after performing the calculation of the measurement bias for harmonization processing through the process described in
[0823] Refer to Fig.38 , when the (Ns + 1)-th measurement location MS.Ns+1 is newly added, the multi-facility subject TS1 is made to tour all of the (Ns + 1) measurement locations again to perform the same process as the above harmonization calculation process, and the measurement bias can be recalculated.
[0824] Through the system that performs the above-described processing, it is possible to adjust and correct the measurement bias at each facility for the measurement data of the brain activity measured at multiple facilities. Thereby, it is possible to adjust the brain function connection-related values based on the measurement data at multiple facilities.
[0825] In addition, according to a system that performs the processing as in this embodiment, it is possible to implement a coordination method for a brain activity classifier, a coordination system for a brain activity classifier, a brain activity analysis system, and a brain activity analysis method that can coordinate measurement data of brain activities measured at multiple facilities to provide data for objectively judging the state of health or disease.
[0826] Alternatively, according to a system that performs the processing as in this embodiment, it is possible to implement a biomarker device based on a brain functional imaging method for neuro / psychiatric diseases and a program for the biomarker device for measurement data of brain activities measured at multiple facilities.
[0827] [Modification Example of Embodiment 1]
[0828] Furthermore, in the above description, the coordination calculation process has been described on the premise that the following subject data has been obtained at each measurement location.
[0829] i) Patient data
[0830] ii) Data of healthy subjects
[0831] iii) Data of multi-facility examinees
[0832] Among them, if it is only for the purpose of evaluating the above-mentioned "measurement bias", it is also possible to evaluate the "measurement bias" based on the data measured for multi-facility examinees by GLMM (Generalized Linear Mixed Model).
[0833] That is, let the number of multi-facility examinees be Nts people.
[0834] By fitting the regression model to the correlation values of the functional connectivity of all participants in the dataset of multi-facility examinees, the participant factor (p) and the measurement bias (m) are evaluated.
[0835] Here too, vectors are represented by lowercase letters (e.g., m), assuming that all vectors are column vectors.
[0836] The elements of the vector are represented by subscripts as in m k in this way.
[0837] The regression model of the functional connectivity vector (assumed to be a column vector) composed of n correlation values between brain regions is expressed by the following formula.
[0838] [Equation 21]
[0839] Connectivity = x m T m + x p T p + const + e
[0840]
[0841] To represent the characteristics of the participants, a 1-of-K (dummy coding) binary code system is used. For the target vector (e.g., xm) for the measurement bias m belonging to location k, all elements other than element k equal to 1 are equal to zero.
[0842] The superscript T represents the transposition of a matrix or vector, and x T represents a row vector.
[0843] With the above structure, it is also possible to set up a structure for performing harmonization processing by calculating the "measurement bias".
[0844] As the difference between the method of Embodiment 1 and the method of the modified example of the first embodiment, when only multi-facility examinees are used as in the method of the modified example of Embodiment 1, there is an advantage that sampling bias can be ignored. On the other hand, if both patient data and healthy subject data are included as in the method of Embodiment 1, there is an advantage that the amount of data used in the estimation increases.
[0845] Therefore, it can be said that there is a trade-off relationship between the improvement in the estimation accuracy of the measurement bias due to ignoring sampling bias and the improvement in the estimation accuracy of the measurement bias due to a large amount of data.
[0846] Therefore, although not particularly limited, for example, it is also possible to optimize the amount of data of the participants (patients and healthy subjects) used at each measurement location in actual operation through experiments. In this case, although not particularly limited, it is also possible to pre-assign the number of participants extracted from each measurement location and randomly extract the data of that number of people at each measurement location.
[0847] [Data of actual brain activity measurement results and their analysis]
[0848] Next, the results obtained by performing the above-described harmonization processing on the data of the actually disclosed multi-disease database will be described.
[0849] Fig.39A and Fig.39B are respectively conceptual diagrams showing the data of the multi-disease database used in the harmonization processing and the data set of multi-facility examinees.
[0850] As Fig.39AAs shown, datasets of patient groups and healthy groups using data publicly available as part of the multi-disease database (https: / / bicr-resource.atr.jp / decnefpro / ) in the Strategic Research Program for Brain Sciences (SRPBS).
[0851] As a multi-disease dataset, data measured at multiple measurement sites (in Fig.39A which typically shows Site 1, Site 2, and Site 3) contains sampling bias for psychiatric patients, sampling bias for healthy individuals, and measurement bias as differences between sites.
[0852] On the other hand, as Fig.39B shown, for datasets regarding multi-facility examinees, data measured at multiple measurement sites (in Fig.39B which typically shows Site 1, Site 2, and Site 3) contains only measurement bias as differences between sites.
[0853] It is possible to partition the between-site differences into measurement bias and sampling bias through simultaneous analysis of the datasets, and quantitatively compare the effect sizes of measurement bias and sampling bias on resting-state functional connectivity with the effect size of mental illness.
[0854] Regarding measurement bias, the effect sizes were quantitatively compared for different imaging parameters, MRI device manufacturers, and the number of receive coils in each MRI scanner.
[0855] To overcome these limitations associated with between-site differences, a harmonization method that can eliminate only measurement bias was performed using datasets of multi-facility examinees.
[0856] Based on the resting-state functional connectivity MRI data from multiple sites harmonized by the newly proposed method and existing methods, regression models of examinee information (attributes) (such as age) and biomarkers for mental illness were constructed.
[0857] The performance of these prediction models was investigated based on how the harmonization method changed.
[0858] (The datasets used)
[0859] Three resting-state functional MRI datasets were used as follows:
[0860] (1) Datasets of multiple diseases in SRPBS;
[0861] (2) 19 multi-facility examinee datasets; and
[0862] (3) Independent validation dataset.
[0863] (Multi-disease dataset of SRPBS)
[0864] Fig.40 ( FIG. 40A to FIG. 40D ) is a figure showing the content of the multi-disease dataset of SRPBS.
[0865] This dataset, which was examined at 9 sites and includes patients with 5 different diseases and healthy control groups (HC), contains a total of 805 participants as follows:
[0866] Data of 482 healthy individuals from 9 sites; data of 161 patients with major depressive disorder (MDD) from 5 sites; data of 49 patients with autism spectrum disorder (ASD) from 1 site; data of 65 patients with obsessive-compulsive disorder (OCD) from 1 site; and data of 48 patients with schizophrenia (SCZ) from 3 sites.
[0867] Fig.41 ( FIG. 41A to FIG. 41D ) is a figure showing the imaging protocol at each measurement site.
[0868] Resting-state functional MRI data were obtained using a unified imaging protocol at all sites except 3 sites. (http: / / www.cns.atr.jp / rs-fmri-protocol-2 / )
[0869] The site-to-site differences in this dataset include both measurement bias and sampling bias.
[0870] Regarding bias evaluation, only data obtained using the unified protocol were used.
[0871] In addition, since no OCD patients were scanned using this unified protocol, no evaluation of disease factors was performed for OCD.
[0872] The abbreviations in the table are as follows.
[0873] ATT: Siemens TimTrio scanner at the National Institute of Information and Communications Technology;
[0874] ATV: Siemens Verio scanner at the National Institute of Information and Communications Technology;
[0875] KUT: Siemens TimTrio scanner at Kyoto University;
[0876] SWA: Showa University;
[0877] HUH: Hiroshima University Hospital;
[0878] HKH: Hiroshima Kajikawa Hospital;
[0879] COI: COI (Hiroshima University);
[0880] KPM: Kyoto Prefectural University of Medicine;
[0881] UTO: The University of Tokyo;
[0882] ASD: Autism Spectrum Disorder.
[0883] MDD: Major Depressive Disorder.
[0884] OCD: Obsessive-Compulsive Disorder.
[0885] SCZ: Schizophrenia.
[0886] SIE: Siemens fMRI device.
[0887] GE: GE fMRI device.
[0888] PHI: Philips fMRI device.
[0889] (Dataset of multi-site subjects)
[0890] To evaluate the measurement bias across the measurement sites in the SRPBS dataset, a dataset of multi-site subjects was obtained.
[0891] Nine healthy participants (all male; age range: 24 - 32 years old; average age: 27 ± 2.6 years old) were imaged at 12 sites, including 9 of the sites where the SRPBS dataset was imaged, and a total of 411 procedures were performed.
[0892] Although an attempt was made to obtain this dataset using the same imaging protocol as the datasets of multiple diseases in SRPBS, there were some differences in the imaging protocols between sites due to parameter settings or conventional scanning conditions at each site.
[0893] For example, there were differences such as two directions of the phase encoding method (P→A and A→P), 3 MRI manufacturers (Siemens, GE, and Philips), 4 types of coil channel numbers (8, 12, 24, and 32), and 7 types of scanner models (TimTrio, Verio, Skyra, Spectral, MR750W, SignaHDxt, and Achieva).
[0894] Since the same 9 participants were imaged at 12 sites, the differences between sites in this dataset only included measurement bias.
[0895] (Independent validation dataset)
[0896] To verify the generalization performance of classifiers for mental disorders and models for predicting the age of participants based on resting-state functional connectivity MRI data, data from an independent validation cohort covering two disorders and seven sites were obtained.
[0897] This data included a total of 625 participants.
[0898] Specifically, data were obtained from 476 healthy controls (HCs) from six sites, 93 patients with MDD from two sites, and 56 patients with SCZ from one site.
[0899] (Visualization of between-site differences and disease effects)
[0900] First, the between-site differences and disease effects in the resting-state functional connectivity MRI datasets of multiple diseases in the SRPBS were visualized by principal component analysis (PCA: principal component analysis).
[0901] Fig.42 Fig. shows a graph visualizing the between-site differences and disease effects obtained by this principal component analysis.
[0902] The principal component analysis here is equivalent to a dimensionality reduction method based on unsupervised learning.
[0903] The functional connectivity of the subjects was calculated as described above as the temporal correlation (using the Pearson correlation coefficient) of the blood oxygenation level-dependent (BOLD) signals of the resting-state functional MRI between two brain regions for each participant.
[0904] Functional connectivity was defined based on a functional brain map consisting of 268 nodes (i.e., brain regions) covering the entire brain.
[0905] The values of the connection strengths (connectivities) of 35,778 (i.e., (268×267) / 2) in the lower triangular array of the matrix representing the correlation of functional connectivity were used.
[0906] As Fig.42 shown, the data of all participants in the datasets of multiple diseases in the SRPBS were all plotted on two axes composed of the first two principal components.
[0907] That is, all participants in the datasets of multiple diseases in the SRPBS were projected into the first two principal components (PCs) in the form of being represented by small light-colored markers.
[0908] For principal component 1, the HUH site was clearly separated, which is a situation that explains most of the variance in the data.
[0909] Patients with ASD were only scanned at the SWA site. Thus, the averages of the ASD (▲) patients and the healthy group HC (●) scanned at this site were plotted at almost the same position.
[0910] In Fig.42 the dimensionality reduction using PCA in the dataset showing data of multiple diseases is shown.
[0911] (Bias evaluation)
[0912] In order to quantitatively investigate the inter-site differences in the functional connectivity MRI data in the resting state, measurement bias, sampling bias, and diagnostic factors were determined.
[0913] As described above, the measurement bias at each site was defined as the deviation from the average of the correlation values of each functional connectivity at all sites. It is assumed that the sampling biases of healthy individuals and patients with mental diseases are different.
[0914] Therefore, for the healthy group and the patient groups with each disease, the sampling bias at each site was calculated independently.
[0915] The disease factor was defined as the deviation from the value of the healthy group.
[0916] Since the patient groups were sampled at multiple sites, the sampling bias was evaluated for the patients with MDD and SCZ.
[0917] As a control, in the dataset of multi-facility examinees, since the participants were fixed, only measurement bias was included.
[0918] By combining the multi-facility examinees with the datasets of multiple diseases of SRPBS, measurement bias and sampling bias were evaluated simultaneously as different factors affected by different sites. In order to evaluate the effects of the two biases and the disease factor on functional connectivity, the "linear mixed effects model" was used as follows.
[0919] (Linear mixed effects model for the datasets of multiple diseases of SRPBS)
[0920] In the linear mixed effects model, the correlation value of the connectivity of each examinee in the datasets of multiple diseases of SRPBS consists of fixed effects and random effects.
[0921] The fixed effects include the average correlation value involving all participants and all sites as the baseline, and the sum of measurement bias, sampling bias, and disease factor.
[0922] The effects generated by the combination of the participant factor (i.e., individual differences) and the variation between scans are regarded as random effects.
[0923] (Details of the evaluation of bias and factors)
[0924] The participant factor (p), measurement bias (m), sampling bias (shc, smdd, sscz), and mental illness factor (d) were evaluated by fitting a regression model to the correlation values of functional connectivity for all participants in datasets of multiple diseases of SRPBS and datasets of multi-facility examinees.
[0925] In this example, vectors are represented by lowercase bold letters (e.g., m), and it is assumed that all vectors are column vectors.
[0926] The elements of the vector such as m k are represented using suffixes as such.
[0927] Similarly to the above, in order to represent the characteristics of the participants, a binary code system of 1-of-K (dummy coding) is used, and for the target vector (e.g., xm) for the measurement bias m belonging to location k, all elements other than the element k equal to 1 are equal to zero.
[0928] Therefore, if the participant does not belong to any category, the target vector is a vector with all elements equal to 0.
[0929] The superscript T represents the transposition of a matrix or vector, and x T represents a row vector.
[0930] Regarding each connectivity, the regression model can be represented as follows.
[0931] [Equation 22]
[0932] Connectivity = x m T m + x Shc T s hc + x Smdd T s mdd + x Sscz T s scz + x d T d + x p T p + const + e
[0933] d1(HC) = 0
[0934] Among them, m represents the measurement bias (12 sites × 1), shc represents the sampling bias of the healthy group (6 sites × 1), smdd represents the sampling bias of MDD patients (3 sites × 1), sscz represents the sampling bias of SCZ patients (3 sites × 1), d represents the disease factor (3 diseases × 1), p represents the participant factor (9 multi-facility examinees × 1), const represents the average of the functional connectivity involving all participants from all sites, and e ∼ N(0, γ -1 ) represents noise.
[0935] For the correlation values of each functional connectivity, each parameter was evaluated using ordinary least squares regression based on the standard of L2 normalization.
[0936] In the case where normalization was not applied, spurious de-correlation between the measurement bias and sampling bias for the healthy group, and spurious correlation between the sampling bias for the healthy group and the sampling bias for patients with mental illness were observed. The hyperparameter λ was adjusted to minimize the absolute average of these spurious correlations.
[0937] (Linear mixed effects model for the dataset of multi-facility examinees)
[0938] In the linear mixed effects model for the dataset of multi-facility examinees, the correlation values of the connectivity of each participant for a specific scan in the dataset of multi-facility examinees include fixed effects and random effects.
[0939] The fixed effects include the average correlation value involving all participants and all sites, the participant factor, and the sum of the measurement biases.
[0940] The variation between scans is regarded as a random effect.
[0941] For each participant, the participant factor is defined as the deviation of the correlation value of the brain functional connectivity from the average involving all participants.
[0942] By simultaneously fitting the correlation values of the functional connectivity of two different datasets to the above two regression models, all biases and factors were evaluated.
[0943] In summary, the biases or each factor were evaluated as vectors in the dimension containing the number of correlation values reflecting connectivity (i.e., 35,778).
[0944] (Analysis of contribution size)
[0945] To quantitatively confirm the magnitude relationship between factors, in the linear mixed effects model, the contribution size was calculated and compared to determine to what extent each type of bias and factor explains the variance of the data.
[0946] [Number 23]
[0947]
[0948] For example, in this model, the contribution size of the measurement bias (i.e., the first term) is calculated as follows.
[0949] [Number 24]
[0950]
[0951] Here, Nm represents the number of elements of each factor, N represents the number of connectivities, Ns represents the number of subjects, and Contribution size m represents the degree of the contribution size of the measurement bias.
[0952] These equations are used to evaluate the contribution size of each factor related to the measurement bias (e.g., phase encoding direction, scanner, coil, and fMRI manufacturer).
[0953] In particular, the measurement bias is decomposed into these factors, and then the associated parameters are evaluated.
[0954] Other parameters are fixed to the same values as the parameters evaluated previously.
[0955] The following can be learned from the evaluation of the contribution size.
[0956] 1) Although the effect size of the measurement bias on functional connectivity shows less than that of the participant factor, it is mostly greater than that of the disease factor, indicating that the measurement bias becomes a significant limitation in the study of mental diseases.
[0957] 2) The largest variance in the sampling bias is significantly greater than the variance of the MDD factor, and the smallest variance in the sampling bias is one-half of the variance of the disease factor.
[0958] This indicates that the sampling bias also becomes a significant limitation in the study of mental diseases.
[0959] 3) The standard deviation of the participant factor is about twice that of the disease factors of SCZ, MDD, and ASD. Therefore, considering all functional connections, the variation among individuals within the healthy population is greater than the variation among patients with SCZ, MDD, and ASD.
[0960] 4) Also, the standard deviation of the measurement bias is mostly greater than the standard deviation of the disease factor. On the other hand, the standard deviation of the sampling bias is at the same level as the standard deviation of the disease factor.
[0961] This relationality makes it very difficult to develop a classifier for mental illness or intellectual disability based on resting-state functional connectivity MRI. A robust and generalizable classifier can be generated for multiple sites only when it is possible to select a very small number of disease-specific or site-independent abnormal functional connections from a large number of connections.
[0962] (Hierarchical clustering analysis for measurement bias)
[0963] For each site k, the Pearson correlation coefficient between the measurement biases mk (N×1, where N is the number of functional connectivities) was calculated, and hierarchical clustering analysis was performed based on the correlation coefficients involving the measurement biases.
[0964] Fig.43 is a dendrogram based on hierarchical clustering analysis.
[0965] The height of each linkage in the dendrogram represents the dissimilarity (1 - r) between the clusters connected by that linkage.
[0966] When investigating the characteristics of the measurement bias, it was investigated for 12 sites whether the similarity of the evaluated measurement bias vectors reflects specific characteristics of the MRI scanner such as the phase-encoding direction, MRI manufacturer, coil type, and scanner type.
[0967] To find clusters of the same pattern for the measurement bias, hierarchical clustering analysis was used.
[0968] As a result, as Fig.43 shown, the measurement biases of the 12 sites were divided into clusters of the phase-encoding direction at the first level.
[0969] The measurement bias was divided into clusters of fMRI manufacturers at the second level, further divided into clusters of coil types, and also divided into clusters of scanner models.
[0970] Fig.44 is a graph showing the contribution sizes of the respective factors.
[0971] As Fig.44 shown, in order to evaluate the contributions of the respective factors, the magnitude relationship of the factors was quantitatively confirmed by using the same model.
[0972] The contribution sizes have the following relationship: the phase-encoding direction (0.0391) is the largest, followed by the fMRI manufacturer (0.0318), the coil type (0.0239), and the scanner model (0.0152).
[0973] These show that the main factor affecting the measurement bias is the phase encoding direction, followed by differences in fMRI manufacturers, coil types, and scanner models.
[0974] (Visualization of harmonization effect)
[0975] Next, a dataset of multi-site subjects is used to illustrate a harmonization method that can only remove measurement bias.
[0976] As described above, a linear mixed effects model was used to evaluate the measurement bias independently of the sampling bias.
[0977] By this method, it is possible to remove only the measurement bias from the datasets of multiple diseases in SRPBS and maintain the sampling bias containing biological information.
[0978] (Harmonization of multi-site subjects)
[0979] To evaluate the measurement bias, the phase encoding direction at the HKH site is different between the datasets of multiple diseases in SRPBS and the dataset of multi-site subjects. Therefore, the phase encoding factor (pa q ) is set to be included separately from the measurement bias in the equation below.
[0980] The measurement bias and the phase encoding factor are evaluated by fitting a regression model to the dataset combined from the datasets of multiple diseases in SRPBS and the dataset of multi-site subjects.
[0981] The regression model is as follows.
[0982] [Equation 25]
[0983] Connectivity = x m T m + x Shc T s hc + x Smdd T s mdd + x Sscz T s scz + x d T d + x p T p + x pa T pa + const + e
[0984] d1(HC) = 0,
[0985] Here, pa represents the phase encoding factor (two phase encoding directions × 1).
[0986] After performing normalization on the correlation values of each functional connectivity using least squares regression based on standard LS normalization and then using them, the parameters were evaluated respectively.
[0987] The difference between locations was removed by subtracting the evaluated difference between locations and the phase encoding factor.
[0988] Therefore, the sampling bias is as follows for the correlation values of the harmonized functional connectivity.
[0989] [Number 26]
[0990]
[0991] Among them, m represents the evaluated measurement bias, and pa (hat) represents the evaluated phase encoding factor.
[0992] Fig.45 is a graph that visualizes the influence of the harmonization process and is a graph for comparison with Fig.42
[0993] In Fig.45 , the data was plotted after only removing the measurement bias from the datasets of multiple diseases in SRPBS.
[0994] In Fig.42 reflects the data before harmonization, compared with the data shown in Fig.42 , in Fig.45 , the HUH location has moved significantly closer to the origin side (that is, the overall average), and no longer deviates significantly from other locations.
[0995] Fig.45 The results shown in Fig.42 indicate that the separation of the HUH location seen in
[0996] is caused by the measurement bias, and this separation can be removed by harmonization.
[0997] Patients with ASD were only scanned at the SWA location. Therefore, the averages of ASD patients (▲) and healthy subjects (○) scanned at this location were plotted at almost the same position in Fig.42 .
[0998] However, in Fig.45 , the two symbols are clearly separated from each other.
[0999] [Embodiment 2]
[1000] In Embodiment 1, as a structure for measuring brain activity data measured at a plurality of measurement sites by a brain activity measurement device (fMRI device) and generating a biomarker and estimating (predicting) a diagnostic label using the biomarker based on the brain activity data, a structure of an example performed by variance processing was described.
[1001] However, it is also possible to adopt a structure in which the following processes are separately distributed and executed in different facilities: i) measurement (data collection) of brain activity data for training a biomarker by machine learning; ii) a generation process of generating a biomarker by machine learning and a process of estimating (predicting) a diagnostic label for a specific subject using the biomarker (estimation process); iii) measurement of brain activity data of the specific subject (measurement of the brain activity of the subject).
[1002] Fig.46 It is a functional block diagram showing an example in the case where data collection, estimation processing, and measurement of the brain activity of the subject are performed separately.
[1003] Refer to Fig.46 , locations 100.1 to 100.N are facilities for measuring data of a patient group and a healthy subject group by a brain activity measurement device, and the data center 200 manages the measurement data from locations 100.1 to 100.Ns.
[1004] The calculation processing system 300 generates an identifier based on the data stored in the data center 200.
[1005] In addition, it is assumed that the coordination calculation unit 3020 of the calculation processing system 300 performs coordination processing including locations 100.1 to 100.Ns and the location of the MRI device 410.
[1006] The MRI device 410 is provided at another location that uses the result of the identifier on the calculation processing system 300, and measures brain activity data of a specific subject.
[1007] The computer 400 is provided at another location where the MRI device 410 is provided, calculates correlation data of the functional connectivity of the brain of a specific subject based on the measurement data of the MRI device 410, sends the correlation data of the functional connectivity to the calculation processing system 300, and uses the result of the returned identifier.
[1008] The data center 200 stores the MRI measurement data 3102 of the patient group and the healthy group sent from locations 100.1 to 100.Ns, as well as the personal attribute information 3104 of the subjects associated with the MRI measurement data 3102, and sends these data to the computing processing system 300 according to the access by the computing processing system 300.
[1009] The computing processing system 300 receives the MRI measurement data 3102 and the personal attribute information 3104 of the subjects from the data center 200 via the communication interface 2090.
[1010] In addition, the hardware structures of the data center 200, the computing processing system 300, and the computer 400 are basically the same as those of the "data processing unit 32" described in Figure 5 and thus the description thereof will not be repeated.
[1011] Return to Fig.46 , regarding the correlation matrix calculation unit 3002, the correlation value correction processing unit 3004, the disease identifier generation unit 3008, the clustering classifier generation unit 3010 and the discrimination value calculation unit 3012, as well as the data 3106 of the correlation matrix of functional connectivity, the measurement bias data 3108, the corrected correlation value data 3110 and the identifier data 3112, they are the same as those described in Embodiment 1 and thus the description thereof will not be repeated.
[1012] The MRI device 410 measures the brain activity data of the subject who is the estimation object of the diagnostic label, and the processing device 4040 of the computer 400 stores the measured MRI measurement data 4102 in the non-volatile storage device 4100.
[1013] Moreover, the processing device 4040 of the computer 400 calculates the data 4106 of the correlation matrix of functional connectivity in the same manner as the correlation matrix calculation unit 3002 based on the MRI measurement data 4102 and stores it in the non-volatile storage device 4100.
[1014] The disease to be diagnosed is specified by the user of the computer 400. According to the instruction sent by the user, the computer 400 sends the data 4106 of the correlation matrix of functional connectivity to the computing processing system 300. In response to this transmission, the computing processing system 300 performs the coordination processing corresponding to the location where the MRI device 410 is set, and the discrimination value calculation unit 3012 calculates the discrimination result regarding the specified diagnostic label and the evaluation result regarding the subtype, and the computing processing system 300 sends them to the computer 400 via the communication interface 2090.
[1015] In the computer 400, the discrimination result is notified to the user via a display device (not shown) or the like.
[1016] By adopting such a structure, it is possible to provide an estimation result of a diagnostic label obtained by an identifier based on data collected from a larger number of subjects.
[1017] In addition, it is also possible to adopt a method in which the data center 200 and the computing and processing system 300 are managed by separate administrators. In this case, by restricting access to the computers that can access the data center 200, the security of the information of the subjects stored in the data center 200 can be further improved.
[1018] Moreover, when viewed from the operating entity of the computing and processing system 300, even if no information about the identifier and information related to "measurement bias" are provided to "the service side (computer 400) that receives discrimination by the identifier", the "service for providing discrimination results" can still be carried out.
[1019] In addition, in the descriptions of the above Embodiment 1 and Embodiment 2, as the brain activity detection device for measuring brain activity in time series using the brain functional imaging method, real-time fMRI is used for illustration. However, as the brain activity detection device, the above-mentioned fMRI, magnetoencephalograph, near-infrared spectroscopy (NIRS), electroencephalograph, or a combination thereof can be used. For example, in the case of using a combination thereof, fMRI and NIRS detect signals associated with blood flow changes in the brain and have high spatial resolution. On the other hand, magnetoencephalographs and electroencephalographs are characterized by high temporal resolution for detecting changes in electromagnetic fields associated with brain activity. Therefore, for example, if fMRI is combined with a magnetoencephalograph, brain activity can be measured with high resolution both spatially and temporally. Or, even if NIRS is combined with an electroencephalograph, it is also possible to construct a system that can measure brain activity with high resolution both spatially and temporally in a small and portable size.
[1020] Through the above structure, for neuro / psychiatric diseases, a brain activity analysis device and a brain activity analysis method that can function as biomarkers based on the brain functional imaging method can be realized.
[1021] In addition, in the above description, the following example is illustrated: in the case of including "diagnostic labels" as attributes of subjects, a classifier is generated by machine learning to make the classifier function as a biomarker. However, the present invention is not necessarily limited to this case. As long as the group of subjects, which is the object for obtaining measurement results to be the object of machine learning, is divided into multiple classes in advance by an objective method, and the correlation (connection) of the activity levels between brain regions (regions of interest) of the subjects is measured, and a classifier for the classes can be generated by performing machine learning on the measurement results, it can also be used for other discriminations.
[1022] In addition, as described above, such discrimination can also display the probability of belonging to a certain attribute as a probability.
[1023] Therefore, for example, by adopting a certain "training" and "action pattern", it is possible to objectively evaluate whether it is helpful to improve the health of the subject. In addition, in fact, even in a state where the subject has not contracted a disease ("not yet ill"), it is possible to objectively evaluate whether certain intakes such as "food" and "drink" and certain activities are effective in approaching a healthier state.
[1024] In addition, in the state of not yet being ill, as described above, for example, as long as a display such as "the probability of being healthy is ○○%" is output, it is possible to display an objective numerical value regarding the health state to the user. At this time, what is output is not necessarily a probability, and it can also be set to display a value obtained by converting a "continuous value of the degree of health, such as the probability of being healthy" into a score. By performing such a display, in addition to being able to use the device of the present embodiment for auxiliary diagnosis, it can also be used as a device for the health management of users.
[1025] [Embodiment 3]
[1026] In the above description, it is configured to evaluate the measurement bias by having the multi-facility examinee move equally to all measurement locations and perform measurements.
[1027] Fig.47 FIG. shows the tour mode of the multi-facility examinee in Embodiment 3.
[1028] As Fig.47 shown, it is premised that the multi-facility examinee TS1 tours the "base measurement locations MS.1 to MS.Ns" that serve as bases.
[1029] In contrast, for example, regarding the base measurement location MS.2 among these base measurement locations MS.1 to MS.Ns, it is preset that there are subordinate MS2.1 to MS2.n. In the case where a "measurement location MS.2.n+1" is newly added, the multi-facility examinee TS2 tours within this subordinate range. The same applies to other base measurement locations.
[1030] That is, it is also possible to configure as follows: the measurement bias of the base measurement location MS.2 is fixed to the value evaluated for the multi-facility examinee TS1, and regarding the measurement locations MS.2, MS2.1, MS2.n, and MS2.n+1, the measurement bias is determined based on the measurement results obtained by the multi-facility examinee TS2 touring.
[1031] For example, assume that the "base measurement locations MS.1 to MS.Ns" exist in predetermined regions, and consider the following structure: one is provided in each of the regions such as Hokkaido, Tohoku, Kanto, ···, Kansai, ···, Kyushu within Japan, and the subordinate location is, for example, the measurement location in Kansai, which is one of the regions.
[1032] Alternatively, the "base measurement locations MS.1 to MS.Ns" can also be determined according to the type of MRI device determined in advance. In this case, the subordinate measurement location refers to the measurement location where an MRI device of the same type as the base location is installed.
[1033] Alternatively, the "base measurement locations MS.1 to MS.Ns" can also be determined according to the predetermined region and the type of MRI device determined. In this case, the subordinate measurement location refers to the measurement location where an MRI device of the same type is installed within the same region as the base location.
[1034] Such a structure can also achieve the same effect as that of Embodiment 1.
[1035] The embodiments disclosed herein are examples of the structures for specifically implementing the present invention and do not limit the protection scope of the present invention. The protection scope of the present invention is not the scope described in the embodiments, but the scope represented by the claims, and is intended to include the scope of the content of the claims and modifications within the scope of equivalent meaning.
[1036] Description of Reference Numerals
[1037] 2: Subject; 6: Display; 10: MRI device; 11: Magnetic field application mechanism; 12: Static magnetic field generation coil; 14: Gradient magnetic field generation coil; 16: RF irradiation unit; 18: Bedding; 20: Receiver coil; 21: Drive unit; 22: Static magnetic field power supply; 24: Gradient magnetic field power supply; 26: Signal transmission unit; 28: Signal reception unit; 30: Bedding drive unit; 32: Data processing unit; 36: Storage unit; 38: Display unit; 40: Input unit; 42: Control unit; 44: Interface unit; 46: Data collection unit; 48: Image processing unit; 50: Network interface.
Claims
1. A clustering device for brain function connection related values, which is used to perform clustering of the subject with at least one specified attribute in the subject based on the measurement results of the brain activity of the subject. The clustering device for brain function connection related values is equipped with a computing processing system, which is used to perform the clustering process for multiple subjects including a first group of subjects with the specified attribute and a second group of subjects without the specified attribute based on the measured values of brain activity. The computing processing system includes a storage device and an arithmetic device. The arithmetic device is configured as follows: i) For each of the multiple subjects, save the feature quantities of multiple brain function connection related values representing the temporal correlation of brain activities between specified multiple pairs of brain regions into the storage device. ii) Based on the feature quantities saved in the storage device, perform machine learning for generating an identifier model for discriminating the presence or absence of the attribute by supervised learning. The arithmetic device performs the following processing in the machine learning for generating the identifier model: According to the first group of subjects and the second group of subjects, perform undersampling and downsampling to generate multiple learning sub-samples. For each of the learning sub-samples of the learning sub-samples, select the feature quantities for clustering from the union of the feature quantities used when generating the identifier model by machine learning according to the importance of the feature quantities belonging to the union. The arithmetic device also performs the following processing: Based on the selected feature quantities for the clustering, perform clustering on the first group of subjects by the co-clustering method of unsupervised learning to generate a classifier of clusters.
2. The clustering device for brain function connection related values according to claim 1, wherein The clustering device for brain function connection related values receives information representing the temporal correlation of brain activities between specified multiple pairs of brain regions of each of the multiple subjects from multiple brain activity measurement devices respectively arranged at multiple measurement locations. The computing processing system includes a coordination computing unit, which corrects the multiple brain function connection related values of each of the multiple subjects to remove the measurement bias of the measurement location, and saves the adjusted values obtained by the correction into the storage device as the feature quantities.
3. The clustering device for brain function connection related values according to claim 1 or 2, wherein The process of generating the identifier model by machine learning is as follows: integrated learning: generate multiple identifier sub-models for the multiple learning sub-samples respectively, and integrate the multiple identifier sub-models to generate the identifier model.
4. The clustering device for brain function connection related values according to claim 2, wherein The attribute is represented by a label of a diagnosis result of a specified mental illness. The clustering is a process of classifying the first group of subjects into at least one subtype cluster by machine learning based on data driving.
5. The clustering device for brain function connection related values according to claim 2, wherein The arithmetic device performs the following processes when generating the recognizer model by machine learning: i) Divide the adjustment value into a training data set for machine learning and a test data set for verification; ii) Perform a prescribed number of undersampling and downsampling on the training data set to generate the prescribed number of learning sub-samples; iii) Generate a recognizer sub-model for each of the learning sub-samples; iv) Integrate the outputs of the recognizer sub-models to generate a recognizer model for the presence or absence of the attribute.
6. The clustering device for brain functional connection-related values according to claim 2, wherein the process of generating the recognizer model by machine learning is a cross-validation with a nested structure having outer cross-validation and inner cross-validation, and the arithmetic device performs the following processes in the process of the nested structure cross-validation: i) Set the outer cross-validation as K-fold cross-validation to divide the adjustment value into a training data set for machine learning and a test data set for verification; ii) Perform a prescribed number of undersampling and downsampling on the training data set to generate the prescribed number of learning sub-samples; iii) In each cycle of the K-fold cross-validation, adjust hyperparameters by the inner cross-validation to generate a recognizer sub-model for each of the learning sub-samples; iv) Generate a recognizer model for the presence or absence of the attribute based on the recognizer sub-model.
7. The clustering device for brain functional connection-related values according to claim 3, wherein the process of generating the recognizer model by machine learning is a machine learning method accompanied by feature quantity selection, in the selection of the feature quantity for the clustering, the importance of the feature quantity is determined according to the ranking of the frequencies of the feature quantities belonging to the union set when generating the recognizer sub-model.
8. The clustering device for brain functional connection-related values according to claim 3, wherein the process of generating the recognizer model by machine learning is the random forest method, in the selection of the feature quantity for the clustering, the importance of the feature quantity belonging to the union set is the importance calculated for each feature quantity based on the Gini impurity in the random forest method.
9. The clustering device for brain functional connection-related values according to claim 3, wherein the process of generating the recognizer model by machine learning is a machine learning method based on L2 regularization, in the selection of the feature quantity for the clustering, the importance of the feature quantity belonging to the union set is determined according to the ranking based on the weights of the feature quantities in the recognizer sub-model calculated by L2 regularization.
10. The clustering device for brain functional connection-related values according to claim 2, wherein the storage device pre-stores the results obtained by measuring brain activities in a plurality of predetermined brain regions for each of a plurality of traveling subjects who are commonly measured objects at the plurality of measurement locations, and the arithmetic device performs the following processes: For each of the travel examinees, a specified element of the brain functional connectivity array is calculated, and the brain functional connectivity array represents the temporal correlation of brain activities of the plurality of brain region pairs; By using the general linear mixed model method, the measurement bias is calculated for each specified element of the brain functional connectivity array as a fixed effect on average for this element at each measurement site with respect to the plurality of measurement sites and the plurality of travel examinees.
11. The clustering device for brain functional connectivity related values according to claim 4, wherein The arithmetic device performs classification processing into the subtype clusters based on measurement data measured for the subject at measurement sites other than the plurality of measurement sites.
12. A clustering system for brain functional connectivity related values, which is configured to perform clustering of the subject having at least one specified attribute among the subjects based on measurement results of brain activities of the subject, the clustering system for brain functional connectivity related values comprising: A plurality of brain activity measurement devices, which are respectively arranged at a plurality of measurement sites to measure the brain activities of a plurality of examinees in time series, the plurality of examinees including a first examinee group having the specified attribute and a second examinee group not having the specified attribute; And A calculation processing system, which is configured to perform the clustering process on the plurality of examinees based on the measured values of brain activities, The calculation processing system includes a storage device and an arithmetic device, The arithmetic device is configured to: i) For each of the plurality of examinees, save a feature quantity based on a plurality of brain functional connectivity related values respectively representing the temporal correlation of brain activities between specified pairs of a plurality of brain regions into the storage device; ii) Based on the feature quantity saved in the storage device, perform machine learning for generating a discriminator model for discriminating the presence or absence of the attribute by supervised learning, The arithmetic device performs the following processing in the machine learning for generating the discriminator model: According to the first examinee group and the second examinee group, perform undersampling and downsampling to generate a plurality of learning sub-samples; For each of the learning sub-samples of the learning sub-samples, select feature quantities for clustering from the union of the feature quantities used when generating the discriminator model by machine learning according to the importance of the feature quantities belonging to the union; The arithmetic device further performs the following processing: Based on the selected feature quantities for the clustering, perform clustering on the first examinee group by the co-clustering method of unsupervised learning to generate a classifier for clusters.
13. The clustering system for brain functional connectivity related values according to claim 12, wherein The calculation processing system receives information representing the temporal correlation of brain activities between specified pairs of a plurality of brain regions of each of the plurality of examinees from a plurality of brain activity measurement devices respectively arranged at a plurality of measurement sites, The computing and processing system includes a coordinated computing unit, which corrects the multiple brain function connection related values of each subject among the multiple subjects to remove the measurement bias of the measurement location, and thus stores the adjusted values obtained by the correction as the feature quantities in the storage device.
14. The clustering system for brain function connection related values according to claim 12 or 13, wherein the attribute is represented by a label of a diagnosis result of a specified mental illness the clustering is a process of classifying the first subject group into at least one subtype cluster by data-driven machine learning.
15. A method for clustering brain function connection related values, which is used for a computing and processing system to perform clustering processing on an object person with at least one specified attribute based on the measurement result of the brain activity of the object person, the computing and processing system includes a storage device and an arithmetic device, the method for clustering brain function connection related values includes the following steps: the arithmetic device stores, for each of a plurality of subjects, a feature quantity based on a brain function connection related value representing the temporal correlation of brain activities between a specified plurality of pairs of brain regions in the storage device, and the plurality of subjects include a first subject group having the specified attribute and a second subject group not having the specified attribute; and the arithmetic device performs machine learning for generating an identifier model for discriminating the presence or absence of the attribute by supervised learning based on the feature quantities stored in the storage device. The step of performing machine learning for generating the identifier model includes the following steps: performing undersampling and downsampling according to the first subject group and the second subject group to generate a plurality of learning sub-samples; and for each of the learning sub-samples of the learning sub-samples, selecting feature quantities for clustering from the union of the feature quantities used when generating an identifier model by machine learning according to the importance of the feature quantities belonging to the union. The method for clustering brain function connection related values further includes the following step: the arithmetic device clusters the first subject group by the multi-co-clustering method of unsupervised learning based on the selected feature quantities for the clustering to generate a classifier of clusters.
16. A computer program product, which includes a classifier program for brain function connection related values, and the classifier program for brain function connection related values is generated by a computing and processing system performing clustering processing on an object person with at least one specified attribute based on the measurement result of the brain activity of the object person, and the classifier program for brain function connection related values is used for a computer to classify input data into each cluster corresponding to the result of the clustering processing. The classifier program has the following classification function: the computer classifies the input data into the cluster with the maximum posterior probability based on the model of the probability distribution of each cluster. The computing and processing system includes a storage device and an arithmetic device, the computing and processing system performs the following steps in the generation process of generating the classifier program based on the clustering processing: For each of a plurality of subjects, the arithmetic device stores a feature quantity of a brain functional connection correlation value indicating a temporal correlation of brain activities between a plurality of prescribed brain region pairs in the storage device, the plurality of subjects including a first subject group having the prescribed attribute and a second subject group not having the prescribed attribute; and Based on the feature quantities stored in the storage device, the arithmetic device performs machine learning for generating an identifier model for discriminating the presence or absence of the attribute by supervised learning. The steps of performing machine learning for generating the identifier model include the following steps: Performing undersampling and downsampling according to the first subject group and the second subject group to generate a plurality of learning sub-samples; and For each of the learning sub-samples of the learning sub-samples, from the union of the feature quantities used when generating an identifier model by machine learning, feature quantities for clustering are selected according to the importance of the feature quantities belonging to the union. The arithmetic device further performs the following steps: Based on the selected feature quantities for the clustering, the first subject group is clustered by the co-clustering method of unsupervised learning to generate a classifier for clusters.
17. The computer program product according to claim 16, wherein The computing and processing system performs the following steps: Receiving information representing the temporal correlation of brain activities between a plurality of specified brain regions of each subject of the plurality of subjects from a plurality of brain activity measurement devices respectively provided at a plurality of measurement locations; and Performing coordination for correcting the plurality of brain functional connection related values of each subject of the plurality of subjects to remove the measurement bias of the measurement location, and storing the adjusted value obtained by the correction as the feature quantity in the storage device.
18. The computer program product according to claim 16 or 17, wherein The attribute is represented by a label of a diagnosis result of a specified mental illness. The clustering is a process of classifying the first subject group into at least one subtype of clusters by machine learning based on data driving.
19. A brain activity marker classification system is generated by performing a clustering process on an object person having at least one specified attribute in the object person based on a measurement result of the brain activity of the object person by a computing and processing system. The brain activity marker classification system is used for a computer to classify input data into each cluster corresponding to the result of the clustering process. The brain activity marker classification system has the following classification function: The computer classifies the input data into a cluster having the maximum posterior probability based on a model of the probability distribution of each cluster. The computing and processing system includes a storage device and an arithmetic device. The computing and processing system performs the following steps in a generation process of generating the brain activity marker classification system based on the clustering process: For each of a plurality of subjects, the arithmetic device stores, in the storage device, a feature quantity based on a brain functional connectivity correlation value representing a temporal correlation of brain activities between a plurality of prescribed brain region pairs, the plurality of subjects including a first subject group having the prescribed attribute and a second subject group not having the prescribed attribute; and Based on the feature quantities stored in the storage device, the arithmetic device performs machine learning for generating an identifier model for discriminating the presence or absence of the attribute by supervised learning. The steps of performing machine learning for generating the identifier model include the following steps: Performing undersampling and downsampling according to the first subject group and the second subject group to generate a plurality of learning sub-samples; and For each of the learning sub-samples of the learning sub-samples, from the union of the feature quantities used when generating the recognizer model through machine learning, according to the importance of the feature quantities belonging to the union, select the feature quantities for clustering. The arithmetic device also performs the following steps: Based on the selected feature quantities for the clustering, perform clustering on the first subject group by the multi-co-clustering method of unsupervised learning to generate a classifier for clusters.
20. The brain activity marker classification system according to claim 19, wherein The computing and processing system performs the following steps: Receive information representing the temporal correlation of brain activities between specified pairs of brain regions of each subject among the multiple subjects from multiple brain activity measurement devices respectively provided at multiple measurement locations; And Perform coordination, where the coordination is used to correct the multiple brain functional connectivity related values representing the temporal correlation of the brain activities of each subject among the multiple subjects to remove the measurement bias of the measurement location, and thus save the adjusted values obtained by the correction as the feature quantities in the storage device.
21. The brain activity marker classification system according to claim 19 or 20, wherein The attribute is represented by a label of a diagnosis result of a specified mental illness, The clustering is a process of classifying the first subject group into at least one subtype cluster through machine learning based on data driving.
22. A clustering classifier program product for brain functional connectivity related values, which is generated by a computing and processing system performing a clustering process on an object person having at least one specified attribute based on a measurement result of the brain activity of the object person, and the clustering classifier program product for brain functional connectivity related values is used by a computer to classify input data into each cluster corresponding to the result of the clustering process. The clustering classifier program product has the following function: For each view obtained by dividing the group of feature quantities representing the object person included in the learning data, according to the information of the feature quantities included in each view and the information of the probability density function for determining each cluster of the object person in each view, classify the input data into the cluster with the maximum posterior probability according to the value of the probability density function calculated for the input data. The computing and processing system includes a storage device and an arithmetic device. The computing and processing system performs the following steps in the generation process of generating the clustering classifier program product based on the clustering process: For each of a plurality of subjects, the arithmetic device stores, in the storage device, a feature quantity based on a brain functional connection correlation value representing a temporal correlation of brain activities between a plurality of prescribed brain region pairs, the plurality of subjects including a first subject group having the prescribed attribute and a second subject group not having the prescribed attribute; And The arithmetic device performs machine learning for generating a recognizer model for discriminating the presence or absence of the attribute with supervised learning based on the feature quantities stored in the storage device. The step of performing machine learning for generating the recognizer model includes the following steps: According to the first subject group and the second subject group, perform undersampling and downsampling to generate multiple learning sub-samples; And For each learning sub-sample of the learning sub-samples, from the union of the feature quantities used when generating the recognizer model by machine learning, feature quantities for clustering are selected according to the importance of the feature quantities belonging to the union. The arithmetic device also performs the following steps: Based on the selected feature quantities for the clustering, the first subject group is clustered by the unsupervised learning multi-co-clustering method, the feature quantities are segmented into the views, and probability density functions of each cluster of the object persons in each view are generated.
Citation Information
Patent Citations
Brain activity analysis device, brain activity analysis method, and biomarker device
JP2015062817A
Medical image processing apparatus and magnetic resonance imaging apparatus
JP2015112474A
Brain activity analyzer, brain activity analysis method, program, and biomarker device
JP2017196523A
Medical image processor, medical image processing method, and medical image processing system
JP2019198376A
Analysis method and array used therefor
JP2019516950A