Method and device for assessing neurodegenerative diseases and precursor disorders
Patent Information
- Application Number
- JP2025536057
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-25
- Estimated Expiration
- 2045-03-17
AI Technical Summary
Clinical differentiation of interrelated neurodegenerative diseases and precursor disorders is difficult due to the invasive nature of pathological diagnosis, especially in early stages, and there is a lack of non-invasive methods to assess the risk of these conditions using saliva samples.
A method using machine learning to analyze the composition of a subject's salivary microbiome through multiple prediction models, trained on microbiome data and disease states, to determine the risk of neurodegenerative diseases and precursor disorders.
Enables accurate determination of neurodegenerative disease risk and precursor disorders using saliva samples without invasive procedures, improving diagnostic accuracy by combining inference results from multiple models.
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a method and apparatus for diagnosing neurodegenerative diseases and precursor disorders. [Background technology]
[0002] Dementia (DE) is a condition in which cognitive function that was once normal is persistently impaired due to acquired brain damage, causing interference with daily and social life. Dementia is a neurodegenerative disease of the central nervous system, and representative types include Alzheimer's disease, dementia with Lewy bodies (DLB), vascular dementia, frontotemporal dementia, and dementia associated with Parkinson's disease (PD).
[0003] Alzheimer's disease, which causes Alzheimer's-type dementia, is a type of proteopathy in which certain proteins become structurally abnormal, condense, or accumulate, disrupting the function of tissues and organs. In Alzheimer's disease, the tau protein typically becomes structurally abnormal and accumulates in the brain, making Alzheimer's disease a tauopathy, a type of proteopathy. Another common type of proteopathy is synucleinopathy, a neurodegenerative disorder characterized by abnormal aggregation of α-synuclein protein. Synucleinopathy includes Parkinson's disease (PD) and multiple system atrophy (MSA), as well as dementia with Lewy bodies. Synucleinopathy can overlap with dementia, as exemplified by dementia with Lewy bodies and some forms of Parkinson's disease.
[0004] Mild cognitive impairment (MCI) is an intermediate state in which cognitive function is not normal but does not (yet) meet the diagnostic criteria for dementia, and is considered a precursor to dementia. Rapid eye movement (REM) sleep behavior disorder (RBD) is a parasomnia characterized by abnormal behavior during REM sleep. REM sleep behavior disorder is known to often progress to synucleinopathies such as Parkinson's disease, dementia with Lewy bodies, and multiple system atrophy. MCI and RBD are considered precursors to neurodegenerative diseases. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Alzheimer's Dementia,2020,vol.12,e12000 [Non-patent document 2] PlosOne 2019 vol.14,No.6,e0218252 Summary of the Invention [Problem to be solved by the invention]
[0006] The aforementioned diseases and disorders are interrelated and often overlap, making clinical differentiation difficult. While the presence of structural abnormalities in the brain that cause disease or disorder can be confirmed by pathological diagnosis, performing pathological diagnosis, such as brain imaging, before death, especially in the early stages, is often difficult not only for technical reasons but also because of the burden it places on individual subjects. Therefore, a technology that can identify which of multiple neurodegenerative diseases and prodromal disorders a subject is likely to have, or not, using a simpler, non-invasive method would be beneficial. Furthermore, a technology that can distinguish the presence of mild cognitive impairment, a precursor to neurodegenerative diseases, from normal events that occur with aging, or that can distinguish between disease states and prodromal disorders, and that can provide these capabilities without subjecting subjects to traditional pathological or clinical diagnostic processes, would be highly beneficial.
[0007] Non-Patent Document 1 reports that logistic regression analysis based on salivary microbiome profiling was used to construct a prediction model based on several specific individual bacterial species, making it possible to distinguish between mild cognitive impairment (MCI) and Alzheimer's disease. Non-Patent Document 2 reports that logistic regression analysis based on microbiome genus and species data obtained by salivary RNA sequencing was able to distinguish between early-stage Parkinson's disease (PD) patients and healthy individuals using data on 11 bacterial species.
[0008] However, there has been no attempt to assess the risk of precursors to neurodegenerative diseases based on the composition pattern of salivary microbiota.
[0009] It is also desirable to be able to more accurately assess the risk of neurodegenerative diseases and precursor disorders of neurodegenerative diseases using saliva samples.
[0010] The present disclosure has been made in light of the above, and aims to determine the risk of neurodegenerative diseases and precursor disorders of neurodegenerative diseases using saliva samples. [Means for solving the problem]
[0011] A method for determining a neurodegenerative disease and its precursor disorder includes a computer inputting data on the composition of a subject's salivary microbiome into each of a plurality of prediction models, obtaining an inference result from each of the plurality of prediction models, and determining a disease state of the subject by combining the plurality of inference results. Each of the plurality of prediction models is trained by machine learning to stratify the disease state of the subject into two groups when the data on the composition of the subject's salivary microbiome is input using the data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data. Each of the plurality of prediction models is trained by machine learning to stratify the disease state of the subject into two groups when the data on the composition of the salivary microbiome of the subject is input using the data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data. Stratification do.
[0012] A determination device according to one aspect of the present disclosure is a determination device for determining neurodegenerative diseases and precursor disorders thereof, comprising an inference unit that inputs data on the composition of a subject's salivary microbiome into each of a plurality of prediction models and obtains an inference result from each of the plurality of prediction models, and a determination unit that combines the inference results to determine the disease state of the subject, wherein each of the plurality of prediction models is trained by machine learning to stratify the disease state of the subject into two groups when the data on the composition of the salivary microbiome of the subject is input using the data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data, and each of the plurality of prediction models predicts a disease state that is different from each other. Stratification do. [Effects of the Invention]
[0013] According to the present disclosure, saliva samples can be used to determine risk for neurodegenerative diseases and precursors to neurodegenerative diseases. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a risk assessment system according to this embodiment. [Figure 2] Figure 2 shows an example of a prediction model for stratifying six disease states, including HC. [Figure 3] FIG. 3 is a diagram illustrating an example of the performance of a prediction model. [Figure 4] FIG. 4 is a diagram illustrating an example of the performance of a prediction model. [Figure 5] FIG. 5 is a diagram showing an example of determining the risk of a disease state of a subject by combining the inference results of prediction models. [Figure 6] FIG. 6 is a diagram showing an example of determining the risk of a disease state of a subject by combining the inference results of prediction models. [Figure 7] FIG. 7 is a diagram showing an example of determining the risk of a disease state of a subject by combining the inference results of prediction models. [Figure 8]FIG. 8 is a flowchart showing an example of the determination method. [Figure 9] FIG. 9 is a flowchart showing an example of a learning method. [Figure 10] FIG. 10 is a diagram illustrating an example of a hardware configuration of a determination device. DETAILED DESCRIPTION OF THE INVENTION
[0015] [System Configuration] An example of the configuration of the risk assessment system of this embodiment will be described with reference to FIG. 1. The systems, devices, and methods described herein are related to one another; for example, an embodiment of an device may be suitable for implementing an embodiment of a method. Therefore, it will be understood by those skilled in the art that the description provided herein can be similarly and reciprocally applied to different embodiments. The risk assessment system 1 includes a sequence analysis device 30 and a determination device 10, and is a system for determining the risk (disease state) of neurodegenerative diseases and precursor disorders of neurodegenerative diseases from a subject's saliva. In this disclosure, the term "disease state" refers to a state of neurodegenerative disease, a state of precursor disorder, and a state that is neither of these, i.e., a healthy state. The term "risk" can be replaced with the term "probability." For example, the expression "determining the risk of neurodegenerative diseases and precursor disorders" can include determining the likelihood of falling into either a neurodegenerative disease or a precursor disorder.
[0016] The sequence analyzer 30 receives and analyzes sequencing data of nucleic acids isolated from a subject's saliva sample to determine the microbiome composition in the saliva sample. Specifically, the sequencing data may be data obtained by amplifying the 16S rRNA gene (e.g., its V1-V2 region or V3-V4 region) by PCR and determining the base sequence of the PCR amplicon using a next-generation sequencer. The sequence analyzer 30 clusters sequencing reads that pass a quality check using a user-specified similarity threshold (e.g., 97%) to identify operational taxonomic units (OTUs). The clustering method is not limited to OTUs. For example, alternative clustering methods known to those skilled in the art include amplicon sequence variants (ASVs), which can correct or remove sequencing errors and PCR amplification errors. The sequence analyzer 30 may compare the OTU or ASV sequences with reference 16S rRNA gene sequences stored in a genome database or 16S rRNA gene sequence database to assign a biological taxonomic name (e.g., genus, species) to each OTU or ASV. As an alternative, those skilled in the art are aware that microbiome composition can be determined based on whole genome shotgun sequencing (WGS) rather than 16S rRNA gene sequencing (also known as metagenomic analysis), and this method may be used in embodiments of the present disclosure. Either method ultimately enables the sequence analyzer 30 to determine a complete picture of the types and proportions of microorganisms present in a saliva sample. It will be understood by those skilled in the art that the database may be updated as more sequence information is acquired.
[0017] In this disclosure, "microbiome composition" refers to information that comprehensively describes the microbial population makeup of each individual, representing the types and proportions of microorganisms present in the microbiome of a given sample. This is distinct from individually determining the abundance of specific, predetermined marker microbial species. The salivary microbiome may include bacteria and archaea. Analysis of the microbiome composition may be performed, for example, at the genus level, species level, or at the OTU (Operational Taxonomic Unit) or ASV (Amplicon Sequence Variant) level determined by methods known to those skilled in the art. Microbial "types" do not necessarily have to be known species, but may be operationally distinguishable types identified by sequence similarity to known species or clustering in OTU analysis, as will be understood by those skilled in the art. OTU analysis of the salivary microbiome typically involves 10 7 or more, for example, 10 8 ~10 9 OTUs may be detected. All detected microbial species may be used to build or input a predictive model, or, for example, only microbial species that account for at least x% of the microbiome composition (where x may be a number greater than or equal to 0.01, greater than or equal to 0.05, or greater than or equal to 0.1), or only microbial species whose abundance in the composition is within the top y ranks (where y may be an integer greater than or equal to 30, greater than or equal to 50, or greater than or equal to 70; for example, y may be less than or equal to 1000, less than or equal to 500, or less than or equal to 100), may be used. This can also provide data on the microbiome composition used in embodiments of the present disclosure, simplifying analysis and reducing noise in the composition that may have little biological meaning.
[0018] A saliva sample collected from a subject, i.e., a test saliva sample, is used as the subject for microbiome analysis. No special fractionation, such as isolating an exosome fraction, is required from the saliva sample; the microbiome composition contained in the entire saliva sample can be determined. The microbiome of a saliva sample can also be distinguished from the microbiome of a so-called buccal swab sample collected from the mucous membrane inside the cheek.
[0019] Optionally, the method may include identifying and generating a list of microbial species that are increased or decreased in proportion in the determined microbiome composition compared to a reference microbiome. For example, known statistical methods can be used to determine which microbial species are significantly increased or decreased in the microbiome composition of the test saliva sample compared to multiple reference microbiomes.
[0020] The determination device 10 inputs data on the composition of the salivary microbiome into multiple prediction models, and combines multiple inference results obtained from the multiple prediction models to determine the disease state of the subject. The determination device 10 may include a learning unit 11, an inference unit 12, a determination unit 13, and a memory unit 14.
[0021] The learning unit 11 uses data on the microbiome composition in each reference saliva sample collected from a group of reference subjects and the disease state of each reference subject as training data. When the saliva microbiome composition data of the subject is input, the learning unit 11 performs machine learning to generate a prediction model that stratifies the risk of a neurodegenerative disease or a precursor to a neurodegenerative disease into two groups. Stratification refers to classification or grouping. The learning unit 11 trains prediction models that can stratify not only between two groups, healthy subjects (referred to as HC, which corresponds to Healthy Control in this disclosure) and those with a neurodegenerative disease (e.g., DE, DLB, PD, and MSA), but also between two groups of neurodegenerative diseases, such as DLB and PD, DLB and MSA, or PD and MSA, as well as between two groups including precursors to neurodegenerative diseases (e.g., MCI or RBD) (e.g., RBD and HC, RBD and DLB, RBD and PD, etc.). The terms "method for diagnosing a neurodegenerative disease and its precursor disorder" and "device for diagnosing a neurodegenerative disease and its precursor disorder" are inclusive and may encompass not only embodiments in which a neurodegenerative disease and a precursor disorder are the subject of diagnosing, but also embodiments in which, for example, a neurodegenerative disease is the subject of diagnosing but a precursor disorder is not the subject of diagnosing. In some embodiments of the diagnosing method, device, or system of the present disclosure, the prediction model stratifies any combination of disease states selected from the group consisting of healthy controls (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups. For example, a prediction model corresponding to all possible combinations of these multiple disease states may be included. Preferably, at least one or more of the combinations of disease states inferred by the prediction model includes REM sleep behavior disorder (RBD). The inventors have confirmed that neurodegenerative diseases, including those mentioned herein, can be statistically significantly distinguished from HC or other disease states based on the salivary microbiome.More notably, it is believed that the ability to statistically significantly distinguish neurodegenerative disease precursor disorders such as RBD and MCI from HC or other disease states based on the salivary microbiome was previously unknown and has been confirmed for the first time by the present inventors. In one aspect, the present disclosure provides a method for inferring or determining whether a subject has RBD (i.e., whether the subject is non-RBD, e.g., healthy, or has another neurodegenerative disease or illness) based on data on the composition of the subject's salivary microbiome using the machine learning techniques described herein. It has also been confirmed that the predictive model can be used to stratify a subject's risk of a neurodegenerative disease or a neurodegenerative disease precursor disorder into three or more groups.
[0022] The inference unit 12 inputs data on the subject's salivary microbiome composition into multiple prediction models and obtains inference results from each of the prediction models. Advantageously, this embodiment does not require the input data to be re-prepared for each type of disease or disorder to be diagnosed. It is possible to implement a method for determining multiple disease states by inputting the same input data, which is a single piece of salivary microbiome composition data determined from the subject, into multiple prediction models. The inference unit 12 may input the data into all prediction models and obtain inference results from all prediction models, or it may input the data into several prediction models and, after obtaining inference results, input the data into another prediction model based on the inference results.
[0023] In addition to data on the composition of the salivary microbiome, the subject's age and gender can also be used to train and infer a predictive model. Predictive models that take age and gender into account achieve better accuracy than models that do not. In one example, the sensitivity of a dementia diagnosis model, which was previously in the 0.7-0.8 range, increased to 0.92 by adding the subject's age and gender data to the training data.
[0024] The determination unit 13 determines the disease state of the subject by combining inference results obtained from multiple prediction models. For example, if the inference results of prediction models that stratify HC and DE, and HC and RBD, are both HC, the determination unit 13 determines the subject as HC, and in other cases, the determination unit 13 may determine the disease state of the subject by combining inference results of other prediction models.
[0025] The determination unit 13 may determine the disease state of the subject using a determination model that outputs a disease state when multiple inference results are input. In this case, the learning unit 11 utilizes the training data used to train the prediction model, and trains the determination model using the multiple inference results obtained from the prediction model and the disease state of the subject as training data.
[0026] The storage unit 14 stores the parameters of the trained prediction model. The storage unit 14 may also store the parameters of the determination model.
[0027] [Prediction model] The prediction model uses training data on the microbiome composition of individual reference saliva samples collected from a group of reference subjects and the known disease states of each reference subject as training data to predict multiple stratified disease states through machine learning. For example, the model stratifies subjects into different disease state combinations, such as healthy control (HC) vs. disease affected (DE), HC vs. recurrent biliary tract disease (RBD), or DE vs. DLB. Machine learning can use classification algorithms such as random forest, LightGBM, Catboost, support vector machines, logistic regression, neural networks, or k-nearest neighbors. The prediction model may also input the subject's age and gender in addition to the microbiome composition data.
[0028] FIG. 2 shows an example of a prediction model for stratifying six disease states including HC. The top row in FIG. 2 is a prediction model for stratifying HC from other disease states. The leftmost column is a prediction model for stratifying DE from other disease states. The second to fifth columns are prediction models for stratifying RBD, DLB, PD, and MSA from other disease states, respectively. When the determination device 10 determines DE, RBD, DLB, PD, and MSA, it is desirable that the determination device 10 has a prediction model for stratifying 15 combinations including healthy (HC). The determination device 10 does not need to have prediction models for all combinations (15 in the example of FIG. 2).
[0029] Next, an example of constructing a prediction model will be described.
[0030] An example of constructing prediction models for stratifying DLB, PD, and MSA (prediction models M8, M9, and M12 in Figure 2) will be described.
[0031] Deoxyribonucleic acid (DNA) was extracted from saliva of multiple patients diagnosed with DLB, PD, or MSA, and the extracted DNA was purified using a phenol-chloroform solution. The V1-V2 region of the 16S rRNA gene was amplified by PCR using the DNA solution as a template, and the PCR amplicon was sequenced using a next-generation sequencer. Sequencing reads that passed a quality check were clustered using a 97% similarity threshold to identify OTUs. The OTU sequences were compared with reference 16S rRNA gene sequences registered in genome databases or 16S rRNA gene sequence databases, and a biological taxonomic name (genus, species, etc.) was assigned to each OTU. The bacterial species and bacterial composition were then analyzed.
[0032] From the bacterial groups (assigned to the phylum, genus, or species level) present at 0.1% or more of the total reads in each of the DLB, PD, and MSA groups, we used machine learning to determine combinations of multiple bacteria that can stratify these diseases with high accuracy from multiple abundant bacterial species (e.g., the top n species; where n can be an integer ≥ 30, ≥ 50, or ≥ 70; n can be, for example, ≤ 1000, ≤ 500, or ≤ 100) and characteristic bacterial species (bacterial species that serve as significant markers) between the DLB and PD, DLB and MSA, and PD and MSA groups. In the example below, a prediction model was constructed by performing machine learning based on the abundance of the most abundant bacteria at the species level (top 30 to 100 species).
[0033] The prediction model was validated using a cohort of patients separate from the one used to build the model. As shown in Figure 3, the Area Under the Curve (AUC) of the prediction model demonstrated excellent performance, enabling highly accurate stratification and prediction of the risk of neurodegenerative diseases, namely DLB and PD, DLB and MSA, and PD and MSA.
[0034] We also constructed and validated prediction models for stratifying RBD from DLB, PD, and MSA (prediction models M7, M11, and M14 in Figure 2). As shown in Figure 4, the AUCs of the prediction models demonstrated good performance, enabling highly accurate stratification and prediction of the risk of neurodegenerative diseases, including RBD and DLB, RBD and PD, and RBD and MSA. In another example (not shown), a prediction model for stratifying RBD from HC (prediction model M2 in Figure 2) achieved a sensitivity of 0.71, a specificity of 0.85, and an AUC of 0.80. Although not shown, stratification of MCI, another neurodegenerative disease precursor, from other disease states was also possible.
[0035] [Judgment method 1] The determination device 10 may combine inference results of multiple predictive models by using them in a stepwise manner. Using inference results in a stepwise manner means obtaining an inference result of at least one predictive model, and selecting which predictive model to further combine inferences from (and in some cases, which predictive model to no longer use inferences from) based on the inference result. Figures 5 to 7 show an example of determining a subject's risk of a disease state by combining inference results of predictive models in a stepwise manner.
[0036] In the example shown in Figure 5, both prediction model M1, which stratifies subjects into HC and DE, and prediction model M2, which stratifies subjects into HC and RBD, predict that the subject is HC. In this case, the subject is determined to be healthy.
[0037] In the example of Figure 6, prediction model M1 infers DE, and prediction model M2 infers HC. In this case, the inference result of prediction model M10, which stratifies the subject into DE (DE other than DLB) and DLB, is further combined. Since the inference result of prediction model M10 is DE, the subject is determined to have DE.
[0038] In the example of Figure 7, prediction model M1 infers DE, and prediction model M2 infers RBD. In this case, the inference result of prediction model M10, which stratifies the subject into DE and DLB, is further combined with the inference result of prediction model M7, which stratifies the subject into RBD and DLB. Since the inference results of prediction models M7 and M10 are DLB, the inference result of prediction model M8, which stratifies the subject into DLB and PD, is further combined. Since the inference result of prediction model M8 is DLB, the subject is diagnosed with DLB.
[0039] In this way, the determination device 10 determines the disease state of the subject by combining the inference results of the prediction models in a stepwise or hierarchical manner (i.e., so that the disease state inferred by the prediction models moves from a larger category (preceding prediction model) to a smaller category (subsequent prediction model)). This enables more accurate diagnosis than when using only one prediction model. Although simplified and fragmented examples are shown in Figures 5 to 7 for the purpose of explaining the principle, the robustness of the determination can be improved by increasing the number of combinations of prediction models.
[0040] [Judgment method 2] The determination device 10 may use a determination model that is machine-learned using the inference results of multiple prediction models and the disease state of the subject as training data, and that outputs a disease state when the inference results of multiple prediction models are input. This is also a form of combining multiple inference results.
[0041] For example, the determination device 10 may have six prediction models M1, M2, M7, M8, M9, and M10 shown in the examples of Figures 5 to 7, and a determination model that determines the disease state of a subject from the inference results of the six prediction models M1, M2, M7, M8, M9, and M10, and may input data on the microbiome composition into the six prediction models M1, M2, M7, M8, M9, and M10, and input the inference results of the six prediction models M1, M2, M7, M8, M9, and M10 into the determination model to determine the disease state of the subject.
[0042] Of course, the determination device 10 may have, for example, the 15 prediction models shown in Fig. 2 and determine the disease state of the subject using all of the inference results of the 15 prediction models. In other words, inference results may be obtained in a "brute force" format from multiple prediction models corresponding to all pairs of multiple disease states to be inferred.
[0043] The inference result may include not only a binary inference result representing any disease state, but also a numerical value representing the reliability of the inference result.
[0044] The specific method for combining the inference results of multiple predictive models to determine a subject's disease state is not limited to the above examples. The inference results of a predictive model can indicate whether the result is biased toward a diseased state or a non-disease state, or toward one disease or another disease, or, in certain embodiments, the degree of bias (or non-biased; whether the result is biased toward one of the two disease states can be determined based on a predetermined threshold value for the predictive model output or the inference result). Thus, typically, obtaining and determining the inference result may include determining the bias toward one of the two stratified disease states based on a threshold value. The presence or degree of bias may be scored. When the inference results of multiple predictive models are combined, for example, a "yes" (probability of occurrence) may be indicated for each of DE, DLB, and PD, and therefore the possibility of each disease occurring concomitantly may be indicated. In a more simplified example, using three prediction models that stratify subjects into two groups, DE / DLB, DE / PD, and DLB / PD, if the bias results are DE=DLB, DE>PD, and DLB>PD, respectively, the subject can be determined to have DE and DLB but not PD. Overall, the determination result supported by the most prediction models can be selected. The determination device 10 may not simply identify a single disease state of the subject, but may instead determine, for example, how closely the subject's microbiota corresponds to each disease state or how likely it is to fall into each disease state, and display this information in an n-sided radar chart (where n is a natural number corresponding to the number of disease states to be inferred). From a binary classification prediction model, the reliability or degree of bias of each class can be obtained, ranging from 0 to 1. The reliability of each disease state may be calculated based on the reliability or degree of bias obtained from each prediction model, and a radar chart including each disease state may be displayed.
[0045] In a typical embodiment, combining the multiple inference results obtained from the multiple predictive models includes statistically processing the multiple inference results. This statistical processing allows quantifying the likelihood that each disease state being assessed exists in the subject. In the simplest example, a judgment score (quantified likelihood) for each disease state can be derived by averaging the inference scores (a number between 0 and 1; 1 indicates a 100% probability of the disease state and 0 indicates a 0% probability) obtained, i.e., output, from each of the multiple predictive models. Alternatively, the judgment score may be derived by combining deviations indicating how biased the raw inference scores obtained from the predictive models are from 0.5 toward the disease state. A higher judgment score indicates a higher likelihood that the subject has the disease state. For each additional predictive model that infers that the subject is biased toward the disease state, the inference score or its average value may be weighted (e.g., multiplied by 1.2). It should be understood that a specific method for aggregating, i.e., combining, the inference results to arrive at a determination result can be appropriately designed by a person skilled in the art based on ordinary knowledge and according to the needs of each application. A person skilled in the art can appropriately combine known statistical methods for use in the determination.
[0046] Table 1 shows simplified assessment method data for the purpose of illustrating the principle. While Table 1 is a small hypothetical dataset for illustrative purposes, inference using multiple prediction models according to the methods described herein can typically output a larger set of such data. In Table 1, Prediction model 1 is a model that stratifies RBD and DLB, Prediction model 2 is a model that stratifies RBD and HC, and Prediction model 3 is a model that stratifies RBD and PD. Each specimen represents a saliva sample from a different subject. The positive probability is the output value of the machine-learned prediction model and can also be called an inference score. In Table 1, the closer the positive probability for each model is to 1, the more likely the subject is to have RBD, or the more biased the subject is toward the possibility of having RBD. In this specific example, the positive probability threshold was set to 0.5, and each model performed a simple binary classification inference of which of two disease states the subject fell into (in terms of focusing on RBD, whether or not the subject had RBD). Therefore, if the positive probability was 0.5 or greater, the model judged the test for RBD positive and assigned a test value of 1; if it was less than 0.5, the model judged the test for RBD negative and assigned a test value of 0. In this example, the positive score, or judgment score, is simply the sum of the test values of the three models. For example, the subject who provided specimen 1 received a positive score of zero, and was judged to have a low probability of having RBD after combining multiple prediction models. In contrast, the subject who provided specimen 5 received a positive score of 3, and was judged to have a high probability of having RBD after combining multiple prediction models. If the judgment score is the average of the inference scores, the judgment scores for the subjects who provided specimen 1 and specimen 5 would be 0.24 and 0.66, respectively. The inference threshold may be more stringent.For example, if the threshold is set to 0.6 to more strictly evaluate the likelihood of RBD, only the subject who provided specimen 5 will receive a positive score of 2 or higher. Table 1 focuses on RBD and displays scores so that the positive probability, which indicates a bias toward RBD, approaches 1 rather than 0. However, it is understood that similar judgment scores can be derived for other disease conditions. In any case, by combining multiple inference results, a disease condition supported (e.g., exceeding a threshold) in the inference results of multiple predictive models can be determined to be likely to exist in the subject. For example, a disease condition supported (e.g., exceeding a threshold) in the inference results of more predictive models can be determined to be more likely to exist in the subject. For example, a disease condition supported (e.g., exceeding a threshold) in the inference results of the most predictive models can be determined to be most likely to exist in the subject. The judgment scores for multiple disease conditions that were the subject of the assessment can also be displayed together in the form of a radar chart, etc.
[0047] [Table 1]
[0048] It is understood that determining can generally include deriving a determination score and / or determining which disease state the subject is likely to have. In certain embodiments, determining can include displaying the determination score and / or the determination result of which disease state the subject is likely to have on an electronic display and / or printing it on a solid medium such as paper.
[0049] In embodiments of the present disclosure, multiple disease states are accurately determined by combining the inference results of predictive models rather than using a single model. By cross-validating multiple machine learning algorithms, the accuracy of the determination is improved, enabling the provision of a robust disease state determination system. Unlike conventional biomarkers, in which the level of a single molecule typically correlates with a single disease state, the salivary microbiome-based biomarkers described in the present disclosure are complex, representing biased fluctuations in the population composition of a very large number of microorganisms and undergoing complex changes depending on the physiological state. These can be called physiological variation markers. The inventors have discovered that these physiological variations contain patterns that correlate with different disease states, and that machine learning can link the status of the physiological variation markers to different disease states. Due to this complexity and non-simplicity, unlike conventional biomarkers, it is advantageous to promote the use of physiological variation markers by performing inference analyses using multiple predictive models, as described in the present disclosure, to derive determination results whose reliability is verified by matching multiple (at least two, three, four, or more) predictive models.
[0050] When a new disease state is determined, a prediction model for stratifying the new disease state from other disease states is added. Prediction models for all combinations of the new disease state with other disease states may be added, or prediction models for some combinations from among all combinations may be added.
[0051] [System Operation] An example of a determination method of the risk determination system will be described with reference to the flowchart of FIG.
[0052] In step S11, the sequence analyzer 30 analyzes the subject's saliva sample and determines the composition of the microbiome in the saliva sample.
[0053] In step S12, the determination device 10 inputs the data on the composition of the saliva microbiome into a plurality of prediction models and obtains an inference result from each of the prediction models.
[0054] In step S13, the determination device 10 determines the disease state of the subject using the multiple inference results.
[0055] An example of a learning method for the risk assessment system will be described with reference to the flowchart of FIG.
[0056] In step S21, the determination device 10 acquires training data.
[0057] In step S22, the determination device 10 trains a prediction model using training data, which includes data on the composition of the saliva microbiome of each reference subject and the disease state of each reference subject, and may further include age and gender.
[0058] In step S23, the determination device 10 trains the determination model. For example, the determination model is trained using a plurality of inference results obtained by inputting data on the saliva microbiome composition of each reference subject into a plurality of prediction models, and the disease states of each reference subject as training data, and the determination device 10 trains the determination model so as to infer the disease state when the plurality of inference results are input.
[0059] As described above, the determination device 10 includes an inference unit 12 that inputs data on the subject's salivary microbiome composition into each of multiple prediction models and obtains inference results from each of the multiple prediction models, and a determination unit 13 that combines the multiple inference results to determine the subject's disease state. Each of the multiple prediction models is trained by machine learning to stratify the subject's disease state into two groups when the subject's salivary microbiome composition data is input, using the salivary microbiome composition data and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data. Each of the multiple prediction models infers a different combination of disease states. This allows multiple disease states to be determined from a salivary sample. Furthermore, by combining the inference results of multiple prediction models, more accurate determination results can be obtained.
[0060] The above-described determination device 10 can be, for example, a general-purpose computer system including a central processing unit (CPU) 901, a memory 902, a storage 903, a communication device 904, an input device 905, and an output device 906, as shown in Fig. 10. In this computer system, the determination device 10 is realized by the CPU 901 executing a predetermined program loaded onto the memory 902. This program can be recorded on a computer-readable non-transitory recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, or can be distributed via a network.
[0061] The present disclosure includes the following embodiments. Embodiment 1 A method for determining a neurodegenerative disease and its precursor disorder, comprising: The computer Inputting data on the composition of the subject's salivary microbiome into each of a plurality of prediction models, and obtaining inference results from each of the plurality of prediction models; combining a plurality of the inference results to determine the disease state of the subject; Each of the plurality of prediction models is trained by machine learning to stratify the disease state of the subject when data on the composition of the salivary microbiome of the subject is input using data on the composition of the salivary microbiome and a disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data; Each of the plurality of predictive models infers a different combination of disease states. Judgment method. Embodiment 2 The determination method according to embodiment 1, determining a disease state of the subject by using the plurality of inference results in a stepwise or hierarchical manner; Judgment method. Embodiment 3 The determination method according to embodiment 1, Combining the multiple inference results obtained from the multiple prediction models includes quantifying the likelihood that each disease state exists in the subject by statistically processing the multiple inference results. Judgment method. Embodiment 4 The determination method according to embodiment 1, A disease state that is supported in the inference results of a larger number of predictive models is determined to be more likely to be present in the subject. Judgment method. Embodiment 5 The determination method according to embodiment 1, determining the disease state of the subject using a trained determination model that outputs the disease state of the subject when the plurality of inference results are input; Judgment method. Embodiment 6 The determination method according to any one of embodiments 1 to 5, The predictive model stratifies any combination of disease states selected from the group consisting of healthy (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups. Judgment method. Embodiment 7 The determination method according to any one of embodiments 1 to 6, At least one or more of the combinations of inferred disease states includes REM sleep behavior disorder (RBD); Judgment method. Embodiment 8 The determination method according to any one of embodiments 1 to 7, A determination method in which the same data on the composition of the subject's salivary microbiome is input into each of the multiple prediction models. Embodiment 9 A device for determining neurodegenerative diseases and precursor disorders thereof, comprising: an inference unit that inputs data on the composition of the subject's salivary microbiome into each of a plurality of prediction models and obtains an inference result from each of the plurality of prediction models; a determination unit that determines the disease state of the subject by combining a plurality of the inference results; Each of the plurality of prediction models is trained by machine learning to stratify the disease state of the subject when data on the composition of the salivary microbiome of the subject is input using data on the composition of the salivary microbiome and a disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data; Each of the plurality of predictive models infers a different combination of disease states. Judgment device. Embodiment 10 The determination device according to embodiment 9, the determination unit determines the disease state of the subject by using the plurality of inference results in a stepwise or hierarchical manner. Judgment device. Embodiment 11 The determination device according to embodiment 9, the determination unit quantifies the degree of possibility that each disease state exists in the subject by statistically processing the plurality of inference results obtained from the plurality of prediction models. Judgment device. Embodiment 12 The determination device according to embodiment 9, The determination unit determines that a disease state supported in the inference results of a larger number of prediction models is more likely to exist in the subject. Judgment device. Embodiment 13 The determination device according to embodiment 9, the determination unit determines the disease state of the subject using a trained determination model that outputs the disease state of the subject when the plurality of inference results are input. Judgment device. Embodiment 14 The determination device according to any one of embodiments 9 to 13, The predictive model stratifies any combination of disease states selected from the group consisting of healthy (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups. Judgment device. Embodiment 15 The determination device according to any one of embodiments 9 to 14, At least one or more of the combinations of inferred disease states includes REM sleep behavior disorder (RBD); Judgment device. [Explanation of symbols]
[0062] 10 Judgment device 11 Learning Department 12 Reasoning part 13 Judgment section 14 Storage section 30 Sequence analyzer
Claims
1. A method for determining a neurodegenerative disease and its precursor disorder, comprising: The computer Inputting data on the composition of the subject's salivary microbiome into each of a plurality of prediction models, and obtaining inference results from each of the plurality of prediction models; combining a plurality of the inference results to determine the disease state of the subject; Each of the plurality of prediction models is trained by machine learning to stratify the disease state of the subject when data on the composition of the salivary microbiome of the subject is input using data on the composition of the salivary microbiome and a disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data; Each of the plurality of prediction models stratifies a different combination of disease states. Judgment method.
2. The determination method according to claim 1, determining a disease state of the subject by using the plurality of inference results in a stepwise or hierarchical manner; Judgment method.
3. The determination method according to claim 1, Combining the multiple inference results obtained from the multiple prediction models includes quantifying the likelihood that each disease state exists in the subject by statistically processing the multiple inference results. Judgment method.
4. The determination method according to claim 1, A disease state that is supported in the inference results of a larger number of predictive models is determined to be more likely to be present in the subject. Judgment method.
5. The determination method according to claim 1, determining the disease state of the subject using a trained determination model that outputs the disease state of the subject when the plurality of inference results are input; Judgment method.
6. The determination method according to claim 1, The predictive model stratifies any combination of disease states selected from the group consisting of healthy controls (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups. Judgment method.
7. The determination method according to claim 1, At least one or more of the combinations of inferred disease states includes REM sleep behavior disorder (RBD); Judgment method.
8. The determination method according to claim 1, A determination method in which the same data on the composition of the subject's salivary microbiome is input into each of the multiple prediction models.
9. A device for determining neurodegenerative diseases and precursor disorders thereof, comprising: an inference unit that inputs data on the composition of the subject's salivary microbiome into each of a plurality of prediction models and obtains an inference result from each of the plurality of prediction models; a determination unit that determines the disease state of the subject by combining a plurality of the inference results; Each of the plurality of prediction models is trained by machine learning to stratify the disease state of the subject when data on the composition of the salivary microbiome of the subject is input using data on the composition of the salivary microbiome and a disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data; Each of the plurality of prediction models stratifies a different combination of disease states. Judgment device.
10. The determination device according to claim 9, the determination unit determines the disease state of the subject by using the plurality of inference results in a stepwise or hierarchical manner. Judgment device.
11. The determination device according to claim 9, the determination unit quantifies the degree of possibility that each disease state exists in the subject by statistically processing the plurality of inference results obtained from the plurality of prediction models. Judgment device.
12. The determination device according to claim 9, The determination unit determines that a disease state supported in the inference results of a larger number of prediction models is more likely to exist in the subject. Judgment device.
13. The determination device according to claim 9, the determination unit determines the disease state of the subject using a trained determination model that outputs the disease state of the subject when the plurality of inference results are input. Judgment device.
14. The determination device according to claim 9, The predictive model stratifies any combination of disease states selected from the group consisting of healthy controls (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups. Judgment device.
15. The determination device according to claim 9, At least one or more of the combinations of inferred disease states includes REM sleep behavior disorder (RBD); Judgment device.
16. The method of claim 1, comprising: Obtaining the inference result includes determining whether the prediction model output is biased toward either of the stratified disease states based on a threshold value of the prediction model output. Judgment method.
17. The determination device according to claim 9, Obtaining the inference result includes determining whether the prediction model output is biased toward either of the stratified disease states based on a threshold value of the prediction model output. Judgment device.