Method and device for assessing neurodegenerative disease and precursor disorder

By employing machine-learning models to analyze salivary microbiome composition, the method effectively distinguishes neurodegenerative diseases and precursor disorders, enhancing diagnostic precision and reducing invasiveness.

WO2025197838A1PCT designated stage Publication Date: 2025-09-25JUNTENDO EDUCATIONAL FOUNDATION
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/010158
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-18
Filing Date
2025-03-17
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Current methods struggle to accurately and non-invasively differentiate between neurodegenerative diseases and their precursor disorders using saliva samples, particularly due to the complexity of interrelated brain abnormalities and the technical challenges of pathological diagnosis before death.

Method used

A method utilizing machine-learning prediction models that analyze the composition of salivary microbiome data to determine the risk of neurodegenerative diseases and precursor disorders by inputting data into multiple models, combining inference results to achieve accurate stratification and diagnosis.

Benefits of technology

Enables precise identification of neurodegenerative diseases and precursor disorders through saliva analysis, improving diagnostic accuracy and reducing the need for invasive procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025010158_25092025_PF_FP_ABST
    Figure JP2025010158_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention assesses the risk of a neurodegenerative disease, and a precursor disorder for the neurodegenerative disease, using a saliva sample. An assessment device 10 comprises an inference unit 12 that inputs data of the saliva microbiome composition of a subject to each of a plurality of prediction models and obtains an inference result from each of the plurality of prediction models, and an assessment unit 13 that combines the plurality of inference results to assess the disease state of the subject. Each of the plurality of prediction models is machine-learned so as to stratify the disease state of the subject when inputting the data of the saliva microbiome composition of the subject using, as teacher data, the data of the saliva microbiome composition and the disease state of a neurodegenerative disease or a precursor disorder of the neurodegenerative disease. Each of the plurality of prediction models infers mutually different combinations of disease states.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for assessing neurodegenerative diseases and precursor disorders

[0001] The present disclosure relates to a method and apparatus for diagnosing neurodegenerative diseases and precursor disorders.

[0002] Dementia (DE) refers to a condition in which cognitive function that was once normal is persistently impaired due to acquired brain damage, causing interference with daily and social life. Dementia is a neurodegenerative disease of the central nervous system, and representative types include Alzheimer's disease, dementia with Lewy bodies (DLB), vascular dementia, frontotemporal dementia, and dementia associated with Parkinson's disease (PD).

[0003] Alzheimer's disease, which causes Alzheimer's-type dementia, is a type of proteopathy in which certain proteins become structurally abnormal, condense, or accumulate, disrupting the function of tissues and organs. In Alzheimer's disease, the tau protein typically becomes structurally abnormal and accumulates in the brain, making Alzheimer's disease a tauopathy, a type of proteopathy. Another representative type of proteopathy is synucleinopathy, a neurodegenerative disease characterized by abnormal aggregation of α-synuclein protein. Synucleinopathy includes Parkinson's disease (PD) and multiple system atrophy (MSA), as well as dementia with Lewy bodies. Synucleinopathy can overlap with dementia, as exemplified by dementia with Lewy bodies and some forms of Parkinson's disease.

[0004] Mild cognitive impairment (MCI) is an intermediate state in which cognitive function cannot be said to be normal but does not (yet) meet the diagnostic criteria for dementia, and is considered to be a precursor to dementia. REM sleep behavior disorder (RBD) is a parasomnia characterized by abnormal behavior during REM sleep. REM sleep behavior disorder is known to often progress to synucleinopathies such as Parkinson's disease, dementia with Lewy bodies, and multiple system atrophy. MCI and RBD are considered precursor disorders of neurodegenerative diseases.

[0005] Alzheimer's Dementia, 2020, vol. 12, e12000PlosOne 2019 vol. 14, No. 6,e0218252

[0006] The aforementioned diseases and disorders are interrelated and often overlap, making clinical differentiation difficult. While the presence of structural abnormalities in the brain that cause disease or disorder can be confirmed by pathological diagnosis, performing pathological diagnosis, such as brain imaging, before death, especially in the early stages, is often difficult not only for technical reasons but also because of the burden it places on individual subjects. Therefore, a technology that can identify which of multiple neurodegenerative diseases and prodromal disorders a subject is likely to have, or not, using a simpler, non-invasive method would be beneficial. Furthermore, a technology that can distinguish the presence of mild cognitive impairment, a precursor to neurodegenerative diseases, from normal events that occur with aging, or that can distinguish between disease states and prodromal disorders, and that can provide these capabilities without subjecting subjects to traditional pathological or clinical diagnostic processes, would be highly beneficial.

[0007] Non-Patent Document 1 reports that a logistic regression analysis based on salivary microbiome profiling was used to construct a prediction model based on several specific individual bacterial species, making it possible to distinguish between mild cognitive impairment (MCI) and Alzheimer's disease. Non-Patent Document 2 reports that a logistic regression analysis method based on microbiome genus and species data obtained by salivary RNA sequencing was used to distinguish between early-stage Parkinson's disease (PD) patients and healthy individuals using data on 11 bacterial species.

[0008] However, there has been no attempt to assess the risk of precursors to neurodegenerative diseases based on the composition pattern of salivary microbiota.

[0009] It is also desirable to be able to more accurately assess the risk of neurodegenerative diseases and precursor disorders of neurodegenerative diseases using saliva samples.

[0010] The present disclosure has been made in light of the above, and aims to determine the risk of neurodegenerative diseases and precursor disorders of neurodegenerative diseases using saliva samples.

[0011] One aspect of the present disclosure is a method for determining neurodegenerative diseases and their precursor disorders, in which a computer inputs data on the composition of a subject's salivary microbiome into each of a plurality of prediction models, obtains inference results from each of the plurality of prediction models, and determines the disease state of the subject by combining the inference results, and each of the plurality of prediction models is machine-learned to stratify the disease state of the subject into two groups when data on the composition of the subject's salivary microbiome is input using data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data, and each of the plurality of prediction models infers a different combination of disease states.

[0012] A determination device of one embodiment of the present disclosure is a determination device for neurodegenerative diseases and their precursor disorders, and includes an inference unit that inputs data on the composition of a subject's salivary microbiome into each of a plurality of prediction models and obtains an inference result from each of the plurality of prediction models, and a determination unit that combines the inference results to determine the disease state of the subject, wherein each of the plurality of prediction models is machine-learned to stratify the disease state of the subject into two groups when data on the composition of the subject's salivary microbiome is input using data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data, and each of the plurality of prediction models infers a different combination of disease states.

[0013] According to the present disclosure, saliva samples can be used to determine risk for neurodegenerative diseases and precursors to neurodegenerative diseases.

[0014] FIG. 1 is a diagram showing an example of the configuration of a risk assessment system of this embodiment. FIG. 2 shows an example of a prediction model for stratifying six disease states including HC. FIG. 3 is a diagram showing an example of the performance of the prediction model. FIG. 4 is a diagram showing an example of the performance of the prediction model. FIG. 5 is a diagram showing an example of determining the risk of a disease state of a subject by combining inference results of the prediction models. FIG. 6 is a diagram showing an example of determining the risk of a disease state of a subject by combining inference results of the prediction models. FIG. 7 is a diagram showing an example of determining the risk of a disease state of a subject by combining inference results of the prediction models. FIG. 8 is a flowchart showing an example of a determination method. FIG. 9 is a flowchart showing an example of a learning method. FIG. 10 is a diagram showing an example of the hardware configuration of a determination device.

[0015] [System Configuration] An example of the configuration of the risk assessment system of this embodiment will be described with reference to FIG. 1. The systems, devices, and methods described herein are related to one another; for example, an embodiment of an device may be suitable for implementing an embodiment of a method. Therefore, it will be understood by those skilled in the art that the description provided herein can be applied equally and reciprocally to different embodiments. The risk assessment system 1 of this embodiment is a system that includes a sequence analysis device 30 and a determination device 10 and determines the risk (disease state) of neurodegenerative diseases and precursor disorders of neurodegenerative diseases from the saliva of a subject. In this disclosure, the term disease state refers to a state of neurodegenerative disease, a state of precursor disorder, and a state that is neither of these, i.e., a healthy state. The term risk can be replaced with the word likelihood. That is, for example, the expression "determining the risk of neurodegenerative diseases and precursor disorders" can include determining the likelihood of falling into either a neurodegenerative disease or a precursor disorder.

[0016] The sequence analysis device 30 receives and analyzes sequencing data of nucleic acids isolated from a subject's saliva sample to determine the microbiome composition in the saliva sample. Specifically, the sequencing data may be data obtained by amplifying the 16S rRNA gene (e.g., its V1-V2 region or V3-V4 region) by PCR and determining the base sequence of the PCR amplicon using a next-generation sequencer. The sequence analysis device 30 clusters the sequencing reads that pass a quality check using a user-specified similarity threshold (e.g., 97%) to identify operational taxonomic units (OTUs). The clustering method is not limited to OTUs. For example, alternative clustering methods known to those skilled in the art include ASVs (amplicon sequence variants), which can correct or remove sequencing errors and PCR amplification errors. The sequence analysis device 30 may compare the OTU or ASV sequences with reference 16S rRNA gene sequences stored in a genome database or 16S rRNA gene sequence database to assign a biological taxonomic name (e.g., genus, species) to each OTU or ASV. As a further alternative, those skilled in the art are aware that microbiome composition can be determined based on whole genome shotgun sequencing (WGS) rather than 16S rRNA gene sequencing (also known as metagenomic analysis), and this method may be used in embodiments of the present disclosure. Either method ultimately allows the sequence analysis device 30 to determine a complete picture of the types and proportions of microorganisms present in a saliva sample. It will be understood by those skilled in the art that the database may be updated as more sequence information is acquired.

[0017] In this disclosure, "microbiome composition" refers to information that comprehensively describes the population makeup of each individual's microorganisms, representing the types and proportions of microorganisms present in the microbiome of a given sample. This is distinct from individually determining the abundance of specific, predetermined marker microbial species. The salivary microbiome may include bacteria and archaea. Analysis of the microbiome composition may be performed, for example, at the genus level, species level, or at the level of operational taxonomic units (OTUs) or amplicon sequence variants (ASVs) determined by methods known to those skilled in the art. Microbial "types" do not necessarily have to be known species, but may be operationally distinguishable types identified by sequence similarity to known species or clustering in OTU analysis, as will be understood by those skilled in the art. OTU analysis of the salivary microbiome typically involves 10 7 More than, for example, 10 8 ~10 9 OTUs of up to 1000 OTUs may be detected. All detected microbial species may be used to build or input a predictive model, or, for example, only microbial species that account for at least x% of the microbiome composition (where x may be a number greater than or equal to 0.01, greater than or equal to 0.05, or greater than or equal to 0.1), or only microbial species whose abundance in the composition is within the top y ranks (where y may be an integer greater than or equal to 30, greater than or equal to 50, or greater than or equal to 70; for example, y may be less than or equal to 1000, less than or equal to 500, or less than or equal to 100), may be used. This can also provide data on the microbiome composition used in embodiments of the present disclosure, simplifying analysis and reducing noise in the composition that may have little biological meaning.

[0018] A saliva sample collected from a subject, i.e., a test saliva sample, is used as the subject for microbiome analysis. No special fractionation, such as isolating an exosome fraction, is required from the saliva sample; the microbiome composition contained in the entire saliva sample can be determined. The microbiome of a saliva sample can also be distinguished from the microbiome of a so-called buccal swab sample collected from the mucous membrane inside the cheek.

[0019] Optionally, the method may include identifying and generating a list of microbial species that are increased or decreased in proportion in the determined microbiome composition compared to a reference microbiome. For example, known statistical methods can be used to determine which microbial species are significantly increased or decreased in the microbiome composition of the test saliva sample compared to multiple reference microbiomes.

[0020] The determination device 10 inputs data on the composition of the salivary microbiome into multiple prediction models, and determines the disease state of the subject by combining multiple inference results obtained from the multiple prediction models. The determination device 10 may include a learning unit 11, an inference unit 12, a determination unit 13, and a memory unit 14.

[0021] The learning unit 11 utilizes data on the microbiome composition in each reference saliva sample collected from a group of reference subjects and the disease state of each reference subject as training data, and when the data on the saliva microbiome composition of the subject is input, performs machine learning to generate a prediction model that stratifies the subject's risk of neurodegenerative disease or precursor disorder of neurodegenerative disease into two groups. Stratification means classification or grouping. The learning unit 11 trains not only prediction models capable of stratifying between two groups, healthy subjects (in the present disclosure, abbreviated as HC, which corresponds to Healthy Control) and those with a neurodegenerative disease (e.g., DE, DLB, PD, and MSA), but also prediction models capable of stratifying between two groups of neurodegenerative diseases, such as DLB and PD, DLB and MSA, or PD and MSA, as well as prediction models capable of stratifying between two groups including precursor disorders of neurodegenerative diseases (MCI or RBD) (e.g., RBD and HC, RBD and DLB, RBD and PD, etc.). The terms "method for determining a neurodegenerative disease and its precursor disorder" and "device for determining a neurodegenerative disease and its precursor disorder" are inclusive and may include, in addition to embodiments in which a neurodegenerative disease and a precursor disorder are the subject of diagnosis, embodiments in which, for example, a neurodegenerative disease is the subject of diagnosis but a precursor disorder is not the subject of diagnosis. In some embodiments of the assessment method, assessment device, or assessment system of the present disclosure, the prediction model stratifies any combination of disease states selected from the group consisting of healthy controls (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups. For example, prediction models corresponding to all possible combinations of these multiple disease states may be included. Preferably, at least one or more of the combinations of disease states inferred by the prediction model includes REM sleep behavior disorder (RBD). The inventors have confirmed that neurodegenerative diseases, including those mentioned herein, can be statistically significantly distinguished from HC or other disease states based on the salivary microbiome.More notably, it is believed that the ability to statistically significantly distinguish neurodegenerative disease precursor disorders such as RBD and MCI from HC or other disease states based on the salivary microbiome was previously unknown and has been confirmed for the first time by the present inventors. In one aspect, the present disclosure provides a method for inferring or determining whether a subject has RBD (i.e., whether the subject is non-RBD, e.g., healthy, or has another neurodegenerative disease or illness) based on data on the composition of the subject's salivary microbiome using the machine learning techniques described herein. It has also been confirmed that the predictive model can be used to stratify a subject's risk of a neurodegenerative disease or a neurodegenerative disease precursor disorder into three or more groups.

[0022] The inference unit 12 inputs data on the subject's salivary microbiome composition into multiple prediction models and obtains inference results from each of the prediction models. Advantageously, in this embodiment, it is not necessary to re-prepare input data for each type of disease or disorder to be diagnosed. This embodiment can implement a method for determining multiple disease states by inputting the same input data, which is a single piece of salivary microbiome composition data determined from the subject, into multiple prediction models. The inference unit 12 may input the data into all prediction models and obtain inference results from all prediction models, or it may input the data into several prediction models, obtain inference results, and then input the data into another prediction model based on the inference results.

[0023] In addition to data on the composition of the salivary microbiome, the subject's age and gender can also be used to train and infer a predictive model. Predictive models that take age and gender into account achieve better accuracy than predictive models that do not. In one example, the sensitivity of a dementia diagnosis model, which had remained in the 0.7-0.8 range, increased to 0.92 by adding data on the subject's age and gender to the training data.

[0024] The determination unit 13 determines the disease state of the subject by combining inference results obtained from multiple prediction models. For example, if the inference results of prediction models that stratify HC and DE, and HC and RBD, are both HC, the determination unit 13 may determine the subject's disease state as HC, and otherwise may determine the disease state of the subject by combining inference results of other prediction models.

[0025] The determination unit 13 may determine the disease state of the subject using a determination model that outputs a disease state when multiple inference results are input. In this case, the learning unit 11 utilizes the training data used to train the prediction model, and trains the determination model using the multiple inference results obtained from the prediction model and the disease state of the subject as training data.

[0026] The storage unit 14 stores the parameters of the trained prediction model. The storage unit 14 may also store the parameters of the determination model.

[0027] [Prediction Model] The prediction model uses data on the microbiome composition in individual reference saliva samples collected from a group of reference subjects and known disease states of each reference subject as training data, and predicts multiple stratified disease states through machine learning. For example, the model stratifies subjects into different disease state combinations, such as healthy control (HC) and disease endemic (DE), HC and recurrent biliary tract disease (RBD), and DE and DLB. Machine learning can use algorithms used for classification, such as random forests, LightGBM, Catboost, support vector machines, logistic regression, neural networks, or k-nearest neighbors. The prediction model may also input the subject's age and gender in addition to the microbiome composition data.

[0028] FIG. 2 shows an example of a prediction model for stratifying six disease states, including HC. The top row in FIG. 2 is a prediction model for stratifying HC from other disease states. The leftmost column is a prediction model for stratifying DE from other disease states. The second to fifth columns are prediction models for stratifying RBD, DLB, PD, and MSA from other disease states, respectively. When the determination device 10 determines DE, RBD, DLB, PD, and MSA, it is desirable that the determination device 10 have a prediction model for stratifying 15 combinations, including healthy (HC). The determination device 10 does not need to have prediction models for all combinations (15 in the example of FIG. 2).

[0029] Next, an example of constructing a prediction model will be described.

[0030] An example of constructing prediction models for stratifying DLB, PD, and MSA (prediction models M8, M9, and M12 in FIG. 2) will be described.

[0031] Deoxyribonucleic acid (DNA) was extracted from the saliva of multiple patients diagnosed with DLB, PD, or MSA, and the extracted DNA was purified using a phenol-chloroform solution. Using the DNA solution as a template, the V1-V2 region of the 16S rRNA gene was amplified by PCR, and the PCR amplicon was subjected to sequence analysis using a next-generation sequencer. Sequencing reads that passed the quality check were clustered using a 97% similarity threshold to identify OTUs. The OTU sequences were compared with reference 16S rRNA gene sequences registered in genome databases or 16S rRNA gene sequence databases, and a biological classification name (genus, species, etc.) was assigned to each OTU. The bacterial species and bacterial composition were then analyzed.

[0032] From the bacterial groups (assigned to the phylum, genus, or species level) present at 0.1% or more of the total reads in each of the DLB, PD, and MSA groups, we used machine learning to determine combinations of multiple bacteria that can stratify these diseases with high accuracy from multiple abundant bacterial species (e.g., the top n species; here, n can be an integer of 30 or more, 50 or more, or 70 or more; n can be, for example, 1000 or less, 500 or less, or 100 or less) and characteristic bacterial species (bacterial species that serve as markers with significant differences) between the DLB and PD, DLB and MSA, and PD and MSA groups. In the example below, a prediction model was constructed by performing machine learning based on the abundance of bacteria (top 30 to 100 species) that are abundant at the species level.

[0033] The prediction model was validated using a separate cohort of patients from the one used to build the model. As shown in Figure 3, the Area Under the Curve (AUC) of the prediction model demonstrated excellent performance, enabling highly accurate stratification and prediction of the risk of neurodegenerative diseases between DLB and PD, DLB and MSA, and PD and MSA.

[0034] Prediction models for stratifying RBD from DLB, PD, and MSA (prediction models M7, M11, and M14 in Figure 2) were also constructed and validated. As shown in Figure 4, the AUCs of the prediction models demonstrated good performance, enabling highly accurate stratification and prediction of neurodegenerative disease risk between RBD and DLB, RBD and PD, and RBD and MSA. In another example (not shown), a prediction model for stratifying RBD from HC (prediction model M2 in Figure 2) achieved a sensitivity of 0.71, a specificity of 0.85, and an AUC of 0.80. Although not shown, stratification between MCI, another neurodegenerative disease precursor, and other disease states was also possible.

[0035] [Determination Method 1] The determination device 10 may combine inference results of multiple prediction models by using them in a stepwise manner. Using inference results in a stepwise manner means obtaining an inference result from at least one prediction model, and then selecting which prediction model to further combine inferences from (and, in some cases, which prediction model to no longer use inferences from) based on the inference result. Figures 5 to 7 show an example of determining a subject's risk of a disease state by combining the inference results of prediction models in a stepwise manner.

[0036] In the example of Fig. 5, both prediction model M1, which stratifies subjects into HC and DE, and prediction model M2, which stratifies subjects into HC and RBD, infer that the subject is HC. In this case, the subject is determined to be healthy.

[0037] In the example of Figure 6, prediction model M1 infers DE, and prediction model M2 infers HC. In this case, the inference results of prediction model M10, which stratifies the subject into DE (DE other than DLB) and DLB, are further combined. Since the inference result of prediction model M10 is DE, the subject is determined to have DE.

[0038] In the example of Figure 7, prediction model M1 infers DE, and prediction model M2 infers RBD. In this case, the inference result of prediction model M10, which stratifies the subject into DE and DLB, is further combined with the inference result of prediction model M7, which stratifies the subject into RBD and DLB. Since the inference results of prediction models M7 and M10 are DLB, the inference result of prediction model M8, which stratifies the subject into DLB and PD, is further combined. Since the inference result of prediction model M8 is DLB, the subject is determined to have DLB.

[0039] In this way, the determination device 10 determines the disease state of the subject by combining the inference results of the prediction models in a stepwise or hierarchical manner (i.e., so that the disease state inferred by the prediction models moves from a larger category (leading prediction model) to a smaller category (subsequent prediction model)). This enables more accurate diagnosis than when using only one prediction model. Although simplified and fragmented examples are shown in Figures 5 to 7 to explain the principle, the robustness of the determination can be improved by increasing the number of combinations of prediction models.

[0040] [Determination Method 2] The determination device 10 may use a determination model that is machine-learned using the inference results of multiple prediction models and the disease state of the subject as training data, and that outputs a disease state when the inference results of multiple prediction models are input. This is also a form of combining multiple inference results.

[0041] For example, the determination device 10 may have six prediction models M1, M2, M7, M8, M9, and M10 shown in the examples of Figures 5 to 7 and a determination model that determines the disease state of a subject from the inference results of the six prediction models M1, M2, M7, M8, M9, and M10, and may input data on the microbiome composition into the six prediction models M1, M2, M7, M8, M9, and M10 and input the inference results of the six prediction models M1, M2, M7, M8, M9, and M10 into the determination model to determine the disease state of the subject.

[0042] Of course, the determination device 10 may have, for example, the 15 prediction models shown in Fig. 2 and determine the disease state of the subject using all of the inference results of the 15 prediction models. In other words, inference results may be obtained in a "brute force" format from multiple prediction models corresponding to all pairs of multiple disease states to be inferred.

[0043] The inference result may include not only a binary inference result representing any disease state, but also a numerical value representing the reliability of the inference result.

[0044] The specific method for combining the inference results of multiple predictive models to determine a subject's disease state is not limited to the above examples. The inference result of a predictive model can indicate whether the result is biased toward a diseased state or a non-disease state, or toward one disease or another disease state, or, in certain embodiments, the degree of bias (or non-biased; whether the result is biased toward either state can be determined based on a predetermined threshold value for the predictive model output or the inference result). Thus, typically, obtaining and determining the inference result may include determining the bias toward either of the two stratified disease states based on a threshold value. The presence or degree of bias may be scored. When the inference results of multiple predictive models are combined, for example, a "Yes" (probability of occurrence) may be indicated for each of DE, DLB, and PD, and therefore the possibility of each disease occurring concomitantly may be indicated. In a more simplified example, using three prediction models that stratify the subject into two groups, DE / DLB, DE / PD, and DLB / PD, and the resulting biases are DE = DLB, DE > PD, and DLB > PD, respectively, the subject can be determined to have DE and DLB but not PD. Overall, the determination result supported by the most prediction models can be selected. Instead of identifying a single disease state of the subject, the determination device 10 may, for example, determine how close the subject's microbiota state is to each disease state or how likely it is to fall into each disease state, and display this information in an n-sided radar chart (where n is a natural number corresponding to the number of disease states to be inferred). From a binary classification prediction model, the reliability or degree of bias of each class can be obtained, ranging from 0 to 1. The reliability of each disease state can be calculated based on the reliability or degree of bias obtained from each prediction model, and a radar chart including each disease state may also be displayed.

[0045] In a typical embodiment, combining the multiple inference results obtained from the multiple predictive models includes statistically processing the multiple inference results. This statistical processing allows quantifying the likelihood that each disease state being assessed exists in the subject. In the simplest example, a judgment score (quantified likelihood) for each disease state can be derived by averaging the inference scores (a numerical value between 0 and 1; 1 indicates a 100% probability of the disease state and 0 indicates a 0% probability) obtained, i.e., output, from each of the multiple predictive models. Alternatively, the judgment score may be derived by combining deviations indicating how biased the raw inference scores obtained from the predictive models are from 0.5 toward the disease state. A higher judgment score indicates a higher likelihood that the subject has the disease state. For each additional predictive model that infers that the subject is biased toward the disease state, the inference score or its average value may be weighted (e.g., multiplied by 1.2). It should be understood that a specific method for aggregating, i.e., combining, the inference results to arrive at a determination result can be appropriately designed by a person skilled in the art based on ordinary knowledge and according to the needs of each application. A person skilled in the art can appropriately combine known statistical methods for use in the determination.

[0046] Table 1 shows simplified assessment method data for the purpose of illustrating the principle. While Table 1 is a small hypothetical dataset for illustrative purposes, inference using multiple prediction models according to the methods described herein can typically output a larger set of such data. In Table 1, Prediction model 1 is a model that stratifies RBD and DLB, Prediction model 2 is a model that stratifies RBD and HC, and Prediction model 3 is a model that stratifies RBD and PD. Each specimen represents a saliva sample from a different subject. The positive probability is the output value of the machine-learned prediction model and can also be called an inference score. In Table 1, the closer the positive probability for each model is to 1, the more likely the subject is to have RBD, or the more biased the subject is toward the possibility of having RBD. In this specific example, the positive probability threshold was set to 0.5, and each model performed a simple binary classification inference of which of two disease states the subject fell into (in terms of focusing on RBD, whether or not the subject fell into RBD). Therefore, if the positive probability was 0.5 or greater, the model determined that the test for RBD was positive and assigned a test value of 1; if it was less than 0.5, the model determined that the test for RBD was negative and assigned a test value of 0. In this example, the positive score, or judgment score, is simply the sum of the test values ​​of the three models. For example, the subject who provided specimen 1 had a positive score of zero, and after combining multiple prediction models, it was determined that the subject had a low probability of having RBD. In contrast, the subject who provided specimen 5 had a positive score of 3, and after combining multiple prediction models, it was determined that the subject had a high probability of having RBD. If the average of the inference scores were used as the judgment score, the judgment scores for the subjects who provided specimen 1 and specimen 5 would be 0.24 and 0.66, respectively. The inference threshold may be stricter.For example, if the threshold is set to 0.6 to more strictly evaluate the likelihood of RBD, only the subject who provided specimen 5 will receive a positive score of 2 or higher. Table 1 focuses on RBD and displays scores so that the positive probability, which indicates a bias toward RBD, approaches 1 rather than 0. However, it is understood that similar judgment scores can be derived for other disease conditions. In any case, by combining multiple inference results, a disease condition supported by the inference results of multiple predictive models (e.g., exceeding a threshold) can be determined to be likely to exist in the subject. For example, a disease condition supported by the inference results of more predictive models (e.g., exceeding a threshold) can be determined to be more likely to exist in the subject. For example, a disease condition supported by the inference results of the most predictive models (e.g., exceeding a threshold) can be determined to be most likely to exist in the subject. The judgment scores for multiple disease conditions to be assessed can also be displayed together in the form of a radar chart, etc.

[0047]

[0048] It is understood that determining may generally include deriving a judgment score and / or determining which disease state the subject is likely to have. In certain embodiments, determining may include displaying the judgment score and / or the determination result of which disease state the subject is likely to have on an electronic display and / or printing it on a solid medium such as paper.

[0049] In embodiments of the present disclosure, multiple disease states are accurately determined by combining the inference results of predictive models rather than using a single model. Mutual validation of multiple machine learning algorithms improves the accuracy of determination, providing a robust disease state determination system. Unlike conventional biomarkers, in which the level of a single molecule typically correlates with a single disease state, the salivary microbiome-based biomarkers described in the present disclosure are complex, representing biased fluctuations in the population composition of a very large number of microorganisms, and undergo complex changes depending on the physiological state. These can be called physiological variation markers. The inventors have discovered that these physiological variations contain patterns that correlate with different disease states, and that machine learning can link the status of the physiological variation marker to different disease states. Due to this complexity and non-simplicity, unlike conventional biomarkers, it is advantageous to promote the use of physiological variation markers by performing inference analyses using multiple predictive models, as described in the present disclosure, to derive determination results whose reliability is verified by matching multiple (at least two, three, four, or more) predictive models.

[0050] When a new disease state is determined, a prediction model for stratifying the new disease state from other disease states is added. Prediction models for all combinations of the new disease state with other disease states may be added, or prediction models for some combinations from among all combinations may be added.

[0051] [System Operation] An example of a determination method of the risk determination system will be described with reference to the flowchart of FIG.

[0052] In step S11, the sequence analyzer 30 analyzes the subject's saliva sample and determines the composition of the microbiome in the saliva sample.

[0053] In step S12, the determination device 10 inputs the data on the composition of the salivary microbiome into multiple prediction models and obtains inference results from each of the prediction models.

[0054] In step S13, the determination device 10 determines the disease state of the subject using the multiple inference results.

[0055] An example of a learning method for the risk assessment system will be described with reference to the flowchart of FIG.

[0056] In step S21, the determination device 10 acquires teacher data.

[0057] In step S22, the determination device 10 trains a prediction model using training data, which includes data on the composition of the salivary microbiome of each reference subject and the disease state of each reference subject, and may also include age and gender.

[0058] In step S23, the determination device 10 trains the determination model. For example, the determination model is trained using a plurality of inference results obtained by inputting data on the salivary microbiome composition of each reference subject into a plurality of prediction models, and the disease states of each reference subject as training data, and the determination device 10 trains the determination model so as to infer the disease state when the plurality of inference results are input.

[0059] As described above, the determination device 10 includes an inference unit 12 that inputs data on the composition of a subject's salivary microbiome into each of multiple prediction models and obtains an inference result from each of the multiple prediction models, and a determination unit 13 that combines the multiple inference results to determine the subject's disease state. Each of the multiple prediction models is trained by machine learning to stratify the subject's disease state into two groups when the subject's salivary microbiome composition data and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease are input, using the data on the salivary microbiome composition and the disease state of a neurodegenerative disease as training data. Each of the multiple prediction models infers a different combination of disease states. This allows multiple disease states to be determined from a salivary sample. Furthermore, by combining the inference results of the multiple prediction models, more accurate determination results can be obtained.

[0060] The above-described determination device 10 may be, for example, a general-purpose computer system including a central processing unit (CPU) 901, a memory 902, a storage 903, a communication device 904, an input device 905, and an output device 906, as shown in Fig. 10. In this computer system, the determination device 10 is realized by the CPU 901 executing a predetermined program loaded onto the memory 902. This program can be recorded on a computer-readable non-transitory recording medium such as a magnetic disk, an optical disk, or a semiconductor memory, or can be distributed via a network.

[0061] The present disclosure includes the following embodiments. Embodiment 1: A method for diagnosing a neurodegenerative disease and its precursor disorder, comprising: a computer inputting data on the composition of a subject's salivary microbiome into each of a plurality of prediction models, obtaining an inference result from each of the plurality of prediction models, and determining a disease state of the subject by combining the plurality of inference results; each of the plurality of prediction models is machine-learned to stratify the disease state of the subject when the data on the composition of the subject's salivary microbiome is input using the data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data; and each of the plurality of prediction models infers a different combination of disease states from each other. Embodiment 2: A method for diagnosing a neurodegenerative disease and its precursor disorder using the plurality of inference results in a stepwise or hierarchical manner. Embodiment 3: A method for diagnosing a neurodegenerative disease and its precursor disorder using the plurality of inference results in the combination of the plurality of inference models. Embodiment 4 The determination method according to embodiment 1, wherein a disease state supported in the inference results of a larger number of prediction models is determined to be more likely to be present in the subject.Embodiment 5 The determination method according to embodiment 1, wherein the disease state of the subject is determined using a trained determination model that receives the plurality of inference results as input and outputs the disease state of the subject.Embodiment 6 The determination method according to any of embodiments 1 to 5, wherein the prediction model stratifies any combination of disease states selected from the group consisting of healthy controls (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups.Embodiment 7 The determination method according to any of embodiments 1 to 6, wherein at least one or more of the combinations of inferred disease states includes REM sleep behavior disorder (RBD).Embodiment 8: The determination method according to any one of Embodiments 1 to 7, wherein the same data on the composition of the subject's salivary microbiome is input to each of the plurality of prediction models.Embodiment 9: A determination device for neurodegenerative diseases and precursor disorders thereof, comprising: an inference unit that inputs data on the composition of the subject's salivary microbiome to each of a plurality of prediction models and obtains an inference result from each of the plurality of prediction models; and a determination unit that combines the plurality of inference results to determine the disease state of the subject, wherein each of the plurality of prediction models is machine-learned to stratify the disease state of the subject when the data on the composition of the subject's salivary microbiome is input using data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data, and each of the plurality of prediction models infers a disease state that is a different combination from each other.Embodiment 10: The determination device according to Embodiment 9, wherein the determination unit determines the disease state of the subject by using the plurality of inference results in a stepwise or hierarchical manner. Embodiment 11: The determination device according to embodiment 9, wherein the determination unit quantifies the magnitude of the possibility that each disease state exists in the subject by statistically processing the inference results obtained from the plurality of prediction models when combining the inference results. Embodiment 12: The determination device according to embodiment 9, wherein the determination unit determines that a disease state supported in the inference results of more prediction models is more likely to exist in the subject. Embodiment 13: The determination device according to embodiment 9, wherein the determination unit determines the disease state of the subject using a trained determination model that outputs the disease state of the subject when the plurality of inference results are input. Embodiment 14: The determination device according to any one of embodiments 9 to 13, wherein the prediction model stratifies any combination of disease states selected from the group consisting of healthy (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups.Embodiment 15 The determination device according to any one of embodiments 9 to 14, wherein at least one or more of the combinations of disease states to be inferred includes REM sleep behavior disorder (RBD).

[0062] 10 Determination device 11 Learning unit 12 Inference unit 13 Determination unit 14 Storage unit 30 Sequence analysis device

Claims

1. A method for diagnosing neurodegenerative diseases and their precursor disorders, comprising: a computer inputting data on the composition of a subject's salivary microbiome into each of a plurality of prediction models, obtaining an inference result from each of the plurality of prediction models, and determining the disease state of the subject by combining the inference results; each of the plurality of prediction models is machine-learned to stratify the disease state of the subject when data on the composition of the subject's salivary microbiome is input, using data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data; and each of the plurality of prediction models infers a different combination of disease states.

2. A method for determining the disease state of the subject according to claim 1, wherein the plurality of inference results are used in a stepwise or hierarchical manner.

3. A method of determination as described in claim 1, wherein combining multiple inference results obtained from the multiple predictive models includes quantifying the likelihood that each disease state exists in the subject by statistically processing the multiple inference results.

4. A method of determination according to claim 1, wherein a disease state that is supported in the inference results of a greater number of predictive models is determined to be more likely to exist in the subject.

5. A method of determining the disease state of the subject according to claim 1, comprising: determining the disease state of the subject using a trained determination model that outputs the disease state of the subject when the multiple inference results are input.

6. The method of determination according to claim 1, wherein the prediction model stratifies any combination of disease states selected from the group consisting of healthy controls (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups.

7. The method of claim 1, wherein at least one or more of the combinations of disease states to be inferred includes REM sleep behavior disorder (RBD).

8. A determination method according to claim 1, wherein the same data on the composition of the subject's salivary microbiome is input into each of the multiple prediction models.

9. A device for diagnosing neurodegenerative diseases and their precursor disorders, comprising: an inference unit that inputs data on the composition of a subject's salivary microbiome into each of a plurality of prediction models and obtains an inference result from each of the plurality of prediction models; and a judgment unit that combines the inference results to judge the disease state of the subject, wherein each of the plurality of prediction models is machine-learned to stratify the disease state of the subject when data on the composition of the subject's salivary microbiome is input using data on the composition of the salivary microbiome and the disease state of a neurodegenerative disease or a precursor disorder of a neurodegenerative disease as training data, and each of the plurality of prediction models infers a different combination of disease states.

10. A determination device according to claim 9, wherein the determination unit determines the disease state of the subject by using the plurality of inference results in a stepwise or hierarchical manner.

11. A determination device as described in claim 9, wherein the determination unit quantifies the likelihood that each disease state exists in the subject by combining multiple inference results obtained from the multiple prediction models and performing statistical processing on the multiple inference results.

12. A determination device according to claim 9, wherein the determination unit determines that a disease state that is supported in the inference results of a greater number of prediction models is more likely to exist in the subject.

13. A determination device according to claim 9, wherein the determination unit determines the disease state of the subject using a trained determination model that outputs the disease state of the subject when the multiple inference results are input.

14. A determination device according to claim 9, wherein the prediction model stratifies any combination of disease states selected from the group consisting of healthy controls (HC), REM sleep behavior disorder (RBD), dementia (DE), dementia with Lewy bodies (DLB), Parkinson's disease (PD), and multiple system atrophy (MSA) into two groups.

15. The determination device according to claim 9, wherein at least one or more of the combinations of disease states to be inferred includes REM sleep behavior disorder (RBD).

Citation Information

Patent Citations

  • Prediction method, prediction device, prediction program, and recording medium

    JP2006221310A

  • Prediction device and prediction method

    JP2019045264A

  • Disease-associated microbiome characterization process

    JP2020530931A

  • Machine learning device, trained model, data structure, periodontal disease testing method, periodontal disease testing system, and periodontal disease testing kit

    JP2022047540A

  • System for predicting clinical outcome of disease, program, and method

    JP2023108532A