Systems and methods for dynamic immunohistochemical profiling of biological disorders and their feature engineering
A non-invasive method for diagnosing biological states in infants and young children using tooth samples stained with C-reactive protein immunohistochemistry and fluorescence analysis, combined with advanced imaging and spectroscopy, achieves high diagnostic accuracy for conditions like autism and ADHD.
Patent Information
- Application Number
- JP2024570573
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-08
- Filing Date
- 2023-06-07
- Publication Date
- 2025-07-15
AI Technical Summary
There is a need for accurate, non-invasive methods and systems for diagnosing biological states, particularly in infants and young children, by profiling biomarkers in human biological samples such as hair shafts, teeth, and nails, to detect conditions like autism spectrum disorder, ADHD, ALS, schizophrenia, and cancer.
A method involving staining a tooth sample, spatially analyzing fluorescence intensity, and using a trained model to predict a diagnostic state based on fluorescence intensity, combined with techniques like C-reactive protein immunohistochemical staining and confocal microscopy, and potentially supplemented with LA-ICP-MS and Raman spectroscopy for enhanced diagnostic accuracy.
The method achieves high sensitivity and specificity in predicting diagnostic states, with models demonstrating at least 70-90% accuracy in diagnosing conditions like autism spectrum disorder and ADHD, utilizing features like recurrence quantification analysis to process fluorescence intensity.
Smart Images

Figure 2025522323000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 350,089, filed on June 8, 2022, the entire contents of which are incorporated herein by reference.
Background Art
[0002] Dynamic biological reactions can represent fundamental biological processes that are structurally and functionally important for humans. For example, abnormal dynamic biological reactions can be associated with many biological conditions such as diseases and disorders. Examples of such biological conditions include neurological conditions (e.g., autism spectrum disorder, schizophrenia, or attention - deficit / hyperactivity disorder (ADHD)), neurodegenerative conditions (e.g., amyotrophic lateral sclerosis (ALS), Alzheimer's disease, Parkinson's disease, and Huntington's disease), and cancer (e.g., pediatric cancer).
Summary of the Invention
[0003] Considering the foregoing background, there is a need for accurate methods and systems for diagnosing biological states, particularly non-invasive diagnosis. Such diagnosis can be based on the accurate profiling of biomarkers detectable by non-invasive methods for diagnosing biological states. The present disclosure provides improved systems and methods for accurately diagnosing biological states based on the analysis of dynamic biological response data from biological samples non-invasively obtained from a subject. Such improved systems and methods for the accurate diagnosis of biological states can be based on a combination of dynamic immunohistochemical profiling of biological samples and artificial intelligence data analysis of such dynamic profiles for the evaluation of disease states. The present disclosure addresses these needs, for example, by providing biomarkers of biological samples for the diagnosis of biological states. Biological samples include human biological samples associated with incremental growth. Such biological samples can be hair shafts, teeth, and nails. The non-invasive biomarkers of the present disclosure can be used for the diagnosis of infants, in some cases, infants less than 1 year old.
[0004] In one aspect, the present disclosure provides a method for predicting a diagnostic state of a subject with respect to a disease or disorder, the method comprising: (a) staining a tooth sample of the subject to produce a stained tooth sample; (b) spatially analyzing fluorescence intensity across the stained tooth sample; and (c) predicting a diagnostic state of the subject with respect to the disease or disorder based at least in part on the analysis of the fluorescence intensity.
[0005] In some embodiments, analyzing determines the temporal dynamics of fundamental biological processes. In some embodiments, analyzing includes obtaining a fluorescence image of a stained tooth sample and analyzing the fluorescence intensity of the fluorescence image. In some embodiments, the fluorescence intensity varies spatially. In some embodiments, obtaining a fluorescence image of a stained tooth sample includes using a confocal microscope, either inverted or non-inverted. In some embodiments, staining the tooth sample includes using C-reactive protein immunohistochemical staining. In some embodiments, the method further includes sectioning the tooth sample. In some embodiments, staining the tooth sample includes (1) cutting the tooth sample, (2) decalcifying the tooth sample, (3) sectioning the decalcified sample, (4) staining the sections of the decalcified tooth with primary and secondary antibodies, (5) measuring spatial antibody fluorescence using a confocal microscope, and / or (6) extracting the temporal profile of the fluorescence intensity.
[0006] In some embodiments, the disease or disorder includes autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder includes ASD. In some embodiments, the subject is human. In some embodiments, the subject is an adult. In some embodiments, the subject is about 12 to about 5 years old. In some embodiments, the subject is about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or less than 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old.
[0007] In some embodiments, analyzing comprises generating a temporal profile of inflammation (e.g., one or more traces) based at least in part on fluorescence intensity and analyzing the temporal profile of inflammation. In some embodiments, at least a portion of the temporal profile of inflammation corresponds to a prenatal period of the subject.
[0008] In some embodiments, predicting the diagnostic state of a subject with respect to a disease or disorder involves processing fluorescence intensity using a trained model. In some embodiments, processing involves extracting features from the fluorescence intensity (e.g., by recurrence quantification analysis) and analyzing the features using a trained model. In some embodiments, the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm (e.g., a gradient boosting implementation of a machine learning algorithm such as a gradient boosted decision tree), and any combination thereof. In some embodiments, the trained model includes a gradient boosted ensemble model. In some embodiments, the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal line length, maximum diagonal line length, divergence, Shannon entropy in diagonal line length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, number of most probable mean diagonal line lengths (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the one or more features are extracted by applying recurrence quantification analysis (RQA) to a fluorescence intensity trace derived from the analysis of a sample. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, mean diagonal line length, maximum diagonal line length, divergence, Shannon entropy in diagonal line length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, number of most probable recurrence counts, mean diagonal line length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof.
[0009] In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters describing the curvature of one or more traces, determination of a sharp change in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantification analysis parameters, determination of one or more cross recurrence quantification analysis parameters, determination of one or more joint recurrence quantification analysis parameters, determination of one or more multi-dimensional recurrence quantification analysis parameters, estimation of a Lyapunov spectrum, or determination of a maximum Lyapunov exponent.
[0010] In some embodiments, the method further includes predicting a diagnostic state of a subject with respect to a disease or disorder using a model having at least about 70%, 75%, 80%, 85%, or 90% sensitivity in predicting a diagnostic state with respect to the disease or disorder across a suitable cohort population (such as those provided in the Examples section below).
[0011] In some embodiments, the method further includes predicting a diagnostic state of a subject with respect to a disease or disorder using a model having a maximum of about 70%, 75%, 80%, 85%, or 90% sensitivity in predicting a diagnostic state with respect to the disease or disorder across a suitable cohort population.
[0012] In some embodiments, the method further includes predicting a diagnostic state of a subject with respect to a disease or disorder using a model having at least about 70%, 75%, 80%, 85%, or 90% specificity in predicting a diagnostic state with respect to the disease or disorder across a suitable cohort population.
[0013] In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder using a model having a specificity of at most about 70%, 75%, 80%, 85%, or 90% when predicting the diagnostic state with respect to the disease or disorder across a suitable cohort population.
[0014] In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder using a model having a positive predictive value of at least about 70%, 75%, 80%, 85%, or 90% when predicting the diagnostic state with respect to the disease or disorder across a suitable cohort population.
[0015] In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder using a model having a positive predictive value of at most about 70%, 75%, 80%, 85%, or 90% when predicting the diagnostic state with respect to the disease or disorder across a suitable cohort population.
[0016] In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder using a model having a negative predictive value of at least about 70%, 75%, 80%, 85%, or 90% when predicting the diagnostic state with respect to the disease or disorder across a suitable cohort population.
[0017] In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder using a model having a negative predictive value of at most about 70%, 75%, 80%, 85%, or 90% when predicting the diagnostic state with respect to the disease or disorder across a suitable cohort population.
[0018] In some embodiments, the method further comprises predicting a diagnostic state of a subject with respect to a disease or disorder using a model that predicts a diagnostic state of the disease or disorder for a suitable cohort population having an area under the receiver operating characteristic curve (AUROC) of at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.82, at least about 0.84, at least about 0.86, at least about 0.88, or at least about 0.90.
[0019] In another aspect, the present disclosure provides a device comprising one or more processors and a memory storing one or more programs for execution by the one or more processors, the one or more programs comprising instructions for: (a) sampling each respective position along a reference line on a biological sample of a subject associated with a c-reactive protein of the subject, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement of the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with the c-reactive protein; (b) analyzing each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by the variation in c-reactive protein fluorescence intensity; and (d) using a trained model to process the features to determine a likelihood that the subject has a disease or disorder associated with the c-reactive protein. In some embodiments, each respective second data set is derived by applying a recurrence quantification analysis or a related method to the corresponding plurality of fluorescence intensity measurements.
[0020] In some embodiments, the plurality of fluorescence intensity measurements are measured using a confocal microscope, either inverted or non-inverted. In some embodiments, the biological sample includes a dental sample. In some embodiments, the dental sample is stained using C-reactive protein immunohistochemical staining. In some embodiments, the instructions further include sectioning the dental sample. In some embodiments, the instructions further include decalcifying the dental sample. In some embodiments, the disease or disorder includes autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder is ASD. In some embodiments, the subject is human. In some embodiments, the human is from about 12 to about 5 years old. In some embodiments, the subject is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, analyzing includes generating a temporal profile of inflammation (e.g., one or more traces) based at least in part on the plurality of fluorescence intensity measurements and analyzing the temporal profile of inflammation. In some embodiments, at least a portion of the temporal profile of inflammation corresponds to the prenatal period of the subject. In some embodiments, predicting the diagnostic status of the subject with respect to a disease or disorder includes processing the plurality of fluorescence intensity measurements using a trained model. In some embodiments, the trained model is selected from the group consisting of neural network algorithms, support vector machine algorithms, decision tree algorithms, unsupervised clustering algorithms, supervised clustering algorithms, regression algorithms, gradient boosting algorithms, and any combination thereof. In some embodiments, the trained model includes a gradient-boosted decision tree.In some embodiments, the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency to recurrence, laminarity, trapping time, maximum vertical length, Shannon entropy in vertical length, mean recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency to recurrence, laminarity, trapping time (TT), maximum vertical length, Shannon entropy in vertical length, mean recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method may apply one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters describing the curvature of one or more traces, determination of a sharp change in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantification analysis parameters, determination of one or more cross recurrence quantification analysis parameters, determination of one or more joint recurrence quantification analysis parameters, determination of one or more multi-dimensional recurrence quantification analysis parameters, estimation of the Lyapunov spectrum, or determination of the maximum Lyapunov exponent.
[0021] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium, the one or more computer programs, when executed by a computer system, cause the computer system to: (a) sample each respective position at a plurality of positions along a reference line on a biological sample associated with a target c-reactive protein, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement in the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample associated with the c-reactive protein; (b) analyze each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (c) derive a respective second data set from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by a continuous variation of the c-reactive protein fluorescence intensity; and (d) use a trained model to process the features to determine the likelihood that the subject has a disease or disorder associated with the c-reactive protein. In some embodiments, each respective second data set is derived by applying a recurrence quantification analysis or a related method to the corresponding plurality of fluorescence intensity measurements.
[0022] In some embodiments, the plurality of fluorescence intensity measurements are measured using an inverted or non-inverted confocal microscope. In some embodiments, the biological sample includes a dental sample. In some embodiments, the dental sample is stained using C-reactive protein immunohistochemical staining. In some embodiments, the method further includes sectioning the dental sample. In some embodiments, the method further includes decalcifying the dental sample. In some embodiments, the disease or disorder includes autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder is ASD. In some embodiments, the subject is human. In some embodiments, the subject is less than 5 years old. In some embodiments, the subject is less than 1 year old. In some embodiments, analyzing comprises generating a temporal profile of inflammation (e.g., one or more traces) based at least in part on the plurality of fluorescence intensity measurements and analyzing the temporal profile of inflammation. In some embodiments, at least a portion of the temporal profile of inflammation corresponds to the prenatal period of the subject. In some embodiments, predicting the diagnostic status of the subject with respect to the disease or disorder comprises processing the plurality of fluorescence intensity measurements using a trained model. In some embodiments, the trained model is selected from the group consisting of neural network algorithms, support vector machine algorithms, decision tree algorithms, unsupervised clustering algorithms, supervised clustering algorithms, regression algorithms, gradient boosting algorithms, and any combination thereof. In some embodiments, the trained model includes a gradient-boosted decision tree.In some embodiments, the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, laminarity, trapping time, maximum vertical length, Shannon entropy in vertical length, mean recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, laminarity, trapping time (TT), maximum vertical length, Shannon entropy in vertical length, mean recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof.
[0023] In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method may apply one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters describing the curvature of one or more traces, determination of a sharp change in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantification analysis parameters, determination of one or more cross recurrence quantification analysis parameters, determination of one or more joint recurrence quantification analysis parameters, determination of one or more multi-dimensional recurrence quantification analysis parameters, estimation of the Lyapunov spectrum, or determination of the maximum Lyapunov exponent.
[0024] In another aspect, the present disclosure provides a method of training a model in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, the method comprising: (a) for each respective training subject among a plurality of training subjects, where a first subset of the training subjects among the plurality of training subjects has a first diagnostic state corresponding to having a first biological state associated with a c-reactive protein, and a second subset of the training subjects among the plurality of training subjects has a second diagnostic state corresponding to not having the first biological state associated with the c-reactive protein, (i) sampling each respective position along a reference line on a biological sample of the subject associated with the c-reactive protein of the subject, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement among the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with the c-reactive protein; (ii) analyzing each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first dataset; and (iii) deriving a respective second dataset from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by the variation in the c-reactive protein fluorescence intensity; and (b) training a model that is not trained or is partially trained using (i) the corresponding sets of features of the respective second datasets of each training subject among the plurality of training subjects, and (ii) the corresponding diagnostic state of each training subject among the plurality of training subjects selected from among the first diagnostic state and the second diagnostic state, thereby obtaining a trained model that provides an indication as to whether a test subject has a first biological state associated with a c-reactive protein based on the values of the features in the set of features obtained from a biological sample associated with the c-reactive protein of the test subject. In some embodiments, each respective second dataset is derived by applying a recurrence quantification analysis or a related method to the corresponding plurality of fluorescence intensity measurements.
[0025] In some embodiments, the trained model is a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering model algorithm, a supervised clustering model algorithm, a regression model, a gradient boosting algorithm (e.g., a gradient boosting implementation of a machine learning algorithm such as a gradient boosted decision tree), or any combination thereof. In some embodiments, the trained model includes a gradient boosted ensemble model. In some embodiments, the trained model predicts a transition for a polynomial distribution. In some embodiments, the trained model predicts a transition for a binomial distribution. In some embodiments, the first biological state associated with c-reactive protein is selected from the group consisting of autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer.
[0026] In some embodiments, evaluating a test subject for a first biological state associated with C-reactive protein further comprises distinguishing between the presence of the first biological state associated with C-reactive protein and the absence of the first biological state associated with C-reactive protein. In some embodiments, evaluating a test subject for a first biological state associated with C-reactive protein further comprises distinguishing between the first biological state associated with C-reactive protein and a second biological state associated with C-reactive protein that is different from the first biological state associated with C-reactive protein. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is typical development, i.e., the absence of a neurodevelopmental disorder. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is attention deficit / hyperactivity disorder. In some embodiments, the test subject is a human. In some embodiments, the human is from about 12 to about 5 years old. In some embodiments, the subject is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, the corresponding biological sample associated with the C-reactive protein of each training subject is selected from the group consisting of hair shafts, teeth, and nails. In some embodiments, the corresponding biological sample associated with the C-reactive protein of each training subject is a hair shaft and the reference line corresponds to the longitudinal direction of the hair shaft. In some embodiments, the corresponding biological sample associated with the C-reactive protein of each training subject is a tooth and the reference line corresponds to a direction across the growth zone including the neonatal line of the tooth. In some embodiments, the corresponding plurality of positions are arranged such that a first position among the corresponding plurality of positions along the corresponding biological sample associated with the C-reactive protein of each training subject corresponds to the position closest to the tip of the corresponding biological sample associated with the C-reactive protein of each training subject.In some embodiments, each trace in a corresponding plurality of fluorescence intensity measurement values includes a plurality of data points, and each data point is an instance of each position at a plurality of positions. In some embodiments, the corresponding set of features is selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, laminarity, trapping time, maximum vertical length, Shannon entropy in vertical length, mean recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the features are derived from recurrence quantitative analysis of fluorescence traces or related computational analysis. In some embodiments, the corresponding plurality of positions includes at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 14000, 16000, 18000, 20000, or more than 20000 positions.
[0027] In some embodiments, the corresponding set of features is selected from the group of temporal dynamic features of one or more traces. In some embodiments, the temporal dynamic features of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method may apply one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters describing the curvature of one or more traces, determination of a sharp change in intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross-recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, or determination of the maximum Lyapunov exponent.
[0028] Another aspect of the present disclosure provides a non-transitory computer-readable medium including machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere in this specification.
[0029] Another aspect of the present disclosure provides a system including one or more computer processors and a computer memory coupled to the computer processors. The computer memory includes machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere in this specification.
[0030] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description. In which only exemplary embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other different embodiments and some of its details are capable of various obvious modifications without departing from the present disclosure in its entirety. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.
[0031] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the incorporated publications and patents or patent applications conflict with the disclosure contained herein, this specification is intended to supersede and / or take precedence over such conflicting material.
[0032] The novel features of the invention are set forth in detail in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description and the accompanying drawings (referred to herein as "Figure" and "FIG.") that describe exemplary embodiments in which the principles of the invention are utilized. BRIEF DESCRIPTION OF THE DRAWINGS
[0033]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
DETAILED DESCRIPTION OF THE INVENTION
[0034] Although various embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Without departing from the present invention, numerous variations, modifications, and substitutions may be contemplated by those skilled in the art here. It should be understood that various alternatives to the embodiments of the present invention described herein may be used.
[0035] Dynamic biological reactions can exhibit fundamental biological processes that are structurally and functionally important for humans. For example, abnormal dynamic biological reactions can be associated with many biological conditions such as diseases and disorders. Examples of such biological conditions can include neurological conditions (e.g., autism spectrum disorder, schizophrenia, or attention deficit / hyperactivity disorder (ADHD)), neurodegenerative conditions (e.g., amyotrophic lateral sclerosis (ALS), Alzheimer's disease, Parkinson's disease, and Huntington's disease), and cancer (e.g., pediatric cancer).
[0036] Considering the above background, there is a need for accurate methods and systems for the diagnosis of biological states, particularly non-invasive diagnosis. Such diagnosis can be based on the accurate profiling of biomarkers detectable by non-invasive methods for the diagnosis of biological states. The present disclosure provides improved systems and methods for accurately diagnosing biological states based on the analysis of dynamic biological response data from biological samples non-invasively obtained from a subject. Such improved systems and methods for the accurate diagnosis of biological states can be based on a combination of dynamic immunohistochemical profiling of biological samples and artificial intelligence data analysis of such dynamic profiles for the evaluation of disease states. The present disclosure addresses these needs, for example, by providing biological sample biomarkers for the diagnosis of biological states. Biological samples include human biological samples associated with incremental growth. Such biological samples can be hair shafts, teeth, and nails. The non-invasive biomarkers of the present disclosure can be used for the diagnosis of infants, in some cases, infants less than 1 year old. In some cases, the pediatric patient is from about 12 to about 5 years old. In some embodiments, the pediatric patient is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the pediatric patient is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old.
[0037] In one aspect, the present disclosure provides a method for predicting a diagnostic state of a subject with respect to a disease or disorder, the method comprising: (a) staining a tooth sample of the subject to generate a stained tooth sample; (b) spatially analyzing fluorescence intensity across the stained tooth sample; and (c) predicting the diagnostic state of the subject with respect to the disease or disorder based at least in part on the analysis of the fluorescence intensity.
[0038] In some embodiments, analyzing comprises obtaining a fluorescence image of a stained tooth sample and analyzing the fluorescence intensity of the fluorescence image. In some embodiments, obtaining a fluorescence image of a stained tooth sample comprises using a confocal microscope, either inverted or non-inverted. In some embodiments, staining the tooth sample comprises using C-reactive protein immunohistochemical staining. In some embodiments, the method further comprises sectioning the tooth sample. In some embodiments, staining the tooth sample comprises decalcifying the tooth sample.
[0039] In some embodiments, the systems and methods disclosed herein may use C-reactive protein fluorescence immunohistochemical staining alone or in combination with other techniques. Such techniques may include laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS), Raman spectroscopy, or any combination thereof. In some embodiments, combining techniques may improve the diagnostic accuracy or precision of a given technique alone. In some embodiments, the addition of LA-ICP-MS provides multiple non-invasive metal metabolism biomarkers of a given biological sample that may complement the diagnostic power of C-reactive protein fluorescence immunohistochemical data. In some embodiments, the metal metabolism biomarkers include zinc, tin, magnesium, copper, iodide, lithium, aluminum, phosphorus, sulfur, calcium, chromium, manganese, iron, cobalt, nickel, arsenic, strontium, cadmium, tin, iodine, barium, mercury, lead, bismuth, molybdenum, or any combination thereof. In some embodiments, the addition of Raman spectroscopy provides multiple spectra indicative of physiological changes induced by a disease or external stressor to complement the diagnostic power of C-reactive protein fluorescence immunohistochemical data. In some embodiments, the multiple metal metabolism biomarkers include at least 2, at least 5, or at least 10 metal metabolism biomarkers. In some embodiments, the multiple metal metabolism biomarkers include 20 or fewer, 10 or fewer, or 5 or fewer metal metabolism biomarkers. In some embodiments, the multiple metal metabolism biomarkers consist of 2 to 5, 3 to 10, or 8 to 20 metal metabolism biomarkers. In some embodiments, the multiple metal metabolism biomarkers are within another range starting from 2 or more metal metabolism biomarkers and ending with 20 or fewer metal metabolism biomarkers. In some embodiments, the multiple spectra include at least 2, at least 5, or at least 10 spectra. In some embodiments, the multiple spectra include 20 or fewer, 10 or fewer, or 5 or fewer spectra. In some embodiments, the multiple spectra consist of 2 to 5, 3 to 10, or 8 to 20 spectra.In some embodiments, the plurality of spectra are within another range that starts with two or more spectra and ends with twenty or fewer spectra.
[0040] In some embodiments, the disease or disorder includes autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder includes ASD. In some embodiments, the subject is human. In some embodiments, the subject is an adult. In some embodiments, the subject is about 12 to about 5 years old. In some embodiments, the subject is about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or less than 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old.
[0041] In some embodiments, analyzing comprises generating a temporal profile of inflammation, at least in part based on fluorescence intensity, and analyzing the temporal profile of inflammation. In some embodiments, at least a portion of the temporal profile of inflammation corresponds to the prenatal period of the subject.
[0042] In some embodiments, predicting the diagnostic state of a subject regarding a disease or disorder involves processing fluorescence intensity using a trained model. In some embodiments, this trained model includes a plurality of parameters, and the term "parameter" can affect (e.g., modify, adjust, and / or tune) one or more inputs, outputs, and / or functions within the model (e.g., if the model is a regressor or classifier), and refers to any coefficient or similarly any value of an internal or external element (e.g., weight and / or hyperparameter) within the model. For example, in some embodiments, the parameters of the model refer to any coefficient, weight, and / or hyperparameter that can be used to control, modify, adjust, and / or tune the behavior, learning, and / or performance of the model. In some examples, the parameters are used to increase or decrease the effect of an input (e.g., a feature) on the model. As a non-limiting example, in some embodiments, the parameters are used to increase or decrease the effect of a node (e.g., a neural network), and the node includes one or more activation functions. The assignment of parameters to specific inputs, outputs, and / or functions of the model is not limited to any one paradigm of a given model, but can be used in any suitable model for desired performance. In some embodiments, the parameters have fixed values. In some embodiments, the values of the parameters are manually and / or automatically adjustable. In some embodiments, the values of the parameters are modified by the model verification and / or training process (e.g., by an error minimization and / or backpropagation method). In some embodiments, the models of the present disclosure include a plurality of parameters.In some embodiments, the plurality of parameters associated with a model (e.g., an untrained, partially trained, or fully trained model) are n parameters, where n≥2, n≥5, n≥10, n≥25, n≥40, n≥50, n≥75, n≥100, n≥125, n≥150, n≥200, n≥225, n≥250, n≥350, n≥500, n≥600, n≥750, n≥1,000, n≥2,000, n≥4,000, n≥5,000, n≥7,500, n≥10,000, n≥20,000, n≥40,000, n≥75,000, n≥100,000, n≥200,000, n≥500,000, n≥1×10. 6 , n≥5×10 6 , or n≥1×10 7 . In some embodiments, n is between 10,000 and 1×10 7 , between 100,000 and 5×10 6 , or between 500,000 and 1×10 6 .
[0043] In some embodiments, the plurality of parameters include at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1×10 6 parameters, at least 5×10 6 parameters, at least 1×10 7 parameters, at least 5×10 7 parameters, or at least 1×10 8 parameters. In some embodiments, the plurality of parameters include 1×10 9 parameters or fewer, 1×10 8 parameters or fewer, 1×10 7 parameters or fewer, 1×10 6 parameters or fewer, 100,000 parameters or fewer, 10,000 parameters or fewer, 1000 parameters or fewer, or 100 parameters or fewer. In some embodiments, the plurality of parameters are between 10 and 10,000, between 100 and 100,000, between 1000 and 1×10 6 parameters, between 100,000 and 1×10 7 parameters, between 1×10 6 and 1×10 8 parameters, or between 1×10 7~1×10 9 consists of parameters. In some embodiments, the plurality of parameters start from 10 or more parameters and are within another range that ends with 1×10 9 or fewer parameters.
[0044] In some embodiments, processing fluorescence intensity (e.g., one or more traces of fluorescence intensity described elsewhere herein) using a trained model includes extracting features from the fluorescence intensity (e.g., by recurrence quantification analysis), and analyzing the features using the trained model. In some embodiments, the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm (e.g., a gradient boosting implementation of a machine learning algorithm such as a gradient boosted decision tree), and any combination thereof. In some embodiments, the trained model includes a gradient boosted ensemble model. In some embodiments, the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal line length, maximum diagonal line length, divergence, Shannon entropy in diagonal line length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, most probable number of recurrences, mean diagonal line length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the one or more features are extracted by applying recurrence quantification analysis (RQA) to a fluorescence intensity trace derived from the analysis of a sample. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, mean diagonal line length, maximum diagonal line length, divergence, Shannon entropy in diagonal line length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, most probable number of recurrences, mean diagonal line length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof.
[0045] In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters that describe the curvature of one or more traces, determination of a sharp change in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantification analysis parameters, determination of one or more cross recurrence quantification analysis parameters, determination of one or more joint recurrence quantification analysis parameters, determination of one or more multi-dimensional recurrence quantification analysis parameters, estimation of a Lyapunov spectrum, or determination of a maximum Lyapunov exponent.
[0046] In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder with at least about 80% sensitivity. In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder with at least about 80% specificity. In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder with at least about 80% positive predictive value. In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder with at least about 80% negative predictive value. In some embodiments, the method further includes predicting the diagnostic state of a subject with respect to a disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.80.
[0047] In another aspect, the present disclosure provides a device comprising one or more processors and a memory storing one or more programs for execution by the one or more processors, the one or more programs comprising instructions for performing: (a) sampling each respective position among a plurality of positions along a baseline on a biological sample of interest associated with a target c-reactive protein, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement among the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of interest associated with the c-reactive protein; (b) analyzing each fluorescence intensity across the baseline on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by a continuous variation in the c-reactive protein fluorescence intensity; and (d) using a trained model to process the features to determine the likelihood that the subject has a disease or disorder associated with the c-reactive protein. In some embodiments, each respective second data set is derived by applying a recurrence quantification analysis or a related method to the corresponding plurality of fluorescence intensity measurements.
[0048] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium, the one or more computer programs, when executed by a computer system, cause the computer system to: (a) sample each respective position at a plurality of positions along a reference line on a biological sample associated with a target c-reactive protein, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement in the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample associated with the c-reactive protein; (b) analyze each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (c) derive a respective second data set from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by a continuous variation in the c-reactive protein fluorescence intensity; and (d) use a trained model to process the features to determine the likelihood that the subject has a disease or disorder associated with the c-reactive protein. In some embodiments, each respective second data set is derived by applying a recurrence quantification analysis or a related method to the corresponding plurality of fluorescence intensity measurements.
[0049] In another aspect, the present disclosure provides a method for training a model in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, the method comprising: (a) for each respective training subject among a plurality of training subjects, where a first subset of the training subjects among the plurality of training subjects has a first diagnostic state corresponding to having a first biological state associated with a c-reactive protein, and a second subset of the training subjects among the plurality of training subjects has a second diagnostic state corresponding to not having the first biological state associated with the c-reactive protein, (i) sampling each respective position along a reference line on a biological sample of the subject associated with the subject's c-reactive protein, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement among the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with the c-reactive protein; (ii) analyzing each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first dataset; and (iii) deriving a respective second dataset from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by the variation in c-reactive protein fluorescence intensity; and (b) training a model that is not trained or is partially trained using (i) the corresponding set of features of each respective second dataset of each training subject among the plurality of training subjects, and (ii) the corresponding diagnostic state of each training subject among the plurality of training subjects selected from among the first diagnostic state and the second diagnostic state, thereby obtaining a trained model that provides an indication as to whether a test subject has a first biological state associated with a c-reactive protein based on the values of the features in the set of features obtained from a biological sample associated with the test subject's c-reactive protein. In some embodiments, each respective second dataset is derived by applying a recurrence quantification analysis or a related method to the corresponding plurality of fluorescence intensity measurements.
[0050] In some embodiments, each subject (e.g., a test subject) is selected from a plurality of subjects. In some embodiments, the plurality of subjects includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, or at least 500 subjects. In some embodiments, the plurality of subjects includes 1,000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, 20 or fewer, or 10 or fewer subjects. In some embodiments, the plurality of subjects consists of 2 to 10, 5 to 20, 10 to 100, or 100 to 1,000 subjects. In some embodiments, the plurality of subjects is within another range starting from 2 or more subjects and ending at 1,000 or fewer subjects.
[0051] In some embodiments, the plurality of training subjects includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1,000, at least 5,000, or at least 100,000 training subjects. In some embodiments, the plurality of training subjects includes 1,000,000 or fewer, 100,000 or fewer, 10,000 or fewer, 1,000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, 20 or fewer, or 10 or fewer training subjects. In some embodiments, the plurality of training subjects consists of 2 to 1,000, 500 to 10,000, 10,000 to 100,000, or 100,000 to 1,000,000 training subjects. In some embodiments, the plurality of training subjects is within another range starting from 2 or more training subjects and ending at 1,000,000 or fewer training subjects.
[0052] In some embodiments, each subset of training subjects among a plurality of training subjects (e.g., the first subset and / or the second subset) includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 training subjects. In some embodiments, each subset of training subjects includes 500,000 or fewer, 10,000 or fewer, 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, 20 or fewer, or 10 or fewer training subjects. In some embodiments, each subset of training subjects consists of 2 to 100, 50 to 2000, 1000 to 10,000, or 10,000 to 500,000 training subjects. In some embodiments, each subset of training subjects is within another range that starts with 2 or more training subjects and ends with 500,000 or fewer training subjects.
[0053] In some embodiments, a set of features (e.g., obtained from a biological sample associated with a c-reactive protein of a subject) includes at least 1, at least 2, at least 3, at least 5, at least 8, at least 10, at least 15, or at least 20 features. In some embodiments, the set of features includes 50 or fewer, 20 or fewer, 10 or fewer, 5 or fewer, or 3 or fewer features. In some embodiments, the set of features consists of 1 to 10, 4 to 15, 8 to 20, or 15 to 50 features. In some embodiments, the set of features is within another range that starts with 1 or more features and ends with 50 or fewer features.
[0054] In some embodiments, the trained model is a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering model algorithm, a supervised clustering model algorithm, a regression model, a gradient boosting algorithm (e.g., a gradient boosting implementation of a machine learning algorithm such as a gradient-boosted decision tree), or any combination thereof. In some embodiments, the trained machine learning model includes a gradient-boosted ensemble model. In some embodiments, the trained model predicts a transition for a polynomial distribution. In some embodiments, the trained model predicts a transition for a binomial distribution. In some embodiments, the first biological state associated with c-reactive protein is selected from the group consisting of autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer.
[0055] In some embodiments, the first biological state and / or the second biological state are selected from a plurality of biological states. In some embodiments, the plurality of biological states include at least two, at least five, or at least ten biological states. In some embodiments, the plurality of biological states include twenty or fewer, ten or fewer, or five or fewer biological states. In some embodiments, the plurality of biological states consist of two to five, three to ten, or eight to twenty biological states. In some embodiments, the plurality of biological states are within another range starting from two or more biological states and ending with twenty or fewer biological states. In some embodiments, the first diagnostic state and / or the second diagnostic state are selected from a plurality of diagnostic states. In some embodiments, the plurality of diagnostic states include at least two, at least five, or at least ten diagnostic states. In some embodiments, the plurality of diagnostic states include twenty or fewer, ten or fewer, or five or fewer diagnostic states. In some embodiments, the plurality of diagnostic states consist of two to five, three to ten, or eight to twenty diagnostic states. In some embodiments, the plurality of diagnostic states are within another range starting from two or more diagnostic states and ending with twenty or fewer diagnostic states.
[0056] In some embodiments, evaluating a subject for a first biological state associated with C-reactive protein further comprises distinguishing the presence of the first biological state associated with C-reactive protein from the absence of the first biological state associated with C-reactive protein. In some embodiments, evaluating a subject for a first biological state associated with C-reactive protein further comprises distinguishing the first biological state associated with C-reactive protein from a second biological state associated with a C-reactive protein different from the first biological state associated with C-reactive protein. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is typical development, i.e., the absence of a neurodevelopmental disorder. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is attention deficit / hyperactivity disorder. In some embodiments, the subject is a human. In some embodiments, the human is from about 12 to about 5 years old. In some embodiments, the human is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the human is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, the human is at least 13 years old, at least 14 years old, at least 15 years old, at least 18 years old, or at least 30 years old. In some embodiments, the human is 40 years old or younger, 30 years old or younger, 18 years old or younger, or 15 years old or younger. In some embodiments, the human is in another age range starting from about 1 year old and ending at about 40 years old or younger. In some embodiments, the corresponding biological sample associated with the C-reactive protein of each subject (e.g., test subject and / or training subject) is selected from the group consisting of hair shafts, teeth, and nails. In some embodiments, the corresponding biological sample associated with the C-reactive protein of each subject (e.g., test subject and / or training subject) is a hair shaft, and the reference line corresponds to the longitudinal direction of the hair shaft.In some embodiments, the corresponding biological sample associated with the c-reactive protein of each subject (e.g., a test subject and / or a training subject) is a tooth, and the reference line corresponds to a direction across a growth zone including the neonatal line of the tooth. In some embodiments, the first position among the corresponding plurality of positions along the corresponding biological sample associated with the c-reactive protein of each subject (e.g., a test subject and / or a training subject) corresponds to the position closest to the tip of the corresponding biological sample associated with the c-reactive protein of each subject, and the corresponding plurality of positions are arranged accordingly. In some embodiments, each trace in the corresponding plurality of fluorescence intensity measurements includes a plurality of data points, and each data point is an instance of each position among the plurality of positions. In some embodiments, the corresponding set of features is selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, laminarity, trapping time, maximum vertical length, Shannon entropy in vertical length, mean recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the features are derived from a recurrence quantitative analysis of the fluorescence trace or a related computational analysis. In some embodiments, the corresponding plurality of positions includes at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 14000, 16000, 18000, 20000, or more than 20000 positions.
[0057] In some embodiments, the plurality of positions along the reference line on the biological sample includes at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, or at least 1×10 6 positions. In some embodiments, the plurality of positions is 1×107 less than, 1×10 6 less than, 100,000 or less, 10,000 or less, 1000 or less, or 100 or less positions. In some embodiments, the plurality of positions is 50 to 1000, 500 to 50,000, 10,000 to 1×10 6 positions, or consists of 1×10 6 to 1×10 7 positions. In some embodiments, the plurality of positions starts at 50 or more positions and is within another range ending at 1×10 7 positions or less. In some embodiments, each respective growth period of the biological sample is selected from a plurality of growth periods. In some embodiments, the plurality of growth periods includes at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, or at least 1×10 6 growth periods. In some embodiments, the plurality of growth periods includes 1×10 7 or less, 1×10 6 or less, 100,000 or less, 10,000 or less, 1000 or less, or 100 or less growth periods. In some embodiments, the plurality of growth periods is 50 to 1000, 500 to 50,000, 10,000 to 1×10 6 growth periods, or consists of 1×10 6 to 1×10 7 growth periods. In some embodiments, the plurality of growth periods starts at 50 or more growth periods and is within another range ending at 1×10 7 growth periods or less.
[0058] In some embodiments, the plurality of fluorescence intensity measurement values is at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, or at least 1×10 6It includes individual fluorescence intensity measurement values. In some embodiments, the plurality of fluorescence intensity measurement values are 1×10 7 or less, 1×10 6 or less, 100,000 or less, 10,000 or less, 1000 or less, or 100 or less fluorescence intensity measurement values. In some embodiments, the plurality of fluorescence intensity measurement values are 50 to 1000, 500 to 50,000, 10,000 to 1×10 6 pieces, or 1×10 6 to 1×10 7 pieces. In some embodiments, the plurality of fluorescence intensity measurement values start from 50 or more fluorescence intensity measurement values and are within another range ending with 1×10 7 or less fluorescence intensity measurement values.
[0059] In some embodiments, each fluorescence intensity measurement value includes one or more traces of fluorescence intensity. In some embodiments, each fluorescence intensity measurement value includes at least 1, at least 2, at least 3, at least 4, at least 5, or at least 10 traces. In some embodiments, each fluorescence intensity measurement value includes 50 or less, 10 or less, 5 or less, or 3 or less traces. In some embodiments, each fluorescence intensity measurement value consists of 1 to 5, 2 to 10, or 10 to 20 traces. In some embodiments, each fluorescence intensity measurement value includes another range of traces starting from one or more traces and ending with 20 or less traces.
[0060] In some embodiments, for each trace in the plurality of fluorescence intensity measurement values, the plurality of data points include at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 data points. In some embodiments, the plurality of data points include 500,000 or fewer, 10,000 or fewer, 1000 or fewer, 500 or fewer, 100 or fewer, 50 or fewer, 20 or fewer, or 10 or fewer data points. In some embodiments, the plurality of data points consist of 2 to 100, 50 to 2000, 1000 to 10,000, or 10,000 to 500,000 data points. In some embodiments, the plurality of data points are within another range that starts with 2 or more data points and ends with 500,000 or fewer data points.
[0061] In some embodiments, the temporal dynamics of one or more traces are analyzed by a data analysis method. In some embodiments, the data analysis method may apply one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters describing the curvature of one or more traces, determination of a sharp change in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantification analysis parameters, determination of one or more cross recurrence quantification analysis parameters, determination of one or more joint recurrence quantification analysis parameters, determination of one or more multi-dimensional recurrence quantification analysis parameters, estimation of the Lyapunov spectrum, or determination of the maximum Lyapunov exponent.
[0062] The details of the exemplary system are described in conjunction with FIG. 1, which shows an example of a block diagram of the computing device 100 of the present disclosure. In some implementations, the device 100 includes one or more processing units CPU 102 (also referred to as processors), one or more network interfaces 104, a user interface 106, non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communication between system components. Non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while persistent memory 112 typically includes CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid storage device. Persistent memory 112 optionally includes one or more storage devices located remotely from the CPU(s) 102. The persistent memory 112 and the non-volatile memory device(s) within the non-persistent memory 112 include non-transitory computer-readable storage media.In some implementations, the non-persistent memory 111 or alternatively the non-transitory computer-readable storage medium stores, optionally along with the persistent memory 112, the following programs, modules, and data structures, or subsets thereof, namely, an optional operating system 116 that addresses various basic system services and includes procedures for performing hardware-dependent tasks, an optional network communication module (or instructions) 118 for connecting the system 100 to other devices and / or communication network 104, an optional classifier training module 120 for a training model for evaluating a subject for a biological state, an optional data store 122 for a dataset for a biological sample from a training subject that includes feature data for one or more training subjects 124, where the feature data includes parameters associated with each of the features 126 and a diagnostic state 128 (e.g., an indication of whether each training subject has been diagnosed as having a biological state or not having a biological state), a data store 122, an optional classifier verification module 130 for verifying a model for distinguishing biological states, an optional data store 132 for a dataset for a biological sample from a verification subject, and an optional patient classification module 134 for classifying a subject having a biological state, such as trained using the classifier training module 120.
[0063] In various implementations, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for performing the above-described functions. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules, and thus, various subsets of these modules and data may be combined in various implementations or otherwise rearranged. In some implementations, the non-transitory memory 111 optionally stores a subset of the modules and data structures identified above. Further, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than the computer system of the visualization system 100, and this computer system is addressable by the visualization system 100, whereby the visualization system 100 can retrieve all or a portion of such data when needed.
[0064] In some embodiments, the system 100 is connected to or includes one or more analytical devices for performing chemical analysis. For example, an optional network communication module (or instructions) 118 is configured to connect the system 100 to one or more analytical devices via, for example, a communication network 104. In some embodiments, the one or more analytical devices include a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer.
[0065] FIG. 1 depicts "System 100", which is more intended as a functional description of various features that may exist in a computer system rather than a structural schematic of the implementation forms described in this specification. In fact, as recognized by those skilled in the art, the separately shown items may be combined and some items may be separated. Further, although FIG. 1 depicts specific data and modules within non-persistent memory 111, some or all of these data and modules may be stored in persistent memory 112.
[0066] In some embodiments, the methods of the present disclosure include obtaining a biological sample (e.g., a single hair including a hair shaft). In some embodiments, the subject is a human. In some embodiments, the subject is a pediatric patient under 5 years of age (e.g., the pediatric patient is 5 years old, 4 years old, 3 years old, 2 years old, 1 year old, 9 months old, 6 months old, 3 months old, or 1 month old or younger). In some embodiments, the subject is an adult. FIG. 2A shows an example of a hair sample from a subject including a hair shaft. In some embodiments, the hair sample is cut from the subject (e.g., with the aid of scissors). In some embodiments, the method of obtaining the hair sample is non-invasive. The obtained hair sample may have a minimum length of 1 cm (e.g., the hair sample is 1 cm, 2 cm, 3 cm, 4 cm, or 5 cm in length). The hair sample may include any part of the hair (e.g., the tip or the part between the tip and the follicle). In particular, there are no special requirements for a hair sample including a hair follicle. FIG. 2B shows an example of a tooth sample from a subject. FIG. 2C shows an example of a nail sample from a subject. In the example of a nail or hair, obtaining a biological sample may refer to positioning the subject so that the nail or hair can be sampled. The nail sample may include the entire nail or a cut nail.
[0067] In some embodiments, the obtained biological sample is pretreated, such as by washing and drying the biological sample using one or more solvents and / or surfactants. When the biological sample is hair, the hair sample is washed in a solution of TRITON X-100 (registered trademark) and ultrapure deionized water (e.g., MILLI-Q (registered trademark) water) and dried overnight in an oven (e.g., at 60 °C). The pretreatment may further include preparing the hair shaft for measurement by placing the hair shaft on a glass slide (e.g., a microscope glass slide) having an adhesive film (e.g., double-sided tape). The hair shaft may be positioned such that the hair shaft is substantially straight. The glass slide containing the hair shaft may be placed in or near a measurement system (e.g., a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer) for analysis. When the biological sample is a tooth or a nail, the surface of the biological sample may be washed (e.g., with a surfactant, water, or one or more solvents). In some cases, the sample is decalcified before, after, immediately before, or in any combination of time frames prior to performing the analysis described elsewhere herein. In some cases, decalcifying the sample includes: (a) immersing the tooth in a solution of ethylenediaminetetraacetic acid (EDTA), wherein the EDTA may have a pH of from about 7.0 to about 7.4 for up to about 5 weeks; (b) weighing the tooth once a week; and (c) removing the sample when the weight of the tooth plateau changes. The sample may be placed in or near a measurement system (e.g., a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer) for analysis.
[0068] FIG. 3 shows a flowchart of a method 300 for evaluating a subject for a biological state, such as a method for predicting a diagnostic state of a subject with respect to a disease or disorder. The method 300 may include staining a dental sample of the subject (such as at operation 302) to produce a stained dental sample. Next, the method 300 may include spatially analyzing fluorescence intensity across the stained dental sample (such as at operation 304). Next, the method 300 may include predicting a diagnostic state of the subject with respect to a disease or disorder, based at least in part on the analysis of the fluorescence intensity (such as at operation 306).
[0069] In some embodiments, analyzing includes obtaining a fluorescence image of the stained dental sample and analyzing the fluorescence intensity of the fluorescence image. In some embodiments, obtaining a fluorescence image of the stained dental sample includes using a confocal microscope, either inverted or non-inverted. In some embodiments, staining the dental sample includes using C-reactive protein immunohistochemical staining. In some embodiments, the method further includes sectioning the dental sample. In some embodiments, staining the dental sample includes decalcifying the dental sample.
[0070] In some embodiments, analyzing includes generating a temporal profile of inflammation and analyzing the temporal profile of inflammation, based at least in part on the fluorescence intensity. In some embodiments, at least a portion of the temporal profile of inflammation corresponds to a prenatal period of the subject.
[0071] In some embodiments, measurement data is collected continuously from a biological sample at a plurality of positions along the biological sample. In some embodiments, the plurality of positions along a reference line of the biological sample includes at least 100 positions (e.g., 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 6000, 7000, 8000, 9000, 10000, 12000, 14000, 16000, 18000, 20000, or more than 20000 positions). In some embodiments, each position is adjacent to each other. By this method, each region corresponding to a different position on the biological sample can thereby be associated with a dynamic (e.g., time-varying) abundance measurement. In some embodiments, each position is separated by a predetermined distance. In some embodiments, sampling is performed along a reference line of the biological sample starting from each position closest to the tip of the biological sample, such as a hair sample (e.g., the position corresponding to the youngest age of the subject). Generally, as long as the direction of sampling is known, sampling can start from each position closest to the tip or the root, and an appropriate trained model is used for the analysis.
[0072] Sampling can generate a set of data points. Each set of data points can correspond to a measurement value (e.g., abundance or concentration) of a substance indicative of a dynamic biological response measured at a plurality of positions along the biological sample. Each position on the reference line of the biological sample can correspond to a particular growth time of the biological sample.
[0073] In some embodiments, in an example of a biological sample of a tooth, the reference line may include growth from 240 to 510 days (e.g., the period of crown formation depending on the type of tooth). In some embodiments, each position along the reference line may correspond to from about 1 to about 0.5 micrometers. Alternatively, or in addition, the biological sample may include a hair shaft, where the position along the reference line corresponds to about one-fifth of the growth (e.g., the period of hair growth calculated using a resolution of 1 micrometer and an average hair growth rate of 1 cm / month). By correlating a plurality of positions along the reference line of the biological sample with the corresponding periods of growth, a first data set including a plurality of traces is obtained. Each trace includes the time-dependent abundance of a substance (e.g., abundance or concentration) that indicates a dynamic biological response measured from the biological sample. For example, the distance between positions may correspond to the estimated growth (e.g., biological time) of the biological sample. For example, the abundance may be measured for a tooth sample along a distance of up to about 8 millimeters (mm), which corresponds to a biological time of about 240 to 510 days. Alternatively, or in addition, the abundance may be measured for a hair sample along a distance of 1.2 cm, which corresponds to a biological time of about 35 days. The biological time can be estimated by using the average rate of hair growth (e.g., 1 cm per month).
[0074] In some embodiments, the data analysis is performed on one or more traces corresponding to the time-dependent abundance(s) (e.g., time-dependent concentration) of a substance that indicates a dynamic biological response measured from the biological sample. This may include customized operations for cleaning up the data (e.g., smoothing the data over a time span and / or removing data points that are higher or lower than a predetermined threshold). In some embodiments, the data analysis includes removing from the trace data points having an average absolute difference between adjacent data points that is at least 1, 2, or 3 times the standard deviation of the average absolute difference between adjacent data points.
[0075] In some embodiments, the temporal dynamics of one or more traces are analyzed by a data analysis method. In some embodiments, the data analysis method applies one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters that describe the curvature of one or more traces, determination of a sharp change in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantification analysis parameters, determination of one or more cross recurrence quantification analysis parameters, determination of one or more joint recurrence quantification analysis parameters, determination of one or more multi-dimensional recurrence quantification analysis parameters, estimation of a Lyapunov spectrum, or determination of a maximum Lyapunov exponent.
[0076] In some embodiments, the data analysis further includes normalizing each trace relative to an internal standard. For example, a measured substance detected in a sample that was evenly incorporated during the development / growth of a biological sample that does not vary with environmental exposure (e.g., diet) can serve as an internal standard.
[0077] In some embodiments, the data analysis further includes performing recurrence quantification analysis (RQA) on the time-dependent trace to obtain a set of features that describe the dynamic periodic characteristics of the trace. RQA measures the variability of the time-dependent trace. RQA involves the estimation of features that describe the periodic characteristics in a given waveform, including determinism, entropy, mean diagonal line length (MDL), laminarity, entropy, trapping time (TT), recurrence time (RT), Vmax, and Lmax, and each of these features captures various aspects such as signal dynamics, determination of linear gradients, determination of multiple non-linear parameters that describe the curvature of one or more traces, determination of abrupt changes in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of changes in the frequency domain representation of one or more traces, determination of changes in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantification analysis parameters, determination of one or more cross recurrence quantification analysis parameters, determination of one or more joint recurrence quantification analysis parameters, determination of one or more multi-dimensional recurrence quantification analysis parameters, estimation of the Lyapunov spectrum, or determination of the maximum Lyapunov exponent, as described in the attached references. The methods and features of RQA are described, for example, in “Simpler Methods Do It Better: Success of Recurrence Quantification Analysis as a General Purpose Data Analysis Tool,” Physics Letters A 373, 3753-3756 (2009) by Webber et al. and “Recurrence Plots for the Analysis of Complex Systems,” Physics Reports 438, 237-239 (2007) by Marwan et al., the contents of each of which are hereby incorporated by reference in their entirety. In some embodiments, the time-dependent trace is analyzed by using other analysis methods such as Fourier transform, wavelet analysis, and cosinol analysis. Applying such techniques can derive similar metrics including spectral analysis of frequency components and the power associated with them.These metrics and related derived measures can be used instead of features derived from RQA for the purpose of predictive classification to analyze time-dependent traces obtained from biological samples.
[0078] RQA involves constructing recurrence plots to visualize and analyze the dynamic temporal structure of each acquired trace. Such recurrence plots can show the reciprocal process in continuous measurements by plotting a given sequence against the time lag of that sequence. From a one-dimensional trace measured from a hair shaft, additional dimensions are computationally derived and the trace is embedded in a high-dimensional space called a phase space, where t refers to the value of the original trace and dimensions (t + τ) and (t + 2τ) are derived from delaying the original time series by an interval τ. Subsequent analysis is then performed on the embedded phase space to construct recurrence plots and conduct recurrence quantification analysis. The recurrence plot can be derived from the phase space by applying a threshold function to each point in the phase space. Typically, on the corresponding recurrence plot consisting of a square binary matrix represented as white or black space, a value of 1 is assigned to a given point at each time interval, and another point in the phase space shares the spatial limit of the assigned threshold boundary. The RQA method is applied to the recurrence plot to examine the interval of delay between states in a given system, and the black points reflect the time intervals when the system revisits the same state. A periodic process in which the system continuously repeats a given pattern of states appears as a black diagonal in the recurrence plot, while the periods of stability appear as square structures, spurious repetitions appear as black points, and unique events appear as white space.
[0079] In some embodiments, recurrence plots are constructed for traces of a single substance or a combination of two substances (e.g., constructed to visualize the interactive periodic pattern of two substances. This can be referred to as cross-recurrence quantification analysis, or joint recurrence quantification analysis). In some embodiments, recurrence plots are constructed for combinations of three or more substances.
[0080] In some embodiments, the data analysis includes analyzing recurrence plots to obtain a set of features associated with the recurrence plots. Features, which may be referred to synonymously with "rhythmic features" or "dynamic features", provide a quantitative measure that describes the periodicity, predictability, and transitionality present in multiple traces. The features are selected from a set including recurrence rate, determinism, mean diagonal line length, maximum diagonal line length, divergence, Shannon entropy in diagonal line length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, mean diagonal line length (MDL), laminarity, entropy, trapping time (TT), recurrence time (RT), Vmax, Lmax, and any combination thereof.
[0081] In some embodiments, the data analysis method may apply one or more of the following operations and / or methods to one or more traces: determination of a linear gradient, determination of a plurality of non-linear parameters that describe the curvature of one or more traces, determination of a sharp change in the intensity of one or more traces, determination of one or more changes in the baseline intensity of one or more traces, determination of a change in the frequency domain representation of one or more traces, determination of a change in the power spectrum domain representation of one or more traces, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, or determination of the maximum Lyapunov exponent.
[0082] In some embodiments, the data analysis further includes inputting the obtained set of features into a trained model. In some embodiments, the trained model includes a predictive calculation algorithm for obtaining the probability of a subject having a biological state. In some embodiments, the predictive calculation algorithm performs the following calculations.
Number
[0083] The weight parameters β1, …, β k can be defined based on model training. The probability p(target) may be provided as a number in the range of 0 to 1, where 1 corresponds to a 100% probability that the target has a biological state.
[0084] In some embodiments, data analysis includes applying a threshold to the obtained probability p(target). If the obtained probability p(target) exceeds the threshold, the target is evaluated as having a biological state. If the obtained probability is below the threshold, the target is evaluated as not having a biological state. In some embodiments, the threshold is between about 0.3 and 0.6 (e.g., a given threshold is about 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, or 0.6). In some embodiments, the threshold is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, or at least 0.7. In some embodiments, the threshold is 0.9 or less, 0.8 or less, 0.7 or less, 0.6 or less, 0.5 or less, or 0.4 or less. In some embodiments, the threshold is within another range starting from 0.1 or more and ending at 0.9 or less. The value assigned to the probabilistic threshold can be established or estimated during the training of the model via the use of a receiver operating characteristic (ROC) chart, and the optimal threshold used corresponds to the value that yields the maximum area under the curve (ROC-AUC). In some embodiments, the obtained probability is represented in terms of the associated odds (e.g., an odds ratio (OR) that can be derived from a probability such as OR = p / (1 - p)). For example, the evaluation includes evaluating the odds that the target has a biological state.
[0085] In some embodiments, data analysis includes distinguishing a first biological state from an alternative state, e.g., a second biological state. In some embodiments, the alternative state is not associated with any known state (e.g., a typically developing state (NT)). In some embodiments, the first biological state is associated with autism spectrum disorder (ASD), and the alternative state is associated with attention deficit / hyperactivity disorder (ADHD). In some embodiments, the alternative state is any other neurodevelopmental state, or a co-diagnosis of two neurodevelopmental states. Thus, data analysis may be capable of distinguishing between two neurodevelopmental states (e.g., distinguishing between autism spectrum disorder and ADHD, or between ASD and a co-morbid (CM) case diagnosed with both ASD and ADHD).
[0086] Healthcare providers such as the patient's physician and treatment team can access patient data (e.g., dynamic biological response data or other health data), and / or predictions or assessments generated from such data. Based on the data analysis results, the healthcare provider can determine clinical decisions or outcomes.
[0087] For example, a physician can, at least in part, instruct a patient to undergo one or more clinical trials at a hospital or other clinical site based on the predicted disease or disorder in the subject. These instructions can be provided when certain predetermined criteria (e.g., a minimum threshold for the likelihood of a disease or disorder) are met.
[0088] Such minimum thresholds are, in some embodiments, for example, at least about 5% likelihood, at least about 10% likelihood, at least about 20% likelihood, at least about 25% likelihood, at least about 30% likelihood, at least about 35% likelihood, at least about 40% likelihood, at least about 45% likelihood, at least about 50% likelihood, at least about 55% likelihood, at least about 60% likelihood, at least about 65% likelihood, at least about 70% likelihood, at least about 75% likelihood, at least about 80% likelihood, at least about 85% likelihood, at least about 90% likelihood, at least about 95% likelihood, at least about 96% likelihood, at least about 97% likelihood, at least about 98% likelihood, or at least about 99% likelihood. In some embodiments, the minimum threshold is 99% or less, 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, or 40% or less. In some embodiments, the minimum threshold is 5% - 20%, 10% - 50%, 30% - 70%, or 60% - 99%. In some embodiments, the minimum threshold is within another range starting at 5% or more and ending at 99% or less.
[0089] As another example, a physician may prescribe a therapeutically effective amount of a therapeutic agent (e.g., a drug), a clinical procedure, or a further clinical trial to be administered to a patient, based at least in part on a predicted disease or disorder in the subject. For example, a physician may prescribe an anti-inflammatory therapeutic agent in response to signs of inflammation in a patient.
[0090] Model The methods and systems of the present disclosure can utilize or access external capabilities of artificial intelligence techniques to develop signatures for various diseases or disorders. Using these signatures, a disease or disorder can be accurately predicted (e.g., months or years earlier than standard clinical care). Using such predictive capabilities, a healthcare provider (e.g., a physician) can make accurate risk-based decisions based on the information, thereby improving the quality of care and monitoring provided to the patient.
[0091] The methods and systems of the present disclosure can analyze dynamic biological response data obtained from a subject (patient) to generate a likelihood that the subject has a disease or disorder. For example, the system may apply a trained (e.g., predictive) algorithm to the obtained dynamic biological response data to generate a likelihood that the subject has a disease or disorder. The trained algorithm may include an artificial intelligence-based model, such as a classifier or a regressor, configured to process the obtained dynamic biological response data to generate a likelihood that the subject has a disease or disorder. The model may be trained using a clinical dataset from one or more cohorts of patients, using, for example, the patient's clinical health data and / or dynamic biological response data as inputs and the patient's known clinical health outcome (e.g., a disease or disorder) as an output to the model.
[0092] The model may include one or more machine learning algorithms. Examples of machine learning algorithms may include support vector machines (SVMs), naive Bayes classification, random forests, neural networks (e.g., deep neural networks (DNNs), recurrent neural networks (RNNs), deep RNNs, long short-term memory (LSTM) recurrent neural networks (RNNs), or gated recurrent units (GRUs)), or other supervised learning algorithms, or unsupervised machine learning, statistical, or deep learning algorithms for classification and regression. The model may similarly include an estimation of an ensemble model consisting of multiple prediction models, for example, in constructing gradient-boosted decision trees, techniques such as gradient boosting may be utilized. The model may be trained using one or more training data sets corresponding to patient data.
[0093] The training data set may be generated, for example, from one or more cohorts of patients having common clinical characteristics (features) and clinical outcomes (labels). The training data set may include a set of features and labels corresponding to the features. The features may correspond to algorithm inputs including dynamic biological response data, patient demographic information derived from electronic medical records (EMRs), and medical observations. The features may include clinical characteristics such as, for example, a particular range or category of dynamic biological response data. The features may include patient information such as the patient's age, the patient's medical history, other medical conditions, current or past medications, and the time since the last observation. For example, a set of features collected from a given patient at a given point in time may act collectively as a signature that can indicate the health state or condition of the patient at the given point in time.
[0094] For example, the ranges of dynamic biological response data and other health measurements can be represented as a plurality of mutually exclusive continuous ranges of continuous measurements, and the categories of dynamic biological response data and other health measurements can be represented as a plurality of mutually exclusive sets of measurements (e.g., {"high", "low"}, {"high", "normal"}, {"low", "normal"}, {"high", "borderline high", "normal", "low"}, etc.). Clinical characteristics can also include clinical markers indicating the patient's health history, such as the diagnosis of a disease or disorder, previous clinical treatments (e.g., drugs, surgical treatments, chemotherapy, radiation therapy, immunotherapy, etc.), behavioral factors, or other health conditions (e.g., a history of hypertension, hyperglycemia, hypercholesterolemia or high blood cholesterol, allergic reactions or other side effects, etc.).
[0095] The markers can include, for example, clinical outcomes such as the presence, absence, diagnosis, or prognosis of a disease or disorder in a subject (e.g., a patient). The clinical outcome can include temporal characteristics associated with the presence, absence, diagnosis, or prognosis of a disease or disorder in the patient. For example, the temporal characteristics can indicate that a disease or disorder occurred within a specific period after the patient had a previous clinical outcome (e.g., was discharged from the hospital, received an administration of a therapeutic agent such as a drug, underwent a clinical procedure such as a surgical operation, etc.). Such a period can be, for example, about 1 hour, about 2 hours, about 3 hours, about 4 hours, about 6 hours, about 8 hours, about 10 hours, about 12 hours, about 14 hours, about 16 hours, about 18 hours, about 20 hours, about 22 hours, about 24 hours, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 10 days, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 6 months, about 8 months, about 10 months, about 1 year, or more than about 1 year. In some embodiments, the period is 5 years or less, 1 year or less, 6 months or less, 3 months or less, 1 month or less, 2 weeks or less, 1 week or less, 1 day or less, or 12 hours or less. In some embodiments, the period is 1 hour to 12 hours, 12 hours to 24 hours, 1 day to 7 days, 1 week to 4 weeks, 1 month to 12 months, or 1 year to 5 years. In some embodiments, the plurality of periods begin at 1 hour or more and end within another range that ends at 5 hours or less.
[0096] The input features can be structured by aggregating data into bins or, alternatively, by using one-hot encoding. The input can also include eigenvalue or vectors derived from the aforementioned inputs, such as cross-correlations calculated between separate dynamic biological response data or other measurements over a fixed period, and finite differences between discrete derivatives or continuous measurements. Such periods can be, for example, about 1 hour, about 2 hours, about 3 hours, about 4 hours, about 6 hours, about 8 hours, about 10 hours, about 12 hours, about 14 hours, about 16 hours, about 18 hours, about 20 hours, about 22 hours, about 24 hours, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 10 days, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 6 months, about 8 months, about 10 months, about 1 year, or more than about 1 year. In some embodiments, the period is 5 years or less, 1 year or less, 6 months or less, 3 months or less, 1 month or less, 2 weeks or less, 1 week or less, 1 day or less, or 12 hours or less. In some embodiments, the period is 1 hour to 12 hours, 12 hours to 24 hours, 1 day to 7 days, 1 week to 4 weeks, 1 month to 12 months, or 1 year to 5 years. In some embodiments, the plurality of periods begins at 1 hour or more and ends within another range that is 5 hours or less.
[0097] The training records can be constructed from a sequence of observations. Such a sequence can include a fixed length to facilitate data processing. For example, the sequence can be zero-padded or selected as an independent subset of a single patient's record.
[0098] The model can process input features to generate output values that include one or more classifications, one or more predictions, or combinations thereof. For example, such classifications or predictions can include binary classifications of healthy / normal health states (e.g., the absence of a disease or disorder) or harmful health states (e.g., the presence of a disease or disorder), classifications among groups of categorical labels (e.g., "no disease or disorder", "apparent disease or disorder", and "suspected disease or disorder"), the likelihood of developing a particular disease or disorder (e.g., relative likelihood or probability), a score indicating the presence of a disease or disorder, a score indicating the level of systemic inflammation experienced by a patient, a "risk factor" indicating the probability of a patient's death, a prediction of the time when a patient is expected to develop a disease or disorder, and confidence intervals for any numerical predictions. Various machine learning techniques can be cascaded such that the output of a machine learning technique can also be used as input features to subsequent layers or subsections of the model.
[0099] The model can be trained using a dataset (e.g., by determining the weights and correlations of the model) to generate real-time classifications or predictions. Such a dataset can be large enough to generate statistically significant classifications or predictions. For example, the dataset can include a database of non-identifying data that includes kinetic biological response data and other measurements, as well as kinetic biological response data and other measurements from a hospital or other clinical setting.
[0100] The dataset can be split into subsets (e.g., individual or overlapping) such as a training dataset, a development dataset, and a test dataset. For example, the dataset can be split into a training dataset that includes 80% of the dataset and a test dataset that includes 20% of the dataset. The training dataset can include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The development dataset can include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The test dataset can include about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. The training set (e.g., the training dataset) can be selected by random sampling of a set of data corresponding to one or more patient cohorts to ensure the independence of sampling. Alternatively, the training set (e.g., the training dataset) can be selected by proportional sampling of a set of data corresponding to one or more patient cohorts to ensure the independence of sampling.
[0101] To improve the accuracy of model prediction and reduce overfitting of the model, the dataset can be augmented to increase the number of samples in the training set. For example, data augmentation can include rearranging the order of observations in the training records. To handle datasets with missing observations, methods for filling in missing data, such as forward fitting, backfitting, linear interpolation, and multi-task Gaussian processes, may be used. The dataset can be filtered to remove confounding factors. For example, within a database, a subset of patients may be excluded.
[0102] The model may include one or more neural networks such as neural networks, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), or deep RNNs. The recurrent neural network may include units that may be long short-term memory (LSTM) units or gated recurrent units (GRUs). For example, the model may include an algorithm architecture that includes a neural network having a set of input features such as vital signs and other measurements, a patient's medical history, and / or a patient's demographics. Neural network techniques, such as dropout or regularization, may be used during training of the model to prevent overfitting. The neural network may include multiple sub-networks, each sub-network being configured to generate a classification or prediction of a different type of output information (e.g., may be combined to form the overall output of the neural network). Alternatively, the model may utilize statistical algorithms or related algorithms including random forests, classification and regression trees, support vector machines, discriminant analysis, regression techniques, and their ensembles and gradient-boosted variations.
[0103] When the model generates a classification or prediction of a disease or disorder, a notification (e.g., an alert or alarm) may be generated and transmitted to a healthcare provider such as a physician, nurse, or other member of the patient's treatment team within a hospital. The notification may be transmitted via an automated phone call, a short message service (SMS) or multimedia message service (MMS) message, an email, or an alert within a dashboard. The notification may include output information such as a prediction of the disease or disorder, the likelihood of the predicted disease or disorder, the time until the expected onset of the disease or disorder, the confidence interval of the likelihood or time, or a recommended course of treatment for the disease or disorder.
[0104] To verify the performance of a model, different performance metrics can be generated. For example, the area under the receiver operating characteristic curve (AUROC) can be used to determine the diagnostic ability of the model. For example, the model can use an adjustable classification threshold, whereby the specificity and sensitivity become adjustable, and the receiver operating characteristic curve (ROC) can be used to identify different operating points corresponding to different values of specificity and sensitivity.
[0105] In some cases, such as when the dataset is not large enough, cross-validation is performed to evaluate the robustness of the model across different training and test datasets.
[0106] To calculate performance metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), AUPRC, AUROC, etc., the following definitions may be used. "False positive" can refer to an outcome where a positive outcome or result is generated incorrectly or prematurely (e.g., before the actual onset of a disease or disorder, or without onset). "True positive" can refer to an outcome where a positive outcome or result is correctly generated when the patient has a disease or disorder (e.g., the patient exhibits symptoms of the disease or disorder, or the patient's record indicates the disease or disorder). "False negative" can refer to an outcome where a negative result or outcome is generated, but the patient has a disease or disorder (e.g., the patient exhibits symptoms of the disease or disorder, or the patient's record indicates the disease or disorder). "True negative" can refer to an outcome where a negative outcome or result is generated (e.g., before the actual onset of a disease or disorder, or in the absence of onset).
[0107] The model can be trained until certain predetermined conditions regarding accuracy or performance are met, such as having a minimum desired value corresponding to the diagnostic accuracy measurement. For example, the diagnostic accuracy measurement can correspond to predicting the likelihood of the occurrence of a disease or disorder in a subject. As another example, the diagnostic accuracy measurement may correspond to predicting the likelihood of deterioration or recurrence of a disease or disorder that the subject has previously been treated for. Examples of diagnostic accuracy measurements can include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, area under the precision-recall curve (AUPRC), and area under the receiver operating characteristic (ROC) curve (AUC) (AUROC) corresponding to the diagnostic accuracy of detecting or predicting a disease or disorder.
[0108] For example, in some embodiments, such a predetermined condition is that the sensitivity of the prediction of a disease or disorder, for example, includes a value of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined condition is that the sensitivity includes a value of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined condition is that the sensitivity includes a value in the range of 50% - 70%, 60% - 80%, 70% - 90%, or 90% - 100%. In some embodiments, the predetermined condition is that the sensitivity is within another range that starts at 50% or more and ends at 100% or less.
[0109] As another example, in some embodiments, such a predetermined condition is that the specificity of the prediction of a disease or disorder is, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined condition is that the specificity is a value including a value of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined condition is that the specificity is a value including a value in the range of 50% - 70%, 60% - 80%, 70% - 90%, or 90% - 100%. In some embodiments, the predetermined condition is that the specificity is a value including a value within another range starting from 50% or more and ending at 100% or less.
[0110] As another example, in some embodiments, such a predetermined condition is that the positive predictive value (PPV) of the prediction of a disease or disorder is, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined condition is that the PPV is a value including a value of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined condition is that the PPV is a value including a value in the range of 50% - 70%, 60% - 80%, 70% - 90%, or 90% - 100%. In some embodiments, the predetermined condition is that the PPV is a value including a value within another range starting from 50% or more and ending at 100% or less.
[0111] As another example, in some embodiments, such a predetermined condition is that the negative predictive value (NPV) of the prediction of a disease or disorder is, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined condition is that the NPV includes a value of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined condition is that the NPV includes a value in the range of 50% - 70%, 60% - 80%, 70% - 90%, or 90% - 100%. In some embodiments, the predetermined condition is that the NPV includes a value within another range starting from 50% or more and ending at 100% or less.
[0112] As another example, in some embodiments, such a predetermined condition is that the area under the curve (AUC) (AUROC) of the receiver operating characteristic (ROC) curve of the prediction of a disease or disorder is at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the predetermined condition is that the AUROC includes a value of 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, or 0.60 or less. In some embodiments, the predetermined condition is that the AUROC includes a value in the range of 0.50 - 0.70, 0.60 - 0.80, 0.70 - 0.90, or 0.90 - 1. In some embodiments, the predetermined condition is that the AUROC includes a value within another range starting from 0.50 or more and ending at 1 or less.
[0113] As another example, in some embodiments, such a predetermined condition is that the area under the precision-recall curve (AUPRC) for predicting a disease or disorder is at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the predetermined condition is that the AUPRC includes a value that is 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, 0.60 or less, or 0.50 or less. In some embodiments, the predetermined condition is that the AUPRC includes a value that is in the range of 0.10 to 0.40, 0.30 to 0.70, 0.60 to 0.90, or 0.80 to 1. In some embodiments, the predetermined condition is that the AUPRC includes a value that is within another range that starts at 0.10 or more and ends at less than 1.
[0114] In some embodiments, the trained model is trained or configured to predict a disease or disorder with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with a sensitivity of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with a sensitivity in the range of 50% - 70%, 60% - 80%, 70% - 90%, or 90% - 100%. In some embodiments, the model is trained or configured to predict a disease or disorder with a sensitivity within another range that starts at 50% or more and ends at 100% or less.
[0115] In some embodiments, the model is trained or configured to predict a disease or disorder with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with a specificity of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with a specificity in the range of 50% - 70%, 60% - 80%, 70% - 90%, or 90% - 100%. In some embodiments, the model is trained or configured to predict a disease or disorder with a specificity within another range that starts at 50% or more and ends at 100% or less.
[0116] In some embodiments, the model is trained or configured to predict a disease or disorder with a positive predictive value (PPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with a PPV of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with a PPV in the range of 50% to 70%, 60% to 80%, 70% to 90%, or 90% to 100%. In some embodiments, the model is trained or configured to predict a disease or disorder with a PPV within another range starting at 50% or more and ending at 100% or less.
[0117] In some embodiments, the model can be trained or configured to predict a disease or disorder with a negative predictive value (NPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with an NPV of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with an NPV in the range of 50% to 70%, 60% to 80%, 70% to 90%, or 90% to 100%. In some embodiments, the model is trained or configured to predict a disease or disorder with an NPV within another range starting at 50% or more and ending at 100% or less.
[0118] In some embodiments, the trained model is trained or configured to predict a disease or disorder with an area under the receiver operating characteristic (ROC) curve (AUC) (AUROC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the trained model is trained or configured to predict a disease or disorder with an AUROC of 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, or 0.60 or less. In some embodiments, the trained model is trained or configured to predict a disease or disorder with an AUROC in the range of 0.50 to 0.70, 0.60 to 0.80, 0.70 to 0.90, or 0.90 to 1. In some embodiments, the trained model is trained or configured to predict a disease or disorder with an AUROC within another range that starts at 0.50 or more and ends at 1 or less.
[0119] In some embodiments, the model may be trained or configured to predict a disease or disorder with an area under the precision-recall curve (AUPRC) of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the model is trained or configured to predict a disease or disorder with an AUPRC of 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, 0.60 or less, or 0.50 or less. In some embodiments, the model is trained or configured to predict a disease or disorder with an AUPRC in the range of 0.10 to 0.40, 0.30 to 0.70, 0.60 to 0.90, or 0.80 to 1. In some embodiments, the model is trained or configured to predict a disease or disorder with an AUPRC within another range starting from 0.10 or more and ending at 1 or less.
[0120] The training dataset can be collected from a training subject (e.g., a human). Each training has a diagnostic status indicating whether the subject has been diagnosed with a biological condition or has not been diagnosed with a biological condition. In some embodiments, the training subject is a pediatric patient 5 years of age or younger (e.g., 5 years, 4 years, 3 years, 2 years, 1 year, 9 months old, 6 months old, 3 months old, or 1 month old or younger). The following training procedure can be performed for each training subject among a plurality of training subjects. In some embodiments, the training subject is 18 years of age or younger, 15 years of age or younger, 12 years of age or younger, 11 years of age or younger, 10 years of age or younger, 9 years of age or younger, 8 years of age or younger, 7 years of age or younger, 5 years of age or younger, 4 years of age or younger, 3 years of age or younger, 2 years of age or younger, or 1 year of age or younger.
[0121] In some embodiments, the training data (e.g., dynamic IHC data) is generated from a biological sample to be trained. For each biological sample, a plurality of positions of a reference line on the biological sample to be trained are sampled to generate measurements therefrom, thereby obtaining a plurality of dynamic biological reaction samples. Each dynamic biological reaction sample in the corresponding plurality of dynamic biological reaction samples corresponds to a different position among the corresponding plurality of positions, and each position among the corresponding plurality of positions represents a different growth period of the corresponding biological sample. Next, each respective position of the biological sample is analyzed (e.g., using a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer) to obtain a plurality of traces. Each trace in the corresponding plurality of traces corresponds to a measurement of the abundance of the corresponding substance, which is determined collectively over time from the corresponding plurality of dynamic biological reaction samples.
[0122] Next, to measure the corresponding set of features, recurrence quantification analysis (RQA) or a related method is applied to the corresponding plurality of traces to obtain respective second data sets, and each respective feature in the corresponding set of features is determined by the variation in the abundance of one or more substances in the corresponding plurality of traces.
[0123] Next, a model that is not trained or partially trained may be generated that has (i) a corresponding set of features of each respective second data set for each of a plurality of training subjects, and (ii) a corresponding diagnostic state for each of the plurality of training subjects, selected from among a first diagnostic state and a second diagnostic state, whereby a trained model is obtained. The trained model provides an indicator as to whether a test subject has a first biological state based on the feature values of a set of features obtained from a biological sample of the test subject. In some embodiments, the trained model is a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering model algorithm, a supervised clustering model algorithm, a regression model, or any combination or variation thereof, and in particular, a gradient boost implementation of the described algorithms, for example, including gradient boosted decision trees. In some embodiments, the trained model predicts a regression for a polynomial or binomial distribution. In some embodiments, the trained model is used to make a binary prediction as to whether a sample was derived from a subject having a first biological state, or a multinomial prediction that distinguishes subjects without a diagnosis from subjects having a first biological state, or a second biological state different from the first biological state.
[0124] In some embodiments, the model is a neural network or a convolutional neural network. See Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408, Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40, and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology. Each of these is hereby incorporated by reference herein.
[0125] SVM is described in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge, Boser et al., 1992, “A training algorithm for optimal margin classifiers,” Proceedings of the 5 thAnnual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152, Vapnik, 1998, Statistical Learning Theory, Wiley, New York, Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265, and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, and Furey et al., 2000, Bioinformatics 16, 906-914, which are hereby incorporated by reference in their entirety. When used for classification, an SVM separates a given set of binary-labeled data from labeled data using a hyperplane that is as far as possible from the data. If linear separation is not possible, the SVM can function in combination with a "kernel" technique that automatically implements a non-linear mapping into the feature space. The hyperplane found by the SVM in the feature space corresponds to a non-linear decision boundary in the input space.
[0126] Decision tree classifiers are outlined by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Decision tree-based methods divide the feature space into a set of rectangles and fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. One particular algorithm that can be used is classification and regression trees (CART). Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forests. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which is incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein in its entirety by reference. Random forests are described in Breiman, 1999, “Random Forests - Random Features,” Technical Report 567, Statistics Department, U.C. Berkeley, September 1999, which is incorporated herein in its entirety by reference.
[0127] Clustering (e.g., unsupervised clustering model algorithms and supervised clustering model algorithms) is described on pages 211 - 256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter "Duda 1973"), which is hereby incorporated by reference in its entirety. As described in Section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groups within a dataset. To identify natural groups, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This metric (similarity measure) is used to ensure that samples within a cluster are more similar to each other than samples in other clusters. Second, a mechanism for partitioning the data into clusters using the similarity measure is determined. Similarity measures are discussed in Section 6.7 of Duda 1973, and one way to initiate a clustering investigation is to define a distance function and compute a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities in the same cluster will be significantly shorter than the distance between reference entities in different clusters. However, as described on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, a non - metric similarity function s(x,x') can be used to compare two vectors x and x'. Conventionally, s(x,x') is a symmetric function that has a large value when x and x' are "similar" in some way. Examples of non - metric similarity functions s(x,x') are provided on page 218 of Duda 1973. Once a method for measuring "similarity" or "dissimilarity" between points in the dataset is selected, clustering requires a criterion function for measuring the clustering quality of any partition of the data. The partition of the dataset that extremizes the criterion function is used to cluster the data. See page 217 of Duda 1973.The reference functions are discussed in Section 6.8 of Duda 1973. More recently, Duda et al., Pattern Classification, 2. nd edition, John Wiley & Sons, Inc., New York was published. It is described in detail about clustering on pages 537 - 563. For details of clustering techniques, see Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, N.Y.; Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, N.Y., and Backer, 1995, Computer - Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey. These are each incorporated herein by reference. Specific exemplary clustering techniques that can be used in the present disclosure include hierarchical clustering (agglomerative clustering using the nearest neighbor algorithm, farthest neighbor algorithm, average linkage algorithm, centroid algorithm, or sum - of - squares algorithm), k - means clustering, fuzzy k - means clustering algorithm, and Jarvis - Patrick clustering, but are not limited thereto. In some embodiments, clustering includes unsupervised clustering, in which case no preconception about which clusters should be formed is imposed when the training set is clustered.
[0128] Regression models such as the multinomial logit model are described in Agresti, An Introduction to Categorical Data Analysis, 1996, John Wiley & Sons, Inc., New York, Chapter 8, which is hereby incorporated by reference in its entirety. In some embodiments, this model utilizes a regressor disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, which is hereby incorporated by reference in its entirety. In some embodiments, gradient boosting models are used, for example, for the classification algorithms described herein, and these gradient boosting models are described in Boehmke, Bradley; Greenwell, Brandon (2019). “Gradient Boosting”. Hands-On Machine Learning with R. Chapman & Hall. pp. 221-245. ISBN 978-1-138-49568-5. This document is hereby incorporated by reference in its entirety. In some embodiments, ensemble modeling techniques are used, for example, for the classification algorithms described herein, and these ensemble modeling techniques are described in the implementation of the classification model herein and are explained in Zhou Zhihua (2012). Ensemble Methods: Foundations and Algorithms. Chapman and Hall / CRC. ISBN 978-1-439-83003-1. This document is hereby incorporated by reference in its entirety.
[0129] In some embodiments, the model is performed by a device that executes one or more programs (e.g., one or more programs stored in the non - persistent memory 111 or the persistent memory 112 of FIG. 1) that include instructions for performing data analysis. In some embodiments, the data analysis is performed by a system that includes at least one processor (e.g., processing core 102) and a memory (e.g., one or more programs stored in the non - persistent memory 111 or the persistent memory 112) that includes instructions for performing data analysis.
[0130] Computer system The present disclosure provides a computer system programmed to implement the methods of the present disclosure. FIG. 4 shows a computer system 401 programmed or otherwise configured to perform, for example, staining a tooth sample, acquiring a fluorescent image of the stained tooth sample, spatially analyzing the fluorescence intensity across the stained tooth sample, generating a temporal profile of inflammation, processing data using a trained model, and determining the risk of a disease or disorder of a subject. The computer system 401 can regulate various aspects of the sensor data analysis of the present disclosure, such as, for example, staining a tooth sample, acquiring a fluorescent image of the stained tooth sample, spatially analyzing the fluorescence intensity across the stained tooth sample, generating a temporal profile of inflammation, measuring the dynamics of the temporal profile, processing process data using a trained model, and predicting the diagnostic state of a subject with respect to a disease or disorder. The computer system 401 can be a user's electronic device or a computer system located remotely with respect to the electronic device. The electronic device can be a mobile electronic device.
[0131] The computer system 401 includes a central processing unit (CPU, also including "processor" and "computer processor" herein) 405, and the central processing unit 405 can be a single-core or multi-core processor, or a plurality of processors for parallel processing. The computer system 401 also includes a memory or memory location 410 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 415 (e.g., hard disk), a communication interface 420 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 425 such as cache, other memory, data storage and / or an electronic display adapter. The memory 410, the storage unit 415, the interface 420, and the peripheral devices 425 communicate with the CPU 405 through a communication bus (solid line) such as a motherboard. The storage unit 415 can be a data storage unit (or data repository) for storing data. The computer system 401 can be operably coupled to a computer network ("network") 430 with the help of the communication interface 420. The network 430 can be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet communicating with the Internet. In some cases, the network 430 is a telecommunications and / or data network. The network 430 can include one or more computer servers, and the computer servers can enable distributed computing such as cloud computing. The network 430 can, in some cases, implement a peer-to-peer network with the help of the computer system 401, and the peer-to-peer network can enable devices coupled to the computer system 401 to operate as clients or servers.
[0132] The CPU 405 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location such as the memory 410. The instructions can direct the CPU 405, and the CPU 405 can then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 405 can include fetch, decode, execute, and write-back.
[0133] The CPU 405 can be part of a circuit such as an integrated circuit. One or more other components of the system 401 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0134] The storage unit 415 can store files such as drivers, libraries, and saved programs. The storage unit 415 can store user data, such as user preferences and user programs. In some cases, the computer system 401 can include one or more additional data storage units external to the computer system 401, such as located on a remote server that communicates with the computer system 401 through an intranet or the Internet.
[0135] The computer system 401 can communicate with one or more remote computer systems through the network 430. For example, the computer system 401 can communicate with a remote computer system of a user (e.g., a medical provider). Examples of remote computer systems include personal computers (e.g., portable PCs), slates or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-compatible devices, Blackberry®), or personal digital assistants. A user can access the computer system 401 via the network 430.
[0136] The methods described herein can be implemented, for example, by machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 401, such as on the memory 410 or the electronic memory unit 415. The machine executable code or machine readable code can be provided in the form of software. In use, the code can be executed by the processor 405. In some cases, the code can be retrieved from the memory unit 415 and stored in the memory 410 for immediate access by the processor 405. In some situations, the electronic memory unit 415 can be excluded and the machine executable instructions can be stored in the memory 410.
[0137] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled during runtime. The code can be provided in a programming language that can be selected to enable the code to be executed in a pre-compiled manner or in a just-in-time compiled manner.
[0138] Aspects of the systems and methods provided herein, such as computer system 401, may be embodied in programming. Various aspects of the technology may typically be considered a "product" or "article of manufacture" in the form of machine (or processor) executable code and / or associated data carried or embodied on a type of machine-readable medium. The machine executable code may be stored in an electronic memory unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of the tangible memory of a computer, processor, etc., or associated modules of tangible memory such as various semiconductor memories, tape drives, disk drives, etc., and these media can provide non-transitory storage at any time for software programming. All or part of the software may be communicated over the Internet or various other electrical communication networks. Such communication can enable, for example, the loading of software from one computer or processor to another, such as from an administrative server or host computer to an application server computer platform. Thus, another type of medium that can carry software elements is used over physical interfaces between local devices, via wired and optical terrestrial communication networks, and via various air links, including light, electrical, and electromagnetic waves. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that carry software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable media" refer to any medium involved in providing instructions to a processor for execution.
[0139] Thus, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks such as any of the storage devices in any computer(s) that can be used to implement, for example, the databases shown in the drawings. Volatile storage media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables, copper wire, and fiber optics including the wires that make up a bus within a computer system. Carrier wave transmission media can take the form of electrical signals or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM, DVD, or DVD-ROM, any other optical media, punch cards, paper tapes, any other physical storage media with patterns of holes, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave carrying data or instructions, a cable or link carrying such a carrier wave, or any other media that a computer can read programming code and / or data from. Many of these forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0140] Computer system 401 can include, or be communicable with, an electronic display 435 that includes, for example, a user interface (UI) 440 that provides fluorescence image data, fluorescence intensity data, temporal profiles of inflammation, and machine learning classifications. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0141] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit 405. The algorithms can, for example, stain a tooth sample, acquire a fluorescence image of the stained tooth sample, spatially analyze the fluorescence intensity across the stained tooth sample, generate a temporal profile of inflammation, process data using a trained model, and determine the risk of a disease or disorder of a subject.
[0142] The methods described elsewhere in this specification show steps or sets of operations according to embodiments, but those skilled in the art will recognize many variations based on the teachings described herein. These steps can be completed in a different order. Steps can be added or omitted. Some steps can include sub-steps. Many steps can be repeated as many times as beneficial.
[0143] One or more steps of each method or set of operations can be performed using one or more of the circuits as described herein, such as a processor or programmable array logic for a field programmable gate array, for example. For example, the circuit can be programmed to provide one or more of each step of the method or set of operations, and the program can include program instructions stored on a computer-readable memory or programmed steps of a logic circuit such as programmable array logic or a field programmable gate array.
Example
[0144] Example 1: Dynamic Molecular Profile in a Tooth Sample for Determining Disease Risk Using the methods and systems of the present disclosure, molecular profiles in dental samples were generated and then analyzed to determine disease risk in a subject. Generally, it has been found that the temporal dynamics of biological reactions (e.g., inflammation) are imprinted in the sample (e.g., dental sample). By analyzing this temporal dynamics, the disease risk of the subject can be determined. A dynamic molecular profile of C-reactive protein (CRP), a marker of inflammation, was generated. Using dental biomarkers, two sets of children, a first set (37 cases) with autism spectrum disorder and a second set (77 controls) without autism spectrum disorder, had dynamic time-series profiles of CRP and inflammation generated during a period including fetal (prenatal) development and infancy. The time-series CRP profiles were analyzed to reveal novel features of the dynamics of the CRP signal, thereby accurately distinguishing autism cases from controls. For example, the inflammatory profiles that were present before the age of 1 year showed large differences between cases and controls. In contrast, clinical diagnosis of autism is usually determined around the age of 3 - 4 years.
[0145] Primary dental samples were obtained from each pediatric subject. The dental samples were sectioned, decalcified, and an immunohistochemical stain (e.g., dentin) was applied to the dental samples. Immunohistochemical staining effectively mapped C-reactive protein (a molecular marker of inflammation) along the growth rings of the dental samples, developing a temporal profile of inflammation over the prenatal and postnatal periods. The temporal profile was analyzed using the machine learning algorithm of the present disclosure to train a very accurate model for determining disease risk (e.g., autism).
[0146] Figure 5 shows an example of the daily C-reactive protein profile of a subject over time, where the y-axis indicates CRP intensity and the x-axis indicates developmental age. The developmental age of the pediatric subject includes the period from the second trimester of pregnancy (e.g., starting 140 days before birth when the subject was at the prenatal stage) to about 6 months after birth. As shown in Figure 5, the inflammatory (indicated by CRP intensity) profile in cases of autistic children was observed to be high prenatally.
[0147] Figures 6A-6B show receiver operating characteristic (ROC) curves for characterizing the sensitivity and specificity of a method for diagnosing autism at various prediction thresholds using features derived from a recurrence quantification analysis of C-reactive protein profiles sampled during prenatal and early childhood (e.g., up to 1 year of age). Figure 6A shows an experimental receiver operating characteristic (ROC) curve for evaluating the accuracy of the disclosed method for assessing subjects with autism spectrum disorder. ROC curves can be used to evaluate the performance of binary classifiers. The ROC curve is plotted as sensitivity (also called true positive rate) against specificity (also called true negative rate). A perfect classifier can have 100% sensitivity and 100% specificity, and an area under the curve (AUC) of 1.0. As shown in Figure 6A, a classifier configured to determine the presence of autism in a subject based on a dynamic C-reactive protein profile had an area under the receiver operating characteristic (ROC) curve (AUC) of 0.86, with a 95% confidence interval (CI) of 0.72-1.00. The receiver operating characteristic (ROC) shows how the sensitivity and specificity values of a classifier change as the threshold applied to the predicted probability of the case state increases or decreases. As the threshold decreases, for example, a more sensitive classification is obtained, but the specificity decreases accordingly. As shown in Figure 6B, the main dynamic features contributing to classifier performance ranked in descending order of feature importance (e.g., as indicated by numerical feature weighting) include laminarity, entropy, TT, MDL, RT1, RT2, Vmax, determinism, and Lmax. For example, laminarity was determined to have a higher feature importance than the others.
[0148] Accordingly, in the analysis of the features obtained from the analysis of the C-reactive protein profile using the methods and systems of the present disclosure, only the features derived from the analysis of the C-reactive protein signature measured on a biological sample (e.g., a dental sample) non-invasively obtained from a pediatric subject were used, and the disease risk of autism with an AUC of 0.86 was successfully determined. These results demonstrate that the dynamics of the initial inflammatory response are later associated with the disease and can be accurately detected and profiled using the methods and systems of the present disclosure.
[0149] Preferred embodiments of the invention are shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited by the specific embodiments provided within the specification. The invention has been described with reference to the foregoing specification, but the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, modifications, and substitutions will occur to those skilled in the art without departing from the invention. Further, it should be understood that all aspects of the invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be used in the practice of the invention. Accordingly, the invention is intended to cover such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be encompassed thereby.
[0150] Embodiments Embodiment 1. A method for predicting the diagnostic status of a subject with respect to a disease or disorder, comprising: (a) staining a dental sample of the subject to produce a stained dental sample; (b) spatially analyzing fluorescence intensity across the stained dental sample; and (c) predicting the diagnostic status of the subject with respect to the disease or disorder, at least in part based on the analysis of the fluorescence intensity. Embodiment 2. The method according to embodiment 1, wherein analyzing comprises obtaining a fluorescence image of the stained tooth sample and analyzing the fluorescence intensity of the fluorescence image. Embodiment 3. The method according to embodiment 2, wherein obtaining a fluorescence image of the stained tooth sample comprises using a confocal microscope, either inverted or non-inverted. Embodiment 4. The method according to any one of embodiments 1 to 3, wherein staining the tooth sample comprises using C-reactive protein immunohistochemical staining. Embodiment 5. The method according to any one of embodiments 1 to 4, further comprising sectioning the tooth sample. Embodiment 6. The method according to any one of embodiments 1 to 5, wherein staining the tooth sample comprises decalcifying the tooth sample. Embodiment 7. The method according to any one of embodiments 1 to 6, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. Embodiment 8. The method according to any one of embodiments 1 to 6, wherein the disease or disorder comprises autism spectrum disorder. Embodiment 9. The method according to any one of embodiments 1 to 8, wherein the subject is a human. Embodiment 10. The method according to embodiment 9, wherein the subject is less than 12 years old. Embodiment 11. The method according to embodiment 9, wherein the subject is less than 1 year old. Embodiment 12. The method according to embodiment 1, wherein analyzing comprises generating a temporal profile of inflammation based at least in part on the fluorescence intensity and analyzing the temporal profile of inflammation. Embodiment 13. The method according to embodiment 12, wherein at least a portion of the temporal profile of inflammation corresponds to the prenatal period of the subject. Embodiment 14. The method according to embodiment 1, wherein predicting the diagnostic state of the subject with respect to the disease or disorder comprises processing the fluorescence intensity using a trained model. Embodiment 15. The method according to embodiment 14, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. Embodiment 16. The method according to embodiment 14, wherein the trained model includes a gradient-boosted decision tree. Embodiment 17. The method according to embodiment 14, wherein the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, layering, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. Embodiment 18. The method according to embodiment 17, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, layering, trapping time (TT), maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. Embodiment 19. The method according to embodiment 18, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% sensitivity. Embodiment 20. The method according to embodiment 18, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% specificity. Embodiment 21. The method according to embodiment 18, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% positive predictive value. Embodiment 22. The method according to embodiment 18, wherein the trained model predicts a diagnostic state regarding a disease or disorder with a negative predictive value of at least about 80%. Embodiment 23. The method according to embodiment 18, wherein the trained model predicts a diagnostic state regarding a disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.80. Embodiment 24. A device comprising one or more processors and a memory storing one or more programs for execution by the one or more processors, wherein the one or more programs are configured to: (a) sample each respective position among a plurality of positions along a reference line on a biological sample of a subject associated with a c-reactive protein of the subject, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement among the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with the c-reactive protein; (b) analyze each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (c) derive a respective second data set from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by continuous variation of the c-reactive protein fluorescence intensity; and (d) use the trained model to process the features to predict a diagnostic state of the subject regarding a disease or disorder associated with the c-reactive protein. Embodiment 25. The device according to embodiment 24, wherein the plurality of fluorescence intensity measurements are measured using a confocal microscope, either inverted or non-inverted. Embodiment 26. The device according to embodiment 24 or 25, wherein the biological sample includes a dental sample. Embodiment 27. The device according to any one of embodiments 24 to 26, wherein the dental sample is stained using c-reactive protein immunohistochemical staining. Embodiment 28. The device according to embodiment 26, wherein the instructions further include slicing the dental sample. Embodiment 29. The device according to embodiment 26, wherein the command further comprises decalcifying the tooth sample. Embodiment 30. The device according to any one of embodiments 24-29, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. Embodiment 31. The device according to any one of embodiments 24-29, wherein the disease or disorder comprises autism spectrum disorder ASD. Embodiment 32. The device according to any one of embodiments 24-31, wherein the subject is a human. Embodiment 33. The device according to any one of embodiments 24-32, wherein the subject is less than 12 years old. Embodiment 34. The device according to any one of embodiments 24-32, wherein the subject is less than 1 year old. Embodiment 35. The device according to any one of embodiments 24-34, wherein analyzing comprises generating a temporal profile of inflammation based at least in part on a plurality of fluorescence intensity measurements and analyzing the temporal profile of inflammation. Embodiment 36. The device according to embodiment 35, wherein at least a portion of the temporal profile of inflammation corresponds to the prenatal period of the subject. Embodiment 37. The device according to any one of embodiments 24-36, wherein predicting the diagnostic state of the subject with respect to the disease or disorder comprises processing a plurality of fluorescence intensity measurements using a trained model. Embodiment 38. The device according to embodiment 37, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. Embodiment 39. The device according to embodiment 37, wherein the trained model comprises a gradient-boosted decision tree. Embodiment 40. The device according to Embodiment 12, wherein the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, laminarity, trapping time, maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the temporal profile, determination of the curvature of the temporal profile, determination of the abrupt change in the intensity of the temporal profile, determination of one or more changes in the baseline intensity of the temporal profile, determination of the change in the frequency domain representation of the temporal profile, determination of the change in the power spectrum domain representation of the temporal profile, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof. Embodiment 41. The device according to Embodiment 12, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, layering, trapping time (TT), maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, the number of most likely recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the temporal profile, determination of the curvature of the temporal profile, determination of the abrupt change in the intensity of the temporal profile, determination of one or more changes in the baseline intensity of the temporal profile, determination of the change in the frequency domain representation of the temporal profile, determination of the change in the power spectrum domain representation of the temporal profile, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof. Embodiment 42. The device according to any one of Embodiments 24 to 41, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% sensitivity. Embodiment 43. The device according to any one of Embodiments 24 to 41, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% specificity. Embodiment 44. The device according to any one of Embodiments 24 to 41, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% positive predictive value. Embodiment 45. The device according to any one of Embodiments 24 to 41, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% negative predictive value. Embodiment 46. The device according to any one of Embodiments 24 to 41, wherein the trained model predicts a diagnostic state regarding a disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.80. Embodiment 47. A non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium, wherein when the one or more computer programs are executed by a computer system, the computer system is caused to: (a) sample each respective position at a plurality of positions along a reference line on a biological sample of a subject associated with a c-reactive protein of the subject, thereby obtaining a plurality of fluorescence intensity measurement values, wherein each fluorescence intensity measurement value in the plurality of fluorescence intensity measurement values corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with the c-reactive protein; (b) analyze each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (c) derive a respective second data set from the corresponding plurality of fluorescence intensity measurement values, wherein each respective feature in the corresponding set of features is determined by the variation in c-reactive protein fluorescence intensity; and (d) use a trained model to process the features to predict a diagnostic state of the subject regarding a disease or disorder associated with the c-reactive protein. A non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium, comprising instructions for causing the method to be performed. Embodiment 48. The non-transitory computer-readable storage medium according to Embodiment 47, wherein the plurality of fluorescence intensity measurement values are measured using a confocal microscope, either inverted or non-inverted. Embodiment 49. The non-transitory computer-readable storage medium according to Embodiment 47 or 48, wherein the biological sample includes a dental sample. Embodiment 50. The non-transitory computer-readable storage medium according to Embodiment 49, wherein the dental sample is stained using C-reactive protein immunohistochemical staining. The non - transitory computer - readable storage medium according to embodiment 49, wherein the method further includes slicing a dental sample. The non - transitory computer - readable storage medium according to any one of embodiments 47 to 51, wherein the method further includes decalcifying a dental sample. The non - transitory computer - readable storage medium according to any one of embodiments 47 to 52, wherein the disease or disorder includes autism spectrum disorder (ASD), attention - deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. The non - transitory computer - readable storage medium according to any one of embodiments 47 to 52, wherein the disease or disorder includes autism spectrum disorder (ASD). The non - transitory computer - readable storage medium according to any one of embodiments 47 to 54, wherein the subject is a human. The non - transitory computer - readable storage medium according to any one of embodiments 47 to 55, wherein the subject is less than 12 years old. The non - transitory computer - readable storage medium according to any one of embodiments 47 to 55, wherein the subject is less than 1 year old. The non - transitory computer - readable storage medium according to any one of embodiments 47 to 57, wherein analyzing includes generating a temporal profile of inflammation based at least in part on a plurality of fluorescence intensity measurements and analyzing the temporal profile of inflammation. The non - transitory computer - readable storage medium according to embodiment 58, wherein at least a portion of the temporal profile of inflammation corresponds to the prenatal period of the subject. The non - transitory computer - readable storage medium according to any one of embodiments 47 to 59, wherein predicting the diagnostic status of the subject with respect to a disease or disorder includes processing a plurality of fluorescence intensity measurements using a trained model. Embodiment 61. The non-transitory computer-readable storage medium according to Embodiment 60, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. Embodiment 62. The non-transitory computer-readable storage medium according to Embodiment 60, wherein the trained model includes a gradient-boosted decision tree. Embodiment 63. The non-transitory computer-readable storage medium according to Embodiment 60, wherein the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, the number of most likely recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of fluorescence intensity across a reference line, determination of a plurality of non-linear parameters describing the curvature of the fluorescence intensity across a reference line, determination of a sharp change in intensity of the fluorescence intensity across a reference line, determination of one or more changes in the baseline intensity of the fluorescence intensity across a reference line, determination of a change in the frequency domain representation of the fluorescence intensity across a reference line, determination of a change in the power spectrum domain representation of the fluorescence intensity across a reference line, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross-recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof. Embodiment 64. The non-transitory computer-readable storage medium according to Embodiment 60, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, layering, trapping time (TT), maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the fluorescence intensity across the reference line, determination of a plurality of non-linear parameters describing the curvature of the fluorescence intensity across the reference line, determination of a sharp change in the intensity of the fluorescence intensity across the reference line, determination of one or more changes in the baseline intensity of the fluorescence intensity across the reference line, determination of a change in the frequency domain representation of the fluorescence intensity across the reference line, determination of a change in the power spectrum domain representation of the fluorescence intensity across the reference line, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross-recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof. Embodiment 65. The non-transitory computer-readable storage medium according to any one of Embodiments 47 to 64, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% sensitivity. Embodiment 66. The non-transitory computer-readable storage medium according to any one of Embodiments 47 to 64, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% specificity. Embodiment 67. The non-transitory computer-readable storage medium according to Embodiment 47, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% positive predictive value. Embodiment 68. The non-transitory computer-readable storage medium according to Embodiment 47, wherein the trained model predicts a diagnostic state regarding a disease or disorder with at least about 80% negative predictive value. The non-transitory computer-readable storage medium of embodiment 47, wherein the trained model predicts a diagnostic state regarding a disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.80. Embodiment 70. A method for training a model, in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, comprising: (a) for each respective training subject among a plurality of training subjects, a first diagnostic state corresponding to the first subset of training subjects among the plurality of training subjects having a first biological state associated with a c-reactive protein, and a second diagnostic state corresponding to the second subset of training subjects among the plurality of training subjects not having the first biological state associated with the c-reactive protein, (i) sampling each respective position among a plurality of positions along a reference line on a biological sample of the subject associated with the subject's c-reactive protein, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement among the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with the c-reactive protein; (ii) analyzing each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (iii) deriving a respective second data set from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by a continuous variation in the c-reactive protein fluorescence intensity; (b) training a model that is not trained or partially trained using (i) the corresponding set of features of each respective second data set of each training subject among the plurality of training subjects, and (ii) the corresponding diagnostic state of each training subject among the plurality of training subjects selected from among the first diagnostic state and the second diagnostic state, thereby obtaining a trained model that provides an indication as to whether a test subject has a first biological state associated with a c-reactive protein based on the values of the features in the set of features obtained from a biological sample associated with the test subject's c-reactive protein. Embodiment 71. The method according to embodiment 70, wherein the trained model is a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering model algorithm, a supervised clustering model algorithm, a regression model, a gradient boosting algorithm, or any combination thereof. Embodiment 72. The method according to embodiment 70, wherein the trained model is a multinomial classifier. Embodiment 73. The method according to embodiment 70, wherein the trained model is a binary classifier. Embodiment 74. The method according to any one of embodiments 70 to 73, wherein the first biological state associated with the c-reactive protein is selected from the group consisting of autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer. Embodiment 75. The method according to embodiment 70, further comprising evaluating a test subject for a first biological state associated with a c-reactive protein by distinguishing the first biological state associated with the c-reactive protein from a second biological state associated with a c-reactive protein that is different from the first biological state associated with the metal metabolism. Embodiment 76. The method according to embodiment 75, wherein the first biological state is autism spectrum disorder and the second biological state is attention deficit / hyperactivity disorder. Embodiment 77. The method according to any one of embodiments 70 to 76, wherein the test subject is a human. Embodiment 78. The method according to embodiment 77, wherein the human is less than 12 years old. Embodiment 79. The method according to embodiment 78, wherein the human is less than 1 year old. Embodiment 80. The method according to any one of embodiments 70 to 79, wherein the corresponding biological sample associated with the c-reactive protein of each training subject is selected from the group consisting of hair shafts, teeth, and nails. Embodiment 81. The method according to Embodiment 80, wherein the corresponding biological sample associated with the c-reactive protein of each training subject is a hair shaft, and the reference line corresponds to the longitudinal direction of the hair shaft. Embodiment 82. The method according to any one of Embodiments 70 to 79, wherein the corresponding biological sample associated with the c-reactive protein of each training subject is a tooth, and the reference line corresponds to the direction across the growth zone including the neonatal line of the tooth. Embodiment 83. The method according to any one of Embodiments 70 to 82, wherein the plurality of corresponding positions are arranged such that the first position at each of the plurality of corresponding positions along the corresponding biological sample associated with the c-reactive protein of each training subject corresponds to the position closest to the tip of the corresponding biological sample associated with the c-reactive protein of each training subject. Embodiment 84. The method according to any one of Embodiments 70 to 79, wherein each trace in the plurality of corresponding fluorescence intensity measurement values includes a plurality of data points, and each data point is an instance of each position at the plurality of positions. Embodiment 85. The method according to any one of Embodiments 70 to 84, wherein the corresponding set of features is selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, layering, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. Embodiment 86. The method according to any one of Embodiments 70 to 85, wherein the plurality of corresponding positions includes at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000, 6000, 7000, 8000, 9000, 10000, 12000, 14000, 16000, 18000, 20000, or more than 20000 positions. Embodiment 87. The method according to any one of Embodiments 70 to 86, wherein the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, layering, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of a plurality of fluorescence intensity measurement values, determination of a plurality of non-linear parameters describing the curvature of a plurality of fluorescence intensity measurement values, determination of a rapid change in intensity of a plurality of fluorescence intensity measurement values, determination of one or more changes in the baseline intensity of a plurality of fluorescence intensity measurement values, determination of a change in the frequency domain representation of a plurality of fluorescence intensity measurement values, determination of a change in the power spectrum domain representation of a plurality of fluorescence intensity measurement values, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof. Embodiment 88. The method according to any one of Embodiments 70 to 86, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, laminarity, trapping time, maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, most probable number of recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of a plurality of fluorescence intensity measurements, determination of a plurality of non-linear parameters describing the curvature of a plurality of fluorescence intensity measurements, determination of a sudden change in intensity of a plurality of fluorescence intensity measurements, determination of one or more changes in the baseline intensity of a plurality of fluorescence intensity measurements, determination of a change in the frequency domain representation of a plurality of fluorescence intensity measurements, determination of a change in the power spectrum domain representation of a plurality of fluorescence intensity measurements, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross-recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.
Claims
1. A method for predicting a diagnostic state of a subject with respect to a disease or disorder, comprising: (a) staining a dental sample of the subject to produce a stained dental sample; (b) spatially analyzing fluorescence intensity across the stained dental sample; and (c) predicting a diagnostic state of the subject with respect to the disease or disorder, based at least in part on the analysis of the fluorescence intensity.
2. The method of claim 1, wherein the analyzing comprises obtaining a fluorescence image of the stained dental sample and analyzing the fluorescence intensity of the fluorescence image.
3. The method of claim 2, wherein obtaining the fluorescence image of the stained dental sample comprises using a confocal microscope, either inverted or non-inverted.
4. The method according to any one of claims 1 to 3, wherein staining the dental sample comprises using C-reactive protein immunohistochemical staining.
5. The method according to any one of claims 1 to 4, further comprising sectioning the dental sample.
6. The method according to any one of claims 1 to 5, wherein staining the dental sample comprises decalcifying the dental sample.
7. The method according to any one of claims 1 to 6, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof.
8. The method according to any one of claims 1 to 6, wherein the disease or disorder comprises autism spectrum disorder.
9. The method according to any one of claims 1 to 8, wherein the subject is a human.
10. The method of claim 9, wherein the subject is less than 12 years old.
11. The method of claim 9, wherein the subject is less than 1 year old.
12. The method according to any one of claims 1 to 11, wherein the analyzing comprises generating a temporal profile of inflammation based at least in part on the fluorescence intensity and analyzing the temporal profile of inflammation.
13. The method of claim 12, wherein at least a portion of the temporal profile of inflammation corresponds to a prenatal period of the subject.
14. The method according to any one of claims 1 to 13, wherein predicting the diagnostic state of a subject regarding the disease or disorder includes processing the fluorescence intensity using a trained model.
15. The method according to claim 14, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof.
16. The method according to claim 14, wherein the trained model includes a gradient-boosted decision tree.
17. The trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the temporal profile, determination of a plurality of non-linear parameters describing the curvature of the temporal profile, determination of a sharp change in the intensity of the temporal profile, determination of one or more changes in the baseline intensity of the temporal profile, determination of a change in the frequency domain representation of the temporal profile, determination of a change in the power spectrum domain representation of the temporal profile, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof. The method according to claim 14.
18. The method according to claim 17, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, layering, trapping time (TT), maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the temporal profile, determination of a plurality of non-linear parameters describing the curvature of the temporal profile, determination of a sharp change in the intensity of the temporal profile, determination of one or more changes in the baseline intensity of the temporal profile, determination of a change in the frequency domain representation of the temporal profile, determination of a change in the power spectrum domain representation of the temporal profile, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.
19. The method according to claim 18, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% sensitivity.
20. The method according to claim 18, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% specificity.
21. The method according to claim 18, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% positive predictive value.
22. The method according to claim 18, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% negative predictive value.
23. The method according to claim 18, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 0.80 area under the receiver operating characteristic curve (AUROC).
24. A device comprising one or more processors and a memory storing one or more programs for execution by the one or more processors, wherein the one or more programs are Sampling each respective position at a plurality of positions along a reference line on the biological sample of the subject associated with the c-reactive protein of interest, thereby obtaining a plurality of fluorescence intensity measurements, wherein each fluorescence intensity measurement in the plurality of fluorescence intensity measurements corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with the c-reactive protein, obtaining; Analyzing each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first dataset; Deriving a respective second dataset from the corresponding plurality of fluorescence intensity measurements, wherein each respective feature in the corresponding set of features is determined by continuous variation of the c-reactive protein fluorescence intensity, deriving; Processing the features using a trained model to predict the diagnostic state of the subject with respect to a disease or disorder associated with the c-reactive protein. A device comprising instructions for performing.
25. The device according to claim 24, wherein the plurality of fluorescence intensity measurements are measured using an inverted or non-inverted confocal microscope.
26. The device according to claim 24 or 25, wherein the biological sample comprises a dental sample.
27. The device according to any one of claims 24 to 26, wherein the dental sample is stained using C-reactive protein immunohistochemical staining.
28. The device according to claim 26, wherein the instructions further comprise sectioning the dental sample.
29. The device according to claim 26, wherein the instructions further comprise decalcifying the dental sample.
30. The device according to any one of claims 24 to 29, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof.
31. The device according to any one of claims 24 to 29, wherein the disease or disorder comprises autism spectrum disorder ASD.
32. The device according to any one of claims 24 to 31, wherein the subject is a human.
33. The device according to any one of claims 24 to 32, wherein the subject is under 12 years old.
34. The device according to any one of claims 24 to 32, wherein the subject is under 1 year old.
35. The device according to any one of claims 24 to 34, wherein the analyzing includes generating a temporal profile of inflammation based at least in part on the plurality of fluorescence intensity measurements and analyzing the temporal profile of inflammation.
36. The device according to claim 35, wherein at least a portion of the temporal profile of inflammation corresponds to a prenatal period of the subject.
37. The device according to any one of claims 24 to 36, wherein predicting the diagnostic state of the subject regarding the disease or disorder includes processing the plurality of fluorescence intensity measurements using the trained model.
38. The device according to claim 37, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof.
39. The device according to claim 37, wherein the trained model includes a gradient-boosted decision tree.
40. The device according to claim 37, wherein the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, laminarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the fluorescence intensity across the reference line, determination of a plurality of non-linear parameters describing the curvature of the fluorescence intensity across the reference line, determination of the abrupt change in the intensity of the fluorescence intensity across the reference line, determination of one or more changes in the baseline intensity of the fluorescence intensity across the reference line, determination of the change in the frequency domain representation of the fluorescence intensity across the reference line, determination of the change in the power spectrum domain representation of the fluorescence intensity across the reference line, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.
41. The device according to claim 37, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, layering, trapping time (TT), maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the fluorescence intensity across the reference line, determination of a plurality of non-linear parameters describing the curvature of the fluorescence intensity across the reference line, determination of a sharp change in the intensity of the fluorescence intensity across the reference line, determination of one or more changes in the baseline intensity of the fluorescence intensity across the reference line, determination of a change in the frequency domain representation of the fluorescence intensity across the reference line, determination of a change in the power spectrum domain representation of the fluorescence intensity across the reference line, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross-recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.
42. The device according to any one of claims 24 to 41, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% sensitivity.
43. The device according to any one of claims 24 to 41, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% specificity.
44. The device according to any one of claims 24 to 41, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% positive predictive value.
45. The device according to any one of claims 24 to 41, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% negative predictive value.
46. The device according to any one of claims 24 to 41, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 0.80 area under the receiver operating characteristic curve (AUROC).
47. A non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium, wherein when the one or more computer programs are executed by a computer system, the computer system is caused to (a) sample each respective position at a plurality of positions along a reference line on the subject biological sample associated with the target c-reactive protein, thereby obtaining a plurality of fluorescence intensity measurement values, wherein each fluorescence intensity measurement value in the plurality of fluorescence intensity measurement values corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the subject biological sample associated with the c-reactive protein; (b) analyze each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (c) derive a respective second data set from the corresponding plurality of fluorescence intensity measurement values, wherein each respective feature in the corresponding set of features is determined by the variation in c-reactive protein fluorescence intensity; (d) use a trained model to process the features to predict a diagnostic state of the subject regarding a disease or disorder associated with the c-reactive protein. A non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium, comprising instructions for causing the method to be performed.
48. The non-transitory computer-readable storage medium according to claim 47, wherein the plurality of fluorescence intensity measurement values are measured using a confocal microscope, either inverted or non-inverted.
49. The non-transitory computer-readable storage medium according to claim 47 or 48, wherein the biological sample includes a dental sample.
50. The non-transitory computer-readable storage medium according to claim 49, wherein the dental sample is stained using c-reactive protein immunohistochemical staining.
51. The non-transitory computer-readable storage medium according to claim 49, wherein the method further includes sectioning the dental sample.
52. The non-transitory computer-readable storage medium according to any one of claims 47 to 51, wherein the method further includes decalcifying the dental sample.
53. The non-transitory computer-readable storage medium according to any one of claims 47 to 52, wherein the disease or disorder includes autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof.
54. The non-transitory computer-readable storage medium according to any one of claims 47 to 52, wherein the disease or disorder includes autism spectrum disorder (ASD).
55. The non-transitory computer-readable storage medium according to any one of claims 47 to 54, wherein the subject is a human.
56. The non-transitory computer-readable storage medium according to any one of claims 47 to 55, wherein the subject is less than 12 years old.
57. The non-transitory computer-readable storage medium according to any one of claims 47 to 55, wherein the subject is less than 1 year old.
58. The non-transitory computer-readable storage medium according to any one of claims 47 to 57, wherein analyzing comprises generating a temporal profile of inflammation based at least in part on the plurality of fluorescence intensity measurements and analyzing the temporal profile of inflammation.
59. The non-transitory computer-readable storage medium according to claim 58, wherein at least a portion of the temporal profile of inflammation corresponds to a prenatal period of the subject.
60. The non-transitory computer-readable storage medium according to any one of claims 47 to 59, wherein predicting a diagnostic state of the subject with respect to the disease or disorder comprises processing the plurality of fluorescence intensity measurements using the trained model.
61. The non-transitory computer-readable storage medium according to claim 60, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof.
62. The non-transitory computer-readable storage medium according to claim 60, wherein the trained model includes a gradient-boosted decision tree.
63. The non-transitory computer-readable storage medium according to claim 60, wherein the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, layering, trapping time, maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the fluorescence intensity across the reference line, determination of a plurality of non-linear parameters describing the curvature of the fluorescence intensity across the reference line, determination of a sharp change in the intensity of the fluorescence intensity across the reference line, determination of one or more changes in the baseline intensity of the fluorescence intensity across the reference line, determination of a change in the frequency domain representation of the fluorescence intensity across the reference line, determination of a change in the power spectrum domain representation of the fluorescence intensity across the reference line, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross-recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.
64. The non - transitory computer - readable storage medium according to claim 60, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, layering, trapping time (TT), maximum vertical length, Shannon entropy in vertical length, mean recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the fluorescence intensity across the reference line, determination of a plurality of non - linear parameters describing the curvature of the fluorescence intensity across the reference line, determination of a sharp change in the intensity of the fluorescence intensity across the reference line, determination of one or more changes in the baseline intensity of the fluorescence intensity across the reference line, determination of a change in the frequency - domain representation of the fluorescence intensity across the reference line, determination of a change in the power - spectrum - domain representation of the fluorescence intensity across the reference line, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross - recurrence quantitative analysis parameters, determination of one or more joint - recurrence quantitative analysis parameters, determination of one or more multi - dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.
65. The non - transitory computer - readable storage medium according to any one of claims 47 to 64, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% sensitivity.
66. The non - transitory computer - readable storage medium according to any one of claims 47 to 64, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% specificity.
67. The non - transitory computer - readable storage medium according to claim 47, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% positive predictive value.
68. The non - transitory computer - readable storage medium according to claim 47, wherein the trained model predicts a diagnostic state regarding the disease or disorder with at least about 80% negative predictive value.
69. The non-transitory computer-readable storage medium of claim 47, wherein the trained model predicts a diagnostic state regarding the disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.
80.
70. A method for training a model, comprising: In a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, (a) For each respective training subject among the plurality of training subjects, a first subset of the training subjects has a first diagnostic state corresponding to having a first biological state associated with C-reactive protein, and a second subset of the training subjects among the plurality of training subjects has a second diagnostic state corresponding to not having the first biological state associated with C-reactive protein. (i) Sampling each respective position among a plurality of positions along a reference line on a biological sample of the subject associated with C-reactive protein of the subject, thereby obtaining a plurality of fluorescence intensity measurement values, wherein each fluorescence intensity measurement value among the plurality of fluorescence intensity measurement values corresponds to a different position among the plurality of positions, and each position among the plurality of positions represents a different growth period of the biological sample of the subject associated with C-reactive protein; (ii) Analyzing each fluorescence intensity across the reference line on the biological sample, thereby obtaining a first data set; (iii) Deriving a respective second data set from the corresponding plurality of fluorescence intensity measurement values, wherein each respective feature in the corresponding set of features is determined by continuous variation in C-reactive protein fluorescence intensity. (b)(i) A corresponding set of features of each respective second data set of each of the plurality of training subjects, and (ii) using the corresponding diagnostic states of each of the plurality of training subjects selected from among the first diagnostic state and the second diagnostic state, training a model that has not been trained or has been partially trained, thereby providing an indication as to whether the test subject has the first biological state associated with C-reactive protein based on the values of the features in the set of features obtained from the biological sample associated with C-reactive protein of the test subject. Obtaining a trained model that does, a method comprising.
71. The method according to claim 70, wherein the trained model is a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering model algorithm, a supervised clustering model algorithm, a regression model, a gradient boosting algorithm, or any combination thereof.
72. The method according to claim 70, wherein the trained model is a multinomial classifier.
73. The method according to claim 70, wherein the trained model is a binary classifier.
74. The first biological state associated with C-reactive protein is selected from the group consisting of autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, inflammatory bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer. The method according to any one of claims 70 to 73.
75. The method further comprises evaluating the test subject for the first biological state associated with C-reactive protein by distinguishing the first biological state associated with C-reactive protein from a second biological state associated with C-reactive protein that is different from the first biological state associated with metal metabolism. The method according to claim 70.
76. The method according to claim 75, wherein the first biological state is autism spectrum disorder and the second biological state is attention deficit / hyperactivity disorder.
77. The method according to any one of claims 70 to 76, wherein the test subject is a human.
78. The method according to claim 77, wherein the human is less than 12 years old.
79. The method according to claim 78, wherein the human is less than 1 year old.
80. The method according to any one of claims 70 to 79, wherein the corresponding biological sample associated with the c-reactive protein of each training subject is selected from the group consisting of hair shafts, teeth, and nails.
81. The method according to claim 80, wherein the corresponding biological sample associated with the c-reactive protein of each training subject is the hair shaft, and the reference line corresponds to the longitudinal direction of the hair shaft.
82. The method according to any one of claims 70 to 79, wherein the corresponding biological sample associated with the c-reactive protein of each training subject is the tooth, and the reference line corresponds to the direction across the growth zone including the neonatal line of the tooth.
83. The method according to any one of claims 70 to 82, wherein the first positions at the corresponding plurality of positions along the corresponding biological sample associated with the c-reactive protein of each training subject are arranged such that they correspond to the positions closest to the tip of the corresponding biological sample associated with the c-reactive protein of each training subject.
84. The method according to any one of claims 70 to 79, wherein each trace in the corresponding plurality of fluorescence intensity measurement values includes a plurality of data points, and each data point is an instance of the respective position at the plurality of positions.
85. The method according to any one of claims 70 to 84, wherein the corresponding set of features is selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, layering, trapping time, maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, most likely number of recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof.
86. The method according to any one of claims 70 to 85, wherein the corresponding plurality of positions includes positions of at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000, 6000, 7000, 8000, 9000, 10000, 12000, 14000, 16000, 18000, 20000, or exceeding 20000.
87. The method according to any one of claims 70 to 86, wherein the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency of recurrence, layering, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the plurality of fluorescence intensity measurement values, determination of a plurality of non-linear parameters describing the curvature of the plurality of fluorescence intensity measurement values, determination of a sharp change in intensity of the plurality of fluorescence intensity measurement values, determination of one or more changes in the baseline intensity of the plurality of fluorescence intensity measurement values, determination of a change in the frequency domain representation of the plurality of fluorescence intensity measurement values, determination of a change in the power spectrum domain representation of the plurality of fluorescence intensity measurement values, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross-recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.
88. The method according to any one of claims 70 to 86, wherein the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, layering, trapping time, maximum vertical length, Shannon entropy in vertical length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, average diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determination of the linear gradient of the plurality of fluorescence intensity measurement values, determination of a plurality of non-linear parameters describing the curvature of the plurality of fluorescence intensity measurement values, determination of a sharp change in intensity of the plurality of fluorescence intensity measurement values, determination of one or more changes in the baseline intensity of the plurality of fluorescence intensity measurement values, determination of a change in the frequency domain representation of the plurality of fluorescence intensity measurement values, determination of a change in the power spectrum domain representation of the plurality of fluorescence intensity measurement values, determination of one or more recurrence quantitative analysis parameters, determination of one or more cross recurrence quantitative analysis parameters, determination of one or more joint recurrence quantitative analysis parameters, determination of one or more multi-dimensional recurrence quantitative analysis parameters, estimation of the Lyapunov spectrum, determination of the maximum Lyapunov exponent, and any combination thereof.