Systems and methods for dynamic Raman profiling of biological diseases and disorders and feature engineering methods thereof
By analyzing Raman spectra from human samples like hair, teeth, and nails, the method provides accurate non-invasive diagnosis of neurological and neurodegenerative disorders and cancer, achieving high diagnostic accuracy through trained models and complementary techniques.
Patent Information
- Application Number
- JP2024570581
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-08
- Filing Date
- 2023-06-07
- Publication Date
- 2025-08-05
AI Technical Summary
There is a need for accurate, non-invasive methods to diagnose biological conditions such as neurological and neurodegenerative disorders and cancer using dynamic biological response data from human biological samples like hair, teeth, and nails.
The method involves exposing a biological sample to a light source, acquiring Raman spectra, generating a spatial map, and predicting diagnostic status using a trained model that processes temporal dynamics of the Raman spectra, potentially combined with techniques like LA-ICP-MS and immunohistochemistry for enhanced diagnostic accuracy.
The method achieves high sensitivity and specificity in diagnosing conditions like ASD, ADHD, ALS, and cancer with accuracy up to 90% and an AUROC of 0.90, enabling non-invasive diagnosis in young children and infants.
Smart Images

Figure 2025525304000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 350,090, filed June 8, 2022, the entire text of which is incorporated herein by reference. [Background technology]
[0002] Dynamic biological responses can represent fundamental biological processes that are structurally and functionally important to humans. For example, abnormal dynamic biological responses can be associated with many biological conditions, such as diseases and disorders. Examples of such biological conditions can include neurological conditions (e.g., autism spectrum disorder, schizophrenia, or attention-deficit / hyperactivity disorder (ADHD)), neurodegenerative conditions (e.g., amyotrophic lateral sclerosis (ALS), Alzheimer's disease, Parkinson's disease, and Huntington's disease), and cancer (e.g., childhood cancer). Summary of the Invention
[0003] Given the above background, there exists a need for accurate methods and systems for diagnosing biological conditions, particularly non-invasive diagnosis. Such diagnosis can be based on accurate profiling of biomarkers detectable by non-invasive methods for diagnosing biological conditions. The present disclosure provides improved systems and methods for accurately diagnosing biological conditions based on analysis of dynamic biological response data from a biological sample non-invasively obtained from a subject. Such improved systems and methods for accurately diagnosing biological conditions may be based on a combination of Raman profiling of the biological sample and artificial intelligence data analysis. The present disclosure addresses these needs, for example, by providing biological sample biomarkers for diagnosing biological conditions. Biological samples include human biological samples associated with incremental growth. Such biological samples can be hair shafts, teeth, and nails. The non-invasive biomarkers of the present disclosure can be used for diagnosing young children, and in some cases, infants under one year of age.
[0004] In one aspect, the present disclosure provides a method for predicting a subject's diagnostic status with respect to a disease or disorder of the subject, the method comprising: (a) exposing a biological sample of the subject to a light source, wherein the biological sample comprises a tooth sample, a hair sample, or a nail sample; (b) acquiring a plurality of Raman spectra from the exposed biological sample; (c) processing the plurality of Raman spectra to generate a spatial map of the plurality of Raman spectra; and (d) predicting the subject's diagnostic status with respect to the disease or disorder based at least in part on the spatial map of the plurality of Raman spectra. In some embodiments, the light source comprises a laser.
[0005] In some embodiments, analyzing determines the temporal dynamics of an underlying biological process. In some embodiments, analyzing includes reducing the dimensionality of the plurality of Raman spectra (e.g., by independent component analysis) before processing. In some embodiments, the optical signal is generated by a light source (e.g., a laser). In some embodiments, the biological sample includes a tooth sample. In some embodiments, the method further includes detecting or monitoring changes in a temporal stress profile (e.g., one or more traces) indicative of a temporal response of the subject. In some embodiments, the temporal response includes a biochemical response. In some embodiments, the temporal response includes a biological response, a physiological response, an anatomical response, a treatment response, a stress-related response, or a combination thereof. In some embodiments, the plurality of Raman spectra includes wavenumbers from about 200 to about 3700. In some embodiments, acquiring includes using a Raman spectroscopic microscope. In some embodiments, the Raman spectroscopic microscope includes a 50x air-coupling objective, a 63x water-immersion coupling objective, or any combination thereof. In some embodiments, the laser includes a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof. In some embodiments, the acquiring is performed using an integration time of about 0.2 seconds to about 0.3 seconds. In some embodiments, the acquiring includes, after acquiring a Raman spectrum of the plurality of Raman spectra, moving the biological sample with a step size of about 2 microns to about 5 microns.
[0006] In some embodiments, the disease or disorder comprises autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder is ASD. In some embodiments, the subject is a human. In some embodiments, the subject is an adult. In some embodiments, the subject is about 12 to about 5 years old. In some embodiments, the subject is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, at least a portion of the temporal Raman profile corresponds to a prenatal period of the subject.
[0007] In some embodiments, predicting a subject's diagnostic status for a disease or disorder includes processing the spatial map using a trained model. In some embodiments, processing includes extracting features from the spatial map (e.g., by recurrence quantification analysis) and analyzing the features using the trained model. In some embodiments, processing includes computational analysis of temporal dynamics derived from the spatial map, for example, by applying dimensionality reduction techniques including independent component analysis (ICA) and / or principal component analysis (PCA), followed by applying recurrence quantification analysis (RQA) to extract computational features describing the dimensionality derived from the ICA / PCA.
[0008] In some embodiments, the processing comprises determining one or more features of the temporal dynamics of the one or more traces. In some embodiments, the temporal dynamics of the one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies to the one or more traces one or more of the following operations and / or methods: determining a linear slope, determining a plurality of non-linear parameters describing the curvature of the one or more traces, determining abrupt changes in intensity of the one or more traces, determining one or more changes in baseline intensity of the one or more traces, determining changes in a frequency domain representation of the one or more traces, determining changes in a power spectral domain representation of the one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent.
[0009] In some embodiments, the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm (e.g., a gradient boosted implementation of a machine learning algorithm such as a gradient boosted decision tree), and any combination thereof. In some embodiments, the trained model comprises a gradient boosted ensemble model. In some embodiments, the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency to recurrence, stratification, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, and / or any combination thereof. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency to recurrence, stratification, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, and / or any combination thereof.
[0010] In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies to one or more traces one or more of the following operations and / or methods: determining a linear slope, determining a plurality of non-linear parameters describing the curvature of one or more traces, determining abrupt changes in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining changes in a frequency domain representation of one or more traces, determining changes in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent.
[0011] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model that has at least about 70%, 75%, 80%, 85%, or 90% sensitivity in predicting the diagnostic status for the disease or disorder across a suitable cohort population (such as, for example, those provided in the Examples section below).
[0012] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model having a sensitivity of up to about 70%, 75%, 80%, 85%, or 90% in predicting the diagnostic status for the disease or disorder across a suitable cohort population.
[0013] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model that has a specificity of at least about 70%, 75%, 80%, 85%, or 90% in predicting the diagnostic status for the disease or disorder across a suitable cohort population.
[0014] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model having a specificity of up to about 70%, 75%, 80%, 85%, or 90% in predicting the diagnostic status for the disease or disorder across a suitable cohort population.
[0015] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model that has a positive predictive value of at least about 70%, 75%, 80%, 85%, or 90% in predicting the diagnostic status for the disease or disorder across a suitable cohort population.
[0016] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model that has a positive predictive value of up to about 70%, 75%, 80%, 85%, or 90% in predicting the diagnostic status for the disease or disorder across a suitable cohort population.
[0017] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model that has a negative predictive value of at least about 70%, 75%, 80%, 85%, or 90% in predicting the diagnostic status for the disease or disorder across a suitable cohort population.
[0018] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model that has a negative predictive value of up to about 70%, 75%, 80%, 85%, or 90% in predicting the diagnostic status for the disease or disorder across a suitable cohort population.
[0019] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder using a model that predicts the diagnostic status for the disease or disorder having an area under the receiver operating characteristic curve (AUROC) for a suitable cohort population of at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.82, at least about 0.84, at least about 0.86, at least about 0.88, or at least about 0.90.
[0020] In another aspect, the present disclosure provides a device including one or more processors and a memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for: (a) sampling each respective position at a plurality of locations along a reference line on a biological sample of the subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, wherein each Raman spectrum in the plurality of Raman spectra corresponds to a different position at the plurality of locations, each position at the plurality of locations representing a different growth period of the biological sample associated with the Raman signature; (b) analyzing each of the plurality of Raman spectra across the reference line on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from corresponding plurality of Raman spectral measurements, wherein each respective feature in a corresponding set of features is determined by successive variations in the Raman spectra; and (d) processing the features using the trained model to predict a diagnostic status of the subject for a disease or disorder associated with the Raman signature. In some embodiments, each second data set is derived by applying recursive quantitative analysis or a related method to the corresponding plurality of Raman spectral measurements. In some embodiments, analyzing the Raman spectra includes cosmic ray removal, background correction, normalization, peak fitting, or any combination thereof.
[0021] In some embodiments, the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof. In some embodiments, the instructions further comprise detecting or monitoring a change in the Raman spectrum across the plurality of locations indicative of a temporal response of the subject. In some embodiments, the temporal response comprises a biological response, a physiological response, an anatomical response, a treatment response, a stress-related response, or a combination thereof. In some embodiments, the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700. In some embodiments, sampling comprises using a Raman spectromicroscope. In some embodiments, the Raman spectromicroscope comprises a 50x air-coupling objective, a 63x water-immersion coupling objective, or any combination thereof. In some embodiments, sampling comprises exposing the biological sample to a light source to generate one of the plurality of Raman spectra at the plurality of locations. In some embodiments, the light source comprises a laser, and the laser comprises a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof. In some embodiments, the instructions further include translating, wherein the translation includes, after acquiring a Raman spectrum of the plurality of Raman spectra, moving the biological sample from a first position to a second position of the plurality of positions with a step size of about 2 microns to about 5 microns. In some embodiments, the translation is performed using an integration time of about 0.2 seconds to about 0.3 seconds. In some embodiments, the disease or disorder comprises autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder is ASD. In some embodiments, predicting the subject's diagnostic status for a disease or disorder comprises processing changes in the Raman spectra across the plurality of positions using a trained model.In some embodiments, the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. In some embodiments, the trained model comprises a gradient boosted ensemble model. In some embodiments, the trained model is configured to process one or more features selected from the group consisting of stratiformity, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of stratiformity, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof.
[0022] In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies to one or more traces one or more of the following operations and / or methods: determining a linear slope, determining a plurality of non-linear parameters describing the curvature of one or more traces, determining abrupt changes in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining changes in a frequency domain representation of one or more traces, determining changes in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent.
[0023] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium for classification, the one or more computer programs, when executed by a computer system, causing the computer system to: (a) sample each respective position in a plurality of positions along a reference line on a biological sample of the subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, wherein each Raman spectrum in the plurality of Raman spectra corresponds to a different position in the plurality of positions, and each position in the plurality of positions corresponds to a Raman signature of the subject; The method includes instructions for performing a method comprising: (a) acquiring a plurality of Raman spectra representing different growth periods of a biological sample associated with a characteristic; (b) analyzing each of a plurality of Raman spectra across a baseline on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from the corresponding plurality of Raman spectral measurements, wherein each respective feature in the corresponding set of features is determined by successive variations in the Raman spectra; and (d) processing the features using the trained model to predict the subject's diagnostic status for a disease or disorder associated with the Raman signature. In some embodiments, each second data set is derived by applying recurrence quantification analysis or a related method to the corresponding plurality of Raman spectral measurements. In some embodiments, analyzing the Raman spectra includes cosmic ray removal, background correction, normalization, peak fitting, or any combination thereof.
[0024] In some embodiments, the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof. In some embodiments, the method further comprises detecting or monitoring changes in the Raman spectrum across the plurality of locations indicative of a temporal response (e.g., one or more traces) of the subject. In some embodiments, the temporal response comprises a biological response, a physiological response, an anatomical response, a treatment response, a stress-related response, or a combination thereof. In some embodiments, the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700. In some embodiments, sampling comprises using a Raman spectromicroscope. In some embodiments, the Raman spectromicroscope comprises a 50x air-coupling objective, a 63x water-immersion coupling objective, or any combination thereof. In some embodiments, sampling comprises exposing the biological sample to a light source to generate one Raman spectrum of the plurality of Raman spectra at the plurality of locations. In some embodiments, the light source comprises a laser, and the laser comprises a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof. In some embodiments, the instructions further include translating, wherein the translation includes, after acquiring a Raman spectrum of the plurality of Raman spectra, moving the biological sample from a first position to a second position of the plurality of positions with a step size of about 2 microns to about 5 microns. In some embodiments, the translation is performed using an integration time of about 0.2 seconds to about 0.3 seconds. In some embodiments, the disease or disorder comprises autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder is ASD. In some embodiments, predicting the subject's diagnostic status for a disease or disorder comprises processing changes in the Raman spectra across the plurality of positions using a trained model.In some embodiments, the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. In some embodiments, the trained model comprises a gradient boosted ensemble model. In some embodiments, the trained model is configured to process one or more features selected from the group consisting of stratiformity, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of stratiformity, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof.
[0025] In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies to one or more traces one or more of the following operations and / or methods: determining a linear slope, determining a plurality of non-linear parameters describing the curvature of one or more traces, determining abrupt changes in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining changes in a frequency domain representation of one or more traces, determining changes in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent.
[0026] In another aspect, the disclosure provides a method of training a model, in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, comprising: (a) for each respective training subject in a plurality of training subjects, wherein a first subset of the training subjects in the plurality of training subjects have a first diagnostic status corresponding to having a first biological state associated with the Raman signature, and a second subset of the training subjects in the plurality of training subjects have a second diagnostic status corresponding to not having the first biological state associated with the Raman signature, (i) sampling each respective position in a plurality of positions along a reference line on a biological sample of the subject associated with the subject's Raman signature, thereby obtaining a plurality of Raman spectra, wherein each Raman spectrum in the plurality of Raman spectra corresponds to a different position in the plurality of positions, and each position in the plurality of positions corresponds to a different growth period of the biological sample of the subject associated with the Raman signature; (ii) analyzing each Raman spectrum across a baseline on the biological sample, thereby obtaining a first data set; and (iii) deriving a respective second data set from a corresponding plurality of Raman spectra, wherein each respective feature in the corresponding set of features is determined by a continuous variation in the Raman spectrum. (b) training an untrained or partially untrained model using (i) the corresponding set of features of each respective second data set for each training subject in the plurality of training subjects and (ii) a corresponding diagnostic state for each training subject in the plurality of training subjects, the corresponding diagnostic state being selected from among a first diagnostic state and a second diagnostic state, thereby obtaining a trained model that provides an indication as to whether the test subject has a first biological state associated with the Raman signature based on the values of the features in the set of features obtained from the biological sample associated with the test subject's Raman signature.In some embodiments, each second data set is derived by applying recursive quantitative analysis or a related method to the corresponding plurality of Raman spectral measurements. In some embodiments, analyzing the Raman spectra includes cosmic ray removal, background correction, normalization, peak fitting, or any combination thereof.
[0027] In some embodiments, the trained model is a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering model algorithm, a supervised clustering model algorithm, a regression model, or a gradient boosting algorithm (e.g., a gradient boosting implementation of a machine learning algorithm such as a gradient boosted decision tree). In some embodiments, the trained model is a multinomial classifier. In some embodiments, the trained model is a binary classifier. In some embodiments, the trained model is a regressor. In some embodiments, the first biological condition is selected from the group consisting of autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer.
[0028] In some embodiments, evaluating the test subject for a first biological state associated with the Raman signature further comprises distinguishing between the presence of the first biological state associated with the Raman signature and the absence of the first biological state associated with the Raman signature. In some embodiments, evaluating the test subject for a first biological state associated with the Raman signature further comprises distinguishing between the first biological state associated with the Raman signature and a second biological state associated with a Raman signature that is different from the first biological state associated with the Raman signature. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is neurotypical, i.e., the absence of a neurodevelopmental disorder. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is attention-deficit / hyperactivity disorder. In some embodiments, the test subject is a human. In some embodiments, the test subject is an adult. In some embodiments, the human is about 12 to about 5 years old. In some embodiments, the subject is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, at least a portion of the temporal profile of the Raman profile corresponds to the subject's prenatal period.
[0029] In some embodiments, the corresponding biological sample associated with each training object's Raman signature is selected from the group consisting of a hair shaft, a tooth, and a nail. In some embodiments, the corresponding biological sample associated with each training object's Raman signature is a hair shaft, and the reference line corresponds to the longitudinal direction of the hair shaft. In some embodiments, the corresponding biological sample associated with each training object's Raman signature is a tooth, and the reference line corresponds to a direction across the growth zone including the tooth's neonatal line. In some embodiments, the corresponding plurality of positions are arranged such that a first position in the corresponding plurality of positions along the corresponding biological sample of each training object corresponds to a position closest to the tip of the corresponding biological sample of each training object. In some embodiments, each trace in the corresponding plurality of Raman spectral measurements includes a plurality of data points, and each data point is an instance of a respective position in the plurality of positions. In some embodiments, the corresponding set of features is selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, propensity for recurrence, stratification, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, and / or any combination thereof. In some embodiments, the corresponding plurality of locations includes at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or more than 10000 locations.
[0030] In some embodiments, the corresponding set of features is selected from a group of temporal dynamic features of one or more traces. In some embodiments, the temporal dynamic features of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies to one or more traces one or more of the following operations and / or methods: determining a linear slope, determining a plurality of non-linear parameters describing the curvature of one or more traces, determining abrupt changes in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining changes in a frequency domain representation of one or more traces, determining changes in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent.
[0031] Another aspect of the present disclosure provides a non-transitory computer-readable medium containing machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0032] Another aspect of the present disclosure provides a system including one or more computer processors and a computer memory coupled to the computer processors, the computer memory including machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.
[0033] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0034] Incorporation by Reference All publications, patents, and patent applications mentioned herein are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or take precedence over such conflicting material.
[0035] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (hereinafter "Figure" and "FIG.") [Brief explanation of the drawings]
[0036] [Figure 1] 1 illustrates an example block diagram of a computing device 100 of the present disclosure. [Figure 2A] Illustrates a diagram of a subject's hair sample. [Figure 2B] 1 shows a diagram of the subject tooth sample. [Figure 2C] Illustrates a diagram of a subject nail sample. [Figure 3] 3 shows a flowchart of a method 300 for assessing a subject for a biological condition. [Figure 4] 1 illustrates a computer system that is programmed or otherwise configured to implement the methods provided herein. [Figure 5]Figure 1 shows an example of model accuracy for predicting autism spectrum disorder (ASD) diagnostic status, utilizing features derived from applying RQA to ICA-derived dimensions of Raman waveforms, as shown by an empirical receiver operating characteristic (ROC) curve to assess the accuracy of the disclosed method for assessing subjects with autism spectrum disorder. Device performance is measured by calculating the area under the curve (AUC) of the ROC plot, which provides a measure of performance at various classification thresholds. Here, the AUC is 0.86, indicating robust and accurate predictive performance. [Figure 6] We present an example of model accuracy for predicting the diagnostic status of amyotrophic lateral sclerosis (ALS) utilizing features derived from applying RQA to dimensions derived by ICA of Raman waveforms, as illustrated by an experimental receiver operating characteristic (ROC) curve to evaluate the accuracy of the disclosed method for assessing subjects for autism spectrum disorder. Device performance is measured by calculating the area under the curve (AUC) of the ROC plot, which provides a measure of performance at various classification thresholds. Here, the AUC is 0.88, indicating robust and accurate predictive performance. DETAILED DESCRIPTION OF THE INVENTION
[0037] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0038] Dynamic biological responses can represent fundamental biological processes that are structurally and functionally important to humans. For example, abnormal dynamic biological responses can be associated with many biological conditions, such as diseases and disorders. Examples of such biological conditions can include neurological conditions (e.g., autism spectrum disorder, schizophrenia, or attention-deficit / hyperactivity disorder (ADHD)), neurodegenerative conditions (e.g., amyotrophic lateral sclerosis (ALS), Alzheimer's disease, Parkinson's disease, and Huntington's disease), and cancer (e.g., childhood cancer).
[0039] Given the above background, there exists a need for accurate methods and systems for diagnosing biological conditions, particularly non-invasive diagnosis. Such diagnosis can be based on accurate profiling of biomarkers detectable by non-invasive methods for diagnosing biological conditions. The present disclosure provides improved systems and methods for accurately diagnosing biological conditions based on analysis of dynamic biological response data from a biological sample non-invasively obtained from a subject. In some embodiments, such improved systems and methods for accurately diagnosing biological conditions are based on a combination of Raman profiling of the biological sample and artificial intelligence data analysis. The present disclosure addresses these needs, for example, by providing biological sample biomarkers for diagnosing biological conditions. Biological samples include human biological samples associated with incremental growth. Such biological samples can be hair shafts, teeth, and nails. The non-invasive biomarkers of the present disclosure can be used to diagnose young children, and in some cases, infants under one year of age. In some cases, the children are between about 12 and about 5 years of age. In some embodiments, the children are less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year of age. In some embodiments, the child is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old.
[0040] In one aspect, the present disclosure provides a method for predicting a subject's diagnostic status with respect to a disease or disorder, the method comprising: (a) exposing a biological sample of the subject to a light source, wherein the biological sample comprises a tooth sample, a hair sample, or a nail sample; (b) acquiring a plurality of Raman spectra from the exposed biological sample; (c) processing the plurality of Raman spectra to generate a spatial map of the plurality of Raman spectra; and (d) predicting the subject's diagnostic status with respect to the disease or disorder based at least in part on the spatial map of the plurality of Raman spectra. In some embodiments, the light source comprises a laser.
[0041] In some embodiments, analyzing determines the temporal dynamics of an underlying biological process. In some embodiments, analyzing includes reducing the dimensionality of the plurality of Raman spectra (e.g., by independent component analysis) before processing. In some embodiments, the optical signal is generated by a light source (e.g., a laser). In some embodiments, the biological sample includes a tooth sample. In some embodiments, the method further includes detecting or monitoring changes in a temporal stress profile (e.g., one or more traces described elsewhere herein) indicative of a temporal response of the subject. In some embodiments, the temporal response includes a biochemical response. In some embodiments, the temporal response includes a biological response, a physiological response, an anatomical response, a therapeutic response, a stress-related response, or a combination thereof. In some embodiments, the plurality of Raman spectra includes wavenumbers from about 200 to about 3700. In some embodiments, acquiring includes using a Raman spectroscopic microscope. In some embodiments, the Raman spectroscopic microscope includes a 50x air-coupled objective, a 63x water-immersion coupled objective, or any combination thereof. In some embodiments, the laser comprises a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof. In some embodiments, the acquiring is performed using an integration time of about 0.2 seconds to about 0.3 seconds. In some embodiments, the acquiring includes, after acquiring a Raman spectrum of the plurality of Raman spectra, moving the biological sample with a step size of about 2 microns to about 5 microns.
[0042] In some embodiments, the systems and methods disclosed herein use Raman spectroscopy alone or in combination with other techniques. In some embodiments, such techniques include laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS), C-reactive immunohistochemistry fluorescence staining, etc. In some embodiments, combining techniques improves the diagnostic accuracy or precision of a given technique alone. In some embodiments, the addition of LA-ICP-MS provides multiple non-invasive metal metabolism biomarkers for a given biological sample that complement the diagnostic power of Raman spectroscopy. In some embodiments, the metal metabolism biomarkers include zinc, tin, magnesium, copper, iodide, lithium, aluminum, phosphorus, sulfur, calcium, chromium, manganese, iron, cobalt, nickel, arsenic, strontium, cadmium, tin, iodine, barium, mercury, lead, bismuth, molybdenum, or any combination thereof. In some embodiments, the addition of C-reactive protein immunohistochemistry fluorescence provides temporal variations in inflammation to complement the diagnostic power of Raman spectroscopy. In some embodiments, the plurality of metal metabolism biomarkers comprises at least two, at least five, or at least 10 metal metabolism biomarkers. In some embodiments, the plurality of metal metabolism biomarkers comprises 20 or fewer, 10 or fewer, or 5 or fewer metal metabolism biomarkers. In some embodiments, the plurality of metal metabolism biomarkers consists of 2-5, 3-10, or 8-20 metal metabolism biomarkers. In some embodiments, the plurality of metal metabolism biomarkers falls within another range beginning with two or more metal metabolism biomarkers and ending with 20 or fewer metal metabolism biomarkers. In some embodiments, the plurality of Raman spectra comprises at least two, at least five, or at least 10 Raman spectra. In some embodiments, the plurality of Raman spectra comprises 20 or fewer, 10 or fewer, or 5 or fewer Raman spectra. In some embodiments, the plurality of Raman spectra consists of 2-5, 3-10, or 8-20 Raman spectra. In some embodiments, the plurality of Raman spectra falls within another range beginning with two or more neurons and ending with 20 or fewer neurons.
[0043] In some embodiments, the disease or disorder comprises autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. In some embodiments, the disease or disorder is ASD. In some embodiments, the subject is a human. In some embodiments, the subject is an adult. In some embodiments, the subject is about 12 to about 5 years old. In some embodiments, the subject is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, at least a portion of the temporal Raman profile corresponds to a prenatal period of the subject.
[0044] In some embodiments, predicting a subject's diagnostic status for a disease or disorder includes processing the spatial map using a trained model. In some embodiments, processing includes extracting features from the spatial map (e.g., by recurrence quantification analysis) and analyzing the features using the trained model. In some embodiments, processing includes computational analysis of temporal dynamics derived from the spatial map, for example, by applying dimensionality reduction techniques including independent component analysis (ICA) and / or principal component analysis (PCA), followed by applying recurrence quantification analysis (RQA) to extract computational features describing the dimensionality derived from the ICA / PCA.
[0045] In some embodiments, the trained model includes multiple parameters, where the term “parameter” refers to any coefficient or similarly any value of an internal or external element (e.g., weight and / or hyperparameter) within a model that can affect (e.g., modify, adjust, and / or tune) one or more inputs, outputs, and / or functions within the model (e.g., if the model is a regressor or classifier). For example, in some embodiments, a parameter of a model refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, adjust, and / or tune the behavior, learning, and / or performance of the model. In some examples, a parameter is used to increase or decrease the influence of an input (e.g., feature) to the model. As a non-limiting example, in some embodiments, a parameter is used to increase or decrease the influence of a node (e.g., a neural network), where a node includes one or more activation functions. The assignment of parameters to particular inputs, outputs, and / or functions of a model is not limited to any one paradigm of a given model, but may be used in any suitable model for desired performance. In some embodiments, a parameter has a fixed value. In some embodiments, the values of the parameters are adjustable manually and / or automatically. In some embodiments, the values of the parameters are modified by a model validation and / or training process (e.g., by error minimization and / or backpropagation methods). In some embodiments, the models of the present disclosure include multiple parameters. In some embodiments, the plurality of parameters associated with a model (e.g., an untrained, partially trained, or fully trained model) is n parameters, where n≧2, n≧5, n≧10, n≧25, n≧40, n≧50, n≧75, n≧100, n≧125, n≧150, n≧200, n≧225, n≧250, n≧350, n≧500, n≧600, n≧750, n≧1,000, n≧2,000, n≧4,000, n≧5,000, n≧7,500, n≧10,000, n≧20,000, n≧40,000, n≧75,000, n≧100,000, n≧200,000, n≧500,000, n≧1×10 6, n ≥ 5 × 10 6 , or n≧1×10 7 In some embodiments, n is between 10,000 and 1×10 7 , 100,000~5×10 6 , or 500,000 to 1×10 6 In some embodiments, the plurality of parameters is at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1 x 10 6 pieces, at least 5 x 10 6 Pieces, at least 1 x 10 7 pieces, at least 5 x 10 7 pieces, or at least 1 x 10 8 In some embodiments, the plurality of parameters comprises 1×10 9 pcs or less, 1×10 8 pcs or less, 1×10 7 pcs or less, 1×10 6 In some embodiments, the plurality of parameters includes 10 to 10,000, 100 to 100,000, 1000 to 1×10 parameters. 6 pieces, 100,000~1×10 7 pieces, 1×10 6 ~1×10 8 pcs or 1 x 10 7 ~1×10 9 In some embodiments, the plurality of parameters starts from 10 or more parameters, and is 1×10 9 It is in another range ending with fewer than 1 parameter.
[0046] In some embodiments, the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm (e.g., a gradient boosted implementation of a machine learning algorithm such as a gradient boosted decision tree), and any combination thereof. In some embodiments, the trained model comprises a gradient boosted ensemble model. In some embodiments, the trained model is configured to process one or more features selected from the group consisting of recurrence rate, determinism, mean diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency to recurrence, stratification, trapping time, maximum vertical line length, Shannon entropy in vertical line length, mean recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, and / or any combination thereof. In some embodiments, the trained model is configured to process two or more features selected from the group consisting of recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency to recurrence, stratification, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, and / or any combination thereof.
[0047] In some embodiments, the trained model is configured to process one or more features of the temporal dynamics of one or more traces. In some embodiments, the temporal dynamics of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies to one or more traces one or more of the following operations and / or methods: determining a linear slope, determining a plurality of non-linear parameters describing the curvature of one or more traces, determining abrupt changes in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining changes in a frequency domain representation of one or more traces, determining changes in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent.
[0048] In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with at least about 80% sensitivity. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with up to about 80% sensitivity. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with at least about 80% specificity. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with up to about 80% specificity. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with at least about 80% positive predictive value. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with up to about 80% positive predictive value. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with at least about 80% negative predictive value. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with up to about 80% negative predictive value. In some embodiments, the method further comprises predicting the subject's diagnostic status for the disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.80.
[0049] In another aspect, the present disclosure provides a device including one or more processors and a memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for: (a) sampling each respective position at a plurality of locations along a reference line on a biological sample of the subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, wherein each Raman spectrum in the plurality of Raman spectra corresponds to a different position at the plurality of locations, each position at the plurality of locations representing a different growth period of the biological sample associated with the Raman signature; (b) analyzing each of the plurality of Raman spectra across the reference line on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from corresponding plurality of Raman spectral measurements, wherein each respective feature in a corresponding set of features is determined by successive variations in the Raman spectra; and (d) processing the features using the trained model to predict a diagnostic status of the subject for a disease or disorder associated with the Raman signature. In some embodiments, each second data set is derived by applying recursive quantitative analysis or a related method to the corresponding plurality of Raman spectral measurements. In some embodiments, analyzing the Raman spectra includes cosmic ray removal, background correction, normalization, peak fitting, or any combination thereof.
[0050] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium for classification, the one or more computer programs, when executed by a computer system, causing the computer system to: (a) sample each respective position in a plurality of positions along a reference line on a biological sample of the subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, wherein each Raman spectrum in the plurality of Raman spectra corresponds to a different position in the plurality of positions, and each position in the plurality of positions corresponds to a Raman signature of the subject; The method includes instructions for performing a method comprising: (a) acquiring a plurality of Raman spectra representing different growth periods of a biological sample associated with a characteristic; (b) analyzing each of a plurality of Raman spectra across a baseline on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from the corresponding plurality of Raman spectral measurements, wherein each respective feature in the corresponding set of features is determined by successive variations in the Raman spectra; and (d) processing the features using the trained model to predict the subject's diagnostic status for a disease or disorder associated with the Raman signature. In some embodiments, each second data set is derived by applying recurrence quantification analysis or a related method to the corresponding plurality of Raman spectral measurements. In some embodiments, analyzing the Raman spectra includes cosmic ray removal, background correction, normalization, peak fitting, or any combination thereof.
[0051] In another aspect, the disclosure provides a method of training a model, in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, comprising: (a) for each respective training subject in a plurality of training subjects, wherein a first subset of the training subjects in the plurality of training subjects have a first diagnostic status corresponding to having a first biological state associated with the Raman signature, and a second subset of the training subjects in the plurality of training subjects have a second diagnostic status corresponding to not having the first biological state associated with the Raman signature, (i) sampling each respective position in a plurality of positions along a reference line on a biological sample of the subject associated with the subject's Raman signature, thereby obtaining a plurality of Raman spectra, wherein each Raman spectrum in the plurality of Raman spectra corresponds to a different position in the plurality of positions, and each position in the plurality of positions corresponds to a different growth period of the biological sample of the subject associated with the Raman signature; (ii) analyzing each Raman spectrum across a baseline on the biological sample, thereby obtaining a first data set; and (iii) deriving a respective second data set from a corresponding plurality of Raman spectra, wherein each respective feature in the corresponding set of features is determined by a continuous variation in the Raman spectrum. (b) training an untrained or partially untrained model using (i) the corresponding set of features of each respective second data set for each training subject in the plurality of training subjects and (ii) a corresponding diagnostic state for each training subject in the plurality of training subjects, the corresponding diagnostic state being selected from among a first diagnostic state and a second diagnostic state, thereby obtaining a trained model that provides an indication as to whether the test subject has a first biological state associated with the Raman signature based on the values of the features in the set of features obtained from the biological sample associated with the test subject's Raman signature.In some embodiments, each second data set is derived by applying recursive quantitative analysis or a related method to the corresponding plurality of Raman spectral measurements. In some embodiments, analyzing the Raman spectra includes cosmic ray removal, background correction, normalization, peak fitting, or any combination thereof.
[0052] In some embodiments, each subject (e.g., test subject) is selected from a plurality of subjects. In some embodiments, the plurality of subjects includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, or at least 500 subjects. In some embodiments, the plurality of subjects includes up to 1,000, up to 500, up to 100, up to 50, up to 20, or up to 10 subjects. In some embodiments, the plurality of subjects consists of 2-10, 5-20, 10-100, or 100-1000 subjects. In some embodiments, the plurality of subjects falls within another range beginning with 2 or more subjects and ending with up to 1000 subjects.
[0053] In some embodiments, the plurality of training subjects comprises at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 100,000 training subjects. In some embodiments, the plurality of training subjects comprises up to 1,000,000, up to 100,000, up to 10,000, up to 1000, 500 or up to 100, 50 or up to 20, or up to 10 training subjects. In some embodiments, the plurality of training subjects consists of between 2 and 1,000, between 500 and 10,000, between 10,000 and 100,000, or between 100,000 and 1,000,000 subjects. In some embodiments, the plurality of training subjects falls within another range beginning with two or more training subjects and ending with 1,000,000 or fewer training subjects.
[0054] In some embodiments, each subset of training subjects (e.g., the first subset and / or the second subset) in the plurality of training subjects includes at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 training subjects. In some embodiments, each subset of training subjects includes no more than 500,000, no more than 10,000, no more than 1000, no more than 500, no more than 100, no more than 50, no more than 20, or no more than 10 training subjects. In some embodiments, each subset of training subjects consists of between 2 and 100, between 50 and 2000, between 1000 and 10,000, or between 10,000 and 500,000 training subjects. In some embodiments, each subset of training subjects falls within a different range, starting with 2 or more training subjects and ending with 500,000 or fewer training subjects.
[0055] In some embodiments, the set of features (e.g., as determined by continuous variation in the Raman spectrum) includes at least 1, at least 2, at least 3, at least 5, at least 8, at least 10, at least 15, or at least 20 features. In some embodiments, the set of features includes 50 or fewer, 20 or fewer, 10 or fewer, 5 or fewer, or 3 or fewer features. In some embodiments, the set of features consists of 1-10, 4-15, 8-20, or 15-50 features. In some embodiments, the set of features falls within another range starting with one or more features and ending with 50 or fewer features.
[0056] In some embodiments, the trained model is a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering model algorithm, a supervised clustering model algorithm, a regression model, or a gradient boosting algorithm (e.g., a gradient boosting implementation of a machine learning algorithm such as a gradient boosted decision tree). In some embodiments, the trained model is a multinomial classifier. In some embodiments, the trained model is a binary classifier. In some embodiments, the first biological condition is selected from the group consisting of autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer.
[0057] In some embodiments, evaluating the test subject for a first biological state associated with the Raman signature further comprises distinguishing between the presence of the first biological state associated with the Raman signature and the absence of the first biological state associated with the Raman signature. In some embodiments, evaluating the test subject for a first biological state associated with the Raman signature further comprises distinguishing between the first biological state associated with the Raman signature and a second biological state associated with a Raman signature that is different from the first biological state associated with the Raman signature. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is neurotypical, i.e., the absence of a neurodevelopmental disorder. In some embodiments, the first biological state is an autism spectrum disorder and the second biological state is attention-deficit / hyperactivity disorder. In some embodiments, the test subject is a human. In some embodiments, the test subject is an adult. In some embodiments, the human is about 12 to about 5 years old. In some embodiments, the subject is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, at least a portion of the temporal profile of the Raman profile corresponds to the subject's prenatal period. In some embodiments, the human is at least 13 years old, at least 14 years old, at least 15 years old, at least 18 years old, or at least 30 years old. In some embodiments, the human is 40 years old or younger, 30 years old or younger, 18 years old or younger, or 15 years old or younger. In some embodiments, the human is about 1 to about 6 years old, about 5 to about 10 years old, or about 10 to about 40 years old. In some embodiments, the human falls within different age ranges starting at about 1 year old or older and ending at about 40 years old or younger.
[0058] In some embodiments, the first biological state and / or the second biological state is selected from a plurality of biological states. In some embodiments, the plurality of biological states includes at least two, at least five, or at least ten biological states. In some embodiments, the plurality of biological states includes 20 or fewer, 10 or fewer, or 5 or fewer biological states. In some embodiments, the plurality of biological states consists of 2-5, 3-10, or 8-20 biological states. In some embodiments, the plurality of biological states falls within another range, beginning with two or more biological states and ending with 20 or fewer biological states. In some embodiments, the first diagnostic state and / or the second diagnostic state is selected from a plurality of diagnostic states. In some embodiments, the plurality of diagnostic states includes at least two, at least five, or at least 10 diagnostic states. In some embodiments, the plurality of diagnostic states includes 20 or fewer, 10 or fewer, or 5 or fewer diagnostic states. In some embodiments, the plurality of diagnostic states consists of 2-5, 3-10, or 8-20 diagnostic states. In some embodiments, the plurality of diagnostic conditions is within another range beginning with two or more diagnostic conditions and ending with 20 or fewer diagnostic conditions.
[0059] In some embodiments, the corresponding biological sample associated with the Raman signature of each subject (e.g., test subject and / or training subject) is selected from the group consisting of a hair shaft, a tooth, and a nail. In some embodiments, the corresponding biological sample associated with the Raman signature of each subject (e.g., test subject and / or training subject) is a hair shaft, and the reference line corresponds to the longitudinal direction of the hair shaft. In some embodiments, the corresponding biological sample associated with the Raman signature of each subject (e.g., test subject and / or training subject) is a tooth, and the reference line corresponds to a direction across the growth zone including the neonatal line of the tooth. In some embodiments, the corresponding plurality of locations are arranged such that a first location in the corresponding plurality of locations along the corresponding biological sample of each subject (e.g., test subject and / or training subject) corresponds to a location closest to the tip of the corresponding biological sample of each training subject. In some embodiments, each trace in the corresponding plurality of Raman spectral measurements includes a plurality of data points, each data point being an instance of a respective location in the plurality of locations. In some embodiments, the corresponding set of features is selected from the group consisting of layeredness, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, and any combination thereof. In some embodiments, the corresponding plurality of locations includes at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or more than 10000 locations.
[0060] In some embodiments, the plurality of locations along the reference line on the biological sample is at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, or at least 1×10 6 In some embodiments, the plurality of locations comprises 1×107 pcs or less, 1×10 6 In some embodiments, the plurality of positions comprises 50-1000, 500-50,000, 10,000-1×10 positions, or 1000-1×10 positions. 6 pcs or 1 x 10 6 ~1×10 7 In some embodiments, the plurality of positions begins with 50 or more positions and is greater than 1×10 7 In some embodiments, each respective growth period of the biological sample is selected from a plurality of growth periods. In some embodiments, the plurality of growth periods is at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, or at least 1 x 10 6 In some embodiments, the plurality of growth periods comprises 1×10 7 pcs or less, 1×10 6 In some embodiments, the number of growth periods is 50-1000, 500-50,000, 10,000-1×10, or 1×10. 6 pcs or 1 x 10 6 ~1×10 7 In some embodiments, the plurality of growth periods consists of 50 or more growth periods, and is greater than 1×10 7 It is within another range that ends with a growth period of less than one.
[0061] In some embodiments, the plurality of Raman spectral measurements is at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, or at least 1×10 6In some embodiments, the plurality of Raman spectral measurements comprises 1×10 7 pcs or less, 1×10 6 In some embodiments, the plurality of Raman spectral measurements includes 50-1000, 500-50,000, 10,000-1×10 Raman spectral measurements. 6 pcs or 1 x 10 6 ~1×10 7 In some embodiments, the plurality of Raman spectral measurements begins with 50 or more Raman spectral measurements and continues through 1×10 7 In some embodiments, each Raman spectral measurement includes one or more traces of fluorescence intensity. In some embodiments, each Raman spectral measurement includes at least one, at least two, at least three, at least four, at least five, or at least 10 traces. In some embodiments, each Raman spectral measurement includes 50 or fewer, 10 or fewer, 5 or fewer, or 3 or fewer traces. In some embodiments, each Raman spectral measurement consists of 1-5, 2-10, or 10-20 traces. In some embodiments, each Raman spectral measurement includes another range of traces starting with one or more traces and ending with 20 or fewer traces.
[0062] In some embodiments, for each trace, the plurality of data points comprises at least 2, at least 5, at least 10, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, or at least 10,000 data points. In some embodiments, the plurality of data points comprises no more than 500,000, no more than 10,000, no more than 1000, no more than 500, no more than 100, no more than 500, no more than 20, or no more than 10 data points. In some embodiments, the plurality of data points consists of 2 to 100, 50 to 2000, 1000 to 10,000, or 10,000 to 500,000 data points. In some embodiments, the plurality of data points falls within another range beginning with two or more data points and ending with no more than 500,000 data points.
[0063] In some embodiments, the corresponding set of features is selected from a group of temporal dynamic features of one or more traces. In some embodiments, the temporal dynamic features of one or more traces are determined by a data analysis method. In some embodiments, the data analysis method applies to one or more traces one or more of the following operations and / or methods: determining a linear slope, determining a plurality of non-linear parameters describing the curvature of one or more traces, determining abrupt changes in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining changes in a frequency domain representation of one or more traces, determining changes in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent.
[0064] Details of an exemplary system are described in conjunction with FIG. 1 , which shows an example block diagram of a computing device 100 of the present disclosure. In some implementations, device 100 includes one or more processing units (CPU(s)) 102 (also referred to as processors), one or more network interfaces 104, a user interface 106, non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communications between system components. Non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while persistent memory 112 typically includes CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. Persistent memory 112 optionally includes one or more storage devices located remotely from CPU(s) 102. Persistent memory 112 and the non-volatile memory device(s) within non-persistent memory 112 include non-transitory computer-readable storage media.In some implementations, the non-persistent memory 111 or alternatively the non-transitory computer-readable storage medium, possibly in conjunction with the persistent memory 112, stores the following programs, modules, and data structures, or a subset thereof: an optional operating system 116 that handles various basic system services and includes procedures for performing hardware-dependent tasks; an optional network communication module (or instructions) 118 for connecting the system 100 with other devices and / or the communication network 104; an optional classifier training module 120 for training models to assess subjects for biological conditions; and an optional classifier training module 120 for training models to assess subjects for biological conditions, including feature data for one or more training objects 124. The system stores an optional data store 122 for a dataset for biological samples, where the feature data includes parameters associated with each of the features 126 and a diagnostic state 128 (e.g., an indication that each training subject has been diagnosed with or has not been diagnosed with a biological condition), an optional classifier validation module 130 for validating a model that distinguishes between biological conditions, an optional data store 132 for a dataset for biological samples from validation subjects, and biological conditions, e.g., as trained using the classifier training module 120, and an optional patient classification module 134 for classifying subjects as having the biological condition.
[0065] In various implementations, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise rearranged in various implementations. In some implementations, non-persistent memory 111 optionally stores a subset of the above-identified modules and data structures. Furthermore, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than that of visualization system 100, which computer system is addressable by visualization system 100, thereby allowing visualization system 100 to retrieve all or a portion of such data when needed.
[0066] In some embodiments, system 100 is connected to or includes one or more analytical devices for performing chemical analyses. For example, optional network communication module (or instructions) 118 is configured to connect system 100 with one or more analytical devices, for example, via communication network 104. In some embodiments, the one or more analytical devices include a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer.
[0067] While Figure 1 depicts "system 100," the diagram is intended more as a functional description of various features that may be present in a computer system, rather than a structural schematic of the implementations described herein. In practice, items shown separately may be combined and some items may be separated, as will be recognized by those skilled in the art. Furthermore, while Figure 1 depicts certain data and modules in non-persistent memory 111, some or all of these data and modules may be stored in persistent memory 112.
[0068] In some embodiments, the methods of the present disclosure include obtaining a biological sample (e.g., a single hair including a hair shaft). In some embodiments, the subject is a human. In some embodiments, the subject is a child 12 years of age or younger (e.g., a child 5, 4, 3, 2, 1, 9, 6, 3, or 1 month of age or younger). In some embodiments, the child is about 12 to about 5 years of age. In some embodiments, the subject is less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year of age. In some embodiments, the subject is at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years of age. In some embodiments, the subject is an adult. In some embodiments, the subject is an adult. Figure 2A shows an example of a hair sample from a subject including a hair shaft. In some embodiments, the hair sample is removed from the subject (e.g., with the aid of scissors). In some embodiments, the method of obtaining the hair sample is non-invasive. In some embodiments, the acquired hair sample has a minimum length of 1 cm (e.g., the hair sample is 1 cm, 2 cm, 3 cm, 4 cm, or 5 cm long). In some embodiments, the acquired hair sample is at least 1 cm, at least 2 cm, at least 3 cm, at least 4 cm, at least 5 cm, at least 6 cm, at least 10 cm, or at least 20 cm long. In some embodiments, the acquired hair sample is 40 cm or less, 20 cm or less, 10 cm or less, or 5 cm or less in length. In some embodiments, the acquired hair sample is 1-5 cm, 4-10 cm, 8-20 cm, or 15-40 cm in length. In some embodiments, the amount of hair sample acquired falls within another range starting from a length of 1 cm or more and ending with a length of 40 cm or less. In some embodiments, the hair sample includes any portion of the hair (e.g., the tip or the portion between the tip and the follicle). There is no special requirement for the hair sample to specifically include the hair follicle. Figure 2B shows an example of a subject's tooth sample. Figure 2C shows an example of a subject's nail sample. In some embodiments, in the example of nails or hair, obtaining a biological sample refers to positioning a subject so that a nail or hair is sampled, hi some embodiments, a nail sample comprises a whole nail or a nail clipping.
[0069] In some embodiments, the obtained biological sample is pretreated, such as by washing and drying the biological sample with one or more solvents and / or surfactants. In some embodiments, if the biological sample is hair, the hair sample is washed in a solution of TRITON X-100® and ultrapure metal-free water (e.g., MILLI-Q® water) and dried in an oven (e.g., at 60°C) overnight. In some embodiments, the pretreatment further includes preparing the hair shaft for measurement by placing it on a glass slide (e.g., a microscope glass slide) with an adhesive film (e.g., double-sided tape). In some embodiments, the hair shaft is positioned so that the hair shaft is substantially straight. In some embodiments, the glass slide containing the hair shaft is placed in or near a measurement system (e.g., a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer) to perform the analysis. In some embodiments, if the biological sample is a tooth or nail, the surface of the biological sample is washed (e.g., with a surfactant, water, or one or more solvents). In some embodiments, the sample is divided into sections and then placed in or near a measurement system (e.g., a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer) for analysis.
[0070] FIG. 3 shows a flowchart of a method 300 for evaluating a subject for a biological condition, such as a method for predicting the subject's diagnostic status with respect to a disease or disorder. In some embodiments, method 300 includes exposing a biological sample of the subject to a light source (as in operation 302). In some cases, the light source includes a laser. In some embodiments, analyzing determines the temporal dynamics of an underlying biological process. In some embodiments, analyzing includes reducing the dimensionality of the multiple Raman spectra (e.g., by independent component analysis) before processing. In some embodiments, the optical signal is generated by a light source (e.g., a laser). In some embodiments, the biological sample includes a tooth sample, a hair sample, or a nail sample. Next, in some embodiments, method 300 includes acquiring multiple Raman spectra from the exposed biological sample (as in operation 304). Next, in some embodiments, method 300 includes processing the multiple Raman spectra (as in operation 306) to generate a spatial map of the multiple Raman spectra. Next, in some embodiments, the method 300 includes predicting a diagnostic status of the subject with respect to a disease or disorder based at least in part on the spatial map of the plurality of Raman spectra (as in operation 308).
[0071] In some embodiments, the plurality of Raman spectra are acquired using a Raman spectroscopic microscope including a 50x air-coupling objective or a 63x water-immersion coupling objective. In some embodiments, the laser includes a wavelength of about 785 nm or a wavelength of about 532 nm. In some embodiments, the acquiring is performed using an integration time of about 0.2 seconds to about 0.3 seconds. In some embodiments, the acquiring includes, after acquiring one Raman spectrum of the plurality of Raman spectra, moving the biological sample with a step size of about 2 microns to about 5 microns.
[0072] In some embodiments, the analyzing comprises generating a temporal Raman profile based at least in part on the acquired Raman spectrum, and analyzing the temporal profile of variation in the Raman spectrum, hi some embodiments, at least a portion of the temporal Raman profile corresponds to a prenatal period of the subject.
[0073] In some embodiments, measurement data is collected from the biological sample sequentially at multiple locations along the biological sample. In some embodiments, the corresponding multiple locations include at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or more than 10000 locations. In some embodiments, the respective locations are adjacent to one another. In this manner, regions corresponding to different locations on the biological sample can be associated with dynamic (e.g., time-varying) abundance measurements. In some embodiments, the respective locations are separated by a predetermined distance. In some embodiments, sampling is performed along a baseline of the biological sample, starting from the respective location closest to the tip of the biological sample, such as a hair sample (e.g., the location corresponding to the subject's youngest age). In general, as long as the sampling direction is known, sampling can start from the respective position closest to the tip or root, and the appropriate trained model is used for analysis.
[0074] In some embodiments, the sampling generates a set of data points. In some embodiments, each set of data points corresponds to measurements (e.g., abundance or concentration) of a substance exhibiting a dynamic biological response measured at multiple locations along the biological sample. In some embodiments, each location on the baseline of the biological sample corresponds to a specific growth time of the biological sample. In some embodiments, in the example of a hair shaft, each location corresponds to an approximately 20-minute hair growth period (e.g., a hair growth period calculated using a laser step size of 5 micrometers and an average rate of hair growth of 1 cm per month). By correlating the multiple locations along the baseline of the biological sample with the corresponding periods of growth, a first data set including multiple traces is obtained. Each trace includes time-dependent abundance measurements (e.g., abundance or concentration) of a substance exhibiting a dynamic biological response measured from the biological sample. For example, the distance between locations can correspond to an estimated growth (e.g., biological time) of the biological sample. For example, abundance may be measured for a hair sample along a distance of 1.2 cm, which corresponds to a biological time of approximately 35 days. In some embodiments, biological time is estimated using the average rate of hair growth (eg, 1 cm per month).
[0075] In some embodiments, data analysis of the Raman spectra includes cosmic ray removal, background correction, spectral normalization, peak fitting, or any combination thereof.
[0076] In some embodiments, data analysis is performed on a trace corresponding to the time-dependent abundance (e.g., time-dependent concentration) of a substance exhibiting a dynamic biological response measured from a biological sample. This may include customized operations to clean up the data (e.g., smoothing the data over a time span and / or removing data points above or below a predetermined threshold). In some embodiments, data analysis includes removing data points from the trace that have a mean absolute difference between adjacent data points that is at least 1, 2, or 3 standard deviations of the mean absolute difference between adjacent data points.
[0077] In some embodiments, the data analysis further comprises a dimensionality reduction step, whereby the high-dimensional array of Raman spectra is decomposed into a lower-dimensional array of derived time-varying components. Methods for dimensionality reduction include independent component analysis (ICA), principal component analysis (PCA), non-negative matrix factorization (NNMF), and related unsupervised and supervised methods.
[0078] In some embodiments, the data analysis further comprises performing recurrence quantification analysis (RQA) on the time-dependent traces or components derived from dimensionality reduction techniques (ICA / PCA) applied to the time-dependent traces to obtain a set of features describing the dynamic periodic properties of the traces, where the RQA measures the variability of the time-dependent traces or components derived from the time-dependent traces. RQA involves estimating features describing periodic characteristics in a given waveform, including recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, recurrence tendency, lamellarity, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, determining a linear slope, determining a plurality of nonlinear parameters describing the curvature of one or more traces, determining an abrupt change in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining a change in a frequency domain representation of one or more traces, determining a change in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, or determining a maximum Lyapunov exponent, and / or any combination thereof. RQA methods and features are described, for example, in Webber et al., “Simpler Methods Do It Better: Success of Recurrence Quantification Analysis as a General Purpose Data Analysis Tool,” Physics Letters A 373, 3753-3756 (2009) and Marwan et al., “Recurrence Plots for the Analysis of Complex Systems,” Physics Reports 438, 237-239 (2007), the contents of each of which are incorporated herein by reference in their entirety.In some embodiments, the time-dependent traces are analyzed using other analytical methods, such as Fourier transforms, wavelet analysis, and cosinor analysis. Such techniques can be applied to derive similar metrics, including spectral analysis of frequency components and their associated power. These metrics and related derivative measures can be used in place of RQA-derived features to analyze time-dependent traces obtained from biological samples for the purpose of predictive classification.
[0079] RQA involves constructing recurrence plots to visualize and analyze the dynamic temporal structure of each acquired trace. Such recurrence plots can illustrate phase dynamic processes in continuous measurements by plotting a given sequence against its time-lagged derivatives. From one-dimensional traces measured from the hair shaft, additional dimensions are computationally derived to embed the traces in a high-dimensional space called a phase diagram, where t refers to the value of the original trace and the dimensions (t + τ) and (t + 2τ) are derived from delaying the original time series by an interval τ. Subsequent analysis is then performed on the embedded phase diagram to construct a recurrence plot and recurrence quantification analysis. Recurrence quantification plots can be derived from phase diagrams by applying a threshold function to each point in the phase diagram. On the corresponding recurrence plot, typically consisting of a square binary matrix represented as white or black space, a given point is assigned a value of 1 at each time interval, and other points in the phase diagram share the spatial limits of the assigned threshold boundary. The RQA method is applied to recurrence plots to examine the intervals of delay between states in a given system, with black dots reflecting the time intervals when the system revisits the same state. Periodic processes in which the system continuously repeats a given pattern of states appear as black diagonal lines in the recurrence plot, while periods of stability appear as square structures, spurious recurrences appear as black dots, and unique events appear as white spaces.
[0080] In some embodiments, recurrence plots are constructed for traces of a single substance or a combination of two substances (e.g., to visualize interactive cyclic patterns of two substances, which may be referred to as cross-recurrence quantitative analysis or joint recurrence quantitative analysis). In some embodiments, recurrence plots are constructed for combinations of three or more substances.
[0081] In some embodiments, the data analysis includes analyzing the recurrence plot to obtain a set of features associated with the recurrence plot. The features, which may be synonymously referred to as "rhythmic features" or "dynamic features," provide quantitative measures describing the periodicity, predictability, and transient nature present in the multiple traces. The features are selected from the set comprising: recurrence rate, determinism, average diagonal length, maximum diagonal length, divergence, Shannon entropy in diagonal length, tendency for recurrence, layering, trapping time, maximum vertical line length, Shannon entropy in vertical line length, average recurrence time, Shannon entropy in recurrence time, number of most likely recurrences, determining a linear slope, determining a plurality of non-linear parameters describing the curvature of one or more traces, determining an abrupt change in intensity of one or more traces, determining one or more changes in baseline intensity of one or more traces, determining a change in a frequency domain representation of one or more traces, determining a change in a power spectral domain representation of one or more traces, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, an estimate of the Lyapunov spectrum, or determining a maximum Lyapunov exponent, and / or any combination thereof.
[0082] In some embodiments, analyzing the data further comprises inputting the obtained set of features into a trained model. In some embodiments, the trained model comprises a predictive computation algorithm for obtaining a probability of the subject having the biological condition. In some embodiments, the predictive computation algorithm performs the following calculations:
number
[0083] In some embodiments, the weighting parameters β,...,β k is defined based on model training. In some embodiments, the probability p(subject) is provided as a number ranging from 0 to 1, where 1 corresponds to a 100% probability that the subject has the biological condition.
[0084] In some embodiments, the data analysis includes applying a threshold to the obtained probability p(subject). If the obtained probability p(subject) is above the predetermined threshold, the subject is assessed as having the biological condition. If the obtained probability is below the threshold, the subject is assessed as not having the biological condition. In some embodiments, the threshold is between about 0.3 and 0.6 (e.g., the predetermined threshold is about 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, or 0.6). In some embodiments, the threshold is at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, or at least 0.7. In some embodiments, the threshold is 0.9 or less, 0.8 or less, 0.7 or less, 0.6 or less, 0.5 or less, or 0.4 or less. In some embodiments, the threshold is within another range starting from 0.1 or more and ending at 0.9 or less. In some embodiments, the value assigned to the probabilistic threshold is determined or estimated during model training through the use of a receiver operating characteristic (ROC) chart, and the optimal threshold used corresponds to the value that results in the largest area under the curve (ROC-AUC). In some embodiments, the obtained probability is expressed in terms of associated odds (e.g., odds ratios (OR), which may be derived from probabilities such as OR=p / (1-p)). For example, assessing includes assessing the odds that a subject has a biological condition.
[0085] In some embodiments, the data analysis includes distinguishing a first biological condition from an alternative condition, e.g., a second biological condition. In some embodiments, the alternative condition is not associated with any known condition (e.g., a neurotypical condition (NT)). In some embodiments, the first biological condition is associated with autism spectrum disorder (ASD) and the alternative condition is associated with attention-deficit / hyperactivity disorder (ADHD). In some embodiments, the alternative condition is any other neurodevelopmental condition, or a comorbid diagnosis of two neurodevelopmental conditions. Thus, in some embodiments, the data analysis is capable of distinguishing between two neurodevelopmental conditions (e.g., distinguishing between autism spectrum disorder and ADHD, or distinguishing between ASD and comorbid (CM) cases diagnosed with both ASD and ADHD).
[0086] Healthcare providers, such as the patient's physician and treatment team, can access patient data (e.g., dynamic biological response data or other health data) and / or predictions or assessments generated from such data. Based on the results of data analysis, the healthcare providers can determine clinical decisions or outcomes.
[0087] For example, a physician may prescribe that a patient undergo one or more clinical trials at a hospital or other clinical setting based, at least in part, on a predicted disease or disorder in the subject. In some embodiments, these prerequisites are provided when certain predetermined criteria (e.g., a minimum threshold for the likelihood of a disease or disorder) are met.
[0088] In some embodiments, such minimum threshold values include, for example, at least about a 5% chance, at least about a 10% chance, at least about a 20% chance, at least about a 25% chance, at least about a 30% chance, at least about a 35% chance, at least about a 40% chance, at least about a 45% chance, at least about a 50% chance, at least about a 55% chance, at least about a 60% chance, at least about a 65% chance, at least about a 70% chance, at least about a 75% chance, at least about a 80% chance, at least about a 85% chance, at least about a 90% chance, at least about a 95% chance, at least about a 96% chance, at least about a 97% chance, at least about a 98% chance, or at least about a 99% chance. In some embodiments, the minimum threshold value is 99% or less, 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, or 40% or less. In some embodiments, the minimum threshold is between 5% and 20%, between 10% and 50%, between 30% and 70%, or between 60% and 99%. In some embodiments, the minimum threshold is within another range starting at or above 5% and ending at or below 99%.
[0089] As another example, a physician may prescribe a therapeutically effective amount of a therapeutic agent (e.g., a drug), clinical procedure, or further clinical trial to be administered to a patient based at least in part on the predicted disease or disorder in the subject. For example, a physician may prescribe an anti-inflammatory therapeutic agent in response to signs of inflammation in a patient.
[0090] Model In some embodiments, the methods and systems of the present disclosure utilize or access external capabilities of artificial intelligence techniques to develop signatures for various diseases or disorders. In some embodiments, these signatures can be used to accurately predict diseases or disorders (e.g., months or years earlier than standard clinical care). Using such predictive capabilities, healthcare providers (e.g., physicians) can make informed and accurate risk-based decisions, thereby improving the quality of care and monitoring provided to patients.
[0091] In some embodiments, the disclosed methods and systems can analyze acquired dynamic biological response data from a subject (patient) to generate a probability that the subject has a disease or disorder. For example, in some embodiments, the system applies a trained (e.g., predictive) algorithm to the acquired dynamic biological response data to generate a probability of the subject having a disease or disorder. In some embodiments, the trained algorithm comprises an artificial intelligence-based model, such as a machine learning-based classifier, configured to process the acquired dynamic biological response data to generate a probability of the subject having a disease or disorder. In some embodiments, the model is trained using clinical datasets from one or more cohorts of patients, e.g., using the patient's clinical health data and / or dynamic biological response data as input and the patient's known clinical health outcome (e.g., disease or disorder) as output to the model.
[0092] In some embodiments, the model comprises one or more machine learning algorithms. Examples of machine learning algorithms include, but are not limited to, support vector machines (SVMs), naive Bayes classification, random forests, neural networks (e.g., deep neural networks (DNNs), recurrent neural networks (RNNs), deep RNNs, long short-term memory (LSTM) recurrent neural networks (RNNs), or gated recurrent units (GRUs), or other supervised or unsupervised machine learning, statistical, or deep learning algorithms for classification and regression. In some embodiments, the model also includes estimation of ensemble models composed of multiple predictive models, utilizing techniques such as gradient boosting in constructing gradient boosted decision trees. In some embodiments, the model is trained using one or more training datasets corresponding to patient data.
[0093] In some embodiments, the training dataset is generated, for example, from one or more cohorts of patients with common clinical characteristics (features) and clinical outcomes (markers). In some embodiments, the training dataset includes a set of features and markers corresponding to the features. In some embodiments, the features correspond to algorithm inputs including dynamic biological response data, patient demographic information derived from electronic medical records (EMRs), and medical observations. In some embodiments, the features include clinical characteristics, such as particular ranges or categories of dynamic biological response data. In some embodiments, the features include patient information such as the patient's age, patient medical history, other medical conditions, current or past medications, and time since last observation. For example, a set of features collected from a given patient at a given time point can collectively act as a signature that can indicate the patient's health state or condition at the given time point.
[0094] For example, in some embodiments, ranges of dynamic biological response data and other health measurements may be represented as multiple disjoint continuous ranges of continuous measurements, and categories of dynamic biological response data and other health measurements may be represented as multiple disjoint sets of measurements (e.g., {"High", "Low"}, {"High", "Normal"}, {"Low", "Normal"}, {"High", "Borderline High", "Normal", "Low"}, etc.). In some embodiments, clinical characteristics also include clinical indicators indicative of the patient's health history, such as a diagnosis of a disease or disorder, previous clinical treatments (e.g., drugs, surgical treatments, chemotherapy, radiation therapy, immunotherapy, etc.), behavioral factors, or other health conditions (e.g., history of high blood pressure, hyperglycemia, hypercholesterolemia or high blood cholesterol, allergic reactions or other adverse reactions, etc.).
[0095] In some embodiments, the indicator comprises a clinical outcome, such as, for example, the presence, absence, diagnosis, or prognosis of a disease or disorder in a subject (e.g., a patient). In some embodiments, the clinical outcome comprises a temporal characteristic associated with the presence, absence, diagnosis, or prognosis of a disease or disorder in a patient. For example, the temporal characteristic may indicate that a patient developed a disease or disorder within a particular period of time after a previous clinical outcome (e.g., discharge from the hospital, administration of a therapeutic agent such as a drug, undergoing a clinical procedure such as surgery, etc.). In some embodiments, such time periods include, for example, about 1 hour, about 2 hours, about 3 hours, about 4 hours, about 6 hours, about 8 hours, about 10 hours, about 12 hours, about 14 hours, about 16 hours, about 18 hours, about 20 hours, about 22 hours, about 24 hours, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 10 days, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 6 months, about 8 months, about 10 months, about 1 year, or more than about 1 year. In some embodiments, the time period is 5 years or less, 1 year or less, 6 months or less, 3 months or less, 1 month or less, 2 weeks or less, 1 week or less, 1 day or less, or 12 hours or less. In some embodiments, the period is between 1 hour and 12 hours, between 12 hours and 24 hours, between 1 day and 7 days, between 1 week and 4 weeks, between 1 month and 12 months, or between 1 year and 5 years. In some embodiments, multiple periods fall within other ranges beginning with greater than 1 hour and ending with less than or equal to 5 hours.
[0096] In some embodiments, the input features are structured by aggregating the data into bins or, alternatively, by using one-hot encoding. In some embodiments, the input also includes feature values or vectors derived from the aforementioned inputs, such as cross-correlations calculated between distinct dynamic biological response data or other measurements over a fixed time period, and discrete derivatives or finite differences between consecutive measurements. In some embodiments, such time periods include, for example, about 1 hour, about 2 hours, about 3 hours, about 4 hours, about 6 hours, about 8 hours, about 10 hours, about 12 hours, about 14 hours, about 16 hours, about 18 hours, about 20 hours, about 22 hours, about 24 hours, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 10 days, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 6 months, about 8 months, about 10 months, about 1 year, or more than about 1 year. In some embodiments, the period is 5 years or less, 1 year or less, 6 months or less, 3 months or less, 1 month or less, 2 weeks or less, 1 week or less, 1 day or less, or 12 hours or less. In some embodiments, the period is 1 hour to 12 hours, 12 hours to 24 hours, 1 day to 7 days, 1 week to 4 weeks, 1 month to 12 months, or 1 year to 5 years. In some embodiments, multiple periods fall within different ranges starting at 1 hour or more and ending at 5 hours or less.
[0097] In some embodiments, training records are constructed from sequences of observations. In some embodiments, such sequences comprise a fixed length to facilitate data processing. For example, the sequences may be zero-padded or selected as an independent subset of a single patient's records.
[0098] In some embodiments, the model processes the input features to generate output values including one or more classifications, one or more predictions, or a combination thereof. For example, such classifications or predictions may include a binary classification of healthy / normal health (e.g., absence of disease or disorder) or adverse health (e.g., presence of disease or disorder), a classification between groups of categorical labels (e.g., “no disease or disorder,” “definite disease or disorder,” and “suspected disease or disorder”), the likelihood (e.g., relative likelihood or probability) of developing a particular disease or disorder, a score indicating the presence of a disease or disorder, a score indicating the level of systemic inflammation experienced by the patient, a “risk factor” indicating the patient's probability of death, a prediction of the time the patient is expected to develop a disease or disorder, and a confidence interval for any numerical prediction. In some embodiments, various machine learning techniques are cascaded such that the outputs of the machine learning techniques are also used as input features to subsequent layers or subsections of the model.
[0099] The model can be trained using a dataset to train the model (e.g., by determining model weights and correlations) to generate real-time classifications or predictions. In some embodiments, such a dataset is large enough to generate statistically significant classifications or predictions. For example, the dataset can include dynamic biological response data and other measurements, as well as a database of de-identified data including dynamic biological response data and other measurements from a hospital or other clinical setting.
[0100] In some embodiments, a dataset is divided into subsets (e.g., separate or overlapping), such as a training dataset, a development dataset, and a test dataset. For example, a dataset may be divided into a training dataset comprising 80% of the dataset, a development dataset comprising 10% of the dataset, and a test dataset comprising 10% of the dataset. In some embodiments, the training dataset comprises about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, the development dataset comprises about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, the test dataset comprises about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of the dataset. In some embodiments, the training set (e.g., training data set) may be selected by random sampling of sets of data corresponding to one or more patient cohorts to ensure independence of sampling. Alternatively, in some embodiments, the training set (e.g., training data set) is selected by proportional sampling of sets of data corresponding to one or more patient cohorts to ensure independence of sampling.
[0101] In some embodiments, the dataset is augmented to increase the number of samples in the training set to improve the accuracy of model predictions and reduce overfitting of the model. For example, data augmentation may include rearranging the order of observations in the training records. To accommodate datasets with missing observations, methods for imputing missing data, such as forward fitting, back fitting, linear interpolation, and multitask Gaussian processes, may be used. The dataset may be filtered to remove confounding factors. For example, a subset of patients may be excluded within the database.
[0102] The model may include one or more neural networks, such as a neural network, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), or a deep RNN. The recurrent neural network may include units, which may be long short-term memory (LSTM) units or gated recurrent units (GRUs). For example, the model may include an algorithmic architecture including a neural network with a set of input features, such as vital signs and other measurements, patient medical history, and / or patient demographics. Neural network techniques, such as dropout or regularization, may be used during model training to prevent overfitting. The neural network may include multiple subnetworks, each configured to generate classifications or predictions of different types of output information (e.g., which may be combined to form the overall output of the neural network). The machine learning model may alternatively utilize statistical or related algorithms, including random forests, classification and regression trees, support vector machines, discriminant analysis, regression techniques, and ensemble and gradient-boosted variations thereof.
[0103] When the model generates a disease or disorder classification or prediction, a notification (e.g., an alert or alarm) may be generated and sent to a healthcare provider, such as a doctor, nurse, or other member of the patient's care team in a hospital. The notification may be sent via an automated phone call, a short message service (SMS) or multimedia message service (MMS) message, an email, or an alert in a dashboard. The notification may include output information such as a disease or disorder prediction, a predicted probability of the disease or disorder, an expected time to onset of the disease or disorder, a confidence interval for the probability or time, or a recommended course of treatment for the disease or disorder.
[0104] To validate the performance of the model, different performance metrics can be generated. For example, the area under the receiver operating curve (AUROC) can be used to determine the diagnostic ability of the model. For example, the model can use an adjustable classification threshold, thereby allowing for adjustable specificity and sensitivity, and the receiver operating curve (ROC) can be used to identify different operating points corresponding to different values of specificity and sensitivity.
[0105] In some cases, such as when the dataset is not large enough, cross-validation may be performed to assess the robustness of the model across different training and testing datasets.
[0106] The following definitions may be used to calculate performance metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), AUPRC, and AUROC. A "false positive" may refer to an outcome in which a positive outcome or result is generated erroneously or prematurely (e.g., before or without the actual onset of a disease or disorder). A "true positive" may refer to an outcome in which a positive outcome or result is correctly generated when the patient has a disease or disorder (e.g., the patient exhibits symptoms of the disease or disorder or the patient's records indicate the disease or disorder). A "false negative" may refer to an outcome in which a negative outcome or result is generated, but the patient has a disease or disorder (e.g., the patient exhibits symptoms of the disease or disorder or the patient's records indicate the disease or disorder). A "true negative" may refer to an outcome in which a negative outcome or result is generated (e.g., before or without the actual onset of a disease or disorder).
[0107] The model may be trained until certain predetermined conditions for accuracy or performance are met, such as having a minimum desired value corresponding to a diagnostic accuracy measure. For example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of a disease or disorder occurring in a subject. As another example, a diagnostic accuracy measure may correspond to a prediction of the likelihood of a disease or disorder for which a subject has previously been treated worsening or recurrence. Examples of diagnostic accuracy measures may include sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, area under the precision-recall curve (AUPRC), and area under the receiver operating characteristic (ROC) curve (AUC) (AUROC), which correspond to diagnostic accuracy in detecting or predicting a disease or disorder.
[0108] For example, in some embodiments, such predetermined conditions include a sensitivity of disease or disorder prediction of, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined conditions include a sensitivity of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined conditions include a sensitivity of 50% to 70%, 60% to 80%, 70% to 90%, or 90% to 100%. In some embodiments, the predetermined conditions include a sensitivity within another range starting at 50% or more and ending at 100% or less.
[0109] As another example, such a predetermined condition, in some embodiments, is a specificity for predicting a disease or disorder that includes, for example, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined condition is a specificity that includes values of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined condition is a specificity that includes values of 50% to 70%, 60% to 80%, 70% to 90%, or 90% to 100%. In some embodiments, the predetermined condition is a specificity that includes values within another range starting from 50% or more and ending at 100% or less.
[0110] As another example, such a predetermined condition, in some embodiments, is a positive predictive value (PPV) for predicting a disease or disorder, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined condition is a PPV of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined condition is a PPV of 50%-70%, 60%-80%, 70%-90%, or 90%-100%. In some embodiments, the predetermined condition is a PPV of 50% or more and 100% or less.
[0111] As another example, such a predetermined condition, in some embodiments, is a negative predictive value (NPV) for predicting a disease or disorder, e.g., at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the predetermined condition is an NPV of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the predetermined condition is an NPV of 50%-70%, 60%-80%, 70%-90%, or 90%-100%. In some embodiments, the predetermined condition is an NPV of 100% or less.
[0112] As another example, such a predetermined condition, in some embodiments, is that the area under the curve (AUC) (AUROC) of a receiver operating characteristic (ROC) curve for predicting a disease or disorder includes a value of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the predetermined condition is that the AUROC includes a value of 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, or 0.60 or less. In some embodiments, the predetermined condition is that the AUROC includes a value of 0.50 to 0.70, 0.60 to 0.80, 0.70 to 0.90, or 0.90 to 1. In some embodiments, the predetermined condition is that the AUROC comprises a value that falls within another range starting from 0.50 or greater and ending at 1 or less.
[0113] As another example, such a predetermined condition, in some embodiments, is that the area under the precision-recall curve (AUPRC) for predicting a disease or disorder comprises a value of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the predetermined condition is that the AUPRC comprises a value of 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, 0.60 or less, or 0.50 or less. In some embodiments, the predetermined condition is that the AUPRC includes a value between 0.10 and 0.40, between 0.30 and 0.70, between 0.60 and 0.90, or between 0.80 and 1. In some embodiments, the predetermined condition is that the AUPRC includes a value that falls within another range starting from greater than or equal to 0.10 and ending at less than or equal to 1.
[0114] In some embodiments, the trained model is trained or configured to predict a disease or disorder with a sensitivity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with a sensitivity of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with a sensitivity of 50% to 70%, 60% to 80%, 70% to 90%, or 90% to 100%. In some embodiments, the model is trained or configured to predict a disease or disorder with a sensitivity within another range starting at 50% or more and ending at 100% or less.
[0115] In some embodiments, the trained model is trained or configured to predict a disease or disorder with a specificity of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with a specificity of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with a specificity of 50% to 70%, 60% to 80%, 70% to 90%, or 90% to 100%. In some embodiments, the model is trained or configured to predict a disease or disorder with a specificity within another range starting at 50% or more and ending at 100% or less.
[0116] In some embodiments, the trained model is trained or configured to predict a disease or disorder with a positive predictive value (PPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with a PPV of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with a PPV of 50%-70%, 60%-80%, 70%-90%, or 90%-100%. In some embodiments, the model is trained or configured to predict diseases or disorders with a PPV that falls within another range, starting at 50% or greater and ending at 100% or less.
[0117] In some embodiments, the trained model is trained or configured to predict a disease or disorder with a negative predictive value (NPV) of at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%. In some embodiments, the model is trained or configured to predict a disease or disorder with an NPV of 100% or less, 99% or less, 90% or less, 80% or less, 70% or less, or 60% or less. In some embodiments, the model is trained or configured to predict a disease or disorder with an NPV of 50%-70%, 60%-80%, 70%-90%, or 90%-100%. In some embodiments, the model is trained or configured to predict diseases or disorders with NPVs that fall within another range, starting at 50% or greater and ending at 100% or less.
[0118] In some embodiments, the trained model is trained or configured to predict a disease or disorder with an area under the receiver operating characteristic (ROC) curve (AUC) (AUROC) of at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the trained model is trained or configured to predict a disease or disorder with an AUROC of 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, or 0.60 or less. In some embodiments, the trained model is trained or configured to predict a disease or disorder with an AUROC of 0.50 to 0.70, 0.60 to 0.80, 0.70 to 0.90, or 0.90 to 1. In some embodiments, the trained model is trained or configured to predict a disease or disorder with an AUROC that falls within another range starting at or above 0.50 and ending at or below 1.
[0119] In some embodiments, the trained model is trained or configured to predict a disease or disorder with an area under the precision-recall curve (AUPRC) of at least about 0.10, at least about 0.15, at least about 0.20, at least about 0.25, at least about 0.30, at least about 0.35, at least about 0.40, at least about 0.45, at least about 0.50, at least about 0.55, at least about 0.60, at least about 0.65, at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.85, at least about 0.90, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99. In some embodiments, the model is trained or configured to predict a disease or disorder with an AUPRC of 1 or less, 0.99 or less, 0.90 or less, 0.80 or less, 0.70 or less, 0.60 or less, or 0.50 or less. In some embodiments, the model is trained or configured to predict diseases or disorders with an AUPRC of 0.10 to 0.40, 0.30 to 0.70, 0.60 to 0.90, or 0.80 to 1. In some embodiments, the model is trained or configured to predict diseases or disorders with an AUPRC that falls within another range starting at or above 0.10 and ending at or below 1.
[0120] In some embodiments, training datasets are collected from training subjects (e.g., humans). Each training subject has a diagnostic status indicating either diagnosed with a biological condition or not diagnosed with a biological condition. In some embodiments, the training subjects are children 12 years of age or younger (e.g., 5, 4, 3, 2, 1, 9, 6, 3, or 1 month or younger). In some embodiments, the children are about 12 to about 5 years old. In some embodiments, the subjects are less than about 12, 11, 10, 9, 8, 7, 5, 4, 3, 2, or 1 year old. In some embodiments, the subjects are at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 years old. In some embodiments, the following training procedure is performed for each training subject in the plurality of training subjects. In some embodiments, the training subject is 18 or younger, 15 or younger, 12 or younger, 11 or younger, 10 or younger, 9 or younger, 8 or younger, 7 or younger, 5 or younger, 4 or younger, 3 or younger, 2 or younger, or 1 or younger.
[0121] In some embodiments, a plurality of baseline locations on the biological sample of the training subject are sampled to thereby obtain a plurality of dynamic biological response samples. Each dynamic biological response sample in the corresponding plurality of dynamic biological response samples is for a different location in the corresponding plurality of locations, and each location in the corresponding plurality of locations represents a different growth period of the corresponding biological sample. Each respective location of the biological response sample is then analyzed (e.g., using a laser ablation inductively coupled plasma mass spectrometer (LA-ICP-MS), a fluorescence image sensor, or a Raman spectrometer) to obtain a plurality of traces. Each trace in the corresponding plurality of traces corresponds to an abundance measurement of a corresponding substance collectively determined over time from the corresponding plurality of dynamic biological response samples.
[0122] Next, respective second data sets are obtained from the corresponding plurality of traces comprising the corresponding set of features, and each respective feature in the corresponding set of features is determined by the change in abundance of one or more substances in the corresponding plurality of traces as assessed by applying recurrence quantification analysis or related methods to the Raman waveforms or dimensions derived from the Raman waveforms via ICA / PCA or related dimensionality reduction techniques.
[0123] Next, an untrained or partially untrained model is generated having (i) a corresponding set of features from the second dataset for each respective training subject in the plurality of training subjects and (ii) a corresponding diagnostic state for each training subject in the plurality of training subjects, the corresponding diagnostic state being selected from among the first diagnostic state and the second diagnostic state, thereby obtaining a trained model. The trained model provides an indication of whether the test subject has the first biological state based on the values of the features from the set of features obtained from the test subject's biological sample. In some embodiments, the trained model is a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, or any combination or variation thereof, particularly including gradient-boosted implementations of the described algorithms, e.g., gradient-boosted decision trees. In some embodiments, the trained machine learning model utilizes a gradient-boosted ensemble algorithm. In some embodiments, the trained model is a multinomial or binary classifier. In some embodiments, the trained model can be used to make a binary prediction as to whether a sample is derived from a subject with a first biological state, or it can be a multinomial prediction that distinguishes subjects without a diagnosis from subjects with the first biological state or a second biological state that is different from the first biological state.
[0124] In some embodiments, the model is a neural network or a convolutional neural network. See Vincent et al., 2010, "Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion," J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, "Exploring strategies for training deep neural networks," J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.
[0125] Independent component analysis (ICA) as described herein for unsupervised dimensionality reduction of Raman waveforms is described in Lee, T.-W. (1998): Independent component analysis: Theory and applications, Boston, Mass: Kluwer Academic Publishers, ISBN 0-7923-8261-7, and Hyvaerinen, A.; Karhunen, J.; Oja, E. (2001): Independent Component Analysis, New York: Wiley, ISBN 978-0-471-40540-5, which are incorporated herein by reference in their entireties.
[0126] Principal component analysis (PCA) as described herein for unsupervised dimensionality reduction of Raman waveforms is described in Jolliffe, IT (2002). Principal Component Analysis. Springer Series in Statistics. New York: Springer-Verlag. doi:10.1007 / b98835. ISBN 978-0-387-95442-4, which is incorporated herein by reference in its entirety.
[0127] SVM is a 5-dimensional model that is used in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge, Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5 thand Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVMs separate a given set of binary labeled data using a hyperplane that is maximally far away from the labeled data. When linear separation is not possible, SVMs can be combined with "kernel" techniques to automatically achieve a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.
[0128] Decision tree classifiers are reviewed by Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Decision tree-based methods divide the feature space into a set of rectangles and fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. One specific algorithm that may be used is a classification and regression tree (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests - Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.
[0129] Clustering (e.g., unsupervised clustering model algorithms and supervised clustering model algorithms) is described in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter "Duda 1973"), pages 211-256, which is incorporated herein by reference in its entirety. As described in Section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings within a data set. To identify natural groupings, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This metric (similarity measure) is used to ensure that samples in one cluster are more similar to each other than samples in other clusters. Second, a mechanism for partitioning the data into clusters using the similarity measure is determined. Similarity measures are discussed in Section 6.7 of Duda 1973, and one way to begin a clustering investigation is to define a distance function and calculate a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities in the same cluster will be significantly shorter than the distance between reference entities in different clusters. However, as described in Duda 1973, p. 215, clustering does not require the use of a distance metric. For example, a non-metric similarity function s(x,x') can be used to compare two vectors x and x'. Traditionally, s(x,x') is a symmetric function whose value is large when x and x' are "similar" in some way. An example of a non-metric similarity function s(x,x') is provided in Duda 1973, p. 218. Once a method for measuring "similarity" or "dissimilarity" between points in a dataset is selected, clustering requires a criterion function to measure the clustering quality of any partition of the data. The partition of the dataset that maximizes the criterion function is used to cluster the data. See Duda 1973, p. 217.Criterion functions are discussed in Section 6.8 of Duda 1973. More recently, Duda et al., Pattern Classification,2. nd edition, John Wiley & Sons, Inc., New York. Clustering is described in detail on pages 537-563. For details on clustering techniques, see Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3rd ed.), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, New Jersey, each of which is incorporated herein by reference. Specific exemplary clustering techniques that can be used in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest neighbor, farthest neighbor, average linkage, centroid, or sum-of-squares algorithms), k-means clustering, fuzzy k-means clustering algorithm, and Jarvis-Patrick clustering. In some embodiments, the clustering comprises unsupervised clustering, where no preconceptions are imposed about which clusters should be formed when the training set is clustered.
[0130] Regression models such as multicategory logit models are described in Agresti, An Introduction to Categorical Data Analysis, 1996, John Wiley & Sons, Inc., New York, Chapter 8, which is incorporated herein by reference in its entirety. In some embodiments, the model utilizes the regressor model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, which is incorporated herein by reference in its entirety. In some embodiments, gradient boosting models are used, for example, for the classification algorithms described herein; these gradient boosting models are described in Boehmke, Bradley; Greenwell, Brandon (2019). "Gradient Boosting." Hands-On Machine Learning with R. Chapman & Hall. pp. 221-245. ISBN 978-1-138-49568-5, which is incorporated herein by reference in its entirety. In some embodiments, ensemble modeling techniques are used, for example, towards the classification algorithms described herein; these ensemble modeling techniques are described in the implementation of the classification models herein and are explained in Zhou Zhihua (2012). Ensemble Methods: Foundations and Algorithms. Chapman and Hall / CRC. ISBN 978-1-439-83003-1, which is incorporated herein by reference in its entirety.
[0131] In some embodiments, the machine learning analysis is performed by a device executing one or more programs (e.g., one or more programs stored in non-persistent memory 111 or persistent memory 112 of FIG. 1 ) that include instructions for performing the data analysis. In some embodiments, the data analysis is performed by a system that includes at least one processor (e.g., processing core 102) and memory (e.g., one or more programs stored in non-persistent memory 111 or persistent memory 112) that include instructions for performing the data analysis.
[0132] Computer Systems The present disclosure provides a computer system programmed to implement the disclosed methods. Figure 4 shows a computer system 401 that is programmed or otherwise configured to, for example, acquire a Raman signature of a tooth sample, analyze the Raman spectrum spatially across the tooth sample, generate a temporal Raman profile, process the data using a trained model, and predict a subject's diagnostic status with respect to a disease or disorder. The computer system 401 can coordinate various aspects of the sensor data analysis of the present disclosure, such as staining the tooth sample, acquiring a fluorescent image of the stained tooth sample, analyzing the fluorescent intensity spatially across the stained tooth sample, generating a temporal Raman profile, measuring the dynamics of the temporal profile, processing the data using a trained model, and predicting a subject's diagnostic status with respect to a disease or disorder. The computer system 401 can be a user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can also be a mobile electronic device.
[0133] The computer system 401 includes a central processing unit (CPU, also including "processor" and "computer processor" herein) 405, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 401 also includes memory or memory locations 410 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 415 (e.g., hard disk), a communication interface 420 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 425, such as cache, other memory, data storage, and / or electronic display adapters. The memory 410, storage unit 415, interface 420, and peripheral devices 425 are in communication with the CPU 405 through a communication bus (solid lines), such as a motherboard. The storage unit 415 may be a data storage unit (or data repository) for storing data. The computer system 401 may be operably coupled to a computer network ("network") 430 with the aid of the communication interface 420. Network 430 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. In some cases, network 430 is a telecommunications and / or data network. Network 430 may include one or more computer servers, which may enable distributed computing such as cloud computing. Network 430, in some cases, may implement a peer-to-peer network with the help of computer system 401, which may enable devices coupled to computer system 401 to act as clients or servers.
[0134] The CPU 405 may execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 410. The instructions may instruct the CPU 405, which may then program or otherwise configure the CPU 405 to implement the methods of the present disclosure. Examples of operations performed by the CPU 405 may include fetch, decode, execute, and write back.
[0135] The CPU 405 may be part of a circuit, such as an integrated circuit. One or more other components of the system 401 may also be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0136] The storage unit 415 may store files such as drivers, libraries, and saved programs. The storage unit 415 may store user data, such as user preferences and user programs. In some cases, the computer system 401 may include one or more additional data storage units external to the computer system 401, such as located on a remote server in communication with the computer system 401 over an intranet or the Internet.
[0137] Computer system 401 can communicate with one or more remote computer systems through network 430. For example, computer system 401 can communicate with a remote computer system of a user (e.g., a healthcare provider). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 401 through network 430.
[0138] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of computer system 401, such as, for example, on memory 410 or electronic storage unit 415. Machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by processor 405. In some cases, the code may be retrieved from storage unit 415 and stored in memory 410 for immediate access by processor 405. In some situations, electronic storage unit 415 may be eliminated, and machine-executable instructions stored in memory 410.
[0139] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or compiled manner.
[0140] Aspects of the systems and methods provided herein, such as computer system 401, may be embodied in programming. Various aspects of the technology may be considered "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in a type of machine-readable medium. The machine-executable code may be stored in an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media may include any or all of the tangible memory of a computer, processor, etc., or associated modules of tangible memory, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory storage for software programming at any time. All or portions of the software may be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another computer, e.g., from a management server or host computer to the computer platform of an application server. Thus, another type of medium that may carry software elements includes optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical terrestrial communication networks, and across various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered software-bearing media. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0141] Thus, a machine-readable medium such as a computer-executable code may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s), which may be used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs, or DVD-ROMs, any other optical media, punched card paper tape, any other physical storage media with a pattern of holes, RAM, ROM, PROMs, and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0142] The computer system 401 can include or communicate with an electronic display 435 that includes a user interface (UI) 440 for providing, for example, fluorescence image data, Raman image data, Raman spectral data, temporal Raman profiles, and models. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0143] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by the central processing unit 405. The algorithms can, for example, acquire Raman images of a tooth sample, analyze the Raman spectra spatially across the tooth sample, generate a temporal Raman profile, process the data using a trained model, and predict the subject's diagnostic status with respect to a disease or disorder. [Example]
[0144] Example 1: Dynamic Raman spectroscopic profiles in dental samples to determine autism spectrum disorder (ASD) disease risk Using the disclosed method and system, dynamic Raman spectroscopic profiles were generated in tooth samples and then analyzed to determine disease risk in subjects. Generally, the temporal dynamics of biological responses (e.g., physiological responses) are found to be imprinted in samples (e.g., tooth samples), and these temporal dynamics can be analyzed to determine disease risk in subjects. Dynamic Raman spectroscopic profiles were generated in two sets of children, one with autism spectrum disorder (ASD) and the other without, during periods including fetal (prenatal) development and early childhood. The dynamic Raman spectroscopic profiles were analyzed to reveal novel features that accurately distinguished autism cases from controls. For example, early-life spectroscopic signatures were found to reveal disease risk for ASD later in life. In contrast, clinical diagnoses of autism are typically made around the age of 3-4 years.
[0145] A primary tooth sample was obtained from each pediatric subject. The tooth sample was segmented, and Raman spectroscopic signals were measured on the tooth sample to develop temporal Raman spectroscopic profiles indicative of prenatal and postnatal physiological responses. The temporal profiles were analyzed using the disclosed machine learning algorithms to train highly accurate classifiers for determining disease risk (e.g., autism).
[0146] Figure 5 shows an example of classifier accuracy for diagnosing autism spectrum disorder (ASD) using features derived from applying RQA to dimensions derived by ICA of Raman waveforms, as illustrated by an empirical receiver operating characteristic (ROC) curve to evaluate the accuracy of the disclosed method for assessing subjects for autism spectrum disorder. ROC curves can be used to evaluate the performance of binary classifiers. ROC curves are plotted as sensitivity (also known as true positive rate) against specificity (also known as true negative rate). A perfect classifier may have 100% sensitivity and specificity, and an area under the curve (AUC) of 1.0. As shown in Figure 5, a classifier configured to determine the presence of ASD in a subject based on dynamic Raman RQA dynamic profiles had an area under the receiver operating characteristic (ROC) curve (AUC) of 0.861, with a 95% confidence interval (CI) of 0.769–0.954. The receiver operating characteristic (ROC) describes how the sensitivity and specificity values of a classifier as varying thresholds are assigned to probabilistic predictions.
[0147] Thus, analysis of Raman spectroscopic signatures using the disclosed methods and systems successfully determined disease risk for autism with an AUC of over 0.86 using only spectroscopic signatures measured on non-invasively obtained biological samples (e.g., dental samples) from pediatric subjects. These results demonstrate that dynamics in human physiology early in life are associated with disease later in life and can be accurately detected and profiled using the disclosed methods and systems.
[0148] Example 2: Dynamic Raman spectroscopic profiles in dental samples to determine amyotrophic lateral sclerosis risk Using the disclosed method and system, dynamic Raman spectroscopic profiles were generated in tooth samples and then analyzed to determine disease risk in subjects. Generally, the temporal dynamics of biological responses (e.g., physiological responses) are found to be imprinted in samples (e.g., tooth samples), and these temporal dynamics can be analyzed to determine disease risk in subjects. Dynamic Raman spectroscopic profiles were generated in two sets of adults, one set with amyotrophic lateral sclerosis (ALS) and the second set without ALS, during periods including childhood and adolescence. The dynamic Raman spectroscopic profiles were analyzed to reveal novel features that accurately distinguished ALS cases from controls. For example, early-life spectroscopic signatures were found to reveal disease risk for ALS later in life.
[0149] Permanent tooth samples were obtained from each adult subject. The tooth samples were segmented, and Raman spectroscopic signals were measured on the tooth samples to develop temporal Raman spectroscopic profiles indicative of physiological responses during childhood and adolescence. The temporal profiles were analyzed using the disclosed machine learning algorithms to train highly accurate classifiers for determining disease risk (e.g., ALS).
[0150] Figure 6 shows an example of classifier accuracy for diagnosing ALS using features derived from applying RQA to dimensions derived by ICA of Raman waveforms, as illustrated by an empirical receiver operating characteristic (ROC) curve to evaluate the accuracy of the disclosed method for assessing subjects with autism spectrum disorder. ROC curves can be used to evaluate the performance of binary classifiers. ROC curves are plotted as sensitivity (also known as true positive rate) against specificity (also known as true negative rate). A perfect classifier would have 100% sensitivity and specificity, and an area under the curve (AUC) of 1.0. As shown in Figure 6, a classifier configured to determine the presence of ASD in a subject based on dynamic Raman RQA dynamic profiles had an area under the ROC curve (AUC) of 0.880, with a 95% confidence interval (CI) of 0.658–1.000. The ROC curve illustrates how the classifier's sensitivity and specificity values are assigned to probabilistic predictions as varying thresholds are used.
[0151] Thus, analysis of Raman spectroscopic signatures using the disclosed methods and systems successfully determined ALS disease risk with an AUC of over 0.88 using only spectroscopic signatures measured on biological samples (e.g., dental samples) from adults. These results demonstrate that dynamics of human physiology early in life are associated with disease later in life and can be accurately detected and profiled using the disclosed methods and systems.
[0152] While methods described elsewhere herein depict steps or sets of operations according to embodiments, one of ordinary skill in the art will recognize many variations based on the teachings provided herein. These steps may be completed in different orders. Steps may be added or omitted. Some steps may include sub-steps. Many steps may be repeated as many times as is beneficial.
[0153] One or more steps of each method or set of operations may be performed using one or more of the circuitry described herein, e.g., logic circuitry such as a processor or programmable array logic for a field programmable gate array. For example, a circuit may be programmed to provide one or more of the steps of each method or set of operations, and the program may include program instructions stored on a computer-readable memory or programmed steps of a logic circuit such as, e.g., programmable array logic or a field programmable gate array.
[0154] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided within the specification. While the present invention has been described with reference to the foregoing specification, the descriptions and illustrations of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention. Accordingly, it is intended that the present invention cover all such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are intended to be covered thereby.
[0155] Embodiment Embodiment 1. A method for predicting a subject's diagnostic status with respect to a disease or disorder, the method comprising: (a) exposing a biological sample of the subject to a light source; (b) acquiring a plurality of Raman spectra from the biological sample; (c) processing the plurality of Raman spectra to generate a spatial map of the plurality of Raman spectra; and (d) predicting the subject's diagnostic status with respect to the disease or disorder based at least in part on the spatial map of the plurality of Raman spectra. Embodiment 2. The method of embodiment 1, wherein the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof. Embodiment 3. The method of embodiment 1 or 2, further comprising detecting or monitoring changes in a temporal stress profile in a spatial map showing the subject's temporal response. Embodiment 4. The method of embodiment 3, wherein the temporal response comprises a biological response, a physiological response, an anatomical response, a therapeutic response, a stress-related response, or a combination thereof. Embodiment 5. The method of any one of embodiments 1-4, wherein the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700. Embodiment 6. The method of any one of embodiments 1 to 5, wherein obtaining comprises using a Raman spectroscopic microscope. Embodiment 7. The method of embodiment 6, wherein the Raman spectroscopic microscope comprises a 50x air-coupled objective, a 63x water-immersion coupled objective, or any combination thereof. Embodiment 8. The method of any one of embodiments 1 to 7, wherein the light source comprises a laser, and the laser comprises a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof. Embodiment 9. The method of any one of embodiments 1 to 8, wherein the acquiring is performed using an integration time of about 0.2 seconds to about 0.3 seconds. Embodiment 10. The method of any one of embodiments 1 to 9, wherein acquiring comprises, after acquiring one Raman spectrum of the plurality of Raman spectra, moving the biological sample with a step size of about 2 microns to about 5 microns. Embodiment 11. The method of any one of embodiments 1-10, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. Embodiment 12. The method of any one of embodiments 1 to 10, wherein the disease or disorder comprises ASD. Embodiment 13. The method of any one of embodiments 1 to 12, wherein predicting the subject's diagnostic status with respect to a disease or disorder comprises processing the spatial map using the trained model. Embodiment 14. The method of embodiment 13, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. Embodiment 15. The method of embodiment 13, wherein the trained model comprises a gradient-boosted ensemble model. Embodiment 16. The method of embodiment 13, wherein the trained model is configured to process one or more features selected from the group consisting of layering, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining the linear slope of the temporal stress profile, determining a plurality of non-linear parameters describing the curvature of the temporal stress profile, determining abrupt changes in the intensity of the temporal stress profile, determining one or more changes in the baseline intensity of the temporal stress profile, determining changes in the frequency domain representation of the temporal stress profile, determining changes in the power spectral domain representation of the temporal stress profile, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating the Lyapunov spectrum, determining the maximum Lyapunov spectrum, and any combination thereof. Embodiment 17. The method of embodiment 3, wherein the trained model is configured to process two or more features selected from the group consisting of layering, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the temporal stress profile, determining a plurality of non-linear parameters describing the curvature of the temporal stress profile, determining abrupt changes in the intensity of the temporal stress profile, determining one or more changes in the baseline intensity of the temporal stress profile, determining changes in a frequency domain representation of the temporal stress profile, determining changes in a power spectral domain representation of the temporal stress profile, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof. Embodiment 18. The method of any one of embodiments 1 to 17, wherein the trained model predicts a diagnostic status for a disease or disorder with a sensitivity of at least about 80%. Embodiment 19. The method of embodiment 1, wherein the trained model predicts the diagnostic status for the disease or disorder with at least about 80% specificity. Embodiment 20. The method of embodiment 1, wherein the trained model predicts a diagnostic status for a disease or disorder with a positive predictive value of at least about 80%. Embodiment 21. The method of embodiment 1, wherein the trained model predicts a diagnostic status for a disease or disorder with a negative predictive value of at least about 80%. Embodiment 22. The method of embodiment 1, wherein the trained model predicts a diagnostic status for a disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.80. Embodiment 23. A device comprising one or more processors and a memory storing one or more programs for execution by the one or more processors, the one or more programs including instructions for: (a) sampling each respective position at a plurality of locations along a reference line on a biological sample of a subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, each Raman spectrum in the plurality of Raman spectra corresponding to a different position in the plurality of locations, each position in the plurality of locations representing a different growth period of the biological sample associated with the Raman signature; (b) analyzing each of the plurality of Raman spectra across the reference line on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from the corresponding plurality of Raman spectral measurements, each respective feature in a corresponding set of features being determined by successive variations in the Raman spectra; and (d) processing the features using a trained model to predict a diagnostic status of the subject for a disease or disorder associated with the Raman signature. Embodiment 24. The device of embodiment 23, wherein the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof. Embodiment 25. The device of embodiment 23 or 24, wherein the instructions further comprise detecting or monitoring changes in the Raman spectrum across multiple locations indicative of a temporal response of the subject. Embodiment 26. The device of embodiment 25, wherein the temporal response comprises a biological response, a physiological response, an anatomical response, a therapeutic response, a stress-related response, or a combination thereof. Embodiment 27. A device according to any one of embodiments 23 to 26, wherein the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700. Embodiment 28. A device described in any one of embodiments 23 to 27, wherein sampling includes using a Raman spectroscopic microscope. Embodiment 29. The device of embodiment 28, wherein the Raman spectroscopic microscope comprises a 50x air-coupled objective, a 63x water-immersion coupled objective, or any combination thereof. Embodiment 30. The device of embodiment 23, wherein sampling comprises exposing the biological sample to a light source to generate one of a plurality of Raman spectra at a plurality of locations. Embodiment 31. The device of embodiment 30, wherein the light source comprises a laser, and the laser comprises a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof. Embodiment 32. A device described in any one of embodiments 23 to 31, wherein the instructions further include translation, and the translation includes moving the biological sample from a first position to a second position among the plurality of positions with a step size of about 2 microns to about 5 microns after acquiring one Raman spectrum among the plurality of Raman spectra. Embodiment 33. A device as described in embodiment 32, wherein the translation is performed using an integration time of about 0.2 seconds to about 0.3 seconds. Embodiment 34. The device of any one of embodiments 23 to 33, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. Embodiment 35. A device described in any one of embodiments 23 to 33, wherein the disease or disorder comprises autism spectrum disorder (ASD). Embodiment 36. A device described in any one of embodiments 23 to 35, wherein predicting the subject's diagnostic status with respect to a disease or disorder comprises processing changes in the Raman spectrum across multiple locations using a trained model. Embodiment 37. The device of embodiment 36, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. Embodiment 38. The device of embodiment 36, wherein the trained model comprises a gradient-boosted ensemble model. Embodiment 39. The device of embodiment 36, wherein the trained model is configured to process one or more features selected from the group consisting of layeredness, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a reference line, determining a plurality of nonlinear parameters describing the curvature of the plurality of Raman spectra across a reference line, determining abrupt changes in the intensity of the plurality of Raman spectra across a reference line, determining one or more changes in baseline intensities of the plurality of Raman spectra across a reference line, determining changes in a frequency domain representation of the plurality of Raman spectra across a reference line, determining changes in a power spectral domain representation of the plurality of Raman spectra across a reference line, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof. Embodiment 40. The device of embodiment 36, wherein the trained model is configured to process two or more features selected from the group consisting of layeredness, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of nonlinear parameters describing the curvature of the plurality of Raman spectra across a baseline, determining abrupt changes in the intensity of the plurality of Raman spectra across a baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across a baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across a baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across a baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof. Embodiment 41. The device of embodiment 23, wherein the trained model predicts a diagnostic status for a disease or disorder with a sensitivity of at least about 80%. Embodiment 42. The device of embodiment 23, wherein the trained model predicts a diagnostic status for a disease or disorder with at least about 80% specificity. Embodiment 43. The device of embodiment 23, wherein the trained model predicts a diagnostic status for a disease or disorder with a positive predictive value of at least about 80%. Embodiment 44. The device of embodiment 23, wherein the trained model predicts a diagnostic status for a disease or disorder with a negative predictive value of at least about 80%. Embodiment 45. The device of embodiment 23, wherein the trained model predicts a diagnostic status for a disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.80. Embodiment 46. A non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium for classification, wherein the one or more computer programs, when executed by a computer system, cause the computer system to: (a) sample each respective position in a plurality of positions along a reference line on a biological sample of the subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, each Raman spectrum in the plurality of Raman spectra corresponding to a different position in the plurality of positions, each position in the plurality of positions representing a different growth period of the biological sample associated with the Raman signature. and (b) analyzing each of a plurality of Raman spectra across a baseline on the biological sample, thereby obtaining a first data set; (c) deriving a respective second data set from a corresponding plurality of Raman spectral measurements, wherein each respective feature in the corresponding set of features is determined by successive variations in the Raman spectra; and (d) processing the features using the trained model to predict a subject's diagnostic status for a disease or disorder associated with the Raman signature, and one or more computer programs embedded in the non-transitory computer-readable storage medium for classification. Embodiment 47. The non-transitory computer-readable storage medium of embodiment 46, wherein the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof. Embodiment 48. A non-transitory computer-readable storage medium as described in embodiment 46 or 47, wherein the method further comprises detecting or monitoring changes in the Raman spectrum across multiple locations indicative of a temporal response of the subject. Embodiment 49. The non-transitory computer-readable storage medium of embodiment 48, wherein the temporal response comprises a biological response, a physiological response, an anatomical response, a therapeutic response, a stress-related response, or a combination thereof. Embodiment 50. The non-transitory computer-readable storage medium of any one of embodiments 46 to 49, wherein the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700. Embodiment 51. A non-transitory computer-readable storage medium described in any one of embodiments 46 to 50, wherein the sampling includes using a Raman spectroscopic microscope. Embodiment 52. A non-transitory computer-readable storage medium as described in embodiment 51, wherein the Raman spectroscopic microscope includes a 50x air-coupled objective, a 63x water-immersion coupled objective, or any combination thereof. Embodiment 53. A non-transitory computer-readable storage medium described in any one of embodiments 46 to 52, wherein sampling includes exposing the biological sample to a light source to generate one Raman spectrum of a plurality of Raman spectra at a plurality of locations. Embodiment 54. A non-transitory computer-readable storage medium as described in embodiment 53, wherein the light source includes a laser, and the laser includes a wavelength of approximately 785 nm, a wavelength of approximately 532 nm, or any combination thereof. Embodiment 55. A non-transitory computer-readable storage medium described in any one of embodiments 46 to 54, wherein the instructions further include translation, and the translation includes moving the biological sample from a first position to a second position among the plurality of positions with a step size of about 2 microns to about 5 microns after acquiring one Raman spectrum among the plurality of Raman spectra. Embodiment 56. A non-transitory computer-readable storage medium as described in embodiment 55, wherein the translation is performed using an integration time of about 0.2 seconds to about 0.3 seconds. Embodiment 57. A non-transitory computer-readable storage medium according to any one of embodiments 46 to 56, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof. Embodiment 58. A non-transitory computer-readable storage medium according to any one of embodiments 46 to 56, wherein the disease or disorder comprises autism spectrum disorder (ASD). Embodiment 59. A non-transitory computer-readable storage medium described in any one of embodiments 46 to 58, wherein predicting a subject's diagnostic status with respect to a disease or disorder comprises processing variations in the Raman spectrum across multiple locations using a trained model. Embodiment 60. The non-transitory computer-readable storage medium of embodiment 59, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. Embodiment 61. A non-transitory computer-readable storage medium as described in embodiment 59, wherein the trained model comprises a gradient-boosted ensemble model. Embodiment 62. The non-transitory computer-readable storage medium of embodiment 59, wherein the trained model is configured to process one or more features selected from the group consisting of layeredness, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across a baseline, determining an abrupt change in the intensity of the plurality of Raman spectra across a baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across a baseline, determining a change in a frequency domain representation of the plurality of Raman spectra across a baseline, determining a change in a power spectral domain representation of the plurality of Raman spectra across a baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof. Embodiment 63. The non-transitory computer-readable storage medium of embodiment 59, wherein the trained model is configured to process two or more features selected from the group consisting of layeredness, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across a baseline, determining abrupt changes in the intensity of the plurality of Raman spectra across a baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across a baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across a baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across a baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof. Embodiment 64. The non-transitory computer-readable storage medium of embodiment 46, wherein the trained model predicts a diagnostic status for a disease or disorder with a sensitivity of at least about 80%. Embodiment 65. The non-transitory computer-readable storage medium of embodiment 46, wherein the trained model predicts a diagnostic status for a disease or disorder with at least about 80% specificity. Embodiment 66. The non-transitory computer-readable storage medium of embodiment 46, wherein the instructions further comprise predicting the subject's diagnostic status for a disease or disorder with a positive predictive value of at least about 80%. Embodiment 67. The non-transitory computer-readable storage medium of embodiment 46, wherein the trained model predicts a diagnostic status for a disease or disorder with a positive predictive value of at least about 80%. Embodiment 68. The non-transitory computer-readable storage medium of embodiment 46, wherein the trained model predicts a diagnostic status for a disease or disorder with a negative predictive value of at least about 80%. Embodiment 69. A method for training a model, in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, comprising: (a) for each respective training subject in a plurality of training subjects, wherein a first subset of the training subjects in the plurality of training subjects have a first diagnostic status corresponding to having a first biological state associated with the Raman signature, and a second subset of the training subjects in the plurality of training subjects have a second diagnostic status corresponding to not having the first biological state associated with the Raman signature, (i) sampling each respective position in a plurality of positions along a reference line on the subject's biological sample associated with the subject's Raman signature, thereby obtaining a plurality of Raman spectra, wherein each Raman spectrum in the plurality of Raman spectra corresponds to a different position in the plurality of positions, and each position in the plurality of positions corresponds to a different growth of the subject's biological sample associated with the Raman signature. (ii) analyzing each Raman spectrum across a baseline on the biological sample, thereby obtaining a first data set; and (iii) deriving respective second data sets from a corresponding plurality of Raman spectra, wherein each respective feature in the corresponding set of features is determined by a continuous variation in the Raman spectrum. (b) training an untrained or partially untrained model using (i) the corresponding set of features of each respective second data set for each training subject in the plurality of training subjects and (ii) a corresponding diagnostic state for each training subject in the plurality of training subjects, the corresponding diagnostic state being selected from among a first diagnostic state and a second diagnostic state, thereby obtaining a trained model that provides an indication as to whether the test subject has a first biological state associated with the Raman signature, based on the values of the features in the set of features acquired from the biological sample associated with the Raman signature of the test subject. Embodiment 70. The method of embodiment 69, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof. Embodiment 71. The method of embodiment 69, wherein the trained model is a multinomial classifier. Embodiment 72. The method of embodiment 69, wherein the trained model is a binary classifier. Embodiment 73. The method of embodiment 69, wherein the first biological condition is selected from the group consisting of autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer. Embodiment 74. The method of any one of embodiments 69 to 73, wherein evaluating the test subject for a first biological state associated with the Raman signature further comprises distinguishing between the first biological state associated with the Raman signature and a second biological state associated with the Raman signature that is different from the first biological state associated with the Raman signature. Embodiment 75. The method of embodiment 74, wherein the first biological condition is an autism spectrum disorder and the second biological condition is attention-deficit / hyperactivity disorder. Embodiment 76. The method of any one of embodiments 69 to 75, wherein the test subject is a human. Embodiment 77. The method of embodiment 76, wherein the human is under 12 years of age. Embodiment 78. The method of embodiment 76, wherein the human is under 1 year old. Embodiment 79. The method of any one of embodiments 69 to 78, wherein the corresponding biological sample associated with each training subject's Raman signature is selected from the group consisting of a hair shaft, a tooth, and a nail. Embodiment 80. The method of embodiment 79, wherein the corresponding biological sample associated with each training subject's Raman signature is a hair shaft, and the reference line corresponds to the longitudinal direction of the hair shaft. Embodiment 81. The method of embodiment 79, wherein the corresponding biological sample associated with each training subject's Raman signature is a tooth, and the reference line corresponds to a direction across the growth zone, including the tooth's neonatal line. Embodiment 82. A method according to any one of embodiments 69 to 81, wherein the corresponding plurality of positions are arranged so that the first position in the corresponding plurality of positions along the corresponding biological sample of each training subject corresponds to the position closest to the tip of the corresponding biological sample of each training subject. Embodiment 83. A method according to any one of embodiments 69 to 82, wherein each trace in a corresponding plurality of Raman spectral measurements includes a plurality of data points, each data point being an instance of a respective position in the plurality of positions. Embodiment 84. The method of any one of embodiments 69 to 83, wherein the corresponding set of features is selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, and Lmax. Embodiment 85. The method of any one of embodiments 69 to 83, wherein the corresponding plurality of positions includes at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or more than 10000 positions. Embodiment 86. The method of any one of embodiments 69 to 85, wherein the trained model is configured to process one or more features selected from the group consisting of layeredness, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining the linear slope of the plurality of Raman spectra across the baseline, determining a plurality of nonlinear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in the intensity of the plurality of Raman spectra across the baseline, determining one or more changes in the baseline intensity of the plurality of Raman spectra across the baseline, determining changes in the frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in the power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating the Lyapunov spectrum, determining the maximum Lyapunov exponent, and any combination thereof. Embodiment 87. The method of any one of embodiments 69 to 85, wherein the trained model is configured to process two or more features selected from the group consisting of layeredness, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining the linear slope of the plurality of Raman spectra across the baseline, determining a plurality of nonlinear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in the intensity of the plurality of Raman spectra across the baseline, determining one or more changes in the baseline intensity of the plurality of Raman spectra across the baseline, determining changes in the frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in the power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating the Lyapunov spectrum, determining the maximum Lyapunov exponent, and any combination thereof.
Claims
1. 1. A method for predicting a subject's diagnostic status with respect to a disease or disorder, comprising: (a) exposing a biological sample of a subject to a light source; (b) acquiring a plurality of Raman spectra from the biological sample; (c) processing the plurality of Raman spectra to generate a spatial map of the plurality of Raman spectra; (d) predicting a diagnostic status of the subject with respect to the disease or disorder based at least in part on the spatial map of the plurality of Raman spectra.
2. 10. The method of claim 1, wherein the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof.
3. 3. The method of claim 1 or 2, further comprising detecting or monitoring changes in a temporal stress profile of the spatial map indicating the subject's temporal response.
4. The method of claim 3 , wherein the temporal response comprises a biological response, a physiological response, an anatomical response, a therapeutic response, a stress-related response, or a combination thereof.
5. 5. The method of claim 1, wherein the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700.
6. The method of any one of claims 1 to 5, wherein obtaining comprises using a Raman spectroscopic microscope.
7. 7. The method of claim 6, wherein the Raman spectromicroscope comprises a 50x air-coupling objective, a 63x water-immersion coupling objective, or any combination thereof.
8. The method of any one of claims 1 to 7, wherein the light source comprises a laser, the laser comprising a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof.
9. The method of any one of claims 1 to 8, wherein said acquiring is performed using an integration time of about 0.2 seconds to about 0.3 seconds.
10. 10. The method of claim 1, wherein the acquiring comprises, after acquiring a Raman spectrum of the plurality of Raman spectra, moving the biological sample with a step size of about 2 microns to about 5 microns.
11. 11. The method of any one of claims 1 to 10, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof.
12. The method of any one of claims 1 to 10, wherein the disease or disorder comprises ASD.
13. 13. The method of any one of claims 1 to 12, wherein predicting the subject's diagnostic status with respect to the disease or disorder comprises processing the spatial map using a trained model.
14. 14. The method of claim 13, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof.
15. The method of claim 13 , wherein the trained model comprises a gradient-boosted ensemble model.
16. 4. The method of claim 3, wherein the trained model is configured to process one or more features selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the temporal stress profile, determining a plurality of non-linear parameters describing the curvature of the temporal stress profile, determining abrupt changes in intensity of the temporal stress profile, determining one or more changes in baseline intensity of the temporal stress profile, determining changes in a frequency domain representation of the temporal stress profile, determining changes in a power spectral domain representation of the temporal stress profile, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.
17. 17. The method of claim 16, wherein the trained model is configured to process two or more features selected from the group consisting of: laminarity, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the temporal stress profile, determining a plurality of non-linear parameters describing the curvature of the temporal stress profile, determining abrupt changes in intensity of the temporal stress profile, determining one or more changes in baseline intensity of the temporal stress profile, determining changes in a frequency domain representation of the temporal stress profile, determining changes in the power spectral domain representation of the temporal stress profile, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.
18. 18. The method of any one of claims 1 to 17, wherein the trained model predicts a diagnostic status for the disease or disorder with a sensitivity of at least about 80%.
19. 10. The method of claim 1, wherein the trained model predicts a diagnostic status for the disease or disorder with at least about 80% specificity.
20. 10. The method of claim 1, wherein the trained model predicts a diagnostic status for the disease or disorder with a positive predictive value of at least about 80%.
21. 10. The method of claim 1, wherein the trained model predicts a diagnostic status for the disease or disorder with a negative predictive value of at least about 80%.
22. 10. The method of claim 1, wherein the trained model predicts a diagnostic status for the disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.
80.
23. 1. A device comprising one or more processors and a memory storing one or more programs for execution by said one or more processors, said one or more programs comprising: (a) sampling each respective position in a plurality of positions along a reference line on a biological sample of the subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, each Raman spectrum in the plurality of positions corresponding to a different position in the plurality of positions, each position in the plurality of positions representing a different growth period of the biological sample associated with the Raman signature; (b) analyzing each of the plurality of Raman spectra across a reference line on the biological sample, thereby obtaining a first data set; (c) deriving respective second data sets from corresponding plurality of Raman spectral measurements, each respective feature in the corresponding set of features being determined by successive variations in the Raman spectra; (d) processing the features using the trained model to predict a subject's diagnostic status for a disease or disorder associated with the Raman signature.
24. 24. The device of claim 23, wherein the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof.
25. 25. The device of claim 23 or 24, wherein the instructions further comprise detecting or monitoring changes in the Raman spectrum across the plurality of locations indicative of a temporal response of the subject.
26. 26. The device of claim 25, wherein the temporal response comprises a biological response, a physiological response, an anatomical response, a therapeutic response, a stress-related response, or a combination thereof.
27. 27. The device of any one of claims 23 to 26, wherein the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700.
28. A device according to any one of claims 23 to 27, wherein sampling comprises using a Raman spectroscopic microscope.
29. 30. The device of claim 28, wherein the Raman spectromicroscope comprises a 50x air-coupling objective, a 63x water-immersion coupling objective, or any combination thereof.
30. 24. The device of claim 23, wherein the sampling comprises exposing the biological sample to a light source to generate the Raman spectrum of the plurality of Raman spectra at the plurality of locations.
31. 31. The device of claim 30, wherein the light source comprises a laser, the laser comprising a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof.
32. 32. The device of claim 23, wherein the instructions further comprise translating, the translating comprising, after acquiring a Raman spectrum of the plurality of Raman spectra, moving the biological sample from a first position to a second position of the plurality of positions with a step size of about 2 microns to about 5 microns.
33. 33. The device of claim 32, wherein the translation is performed using an integration time of about 0.2 seconds to about 0.3 seconds.
34. 34. The device of any one of claims 23 to 33, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof.
35. The device of any one of claims 23 to 33, wherein the disease or disorder comprises autism spectrum disorder (ASD).
36. 36. The device of any one of claims 23 to 35, wherein predicting the subject's diagnostic status with respect to the disease or disorder comprises processing variations in the Raman spectrum across the plurality of locations with a trained model.
37. 37. The device of claim 36, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof.
38. 37. The device of claim 36, wherein the trained model comprises a gradient-boosted ensemble model.
39. 37. The device of claim 36, wherein the trained model is configured to process one or more features selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in intensity of the plurality of Raman spectra across the baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across the baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.
40. 37. The device of claim 36, wherein the trained model is configured to process two or more features selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in intensity of the plurality of Raman spectra across the baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across the baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.
41. 24. The device of claim 23, wherein the trained model predicts a diagnostic status for the disease or disorder with a sensitivity of at least about 80%.
42. 24. The device of claim 23, wherein the trained model predicts a diagnostic status for the disease or disorder with at least about 80% specificity.
43. 24. The device of claim 23, wherein the trained model predicts a diagnostic status for the disease or disorder with a positive predictive value of at least about 80%.
44. 24. The device of claim 23, wherein the trained model predicts a diagnostic status for the disease or disorder with a negative predictive value of at least about 80%.
45. 24. The device of claim 23, wherein the trained model predicts a diagnostic status for the disease or disorder with an area under the receiver operating characteristic curve (AUROC) of at least about 0.
80.
46. A non-transitory computer-readable storage medium and one or more computer programs embedded in the non-transitory computer-readable storage medium for classification, the one or more computer programs, when executed by a computer system, causing the computer system to: (a) sampling each respective position in a plurality of positions along a reference line on a biological sample of the subject associated with a Raman signature of the subject, thereby obtaining a plurality of Raman spectra, each Raman spectrum in the plurality of positions corresponding to a different position in the plurality of positions, each position in the plurality of positions representing a different growth period of the biological sample associated with the Raman signature; (b) analyzing each of the plurality of Raman spectra across a reference line on the biological sample, thereby obtaining a first data set; (c) deriving respective second data sets from corresponding plurality of Raman spectral measurements, each respective feature in the corresponding set of features being determined by successive variations in the Raman spectra; and (d) processing the features using the trained model to predict a subject's diagnostic status for a disease or disorder associated with the Raman signature.
47. 47. The non-transitory computer-readable storage medium of claim 46, wherein the biological sample comprises a tooth sample, a hair sample, a nail sample, or any combination thereof.
48. 48. The non-transitory computer-readable storage medium of claim 46 or 47, wherein the method further comprises detecting or monitoring changes in the Raman spectrum across the plurality of locations indicative of a temporal response of the subject.
49. 49. The non-transitory computer-readable storage medium of claim 48, wherein the temporal response comprises a biological response, a physiological response, an anatomical response, a therapeutic response, a stress-related response, or a combination thereof.
50. 50. The non-transitory computer-readable storage medium of any one of claims 46 to 49, wherein the plurality of Raman spectra comprises wavenumbers from about 200 to about 3700.
51. 51. The non-transitory computer-readable storage medium of any one of claims 46 to 50, wherein sampling comprises using a Raman spectroscopic microscope.
52. 52. The non-transitory computer-readable storage medium of claim 51, wherein the Raman spectroscopic microscope comprises a 50x air-coupling objective, a 63x water-immersion coupling objective, or any combination thereof.
53. 53. The non-transitory computer-readable storage medium of any one of claims 46-52, wherein sampling comprises exposing the biological sample to a light source to generate the Raman spectrum of the plurality of Raman spectra at the plurality of locations.
54. 54. The non-transitory computer-readable storage medium of claim 53, wherein the light source comprises a laser, the laser comprising a wavelength of about 785 nm, a wavelength of about 532 nm, or any combination thereof.
55. 55. The non-transitory computer-readable storage medium of any one of claims 46-54, wherein the instructions further comprise translating, the translating comprising, after acquiring a Raman spectrum of the plurality of Raman spectra, moving the biological sample from a first position to a second position of the plurality of positions with a step size of about 2 microns to about 5 microns.
56. 56. The non-transitory computer-readable storage medium of claim 55, wherein the translation is performed using an integration time of about 0.2 seconds to about 0.3 seconds.
57. 57. The non-transitory computer-readable storage medium of any one of claims 46 to 56, wherein the disease or disorder comprises autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, pediatric cancer, or any combination thereof.
58. 57. The non-transitory computer-readable storage medium of any one of claims 46 to 56, wherein the disease or disorder comprises an autism spectrum disorder (ASD).
59. 59. The non-transitory computer-readable storage medium of any one of claims 46 to 58, wherein predicting the subject's diagnostic status with respect to the disease or disorder comprises processing variations in the Raman spectrum across the plurality of locations with a trained model.
60. 60. The non-transitory computer-readable storage medium of claim 59, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof.
61. 60. The non-transitory computer-readable storage medium of claim 59, wherein the trained model comprises a gradient-boosted ensemble model.
62. 60. The non-transitory computer-readable storage medium of claim 59, wherein the trained model is configured to process one or more features selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in the intensities of the plurality of Raman spectra across the baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across the baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.
63. 60. The non-transitory computer-readable storage medium of claim 59, wherein the trained model is configured to process two or more features selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in intensity of the plurality of Raman spectra across the baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across the baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross-recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.
64. 47. The non-transitory computer-readable storage medium of claim 46, wherein the trained model predicts a diagnostic status for the disease or disorder with at least about 80% sensitivity.
65. 47. The non-transitory computer-readable storage medium of claim 46, wherein the trained model predicts a diagnostic status for the disease or disorder with at least about 80% specificity.
66. 47. The non-transitory computer-readable storage medium of claim 46, wherein the instructions further comprise predicting the subject's diagnostic status for the disease or disorder with a positive predictive value of at least about 80%.
67. 47. The non-transitory computer-readable storage medium of claim 46, wherein the trained model predicts a diagnostic status for the disease or disorder with a positive predictive value of at least about 80%.
68. 47. The non-transitory computer-readable storage medium of claim 46, wherein the trained model predicts a diagnostic status for the disease or disorder with a negative predictive value of at least about 80%.
69. 1. A method for training a model, comprising:
1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, (a) for each respective training object in a plurality of training objects, a first subset of the training objects in the plurality of training objects having a first diagnostic status corresponding to having a first biological state associated with a Raman signature, and a second subset of the training objects in the plurality of training objects having a second diagnostic status corresponding to not having the first biological state associated with the Raman signature; (i) sampling each respective position in a plurality of positions along a reference line on the biological sample of the object associated with the Raman signature of the object, thereby obtaining a plurality of Raman spectra, each Raman spectrum in the plurality of Raman spectra corresponding to a different position in the plurality of positions, each position in the plurality of positions representing a different growth period of the biological sample of the object associated with the Raman signature; (ii) analyzing each Raman spectrum across a reference line on the biological sample, thereby obtaining a first data set; (iii) deriving respective second data sets from the corresponding plurality of Raman spectra, each respective feature in the corresponding set of features being determined by a successive variation in the Raman spectra; (b) training an untrained or partially untrained model using (i) the corresponding set of features of each respective second data set for each training subject in the plurality of training subjects, and (ii) a corresponding diagnostic state for each training subject in the plurality of training subjects, the corresponding diagnostic state being selected from among the first diagnostic state and the second diagnostic state, thereby obtaining a trained model that provides an indication as to whether the test subject has the first biological state associated with the Raman signature based on feature values in a set of features acquired from a biological sample associated with the Raman signature of a test subject.
70. 70. The method of claim 69, wherein the trained model is selected from the group consisting of a neural network algorithm, a support vector machine algorithm, a decision tree algorithm, an unsupervised clustering algorithm, a supervised clustering algorithm, a regression algorithm, a gradient boosting algorithm, and any combination thereof.
71. 70. The method of claim 69, wherein the trained model is a multinomial classifier.
72. 70. The method of claim 69, wherein the trained model is a binary classifier.
73. 70. The method of claim 69, wherein the first biological condition is selected from the group consisting of autism spectrum disorder (ASD), attention-deficit / hyperactivity disorder (ADHD), amyotrophic lateral sclerosis (ALS), schizophrenia, irritable bowel disease (IBD), pediatric kidney disease, kidney transplant rejection, and pediatric cancer.
74. 74. The method of any one of claims 69-73, wherein evaluating the test subject for the first biological state associated with a Raman signature further comprises distinguishing between the first biological state associated with the Raman signature and a second biological state associated with the Raman signature that is different from the first biological state associated with the Raman signature.
75. 75. The method of claim 74, wherein the first biological condition is an autism spectrum disorder and the second biological condition is attention-deficit / hyperactivity disorder.
76. 76. The method of any one of claims 69 to 75, wherein the test subject is a human.
77. 77. The method of claim 76, wherein the human is under 12 years of age.
78. 77. The method of claim 76, wherein the human is less than 1 year old.
79. 79. The method of any one of claims 69 to 78, wherein the corresponding biological sample associated with the Raman signature of each training subject is selected from the group consisting of a hair shaft, a tooth, and a nail.
80. 80. The method of claim 79, wherein the corresponding biological sample associated with the Raman signature of each training subject is the hair shaft, and the reference line corresponds to the longitudinal direction of the hair shaft.
81. 80. The method of claim 79, wherein the corresponding biological sample associated with the Raman signature of each training subject is the tooth, and the reference line corresponds to a direction across a growth zone, including a neonatal line, of the tooth.
82. 82. The method of any one of claims 69-81, wherein the corresponding plurality of locations along the corresponding biological specimen of each training object is arranged such that a first location in the corresponding plurality of locations along the corresponding biological specimen of each training object corresponds to a location closest to a leading edge of the corresponding biological specimen of each training object.
83. 83. The method of any one of claims 69 to 82, wherein each trace in a corresponding plurality of Raman spectral measurements comprises a plurality of data points, each data point being an instance of a respective location in the plurality of locations.
84. 84. The method of any one of claims 69 to 83, wherein the corresponding set of features is selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, and Lmax.
85. 84. The method of any one of claims 69 to 83, wherein the corresponding plurality of locations comprises at least 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, or more than 10000 locations.
86. 86. The method of any one of claims 69 to 85, wherein the trained model is configured to process one or more features selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in the intensities of the plurality of Raman spectra across the baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across the baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.
87. 86. The method of any one of claims 69 to 85, wherein the trained model is configured to process two or more features selected from the group consisting of stratification, entropy, trapping time (TT), mean diagonal length (MDL), recurrence time (RT), Vmax, determinism, Lmax, determining a linear slope of the plurality of Raman spectra across a baseline, determining a plurality of non-linear parameters describing the curvature of the plurality of Raman spectra across the baseline, determining abrupt changes in the intensities of the plurality of Raman spectra across the baseline, determining one or more changes in baseline intensities of the plurality of Raman spectra across the baseline, determining changes in a frequency domain representation of the plurality of Raman spectra across the baseline, determining changes in a power spectral domain representation of the plurality of Raman spectra across the baseline, determining one or more recurrence quantification analysis parameters, determining one or more cross recurrence quantification analysis parameters, determining one or more joint recurrence quantification analysis parameters, determining one or more multidimensional recurrence quantification analysis parameters, estimating a Lyapunov spectrum, determining a maximum Lyapunov exponent, and any combination thereof.