Methods and systems for sample classification

A computer-implemented method using microsatellite allele analysis and machine learning generates a genetic classifier for precise phenotype prediction, addressing the limitations of existing methods by achieving high sensitivity and specificity in identifying diseases and predicting therapeutic responses.

WO2025199499A1PCT designated stage Publication Date: 2025-09-25ORBIT GENOMICS INC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/021017
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-22
Filing Date
2025-03-21
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Current methods are inadequate in accurately identifying genetic markers for phenotypes such as diseases or conditions using microsatellite alleles, particularly in predicting therapeutic responses and genomic instability, with low sensitivity and specificity.

Method used

A computer-implemented method involving microsatellite allele analysis, including statistical testing and machine learning algorithms, to generate a genetic classifier that differentiates subjects with and without a phenotype based on microsatellite allele lengths, using an optimal cutoff value derived from ROC analysis and weight adjustments for allele significance.

Benefits of technology

The method achieves high sensitivity and specificity in predicting phenotypes like cancer, cardiac diseases, and neurological disorders, enabling targeted therapies with positive predictive values above 60% and specificities above 70%, and allows monitoring therapeutic responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025021017_25092025_PF_FP_ABST
    Figure US2025021017_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are methods and systems for determining whether a subject has or will develop a phenotype using lengths of microsatellite alleles at microsatellite loci measured in a sample obtained from the subject. The present disclosure also provides methods and systems for the discovery and validation of tests for identifying novel genetic classifiers for clinical use. Such genetic classifiers are useful for predicting a likelihood that a subject has or will develop a phenotype, such as a disease or a condition. The genetic classifiers may also be used for treating, selecting a subject for treatment, or optimizing treatment for a subject to treat a disease or a condition, based at least in part on the lengths of the microsatellites present in a biological sample obtained from the subject.
Need to check novelty before this filing date? Find Prior Art

Description

WSGR Docket No.56034-703.601 METHODS AND SYSTEMS FOR SAMPLE CLASSIFICATION CROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 568,802, filed on March 22, 2024, which is incorporated herein by reference, in its entirety. INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 56034-703_601_SL.xml, created March 20, 2025, which is 651,264 bytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety. SUMMARY

[0003] Aspects disclosed herein provide, in some embodiments, computer-implemented methods of identifying a genetic classifier for determining whether a subject has a phenotype, the method comprising: (a) receiving sequence data from a plurality of samples obtained from a plurality of subjects comprising subjects that have the phenotype and subjects that do not have the phenotype, wherein the sequence data comprises a plurality of microsatellite alleles at microsatellite loci; (b) identifying a distribution of lengths of the plurality of microsatellite alleles at microsatellite loci; (c) performing a statistical test comprising tabulating the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci to determine which of the plurality of microsatellite alleles at microsatellite loci is predictive of the phenotype; (d) generating a score for a subset of microsatellite alleles at microsatellite loci of the plurality of microsatellite loci determined to be predictive of the phenotype based, at least in part, on the performing the statistical significance test in (c); (e) determining an optimal cutoff value of the score for the subset of microsatellite loci that differentiates the subjects that have the phenotype from the subjects that do not have the phenotype; (f) applying the optimal cutoff value to the score for the subset of microsatellite alleles at microsatellite loci, thereby producing the genetic classifier. In some embodiments, the optimal cutoff value is derived from a second statistical test performed for the subset of microsatellite alleles at microsatellite loci. In some embodiments, the optimal cutoff value has a combined P-value of at most about 0.05 for differentiating the subjects that have theWSGR Docket No.56034-703.601 phenotype from the subjects that do not have the phenotype. In some embodiments, the second statistical test comprises performing a receiver operating characteristic (ROC) analysis. In some embodiments, the optimal cutoff value is derived from Youden’s index. In some embodiments, the determining the optimal cutoff value in (e) and the applying the optimal cutoff value in (f) are performed with a machine learning algorithm. In some embodiments, the machine learning algorithm comprises an artificial neural network or a random forest algorithm. In some embodiments, method further comprise training the machine learning algorithm with training sequence data from a training cohort of subjects having the phenotype and not having the phenotype. In some embodiments, methods further comprise validating the machine learning algorithm with validation sequence data from a validation cohort of subjects having the phenotype and not having the phenotype. In some embodiments, the identifying the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci in (b) comprises aligning the plurality of microsatellite alleles at microsatellite loci with reference to a human genome using an alignment algorithm comprising a scoring matrix configured to align flanking regions of the plurality of microsatellite alleles at microsatellite loci, and wherein the flanking regions have a length that is greater than or equal to about 25 contiguous base pairs. In some embodiments, the length is greater than or equal to about 50 contiguous base pairs. In some embodiments, the length is about 100 contiguous base pairs in length. In some embodiments, the alignment algorithm comprises improved gap open penalties (GOP), gap extension penalties (GEP), or any combination thereof, relative to a reference alignment algorithm with a scoring matrix configured to align flanking regions of the plurality of microsatellite alleles at microsatellite loci having a length of fewer or equal to 20 contiguous base pairs. In some embodiments, the reference algorithm is a Needleman–Wunsch algorithm. In some embodiments, the reference algorithm is a Smith-Waterman algorithm. In some embodiments, the alignment algorithm is a hybrid algorithm comprising components of a Needleman–Wunsch algorithm and a Smith- Waterman algorithm. In some embodiments, methods further comprise applying a weight to microsatellite alleles at a microsatellite locus of the subset of microsatellite alleles at microsatellite loci. In some embodiments, the weight is based on a strength of an association of microsatellite alleles at the microsatellite locus to the phenotype relative to other microsatellite alleles at microsatellite loci of the subset. In some embodiments, the weight is higher if the microsatellite allele at the microsatellite locus is a primary microsatellite allele than if the microsatellite allele at the microsatellite locus is a minor microsatellite allele. In some embodiments, the weight is higher if the microsatellite allele at the microsatellite locusWSGR Docket No.56034-703.601 comprises a single nucleotide polymorphism (SNP) than if the microsatellite allele at the microsatellite locus lacks the SNP. In some embodiments, the weight is higher if the microsatellite allele at the microsatellite locus comprises a CpG dinucleotide than if the microsatellite allele at the microsatellite locus lacks the CpG dinucleotide, and wherein the CpG dinucleotide is a potential methylation site. In some embodiments, the subset of microsatellite alleles at microsatellite loci comprises one or more minor microsatellite alleles; and / or the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci comprises the one or more minor microsatellite alleles. In some embodiments, the genetic classifier additionally differentiates subjects having instability at one or more microsatellite loci of the one or more minor microsatellite alleles from other subjects of the plurality of subjects. In some embodiments, the subset of microsatellite alleles at microsatellite loci comprises one or more primary microsatellite alleles and one or more minor microsatellite alleles; and / or the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci comprises of the one or more primary microsatellite alleles and the one or more minor microsatellite alleles. In some embodiments, the subset of microsatellite alleles at microsatellite loci comprises one or more microsatellite loci that comprises a SNP; and / or the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci comprises of the one or more microsatellite loci that comprises the SNP. In some embodiments, the SNP is associated with a risk that a subject has or will develop the phenotype relative to another subject that lacks the SNP. In some embodiments, the statistical test is a Fisher test. In some embodiments, the statistical test is a regression analysis. In some embodiments, the statistical test is a chi-squared test. In some embodiments, the tabulating the lengths of the plurality of microsatellite alleles at microsatellite loci comprises coercing sequence data for the plurality of microsatellite alleles at microsatellite loci into a 2 x 2 table for each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci. In some embodiments, the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprises tabulating genotypes for each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci based on which genotypes are most common in the subjects that have the phenotype. In some embodiments, the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprises tabulating genotypes for each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci based on an average length of microsatellite alleles at microsatellite loci in each microsatellite locus. In some embodiments, the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprisesWSGR Docket No.56034-703.601 tabulating genotypes for each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci based on a median length of microsatellite alleles at microsatellite loci in each microsatellite locus. In some embodiments, the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprises: aggregating the plurality of microsatellite alleles at microsatellite loci for each microsatellite locus into long alleles and short alleles relative to a cutoff length; and tabulating the plurality of microsatellite alleles at microsatellite loci for each microsatellite locus into the 2 x 2 table based on whether a microsatellite locus of the plurality of microsatellite alleles at microsatellite loci is longer or shorter than the cutoff length. In some embodiments, the cutoff length is determined using an ROC analysis per microsatellite locus. In some embodiments, the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a positive predictive value (PPV) of at least about 51%. In some embodiments, the PPV is at least about 60%. In some embodiments, the PPV is at least about 70%. In some embodiments, the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a sensitivity of at least about 70%. In some embodiments, the sensitivity is at least about 80%. In some embodiments, the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a specificity value of at least about 70%. In some embodiments, the specificity is at least about 80%. In some embodiments, the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a Confidence Interval (CI) of at least about 0.90. In some embodiments, the phenotype is a disease or a condition. In some embodiments, the disease is: a cancer; a neoplasm; a cardiac disease, or a neurological disease. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the neurological disease is schizophrenia, bipolar disorder, autism, or Alzheimer’s disease. In some embodiments, the plurality of subjects comprises greater than or equal to about 20,000 subjects. In some embodiments, about half of the plurality of subjects have the phenotype, and the other half of the plurality of subjects do not have the phenotype; or a distribution of the plurality of subjects between a first fraction of the plurality of subjects that have the phenotype and a second fraction of the plurality of subjects that do not have the phenotype reflects the distributions in the general population. In some embodiments, methods further comprise providing: the optimal cutoff value; the subset of the microsatellite alleles at microsatellite loci; and an odds ratio for eachWSGR Docket No.56034-703.601 microsatellite allele at each microsatellite locus of the subset of microsatellite alleles at microsatellite loci. In some embodiments, methods further comprise excluding microsatellite alleles at microsatellite loci from the subset of microsatellite alleles at microsatellite loci having a call rate of less than 80%. In some embodiments, the call rate is calculated by dividing a total number of genotypes characterized by a microsatellite allele at a microsatellite locus length by a total number of microsatellite alleles at microsatellite locus lengths observed for the plurality of microsatellite alleles at microsatellite loci. In some embodiments, the phenotype comprises a condition, and wherein the condition comprises: a rate of metabolism of a therapeutic agent in the subject; a positive therapeutic response of the subject to the therapeutic agent; a therapeutic non-response or loss of response of the subject to the therapeutic agent; an allergy; an adverse effect to a therapeutic agent; a sensitivity to a chemical or environmental agent associated with cancer; obesity; or substance abuse.

[0004] Aspects disclosed herein provide, in some embodiments, computer-implemented systems comprising a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the computing device to perform a method of the present disclosure.

[0005] Aspects disclosed herein provide, in some embodiments, systems comprising: the computer-implemented system of the present disclosure; and a nucleic acid sequencer configured to transmit the sequence data to the at least one processor.

[0006] Aspects disclosed herein provide, in some embodiments, non-transitory computer- readable storage media encoded with a computer program including instructions executable by one or more processors to determine a probability of a subject as having or developing a phenotype, the non-transitory computer-readable storage media comprising: a database, in a computer memory, of the instructions comprising a method of the present disclosure; and a software module configured to perform the instructions.

[0007] Aspects disclosed herein provide, in some embodiments, methods of treating the disease or the condition in a subject, the method comprising: administering a therapy to the subject to treat the disease or the condition, wherein the subject is predicted to have the phenotype with the genetic classifier identified by a method of the present disclosure, wherein the phenotype comprises a positive therapeutic response to the therapy. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is aWSGR Docket No.56034-703.601 cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolism of the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic response to the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In some embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample, and wherein the one or more microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles at microsatellite loci per microsatellite locus; (b) a median length of all microsatellite alleles at microsatellite loci per microsatellite locus; (c) a longest length of all microsatellite alleles at microsatellite loci per microsatellite locus; (d) a shortest length of all microsatellite alleles at microsatellite loci per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering aWSGR Docket No.56034-703.601 therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy, a gene editing therapy, or a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is a cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is a cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0008] Aspects disclosed herein provide, in some embodiments, methods of predicting whether a subject has or will develop the disease or the condition, the method comprising: determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; and applying the genetic classifier identified by a method of the present disclosure to the lengths of microsatellite alleles at microsatellite loci to predict whether the subject has or will develop the phenotype, wherein the phenotype comprises the disease or the condition. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is a cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolismWSGR Docket No.56034-703.601 of the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic response to the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In some embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample, and wherein the one or more microsatellite alleles at microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles per microsatellite locus; (b) a median length of all microsatellite alleles per microsatellite locus; (c) a longest length of all microsatellite alleles per microsatellite locus; (d) a shortest length of all microsatellite alleles per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering a therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy,WSGR Docket No.56034-703.601 a gene editing therapy, or a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is a cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is a cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0009] Aspects disclosed herein provide, in some embodiments, methods of monitoring a treatment for a disease or a condition in a subject, the method comprising: determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject, wherein the subject has been administered a therapy to treat the disease or the condition; and applying the genetic classifier identified by a method of the present disclosure to the lengths of microsatellite alleles at microsatellite loci to predict whether the subject will exhibit the phenotype, wherein the phenotype comprises a positive therapeutic response or the loss of therapeutic response to the therapy. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is a cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolism of the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic response to the therapy to treat theWSGR Docket No.56034-703.601 disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In some embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample, and wherein the one or more microsatellite alleles at microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles per microsatellite locus; (b) a median length of all microsatellite alleles per microsatellite locus; (c) a longest length of all microsatellite alleles per microsatellite locus; (d) a shortest length of all microsatellite alleles per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering a therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy, a gene editing therapy, or a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced ShortWSGR Docket No.56034-703.601 Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is a cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is a cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0010] Aspects disclosed herein provide, in some embodiments methods of predicting a rate of metabolism of a therapy in a subject, the method comprising: determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; and applying the genetic classifier identified by a method of the present disclosure to the lengths of microsatellite alleles at microsatellite loci to predict whether the subject has or will develop the phenotype, wherein the phenotype comprises a rate of the metabolism of the therapy in the subject that is faster than a control rate. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is a cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolism of the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic response to the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In some embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In someWSGR Docket No.56034-703.601 embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite alleles at microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample, and wherein the one or more microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles per microsatellite locus; (b) a median length of all microsatellite alleles per microsatellite locus; (c) a longest length of all microsatellite alleles per microsatellite locus; (d) a shortest length of all microsatellite alleles per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering a therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy, a gene editing therapy, or a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is a cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is aWSGR Docket No.56034-703.601 cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0011] Aspects disclosed herein provide, in some embodiments, methods of treating a disease or a condition in a subject, the method comprising: administering a therapy to the subject to treat the disease or the condition, wherein the subject is predicted to have or develop a phenotype with a test having a positive predictive value of at least about 51%, based on lengths of a plurality of microsatellite alleles at microsatellite loci detected in a biological sample obtained from the subject, and wherein the phenotype comprises a therapeutic response to the therapy. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is a cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolism of the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic response to the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In some embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with aWSGR Docket No.56034-703.601 confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite alleles at microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample, and wherein the one or more microsatellite alleles at microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles per microsatellite locus; (b) a median length of all microsatellite alleles per microsatellite alleles; (c) a longest length of all microsatellite alleles per microsatellite locus; (d) a shortest length of all microsatellite alleles per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering a therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy, a gene editing therapy, or a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is a cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is a cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channelWSGR Docket No.56034-703.601 blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0012] Aspects disclosed herein provide, in some embodiments, methods of predicting whether a subject has or will develop a disease or a condition, the method comprising: determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; comparing the lengths of microsatellite alleles at microsatellite loci to an optimal cutoff value, wherein the optimal cutoff value is determined by performing a statistical analysis on a subset of microsatellite alleles at microsatellite loci determined to be statistically significantly associated with incidence of the disease or the condition in a reference population of subjects having the disease or the condition and subjects not having the disease or the condition; and predicting whether the subject has or will develop a phenotype based, at least in part, on the comparing in (b) with a positive predictive value of at least about 51%, wherein the phenotype comprises the disease or the condition. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is a cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolism of the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic response to the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In some embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, theWSGR Docket No.56034-703.601 subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite alleles at microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample, and wherein the one or more microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles per microsatellite locus; (b) a median length of all microsatellite alleles per microsatellite locus; (c) a longest length of all microsatellite alleles per microsatellite locus; (d) a shortest length of all microsatellite alleles per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering a therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy, a gene editing therapy, or a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is a cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is a cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACEWSGR Docket No.56034-703.601 inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0013] Aspects disclosed herein provide, in some embodiments, methods of monitoring a treatment for a disease or a condition in a subject, the method comprising: determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject, wherein the subject has been administered a therapy to treat the disease or the condition; comparing the lengths of microsatellite alleles at microsatellite loci to an optimal cutoff value, wherein the optimal cutoff value is determined by performing a statistical analysis on a subset of microsatellite alleles at microsatellite loci determined to be statistically significantly associated with a positive therapeutic response or a loss of therapeutic response to the therapy to treat the disease or the condition in a reference population having the disease or the condition; and predicting that the subject will has or will develop a phenotype with a positive predictive value of at least about 51%, wherein the phenotype comprises the positive therapeutic response or the loss of therapeutic response to the therapy. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is a cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolism of the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic response to the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In someWSGR Docket No.56034-703.601 embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite alleles at microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample, and wherein the one or more microsatellite alleles at microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles per microsatellite locus; (b) a median length of all microsatellite alleles per microsatellite locus; (c) a longest length of all microsatellite alleles per microsatellite locus; (d) a shortest length of all microsatellite alleles per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering a therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy, a gene editing therapy, or a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is aWSGR Docket No.56034-703.601 cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is a cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0014] Aspects disclosed herein provide, in some embodiments methods of predicting a rate of metabolism of a therapy in a subject, the method comprising: determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; comparing the lengths of microsatellite alleles at microsatellite loci to an optimal cutoff value, wherein the optimal cutoff value is determined by performing a statistical analysis on a subset of microsatellite alleles at microsatellite loci determined to be statistically significantly associated with an aberrant rate of metabolism of the therapy to treat the disease or the condition in a reference population having the disease or the condition; and predicting whether the subject has or will develop the phenotype based, at least in part, on the comparing in (b) with a positive predictive value of at least about 51%, wherein the phenotype comprises the rate of the metabolism of the therapy in the subject that is faster than a control rate. In some embodiments, the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer metastasized into lung tissue of the subject. In some embodiments, the disease is a cardiac disease. In some embodiments, the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease. In some embodiments, the condition comprises a rate of metabolism of the therapy to treat the disease. In some embodiments, the condition comprises a therapeuticWSGR Docket No.56034-703.601 response to the therapy to treat the disease. In some embodiments, the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease. In some embodiments, the positive predictive value is at least about 60% In some embodiments, the positive predictive value is at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a specificity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a sensitivity of at least about 70%. In some embodiments, the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.90. In some embodiments, the predicting is performed with a specificity of at least about 70%. In some embodiments, the predicting is performed with a sensitivity of at least about 70%. In some embodiments, the predicting is performed with a CI of at least about 0.90. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles. In some embodiments, the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles. In some embodiments, methods further comprise predicting genomic instability at one or more microsatellite alleles at microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample. In some embodiments, the microsatellite alleles at microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs. In some embodiments, methods further comprise determining one or more genotypes for one or more microsatellite alleles at microsatellite loci of the microsatellite alleles at microsatellite loci present in the biological sample, and wherein the one or more microsatellite alleles at microsatellite loci is present in the biological sample. In some embodiments, the lengths of microsatellite alleles at microsatellite loci that are determined comprise: (a) an average length of all microsatellite alleles per microsatellite locus; (b) a median length of all microsatellite alleles per microsatellite locus; (c) a longest length of all microsatellite alleles per microsatellite locus; (d) a shortest length of all microsatellite alleles per microsatellite locus; (e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or (f) any combination of (a) to (e). In some embodiments, methods further comprise treating the disease in the subject by delivering a therapy to the subject. In some embodiments, methods further comprise treating the disease in the subject by delivering the therapy to the subject. In some embodiments, the therapy is a cancer therapy provided in Table 3. In some embodiments, the therapy is an immunotherapy, a gene editing therapy, orWSGR Docket No.56034-703.601 a surgery. In some embodiments, the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy. In some embodiments, the immunotherapy is a cancer therapy. In some embodiments, the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy. In some embodiments, the therapy is a cardiac therapy. In some embodiments, the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator. In some embodiments, the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample. In some embodiments, the whole blood sample is a peripheral blood sample. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype. In some embodiments, methods further comprise calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

[0015] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.

[0016] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein. INCORPORATION BY REFERENCE

[0017] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.WSGR Docket No.56034-703.601 BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The novel features of the inventive concepts are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present inventive concepts will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the inventive concepts are utilized, and the accompanying drawings of which:

[0019] FIG.1 shows a non-limiting example of a computing device; in this case, a device with one or more processors, memory, storage, and a network interface.

[0020] FIG.2 shows a non-limiting example of a web / mobile application provision system; in this case, a system providing browser-based and / or native mobile user interfaces.

[0021] FIG.3 shows a non-limiting example of a cloud-based web / mobile application provision system; in this case, a system comprising an elastically load balanced, auto-scaling web server and application server resources as well synchronously replicated databases.

[0022] FIG.4 shows a non-limiting example of a codebase for practicing the methods described herein.

[0023] FIG.5 depicts a non-limiting example of a flowchart for the discovery stage.

[0024] FIG.6 depicts a non-limiting example of a flowchart for the sequencing stage.

[0025] FIG.7 depicts a non-limiting example of a flowchart for the validation stage.

[0026] FIG.8 depicts non-limiting example methods for tabulating single nucleotide polymorphism (SNP) or microsatellite alleles at microsatellite loci and subjects having the disease or not.

[0027] FIG.9 depicts non-limiting exemplary methods for sequence alignments.

[0028] FIG.10A depicts a non-limiting example of a method for sequence alignments of the microsatellite alleles at microsatellite locus.

[0029] FIG.10B depicts another non-limiting example of a method for sequence alignments of the microsatellite alleles at microsatellite locus.

[0030] FIG.11 shows that pooled tallies of reads (including minor alleles) for a single microsatellite in healthy and cancer cohorts.

[0031] FIG.12A shows a non-limiting example of a receiver operating characteristic (ROC) analysis of the training set using minor alleles of the microsatellite loci as described herein.

[0032] FIG.12B shows a non-limiting example of a ROC analysis of the validation set using minor alleles of the microsatellite alleles at microsatellite loci as described herein.WSGR Docket No.56034-703.601

[0033] FIG.13 depicts a non-limiting example of disease classification criteria for distinguishing lung cancer samples from healthy controls using minor alleles.

[0034] FIG.14 depicts results for training and validation samples using a non-limiting exemplary minor allele classification method as described herein.

[0035] FIG.15 depicts non-limiting example of methods for tabulating genotypes.

[0036] FIG.16 depicts a visual illustration of a non-limiting example of metrics used to assign sample scores based on Fisher tables.

[0037] FIG.17 depicts a non-limiting example analysis of SNP allele burden.

[0038] FIG.18 depicts a non-limiting example method for disease classification combining SNPs and microsatellite.

[0039] FIG.19 depicts a non-limiting example gene enrichment analysis using the methods described herein.

[0040] FIG.20 depicts an example of a single sample’s allele count information represented as an image. Normalized counts at each length are color coded. Individual candidate microsatellite loci are depicted on the vertical, with the allele length given on the horizontal. The count of supporting sequence reads for each allele is represented as the intensity of a spot for a given allele.

[0041] FIG.21 depicts a model training example Showing Train / Test accuracy for the training set and loss evaluated with the testing set at every Epoch. The binary accuracy converges smoothly as shown in this training run.

[0042] FIG.22 depicts typical fold results of the binary benign / malignant classifier with ~200 informative loci on independent test validation samples (n=138) conducted with 10 kfold random sample sets.

[0043] FIG.23 depicts identification and classification of samples using 41 microsatellite loci with final classification done using an image analysis approach. DETAILED DESCRIPTION

[0044] Disclosed herein, in some embodiments are methods and systems for generating and using a genetic classifier capable of classifying a sample from a subject as having or not having a phenotype. The phenotype can be a disease, a trait, or a condition. Non-limiting examples of diseases include cardiac disease, neurological disease, cancer or a neoplasm. Non-limiting examples of conditions include a rate of metabolism of a therapeutic agent inWSGR Docket No.56034-703.601 the subject, a positive therapeutic response of the subject to the therapeutic agent, a therapeutic non-response or loss of response of the subject to the therapeutic agent, an allergy, an adverse effect to a therapeutic agent, a sensitivity to a chemical or environmental agent associated with cancer, obesity, or substance abuse. The systems and methods or the present disclosure leverage the discovery that microsatellite lengths relative to a distribution of microsatellite lengths representative of the general population may be used as a marker alone, or in combination with other markers (e.g., polymorphisms at the microsatellite locus) to classify a subject has having or likely to develop a phenotype. Unlike existing methods, the methods and systems of the present disclosure do not necessarily discriminate or filter out minor microsatellite alleles in the sample in recognition that the lengths of minor microsatellite alleles may provide powerful predictive capability when measured against the distribution of microsatellite alleles representative of the general population. In some embodiments, the methods comprise identifying a distribution of length of microsatellite alleles at microsatellite loci in samples from subjects that have or do not have the phenotype and performing a statistical test that tabulates the distribution of lengths of microsatellites to determine which of the microsatellite alleles at microsatellite locus lengths are predictive of the phenotype. In some embodiments, an optimal cutoff of microsatellite length is determined using a statistical analysis that can be used to generate a score. The genetic classifiers of the present disclosure can be used to predict whether a subject or a patient has or will develop the phenotype.

[0045] Also provided are methods and systems for treating, selecting for treatment, optimizing a treatment, predicting a positive therapeutic response to a treatment, predicting a loss of response to a treatment, or diagnosing of a disease to a disease or a disorder in a subject based, at least in part, on a score calculated for the subject using the genetic classifier of the present disclosure. In some embodiments, the subject is human. In some embodiments, the methods comprise classifying a patient who was previously diagnosed with a disease or a condition as having a high likelihood of developing a loss of response to a treatment. In some embodiments, the methods comprise classifying a patient who was previously diagnosed with a disease or a condition as having a high likelihood of exhibiting a positive therapeutic response to a treatment. In some embodiments, the methods comprise classifying a subject who was not yet diagnosed with a disease or a condition as having a high likelihood of developing the disease or the condition.WSGR Docket No.56034-703.601 METHODS

[0046] Disclosed herein, in some embodiments are methods and systems for generating and using a genetic classifier capable of classifying a sample obtained from a subject as having or not having a phenotype. The phenotype can be a disease or a condition. Non- limiting examples of diseases include cardiac disease, neurological disease, cancer or a neoplasm. Non-limiting examples of conditions include a rate of metabolism of a therapeutic agent in the subject, a positive therapeutic response of the subject to the therapeutic agent, a therapeutic non-response or loss of response of the subject to the therapeutic agent, an allergy, an adverse effect to a therapeutic agent, a sensitivity to a chemical or environmental agent associated with cancer, obesity, or substance abuse. The systems and methods or the present disclosure identify a distribution of length of microsatellite alleles in samples from subjects that have or do not have the phenotype and perform a statistical test that tabulates the distribution of lengths of microsatellites to determine which of the microsatellite allele lengths are predictive of the phenotype. In some embodiments, an optimal cutoff of microsatellite length is determined using a statistical analysis that can be used to generate a score. The genetic classifiers of the present disclosure can be used to predict whether a subject or a patient has or will develop the phenotype.

[0047] Also provided are methods of treating, selecting for treatment, optimizing a treatment, predicting a positive therapeutic response to a treatment, predicting a loss of response to a treatment, or diagnosing of a disease to a disease or a disorder in a subject based, at least in part, on a score calculated for the subject using the genetic classifier of the present disclosure. In some embodiments, the subject is human. In some embodiments, the methods comprise classifying a patient who was previously diagnosed with a disease or a condition as having a high likelihood of developing a loss of response to a treatment. In some embodiments, the methods comprise classifying a patient who was previously diagnosed with a disease or a condition as having a high likelihood of exhibiting a positive therapeutic response to a treatment. In some embodiments, the methods comprise classifying a subject who was not yet diagnosed with a disease or a condition as having a high likelihood of developing the disease or the condition. Methods of Discovery

[0048] Provided herein are methods of discovering novel genetic markers for determining whether a subject as a phenotype. In some embodiments, the phenotype is a disease or a condition. In some embodiments, the method is a computer-implemented method. In someWSGR Docket No.56034-703.601 embodiments, the method comprises (a) receiving sequence data from a plurality of samples obtained from a plurality of subjects comprising subjects that have the phenotype and subjects that do not have the phenotype, wherein the sequence data comprises a plurality of microsatellite alleles at microsatellite loci; (b) identifying a distribution of lengths of the plurality of microsatellite alleles at microsatellite loci; (c) performing a statistical test comprising tabulating the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci to determine which of the plurality of microsatellite alleles at microsatellite loci is predictive of the phenotype; and (d) generating a score for a subset of microsatellite loci of the plurality of microsatellite loci determined to be predictive of the phenotype based, at least in part, on the performing the statistical significance test in (c). In some embodiments, the method comprises determining an optimal cutoff value of the score for the subset of microsatellite alleles at microsatellite loci that differentiates the subjects that have the phenotype from the subjects that do not have the phenotype. In some embodiments, the methods comprise applying the optimal cutoff value to the score for the subset of microsatellite alleles at microsatellite loci, thereby producing the genetic classifier. In some embodiments, the plurality of subjects is representative of the general population. The plurality of microsatellite loci may comprise a plurality of microsatellite alleles.

[0049] In some embodiments, the distribution of lengths of the microsatellite alleles at microsatellite loci comprises minor alleles, primary alleles, or a combination thereof. In some embodiments, the distribution of lengths of the microsatellite alleles at microsatellite loci comprises only minor alleles. In some embodiments, the distribution of lengths of the microsatellite alleles at microsatellite loci comprises only primary alleles. In some embodiments, the distribution of lengths of the microsatellite alleles at microsatellite loci comprises microsatellite alleles at microsatellite loci having one or more single nucleotide polymorphism (SNP) present in a microsatellite allele at a microsatellite locus. In some embodiments, the one or more SNPs is significantly associated with the phenotype. In some embodiments, the distribution of lengths of the microsatellite loci comprises one or more microsatellite alleles at microsatellite loci that result in a frameshift mutation thereby effecting protein expression (e.g., encoding a stop codon).

[0050] In some embodiments, microsatellite alleles at a microsatellite locus may be identified from the sequencing data of a subject. In some embodiments, the method may comprise aligning the sequencing data with reference to a human genome using an alignment algorithm as described herein. In some embodiments, the method may comprise aligning the sequencing data for a plurality of microsatellite loci with reference to a reference genomeWSGR Docket No.56034-703.601 using the alignment algorithm. The reference genome may comprise a mammalian genome. The reference genome may comprise a primate genome. In some embodiments, a reference genome may comprise a human genome. Other reference genomes can comprise a feline, canine, bovine, porcine, or a rodent genome.

[0051] The alignment algorithm may comprise a scoring matrix configured to align flanking regions of the one or more microsatellite loci. In some embodiments, the flanking region may have a length of at least about 1 contiguous base pair (bp; used interchangeably with nucleotides when describing a length of a single- or double-stranded polynucleotide or nucleic acid), at least about 2 contiguous bp, at least about 3 contiguous bp, at least about 4 contiguous bp, at least about 5 contiguous bp, at least about 6 contiguous bp, at least about 7 contiguous bp, at least about 8 contiguous bp, at least about 9 contiguous bp, at least about 10 contiguous bp, at least about 11 contiguous bp, at least about 12 contiguous bp, at least about 13 contiguous bp, at least about 14 contiguous bp, at least about 15 contiguous bp, at least about 16 contiguous bp, at least about 17 contiguous bp, at least about 18 contiguous bp, at least about 19 contiguous bp, at least about 20 contiguous bp, at least about 21 contiguous bp, at least about 22 contiguous bp, at least about 23 contiguous bp, at least about 24 contiguous bp, at least about 25 contiguous bp, at least about 26 contiguous bp, at least about 27 contiguous bp, at least about 28 contiguous bp, at least about 29 contiguous bp, at least about 30 contiguous bp, at least about 31 contiguous bp, at least about 32 contiguous bp, at least about 33 contiguous bp, at least about 34 contiguous bp, at least about 35 contiguous bp, at least about 36 contiguous bp, at least about 37 contiguous bp, at least about 38 contiguous bp, at least about 39 contiguous bp, at least about 40 contiguous bp, at least about 41 contiguous bp, at least about 42 contiguous bp, at least about 43 contiguous bp, at least about 44 contiguous bp, at least about 45 contiguous bp, at least about 46 contiguous bp, at least about 47 contiguous bp, at least about 48 contiguous bp, at least about 49 contiguous bp, at least about 50 contiguous bp, at least about 51 contiguous bp, at least about 52 contiguous bp, at least about 53 contiguous bp, at least about 54 contiguous bp, at least about 55 contiguous bp, at least about 56 contiguous bp, at least about 57 contiguous bp, at least about 58 contiguous bp, at least about 59 contiguous bp, at least about 60 contiguous bp, at least about 61 contiguous bp, at least about 62 contiguous bp, at least about 63 contiguous bp, at least about 64 contiguous bp, at least about 65 contiguous bp, at least about 66 contiguous bp, at least about 67 contiguous bp, at least about 68 contiguous bp, at least about 69 contiguous bp, at least about 70 contiguous bp, at least about 71 contiguous bp, at least about 72 contiguous bp, at least about 73 contiguous bp, at least about 74 contiguous bp, at least about 75 contiguousWSGR Docket No.56034-703.601 bp, at least about 76 contiguous bp, at least about 77 contiguous bp, at least about 78 contiguous bp, at least about 79 contiguous bp, at least about 80 contiguous bp, at least about 81 contiguous bp, at least about 82 contiguous bp, at least about 83 contiguous bp, at least about 84 contiguous bp, at least about 85 contiguous bp, at least about 86 contiguous bp, at least about 87 contiguous bp, at least about 88 contiguous bp, at least about 89 contiguous bp, at least about 90 contiguous bp, at least about 91 contiguous bp, at least about 92 contiguous bp, at least about 93 contiguous bp, at least about 94 contiguous bp, at least about 95 contiguous bp, at least about 96 contiguous bp, at least about 97 contiguous bp, at least about 98 contiguous bp, at least about 99 contiguous bp, at least about 100 contiguous bp, at least about 101 contiguous bp, at least about 102 contiguous bp, at least about 103 contiguous bp, at least about 104 contiguous bp, at least about 105 contiguous bp, at least about 106 contiguous bp, at least about 107 contiguous bp, at least about 108 contiguous bp, at least about 109 contiguous bp, at least about 110 contiguous bp, at least about 111 contiguous bp, at least about 112 contiguous bp, at least about 113 contiguous bp, at least about 114 contiguous bp, at least about 115 contiguous bp, at least about 116 contiguous bp, at least about 117 contiguous bp, at least about 118 contiguous bp, at least about 119 contiguous bp, at least about 120 contiguous bp, at least about 121 contiguous bp, at least about 122 contiguous bp, at least about 123 contiguous bp, at least about 124 contiguous bp, at least about 125 contiguous bp, at least about 126 contiguous bp, at least about 127 contiguous bp, at least about 128 contiguous bp, at least about 129 contiguous bp, at least about 130 contiguous bp, at least about 131 contiguous bp, at least about 132 contiguous bp, at least about 133 contiguous bp, at least about 134 contiguous bp, at least about 135 contiguous bp, at least about 136 contiguous bp, at least about 137 contiguous bp, at least about 138 contiguous bp, at least about 139 contiguous bp, at least about 140 contiguous bp, at least about 141 contiguous bp, at least about 142 contiguous bp, at least about 143 contiguous bp, at least about 144 contiguous bp, at least about 145 contiguous bp, at least about 146 contiguous bp, at least about 147 contiguous bp, at least about 148 contiguous bp, at least about 149 contiguous bp, at least about 150 contiguous bp, at least about 200 contiguous bp, at least about 500 contiguous bp or more. In some embodiments, the flanking region may have a length of at most about 1 contiguous bp, at most about 2 contiguous bp, at most about 3 contiguous bp, at most about 4 contiguous bp, at most about 5 contiguous bp, at most about 6 contiguous bp, at most about 7 contiguous bp, at most about 8 contiguous bp, at most about 9 contiguous bp, at most about 10 contiguous bp, at most about 11 contiguous bp, at most about 12 contiguous bp, at most about 13 contiguous bp, at most about 14 contiguous bp, atWSGR Docket No.56034-703.601 most about 15 contiguous bp, at most about 16 contiguous bp, at most about 17 contiguous bp, at most about 18 contiguous bp, at most about 19 contiguous bp, at most about 20 contiguous bp, at most about 21 contiguous bp, at most about 22 contiguous bp, at most about 23 contiguous bp, at most about 24 contiguous bp, at most about 25 contiguous bp, at most about 26 contiguous bp, at most about 27 contiguous bp, at most about 28 contiguous bp, at most about 29 contiguous bp, at most about 30 contiguous bp, at most about 31 contiguous bp, at most about 32 contiguous bp, at most about 33 contiguous bp, at most about 34 contiguous bp, at most about 35 contiguous bp, at most about 36 contiguous bp, at most about 37 contiguous bp, at most about 38 contiguous bp, at most about 39 contiguous bp, at most about 40 contiguous bp, at most about 41 contiguous bp, at most about 42 contiguous bp, at most about 43 contiguous bp, at most about 44 contiguous bp, at most about 45 contiguous bp, at most about 46 contiguous bp, at most about 47 contiguous bp, at most about 48 contiguous bp, at most about 49 contiguous bp, at most about 50 contiguous bp, at most about 51 contiguous bp, at most about 52 contiguous bp, at most about 53 contiguous bp, at most about 54 contiguous bp, at most about 55 contiguous bp, at most about 56 contiguous bp, at most about 57 contiguous bp, at most about 58 contiguous bp, at most about 59 contiguous bp, at most about 60 contiguous bp, at most about 61 contiguous bp, at most about 62 contiguous bp, at most about 63 contiguous bp, at most about 64 contiguous bp, at most about 65 contiguous bp, at most about 66 contiguous bp, at most about 67 contiguous bp, at most about 68 contiguous bp, at most about 69 contiguous bp, at most about 70 contiguous bp, at most about 71 contiguous bp, at most about 72 contiguous bp, at most about 73 contiguous bp, at most about 74 contiguous bp, at most about 75 contiguous bp, at most about 76 contiguous bp, at most about 77 contiguous bp, at most about 78 contiguous bp, at most about 79 contiguous bp, at most about 80 contiguous bp, at most about 81 contiguous bp, at most about 82 contiguous bp, at most about 83 contiguous bp, at most about 84 contiguous bp, at most about 85 contiguous bp, at most about 86 contiguous bp, at most about 87 contiguous bp, at most about 88 contiguous bp, at most about 89 contiguous bp, at most about 90 contiguous bp, at most about 91 contiguous bp, at most about 92 contiguous bp, at most about 93 contiguous bp, at most about 94 contiguous bp, at most about 95 contiguous bp, at most about 96 contiguous bp, at most about 97 contiguous bp, at most about 98 contiguous bp, at most about 99 contiguous bp, at most about 100 contiguous bp, at most about 101 contiguous bp, at most about 102 contiguous bp, at most about 103 contiguous bp, at most about 104 contiguous bp, at most about 105 contiguous bp, at most about 106 contiguous bp, at most about 107 contiguous bp, at most about 108 contiguous bp, at most about 109 contiguous bp,WSGR Docket No.56034-703.601 at most about 110 contiguous bp, at most about 111 contiguous bp, at most about 112 contiguous bp, at most about 113 contiguous bp, at most about 114 contiguous bp, at most about 115 contiguous bp, at most about 116 contiguous bp, at most about 117 contiguous bp, at most about 118 contiguous bp, at most about 119 contiguous bp, at most about 120 contiguous bp, at most about 121 contiguous bp, at most about 122 contiguous bp, at most about 123 contiguous bp, at most about 124 contiguous bp, at most about 125 contiguous bp, at most about 126 contiguous bp, at most about 127 contiguous bp, at most about 128 contiguous bp, at most about 129 contiguous bp, at most about 130 contiguous bp, at most about 131 contiguous bp, at most about 132 contiguous bp, at most about 133 contiguous bp, at most about 134 contiguous bp, at most about 135 contiguous bp, at most about 136 contiguous bp, at most about 137 contiguous bp, at most about 138 contiguous bp, at most about 139 contiguous bp, at most about 140 contiguous bp, at most about 141 contiguous bp, at most about 142 contiguous bp, at most about 143 contiguous bp, at most about 144 contiguous bp, at most about 145 contiguous bp, at most about 146 contiguous bp, at most about 147 contiguous bp, at most about 148 contiguous bp, at most about 149 contiguous bp, at most about 150 contiguous bp, at most about 200 contiguous bp, or at most about 500 contiguous bp. In some embodiments, the flanking region may be upstream or downstream of the microsatellite allele at the microsatellite locus. The upstream or downstream flanking region may comprise the number of contiguous base pairs or nucleotides as described herein.

[0052] In some embodiments, the alignment algorithm described herein may improve the alignment performance of a reference alignment algorithm with a scoring matrix configured to align flanking regions having a length of fewer or equal to 24 contiguous bp, fewer or equal to 23 contiguous bp, fewer or equal to 22 contiguous bp, fewer or equal to 21 contiguous bp, fewer or equal to 20 contiguous bp, fewer or equal to 19 contiguous bp, fewer or equal to 18 contiguous bp, fewer or equal to 17 contiguous bp, fewer or equal to 16 contiguous bp, fewer or equal to 15 contiguous bp, fewer or equal to 14 contiguous bp, fewer or equal to 13 contiguous bp, fewer or equal to 12 contiguous bp, fewer or equal to 11 contiguous bp, fewer or equal to 10 contiguous bp, fewer or equal to 9 contiguous bp, fewer or equal to 8 contiguous bp, fewer or equal to 7 contiguous bp, fewer or equal to 6 contiguous bp, or fewer or equal to 5 contiguous bp. The alignment performance can be measured by gap open penalties (GOP), gap extension penalties (GEP), or any combination thereof. The alignment performance can be measured by GOP. The alignment performance can be measured by GEP. The alignment performance can be measured by GOP and GEP. In some embodiments, the alignment algorithm may comprise a Needleman–Wunsch algorithm, aWSGR Docket No.56034-703.601 Smith-Waterman algorithm, a hybrid algorithm comprising a subset of components of the Needleman–Wunsch algorithm and the Smith-Waterman algorithm, a hybrid algorithm comprising all components of the Needleman–Wunsch algorithm and the Smith-Waterman algorithm, or a combination thereof. In some embodiments, the alignment algorithm may comprise the Needleman–Wunsch algorithm. In some embodiments, the alignment algorithm may comprise the Smith-Waterman algorithm. In some embodiments, the alignment algorithm may comprise the hybrid algorithm comprising components of the Needleman– Wunsch algorithm and the Smith-Waterman algorithm. In some embodiments, the alignment algorithm may comprise the hybrid algorithm comprising all components of the Needleman– Wunsch algorithm and the Smith-Waterman algorithm. In some embodiments, the hybrid algorithm comprising a subset of components of the Needleman–Wunsch algorithm and a subset of components of the Smith-Waterman algorithm. In some embodiments, the hybrid algorithm comprising a subset of components of the Needleman–Wunsch algorithm and all components of the Smith-Waterman algorithm. In some embodiments, the hybrid algorithm comprising all components of the Needleman–Wunsch algorithm and a subset of components of the Smith-Waterman algorithm. Other alignment algorithms can comprise a Bayesian model selection guided by an empirically derived error model, or a Discretized Gaussian Mixture (e.g., GenoTan). The algorithm can be, e.g., Repeatseq. A dynamic programming based approach or heuristic method can be used to genotype microsatellites. Other tools for microsatellite genotyping include PHOBOS, MISA, Tandem Repeats Finder, FullSSR, or bMSISEA.

[0053] In some embodiments, a sequence read can be mapped to a sequence, loci, or region of the reference genome using the alignment algorithm described herein, with the score matrix configuration as described herein. When a sequence read is mapped to the sequence, loci, or region of the reference genome that represents a microsatellite allele at a microsatellite locus, the sequence read can be identified as the sequence read of the microsatellite allele at the microsatellite locus.

[0054] Microsatellite alleles at a microsatellite locus can be filtered to control for any number of factors, e.g., age, ethnicity, gender, sequencing protocol (e.g., whole genome sequencing, whole exome sequencing, or targeted sequencing), if e.g., the samples from subjects with a condition and samples from subjects without a condition are not matched for the factor. Microsatellite alleles at a microsatellite loci with potential bias can be excluded from subsequent analysis. Additional filters for filtering microsatellite alleles at microsatellite loci can include length of the microsatellite allele at the microsatellite locus repeat motif, theWSGR Docket No.56034-703.601 total length of the microsatellite allele at the microsatellite locus (e.g., number of copies of the motif), the sequence of the motif (for example, using only those with high GC content), and on the purity of the microsatellite allele at the microsatellite locus, e.g., if it has any bases that can interrupt a perfect set of copies of the motif. In some embodiments, the microsatellite alleles at microsatellite loci can be filtered by their positions in the genome, e.g., exome, intron, intergenic regions, or untranslated regions. Filtering can include filtering by genes or functional elements that are in proximity to the microsatellite loci.

[0055] The methods described herein can further comprise identifying another marker that is associated with a phenotype (e.g., a disease) or the risk of developing the phenotype. In some embodiments, the marker is or comprises a genetic marker, such as a single nucleotide polymorphism (SNP), copy number variant (CNV) or an indel. In some embodiments, the marker may be present at a microsatellite locus or one or more microsatellite loci. In some embodiments, the microsatellite locus comprises greater than or equal to 2, 3, 4, 5, 6, 7, 8, 9, or 10 markers. In some embodiments, the microsatellite locus comprises fewer than or equal to 2, 3, 4, 5, 6, 7, 8, 9, or 10 markers.

[0056] The method described herein may comprise comparing a length of a microsatellite allele at a microsatellite locus to a threshold length to determine if a subject has the phenotype or is at risk of developing the phenotype. In some embodiments, the methods may comprise identifying whether a particular microsatellite alleles at microsatellite locus(s) has a length that is above or below the threshold length that can indicate or be used to determine if a subject has or is at risk of developing the phenotype. By comparing the microsatellite alleles at microsatellite locus(s) has a length that is above or below the threshold length, the method may identify the genotype of the microsatellite alleles at microsatellite locus(s) of a subject. In some embodiments, a statistical analysis may be used to determine a threshold length of the microsatellite allele at the microsatellite locus. The statistical analysis may be used to determine the correlation between any two of the following parameters: (1) sensitivity, recall, hit rate, or true positive rate (TPR); (2) specificity, selectivity or true negative rate (TNR); (3) precision or positive predictive value (PPV); (4) negative predictive value (NPV); (5); negative predictive value (NPV); (6) miss rate or false negative rate (FNR); (7) fall-out or false positive rate (FPR); (8) false discovery rate (FDR); (9) false omission rate (FOR); (10) Positive likelihood ratio (LR+); (11) Negative likelihood ratio (LR-); (12) prevalence threshold (PT); or (13) threat score (TS) or critical success index (CSI) of various threshold lengths of a same microsatellite allele at a microsatellite locus identified from a population of reference subjects having (or is at risk of) a disease and a population ofWSGR Docket No.56034-703.601 reference subjects not having (or is not at risk of) the phenotype. In some embodiments, the statistical analysis may comprise analyzing the correlation between TPR and FPR using various lengths of the microsatellite alleles at the microsatellite locus as the threshold length. In some embodiments, the statistical analysis may comprise a receiver operating characteristic (ROC) analysis. In some embodiments, when determining a threshold length of a microsatellite allele at a microsatellite locus, the method may comprise determining the TPR and FPR for each of the various threshold length in identifying each microsatellite allele at a microsatellite locus as having (or is at risk of) or not having (or is not at risk of) the disease. In some embodiments, the TPR of the threshold length of a microsatellite allele at a microsatellite locus is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more. In some embodiments, the TPR of the threshold length is at most about at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99% or more. In some embodiments, the TPR of the threshold length is at most about 50%, at most about 55%, at most about 60%, at most about 65%, at most about 70%, at most about 75%, at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99% or more. In some embodiments, the TPR of the threshold length is at most about at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, or at most about 99%. In some embodiments, the FPR of the threshold length of a microsatellite allele at a microsatellite locus is at least about 1 %, at least about 2 %, at least about 3 %, at least about 5 %, at least about 6 %, at least about 7 %, at least about 8 %, at least about 9 %, at least about 10 %, at least about 15 %, at least about 20 %, at least about 25 %, at least about 30 %, at least about 35 %, at least about 40 %, at least about 45 %, or at least about 50 %. In some embodiments, the FPR of the threshold length of a microsatellite allele at a microsatellite locus is at most about 1 %, at most about 2 %, at most about 3 %, at most about 5 %, at most about 6 %, at most about 7 %, at most about 8 %, at most about 9 %, at most about 10 %, at most about 15 %, at most about 20 %, at most about 25 %, at most about 30 %, at most about 35 %, at most about 40 %, at most about 45WSGR Docket No.56034-703.601 %, or at most about 50 %. The percentages of TPR or FPR described herein can be applied to other parameters for determining a threshold length of a microsatellite allele at a microsatellite locus.

[0057] In some embodiments, the threshold length of a microsatellite allele at a microsatellite locus may be at least about 1 bp, at least about 2 bp, at least about 3 bp, at least about 5 bp, at least about 6 bp, at least about 7 bp, at least about 8 bp, at least about 9 bp, at least about 10 bp, at least about 20 bp, at least about 50 bp, at least about 100 bp, at least about 500 bp, at least about 1000 bp, at least about 5000 bp, at least about 10000 bp or more. In some embodiments, the threshold length of a microsatellite allele at a microsatellite locus may be at most about 1 bp, at most about 2 bp, at most about 3 bp, at most about 5 bp, at most about 6 bp, at most about 7 bp, at most about 8 bp, at most about 9 bp, at most about 10 bp, at most about 20 bp, at most about 50 bp, at most about 100 bp, at most about 500 bp, at most about 1000 bp, at most about 5000 bp, or at most about 10000 bp.

[0058] In some embodiments, the threshold length can be determined by determining the most common length of microsatellite alleles at a microsatellite locus present within a population of subjects having or is at risk of the disease. For example, the most common length of microsatellite alleles at a microsatellite locus can be a modal length of the microsatellite alleles at the microsatellite locus present within a population of subjects having or is at risk of the disease. In other cases, the threshold length of microsatellite alleles at a microsatellite locus can be a maximal length; mean length; median length; minimal length; a length that is at 25 % or 75 % percentile; or a combination thereof of microsatellite alleles at the microsatellite locus present within a population of subjects having or is at risk of the phenotype (e.g., disease). For example, the maximal length; mean length; median length; minimal length; a length that is at 25 % or 75 % percentile; an aberrant length (as described herein); or a combination thereof of microsatellite alleles at a microsatellite locus of the subjects having the disease or the risk thereof may be used to calculate the threshold length. In some embodiments, the maximal length; mean length; median length; minimal length; a length that is at 25 % or 75 % percentile; an aberrant length; or a combination thereof of a microsatellite locus of the subjects not having the disease or the risk thereof may be used to calculate the threshold length.

[0059] In some embodiments, an aberrant length of microsatellite alleles at a microsatellite locus may be a length that exceeds a specific length. The specific length of microsatellite alleles at the microsatellite locus may be one observed among subjects that have the disease or the risk thereof. The specific length may be at least about 1 bp, at leastWSGR Docket No.56034-703.601 about 2 bp, at least about 3 bp, at least about 4 bp, at least about 5 bp, at least about 6 bp, at least about 7 bp, at least about 8 bp, at least about 9 bp, at least about 10 bp, at least about 11 bp, at least about 12 bp, at least about 13 bp, at least about 14 bp, at least about 15 bp, at least about 16 bp, at least about 17 bp, at least about 18 bp, at least about 19 bp, at least about 20 bp, at least about 21 bp, at least about 22 bp, at least about 23 bp, at least about 24 bp, at least about 25 bp, at least about 50 bp, at least about 100 bp, at least about 200 bp, at least about 300 bp, at least about 400 bp, at least about 500 bp, at least about 600 bp, at least about 700 bp, at least about 800 bp, at least about 900 bp, at least about 1000 bp or more. The specific length may be at most about 1 bp, at most about 2 bp, at most about 3 bp, at most about 4 bp, at most about 5 bp, at most about 6 bp, at most about 7 bp, at most about 8 bp, at most about 9 bp, at most about 10 bp, at most about 11 bp, at most about 12 bp, at most about 13 bp, at most about 14 bp, at most about 15 bp, at most about 16 bp, at most about 17 bp, at most about 18 bp, at most about 19 bp, at most about 20 bp, at most about 21 bp, at most about 22 bp, at most about 23 bp, at most about 24 bp, at most about 25 bp, at most about 50 bp, at most about 100 bp, at most about 200 bp, at most about 300 bp, at most about 400 bp, at most about 500 bp, at most about 600 bp, at most about 700 bp, at most about 800 bp, at most about 900 bp, or at most about 1000 bp. The specific length of microsatellite alleles at the microsatellite locus may be a length that is at least about 5 %, at least about 6 %, at least about 7 %, at least about 8 %, at least about 9 %, at least about 10 %, at least about 20 %, at least about 30 %, at least about 40 %, at least about 50 %, at least about 60 %, at least about 70 %, at least about 80 %, at least about 90 %, at least about 100 %, at least about 150 %, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least about 100-fold, at least about 1000-fold longer than a mean, model, median, maximal, minimal, 25 % percentile, or 75 % percentile length of microsatellite alleles at the microsatellite locus of the subjects not having the disease or the risk thereof. The specific length of microsatellite alleles at the microsatellite locus may be a length that is at most about 5 %, at most about 6 %, at most about 7 %, at most about 8 %, at most about 9 %, at most about 10 %, at most about 20 %, at most about 30 %, at most about 40 %, at most about 50 %, at most about 60 %, at most about 70 %, at most about 80 %, at most about 90 %, at most about 100 %, at most about 150 %, at most about 2-fold, at most about 3-fold, at most about 4-fold, at most about 5-fold, at most about 6-fold, at most about 7-fold, at most about 8-fold, at most about 9-fold, at most about 10-fold, at most about 100-fold, at most about 1000-fold longer than a mean, model, median, maximal, minimal, 25 % percentile, or 75 % percentileWSGR Docket No.56034-703.601 length of microsatellite alleles at the microsatellite locus of the subjects not having the disease or the risk thereof. The specific length of microsatellite alleles at the microsatellite locus may be a length that is at least about 0.1 %, 0.5 %, 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, or 50 % of a mean, model, median, maximal, minimal, 25 % percentile, or 75 % percentile length of microsatellite alleles at the microsatellite locus of the subjects having the disease or the risk thereof. The specific length of microsatellite alleles at the microsatellite locus may be a length that is at most about 0.1 %, 0.5 %, 1 %, 2 %, 3 %, 4 %, 5 %, 6 %, 7 %, 8 %, 9 %, 10 %, 20 %, 30 %, 40 %, or 50 % of a mean, model, median, maximal, minimal, 25 % percentile, or 75 % percentile length of microsatellite alleles at the microsatellite locus of the subjects having the disease or the risk thereof.

[0060] In some embodiments, the threshold length can be determined by a classification algorithm. The classifier may comprise decision trees, random forests, Bayesian networks, support vector machines, neural networks, logistic regression, probit model, genetic programming, multi expression programming, linear genetic programming, or a combination thereof. Subsequent to inputting the lengths of a microsatellite locus present within a population of subjects having or is at risk of the phenotype, the classifier can determine the threshold length that can identify a subject of having or not having the phenotype (or is at risk of the phenotype) with any of the following parameters as described herein, such as (1) sensitivity, recall, hit rate, or true positive rate (TPR); (2) specificity, selectivity or true negative rate (TNR); (3) precision or positive predictive value (PPV); (4) negative predictive value (NPV); (5); negative predictive value (NPV); (6) miss rate or false negative rate (FNR); (7) fall-out or false positive rate (FPR); (8) false discovery rate (FDR); (9) false omission rate (FOR); (10) Positive likelihood ratio (LR+); (11) Negative likelihood ratio (LR-); (12) prevalence threshold (PT); or (13) threat score (TS) or critical success index (CSI) of various threshold lengths of a same microsatellite allele at a microsatellite locus identified from a population of reference subjects having (or is at risk of) a disease and a population of reference subjects not having (or is not at risk of) the disease.

[0061] In some embodiments, the method described herein may comprise determining the statistical significance of the association a genotype of a marker (such as a microsatellite allele at a microsatellite locus or others as described herein) to a phenotype or a risk of developing the phenotype. The statistical significance can be calculated by tests such as t-test, Z-test, ANOVA, regression analysis, Mann-Whitney-Wilcoxon, Chi-squared test, correlation, Fisher’s exact test, Bonferroni correction, and Benjamini-Hochberg test. In someWSGR Docket No.56034-703.601 embodiments, statistical differences are quantified using a generalized Fisher’s exact test. In some embodiments, a Benjamini-Hochberg multiple testing correction is applied to control false discovery rate. In some embodiments, the threshold length can be determined by a classifier analysis. In some embodiments, determining the statistical significance of the association of a genotype of a marker to a phenotype or a risk of developing the phenotype is performed by a decision tree algorithm, a random forest algorithm, a Bayesian network, a support vector machine, an artificial neural network, a logistic regression algorithm, a probit model, genetic programming, multi expression programming, linear genetic programming, or a combination thereof.

[0062] In some embodiments, the methods can comprise genotyping microsatellite alleles at a microsatellite locus of a subject. The genotyping can comprise determining a length of microsatellite alleles at a microsatellite locus of the subject. In some embodiments, a flanking sequence upstream a downstream of the microsatellites are sequenced and the distance between the flanking sequences corresponds to the microsatellite alleles at the microsatellite locus length. The lengths of the microsatellites can be tabulated, and genotypes inferred in accordance with the methods of the present disclosure. In some embodiments, the flanking sequences are predetermined. In some embodiments, the length of microsatellite alleles at the microsatellite locus is determined using the alignment algorithm and method disclosed elsewhere herein.

[0063] The method may comprise tabulating the genotype of microsatellite alleles at a microsatellite locus with respective to its association with the phenotype or a risk of developing the phenotype and carrying out a statistical analysis to determine the statistical significance of the association. In some embodiments, the tabulation may be a 2x2 tabulation, wherein one variable of the table comprises the genotypes (2 different genotypes of microsatellite alleles at the microsatellite locus), and another variable of the table comprises whether one of the particular genotypes is presented within a subject having a disease / a risk thereof or not. In some embodiments, the statistical test may comprise a chi-squared test. In some embodiments, the tabulation may comprise Nx2 tabulation; wherein N may indicate the number of genotypes of microsatellite alleles at the microsatellite locus (or other marker as described herein) and can be 1, 2, 3, 4, 5 or more. For example, in some embodiments, the genotype of microsatellite alleles at the microsatellite locus may not be identified as binary (such as longer or shorter than a particular threshold). In such case, other statistical analysis can be used to calculate the statistical significance of the association of the genotypes and the disease or the risk thereof. For example, a Fisher’s exact test can be used to calculate theWSGR Docket No.56034-703.601 statistical significance. In some embodiments, the genotype may be heterozygous or homozygous for microsatellite alleles at a microsatellite locus of the one or more microsatellite alleles at microsatellite loci.

[0064] In some embodiments, the microsatellite locus (or other markers as described herein) of a plurality of reference subjects may be used to calculate the statistical significance of the association. The reference subjects may comprise subjects that have the phenotype and subjects that do not have the phenotype. The plurality of reference subjects used to calculate the statistical significance of the association can be referred to as a training set in this disclosure. For example, at least about 10 reference subjects, at least about 50 reference subjects, at least about 100 reference subjects, at least about 200 reference subjects, at least about 300 reference subjects, at least about 400 reference subjects, at least about 500 reference subjects, at least about 600 reference subjects, at least about 700 reference subjects, at least about 800 reference subjects, at least about 900 reference subjects, at least about 1000 reference subjects, at least about 5000 reference subjects, at least about 10000 reference subjects, at least about 50000 reference subjects, at least about 100000 reference subjects, at least about 500000 reference subjects or more can be used in a training set. In some embodiments, at most about 10 reference subjects, at most about 50 reference subjects, at most about 100 reference subjects, at most about 200 reference subjects, at most about 300 reference subjects, at most about 400 reference subjects, at most about 500 reference subjects, at most about 600 reference subjects, at most about 700 reference subjects, at most about 800 reference subjects, at most about 900 reference subjects, at most about 1000 reference subjects, at most about 5000 reference subjects, at most about 10000 reference subjects, at most about 50000 reference subjects, at most about 100000 reference subjects, or at most about 500000 reference subjects can be used in a training set. In some embodiments, a ratio of the reference subjects having the disease or the risks thereof and those that not having the disease or the risks thereof may be about 1:1, 1:2, 1:3.1:5, 1:10, 1:00, 1:1000, 2:1, 3:1, 4:1, 5:1, 10:1, 100:1, or 1000:1.

[0065] In some embodiments, the training set may comprise publicly available database. For example, the training set may comprise 1000 genomes project (KJGP) or The Cancer Genome Atlas (TCGA).

[0066] In some embodiments, a plurality (or panel) of microsatellite alleles at microsatellite loci (or other markers as described herein) may be selected for testing or determining if a subject having the disease or the risk thereof. The plurality of microsatellite alleles at microsatellite loci may be tested for the methods described herein, wherein at leastWSGR Docket No.56034-703.601 one of a genotype of a microsatellite locus of the panel may be statistically significantly associated with the disease or the risk thereof, using the method described herein. In some embodiments, the panel may comprise at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 5000 or more markers as described herein. In some embodiments, the panel may comprise at most about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 5000 markers as described herein. In some embodiments, the panel may comprise at least about 10 , 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 5000 or more microsatellite alleles at microsatellite loci. In some embodiments, the panel may comprise at most about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 5000 microsatellite alleles at microsatellite loci. In some embodiments, the plurality (or panel) of microsatellite alleles at microsatellite loci for uses in the methods described herein may be selected based on a call rate. A call rate may comprise the proportion or percentage of sequence reads or a sample in which its genotype can be determined or identified. For a marker (such as a microsatellite allele at a microsatellite locus or others as described herein to be selected, its call rate may be at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.9% or more. For a marker (such as a microsatellite allele at a microsatellite locus or others as described herein to be selected, its call rate may be at most about 50%, at most about 55%, at most about 60%, at most about 65%, at most about 70%, at most about 75%, at most about 80%, at most about 85%, at most about 90%, at most about 91%, at most about 92%, at most about 93%, at most about 94%, at most about 95%, at most about 96%, at most about 97%, at most about 98%, at most about 99%, at most about 99.9% or about 100 %.

[0067] The methods described herein may generate a score that indicates the probability of the subject having or developing the phenotype. In some embodiments, the threshold length(s) can be used to determine the number or quantity of the plurality of microsatellite alleles at microsatellite locus(s) present within a subject that are above (or higher) or below (or lower) the respective threshold lengths of the microsatellite alleles at microsatellite locus(s). The number or quantity of the microsatellite alleles at microsatellite locus(s) within the subject that are above (or higher) or below (or lower) the respective threshold lengths of the microsatellite alleles at microsatellite locus(s) may be used to generate a score. In some embodiments, the score may indicate a probability that the subject has or is at risk of developing a phenotype.WSGR Docket No.56034-703.601

[0068] In some embodiments, the method may comprise applying a weight to numerical value assigned to microsatellite alleles at a microsatellite locus of the plurality (or panel) of microsatellite alleles at microsatellite loci for uses in the methods described herein. The weight may be the relative association of the one or more microsatellite alleles at microsatellite loci and the disease, the odd ratio of microsatellite alleles at the microsatellite locus, the difference of an allele in frequency determined for microsatellite alleles at a microsatellite locus, a derivative thereof (summation, subtraction, division, multiplication, logarithm transformation) or a combination thereof. The weight may be the relative association of the one or more microsatellite alleles at microsatellite loci and the disease or the derivative of relative association (e.g., summation, subtraction, division, multiplication, logarithm transformation). The weight may be the odd ratio of the microsatellite alleles at microsatellite locus or the derivative of the odd ratio (e.g., summation, subtraction, division, multiplication, logarithm transformation). The weight may be the difference of an allele in frequency determined for microsatellite alleles at a microsatellite locus or the derivative of the difference of an allele in frequency. In some embodiments, the method may further comprise calculating a weight for each genotype for microsatellite alleles at a microsatellite locus(s). In some embodiments, the weight can comprise or be derived from a difference between a normalized frequency of each genotype for the microsatellite alleles at the microsatellite locus in reference subjects with the disease and reference subjects without the disease. In some embodiments, microsatellite alleles at microsatellite loci or genotypes are weighted more heavily if a SNP or other marker is detected at the microsatellite locus. In some embodiments, microsatellite alleles at microsatellite loci or genotypes are weighted more heavily if the microsatellite results in a frameshift mutation resulting in aberrant protein expression or lack thereof. In some embodiments, the weight for each genotype together of microsatellite alleles at a microsatellite locus can be added together to generate the score.

[0069] The score described herein may be compared to a threshold score for determining if a subject has or is at risk of a disease as described herein. In some embodiments, if the score is above (exceeds or is higher) the threshold score, the subject has the disease. In some embodiments, if the score is above the threshold score, the subject will develop the disease. In some embodiments, if the score is below (lower or under) the threshold score, the subject does not have the disease. In some embodiments, if the score is below the threshold score, the subject will not develop the disease. In some embodiments, if the score is below the threshold score, the subject has the disease. In some embodiments, if the score is below the threshold score, the subject will develop the disease. In some embodiments, if the score is above theWSGR Docket No.56034-703.601 threshold score, the subject does not have the disease. In some embodiments, if the score is above the threshold score, the subject will not develop the disease. In some embodiments, if the score is equal to the threshold score, the probability that the subject has or will develop the disease is indeterminate. In some embodiments, if the score is equal to the threshold score, the subject has the disease. In some embodiments, if the score is equal to the threshold score, the subject will develop the disease. In some embodiments, if the score is equal the threshold score, the subject does not have the disease. In some embodiments, if the score is equal the threshold score, the subject will not develop the disease.

[0070] In some embodiments, the threshold score described herein can be determined using a statistical analysis or a classifier analysis as described herein of a plurality of scores from a population of subjects or reference subjects, wherein the subjects or reference subjects comprise subjects that have or are at risk of the phenotype (e.g., disease) and subjects that do not have or are not at risk of the phenotype (e.g., disease). For example, the threshold score can be determined using ROC analysis, decision trees, random forests, Bayesian networks, support vector machines, neural networks, logistic regression, probit model, genetic programming, multi expression programming, linear genetic programming, or a combination thereof. In some embodiments, the method may comprise analyzing various scores using the statistical (or classifier) and determine which score can provide the: (1) sensitivity, recall, hit rate, or true positive rate (TPR); (2) specificity, selectivity or true negative rate (TNR); (3) precision or positive predictive value (PPV); (4) negative predictive value (NPV); (5); negative predictive value (NPV); (6) miss rate or false negative rate (FNR); (7) fall-out or false positive rate (FPR); (8) false discovery rate (FDR); (9) false omission rate (FOR); (10) Positive likelihood ratio (LR+); (11) Negative likelihood ratio (LR-); (12) prevalence threshold (PT); or (13) threat score (TS) or critical success index (CSI) as described herein, for determining if a subject has or will develop the disease, wherein the score determined may become the threshold score. For example, the method may comprise analyzing various scores using the ROC analysis. The score determined to be the threshold score may comprise a score that has or be associated with the maximal value of the Youden index. The score determined to be the threshold score may comprise a score that has or be associated with at least about 75 %, at least about 80 %, at least about 85 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, at least about 99.9 %, at least about 99.99 %, at least about 99.999 % or more of the Youden index.WSGR Docket No.56034-703.601

[0071] In some embodiments, the sensitivity, specificity, accuracy, positive predictive value, or negative predictive value of the methods of the present disclosure are improved relative to existing microsatellite or SNP-based tests. Other microsatellite or SNP-based tests may include MayoComplete Lung Cancer-Targeted Gene Panel with Rearrangement, MayoComplete Lung Cancer Mutations Next Generation Sequencing test, the MayoComplete Lung Rearrangements, Rapid Test, or the tests described in US-20220189583-A1. The MayoComplete Lung Cancer-Targeted Gene Panel with Rearrangement uses uses targeted next-generation sequencing to determine microsatellite instability status and to evaluate for somatic mutations within the ALK, BRAF, EGFR, ERBB2, HRAS, KRAS, MDM2, MET, NRAS, RET, ROS1, and STK11 genes, and activating exon 14 skipping mutations in MET, as well as reverse transcription polymerase chain reaction (rtPCR) to detect gene fusions by identifying specific rearrangements within the ALK, ROS1 and RET genes and expression imbalance for ALK, ROS1, RET, NTRK1, NTRK2, and NTRK3 genes. The MayoComplete Lung Cancer Mutations Next Generation Sequencing test, uses sequencing to identify microsatellites and somatic mutations in the ALK, BRAF, EGFR, ERBB2, HRAS, KRAS, MDM2, MET, NRAS, RET, ROS1, and STK11 genes. The MayoComplete Lung Rearrangements, Rapid Test uses polymerase chain reactions (PCR) to identify specific gene fusions (rearrangements) involving the ALK, ROS1, and RET genes, MET exon 14 skipping, and expression imbalance for ALK, ROS1, RET, NTRK1, NTRK2, and NTRK3 genes. Other tests may use multiplex rtPCR or PCR to detect (i) rearrangements within no more than 1, 2, 3, 4, or 5 genes, (ii) expression imbalance for no more than 6, 7, 8, 9, or 10 genes, or (iii) a combination thereof. In some embodiments, the sensitivity, accuracy, positive predictive value, or negative predictive value of the methods of the present disclosure are improved relative to methods that do not incorporate minor alleles of microsatellite loci.

[0072] In some embodiments, the method can have the sensitivity as described herein. In some embodiments, the sensitivity of the method may be the true positive rate. In some embodiments, the sensitivity of the method may be calculated by: Sensitivity = Number of subjects correctly identified as having the disease or the risk of the disease (true positive / TP) / (TP + Number of subjects incorrectly identified as not having the disease or the risk of the disease (false negative / FN). In some embodiments, the method is identified with a sensitivity of at least about 50 %, at least about 55 %, at least about 60 %, at least about 65 %, at least about 70 %, at least about 75 %, at least about 80 %, at least about 85 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %,WSGR Docket No.56034-703.601 at least about 99.9 %, at least about 99.99 %, at least about 99.999 % or more. In some embodiments, the method is identified with a sensitivity of at most about 50 %, at most about 55 %, at most about 60 %, at most about 65 %, at most about 70 %, at most about 75 %, at most about 80 %, at most about 85 %, at most about 90 %, at most about 91 %, at most about 92 %, at most about 93 %, at most about 94 %, at most about 95 %, at most about 96 %, at most about 97 %, at most about 98 %, at most about 99 %, at most about 99.9 %, at most about 99.99 %, at most about 99.999 %, or about 100 %.

[0073] In some embodiments, the method can have the specificity as described herein. In some embodiments, the specificity of the method may be calculated by: Specificity = Number of subjects correctly identified as not having the disease or the risk of the disease (true negative / TN) / (TN + Number of subjects incorrectly identified as having the disease or the risk of the disease (false positive / FP). In some embodiments, the method is identified with a specificity of at least about 50 %, at least about 55 %, at least about 60 %, at least about 65 %, at least about 70 %, at least about 75 %, at least about 80 %, at least about 85 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, at least about 99.9 %, at least about 99.99 %, at least about 99.999 % or more. In some embodiments, the method is identified with a specificity of at most about 50 %, at most about 55 %, at most about 60 %, at most about 65 %, at most about 70 %, at most about 75 %, at most about 80 %, at most about 85 %, at most about 90 %, at most about 91 %, at most about 92 %, at most about 93 %, at most about 94 %, at most about 95 %, at most about 96 %, at most about 97 %, at most about 98 %, at most about 99 %, at most about 99.9 %, at most about 99.99 %, at most about 99.999 %, or about 100 %.

[0074] In some embodiments, the method has the accuracy as described herein. In some embodiments, the accuracy of the method may be calculated by: Accuracy = (TP + TN) / (TP + TN + FP +FN). In some embodiments, the method is identified with a accuracy of at least about 50 %, at least about 55 %, at least about 60 %, at least about 65 %, at least about 70 %, at least about 75 %, at least about 80 %, at least about 85 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, at least about 99.9 %, at least about 99.99 %, at least about 99.999 % or more. In some embodiments, the method is identified with a accuracy of at most about 50 %, at most about 55 %, at most about 60 %, at most about 65 %, at most about 70 %, at most about 75 %, at most about 80 %, at most about 85 %, at most about 90 %, at most about 91 %, at most about 92 %, at mostWSGR Docket No.56034-703.601 about 93 %, at most about 94 %, at most about 95 %, at most about 96 %, at most about 97 %, at most about 98 %, at most about 99 %, at most about 99.9 %, at most about 99.99 %, at most about 99.999 %, or about 100 %.

[0075] In some embodiments, the method can have the Positive predictive value (PPV) as described herein. In some embodiments, the PPV of the method may be calculated by: PPV = TP / (TP+FP). In some embodiments, the method is identified with a PPV of at least about 50 %, at least about 55 %, at least about 60 %, at least about 65 %, at least about 70 %, at least about 75 %, at least about 80 %, at least about 85 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, at least about 99.9 %, at least about 99.99 %, at least about 99.999 % or more. In some embodiments, the method is identified with a PPV of at most about 50 %, at most about 55 %, at most about 60 %, at most about 65 %, at most about 70 %, at most about 75 %, at most about 80 %, at most about 85 %, at most about 90 %, at most about 91 %, at most about 92 %, at most about 93 %, at most about 94 %, at most about 95 %, at most about 96 %, at most about 97 %, at most about 98 %, at most about 99 %, at most about 99.9 %, at most about 99.99 %, at most about 99.999 %, or about 100 %.

[0076] In some embodiments, the method can have the Negative predictive value (NPV) as described herein. In some embodiments, the NPV of the method may be calculated by: NPV = TN / (TN+FN). In some embodiments, the method is identified with a NPV of at least about 50 %, at least about 55 %, at least about 60 %, at least about 65 %, at least about 70 %, at least about 75 %, at least about 80 %, at least about 85 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, at least about 99.9 %, at least about 99.99 %, at least about 99.999 % or more. In some embodiments, the method is identified with a NPV of at most about 50 %, at most about 55 %, at most about 60 %, at most about 65 %, at most about 70 %, at most about 75 %, at most about 80 %, at most about 85 %, at most about 90 %, at most about 91 %, at most about 92 %, at most about 93 %, at most about 94 %, at most about 95 %, at most about 96 %, at most about 97 %, at most about 98 %, at most about 99 %, at most about 99.9 %, at most about 99.99 %, at most about 99.999 %, or about 100 %.

[0077] In some embodiments, the method can be performed with a confidence level (CI). In some embodiments, the method is performed with a CI of at least about 50 %, at least about 55 %, at least about 60 %, at least about 65 %, at least about 70 %, at least about 75 %,WSGR Docket No.56034-703.601 at least about 80 %, at least about 85 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, at least about 99.9 %, at least about 99.99 %, at least about 99.999 % or more. In some embodiments, the method is performed with a CI of at most about 50 %, at most about 55 %, at most about 60 %, at most about 65 %, at most about 70 %, at most about 75 %, at most about 80 %, at most about 85 %, at most about 90 %, at most about 91 %, at most about 92 %, at most about 93 %, at most about 94 %, at most about 95 %, at most about 96 %, at most about 97 %, at most about 98 %, at most about 99 %, at most about 99.9 %, at most about 99.99 %, at most about 99.999 %, or about 100 %.

[0078] In some embodiments, the method may further comprise (a) identifying a genotype of the sample for the one or more microsatellite loci; (b) inputting the genotype into a trained neural network; and (c) receiving the probability of the subject having or developing the phenotype (e.g., disease). In some embodiments, the phenotype is a disease. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer in lung tissue. In some embodiments, the methods further comprises identifying a plurality of genotypes and minor alleles at a plurality of microsatellite loci associated with a plurality of cancer types (e.g., ow-grade glioblastoma, high-grade glioblastoma, melanoma, ovarian, prostate and pancreatic cancers).

[0079] In some embodiments, the method may further comprise classifying the sample, wherein the sample is further classified by a method comprising: (a) identifying a genotype of the sample for the one or more microsatellite loci; (b) inputting the genotype into a random forest algorithm; and (c) receiving the probability of the subject having or developing the phenotype (e.g., disease). In some embodiments, the phenotype is a disease. In some embodiments, the disease is cancer. In some embodiments, the cancer is lung cancer or cancer in lung tissue. In some embodiments, the methods further comprises identifying a plurality of genotypes and minor alleles at a plurality of microsatellite loci associated with a plurality of cancer types (e.g., ow-grade glioblastoma, high-grade glioblastoma, melanoma, ovarian, prostate and pancreatic cancers).

[0080] In some embodiments, the method may further comprise determining a presence of one or more single nucleotide polymorphisms (SNPs) in the sequencing data of the subject, wherein the one or more SNPs comprises a risk allele associated with the high probability that the subject has or will develop the disease relative to another reference subject that is not a carrier of the risk allele.WSGR Docket No.56034-703.601

[0081] In some embodiments, the method may further comprise (a) determining an allele burden for a plurality of the one or more SNPs in subjects with the disease and subjects without the disease, wherein the allele burden is on a range of 0 to twice a number of the plurality of the one or more SNPs; (b) performing a third ROC analysis to identify a cut off number of risk alleles that differentiates the subjects with the disease and the subjects without the disease.

[0082] In some embodiments, determining the score described herein may further comprises: (i) determining a total number of risk alleles for the one or more SNPs in the sequencing data of the subject; (ii) comparing the total number of risk alleles for the one or more SNPs to the cut off number of risk alleles to produce a second score; and (iii) altering the score based on the second score. In some embodiments, the method may further comprise any one or all the steps of: identifying additional microsatellite regions within genes having one or more single nucleotide polymorphisms associated with the cancer with a p-value of greater than or equal to about 0.05; removing duplicates between the additional microsatellite regions and the one or more microsatellite regions identified in (b) to identify unique additional microsatellite regions; and providing (i) the unique additional microsatellite regions, (ii) the cut off for each of the unique additional microsatellite regions, and (iii) an odds ratio for each of the unique additional microsatellite regions

[0083] In some embodiments, the marker or microsatellite region may comprise or is a minor allele, as described herein. Minor alleles can be distinct from the primary alleles of the genotype; they can be somatically acquired in normal tissues as one ages. Minor alleles can be used as an indication of microsatellite mutability. For example, the sequence reads of a subject of a microsatellite region may comprise at least three distinct sequences or lengths. In some embodiments, at least one sequence read with one distinct sequence or length may be present in the minority of all the sequence reads obtained from the microsatellite region from the subject. In such a case, the at least one sequence read with one distinct sequence or length present in the minority of all the sequence reads may be a minor allele(s)

[0084] In some embodiments, a minor allele may comprise a subset of the one or more microsatellite regions is present in the plurality of samples with a minor allele frequency of less than or equal to about 5%. In some embodiments, a minor allele may comprise a subset of the one or more microsatellite regions is present in the plurality of samples with a minor allele frequency of at most about 1%, at most about 2%, at most about 3%, at most about 4%, at most about 5%, at most about 10%, at most about 15%, at most about 20%, at most aboutWSGR Docket No.56034-703.601 25%, at most about 30%, at most about 35%, at most about 40%, at most about 45%, or at most about 50%. Samples

[0085] Disclosed herein, in some embodiments, are methods and systems that analyze one or more samples. A sample can comprise a body fluid. In some embodiments, a sample may comprise a blood sample. In some embodiments, a sample may comprise a serum sample. In some embodiments, a sample may comprise a plasma sample. In some embodiments, a body fluid may comprise an intracellular body fluid or an extracellular body fluid. Non-limiting examples of extracellular body fluid can include intravascular fluid, interstitial fluid, lymphatic fluid, transcellular fluid, or a combination thereof. In some embodiments, a body fluid can also comprise ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, cerumen, breast milk, broncheoalveolar lavage fluid, semen (including prostatic fluid), Cowper's fluid or pre- ejaculatory fluid, female ejaculate, sweat, fecal matter, hair, tears, cyst fluid, pleural and peritoneal fluid, pericardial fluid, lymph, chyme, chyle, bile, interstitial fluid, menses, pus, sebum, vomit, vaginal secretions, mucosal secretion, stool water, pancreatic juice, lavage fluids from sinus cavities, bronchopulmonary aspirates or other lavage fluids. Body fluid may also comprise the blastocyl cavity, umbilical cord blood, or maternal circulation which may be of fetal or maternal origin. In some embodiments, the fluid sample is derived from a body fluid selected from among whole blood, sputum, serum, plasma, urine, cerebrospinal fluid, nipple aspirate, saliva, fine needle aspirate. In some embodiments, a body fluid may be used as a sample for determining if a subject (that the body fluid is derived from) has cancer or the risk thereof. In some embodiments, blood or serum may be used as the sample for determining if the subject (that the blood is derived from) has lung cancer or the risk thereof.

[0086] A sample for uses in the methods described herein can comprise a biopsy sample. Non-limiting examples of biopsy samples can include bone biopsy, a bone marrow biopsy, a breast biopsy, a gastrointestinal biopsy, a lung biopsy, a liver biopsy, a prostate biopsy, a nervous system biopsy, a urogenital biopsy, a lymph node biopsy, a muscle biopsy, a skin biopsy, a blood biopsy, a bodily fluid biopsy, a cardiac biopsy, an endometrial biopsy, an open biopsy, a sentinel lymph node biopsy, or any combinations thereof. In some embodiments, a biopsy can comprise a fine needle aspiration biopsy, a core needle biopsy, a vacuum-assisted biopsy, an excisional biopsy, a shave biopsy, a punch biopsy, an endoscopic biopsy, a laparoscopic biopsy, a bone marrow aspiration biopsy, a liquid biopsy, anyWSGR Docket No.56034-703.601 derivatives herein and thereof, or any combinations herein and thereof. In some embodiments, a biopsy can comprise an incisional biopsy or an excisional biopsy.

[0087] A sample can be a biological sample from a subject. The sample can be whole blood, peripheral blood, plasma, serum, saliva, mucus, urine, semen, lymph, amniotic fluid, fecal extract, cheek swab, cells or other bodily fluid or tissue, including tissue obtained through surgical biopsy or surgical resection. In some embodiments, a sample can be a primary subject (e.g., patient) derived cell line or an archived subject (e.g., patient) sample, e.g., a preserved sample, e.g., a formalin fixed paraffin embedded (FFPE) sample, or fresh frozen sample. The sample, e.g., a biological sample, can be obtained or derived from a subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a DNA or RNA collection tube, or a cell-free DNA or cell-free RNA collection tube. The sample, e.g., biological sample, can be derived from a whole blood sample by fractionation. The sample, e.g., biological sample, or derivative thereof can comprise cells. The sample, e.g., biological sample, can be a blood sample or a derivative thereof (e.g., blood collected from a collection tube or blood drops).

[0088] A sample can be a cell-free sample. A cell-free sample can be a biological sample that is substantially devoid of intact cells. The cell-free sample can be a biological sample that is itself substantially devoid of cells or can be derived from a sample from which cells have been removed. Examples of cell-free samples include those derived from blood, such as serum or plasma; urine; or samples derived from other sources, such as semen, sputum, feces, ductal exudate, lymph, or recovered lavage.

[0089] The sample can contain one or more analytes capable of being assayed. The sample can comprise one or more nucleic acid molecules. The one or more nucleic acid molecules (or any nucleic acid molecule disclosed herein, including primers and probes) can be a polymeric form a nucleotides of any length, e.g., either deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs). The nucleic acid molecules can comprise deoxyribonucleic acid (DNA). The DNA can be genomic DNA, mitochondrial DNA, amplified DNA, circular DNA, circulating DNA, cell-free DNA, or exosomal DNA. In some embodiments, the DNA is single-stranded DNA (ssDNA), double-stranded DNA, denatured double-stranded DNA, synthetic DNA, and combinations thereof. The circular DNA can be cleaved or fragmented. The DNA can comprise a coding or non-coding region of a gene or gene fragment of interest, loci (locus) defined from linkage analysis, exon, or intron. The DNA can be complementary DNA (cDNA). The nucleic acid molecule can be a recombinant nucleic acid, branched nucleic acid, plasmid, vector, or isolated DNA. A nucleic acid molecule can comprise one orWSGR Docket No.56034-703.601 more modified nucleotides, e.g., methylated nucleotides or nucleotide analogs. Modifications to the nucleotide structure can be made before or after assembly of the nucleic acid molecule. A sequence of a nucleotides of a nucleic acid molecule can be interrupted by non-nucleotide components. A nucleic acid molecule can be further modified after polymerization, such as by conjugation or binding with a reporter agent.

[0090] The nucleic acid molecule can comprise a locus, genetic locus, or genomic region, which can be identified by its location in a genome or chromosome. In some examples, a locus can be referred to by a gene name and encompass coding and non-coding regions associated with that physical region of nucleic acid. A gene can comprise coding regions (exons), non-coding regions (introns), transcriptional control or other regulatory regions, and promoters. In another example, the genomic region can incorporate an intron or exon or an intron / exon boundary within a named gene.

[0091] In some embodiments, the nucleic acid molecules comprise ribonucleic acid (RNA). The RNA can be fragmented RNA. The RNA can be degraded RNA. The RNA can be microRNA or portion thereof. The RNA can be an RNA molecule or a fragmented RNA molecule (RNA fragments) selected from: a microRNA (miRNA), a pre-miRNA, a pri- miRNA, a messenger RNA (mRNA), a pre-mRNA, circular RNA (circRNA), a ribosomal RNA (rRNA), a transfer RNA (tRNA), a pre-tRNA, a long non-coding RNA (lncRNA), a small nuclear RNA (snRNA), a circulating RNA, a cell-free RNA, an exosomal RNA, an RNA transcript, cell-free RNA, and combinations thereof.

[0092] In some embodiments, the sample comprises cell-free nucleic acid molecules. Cell-free nucleic acid molecules can include, for example, all non-encapsulated nucleic acid molecules sourced from a bodily fluid from a subject. A cell-free nucleic acid (cfNA) molecule can be a nucleic acid (e.g., cell-free RNA (cfRNA) molecule or cell-free DNA (cfDNA) molecule in a biological sample that is not contained in a cell. A cfDNA molecule can circulate freely in in a bodily fluid, such as in the bloodstream. The cell-free DNA molecule can be circulating tumor DNA, e.g., cfDNA originating from a tumor.

[0093] The sample can comprise germline nucleic acid molecules (e.g., nucleic acid from a non-diseased cell or tissue, e.g., tumor). The sample can comprise nucleic acid molecules from a tumor. In some embodiments, the sample can comprise germline nucleic acid molecules (e.g., from a non-diseased tissue) and nucleic acid molecules from a diseased tissue (e.g., a tumor).WSGR Docket No.56034-703.601

[0094] The sample can comprise a target nucleic acid molecule. A target nucleic acid molecule can be a nucleic acid molecule having a nucleotide sequence whose presence, amount, and / or sequence, or changes in one or more of these, are desired to be determined.

[0095] Processing the sample obtained from the subject can comprise subjecting the sample to conditions that are sufficient to isolate, enrich, or extract a plurality of nucleic acid molecules, and assaying the plurality of nucleic acid molecules to generate the dataset.

[0096] Nucleic acid molecules (e.g., RNA or DNA) can be extracted from a sample, e.g., using Qiagen QIAmp DNA Blood Mini Kit, FastDNA Kit protocol from MP Biomedicals, or a cell-free biological DNA isolation kit protocol from Norgen Biotek. The extraction method can extract all RNA or DNA molecules from a sample. The extract method can selectively extract a portion of RNA or DNA molecules from a sample. Extracted RNA molecules from a sample can be converted to DNA molecules by reverse transcription (RT). Reverse transcription can be the generation of deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template via the action of a reverse transcriptase.

[0097] The quality of the extracted nucleic acid can be analyzed, e.g., using BIOANALYZER or NANODROP systems. Biomarkers

[0098] Provided herein are one or more biomarkers used in the methods of the present disclosure. In some embodiments, the methods may comprise identifying a marker for determining if a subject has a phenotype or a risk thereof, or using the marker to determine if the subject has the phenotype or the risk thereof. The markers may comprise microsatellite alleles at a microsatellite locus, copy number variants (CNVs), single nucleotide polymorphisms (SNPs), indels (e.g., insertion, deletion, ratio of insertion and deletion, or any combination thereof), or a combination thereof, with one or more alleles associated with the phenotype (e.g., disease).

[0099] The marker may be used for identifying a subject having a phenotype or risk of developing a phenotype. For example, a marker may be statistically predominantly present within a population of subjects having the phenotype, relative to a population of subjects not having the phenotype. In other cases, a marker may be statistically predominantly present within a population of subjects not having the phenotype, relative to a population of subjects having the phenotype.

[0100] In some embodiments, the marker is a microsatellite alleles at microsatellite loci or genotype thereof. In some embodiments, the marker is a length of microsatellite alleles atWSGR Docket No.56034-703.601 the microsatellite locus, or the length of the microsatellite alleles at microsatellite loci making up the genotype. In some embodiments, microsatellite alleles at a microsatellite locus comprises two or more microsatellites at a locus of a nucleotide or nucleic acid sequence. In some embodiments, the microsatellite has 1 to 6 nucleotides, 2 to 5 nucleotides, or 3 to 4 nucleotides. In some embodiments, the microsatellite has at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 contiguous nucleotides. In some embodiments, the microsatellite has at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 contiguous nucleotides. In some embodiments, microsatellite loci the microsatellite comprises more than 6 nucleotides. The one or more microsatellites are upstream of an exon, downstream of an exon, in an exon, in an intergenic sequence, in an intron, in a region spanning an exon and an intron, in a 3’ untranslated region (UTR), in a 5’ UTR, or any other region in a genome. Microsatellites can occur at thousands of loci within a genome. In some embodiments, microsatellite loci can have a higher mutation rate than other areas of DNA and lead to high genetic diversity or vulnerabilities in the genetic code. In some embodiments, the microsatellites are present on a human chromosome, e.g., human chromosome 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X, or Y. In some embodiments, the microsatellites are not on an X chromosome and / or a Y chromosome.

[0101] In some embodiments, one or more markers (or the genotype thereof) may be statistically significantly associated with a phenotype. The statistical significance can be determined or measured by using the statistical analysis as described herein. In some embodiments, the marker or genotype thereof is associated with the phenotype with a P-value that is at most about 1x10^-20, 1x10^-19, 1x10^-18, 1x10^-17, 1x10^-16, 1x10^-15, 1x10^- 14, 1x10^-13, 1x10^-12, 1x10^-11, 1x10^-10, 1x10^-9, 1x10^-8, 1x10^-7, 1x10^-6, 1x10^-5, 1x10^-4, 1x10^-3, 1x10^-2, 1x10^-1, or 0.05. In some embodiments, the marker is associated with the phenotype with a P-value that is at most about 1x10^-20, 1x10^-19, 1x10^-18, 1x10^-17, 1x10^-16, 1x10^-15, 1x10^-14, 1x10^-13, 1x10^-12, 1x10^-11, 1x10^-10, 1x10^- 9, 1x10^-8, 1x10^-7, 1x10^-6, 1x10^-5, 1x10^-4, 1x10^-3, 1x10^-2, 1x10^-1, or 0.05. A marker (or the genotype thereof) may be present within a population of subjects having the phenotype with a frequency of at least about 0.001 %, 0.002 %, 0.003 %, 0.004 %, 0.005 %, 0.006 %, 0.007 %, 0.008 %, 0.009 %, 0.01 %, 0.02 %, 0.03 %, 0.04 %, 0.05 %, 0.06 %, 0.07 %, 0.08 %, 0.09 %, 0.1 %, 0.2 %, 0.3 %, 0.4 %, 0.5 %, 0.6 %, 0.7 %, 0.8 %, 0.9 %, 1 %, 1.5 %, 2 %, 2.5 %, 3 %, 3.5 %, 4 %, 4.5 %, 5 %, 10 %, 15 %, 20 %, 25 %, 30 %, 35 %, 40 %, 45 %, 50 %, 55 %, 60 %, 65 %, 70 %, 75 %, 80 %, 85 %, 90 %, 95 %, 99 %, or 100 %. In some embodiments, the marker (or the genotype thereof) may be present within a population ofWSGR Docket No.56034-703.601 subjects having the phenotype with a frequency of at most about 0.001 %, 0.002 %, 0.003 %, 0.004 %, 0.005 %, 0.006 %, 0.007 %, 0.008 %, 0.009 %, 0.01 %, 0.02 %, 0.03 %, 0.04 %, 0.05 %, 0.06 %, 0.07 %, 0.08 %, 0.09 %, 0.1 %, 0.2 %, 0.3 %, 0.4 %, 0.5 %, 0.6 %, 0.7 %, 0.8 %, 0.9 %, 1 %, 1.5 %, 2 %, 2.5 %, 3 %, 3.5 %, 4 %, 4.5 %, 5 %, 10 %, 15 %, 20 %, 25 %, 30 %, 35 %, 40 %, 45 %, 50 %, 55 %, 60 %, 65 %, 70 %, 75 %, 80 %, 85 %, 90 %, 95 %, 99 %, or 100 %. In some case, the marker (or the genotype thereof) may be presented within a population of subjects not having the disease with a frequency of at least about 0.001 %, 0.002 %, 0.003 %, 0.004 %, 0.005 %, 0.006 %, 0.007 %, 0.008 %, 0.009 %, 0.01 %, 0.02 %, 0.03 %, 0.04 %, 0.05 %, 0.06 %, 0.07 %, 0.08 %, 0.09 %, 0.1 %, 0.2 %, 0.3 %, 0.4 %, 0.5 %, 0.6 %, 0.7 %, 0.8 %, 0.9 %, 1 %, 1.5 %, 2 %, 2.5 %, 3 %, 3.5 %, 4 %, 4.5 %, 5 %, 10 %, 15 %, 20 %, 25 %, 30 %, 35 %, 40 %, 45 %, 50 %, 55 %, 60 %, 65 %, 70 %, 75 %, 80 %, 85 %, 90 %, 95 %, 99 %, or 100 %. In some case, the marker (or the genotype thereof) may be presented within a population of subjects not having the disease with a frequency of at most about 0.001 %, 0.002 %, 0.003 %, 0.004 %, 0.005 %, 0.006 %, 0.007 %, 0.008 %, 0.009 %, 0.01 %, 0.02 %, 0.03 %, 0.04 %, 0.05 %, 0.06 %, 0.07 %, 0.08 %, 0.09 %, 0.1 %, 0.2 %, 0.3 %, 0.4 %, 0.5 %, 0.6 %, 0.7 %, 0.8 %, 0.9 %, 1 %, 1.5 %, 2 %, 2.5 %, 3 %, 3.5 %, 4 %, 4.5 %, 5 %, 10 %, 15 %, 20 %, 25 %, 30 %, 35 %, 40 %, 45 %, 50 %, 55 %, 60 %, 65 %, 70 %, 75 %, 80 %, 85 %, 90 %, 95 %, 99 %, or 100 %. In some embodiments, the marker is a microsatellite alleles at microsatellite locus or genotype thereof, a SNP, a CNV or an indel. In some embodiments, the marker is a length of a microsatellite alleles at microsatellite locus or a genotype thereof. In some embodiments, the marker is a minor microsatellite allele, a primary microsatellite allele or a combination thereof.

[0102] In some embodiments, the marker is a microsatellite allele at a microsatellite locus listed in Table 1A. In some embodiments, the marker comprises a microsatellite locus at chr1:10442251-10442260, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 10. In some embodiments, the marker comprises a microsatellite locus at chr1:36009426-36009435, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 11. In some embodiments, the marker comprises a microsatellite locus at chr1:43198649- 43198658, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 12. In some embodiments, the marker comprises a microsatellite locus at chr1:85045375-85045392, wherein the microsatellite locus comprises a repeat comprising "AAC" starting at nucleobase position 101 in SEQ ID NO: 13. In some embodiments, the marker comprises a microsatellite locus at chr1:94108787-94108798,WSGR Docket No.56034-703.601 wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 14. In some embodiments, the marker comprises a microsatellite locus at chr1:111489033-111489068, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 15. In some embodiments, the marker comprises a microsatellite locus at chr1:114398011-114398022, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 16. In some embodiments, the marker comprises a microsatellite locus at chr1:120641781-120641790, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 17. In some embodiments, the marker comprises a microsatellite locus at chr1:121088984-121088996, wherein the microsatellite locus comprises a repeat comprising "TG" starting at nucleobase position 101 in SEQ ID NO: 18. In some embodiments, the marker comprises a microsatellite locus at chr1:152223253-152223263, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 19. In some embodiments, the marker comprises a microsatellite locus at chr1:205202118-205202154, wherein the microsatellite locus comprises a repeat comprising "GGGAA" starting at nucleobase position 101 in SEQ ID NO: 20. In some embodiments, the marker comprises a microsatellite locus at chr1:235663083-235663095, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 21. In some embodiments, the marker comprises a microsatellite locus at chr1:236574575-236574585, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 22. In some embodiments, the marker comprises a microsatellite locus at chr10:12235131- 12235149, wherein the microsatellite locus comprises a repeat comprising "TTTG" starting at nucleobase position 101 in SEQ ID NO: 23. In some embodiments, the marker comprises a microsatellite locus at chr10:17909285-17909297, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 24. In some embodiments, the marker comprises a microsatellite locus at chr10:46772386-46772416, wherein the microsatellite locus comprises a repeat comprising "CAAAAA" starting at nucleobase position 101 in SEQ ID NO: 25. In some embodiments, the marker comprises a microsatellite locus at chr10:47568021-47568045, wherein the microsatellite locus comprises a repeat comprising "GTTTTT" starting at nucleobase position 101 in SEQ ID NO: 26. In some embodiments, the marker comprises a microsatellite locus at chr10:73454380- 73454390, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 27. In some embodiments, the marker comprises aWSGR Docket No.56034-703.601 microsatellite locus at chr10:75022148-75022171, wherein the microsatellite locus comprises a repeat comprising "GAA" starting at nucleobase position 101 in SEQ ID NO: 28. In some embodiments, the marker comprises a microsatellite locus at chr10:93355583-93355605, wherein the microsatellite locus comprises a repeat comprising "AAAAT" starting at nucleobase position 101 in SEQ ID NO: 29. In some embodiments, the marker comprises a microsatellite locus at chr11:17089651-17089679, wherein the microsatellite locus comprises a repeat comprising "GT" starting at nucleobase position 101 in SEQ ID NO: 30. In some embodiments, the marker comprises a microsatellite locus at chr11:65870267-65870288, wherein the microsatellite locus comprises a repeat comprising "GAGGGCA" starting at nucleobase position 101 in SEQ ID NO: 31. In some embodiments, the marker comprises a microsatellite locus at chr11:121304570-121304579, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 32. In some embodiments, the marker comprises a microsatellite locus at chr11:121496849- 121496861, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 33. In some embodiments, the marker comprises a microsatellite locus at chr12:1800491-1800500, wherein the microsatellite locus comprises a repeat comprising "C" starting at nucleobase position 101 in SEQ ID NO: 34. In some embodiments, the marker comprises a microsatellite locus at chr12:6867711-6867732, wherein the microsatellite locus comprises a repeat comprising "GGGCCG" starting at nucleobase position 101 in SEQ ID NO: 35. In some embodiments, the marker comprises a microsatellite locus at chr12:15549269-15549284, wherein the microsatellite locus comprises a repeat comprising "AT" starting at nucleobase position 101 in SEQ ID NO: 36. In some embodiments, the marker comprises a microsatellite locus at chr12:43432475-43432486, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 37. In some embodiments, the marker comprises a microsatellite locus at chr12:46788623-46788637, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 38. In some embodiments, the marker comprises a microsatellite locus at chr12:54961093-54961104, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 39. In some embodiments, the marker comprises a microsatellite locus at chr12:107544148-107544164, wherein the microsatellite locus comprises a repeat comprising "TGCCT" starting at nucleobase position 101 in SEQ ID NO: 40. In some embodiments, the marker comprises a microsatellite locus at chr12:111263683-111263692, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobaseWSGR Docket No.56034-703.601 position 101 in SEQ ID NO: 41. In some embodiments, the marker comprises a microsatellite locus at chr12:111447548-111447572, wherein the microsatellite locus comprises a repeat comprising "TGGGG" starting at nucleobase position 101 in SEQ ID NO: 42. In some embodiments, the marker comprises a microsatellite locus at chr12:113180951-113180965, wherein the microsatellite locus comprises a repeat comprising "CTT" starting at nucleobase position 101 in SEQ ID NO: 43. In some embodiments, the marker comprises a microsatellite locus at chr13:40808422-40808431, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 44. In some embodiments, the marker comprises a microsatellite locus at chr13:99864924-99864956, wherein the microsatellite locus comprises a repeat comprising "TC" starting at nucleobase position 101 in SEQ ID NO: 45. In some embodiments, the marker comprises a microsatellite locus at chr14:36320608-36320625, wherein the microsatellite locus comprises a repeat comprising "CCA" starting at nucleobase position 101 in SEQ ID NO: 46. In some embodiments, the marker comprises a microsatellite locus at chr14:68979142-68979163, wherein the microsatellite locus comprises a repeat comprising "GGGCT" starting at nucleobase position 101 in SEQ ID NO: 47. In some embodiments, the marker comprises a microsatellite locus at chr14:75041563-75041587, wherein the microsatellite locus comprises a repeat comprising "AAG" starting at nucleobase position 101 in SEQ ID NO: 48. In some embodiments, the marker comprises a microsatellite locus at chr15:41844823-41844841, wherein the microsatellite locus comprises a repeat comprising "TGGCTT" starting at nucleobase position 101 in SEQ ID NO: 49. In some embodiments, the marker comprises a microsatellite locus at chr15:43492238-43492248, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 50. In some embodiments, the marker comprises a microsatellite locus at chr15:44736722-44736732, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 51. In some embodiments, the marker comprises a microsatellite locus at chr15:58599751-58599762, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 52. In some embodiments, the marker comprises a microsatellite locus at chr15:59639174-59639191, wherein the microsatellite locus comprises a repeat comprising "AAAGT" starting at nucleobase position 101 in SEQ ID NO: 53. In some embodiments, the marker comprises a microsatellite locus at chr15:72163493-72163502, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 54. In some embodiments, the marker comprises a microsatellite locus at chr15:83858896-83858916,WSGR Docket No.56034-703.601 wherein the microsatellite locus comprises a repeat comprising "AAAAAAC" starting at nucleobase position 101 in SEQ ID NO: 55. In some embodiments, the marker comprises a microsatellite locus at chr16:3758052-3758064, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 56. In some embodiments, the marker comprises a microsatellite locus at chr16:71175496-71175523, wherein the microsatellite locus comprises a repeat comprising "ACC" starting at nucleobase position 101 in SEQ ID NO: 57. In some embodiments, the marker comprises a microsatellite locus at chr17:15230798-15230812, wherein the microsatellite locus comprises a repeat comprising "GTTTG" starting at nucleobase position 101 in SEQ ID NO: 58. In some embodiments, the marker comprises a microsatellite locus at chr17:21302047-21302066, wherein the microsatellite locus comprises a repeat comprising "GGGGCT" starting at nucleobase position 101 in SEQ ID NO: 59. In some embodiments, the marker comprises a microsatellite locus at chr17:44172350-44172362, wherein the microsatellite locus comprises a repeat comprising "TG" starting at nucleobase position 101 in SEQ ID NO: 60. In some embodiments, the marker comprises a microsatellite locus at chr17:46332539-46332549, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 61. In some embodiments, the marker comprises a microsatellite locus at chr17:50199688-50199738, wherein the microsatellite locus comprises a repeat comprising "CCCCAG" starting at nucleobase position 101 in SEQ ID NO: 62. In some embodiments, the marker comprises a microsatellite locus at chr18:12963148-12963157, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 63. In some embodiments, the marker comprises a microsatellite locus at chr19:2808459-2808492, wherein the microsatellite locus comprises a repeat comprising "GGGCAG" starting at nucleobase position 101 in SEQ ID NO: 64. In some embodiments, the marker comprises a microsatellite locus at chr19:38922117-38922126, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 65. In some embodiments, the marker comprises a microsatellite locus at chr19:45409232-45409251, wherein the microsatellite locus comprises a repeat comprising "AAG" starting at nucleobase position 101 in SEQ ID NO: 66. In some embodiments, the marker comprises a microsatellite locus at chr19:49045346-49045357, wherein the microsatellite locus comprises a repeat comprising "CA" starting at nucleobase position 101 in SEQ ID NO: 67. In some embodiments, the marker comprises a microsatellite locus at chr19:53223893-53223911, wherein the microsatellite locus comprises a repeat comprising "GTTTT" starting at nucleobase position 101 in SEQ ID NO: 68. In someWSGR Docket No.56034-703.601 embodiments, the marker comprises a microsatellite locus at chr19:54805834-54805850, wherein the microsatellite locus comprises a repeat comprising "GA" starting at nucleobase position 101 in SEQ ID NO: 69. In some embodiments, the marker comprises a microsatellite locus at chr2:9412255-9412266, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 70. In some embodiments, the marker comprises a microsatellite locus at chr2:9630461-9630478, wherein the microsatellite locus comprises a repeat comprising "CGGGGC" starting at nucleobase position 101 in SEQ ID NO: 71. In some embodiments, the marker comprises a microsatellite locus at chr2:25434021-25434031, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 72. In some embodiments, the marker comprises a microsatellite locus at chr2:61311469-61311489, wherein the microsatellite locus comprises a repeat comprising "AAAGAG" starting at nucleobase position 101 in SEQ ID NO: 73. In some embodiments, the marker comprises a microsatellite locus at chr2:72961093-72961111, wherein the microsatellite locus comprises a repeat comprising "CCCTG" starting at nucleobase position 101 in SEQ ID NO: 74. In some embodiments, the marker comprises a microsatellite locus at chr2:91699880-91699892, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 75. In some embodiments, the marker comprises a microsatellite locus at chr2:160016486-160016499, wherein the microsatellite locus comprises a repeat comprising "GA" starting at nucleobase position 101 in SEQ ID NO: 76. In some embodiments, the marker comprises a microsatellite locus at chr2:161993435-161993445, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 77. In some embodiments, the marker comprises a microsatellite locus at chr2:165162255-165162264, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 78. In some embodiments, the marker comprises a microsatellite locus at chr2:170848648-170848658, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 79. In some embodiments, the marker comprises a microsatellite locus at chr2:182967373-182967384, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 80. In some embodiments, the marker comprises a microsatellite locus at chr2:227310700-227310709, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 81. In some embodiments, the marker comprises a microsatellite locus at chr20:2616093- 2616115, wherein the microsatellite locus comprises a repeat comprising "T" starting atWSGR Docket No.56034-703.601 nucleobase position 101 in SEQ ID NO: 82. In some embodiments, the marker comprises a microsatellite locus at chr20:47169102-47169122, wherein the microsatellite locus comprises a repeat comprising "TC" starting at nucleobase position 101 in SEQ ID NO: 83. In some embodiments, the marker comprises a microsatellite locus at chr20:47642168-47642178, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 84. In some embodiments, the marker comprises a microsatellite locus at chr20:53571861-53571870, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 85. In some embodiments, the marker comprises a microsatellite locus at chr21:35599704-35599721, wherein the microsatellite locus comprises a repeat comprising "GAG" starting at nucleobase position 101 in SEQ ID NO: 86. In some embodiments, the marker comprises a microsatellite locus at chr21:37122954-37122969, wherein the microsatellite locus comprises a repeat comprising "TG" starting at nucleobase position 101 in SEQ ID NO: 87. In some embodiments, the marker comprises a microsatellite locus at chr21:42129699-42129709, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 88. In some embodiments, the marker comprises a microsatellite locus at chr3:64540976-64541000, wherein the microsatellite locus comprises a repeat comprising "CAATGTG" starting at nucleobase position 101 in SEQ ID NO: 89. In some embodiments, the marker comprises a microsatellite locus at chr3:70959191-70959203, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 90. In some embodiments, the marker comprises a microsatellite locus at chr3:108437406-108437415, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 91. In some embodiments, the marker comprises a microsatellite locus at chr3:136899869-136899878, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 92. In some embodiments, the marker comprises a microsatellite locus at chr4:127896751-127896763, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 93. In some embodiments, the marker comprises a microsatellite locus at chr4:145155481-145155491, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 94. In some embodiments, the marker comprises a microsatellite locus at chr4:147906547-147906556, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 95. In some embodiments, the marker comprises a microsatellite locus at chr5:33963627-33963643, wherein the microsatelliteWSGR Docket No.56034-703.601 locus comprises a repeat comprising "AAC" starting at nucleobase position 101 in SEQ ID NO: 96. In some embodiments, the marker comprises a microsatellite locus at chr5:34042762-34042777, wherein the microsatellite locus comprises a repeat comprising "AAC" starting at nucleobase position 101 in SEQ ID NO: 97. In some embodiments, the marker comprises a microsatellite locus at chr5:56882022-56882047, wherein the microsatellite locus comprises a repeat comprising "CAA" starting at nucleobase position 101 in SEQ ID NO: 98. In some embodiments, the marker comprises a microsatellite locus at chr5:116051838-116051860, wherein the microsatellite locus comprises a repeat comprising "TTTTTA" starting at nucleobase position 101 in SEQ ID NO: 99. In some embodiments, the marker comprises a microsatellite locus at chr5:163439610-163439619, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 100. In some embodiments, the marker comprises a microsatellite locus at chr6:20758578-20758590, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 101. In some embodiments, the marker comprises a microsatellite locus at chr6:31678970-31678980, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 102. In some embodiments, the marker comprises a microsatellite locus at chr6:131582815-131582831, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 103. In some embodiments, the marker comprises a microsatellite locus at chr7:29504722-29504738, wherein the microsatellite locus comprises a repeat comprising "TC" starting at nucleobase position 101 in SEQ ID NO: 104. In some embodiments, the marker comprises a microsatellite locus at chr7:74753041-74753054, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 105. In some embodiments, the marker comprises a microsatellite locus at chr7:114663437-114663446, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 106. In some embodiments, the marker comprises a microsatellite locus at chr9:33912105-33912115, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 107. In some embodiments, the marker comprises a microsatellite locus at chr9:39817736-39817764, wherein the microsatellite locus comprises a repeat comprising "AAAC" starting at nucleobase position 101 in SEQ ID NO: 108. In some embodiments, the marker comprises a microsatellite locus at chr9:100498685-100498698, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 109. In someWSGR Docket No.56034-703.601 embodiments, the marker comprises a microsatellite locus at chr9:124524880-124524903, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 110. In some embodiments, the marker comprises a microsatellite locus at chr9:135096053-135096069, wherein the microsatellite locus comprises a repeat comprising "TCC" starting at nucleobase position 101 in SEQ ID NO: 111. In some embodiments, the marker comprises a microsatellite locus at chrX:38561513- 38561534, wherein the microsatellite locus comprises a repeat comprising "GCC" starting at nucleobase position 101 in SEQ ID NO: 112. In some embodiments, the marker comprises a microsatellite locus at chrX:49183661-49183679, wherein the microsatellite locus comprises a repeat comprising "GT" starting at nucleobase position 101 in SEQ ID NO: 113. In some embodiments, the marker comprises a microsatellite locus at chrX:116963542-116963562, wherein the microsatellite locus comprises a repeat comprising "GAT" starting at nucleobase position 101 in SEQ ID NO: 114. In some embodiments, the marker comprises a microsatellite locus at chr1:26189039-26189060, wherein the microsatellite locus comprises a repeat comprising "CCAGACT" starting at nucleobase position 101 in SEQ ID NO: 115. In some embodiments, the marker comprises a microsatellite locus at chr1:31433043-31433057, wherein the microsatellite locus comprises a repeat comprising "CAG" starting at nucleobase position 101 in SEQ ID NO: 116. In some embodiments, the marker comprises a microsatellite locus at chr1:35738059-35738073, wherein the microsatellite locus comprises a repeat comprising "TTC" starting at nucleobase position 101 in SEQ ID NO: 117. In some embodiments, the marker comprises a microsatellite locus at chr1:64673105-64673117, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 118. In some embodiments, the marker comprises a microsatellite locus at chr1:110517191-110517222, wherein the microsatellite locus comprises a repeat comprising "GA" starting at nucleobase position 101 in SEQ ID NO: 119. In some embodiments, the marker comprises a microsatellite locus at chr1:244857990- 244858002, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 120. In some embodiments, the marker comprises a microsatellite locus at chr10:12198693-12198702, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 121. In some embodiments, the marker comprises a microsatellite locus at chr10:32804931-32804952, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 122. In some embodiments, the marker comprises a microsatellite locus at chr10:45139349-45139393, wherein the microsatellite locus comprisesWSGR Docket No.56034-703.601 a repeat comprising "TCTT" starting at nucleobase position 101 in SEQ ID NO: 123. In some embodiments, the marker comprises a microsatellite locus at chr10:133402253-133402279, wherein the microsatellite locus comprises a repeat comprising "GGGGCT" starting at nucleobase position 101 in SEQ ID NO: 124. In some embodiments, the marker comprises a microsatellite locus at chr11:18105984-18106001, wherein the microsatellite locus comprises a repeat comprising "CCT" starting at nucleobase position 101 in SEQ ID NO: 125. In some embodiments, the marker comprises a microsatellite locus at chr11:112182842-112182853, wherein the microsatellite locus comprises a repeat comprising "TA" starting at nucleobase position 101 in SEQ ID NO: 126. In some embodiments, the marker comprises a microsatellite locus at chr11:130295631-130295641, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 127. In some embodiments, the marker comprises a microsatellite locus at chr12:122204563- 122204579, wherein the microsatellite locus comprises a repeat comprising "CCG" starting at nucleobase position 101 in SEQ ID NO: 128. In some embodiments, the marker comprises a microsatellite locus at chr13:32741052-32741064, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 129. In some embodiments, the marker comprises a microsatellite locus at chr13:49792545-49792562, wherein the microsatellite locus comprises a repeat comprising "GCG" starting at nucleobase position 101 in SEQ ID NO: 130. In some embodiments, the marker comprises a microsatellite locus at chr13:72719112-72719125, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 131. In some embodiments, the marker comprises a microsatellite locus at chr15:20458509-20458521, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 132. In some embodiments, the marker comprises a microsatellite locus at chr15:40035007-40035022, wherein the microsatellite locus comprises a repeat comprising "TCTTT" starting at nucleobase position 101 in SEQ ID NO: 133. In some embodiments, the marker comprises a microsatellite locus at chr15:51458895- 51458919, wherein the microsatellite locus comprises a repeat comprising "AGC" starting at nucleobase position 101 in SEQ ID NO: 134. In some embodiments, the marker comprises a microsatellite locus at chr15:69417297-69417315, wherein the microsatellite locus comprises a repeat comprising "TGAA" starting at nucleobase position 101 in SEQ ID NO: 135. In some embodiments, the marker comprises a microsatellite locus at chr16:67438474- 67438495, wherein the microsatellite locus comprises a repeat comprising "CA" starting at nucleobase position 101 in SEQ ID NO: 136. In some embodiments, the marker comprises aWSGR Docket No.56034-703.601 microsatellite locus at chr17:28468443-28468454, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 137. In some embodiments, the marker comprises a microsatellite locus at chr17:81700634-81700665, wherein the microsatellite locus comprises a repeat comprising "CCAGCC" starting at nucleobase position 101 in SEQ ID NO: 138. In some embodiments, the marker comprises a microsatellite locus at chr17:81719696-81719720, wherein the microsatellite locus comprises a repeat comprising "CACACC" starting at nucleobase position 101 in SEQ ID NO: 139. In some embodiments, the marker comprises a microsatellite locus at chr18:41962443- 41962454, wherein the microsatellite locus comprises a repeat comprising "AT" starting at nucleobase position 101 in SEQ ID NO: 140. In some embodiments, the marker comprises a microsatellite locus at chr18:58745651-58745661, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 141. In some embodiments, the marker comprises a microsatellite locus at chr18:79715188-79715208, wherein the microsatellite locus comprises a repeat comprising "GGA" starting at nucleobase position 101 in SEQ ID NO: 142. In some embodiments, the marker comprises a microsatellite locus at chr19:6908665-6908677, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 143. In some embodiments, the marker comprises a microsatellite locus at chr19:19337125-19337136, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 144. In some embodiments, the marker comprises a microsatellite locus at chr19:38408862-38408879, wherein the microsatellite locus comprises a repeat comprising "AAG" starting at nucleobase position 101 in SEQ ID NO: 145. In some embodiments, the marker comprises a microsatellite locus at chr19:48836426-48836440, wherein the microsatellite locus comprises a repeat comprising "TC" starting at nucleobase position 101 in SEQ ID NO: 146. In some embodiments, the marker comprises a microsatellite locus at chr19:55014736-55014753, wherein the microsatellite locus comprises a repeat comprising "CAGA" starting at nucleobase position 101 in SEQ ID NO: 147. In some embodiments, the marker comprises a microsatellite locus at chr19:55279519- 55279533, wherein the microsatellite locus comprises a repeat comprising "GCC" starting at nucleobase position 101 in SEQ ID NO: 148. In some embodiments, the marker comprises a microsatellite locus at chr2:9335072-9335084, wherein the microsatellite locus comprises a repeat comprising "TG" starting at nucleobase position 101 in SEQ ID NO: 149. In some embodiments, the marker comprises a microsatellite locus at chr2:38300256-38300269, wherein the microsatellite locus comprises a repeat comprising "CT" starting at nucleobaseWSGR Docket No.56034-703.601 position 101 in SEQ ID NO: 150. In some embodiments, the marker comprises a microsatellite locus at chr2:41850169-41850184, wherein the microsatellite locus comprises a repeat comprising "TG" starting at nucleobase position 101 in SEQ ID NO: 151. In some embodiments, the marker comprises a microsatellite locus at chr2:63752198-63752208, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 152. In some embodiments, the marker comprises a microsatellite locus at chr2:134946081-134946093, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 153. In some embodiments, the marker comprises a microsatellite locus at chr2:149213132- 149213143, wherein the microsatellite locus comprises a repeat comprising "GT" starting at nucleobase position 101 in SEQ ID NO: 154. In some embodiments, the marker comprises a microsatellite locus at chr2:215052616-215052631, wherein the microsatellite locus comprises a repeat comprising "CA" starting at nucleobase position 101 in SEQ ID NO: 155. In some embodiments, the marker comprises a microsatellite locus at chr2:237087093- 237087109, wherein the microsatellite locus comprises a repeat comprising "TTGTT" starting at nucleobase position 101 in SEQ ID NO: 156. In some embodiments, the marker comprises a microsatellite locus at chr22:30345242-30345262, wherein the microsatellite locus comprises a repeat comprising "AGGGG" starting at nucleobase position 101 in SEQ ID NO: 157. In some embodiments, the marker comprises a microsatellite locus at chr22:43972849-43972860, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 158. In some embodiments, the marker comprises a microsatellite locus at chr3:48177831-48177848, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 159. In some embodiments, the marker comprises a microsatellite locus at chr3:184711346-184711366, wherein the microsatellite locus comprises a repeat comprising "TCC" starting at nucleobase position 101 in SEQ ID NO: 160. In some embodiments, the marker comprises a microsatellite locus at chr3:196369530-196369539, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 161. In some embodiments, the marker comprises a microsatellite locus at chr4:103196256-103196280, wherein the microsatellite locus comprises a repeat comprising "AAC" starting at nucleobase position 101 in SEQ ID NO: 162. In some embodiments, the marker comprises a microsatellite locus at chr4:146707554-146707563, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 163. In some embodiments, the marker comprises a microsatellite locus atWSGR Docket No.56034-703.601 chr5:155492237-155492249, wherein the microsatellite locus comprises a repeat comprising "AG" starting at nucleobase position 101 in SEQ ID NO: 164. In some embodiments, the marker comprises a microsatellite locus at chr5:177212217-177212229, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 165. In some embodiments, the marker comprises a microsatellite locus at chr6:7727290-7727310, wherein the microsatellite locus comprises a repeat comprising "AGC" starting at nucleobase position 101 in SEQ ID NO: 166. In some embodiments, the marker comprises a microsatellite locus at chr6:39079260-39079276, wherein the microsatellite locus comprises a repeat comprising "AG" starting at nucleobase position 101 in SEQ ID NO: 167. In some embodiments, the marker comprises a microsatellite locus at chr6:52535687-52535698, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 168. In some embodiments, the marker comprises a microsatellite locus at chr6:162814127-162814137, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 169. In some embodiments, the marker comprises a microsatellite locus at chr7:16308638-16308654, wherein the microsatellite locus comprises a repeat comprising "AAC" starting at nucleobase position 101 in SEQ ID NO: 170. In some embodiments, the marker comprises a microsatellite locus at chr7:16682768-16682777, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 171. In some embodiments, the marker comprises a microsatellite locus at chr7:77254490-77254501, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 172. In some embodiments, the marker comprises a microsatellite locus at chr7:87375972-87375982, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 173. In some embodiments, the marker comprises a microsatellite locus at chr7:99040294-99040308, wherein the microsatellite locus comprises a repeat comprising "CG" starting at nucleobase position 101 in SEQ ID NO: 174. In some embodiments, the marker comprises a microsatellite locus at chr7:137690016-137690036, wherein the microsatellite locus comprises a repeat comprising "AAAAAC" starting at nucleobase position 101 in SEQ ID NO: 175. In some embodiments, the marker comprises a microsatellite locus at chr8:35705907-35705925, wherein the microsatellite locus comprises a repeat comprising "TTTC" starting at nucleobase position 101 in SEQ ID NO: 176. In some embodiments, the marker comprises a microsatellite locus at chr8:103910292-103910305, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobaseWSGR Docket No.56034-703.601 position 101 in SEQ ID NO: 177. In some embodiments, the marker comprises a microsatellite locus at chrX:9654411-9654428, wherein the microsatellite locus comprises a repeat comprising "AGA" starting at nucleobase position 101 in SEQ ID NO: 178. In some embodiments, the marker comprises a microsatellite locus at chrX:17725235-17725245, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 179. In some embodiments, the marker comprises a microsatellite locus at chrX:19355530-19355549, wherein the microsatellite locus comprises a repeat comprising "GGCCAA" starting at nucleobase position 101 in SEQ ID NO: 180. In some embodiments, the marker comprises a microsatellite locus at chrX:35925961- 35925970, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 181. In some embodiments, the marker comprises a microsatellite locus at chrX:71389519-71389528, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 182. In some embodiments, the marker comprises a microsatellite locus at chrX:100915089-100915113, wherein the microsatellite locus comprises a repeat comprising "AAAAAC" starting at nucleobase position 101 in SEQ ID NO: 183. In some embodiments, the marker comprises a microsatellite locus at chrX:120253979-120253998, wherein the microsatellite locus comprises a repeat comprising "TGA" starting at nucleobase position 101 in SEQ ID NO: 184. In some embodiments, the marker comprises a microsatellite locus at chr1:43592645- 43592659, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 185. In some embodiments, the marker comprises a microsatellite locus at chr1:89193452-89193473, wherein the microsatellite locus comprises a repeat comprising "TTCA" starting at nucleobase position 101 in SEQ ID NO: 186. In some embodiments, the marker comprises a microsatellite locus at chr1:91981405-91981428, wherein the microsatellite locus comprises a repeat comprising "TGTA" starting at nucleobase position 101 in SEQ ID NO: 187. In some embodiments, the marker comprises a microsatellite locus at chr1:95150193-95150204, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 188. In some embodiments, the marker comprises a microsatellite locus at chr1:108923941-108923967, wherein the microsatellite locus comprises a repeat comprising "TTTTTTG" starting at nucleobase position 101 in SEQ ID NO: 189. In some embodiments, the marker comprises a microsatellite locus at chr1:154869855-154869879, wherein the microsatellite locus comprises a repeat comprising "TGC" starting at nucleobase position 101 in SEQ ID NO: 190. In some embodiments, the marker comprises a microsatellite locus at chr1:158699553-WSGR Docket No.56034-703.601 158699579, wherein the microsatellite locus comprises a repeat comprising "GTTTG" starting at nucleobase position 101 in SEQ ID NO: 191. In some embodiments, the marker comprises a microsatellite locus at chr1:186676720-186676738, wherein the microsatellite locus comprises a repeat comprising "AAAT" starting at nucleobase position 101 in SEQ ID NO: 192. In some embodiments, the marker comprises a microsatellite locus at chr1:200850297-200850306, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 193. In some embodiments, the marker comprises a microsatellite locus at chr1:216190380-216190404, wherein the microsatellite locus comprises a repeat comprising "AAAG" starting at nucleobase position 101 in SEQ ID NO: 194. In some embodiments, the marker comprises a microsatellite locus at chr1:223363361-223363382, wherein the microsatellite locus comprises a repeat comprising "TGC" starting at nucleobase position 101 in SEQ ID NO: 195. In some embodiments, the marker comprises a microsatellite locus at chr1:228103091-228103115, wherein the microsatellite locus comprises a repeat comprising "CCCGGAG" starting at nucleobase position 101 in SEQ ID NO: 196. In some embodiments, the marker comprises a microsatellite locus at chr1:245017338-245017348, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 197. In some embodiments, the marker comprises a microsatellite locus at chr10:59814600- 59814623, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 198. In some embodiments, the marker comprises a microsatellite locus at chr10:68422212-68422222, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 199. In some embodiments, the marker comprises a microsatellite locus at chr10:72274098-72274116, wherein the microsatellite locus comprises a repeat comprising "GGTCT" starting at nucleobase position 101 in SEQ ID NO: 200. In some embodiments, the marker comprises a microsatellite locus at chr10:87715884-87715894, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 201. In some embodiments, the marker comprises a microsatellite locus at chr10:99698498-99698518, wherein the microsatellite locus comprises a repeat comprising "TGTT" starting at nucleobase position 101 in SEQ ID NO: 202. In some embodiments, the marker comprises a microsatellite locus at chr11:281785-281799, wherein the microsatellite locus comprises a repeat comprising "AGA" starting at nucleobase position 101 in SEQ ID NO: 203. In some embodiments, the marker comprises a microsatellite locus at chr11:3360569-3360583, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobaseWSGR Docket No.56034-703.601 position 101 in SEQ ID NO: 204. In some embodiments, the marker comprises a microsatellite locus at chr11:58612720-58612731, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 205. In some embodiments, the marker comprises a microsatellite locus at chr11:59600746-59600761, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 206. In some embodiments, the marker comprises a microsatellite locus at chr11:61955085-61955105, wherein the microsatellite locus comprises a repeat comprising "CCACCC" starting at nucleobase position 101 in SEQ ID NO: 207. In some embodiments, the marker comprises a microsatellite locus at chr11:64928256- 64928280, wherein the microsatellite locus comprises a repeat comprising "CCCTGA" starting at nucleobase position 101 in SEQ ID NO: 208. In some embodiments, the marker comprises a microsatellite locus at chr11:70325462-70325476, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 209. In some embodiments, the marker comprises a microsatellite locus at chr11:72597468- 72597493, wherein the microsatellite locus comprises a repeat comprising "CCCTGC" starting at nucleobase position 101 in SEQ ID NO: 210. In some embodiments, the marker comprises a microsatellite locus at chr11:75566996-75567010, wherein the microsatellite locus comprises a repeat comprising "TCC" starting at nucleobase position 101 in SEQ ID NO: 211. In some embodiments, the marker comprises a microsatellite locus at chr11:89662136-89662146, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 212. In some embodiments, the marker comprises a microsatellite locus at chr11:108325986-108326000, wherein the microsatellite locus comprises a repeat comprising "ATT" starting at nucleobase position 101 in SEQ ID NO: 213. In some embodiments, the marker comprises a microsatellite locus at chr11:116820796-116820830, wherein the microsatellite locus comprises a repeat comprising "GACA" starting at nucleobase position 101 in SEQ ID NO: 214. In some embodiments, the marker comprises a microsatellite locus at chr11:117410000-117410015, wherein the microsatellite locus comprises a repeat comprising "CCT" starting at nucleobase position 101 in SEQ ID NO: 215. In some embodiments, the marker comprises a microsatellite locus at chr11:119664968-119664992, wherein the microsatellite locus comprises a repeat comprising "CCT" starting at nucleobase position 101 in SEQ ID NO: 216. In some embodiments, the marker comprises a microsatellite locus at chr12:1953158-1953186, wherein the microsatellite locus comprises a repeat comprising "TGC" starting at nucleobase position 101 in SEQ ID NO: 217. In some embodiments, the marker comprises a microsatellite locus atWSGR Docket No.56034-703.601 chr12:9420667-9420698, wherein the microsatellite locus comprises a repeat comprising "AAGG" starting at nucleobase position 101 in SEQ ID NO: 218. In some embodiments, the marker comprises a microsatellite locus at chr12:15899834-15899845, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 219. In some embodiments, the marker comprises a microsatellite locus at chr12:53669017-53669038, wherein the microsatellite locus comprises a repeat comprising "AAAC" starting at nucleobase position 101 in SEQ ID NO: 220. In some embodiments, the marker comprises a microsatellite locus at chr12:88532514-88532524, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 221. In some embodiments, the marker comprises a microsatellite locus at chr12:121750485-121750505, wherein the microsatellite locus comprises a repeat comprising "CCCACA" starting at nucleobase position 101 in SEQ ID NO: 222. In some embodiments, the marker comprises a microsatellite locus at chr13:51908303-51908321, wherein the microsatellite locus comprises a repeat comprising "GTGGGG" starting at nucleobase position 101 in SEQ ID NO: 223. In some embodiments, the marker comprises a microsatellite locus at chr13:75713297-75713308, wherein the microsatellite locus comprises a repeat comprising "TG" starting at nucleobase position 101 in SEQ ID NO: 224. In some embodiments, the marker comprises a microsatellite locus at chr13:110211628-110211642, wherein the microsatellite locus comprises a repeat comprising "AAAAT" starting at nucleobase position 101 in SEQ ID NO: 225. In some embodiments, the marker comprises a microsatellite locus at chr14:39150384-39150394, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 226. In some embodiments, the marker comprises a microsatellite locus at chr14:92000141-92000155, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 227. In some embodiments, the marker comprises a microsatellite locus at chr15:25928097-25928106, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 228. In some embodiments, the marker comprises a microsatellite locus at chr15:42082404-42082430, wherein the microsatellite locus comprises a repeat comprising "TCAT" starting at nucleobase position 101 in SEQ ID NO: 229. In some embodiments, the marker comprises a microsatellite locus at chr15:53613624-53613633, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 230. In some embodiments, the marker comprises a microsatellite locus at chr15:57247153-57247167, wherein the microsatellite locus comprises a repeat comprising "CCA" starting at nucleobaseWSGR Docket No.56034-703.601 position 101 in SEQ ID NO: 231. In some embodiments, the marker comprises a microsatellite locus at chr15:63279081-63279093, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 232. In some embodiments, the marker comprises a microsatellite locus at chr15:90487409-90487430, wherein the microsatellite locus comprises a repeat comprising "TGC" starting at nucleobase position 101 in SEQ ID NO: 233. In some embodiments, the marker comprises a microsatellite locus at chr16:74893610-74893622, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 234. In some embodiments, the marker comprises a microsatellite locus at chr17:20776957-20776979, wherein the microsatellite locus comprises a repeat comprising "AAC" starting at nucleobase position 101 in SEQ ID NO: 235. In some embodiments, the marker comprises a microsatellite locus at chr17:35834645-35834654, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 236. In some embodiments, the marker comprises a microsatellite locus at chr17:38963112-38963122, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 237. In some embodiments, the marker comprises a microsatellite locus at chr17:42118023-42118034, wherein the microsatellite locus comprises a repeat comprising "GA" starting at nucleobase position 101 in SEQ ID NO: 238. In some embodiments, the marker comprises a microsatellite locus at chr17:50699860-50699882, wherein the microsatellite locus comprises a repeat comprising "TCA" starting at nucleobase position 101 in SEQ ID NO: 239. In some embodiments, the marker comprises a microsatellite locus at chr17:68045756-68045769, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 240. In some embodiments, the marker comprises a microsatellite locus at chr17:68976224-68976233, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 241. In some embodiments, the marker comprises a microsatellite locus at chr17:78802467-78802484, wherein the microsatellite locus comprises a repeat comprising "TTTTTC" starting at nucleobase position 101 in SEQ ID NO: 242. In some embodiments, the marker comprises a microsatellite locus at chr19:30009212- 30009238, wherein the microsatellite locus comprises a repeat comprising "TGA" starting at nucleobase position 101 in SEQ ID NO: 243. In some embodiments, the marker comprises a microsatellite locus at chr19:49808545-49808562, wherein the microsatellite locus comprises a repeat comprising "CCCCTG" starting at nucleobase position 101 in SEQ ID NO: 244. In some embodiments, the marker comprises a microsatellite locus at chr19:58059347-WSGR Docket No.56034-703.601 58059368, wherein the microsatellite locus comprises a repeat comprising "GGCCG" starting at nucleobase position 101 in SEQ ID NO: 245. In some embodiments, the marker comprises a microsatellite locus at chr2:47839094-47839104, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 246. In some embodiments, the marker comprises a microsatellite locus at chr2:62955477-62955486, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 247. In some embodiments, the marker comprises a microsatellite locus at chr2:66568967-66568976, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 248. In some embodiments, the marker comprises a microsatellite locus at chr2:73075764-73075782, wherein the microsatellite locus comprises a repeat comprising "GCCT" starting at nucleobase position 101 in SEQ ID NO: 249. In some embodiments, the marker comprises a microsatellite locus at chr2:88428657-88428669, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 250. In some embodiments, the marker comprises a microsatellite locus at chr2:128013158-128013168, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 251. In some embodiments, the marker comprises a microsatellite locus at chr2:151379532-151379545, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 252. In some embodiments, the marker comprises a microsatellite locus at chr2:163609530- 163609539, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 253. In some embodiments, the marker comprises a microsatellite locus at chr2:208488932-208488941, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 254. In some embodiments, the marker comprises a microsatellite locus at chr2:217848164- 217848192, wherein the microsatellite locus comprises a repeat comprising "GCT" starting at nucleobase position 101 in SEQ ID NO: 255. In some embodiments, the marker comprises a microsatellite locus at chr2:230255541-230255550, wherein the microsatellite locus comprises a repeat comprising "G" starting at nucleobase position 101 in SEQ ID NO: 256. In some embodiments, the marker comprises a microsatellite locus at chr2:236507156- 236507171, wherein the microsatellite locus comprises a repeat comprising "AGAAA" starting at nucleobase position 101 in SEQ ID NO: 257. In some embodiments, the marker comprises a microsatellite locus at chr20:13110087-13110097, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO:WSGR Docket No.56034-703.601 258. In some embodiments, the marker comprises a microsatellite locus at chr20:23492886- 23492897, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 259. In some embodiments, the marker comprises a microsatellite locus at chr20:33714516-33714542, wherein the microsatellite locus comprises a repeat comprising "CAA" starting at nucleobase position 101 in SEQ ID NO: 260. In some embodiments, the marker comprises a microsatellite locus at chr22:46356688-46356707, wherein the microsatellite locus comprises a repeat comprising "CAG" starting at nucleobase position 101 in SEQ ID NO: 261. In some embodiments, the marker comprises a microsatellite locus at chr3:3175293-3175303, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 262. In some embodiments, the marker comprises a microsatellite locus at chr3:9977056-9977069, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 263. In some embodiments, the marker comprises a microsatellite locus at chr3:94035443-94035458, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 264. In some embodiments, the marker comprises a microsatellite locus at chr3:96777407-96777420, wherein the microsatellite locus comprises a repeat comprising "GT" starting at nucleobase position 101 in SEQ ID NO: 265. In some embodiments, the marker comprises a microsatellite locus at chr3:121493431-121493440, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 266. In some embodiments, the marker comprises a microsatellite locus at chr3:179580797-179580808, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 267. In some embodiments, the marker comprises a microsatellite locus at chr3:196488581-196488591, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 268. In some embodiments, the marker comprises a microsatellite locus at chr3:197065486-197065495, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 269. In some embodiments, the marker comprises a microsatellite locus at chr3:197835635-197835651, wherein the microsatellite locus comprises a repeat comprising "GT" starting at nucleobase position 101 in SEQ ID NO: 270. In some embodiments, the marker comprises a microsatellite locus at chr4:15824793-15824802, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 271. In some embodiments, the marker comprises a microsatellite locus at chr4:78870929-78870954, wherein the microsatellite locus comprises a repeat comprisingWSGR Docket No.56034-703.601 "CAG" starting at nucleobase position 101 in SEQ ID NO: 272. In some embodiments, the marker comprises a microsatellite locus at chr4:127891076-127891089, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 273. In some embodiments, the marker comprises a microsatellite locus at chr4:169107210-169107219, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 274. In some embodiments, the marker comprises a microsatellite locus at chr4:183253744-183253754, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 275. In some embodiments, the marker comprises a microsatellite locus at chr5:13884915-13884934, wherein the microsatellite locus comprises a repeat comprising "CA" starting at nucleobase position 101 in SEQ ID NO: 276. In some embodiments, the marker comprises a microsatellite locus at chr5:119473839-119473856, wherein the microsatellite locus comprises a repeat comprising "CA" starting at nucleobase position 101 in SEQ ID NO: 277. In some embodiments, the marker comprises a microsatellite locus at chr5:179723788-179723798, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 278. In some embodiments, the marker comprises a microsatellite locus at chr6:29659154-29659171, wherein the microsatellite locus comprises a repeat comprising "GAA" starting at nucleobase position 101 in SEQ ID NO: 279. In some embodiments, the marker comprises a microsatellite locus at chr6:31701373-31701390, wherein the microsatellite locus comprises a repeat comprising "AC" starting at nucleobase position 101 in SEQ ID NO: 280. In some embodiments, the marker comprises a microsatellite locus at chr6:33398464-33398478, wherein the microsatellite locus comprises a repeat comprising "ATTTT" starting at nucleobase position 101 in SEQ ID NO: 281. In some embodiments, the marker comprises a microsatellite locus at chr6:55874914-55874927, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 282. In some embodiments, the marker comprises a microsatellite locus at chr6:72392821-72392830, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 283. In some embodiments, the marker comprises a microsatellite locus at chr6:80127526-80127535, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 284. In some embodiments, the marker comprises a microsatellite locus at chr6:83188582-83188592, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 285. In some embodiments, the marker comprises a microsatellite locus atWSGR Docket No.56034-703.601 chr6:90008859-90008895, wherein the microsatellite locus comprises a repeat comprising "GAAA" starting at nucleobase position 101 in SEQ ID NO: 286. In some embodiments, the marker comprises a microsatellite locus at chr6:112206947-112206974, wherein the microsatellite locus comprises a repeat comprising "CAAAA" starting at nucleobase position 101 in SEQ ID NO: 287. In some embodiments, the marker comprises a microsatellite locus at chr6:117414554-117414564, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 288. In some embodiments, the marker comprises a microsatellite locus at chr6:152444592-152444602, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 289. In some embodiments, the marker comprises a microsatellite locus at chr7:5294842-5294855, wherein the microsatellite locus comprises a repeat comprising "CT" starting at nucleobase position 101 in SEQ ID NO: 290. In some embodiments, the marker comprises a microsatellite locus at chr7:5730643-5730665, wherein the microsatellite locus comprises a repeat comprising "AAC" starting at nucleobase position 101 in SEQ ID NO: 291. In some embodiments, the marker comprises a microsatellite locus at chr7:7532737-7532749, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 292. In some embodiments, the marker comprises a microsatellite locus at chr7:15686173-15686201, wherein the microsatellite locus comprises a repeat comprising "TGG" starting at nucleobase position 101 in SEQ ID NO: 293. In some embodiments, the marker comprises a microsatellite locus at chr7:54542847-54542858, wherein the microsatellite locus comprises a repeat comprising "GT" starting at nucleobase position 101 in SEQ ID NO: 294. In some embodiments, the marker comprises a microsatellite locus at chr7:75193020-75193032, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 295. In some embodiments, the marker comprises a microsatellite locus at chr7:76021996-76022018, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 296. In some embodiments, the marker comprises a microsatellite locus at chr7:121372143-121372153, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 297. In some embodiments, the marker comprises a microsatellite locus at chr8:35766806-35766826, wherein the microsatellite locus comprises a repeat comprising "TCA" starting at nucleobase position 101 in SEQ ID NO: 298. In some embodiments, the marker comprises a microsatellite locus at chr8:85135766-85135776, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 inWSGR Docket No.56034-703.601 SEQ ID NO: 299. In some embodiments, the marker comprises a microsatellite locus at chr9:8527274-8527290, wherein the microsatellite locus comprises a repeat comprising "TAA" starting at nucleobase position 101 in SEQ ID NO: 300. In some embodiments, the marker comprises a microsatellite locus at chr9:35812102-35812120, wherein the microsatellite locus comprises a repeat comprising "ACCGCG" starting at nucleobase position 101 in SEQ ID NO: 301. In some embodiments, the marker comprises a microsatellite locus at chr9:111412160-111412169, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 302. In some embodiments, the marker comprises a microsatellite locus at chr9:111914767- 111914793, wherein the microsatellite locus comprises a repeat comprising "TTTTG" starting at nucleobase position 101 in SEQ ID NO: 303. In some embodiments, the marker comprises a microsatellite locus at chr9:114002604-114002617, wherein the microsatellite locus comprises a repeat comprising "CT" starting at nucleobase position 101 in SEQ ID NO: 304. In some embodiments, the marker comprises a microsatellite locus at chrX:16670714- 16670723, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 305. In some embodiments, the marker comprises a microsatellite locus at chrX:31929468-31929484, wherein the microsatellite locus comprises a repeat comprising "AAAAC" starting at nucleobase position 101 in SEQ ID NO: 306. In some embodiments, the marker comprises a microsatellite locus at chrX:55001532- 55001541, wherein the microsatellite locus comprises a repeat comprising "T" starting at nucleobase position 101 in SEQ ID NO: 307. In some embodiments, the marker comprises a microsatellite locus at chrX:106866054-106866063, wherein the microsatellite locus comprises a repeat comprising "A" starting at nucleobase position 101 in SEQ ID NO: 308. In some embodiments, the marker comprises a microsatellite locus at chrX:156010134- 156010158, wherein the microsatellite locus comprises a repeat comprising "AGC" starting at nucleobase position 101 in SEQ ID NO: 309. Table 1A: Microsatellite lociWSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601embodiments, the marker comprises a SNP at chr1:1719368-1719368, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 310 comprising anWSGR Docket No.56034-703.601 allele at nucleobase position 101 in SEQ ID NO: 310. In some embodiments, the marker comprises a SNP at chr1:1719406-1719406, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 311 comprising an allele at nucleobase position 101 in SEQ ID NO: 311. In some embodiments, the marker comprises a SNP at chr1:12848032-12848032, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 312 comprising an allele at nucleobase position 101 in SEQ ID NO: 312. In some embodiments, the marker comprises a SNP at chr1:16049737-16049737, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 313 comprising an allele at nucleobase position 101 in SEQ ID NO: 313. In some embodiments, the marker comprises a SNP at chr1:46643156-46643156, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 314 comprising an allele at nucleobase position 101 in SEQ ID NO: 314. In some embodiments, the marker comprises a SNP at chr1:180196564-180196564, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 315 comprising an allele at nucleobase position 101 in SEQ ID NO: 315. In some embodiments, the marker comprises a SNP at chr1:227493872-227493872, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 316 comprising an allele at nucleobase position 101 in SEQ ID NO: 316. In some embodiments, the marker comprises a SNP at chr10:17129266-17129266, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 317 comprising an allele at nucleobase position 101 in SEQ ID NO: 317. In some embodiments, the marker comprises a SNP at chr10:17157420-17157420, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 318 comprising an allele at nucleobase position 101 in SEQ ID NO: 318. In some embodiments, the marker comprises a SNP at chr10:20852540-20852540, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 319 comprising an allele at nucleobase position 101 in SEQ ID NO: 319. In some embodiments, the marker comprises a SNP at chr10:23119210-23119210, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 320 comprising an allele at nucleobase position 101 in SEQ ID NO: 320. In some embodiments, the marker comprises a SNP at chr10:96318244-96318244, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 321 comprising an allele at nucleobase position 101 in SEQ ID NO: 321. In some embodiments, the marker comprises a SNP at chr11:17150586-17150586, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 322 comprising an allele at nucleobase position 101 in SEQ ID NO: 322. In some embodiments, the markerWSGR Docket No.56034-703.601 comprises a SNP at chr11:20407908-20407908, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 323 comprising an allele at nucleobase position 101 in SEQ ID NO: 323. In some embodiments, the marker comprises a SNP at chr11:27506807-27506807, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 324 comprising an allele at nucleobase position 101 in SEQ ID NO: 324. In some embodiments, the marker comprises a SNP at chr11:45904894-45904894, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 325 comprising an allele at nucleobase position 101 in SEQ ID NO: 325. In some embodiments, the marker comprises a SNP at chr11:69057569-69057569, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 326 comprising an allele at nucleobase position 101 in SEQ ID NO: 326. In some embodiments, the marker comprises a SNP at chr11:69083946-69083946, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 327 comprising an allele at nucleobase position 101 in SEQ ID NO: 327. In some embodiments, the marker comprises a SNP at chr11:69085565-69085565, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 328 comprising an allele at nucleobase position 101 in SEQ ID NO: 328. In some embodiments, the marker comprises a SNP at chr11:120228949- 120228949, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 329 comprising an allele at nucleobase position 101 in SEQ ID NO: 329. In some embodiments, the marker comprises a SNP at chr12:12330993-12330993, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 330 comprising an allele at nucleobase position 101 in SEQ ID NO: 330. In some embodiments, the marker comprises a SNP at chr12:19306657-19306657, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 331 comprising an allele at nucleobase position 101 in SEQ ID NO: 331. In some embodiments, the marker comprises a SNP at chr13:50361705-50361705, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 332 comprising an allele at nucleobase position 101 in SEQ ID NO: 332. In some embodiments, the marker comprises a SNP at chr13:89361193-89361193, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 333 comprising an allele at nucleobase position 101 in SEQ ID NO: 333. In some embodiments, the marker comprises a SNP at chr13:89362688-89362688, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 334 comprising an allele at nucleobase position 101 in SEQ ID NO: 334. In some embodiments, the marker comprises a SNP at chr13:89363038-89363038, wherein the SNP comprises at least 10WSGR Docket No.56034-703.601 contiguous nucleic acid molecules of SEQ ID NO: 335 comprising an allele at nucleobase position 101 in SEQ ID NO: 335. In some embodiments, the marker comprises a SNP at chr14:65879728-65879728, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 336 comprising an allele at nucleobase position 101 in SEQ ID NO: 336. In some embodiments, the marker comprises a SNP at chr14:74549779-74549779, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 337 comprising an allele at nucleobase position 101 in SEQ ID NO: 337. In some embodiments, the marker comprises a SNP at chr14:96305622-96305622, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 338 comprising an allele at nucleobase position 101 in SEQ ID NO: 338. In some embodiments, the marker comprises a SNP at chr14:96315575-96315575, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 339 comprising an allele at nucleobase position 101 in SEQ ID NO: 339. In some embodiments, the marker comprises a SNP at chr14:96331387-96331387, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 340 comprising an allele at nucleobase position 101 in SEQ ID NO: 340. In some embodiments, the marker comprises a SNP at chr14:96347366-96347366, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 341 comprising an allele at nucleobase position 101 in SEQ ID NO: 341. In some embodiments, the marker comprises a SNP at chr15:42228631-42228631, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 342 comprising an allele at nucleobase position 101 in SEQ ID NO: 342. In some embodiments, the marker comprises a SNP at chr15:45147712-45147712, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 343 comprising an allele at nucleobase position 101 in SEQ ID NO: 343. In some embodiments, the marker comprises a SNP at chr15:52223481-52223481, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 344 comprising an allele at nucleobase position 101 in SEQ ID NO: 344. In some embodiments, the marker comprises a SNP at chr15:52239744-52239744, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 345 comprising an allele at nucleobase position 101 in SEQ ID NO: 345. In some embodiments, the marker comprises a SNP at chr15:75207278-75207278, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 346 comprising an allele at nucleobase position 101 in SEQ ID NO: 346. In some embodiments, the marker comprises a SNP at chr15:75210806-75210806, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 347 comprising an allele at nucleobaseWSGR Docket No.56034-703.601 position 101 in SEQ ID NO: 347. In some embodiments, the marker comprises a SNP at chr15:78739131-78739131, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 348 comprising an allele at nucleobase position 101 in SEQ ID NO: 348. In some embodiments, the marker comprises a SNP at chr15:78739143-78739143, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 349 comprising an allele at nucleobase position 101 in SEQ ID NO: 349. In some embodiments, the marker comprises a SNP at chr15:80942400-80942400, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 350 comprising an allele at nucleobase position 101 in SEQ ID NO: 350. In some embodiments, the marker comprises a SNP at chr15:89669381-89669381, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 351 comprising an allele at nucleobase position 101 in SEQ ID NO: 351. In some embodiments, the marker comprises a SNP at chr15:89669998-89669998, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 352 comprising an allele at nucleobase position 101 in SEQ ID NO: 352. In some embodiments, the marker comprises a SNP at chr15:100573657- 100573657, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 353 comprising an allele at nucleobase position 101 in SEQ ID NO: 353. In some embodiments, the marker comprises a SNP at chr16:2765236-2765236, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 354 comprising an allele at nucleobase position 101 in SEQ ID NO: 354. In some embodiments, the marker comprises a SNP at chr16:22116731-22116731, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 355 comprising an allele at nucleobase position 101 in SEQ ID NO: 355. In some embodiments, the marker comprises a SNP at chr16:22116732-22116732, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 356 comprising an allele at nucleobase position 101 in SEQ ID NO: 356. In some embodiments, the marker comprises a SNP at chr16:22149815-22149815, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 357 comprising an allele at nucleobase position 101 in SEQ ID NO: 357. In some embodiments, the marker comprises a SNP at chr16:87403158-87403158, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 358 comprising an allele at nucleobase position 101 in SEQ ID NO: 358. In some embodiments, the marker comprises a SNP at chr16:89919746-89919746, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 359 comprising an allele at nucleobase position 101 in SEQ ID NO: 359. In some embodiments, the marker comprises a SNP atWSGR Docket No.56034-703.601 chr16:90095184-90095184, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 360 comprising an allele at nucleobase position 101 in SEQ ID NO: 360. In some embodiments, the marker comprises a SNP at chr17:36104491-36104491, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 361 comprising an allele at nucleobase position 101 in SEQ ID NO: 361. In some embodiments, the marker comprises a SNP at chr17:36104568-36104568, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 362 comprising an allele at nucleobase position 101 in SEQ ID NO: 362. In some embodiments, the marker comprises a SNP at chr19:9854147-9854147, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 363 comprising an allele at nucleobase position 101 in SEQ ID NO: 363. In some embodiments, the marker comprises a SNP at chr19:17219251-17219251, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 364 comprising an allele at nucleobase position 101 in SEQ ID NO: 364. In some embodiments, the marker comprises a SNP at chr19:17403631-17403631, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 365 comprising an allele at nucleobase position 101 in SEQ ID NO: 365. In some embodiments, the marker comprises a SNP at chr19:17423884-17423884, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 366 comprising an allele at nucleobase position 101 in SEQ ID NO: 366. In some embodiments, the marker comprises a SNP at chr19:17424737-17424737, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 367 comprising an allele at nucleobase position 101 in SEQ ID NO: 367. In some embodiments, the marker comprises a SNP at chr19:32953550-32953550, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 368 comprising an allele at nucleobase position 101 in SEQ ID NO: 368. In some embodiments, the marker comprises a SNP at chr19:54102901-54102901, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 369 comprising an allele at nucleobase position 101 in SEQ ID NO: 369. In some embodiments, the marker comprises a SNP at chr2:47803553-47803553, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 370 comprising an allele at nucleobase position 101 in SEQ ID NO: 370. In some embodiments, the marker comprises a SNP at chr2:115739903-115739903, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 371 comprising an allele at nucleobase position 101 in SEQ ID NO: 371. In some embodiments, the marker comprises a SNP at chr2:219540599-219540599, wherein the SNP comprises at least 10 contiguous nucleic acidWSGR Docket No.56034-703.601 molecules of SEQ ID NO: 372 comprising an allele at nucleobase position 101 in SEQ ID NO: 372. In some embodiments, the marker comprises a SNP at chr21:41472033-41472033, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 373 comprising an allele at nucleobase position 101 in SEQ ID NO: 373. In some embodiments, the marker comprises a SNP at chr21:41914787-41914787, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 374 comprising an allele at nucleobase position 101 in SEQ ID NO: 374. In some embodiments, the marker comprises a SNP at chr21:43687637-43687637, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 375 comprising an allele at nucleobase position 101 in SEQ ID NO: 375. In some embodiments, the marker comprises a SNP at chr21:44600805-44600805, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 376 comprising an allele at nucleobase position 101 in SEQ ID NO: 376. In some embodiments, the marker comprises a SNP at chr21:44600806-44600806, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 377 comprising an allele at nucleobase position 101 in SEQ ID NO: 377. In some embodiments, the marker comprises a SNP at chr21:44601173-44601173, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 378 comprising an allele at nucleobase position 101 in SEQ ID NO: 378. In some embodiments, the marker comprises a SNP at chr3:50646378-50646378, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 379 comprising an allele at nucleobase position 101 in SEQ ID NO: 379. In some embodiments, the marker comprises a SNP at chr4:5576157-5576157, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 380 comprising an allele at nucleobase position 101 in SEQ ID NO: 380. In some embodiments, the marker comprises a SNP at chr4:80388302-80388302, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 381 comprising an allele at nucleobase position 101 in SEQ ID NO: 381. In some embodiments, the marker comprises a SNP at chr5:1409012-1409012, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 382 comprising an allele at nucleobase position 101 in SEQ ID NO: 382. In some embodiments, the marker comprises a SNP at chr5:31532322-31532322, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 383 comprising an allele at nucleobase position 101 in SEQ ID NO: 383. In some embodiments, the marker comprises a SNP at chr5:33951588-33951588, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 384 comprising an allele at nucleobase position 101 in SEQ IDWSGR Docket No.56034-703.601 NO: 384. In some embodiments, the marker comprises a SNP at chr5:36201234-36201234, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 385 comprising an allele at nucleobase position 101 in SEQ ID NO: 385. In some embodiments, the marker comprises a SNP at chr5:39389959-39389959, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 386 comprising an allele at nucleobase position 101 in SEQ ID NO: 386. In some embodiments, the marker comprises a SNP at chr5:59193558-59193558, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 387 comprising an allele at nucleobase position 101 in SEQ ID NO: 387. In some embodiments, the marker comprises a SNP at chr5:76832696-76832696, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 388 comprising an allele at nucleobase position 101 in SEQ ID NO: 388. In some embodiments, the marker comprises a SNP at chr5:141157782-141157782, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 389 comprising an allele at nucleobase position 101 in SEQ ID NO: 389. In some embodiments, the marker comprises a SNP at chr5:143041850-143041850, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 390 comprising an allele at nucleobase position 101 in SEQ ID NO: 390. In some embodiments, the marker comprises a SNP at chr6:33068611-33068611, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 391 comprising an allele at nucleobase position 101 in SEQ ID NO: 391. In some embodiments, the marker comprises a SNP at chr6:33068617-33068617, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 392 comprising an allele at nucleobase position 101 in SEQ ID NO: 392. In some embodiments, the marker comprises a SNP at chr6:33068624-33068624, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 393 comprising an allele at nucleobase position 101 in SEQ ID NO: 393. In some embodiments, the marker comprises a SNP at chr7:73297616-73297616, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 394 comprising an allele at nucleobase position 101 in SEQ ID NO: 394. In some embodiments, the marker comprises a SNP at chr7:75501512-75501512, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 395 comprising an allele at nucleobase position 101 in SEQ ID NO: 395. In some embodiments, the marker comprises a SNP at chr7:75501861-75501861, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 396 comprising an allele at nucleobase position 101 in SEQ ID NO: 396. In some embodiments, the marker comprises a SNP at chr7:75501927-75501927,WSGR Docket No.56034-703.601 wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 397 comprising an allele at nucleobase position 101 in SEQ ID NO: 397. In some embodiments, the marker comprises a SNP at chr7:75516085-75516085, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 398 comprising an allele at nucleobase position 101 in SEQ ID NO: 398. In some embodiments, the marker comprises a SNP at chr8:143992718-143992718, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 399 comprising an allele at nucleobase position 101 in SEQ ID NO: 399. In some embodiments, the marker comprises a SNP at chr8:144440133-144440133, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 400 comprising an allele at nucleobase position 101 in SEQ ID NO: 400. In some embodiments, the marker comprises a SNP at chr8:144440254-144440254, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 401 comprising an allele at nucleobase position 101 in SEQ ID NO: 401. In some embodiments, the marker comprises a SNP at chr8:144442022-144442022, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 402 comprising an allele at nucleobase position 101 in SEQ ID NO: 402. In some embodiments, the marker comprises a SNP at chr8:144464532-144464532, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 403 comprising an allele at nucleobase position 101 in SEQ ID NO: 403. In some embodiments, the marker comprises a SNP at chr8:144464925-144464925, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 404 comprising an allele at nucleobase position 101 in SEQ ID NO: 404. In some embodiments, the marker comprises a SNP at chr8:144467429-144467429, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 405 comprising an allele at nucleobase position 101 in SEQ ID NO: 405. In some embodiments, the marker comprises a SNP at chr8:144468337-144468337, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 406 comprising an allele at nucleobase position 101 in SEQ ID NO: 406. In some embodiments, the marker comprises a SNP at chr8:144471861-144471861, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 407 comprising an allele at nucleobase position 101 in SEQ ID NO: 407. In some embodiments, the marker comprises a SNP at chr8:144512253-144512253, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 408 comprising an allele at nucleobase position 101 in SEQ ID NO: 408. In some embodiments, the marker comprises a SNP at chr8:144517130-144517130, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO:WSGR Docket No.56034-703.601 409 comprising an allele at nucleobase position 101 in SEQ ID NO: 409. In some embodiments, the marker comprises a SNP at chr8:144519798-144519798, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 410 comprising an allele at nucleobase position 101 in SEQ ID NO: 410. In some embodiments, the marker comprises a SNP at chr9:89405518-89405518, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 411 comprising an allele at nucleobase position 101 in SEQ ID NO: 411. In some embodiments, the marker comprises a SNP at chr9:89450019-89450019, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 412 comprising an allele at nucleobase position 101 in SEQ ID NO: 412. In some embodiments, the marker comprises a SNP at chr9:89450169-89450169, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 413 comprising an allele at nucleobase position 101 in SEQ ID NO: 413. In some embodiments, the marker comprises a SNP at chr9:124857643-124857643, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 414 comprising an allele at nucleobase position 101 in SEQ ID NO: 414. In some embodiments, the marker comprises a SNP at chr1:16039651-16039651, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 415 comprising an allele at nucleobase position 101 in SEQ ID NO: 415. In some embodiments, the marker comprises a SNP at chr1:21981045-21981045, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 416 comprising an allele at nucleobase position 101 in SEQ ID NO: 416. In some embodiments, the marker comprises a SNP at chr1:21983769-21983769, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 417 comprising an allele at nucleobase position 101 in SEQ ID NO: 417. In some embodiments, the marker comprises a SNP at chr1:21984331-21984331, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 418 comprising an allele at nucleobase position 101 in SEQ ID NO: 418. In some embodiments, the marker comprises a SNP at chr1:35095909-35095909, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 419 comprising an allele at nucleobase position 101 in SEQ ID NO: 419. In some embodiments, the marker comprises a SNP at chr1:35751234-35751234, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 420 comprising an allele at nucleobase position 101 in SEQ ID NO: 420. In some embodiments, the marker comprises a SNP at chr1:35850970-35850970, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 421 comprising an allele at nucleobase position 101 in SEQ ID NO: 421. In someWSGR Docket No.56034-703.601 embodiments, the marker comprises a SNP at chr1:62473963-62473963, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 422 comprising an allele at nucleobase position 101 in SEQ ID NO: 422. In some embodiments, the marker comprises a SNP at chr1:98760071-98760071, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 423 comprising an allele at nucleobase position 101 in SEQ ID NO: 423. In some embodiments, the marker comprises a SNP at chr1:156905300-156905300, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 424 comprising an allele at nucleobase position 101 in SEQ ID NO: 424. In some embodiments, the marker comprises a SNP at chr1:156905315-156905315, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 425 comprising an allele at nucleobase position 101 in SEQ ID NO: 425. In some embodiments, the marker comprises a SNP at chr10:15531051-15531051, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 426 comprising an allele at nucleobase position 101 in SEQ ID NO: 426. In some embodiments, the marker comprises a SNP at chr10:87365105-87365105, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 427 comprising an allele at nucleobase position 101 in SEQ ID NO: 427. In some embodiments, the marker comprises a SNP at chr10:106706654-106706654, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 428 comprising an allele at nucleobase position 101 in SEQ ID NO: 428. In some embodiments, the marker comprises a SNP at chr11:60306124-60306124, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 429 comprising an allele at nucleobase position 101 in SEQ ID NO: 429. In some embodiments, the marker comprises a SNP at chr11:64804546-64804546, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 430 comprising an allele at nucleobase position 101 in SEQ ID NO: 430. In some embodiments, the marker comprises a SNP at chr11:66020601-66020601, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 431 comprising an allele at nucleobase position 101 in SEQ ID NO: 431. In some embodiments, the marker comprises a SNP at chr11:75151346-75151346, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 432 comprising an allele at nucleobase position 101 in SEQ ID NO: 432. In some embodiments, the marker comprises a SNP at chr11:94070551-94070551, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 433 comprising an allele at nucleobase position 101 in SEQ ID NO: 433. In some embodiments, the marker comprises a SNP at chr11:125961075-125961075, wherein theWSGR Docket No.56034-703.601 SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 434 comprising an allele at nucleobase position 101 in SEQ ID NO: 434. In some embodiments, the marker comprises a SNP at chr12:3018486-3018486, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 435 comprising an allele at nucleobase position 101 in SEQ ID NO: 435. In some embodiments, the marker comprises a SNP at chr12:21638338-21638338, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 436 comprising an allele at nucleobase position 101 in SEQ ID NO: 436. In some embodiments, the marker comprises a SNP at chr12:21643880-21643880, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 437 comprising an allele at nucleobase position 101 in SEQ ID NO: 437. In some embodiments, the marker comprises a SNP at chr12:21647065-21647065, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 438 comprising an allele at nucleobase position 101 in SEQ ID NO: 438. In some embodiments, the marker comprises a SNP at chr12:21654498-21654498, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 439 comprising an allele at nucleobase position 101 in SEQ ID NO: 439. In some embodiments, the marker comprises a SNP at chr12:21654706-21654706, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 440 comprising an allele at nucleobase position 101 in SEQ ID NO: 440. In some embodiments, the marker comprises a SNP at chr12:45301952-45301952, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 441 comprising an allele at nucleobase position 101 in SEQ ID NO: 441. In some embodiments, the marker comprises a SNP at chr12:66392311-66392311, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 442 comprising an allele at nucleobase position 101 in SEQ ID NO: 442. In some embodiments, the marker comprises a SNP at chr12:66392864-66392864, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 443 comprising an allele at nucleobase position 101 in SEQ ID NO: 443. In some embodiments, the marker comprises a SNP at chr12:124374499-124374499, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 444 comprising an allele at nucleobase position 101 in SEQ ID NO: 444. In some embodiments, the marker comprises a SNP at chr12:132140773- 132140773, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 445 comprising an allele at nucleobase position 101 in SEQ ID NO: 445. In some embodiments, the marker comprises a SNP at chr12:132147600-132147600, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 446 comprisingWSGR Docket No.56034-703.601 an allele at nucleobase position 101 in SEQ ID NO: 446. In some embodiments, the marker comprises a SNP at chr13:19395372-19395372, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 447 comprising an allele at nucleobase position 101 in SEQ ID NO: 447. In some embodiments, the marker comprises a SNP at chr13:20949602-20949602, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 448 comprising an allele at nucleobase position 101 in SEQ ID NO: 448. In some embodiments, the marker comprises a SNP at chr13:24883069-24883069, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 449 comprising an allele at nucleobase position 101 in SEQ ID NO: 449. In some embodiments, the marker comprises a SNP at chr14:21154851-21154851, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 450 comprising an allele at nucleobase position 101 in SEQ ID NO: 450. In some embodiments, the marker comprises a SNP at chr14:31066397-31066397, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 451 comprising an allele at nucleobase position 101 in SEQ ID NO: 451. In some embodiments, the marker comprises a SNP at chr14:60780471-60780471, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 452 comprising an allele at nucleobase position 101 in SEQ ID NO: 452. In some embodiments, the marker comprises a SNP at chr14:60984871-60984871, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 453 comprising an allele at nucleobase position 101 in SEQ ID NO: 453. In some embodiments, the marker comprises a SNP at chr14:61043206-61043206, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 454 comprising an allele at nucleobase position 101 in SEQ ID NO: 454. In some embodiments, the marker comprises a SNP at chr15:39992919-39992919, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 455 comprising an allele at nucleobase position 101 in SEQ ID NO: 455. In some embodiments, the marker comprises a SNP at chr15:42226938-42226938, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 456 comprising an allele at nucleobase position 101 in SEQ ID NO: 456. In some embodiments, the marker comprises a SNP at chr15:42239570-42239570, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 457 comprising an allele at nucleobase position 101 in SEQ ID NO: 457. In some embodiments, the marker comprises a SNP at chr15:42272172-42272172, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 458 comprising an allele at nucleobase position 101 in SEQ ID NO: 458. In some embodiments, the markerWSGR Docket No.56034-703.601 comprises a SNP at chr15:42276349-42276349, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 459 comprising an allele at nucleobase position 101 in SEQ ID NO: 459. In some embodiments, the marker comprises a SNP at chr15:42276432-42276432, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 460 comprising an allele at nucleobase position 101 in SEQ ID NO: 460. In some embodiments, the marker comprises a SNP at chr15:42287624-42287624, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 461 comprising an allele at nucleobase position 101 in SEQ ID NO: 461. In some embodiments, the marker comprises a SNP at chr15:42292711-42292711, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 462 comprising an allele at nucleobase position 101 in SEQ ID NO: 462. In some embodiments, the marker comprises a SNP at chr15:52247476-52247476, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 463 comprising an allele at nucleobase position 101 in SEQ ID NO: 463. In some embodiments, the marker comprises a SNP at chr15:52389376-52389376, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 464 comprising an allele at nucleobase position 101 in SEQ ID NO: 464. In some embodiments, the marker comprises a SNP at chr15:52397329-52397329, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 465 comprising an allele at nucleobase position 101 in SEQ ID NO: 465. In some embodiments, the marker comprises a SNP at chr15:62675143-62675143, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 466 comprising an allele at nucleobase position 101 in SEQ ID NO: 466. In some embodiments, the marker comprises a SNP at chr15:62675355-62675355, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 467 comprising an allele at nucleobase position 101 in SEQ ID NO: 467. In some embodiments, the marker comprises a SNP at chr15:62697576-62697576, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 468 comprising an allele at nucleobase position 101 in SEQ ID NO: 468. In some embodiments, the marker comprises a SNP at chr15:99729591-99729591, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 469 comprising an allele at nucleobase position 101 in SEQ ID NO: 469. In some embodiments, the marker comprises a SNP at chr16:90094626-90094626, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 470 comprising an allele at nucleobase position 101 in SEQ ID NO: 470. In some embodiments, the marker comprises a SNP at chr17:6589736-6589736, wherein the SNP comprises at least 10WSGR Docket No.56034-703.601 contiguous nucleic acid molecules of SEQ ID NO: 471 comprising an allele at nucleobase position 101 in SEQ ID NO: 471. In some embodiments, the marker comprises a SNP at chr17:10860688-10860688, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 472 comprising an allele at nucleobase position 101 in SEQ ID NO: 472. In some embodiments, the marker comprises a SNP at chr17:36070917-36070917, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 473 comprising an allele at nucleobase position 101 in SEQ ID NO: 473. In some embodiments, the marker comprises a SNP at chr17:75569181-75569181, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 474 comprising an allele at nucleobase position 101 in SEQ ID NO: 474. In some embodiments, the marker comprises a SNP at chr17:78974797-78974797, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 475 comprising an allele at nucleobase position 101 in SEQ ID NO: 475. In some embodiments, the marker comprises a SNP at chr18:23915519-23915519, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 476 comprising an allele at nucleobase position 101 in SEQ ID NO: 476. In some embodiments, the marker comprises a SNP at chr18:45866769-45866769, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 477 comprising an allele at nucleobase position 101 in SEQ ID NO: 477. In some embodiments, the marker comprises a SNP at chr19:372551-372551, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 478 comprising an allele at nucleobase position 101 in SEQ ID NO: 478. In some embodiments, the marker comprises a SNP at chr19:7741837-7741837, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 479 comprising an allele at nucleobase position 101 in SEQ ID NO: 479. In some embodiments, the marker comprises a SNP at chr19:44392323-44392323, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 480 comprising an allele at nucleobase position 101 in SEQ ID NO: 480. In some embodiments, the marker comprises a SNP at chr19:53811331-53811331, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 481 comprising an allele at nucleobase position 101 in SEQ ID NO: 481. In some embodiments, the marker comprises a SNP at chr2:143029705-143029705, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 482 comprising an allele at nucleobase position 101 in SEQ ID NO: 482. In some embodiments, the marker comprises a SNP at chr2:185805377-185805377, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 483 comprising an allele at nucleobaseWSGR Docket No.56034-703.601 position 101 in SEQ ID NO: 483. In some embodiments, the marker comprises a SNP at chr2:185833149-185833149, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 484 comprising an allele at nucleobase position 101 in SEQ ID NO: 484. In some embodiments, the marker comprises a SNP at chr2:186747021-186747021, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 485 comprising an allele at nucleobase position 101 in SEQ ID NO: 485. In some embodiments, the marker comprises a SNP at chr2:219635701-219635701, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 486 comprising an allele at nucleobase position 101 in SEQ ID NO: 486. In some embodiments, the marker comprises a SNP at chr20:2075712-2075712, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 487 comprising an allele at nucleobase position 101 in SEQ ID NO: 487. In some embodiments, the marker comprises a SNP at chr3:50640570-50640570, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 488 comprising an allele at nucleobase position 101 in SEQ ID NO: 488. In some embodiments, the marker comprises a SNP at chr3:50787815-50787815, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 489 comprising an allele at nucleobase position 101 in SEQ ID NO: 489. In some embodiments, the marker comprises a SNP at chr3:69176437-69176437, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 490 comprising an allele at nucleobase position 101 in SEQ ID NO: 490. In some embodiments, the marker comprises a SNP at chr3:69180910-69180910, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 491 comprising an allele at nucleobase position 101 in SEQ ID NO: 491. In some embodiments, the marker comprises a SNP at chr3:69181650-69181650, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 492 comprising an allele at nucleobase position 101 in SEQ ID NO: 492. In some embodiments, the marker comprises a SNP at chr3:69195011-69195011, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 493 comprising an allele at nucleobase position 101 in SEQ ID NO: 493. In some embodiments, the marker comprises a SNP at chr3:108218439-108218439, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 494 comprising an allele at nucleobase position 101 in SEQ ID NO: 494. In some embodiments, the marker comprises a SNP at chr3:130103796-130103796, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 495 comprising an allele at nucleobase position 101 in SEQ ID NO: 495. In some embodiments, the marker comprises a SNP atWSGR Docket No.56034-703.601 chr3:187236497-187236497, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 496 comprising an allele at nucleobase position 101 in SEQ ID NO: 496. In some embodiments, the marker comprises a SNP at chr4:44682438-44682438, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 497 comprising an allele at nucleobase position 101 in SEQ ID NO: 497. In some embodiments, the marker comprises a SNP at chr4:80388596-80388596, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 498 comprising an allele at nucleobase position 101 in SEQ ID NO: 498. In some embodiments, the marker comprises a SNP at chr4:86913609-86913609, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 499 comprising an allele at nucleobase position 101 in SEQ ID NO: 499. In some embodiments, the marker comprises a SNP at chr4:189955040-189955040, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 500 comprising an allele at nucleobase position 101 in SEQ ID NO: 500. In some embodiments, the marker comprises a SNP at chr5:35033500-35033500, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 501 comprising an allele at nucleobase position 101 in SEQ ID NO: 501. In some embodiments, the marker comprises a SNP at chr5:35037010-35037010, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 502 comprising an allele at nucleobase position 101 in SEQ ID NO: 502. In some embodiments, the marker comprises a SNP at chr5:36207322-36207322, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 503 comprising an allele at nucleobase position 101 in SEQ ID NO: 503. In some embodiments, the marker comprises a SNP at chr5:77784916-77784916, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 504 comprising an allele at nucleobase position 101 in SEQ ID NO: 504. In some embodiments, the marker comprises a SNP at chr5:77785784-77785784, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 505 comprising an allele at nucleobase position 101 in SEQ ID NO: 505. In some embodiments, the marker comprises a SNP at chr5:147591690-147591690, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 506 comprising an allele at nucleobase position 101 in SEQ ID NO: 506. In some embodiments, the marker comprises a SNP at chr5:160424204-160424204, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 507 comprising an allele at nucleobase position 101 in SEQ ID NO: 507. In some embodiments, the marker comprises a SNP at chr5:180614001-180614001, wherein the SNP comprises at least 10 contiguous nucleic acidWSGR Docket No.56034-703.601 molecules of SEQ ID NO: 508 comprising an allele at nucleobase position 101 in SEQ ID NO: 508. In some embodiments, the marker comprises a SNP at chr6:10955175-10955175, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 509 comprising an allele at nucleobase position 101 in SEQ ID NO: 509. In some embodiments, the marker comprises a SNP at chr6:17629385-17629385, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 510 comprising an allele at nucleobase position 101 in SEQ ID NO: 510. In some embodiments, the marker comprises a SNP at chr6:100050817-100050817, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 511 comprising an allele at nucleobase position 101 in SEQ ID NO: 511. In some embodiments, the marker comprises a SNP at chr7:16864739-16864739, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 512 comprising an allele at nucleobase position 101 in SEQ ID NO: 512. In some embodiments, the marker comprises a SNP at chr7:21705443-21705443, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 513 comprising an allele at nucleobase position 101 in SEQ ID NO: 513. In some embodiments, the marker comprises a SNP at chr7:75992343-75992343, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 514 comprising an allele at nucleobase position 101 in SEQ ID NO: 514. In some embodiments, the marker comprises a SNP at chr7:85381348-85381348, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 515 comprising an allele at nucleobase position 101 in SEQ ID NO: 515. In some embodiments, the marker comprises a SNP at chr7:102382706-102382706, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 516 comprising an allele at nucleobase position 101 in SEQ ID NO: 516. In some embodiments, the marker comprises a SNP at chr7:149054118-149054118, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 517 comprising an allele at nucleobase position 101 in SEQ ID NO: 517. In some embodiments, the marker comprises a SNP at chr8:39157613-39157613, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 518 comprising an allele at nucleobase position 101 in SEQ ID NO: 518. In some embodiments, the marker comprises a SNP at chr8:101492746-101492746, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 519 comprising an allele at nucleobase position 101 in SEQ ID NO: 519. In some embodiments, the marker comprises a SNP at chr8:138630555-138630555, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 520 comprising an allele at nucleobase position 101 in SEQ IDWSGR Docket No.56034-703.601 NO: 520. In some embodiments, the marker comprises a SNP at chr8:138679683-138679683, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 521 comprising an allele at nucleobase position 101 in SEQ ID NO: 521. In some embodiments, the marker comprises a SNP at chr8:144467002-144467002, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 522 comprising an allele at nucleobase position 101 in SEQ ID NO: 522. In some embodiments, the marker comprises a SNP at chr9:738434-738434, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 523 comprising an allele at nucleobase position 101 in SEQ ID NO: 523. In some embodiments, the marker comprises a SNP at chr9:15479637- 15479637, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 524 comprising an allele at nucleobase position 101 in SEQ ID NO: 524. In some embodiments, the marker comprises a SNP at chr9:17464404-17464404, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 525 comprising an allele at nucleobase position 101 in SEQ ID NO: 525. In some embodiments, the marker comprises a SNP at chr9:35546738-35546738, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 526 comprising an allele at nucleobase position 101 in SEQ ID NO: 526. In some embodiments, the marker comprises a SNP at chr9:124856497-124856497, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 527 comprising an allele at nucleobase position 101 in SEQ ID NO: 527. In some embodiments, the marker comprises a SNP at chr9:124874936-124874936, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 528 comprising an allele at nucleobase position 101 in SEQ ID NO: 528. In some embodiments, the marker comprises a SNP at chr9:127953643-127953643, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 529 comprising an allele at nucleobase position 101 in SEQ ID NO: 529. In some embodiments, the marker comprises a SNP at chr9:129634376-129634376, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 530 comprising an allele at nucleobase position 101 in SEQ ID NO: 530. In some embodiments, the marker comprises a SNP at chr9:129640629-129640629, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 531 comprising an allele at nucleobase position 101 in SEQ ID NO: 531. In some embodiments, the marker comprises a SNP at chr9:136502238-136502238, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 532 comprising an allele at nucleobase position 101 in SEQ ID NO: 532. In some embodiments, the marker comprises a SNP at chr1:1313807-1313807, wherein the SNPWSGR Docket No.56034-703.601 comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 533 comprising an allele at nucleobase position 101 in SEQ ID NO: 533. In some embodiments, the marker comprises a SNP at chr1:35739565-35739565, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 534 comprising an allele at nucleobase position 101 in SEQ ID NO: 534. In some embodiments, the marker comprises a SNP at chr1:47142179-47142179, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 535 comprising an allele at nucleobase position 101 in SEQ ID NO: 535. In some embodiments, the marker comprises a SNP at chr1:74748706-74748706, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 536 comprising an allele at nucleobase position 101 in SEQ ID NO: 536. In some embodiments, the marker comprises a SNP at chr1:156879203-156879203, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 537 comprising an allele at nucleobase position 101 in SEQ ID NO: 537. In some embodiments, the marker comprises a SNP at chr10:20147909-20147909, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 538 comprising an allele at nucleobase position 101 in SEQ ID NO: 538. In some embodiments, the marker comprises a SNP at chr10:72340884-72340884, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 539 comprising an allele at nucleobase position 101 in SEQ ID NO: 539. In some embodiments, the marker comprises a SNP at chr10:93283990-93283990, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 540 comprising an allele at nucleobase position 101 in SEQ ID NO: 540. In some embodiments, the marker comprises a SNP at chr10:93284010-93284010, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 541 comprising an allele at nucleobase position 101 in SEQ ID NO: 541. In some embodiments, the marker comprises a SNP at chr10:93284012-93284012, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 542 comprising an allele at nucleobase position 101 in SEQ ID NO: 542. In some embodiments, the marker comprises a SNP at chr10:93284086-93284086, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 543 comprising an allele at nucleobase position 101 in SEQ ID NO: 543. In some embodiments, the marker comprises a SNP at chr11:45927983-45927983, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 544 comprising an allele at nucleobase position 101 in SEQ ID NO: 544. In some embodiments, the marker comprises a SNP at chr11:120226789-120226789, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 545 comprisingWSGR Docket No.56034-703.601 an allele at nucleobase position 101 in SEQ ID NO: 545. In some embodiments, the marker comprises a SNP at chr12:52547745-52547745, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 546 comprising an allele at nucleobase position 101 in SEQ ID NO: 546. In some embodiments, the marker comprises a SNP at chr12:111446804-111446804, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 547 comprising an allele at nucleobase position 101 in SEQ ID NO: 547. In some embodiments, the marker comprises a SNP at chr12:121957015- 121957015, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 548 comprising an allele at nucleobase position 101 in SEQ ID NO: 548. In some embodiments, the marker comprises a SNP at chr15:27983498-27983498, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 549 comprising an allele at nucleobase position 101 in SEQ ID NO: 549. In some embodiments, the marker comprises a SNP at chr15:28111713-28111713, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 550 comprising an allele at nucleobase position 101 in SEQ ID NO: 550. In some embodiments, the marker comprises a SNP at chr15:45153301-45153301, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 551 comprising an allele at nucleobase position 101 in SEQ ID NO: 551. In some embodiments, the marker comprises a SNP at chr15:45253280-45253280, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 552 comprising an allele at nucleobase position 101 in SEQ ID NO: 552. In some embodiments, the marker comprises a SNP at chr15:45262069-45262069, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 553 comprising an allele at nucleobase position 101 in SEQ ID NO: 553. In some embodiments, the marker comprises a SNP at chr15:45265258-45265258, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 554 comprising an allele at nucleobase position 101 in SEQ ID NO: 554. In some embodiments, the marker comprises a SNP at chr15:52242147-52242147, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 555 comprising an allele at nucleobase position 101 in SEQ ID NO: 555. In some embodiments, the marker comprises a SNP at chr16:653721-653721, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 556 comprising an allele at nucleobase position 101 in SEQ ID NO: 556. In some embodiments, the marker comprises a SNP at chr16:655360-655360, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 557 comprising an allele at nucleobase position 101 in SEQ ID NO: 557. In some embodiments, the markerWSGR Docket No.56034-703.601 comprises a SNP at chr16:658419-658419, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 558 comprising an allele at nucleobase position 101 in SEQ ID NO: 558. In some embodiments, the marker comprises a SNP at chr16:659001- 659001, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 559 comprising an allele at nucleobase position 101 in SEQ ID NO: 559. In some embodiments, the marker comprises a SNP at chr16:659157-659157, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 560 comprising an allele at nucleobase position 101 in SEQ ID NO: 560. In some embodiments, the marker comprises a SNP at chr16:661326-661326, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 561 comprising an allele at nucleobase position 101 in SEQ ID NO: 561. In some embodiments, the marker comprises a SNP at chr16:661582- 661582, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 562 comprising an allele at nucleobase position 101 in SEQ ID NO: 562. In some embodiments, the marker comprises a SNP at chr16:661712-661712, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 563 comprising an allele at nucleobase position 101 in SEQ ID NO: 563. In some embodiments, the marker comprises a SNP at chr16:661905-661905, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 564 comprising an allele at nucleobase position 101 in SEQ ID NO: 564. In some embodiments, the marker comprises a SNP at chr16:665847- 665847, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 565 comprising an allele at nucleobase position 101 in SEQ ID NO: 565. In some embodiments, the marker comprises a SNP at chr16:665990-665990, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 566 comprising an allele at nucleobase position 101 in SEQ ID NO: 566. In some embodiments, the marker comprises a SNP at chr16:666428-666428, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 567 comprising an allele at nucleobase position 101 in SEQ ID NO: 567. In some embodiments, the marker comprises a SNP at chr16:667523- 667523, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 568 comprising an allele at nucleobase position 101 in SEQ ID NO: 568. In some embodiments, the marker comprises a SNP at chr16:672548-672548, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 569 comprising an allele at nucleobase position 101 in SEQ ID NO: 569. In some embodiments, the marker comprises a SNP at chr16:673647-673647, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 570 comprising an allele at nucleobase position 101WSGR Docket No.56034-703.601 in SEQ ID NO: 570. In some embodiments, the marker comprises a SNP at chr16:20341434- 20341434, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 571 comprising an allele at nucleobase position 101 in SEQ ID NO: 571. In some embodiments, the marker comprises a SNP at chr16:89769915-89769915, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 572 comprising an allele at nucleobase position 101 in SEQ ID NO: 572. In some embodiments, the marker comprises a SNP at chr16:89792117-89792117, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 573 comprising an allele at nucleobase position 101 in SEQ ID NO: 573. In some embodiments, the marker comprises a SNP at chr17:7830515-7830515, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 574 comprising an allele at nucleobase position 101 in SEQ ID NO: 574. In some embodiments, the marker comprises a SNP at chr17:9643234-9643234, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 575 comprising an allele at nucleobase position 101 in SEQ ID NO: 575. In some embodiments, the marker comprises a SNP at chr17:19310022-19310022, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 576 comprising an allele at nucleobase position 101 in SEQ ID NO: 576. In some embodiments, the marker comprises a SNP at chr17:19343762-19343762, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 577 comprising an allele at nucleobase position 101 in SEQ ID NO: 577. In some embodiments, the marker comprises a SNP at chr17:36424564-36424564, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 578 comprising an allele at nucleobase position 101 in SEQ ID NO: 578. In some embodiments, the marker comprises a SNP at chr17:36432760-36432760, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 579 comprising an allele at nucleobase position 101 in SEQ ID NO: 579. In some embodiments, the marker comprises a SNP at chr19:35228354-35228354, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 580 comprising an allele at nucleobase position 101 in SEQ ID NO: 580. In some embodiments, the marker comprises a SNP at chr19:39886005-39886005, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 581 comprising an allele at nucleobase position 101 in SEQ ID NO: 581. In some embodiments, the marker comprises a SNP at chr19:43269447-43269447, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 582 comprising an allele at nucleobase position 101 in SEQ ID NO: 582. In some embodiments, the marker comprises a SNP at chr19:43269474-43269474,WSGR Docket No.56034-703.601 wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 583 comprising an allele at nucleobase position 101 in SEQ ID NO: 583. In some embodiments, the marker comprises a SNP at chr2:159217384-159217384, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 584 comprising an allele at nucleobase position 101 in SEQ ID NO: 584. In some embodiments, the marker comprises a SNP at chr2:178780128-178780128, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 585 comprising an allele at nucleobase position 101 in SEQ ID NO: 585. In some embodiments, the marker comprises a SNP at chr20:32786714-32786714, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 586 comprising an allele at nucleobase position 101 in SEQ ID NO: 586. In some embodiments, the marker comprises a SNP at chr20:32798541-32798541, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 587 comprising an allele at nucleobase position 101 in SEQ ID NO: 587. In some embodiments, the marker comprises a SNP at chr20:32798643-32798643, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 588 comprising an allele at nucleobase position 101 in SEQ ID NO: 588. In some embodiments, the marker comprises a SNP at chr20:45685970-45685970, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 589 comprising an allele at nucleobase position 101 in SEQ ID NO: 589. In some embodiments, the marker comprises a SNP at chr20:46012320-46012320, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 590 comprising an allele at nucleobase position 101 in SEQ ID NO: 590. In some embodiments, the marker comprises a SNP at chr20:46034836-46034836, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 591 comprising an allele at nucleobase position 101 in SEQ ID NO: 591. In some embodiments, the marker comprises a SNP at chr21:41452806-41452806, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 592 comprising an allele at nucleobase position 101 in SEQ ID NO: 592. In some embodiments, the marker comprises a SNP at chr3:58109125-58109125, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 593 comprising an allele at nucleobase position 101 in SEQ ID NO: 593. In some embodiments, the marker comprises a SNP at chr3:58109739-58109739, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 594 comprising an allele at nucleobase position 101 in SEQ ID NO: 594. In some embodiments, the marker comprises a SNP at chr3:58109773-58109773, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO:WSGR Docket No.56034-703.601 595 comprising an allele at nucleobase position 101 in SEQ ID NO: 595. In some embodiments, the marker comprises a SNP at chr3:58126761-58126761, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 596 comprising an allele at nucleobase position 101 in SEQ ID NO: 596. In some embodiments, the marker comprises a SNP at chr3:108453970-108453970, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 597 comprising an allele at nucleobase position 101 in SEQ ID NO: 597. In some embodiments, the marker comprises a SNP at chr3:123692742-123692742, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 598 comprising an allele at nucleobase position 101 in SEQ ID NO: 598. In some embodiments, the marker comprises a SNP at chr3:140562881-140562881, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 599 comprising an allele at nucleobase position 101 in SEQ ID NO: 599. In some embodiments, the marker comprises a SNP at chr4:38797027-38797027, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 600 comprising an allele at nucleobase position 101 in SEQ ID NO: 600. In some embodiments, the marker comprises a SNP at chr6:144827554-144827554, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 601 comprising an allele at nucleobase position 101 in SEQ ID NO: 601. In some embodiments, the marker comprises a SNP at chr6:144840865-144840865, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 602 comprising an allele at nucleobase position 101 in SEQ ID NO: 602. In some embodiments, the marker comprises a SNP at chr7:24718946-24718946, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 603 comprising an allele at nucleobase position 101 in SEQ ID NO: 603. In some embodiments, the marker comprises a SNP at chr7:24719026-24719026, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 604 comprising an allele at nucleobase position 101 in SEQ ID NO: 604. In some embodiments, the marker comprises a SNP at chr7:24719176-24719176, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 605 comprising an allele at nucleobase position 101 in SEQ ID NO: 605. In some embodiments, the marker comprises a SNP at chr7:142752462-142752462, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 606 comprising an allele at nucleobase position 101 in SEQ ID NO: 606. In some embodiments, the marker comprises a SNP at chr8:98044922-98044922, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 607 comprising an allele at nucleobase position 101 in SEQ ID NO: 607. In someWSGR Docket No.56034-703.601 embodiments, the marker comprises a SNP at chr8:98045298-98045298, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 608 comprising an allele at nucleobase position 101 in SEQ ID NO: 608. In some embodiments, the marker comprises a SNP at chr9:123158097-123158097, wherein the SNP comprises at least 10 contiguous nucleic acid molecules of SEQ ID NO: 609 comprising an allele at nucleobase position 101 in SEQ ID NO: 609. Table 1B: SNPSWSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601WSGR Docket No.56034-703.601

[0104] Disclosed herein in the following embodiments, are the subsets of biomarkers disclosed herein: 1. A subset comprising at least one microsatellite allele or a SNP at a genetic locus. 2. The subset of embodiment 1 comprising a microsatellite locus provided in Table 1A or a SNP provided in Table 1B. 3. The subset of embodiments 1-2 wherein the subset comprises at least 100 microsatellite loci or SNPs. 4. The subset of embodiments 1-3, wherein the subset comprises at least 150 microsatellite loci or SNPs. 5. The subset of embodiments 1-4, wherein the subset comprises at least 200 microsatellite loci or SNPs. 6. The subset of embodiments 1-5, wherein the subset comprises at least 250 microsatellite loci or SNPs. 7. The subset of embodiments 1-6, wherein the subset comprises at least 300 microsatellite loci or SNPs. 8. The subset of embodiments 1-7, wherein the subset comprises at least 400 microsatellite loci or SNPs. 9. The subset of embodiments 1-8, wherein the subset comprises at least 500 microsatellite loci or SNPs. 10. The subset of embodiments 1-9, wherein the subset comprises 600 microsatellite loci or SNPs. 11. The subset of embodiments 1-10, wherein the microsatellite locus or SNP is selected from the group consisting of: ABCA12,AC072062.3; ABCA4; ABCA9,ABCA9-WSGR Docket No.56034-703.601 AS1; AC027612.3; ACTL6A; ACTN1; ADAM10; ADAM32,RPL3P10; ADAMTS20; ADAMTS9; ADAMTSL3; ADORA3; AGO3,RP4-665N4.8; AGO4; AGR3,RAD17P1; AGXT2; ALG1L2,FAM86HP; ANKRD40; ANO6; AP4S1; APEX2; APH1B; APITD1,APITD1-CORT; APOA4; ARHGAP10; ARHGAP26; ARL13B; ARL14EPL,CTD-2287O16.3; ARPC5L; ASAP2; ASB16; ASB6,RP11- 483H20.4; ATG2B; ATL2; ATM,C11orf65; ATP5G2; ATP6V0D1; ATPAF1; B3GNT4; B9D1; BACH2; BCKDHB; BCO2,KCTD9P4,AP002884.2; BEST1; BMP2K; BMP5; BMP6; BRDT; BST2; BTBD11; BZW2; C15orf39; C1orf35; C1QTNF3,C1QTNF3-AMACR; C2CD2; C4orf22,RP11-395C17.1; C4orf36,RP11- 397E7.4; C5orf22; C9,DAB2; CACHD1; CACNA2D4; CAMSAP2; CANX; CATSPER1; CCDC146,AC073635.5; CCDC6; CCDC7; CCL18; CCL4; CCNG1; CCNT2; CD209; CD38; CDC123; CDC25A; CDK11A,CDK11B,RP1-283E3.8; CDKAL1,RP3-348I23.2; CDON; CDV3P1; CELA3B; CEMIP,RP11-351M8.2; CENPE; CENPJ; CEP164; CEP89; CFAP52,RP11-55L4.2; CHN2; CHPF; CLCNKB; CLSPN; CLSPN,RP11-435D7.3; CLSTN2; CLYBL; CNKSR1; CNTLN; COL1A1; COL22A1; COL28A1; COL4A1; COL4A3,AC097662.2; COPS8; COX6B1P1; CPSF3L; CRBN; CREBBP; CROT; CST8,RP3-333B15.4; CTAGE3P; CTB- 95D12.1; CTC-512J12.6,ZNF285; CTD-2014B16.1; CTDP1; CTPS2; CUBN; CUX2; CXorf22; CYHR1,CTD-2517M22.16; CYP4A22,CYP4A22-AS1; DCP1B; DDIT4; DDX12P; DDX51; DDX54; DFNA5; DGKI; DLEU1,RPL34P26; DLG1; DMD; DMXL2,RP11-707P17.2,RP11-707P17.1; DNA2; DNAH11; DNAH2; DNAH5,CTB-51A17.1; DNAJB12; DNMT3B; DNTT; DOCK3,ZNF652P1; DOCK7; DPP10; DPP4; DSTYK; DTNB; DUOX1; DUOX1,CTD-2651B20.1; DYNLL1P7; EBNA1BP2,CFAP57; EFCAB2; EFEMP2; EHBP1; EIF2AK4; EMC3; EMR1; ENTPD7; EPG5; EPN2; ERCC1,CD3EAP; EVC2; EYA2; F2RL1; FAM102A; FAM171B; FAM187B; FAM3C; FAM98C; FANCA; FBXO11; FBXO47; FCGBP; FCGR1B,RP11-439A17.10; FIGN; FLNB; FLT4; FOLH1B; FOXP1; FOXP2; FRG1; FRMD4B; FSIP2; FUZ; GAD1; GANC; GBP4; GCNT3,GTF2A2; GLP1R; GLUD1P7; GLUD1P8; GOPC,ROS1; GP6,CTC- 550B14.7; GPSM2,RP11-475E11.9; GRAMD2; GRHL2; GRIP1; GTF2I; GTF2IP1; GUF1,GNPDA2; GYLTL1B; HEATR1; HEPHL1; HERC2; HERC2P3; HGS; HLA- DPA1; HNRNPCL1; HNRNPU; HRNR,FLG-AS1; HSD17B14; HSD17B4; HSPBP1; HYDIN; IFT57; IL9R,AJ271736.10; IQCA1; IQGAP1; ISPD; ITGA8; ITGB1BP1; JAKMIP2,JAKMIP2-AS1; JMJD7-PLA2G4B,PLA2G4B; KANK1;WSGR Docket No.56034-703.601 KAT2A; KAT6B,RP11-77G23.2; KCNN3; KIAA0368; KIAA0753; KIF23; KIFC1; KIFC2; KIR3DL1,KIR2DL4; KITLG; KPNA2; KPNA3; KRT71; KYNU; LAMA3; LAMA4; LDHB; LGALS3BP; LINS; LLGL2; LMO7,RP11-29G8.3; LRCH3; LRRC14; LRRC37A,ARL17B; LRRCC1; LTBP2; LY6G5C; LYPD6B; LYSMD4; LYST; MAGEF1; MALT1,RP11-126O1.4; MANSC1; MAP1LC3B; MAP2K3; MAP3K1; MAPK8IP1; MAPKAPK3; MASP1; MAU2; MBIP; MCOLN3,WDR63; MED23,ARG1; MEIS1; MEN1; MEOX2; MLH3; MMP9; MNAT1,RP11-307P22.1; MOG; MORC4,TBC1D8B; MRC1; MS4A4A; MSH6,FBXO11; MSLN,RHOT2; MSLN,WDR90; MSLN,WDR90,LA16c-349E10.1; MSRB2; MTG1,RP11-108K14.8; MVB12A; MX1; MYH15; MYLK; MYO5A; MYO5C; MYOF; MZT1; NADK2; NCK1,RAD51AP1P1; NCKAP1; NCOA3; NCOR2; NDUFA3; NEBL; NHS; NLRP12; NLRP6; NOC4L; NOTCH1; NPM1P38,MCHR2-AS1; NR6A1; NSD1; NTMT1; NTRK1; NUP153; NUTM2D; OAF; OCA2; OLFM1; OLFM2; OSBP; OTUD4; PACRG; PAPSS2; PARP10,GRINA; PARP4P2; PCDHB17P,AC005754.7,CH17-140K24.2; PDE2A; PDE4D; PDHA1; PDS5B; PEAR1; PGM3; PIK3C2A; PIK3C3; PLA2G4D; PLA2R1; PLEKHA5; PLIN1; PLK4; PLXDC2; PMP22; PMS2P3; POLQ; POM121B; PPFIA1,AP000487.6,CTA- 797E19.2; PPP2R5B; PPP3CB; PRICKLE3; PRKRIP1,RP11-163E9.1; PRMT3; PRSS1,TRBV25-1,TRBV28,TRBC2; PSG9; PSIP1; PTGS2; PTH2R,AC019185.4; PTPRD; PTPRF; PTPRO; PTTG1; PVRL1; PXMP4; QSOX1; RAB11FIP5; RECQL4; RIMS1; RIMS2; RNF168; RNF216; RP11-1084I9.1; RP11-192H23.4; RP11-204M4.2; RP11-5P18.10; RP11-624L12.1; RPL15P21,RP11-963H4.3; RPL30; RPL7P2; RRP1B; RSU1P2,CUBNP3,RP11-285G1.15; RUNX1; RUSC2; SAAL1; SAMM50; SAP130; SARS2,CTC-360G5.8; SC5D; SCN3A,AC013463.2; SEH1L; SEMA4D; SEMA4D,RP13-93L13.2; SERINC2; SERPINH1; SETP8; SF3A1; SFXN5; SH2B3; SH3RF1; SLC12A5; SLC25A10,SLC25A10; SLC25A15; SLC28A2,CTD-2651B20.3; SLC29A4; SLC38A4,RP11-96H19.1; SLC38A6; SLC45A2; SLC4A3; SLC6A3; SLCO2B1; SMURF1,AC004893.11; SNX7; SORCS1; SORL1; SP140; SP3P; SPAG8; SPDYE5; SPTLC3; SRRM2; STRAP; STRBP; STYXL1; SUSD4; SYCP2L; SYNE1; TAF1; TAF15; TANC1; TBC1D3H,TBC1D3I,TBC1D3F; TBCA,ACTBP2; TBL1X; TCF12,HNRNPA3P11; TEAD4; TESPA1; THEG; THOP1; TLN2; TLR1; TMC2; TMEFF1,MSANTD3- TMEFF1; TMEM120A; TMEM120B; TMEM56,TMEM56-RWDD3; TMEM87A; TMEM87A,RP11-546B15.1; TMPRSS2; TNFAIP6; TNS1; TONSL; TP53BP1;WSGR Docket No.56034-703.601 TPCN2; TPI1; TRAM2; TRAPPC6B; TRDMT1; TRIM33; TRIM69; TRIP11; TRMU; TSPAN7,RP5-972B16.2; TSPEAR,KRTAP10-7; TTC29; TTC3; TTN; TUBB3,MC1R; TUBB8P10; TUBB8P7; TYW3; UBE2R2; UBXN7; UGCG; UMOD; UMODL1; UNC5D; URI1; USE1; USH2A; USP34; USP36; UTRN; VSTM2A; VWA3A; WDPCP,HNRNPA1P66; WDR38; WDR59; WDR66; WDR72; WFDC10B; WWC2; XKRX; XXbac-BPG32J3.20,ABHD16A; YWHAQ; ZBTB33; ZBTB44,DDX18P5; ZFP91,ZFP91-CNTF; ZMYM1; ZNF135; ZNF195; ZNF217,RP4-724E16.2; and ZNF618. 12. The subset of embodiments 1-11, wherein at least one microsatellite locus comprises a major allele. 13. The subset of embodiments 1-12, wherein at least one microsatellite locus comprises a minor allele. 14. The subset of embodiments 1-13, wherein the subset comprises a combination of microsatellite loci and SNPs described in Table 2 or a microsatellite loci or SNP in linkage disequilibrium therewith as determined with an R2of at least about 0.80. 15. The subset of embodiments 1-13, wherein the subset comprises a combination of microsatellite loci and SNPs described in Combination 1 or a microsatellite loci or SNP in linkage disequilibrium therewith as determined with an R2of at least about 0.80. 16. The subset of embodiments 1-13, wherein the subset comprises a combination of microsatellite loci and SNPs described in Combination 2 or a microsatellite loci or SNP in linkage disequilibrium therewith as determined with an R2of at least about 0.80. 17. The subset of embodiments 1-13, wherein the subset comprises a combination of microsatellite loci and SNPs described in ...

Claims

WSGR Docket No.56034-703.601 CLAIMS What is Claimed:

1. A computer-implemented method of identifying a genetic classifier for determining whether a subject has a phenotype, the method comprising: (a) receiving sequence data from a plurality of samples obtained from a plurality of subjects comprising subjects that have the phenotype and subjects that do not have the phenotype, wherein the sequence data comprises a plurality of microsatellite alleles at microsatellite loci; (b) identifying a distribution of lengths of the plurality of microsatellite alleles at microsatellite loci; (c) performing a statistical test comprising tabulating the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci to determine which of the plurality of microsatellite alleles at microsatellite loci is predictive of the phenotype; (d) generating a score for a subset of microsatellite alleles at microsatellite loci of the plurality of microsatellite alleles at microsatellite loci determined to be predictive of the phenotype based, at least in part, on the performing the statistical significance test in (c); (e) determining an optimal cutoff value of the score for the subset of microsatellite alleles at microsatellite loci that differentiates the subjects that have the phenotype from the subjects that do not have the phenotype; and (f) applying the optimal cutoff value to the score for the subset of microsatellite alleles at microsatellite loci, thereby producing the genetic classifier.

2. The computer-implemented method of claim 1, wherein the optimal cutoff value is derived from a second statistical test performed for the subset of microsatellite alleles at microsatellite loci.

3. The computer-implemented method of claim 2, wherein the optimal cutoff value has a combined P-value of at most about 0.05 for differentiating the subjects that have the phenotype from the subjects that do not have the phenotype.

4. The computer-implemented method of claim 2, wherein the second statistical test comprises performing a receiver operating characteristic (ROC) analysis.

5. The computer-implemented method of claim 4, wherein the optimal cutoff value is derived from Youden’s index.WSGR Docket No.56034-703.601 6. The computer-implemented method of claim 1, wherein the determining the optimal cutoff value in (e) and the applying the optimal cutoff value in (f) are performed with a machine learning algorithm.

7. The computer-implemented method of claim 6, wherein the machine learning algorithm comprises an artificial neural network or a random forest algorithm.

8. The computer-implemented method of claim 7, further comprising training the machine learning algorithm with training sequence data from a training cohort of subjects having the phenotype and not having the phenotype.

9. The computer-implemented method of claim 7 or 8, further comprising validating the machine learning algorithm with validation sequence data from a validation cohort of subjects having the phenotype and not having the phenotype.

10. The computer-implemented method of claim 1, wherein the identifying the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci in (b) comprises aligning the plurality of microsatellite alleles at microsatellite loci with reference to a human genome using an alignment algorithm comprising a scoring matrix configured to align flanking regions of the plurality of microsatellite alleles at microsatellite loci, and wherein the flanking regions have a length that is greater than or equal to about 25 contiguous base pairs.

11. The computer-implemented method of claim 10, wherein the length is greater than or equal to about 50 contiguous base pairs.

12. The computer-implemented method of claim 11, wherein the length is about 100 contiguous base pairs in length.

13. The method of claim 1, further comprising: a) generating an image based on the distribution of lengths that were tabulated; and b) analyzing, by a machine learning model, the image, wherein the machine learning model generates the score in (d).

14. The method of claim 2, further comprising generating a second image based SNPs within the sequence data.

15. The method of claim 3, wherein the score is generated further based, at least in part, on the second image.

16. The method of claim 13, wherein the image comprises: a) an indication of one or more microsatellite alleles at one or more microsatellite loci, wherein the indication: i. comprises an intensity for the one or more microsatellite alleles at the one or more microsatellite loci; andWSGR Docket No.56034-703.601 ii. indicates a length for the one or more microsatellite alleles at the one or more microsatellite loci.

17. The method of claim 16, wherein the intensity for the one or more microsatellite alleles in the one or more microsatellite loci indicates a frequency of the one or more microsatellite alleles within the one or more microsatellite loci.

18. The method of claim 13, wherein the image is associated with a sample of the plurality of samples.

19. The method of claim 13, wherein an image is generated for each sample of the plurality of samples.

20. The computer-implemented method of claim 1, wherein the genetic classifier comprises greater than or equal to about 100 different microsatellite alleles.

21. The computer-implemented method of claim 20, wherein the genetic classifier comprises a combination of microsatellite alleles and single nucleotide polymorphisms.

22. The computer-implemented method of claim 21, wherein the combination of microsatellite alleles and single nucleotide polymorphisms comprises any one of Combination 1-13 in Table 2, or a microsatellite allele and single nucleotide polymorphism in linkage disequilibrium therewith as determined by a r2of greater than or equal to about 0.

80.

23. The computer-implemented method of any one of claims 20-22, wherein the alignment algorithm comprises improved gap open penalties (GOP), gap extension penalties (GEP), or any combination thereof, relative to a reference alignment algorithm with a scoring matrix configured to align flanking regions of the plurality of microsatellite alleles at microsatellite loci having a length of fewer or equal to 20 contiguous base pairs.

24. The computer-implemented method of claim 23, wherein the reference algorithm is a Needleman–Wunsch algorithm.

25. The computer-implemented method of claim 23, wherein the reference algorithm is a Smith-Waterman algorithm.

26. The computer-implemented method of any one of claims 10-25, wherein the alignment algorithm is a hybrid algorithm comprising components of a Needleman–Wunsch algorithm and a Smith-Waterman algorithm.

27. The computer-implemented method of any one of claims 1-26, further comprising applying a weight to a microsatellite allele at a microsatellite locus of the subset of microsatellite alleles at microsatellite loci.WSGR Docket No.56034-703.601 28. The computer-implemented method of claim 27, wherein the weight is based on a strength of an association of the microsatellite allele at the microsatellite locus to the phenotype relative to other microsatellite alleles at microsatellite loci of the subset.

29. The computer-implemented method of claim 27, wherein the weight is higher if the microsatellite allele at the microsatellite locus is a primary microsatellite allele than if the microsatellite locus is a minor microsatellite allele.

30. The computer-implemented method of claim 27, wherein the weight is higher if the microsatellite allele at the microsatellite locus comprises a single nucleotide polymorphism (SNP) than if the microsatellite allele at the microsatellite locus lacks the SNP.

31. The computer-implemented method of claim 27, wherein the weight is higher if the microsatellite allele at the microsatellite locus comprises a CpG dinucleotide than if the microsatellite allele at the microsatellite locus lacks the CpG dinucleotide, and wherein the CpG dinucleotide is a potential methylation site.

32. The computer-implemented method of any one of claims 1-31, wherein: a) the subset of microsatellite alleles at microsatellite loci comprises one or more minor microsatellite alleles; and / or b) the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci comprises the one or more minor microsatellite alleles.

33. The computer-implemented method of claim 32, wherein the genetic classifier additionally differentiates subjects having instability at one or more microsatellite loci of the one or more minor microsatellite alleles from other subjects of the plurality of subjects.

34. The computer-implemented method of any one of claims 1-33, wherein: a) the subset of microsatellite alleles at microsatellite loci comprises one or more primary microsatellite alleles and one or more minor microsatellite alleles; and / or b) the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci comprises of the one or more primary microsatellite alleles and the one or more minor microsatellite alleles.

35. The computer-implemented method of any one of claims 1-34, wherein: a) the subset of microsatellite alleles at microsatellite loci comprises one or more microsatellite alleles at microsatellite loci that comprises a SNP; and / or b) the distribution of lengths of the plurality of microsatellite alleles at microsatellite loci comprises of the one or morse microsatellite alleles at microsatellite loci that comprises the SNP.WSGR Docket No.56034-703.601 36. The computer-implemented method of claim 35, wherein the SNP is associated with a risk that a subject has or will develop the phenotype relative to another subject that lacks the SNP.

37. The computer-implemented method of any one of claims 1-36, wherein the statistical test is a Fisher test.

38. The computer-implemented method of any one of claims 1-36, wherein the statistical test is a regression analysis.

39. The computer-implemented method of any one of claims 1-36, wherein the statistical test is a chi-squared test.

40. The computer-implemented method of any one of claims 1-39, wherein the tabulating the lengths of the plurality of microsatellite alleles at microsatellite loci comprises coercing sequence data for the plurality of microsatellite alleles at microsatellite loci into a 2 x 2 table for each microsatellite allele at each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci.

41. The computer-implemented method of claim 40, wherein the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprises tabulating genotypes for each microsatellite allele at each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci based on which genotypes are most common in the subjects that have the phenotype.

42. The computer-implemented method of claim 40, wherein the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprises tabulating genotypes for each microsatellite allele at each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci based on an average length of microsatellite alleles at microsatellite loci in each microsatellite locus.

43. The computer-implemented method of claim 40, wherein the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprises tabulating genotypes for each microsatellite allele at each microsatellite locus of the plurality of microsatellite alleles at microsatellite loci based on a median length of microsatellite alleles at microsatellite loci in each microsatellite locus.

44. The computer-implemented method of claim 40, wherein the coercing the sequence data for the plurality of microsatellite alleles at microsatellite loci into the 2 x 2 table comprises: a) aggregating the plurality of microsatellite alleles at microsatellite loci for each microsatellite allele at each microsatellite locus into long alleles and short alleles relative to a cutoff length; andWSGR Docket No.56034-703.601 b) tabulating the plurality of microsatellite alleles at microsatellite loci for each microsatellite allele at each microsatellite locus into the 2 x 2 table based on whether a microsatellite allele at a microsatellite locus of the plurality of microsatellite alleles at microsatellite loci is longer or shorter than the cutoff length.

45. The computer-implemented method of any one of claims 32-44, wherein the cutoff length is determined using an ROC analysis per microsatellite locus.

46. The computer-implemented method of any one of claims 1-45, wherein the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a positive predictive value (PPV) of at least about 51%.

47. The computer-implemented method of claim 46, wherein the PPV is at least about 60%.

48. The computer-implemented method of claim 46, wherein the PPV is at least about 70%.

49. The computer-implemented method of any one of claims 1-48, wherein the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a sensitivity of at least about 70%.

50. The computer-implemented method of claim 49, wherein the sensitivity is at least about 80%.

51. The computer-implemented method of any one of claims 1-50, wherein the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a specificity value of at least about 70%.

52. The computer-implemented method of claim 51, wherein the specificity is at least about 80%.

53. The computer-implemented method of any one of claims 1-52, wherein the genetic classifier differentiates the subjects that have the phenotype from the subjects that do not have the phenotype with a Confidence Interval (CI) of at least about 0.

90.

54. The computer-implemented method of any one of claims 1-53, wherein the phenotype is a disease or a condition.

55. The computer-implemented method of claim 54, wherein the disease is: a) a cancer; b) a neoplasm; c) a cardiac disease, or d) a neurological disease.

56. The computer-implemented method of claim 55, wherein the cancer is lung cancer or cancer metastasized into lung tissue.WSGR Docket No.56034-703.601 57. The computer-implemented method of claim 55, wherein the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease.

58. The computer-implemented method of claim 55, wherein the neurological disease is schizophrenia, bipolar disorder, autism, or Alzheimer’s disease.

59. The computer-implemented method of any one of claims 1-58, wherein the plurality of subjects comprises greater than or equal to about 20,000 subjects.

60. The computer-implemented method of claim 59, wherein: a) about half of the plurality of subjects have the phenotype, and the other half of the plurality of subjects do not have the phenotype; or b) a distribution of the plurality of subjects between a first fraction of the plurality of subjects that have the phenotype and a second fraction of the plurality of subjects that do not have the phenotype reflects the distributions in the general population.

61. The computer-implemented method of any one of claims 1-60, further comprising providing: a) the optimal cutoff value; b) the subset of the microsatellite alleles at microsatellite loci; and c) an odds ratio for each microsatellite alleles at each microsatellite locus of the subset of microsatellite alleles at microsatellite loci.

62. The computer-implemented method of any one of claims 1-61, further comprising excluding microsatellite alleles at microsatellite loci from the subset of microsatellite alleles at microsatellite loci having a call rate of less than 80%.

63. The computer-implemented method of claim 62, wherein the call rate is calculated by dividing a total number of genotypes characterized by a microsatellite allele at microsatellite locus length by a total number of microsatellite alleles at microsatellite locus lengths observed for the plurality of microsatellite alleles at microsatellite loci.

64. The computer-implemented method of any one of claims 1-63, wherein the phenotype comprises a condition, and wherein the condition comprises: a) a rate of metabolism of a therapeutic agent in the subject; b) a positive therapeutic response of the subject to the therapeutic agent; c) a therapeutic non-response or loss of response of the subject to the therapeutic agent d) an allergy; e) an adverse effect to a therapeutic agent;WSGR Docket No.56034-703.601 f) a sensitivity to a chemical or environmental agent associated with cancer; g) obesity; or h) substance abuse.

65. A computer-implemented system comprising a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the computing device to perform a method in any one of claims 1-64.

66. A system comprising: a) the computer-implemented system of claim 65; and b) a nucleic acid sequencer configured to transmit the sequence data to the at least one processor.

67. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by one or more processors to determine a probability of a subject as having or developing a phenotype, the non-transitory computer-readable storage media comprising: a) a database, in a computer memory, of the instructions comprising a method in any one of claims 1-64; and b) a software module configured to perform the instructions.

68. A method of treating the disease or the condition in a subject, the method comprising: administering a therapy to the subject to treat the disease or the condition, wherein the subject is predicted to have the phenotype with the genetic classifier identified by a method of any one or claims 54-64, wherein the phenotype comprises a positive therapeutic response to the therapy.

69. A method of predicting whether a subject has or will develop the disease or the condition, the method comprising: a) determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; and b) applying the genetic classifier identified by a method of any one or claims 54-64 to the lengths of microsatellite alleles at microsatellite loci to predict whether the subject has or will develop the phenotype, wherein the phenotype comprises the disease or the condition.

70. A method of optimizing a treatment for a disease or a condition in a subject, the method comprising:WSGR Docket No.56034-703.601 a) determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject, wherein the subject has been administered a therapy to treat the disease or the condition; and b) applying the genetic classifier identified by a method of any one of claims 54-64 to the lengths of microsatellite alleles at microsatellite loci to predict whether the subject will exhibit the phenotype, wherein the phenotype comprises a positive therapeutic response or the loss of therapeutic response to the therapy.

71. A method of predicting a rate of metabolism of a therapy in a subject, the method comprising: a) determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; and b) applying the genetic classifier identified by a method of any one of claims 54-64 to the lengths of microsatellite alleles at microsatellite loci to predict whether the subject has or will develop the phenotype, wherein the phenotype comprises a rate of the metabolism of the therapy in the subject that is faster than a control rate.

72. A method of treating a disease or a condition in a subject, the method comprising: administering a therapy to the subject to treat the disease or the condition, wherein the subject is predicted to have or develop a phenotype with a test having a positive predictive value of at least about 51%, based on lengths of a plurality of microsatellite alleles at a plurality of microsatellite loci detected in a biological sample obtained from the subject, and wherein the phenotype comprises a therapeutic response to the therapy.

73. A method of predicting whether a subject has or will develop a disease or a condition, the method comprising: a) determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; b) comparing the lengths of microsatellite alleles at microsatellite loci to an optimal cutoff value, wherein the optimal cutoff value is determined by performing a statistical analysis on a subset of microsatellite alleles at microsatellite loci determined to be statistically significantly associated with incidence of the disease or the condition in a reference population of subjects having the disease or the condition and subjects not having the disease or the condition; and c) predicting whether the subject has or will develop a phenotype based, at least in part, on the comparing in (b) with a positive predictive value of at least about 51%, wherein the phenotype comprises the disease or the condition.WSGR Docket No.56034-703.601 74. A method of monitoring a treatment for a disease or a condition in a subject, the method comprising: a) determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject, wherein the subject has been administered a therapy to treat the disease or the condition; b) comparing the lengths of microsatellite alleles at microsatellite loci to an optimal cutoff value, wherein the optimal cutoff value is determined by performing a statistical analysis on a subset of microsatellite alleles at microsatellite loci determined to be statistically significantly associated with a positive therapeutic response or a loss of therapeutic response to the therapy to treat the disease or the condition in a reference population having the disease or the condition; and c) predicting that the subject will has or will develop a phenotype with a positive predictive value of at least about 51%, wherein the phenotype comprises the positive therapeutic response or the loss of therapeutic response to the therapy.

75. A method of predicting a rate of metabolism of a therapy in a subject, the method comprising: a) determining lengths of microsatellite alleles at microsatellite loci present in a biological sample obtained from the subject; b) comparing the lengths of microsatellite alleles at microsatellite loci to an optimal cutoff value, wherein the optimal cutoff value is determined by performing a statistical analysis on a subset of microsatellite alleles at microsatellite loci determined to be statistically significantly associated with an aberrant rate of metabolism of the therapy to treat the disease or the condition in a reference population having the disease or the condition; and c) predicting whether the subject has or will develop the phenotype based, at least in part, on the comparing in (b) with a positive predictive value of at least about 51%, wherein the phenotype comprises the rate of the metabolism of the therapy in the subject that is faster than a control rate.

76. The method of any one of claims 72-75, wherein the statistical analysis is an ROC analysis, a Fisher test, or a chi-squared test.

77. The method of any one of claims 72-76, wherein the disease is cancer.

78. The method of claim 76, wherein the cancer is lung cancer or cancer metastasized into lung tissue of the subject.

79. The method of any one of claims 72-76, wherein the disease is a cardiac disease.WSGR Docket No.56034-703.601 80. The method of claim 79, wherein the cardiac disease comprises an atrial fibrilization, a myocardial infarction, angina, heart failure, or coronary heart disease.

81. The method of claim 73, wherein the condition comprises a rate of metabolism of the therapy to treat the disease.

82. The method of claim 73, wherein the condition comprises a therapeutic response to the therapy to treat the disease.

83. The method of claim 73, wherein the condition comprises a therapeutic non-response or loss of response to the therapy to treat the disease.

84. The method of any one of claims 73-83. wherein the positive predictive value is at least about 60% 85. The method of any one of claims 73-84, wherein the positive predictive value is at least about 70%.

86. The method of any one of claims 68, 72 and 76-85, wherein the subject is predicted to have the phenotype with a specificity of at least about 70%.

87. The method of any one of claims 68, 72 and 76-85, wherein the subject is predicted to have the phenotype with a sensitivity of at least about 70%.

88. The method of any one of claims 68, 72 and 76-85, wherein the subject is predicted to have the phenotype with a confidence interval (CI) of at least about 0.

90.

89. The method of any one of claims 69-72 and 73-88, wherein the predicting is performed with a specificity of at least about 70%.

90. The method of any one of claims 69-71 and 73-89, wherein the predicting is performed with a sensitivity of at least about 70%.

91. The method of any one of claims 69-71 and 73-90, wherein the predicting is performed with a CI of at least about 0.

90.

92. The method of any one of claims 68-91, wherein the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles.

93. The method of any one of claims 68-92, wherein the microsatellite alleles at microsatellite loci present in the biological sample comprise one or more minor microsatellite alleles and one or more major microsatellite alleles.

94. The method of any claim 92 or 93, further comprising predicting genomic instability at one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample.WSGR Docket No.56034-703.601 95. The method of any one of claims 68-94, wherein the microsatellite loci present in a biological sample comprise one or more microsatellite alleles at microsatellite loci comprising one or more SNPs.

96. The method of any one of claims 58-95, further comprising detecting a presence or an absence of one or more SNPs in the biological sample to produce a combination of microsatellite alleles and SNPs.

97. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 1 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

98. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 2 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

99. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 3 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

100. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 4 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

101. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 5 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

102. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 6 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

103. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 7 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

104. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 8 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

105. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 9 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.80.WSGR Docket No.56034-703.601 106. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 10 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

107. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 11 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

108. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 12 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

109. The method of claim 96, wherein the combination of microsatellite alleles and SNPs comprises Combination 13 in Table 2 or a microsatellite allele or SNP in linkage disequilibrium therewith as determined by an r2value of at least about 0.

80.

110. The method of any one of claims 68-95, further comprising determining one or more genotypes for one or more microsatellite alleles at microsatellite loci of the microsatellite loci present in the biological sample, and wherein the one or more microsatellite alleles at microsatellite loci is present in the biological sample.

111. The method of claim 103, wherein the lengths of microsatellite alleles at microsatellite loci that are determined comprise: a) an average length of all microsatellite alleles per microsatellite locus; b) a median length of all microsatellite alleles per microsatellite locus; c) a longest length of all microsatellite alleles per microsatellite locus; d) a shortest length of all microsatellite alleles per microsatellite locus; e) any one of (a) to (d) for genotypes at a given microsatellite locus determined to be the most prevalent in a reference subject population of subjects having the disease or the condition; or f) any combination of (a) to (e).

112. The method of any one of claims 69, 73, and 76-111, further comprising treating the disease in the subject by delivering a therapy to the subject.

113. The method of any one of claims 68, 70-71 and 74-112, further comprising treating the disease in the subject by delivering the therapy to the subject.

114. The method of any one of claims 68, 70-72 and 74-107, wherein the therapy is a cancer therapy provided in Table 3.

115. The method of any one of claims 67, 70-72 and 74-107, wherein the therapy is an immunotherapy, a gene editing therapy, or a surgery.WSGR Docket No.56034-703.601 116. The method of claim 115, wherein the gene editing therapy is a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) therapy.

117. The method of claim 115, wherein the immunotherapy is a cancer therapy.

118. The method of claim 117, wherein the cancer therapy is an immune checkpoint inhibitor, a cancer vaccine, or an adoptive T cell therapy.

119. The method of any one of claims 68, 70-72 and 74-118, wherein the therapy is a cardiac therapy.

120. The method of claim 119, wherein the cardiac therapy comprises an anticoagulant, an antiplatelet agent, a dual antiplatelet therapy, an ACE inhibitor, an Angiotensin II receptor blocker, an Angiotensin receptor-neprilysin inhibitor, a Beta blocker, a calcium channel blocker, a cholesterol-lowering medication, a digitalis preparation, a diuretic, or a vasodilator.

121. The method of any one of claims 69-120, wherein the biological sample is a whole blood sample, a plasma sample, a saliva sample, or a solid tissue sample.

122. The method of claim 121, wherein the whole blood sample is a peripheral blood sample.

123. The method of any one of claims 68-122, further comprising calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a percentile risk relative to a distribution of scores calculated for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

124. The method of any one of claims 68-122, further comprising calculating a score for the subject indicative of a risk that the subject has or will develop the phenotype, wherein the score is expressed as a Z score of scores for a reference subject population of subjects having the phenotype and subjects not having the phenotype.

Citation Information

Patent Citations

  • Oral cancer markers and their detection

    US20090023138A1

  • Methods and compositions for correlating genetic markers with prostate cancer risk

    US20090226912A1

  • Bioinformatics Systems, Apparatuses, and Methods for Generating a De Bruijn Graph

    US20190130998A1

  • System and method for single channel whole cell segmentation

    US20190147215A1

  • Systems and methods for processing electronic images for biomarker localization

    US20220130041A1