Method and system for microsatellite analysis

By constructing an optimized microsatellite classifier and analyzing a subset of microsatellites, the unreliability of microsatellite analysis in the early stages of detection and diagnosis in existing technologies has been solved, enabling early and accurate detection and characterization of complex multigene health states.

CN120913652APending Publication Date: 2025-11-07ORBIT GENOMICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511020251.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-04-22
Filing Date
2020-04-21
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively predict, detect, and characterize complex polygenic health conditions, such as cancer, neurological disorders, or cardiovascular diseases, in the early stages of microsatellite analysis, leading to unreliability and difficulty in detection and diagnosis.

Method used

An optimized microsatellite classifier was constructed using computer-based methods. By employing subset sorting of microsatellites and genetic algorithms, a microsatellite population associated with the disease was identified. Multiple parameters and minor allele characteristics were then used for analysis to determine the genomic age and disease type of the subjects.

Benefits of technology

It has improved the accuracy and reliability of early prediction, detection and characterization of complex polygenic health status, and enhanced the basis for diagnostic and treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913652A_ABST
    Figure CN120913652A_ABST
Patent Text Reader

Abstract

The present disclosure provides methods and systems for microsatellite analysis, in particular methods and systems for classifying microsatellites and secondary alleles in a sample. In addition, the present disclosure provides methods and systems for generating classifiers for conditions based on microsatellite sites and for performing pan cancer assays. The methods and systems may involve next generation sequencing of a nucleic acid sample from a subject and genotyping of microsatellite sites in the sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 202080045548.1, entitled "Method and System for Microsatellite Analysis" (the corresponding PCT application was filed on April 21, 2020, with application number 202080045548.1).

[0002] Cross-referencing

[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 837,109, filed April 22, 2019, the entire contents of which are incorporated herein by reference. Background Technology

[0004] Microsatellites (MS) and their alterations and instabilities can be the genetic drivers behind numerous complex polygenic health states, including cancer, neurological disorders, and cardiovascular diseases. Currently, predicting, detecting, diagnosing, and characterizing these health states via microsatellites involves matching a patient's microsatellite profile with a database of microsatellites associated with these health states. Such methods are only applicable to later stages of health state progression, which can lead to unreliability and difficulties in detection, prognosis, diagnosis, treatment selection, and treatment outcomes. Therefore, improved methods for predicting, detecting, and characterizing these health states at both early and late stages remain needed through the analysis of microsatellite loci. Summary of the Invention

[0005] In one aspect, this disclosure provides a computer-implemented method for constructing an optimized classifier for a disease, the method comprising sorting a subset of multiple microsatellites into a classifier for the disease in multiple optimization cycles, wherein the subset of multiple microsatellites includes microsatellites in an initial population of microsatellites associated with the disease, thereby identifying an optimized subset of the subset of multiple microsatellites as the optimized classifier for the disease. In some aspects, the computer-implemented method further includes comparing microsatellites from a first group of samples from subjects suffering from the disease and microsatellites from a second group of samples from subjects not suffering from the disease, thereby identifying an initial population of microsatellites.

[0006] Ranking may include comparing microsatellites from a first set of samples from subjects with the condition and microsatellites from a second set of samples from subjects without the condition to identify an initial population of microsatellites. A computer-implemented method may include initialization, wherein initialization includes randomly selecting an initial subset of microsatellites from the initial population of microsatellites for ranking in optimization cycles across multiple optimization cycles. A population of at least about 100 subsets of the initial population of microsatellites may be used in multiple optimization cycles. The minimum number of microsatellites in a subset of microsatellite subsets may be 8. The maximum number of microsatellites in a subset of microsatellite subsets may be 64. In some cases, duplicate microsatellites are not allowed in a subset of microsatellite subsets. Ranking may include performing receiver operating characteristic (ROC) analysis using (i) a subset of microsatellites, (ii) microsatellites from samples from subjects with the condition, and (iii) microsatellites from samples from subjects without the condition. Ranking in optimization cycles across multiple optimization cycles may include determining the sum of the sensitivity and specificity of microsatellites in each subset of a subset serving as a classifier for the condition. The optimization cycle of multiple optimization cycles may include adding 10 new subsets of the initial population of microsatellites to subsets from previous optimization cycles. Seven of the 10 new subsets may be generated by randomly splitting and recombinating two subsets randomly selected from previous optimization cycles, and three of the 10 new subsets may be generated by randomly selecting microsatellites from the initial population of said microsatellites. The method may also include at least in part discarding 10 subsets of the subsets in the optimization cycle based on the lowest ranking in the optimization cycle. In some cases, the symptom may be the presence or absence of a subject's health status. The symptom may be an increased or decreased likelihood of a subject developing a health status. The symptom may be an increased or decreased likelihood of a subject benefiting from treatment for a health status. In some cases, the symptom may be an increased or decreased likelihood of a subject having an increased risk of adverse effects due to treatment for a health status. The symptom may be a subject's responsiveness to treatment for a health status. In some cases, the symptom may be a prognosis of a subject's health status. In some cases, a health status may be cancer. Cancer may be lung cancer. In other cases, a health status may be a neurological disease or a cardiovascular disease.

[0007] In another aspect, this disclosure provides a computer-implemented method comprising determining the value of a classifier for a condition from samples of a subject using a plurality of parameters, wherein each of the plurality of parameters is a statistical measure of the correlation of each of a plurality of microsatellites from samples of a subject having said condition and / or from samples of a subject not having said condition.

[0008] Multiple weights may include multiple optimal weights. In some aspects, computer-implemented methods may include determining multiple optimal weights. Determining multiple optimal weights may include applying standardized regression analysis to the multiple weights. Determining multiple optimal weights may include using a genetic algorithm. Determining a classifier may include using minor allele frequency data. Multiple microsatellites may include at least 10 microsatellites. In some instances, each of the multiple microsatellites is associated with the presence of a disease. The value of the classifier may also include comparing the classifier to a threshold. In some aspects, the disease may be the presence or absence of a subject's health status, an increased or decreased likelihood of a subject developing a health status, an increased or decreased likelihood of a subject benefiting from treatment for a health status, an increased or decreased likelihood of a subject having an increased risk of adverse effects from treatment for a health status, a subject's responsiveness to treatment for a health status, or a combination thereof. In some cases, the health status is cancer, cardiovascular disease, or a neurological disease. When the health status is cancer, the cancer may be lung cancer.

[0009] In another aspect, this disclosure provides a computer-implemented method for determining the genomic age of a subject, the method comprising: determining microsatellite minor allele characteristics in a first sample from the subject; processing the microsatellite minor allele characteristics with a reference; and determining the genomic age of the subject based on the processing.

[0010] In some cases, the process includes comparing microsatellite minor allele characteristics with a reference. Minor allele characteristics can be many important alleles at a locus. The number of minor alleles can be supported by at least three next-generation sequencing reads. Minor allele characteristics can be a normalized total of minor allele reads to the total number of major allele reads at the locus. The method may also include performing next-generation sequencing on a first sample from the subject to generate microsatellite reads for the subject. The first sample may include blood, saliva, or a tumor. The method may also include determining minor allele characteristics in a second sample from the subject after determining a first genomic age. The method may include evaluating minor allele characteristics in the first sample and the second sample from the subject, and determining the subject's genomic aging rate based on the evaluation.

[0011] In another aspect, this disclosure provides a computer-implemented method comprising: using microsatellites from a sample from a subject to determine multiple classifiers from the sample from the subject; processing the multiple classifiers with multiple reference classifiers for multiple conditions; and, based on the processing, determining at least one condition for the subject from the multiple conditions.

[0012] The processing may include comparing multiple classifiers with multiple reference classifiers for multiple conditions. In some cases, at least one condition among multiple conditions includes the presence or absence of at least one of the subject's multiple health states. In some cases, at least one condition among multiple conditions includes an increased or decreased likelihood of developing at least one health state from the subject's multiple health states. At least one condition among multiple conditions may include an increased or decreased likelihood that the subject benefits from treatment for at least one of the subject's multiple health states. At least one condition among multiple conditions may include an increased or decreased likelihood that treatment for at least one of the subject's multiple health states results in an increased risk of adverse effects on the subject. At least one condition among multiple conditions may include the subject's responsiveness to treatment for at least one of the subject's multiple health states. Multiple health states may include multiple cancers, including ovarian cancer, breast cancer, low-grade glioma, glioblastoma, lung cancer, prostate cancer, or melanoma. In some cases, multiple health states may include multiple neurological disorders or multiple cardiovascular diseases.

[0013] In one aspect, this disclosure provides a non-transitory computer-readable medium including executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method for constructing an optimized classifier for a disease. The method includes, in multiple optimization cycles, sorting subsets of a plurality of microsatellites into a classifier for the disease, wherein the subset of microsatellites includes microsatellites from an initial population of microsatellites associated with the disease, thereby identifying an optimized subset of the subset of microsatellites as the optimized classifier for the disease. The computer-implemented method may further include comparing microsatellites from a first set of samples from subjects suffering from the disease with microsatellites from a second set of samples from subjects not suffering from the disease, thereby identifying an initial population of microsatellites.

[0014] Ranking may include comparing microsatellites from a first set of samples from subjects with the disease and microsatellites from a second set of samples from subjects without the disease to identify an initial population of microsatellites. A computer-implemented method may include initialization, wherein initialization includes randomly selecting an initial subset of microsatellites from the initial population of microsatellites for ranking in optimization cycles across multiple optimization cycles. A population of at least about 100 subsets of the initial population of microsatellites may be used in multiple optimization cycles. The minimum number of microsatellites in a subset of microsatellite subsets may be 8. The maximum number of microsatellites in a subset of microsatellite subsets may be 64. In some embodiments, duplicate microsatellites are not allowed in a subset of microsatellite subsets. Ranking may include performing receiver operating characteristic (ROC) analysis using (i) a subset of microsatellites, (ii) microsatellites from samples from subjects with the disease, and (iii) microsatellites from samples from subjects without the disease. Ranking in optimization cycles across multiple optimization cycles may include determining the sum of the sensitivity and specificity of microsatellites in each subset that serves as a classifier for the disease. The optimization cycle of multiple optimization cycles may include adding 10 new subsets of the initial population of microsatellites to subsets from previous optimization cycles. Seven of the 10 new subsets may be generated by randomly splitting and recombinating two subsets randomly selected from previous optimization cycles, and three of the 10 new subsets may be generated by randomly selecting microsatellites from the initial population of said microsatellites. The method may also include at least in part discarding 10 subsets of the subset in the optimization cycle based on the lowest ranking in the optimization cycle. The condition may be the presence or absence of a subject's health status. The condition may be an increased or decreased likelihood of a subject developing a health status. The condition may be an increased or decreased likelihood of a subject benefiting from treatment for a health status. The condition may be an increased or decreased likelihood of a subject having an increased risk of adverse effects due to treatment for a health status. The condition may be a subject's responsiveness to treatment for a health status. The condition may be a prognosis of a subject's health status. A health status may be cancer. Cancer may be lung cancer. A health status may be a neurological disease or a cardiovascular disease.

[0015] In another aspect, this disclosure provides a non-transitory computer-readable medium including executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising using a plurality of parameters to determine the value of a classifier for a condition from samples of a subject, wherein each of the plurality of parameters is a statistical measure of the correlation of each of a plurality of microsatellites from samples of a subject having the condition and / or from samples of a subject not having the condition.

[0016] Multiple weights may include multiple optimal weights. Computer-implemented methods may include determining multiple optimal weights. Determining multiple optimal weights may include applying standardized regression analysis to the multiple weights. Determining multiple optimal weights may include using a genetic algorithm. Determining a classifier may include using minor allele frequency data. Multiple microsatellites may include at least 10 microsatellites. Each of the multiple microsatellites may be associated with the presence of a condition. The value of the classifier may also include comparing the classifier to a threshold. The condition may be the presence or absence of a subject's health status, an increased or decreased likelihood of a subject developing a health status, an increased or decreased likelihood of a subject benefiting from treatment for a health status, an increased or decreased likelihood of a subject having an increased risk of adverse effects from treatment for a health status, a subject's responsiveness to treatment for a health status, or a combination thereof. A health status may be cancer, cardiovascular disease, or a neurological disease. Cancer may be lung cancer.

[0017] In another aspect, this disclosure provides a non-transitory computer-readable medium including executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method for determining the genomic age of a subject, the method comprising: determining microsatellite minor allele characteristics in a first sample from the subject; processing the microsatellite minor allele characteristics with a reference; and determining the genomic age of the subject based on the processing.

[0018] The process may include comparing microsatellite minor allele characteristics with a reference. Minor allele characteristics may be a number of major alleles at a locus. The number of minor alleles may be supported by at least three next-generation sequencing reads. Minor allele characteristics may be normalized to the total number of reads of major alleles at the locus. The method may further include performing next-generation sequencing on a first sample from the subject to generate microsatellite reads for the subject. The first sample may include blood, saliva, or a tumor. The method may further include determining minor allele characteristics in a second sample from the subject after determining a first genomic age. The method may include evaluating minor allele characteristics in the first sample and the second sample from the subject, and determining the subject's genomic aging rate based on the evaluation.

[0019] In another aspect, this disclosure provides a non-transitory computer-readable medium including executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising: using microsatellites from a sample from a subject to determine a plurality of classifiers from the sample from the subject; processing the plurality of classifiers with a plurality of reference classifiers for a plurality of conditions; and, based on the processing, determining at least one condition for the subject from the plurality of conditions.

[0020] The processing may include comparing multiple classifiers with multiple reference classifiers for multiple conditions. At least one condition among multiple conditions may include the presence or absence of at least one of the subject's multiple health states. At least one condition among multiple conditions may include an increased or decreased likelihood of developing at least one health state from the subject's multiple health states. At least one condition among multiple conditions may include an increased or decreased likelihood that the subject benefits from treatment for at least one of the subject's multiple health states. At least one condition among multiple conditions may include an increased or decreased likelihood that treatment for at least one of the subject's multiple health states results in an increased risk of adverse effects on the subject. At least one condition among multiple conditions may include the subject's responsiveness to treatment for at least one of the subject's multiple health states. Multiple health states may include multiple cancers, including ovarian cancer, breast cancer, low-grade glioma, glioblastoma, lung cancer, prostate cancer, or melanoma. Multiple health states may include multiple neurological disorders or multiple cardiovascular diseases.

[0021] Another aspect of this disclosure provides a non-transitory computer-readable medium including machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0022] Another aspect of this disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory includes machine-executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.

[0023] Further aspects and advantages of this disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the disclosure are shown and described. As will be appreciated, this disclosure is capable of other and different embodiments, and certain details thereof can be modified in various obvious ways without departing from this disclosure. Therefore, the drawings and descriptions are to be considered illustrative in nature and not restrictive.

[0024] Incorporation

[0025] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent, or patent application is specifically and individually indicated as being incorporated by reference. Where a publication, patent, or patent application incorporated by reference contradicts the disclosure contained in this specification, the specification is intended to supersede and / or take precedence over any such contradictory material. Attached Figure Description

[0026] The novel features of the invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description, which illustrates illustrative embodiments in which the principles of the invention are utilized, and is shown in the accompanying drawings:

[0027] Figure 1 The illustration shows an example of the workflow of a computer-implemented method for generating a microsatellite classifier.

[0028] Figure 2 The illustration shows an example of the development process of using a computer-implemented method to identify informative microsatellite loci and generate a classifier for the disease.

[0029] Figure 3 The illustration shows an example of the validation process for lung cancer testing.

[0030] Figure 4 The illustration shows an example of validating a pan-cancer assay.

[0031] Figure 5 The diagram illustrates an example of a workflow used to analyze patient samples.

[0032] Figure 6 The illustration shows a method for identifying and validating medulloblastoma (MB)-associated microsatellite metastases (MS). The method comprises three stages: calculating informative MS loci using a training set, validating microsatellite markers in an independent validation cohort, and performing downstream analysis of genes associated with these MSs. The first stage includes a filter to eliminate MSs that vary with age, ethnicity, and sequencing technology.

[0033] Figures 7A to 7D An example of validation and training data is illustrated. Figure 7A The diagram illustrates the distribution of metric scores in the training queue. Figure 7C The figure illustrates the distribution of indicator scores in the validation cohort. (For training (120MB of subjects and 425 control subjects)) Figure 7B ) and validation (102MB subjects and 428 control subjects) cohort ( Figure 7D Perform ROC analysis.

[0034] Figure 8A The illustration shows a pie chart displaying the genomic locations of 139 MS informative sites of MB. Figure 8B The illustration shows the gene ontology analysis of MS loci in informative medulloblastoma. Figure 8C The protein-protein interaction (PPI) network of 124 genes associated with informative MS sites is illustrated. The PPI contains 129 nodes and 49 edges, resulting in a network enriched with a p-value of 0.0007.

[0035] Figure 9 The illustration shows an example of genotype distribution and contingency tables used in the study described in this paper. Genotype distribution of microsatellite marker 242626 on chromosome 1 base pair 153645035. The p-value for this example is 3.5e. -4 The table on the right is a contingency table of the same microsatellite marker.

[0036] Figure 10 The diagram illustrates a workflow outline for identifying age-sensitive MSs.

[0037] Figure 11 The diagram illustrates a workflow outline for identifying MSs that are sensitive to sequencing technologies.

[0038] Figure 12 The diagram illustrates a workflow outline for identifying race-sensitive MSs.

[0039] Figure 13 The illustration shows an example of an indicator used to assign scores to a sample. Consider hypothetical samples with genotypes 22|22, 12|12, and 13|13 respectively, for the above labeling. To apply the indicator to this sample, the frequency differences of each genotype in the MB group and the healthy group are summed: the result is a score of 0.95. In other words, for each genotype, its frequency in the MB group is subtracted from its frequency in the normal group; then the differences are summarized. Therefore, healthy controls primarily have negative scores, while affected individuals have positive scores.

[0040] Figure 14 The diagram illustrates the Youden index, used to determine the criterion for distinguishing between MB and healthy samples. The Youden index is also used to determine the cutoff value for the ROC curve in the training set. The optimal criterion for the 43-label list is 0.155. The same criterion is used to calculate the specificity and sensitivity of the validation cohort.

[0041] Figure 15 The diagram shows the chromosomal locations of 43 informative sites indicating MB.

[0042] Figure 16The diagram illustrates the genotypic distribution of microsatellite markers 166663 (an exon microsatellite located in the RAI gene) and 164048 (an exon microsatellite located in the BLC6B gene). Adding a CAG triplet can alter protein structure and weaken its function, similar to a missense mutation.

[0043] Figure 17 The illustration shows an example of an output from a computer-based method for reporting microsatellite analysis results to assess a subject's risk of developing cancer.

[0044] Figure 18 The illustration depicts a computer system that has been programmed or otherwise configured to implement the methods provided herein.

[0045] Figure 19 The diagram illustrates a list of 139 informative phylogenetic MSs associated with MB.

[0046] Figure 20 The diagram illustrates a list of 43 microsatellite sites in the MB signature set.

[0047] Figure 21 The illustration shows the ingenuity pathway analysis of informative MBMS sites.

[0048] Figure 22 The illustration shows mutations in genes associated with informative MBMS sites in the cBioportalMB cohort.

[0049] Figure 23 The illustration shows an analysis of the cBioportal MB cancer study, which revealed that mutations in 135 gene pairs tend to co-occur significantly in the MB cancer risk classifier.

[0050] Figure 24 The diagram illustrates a threshold with a confidence interval of 1 standard deviation. Classifiers outside the interval indicate whether a subject has the condition (above 0.5) or does not have the condition (below 0.1). Classifier values ​​further away from the threshold carry a stronger indication. Detailed Implementation

[0051] I. Overview

[0052] This disclosure provides a computer-implemented method for generating disease classifiers using, for example, microsatellites. Figure 1The diagram illustrates an example of how a computer-implemented method generates a classifier. Deoxyribonucleic acid (DNA) sequences are obtained from a sequence information database (101) of samples from subjects with the disease and a sequence information database (102) of reference subjects without the disease. Microsatellite loci from 101 and 102 are identified (genotyped) and compared to reveal microsatellite populations that are only associated with or related to the disease (103). These microsatellite populations are then further analyzed and weighted to obtain an initial set of microsatellite loci (104) for classifier optimization (105). The optimization iteratively arranges how the microsatellites are associated with or related to the disease. Optimization can be repeated with additional microsatellite sets for additional optimization cycles. In some cases, microsatellite sets are randomly split and recombine to generate new initial microsatellite sets for additional optimization cycles (106). After optimization, the computer-implemented method identifies the microsatellite set most informative for generating the classifier (107). Further validation or optimization steps can be performed by analyzing additional samples (e.g., from a database) of subjects who are known to have or not have the condition. (108) After 108, a computer-implemented method can be used to generate the final classifier (109).

[0053] In one aspect, this disclosure provides an improved computer-implemented method for identifying a set of microsatellites as markers (classifiers) for a disease. The method may further include comparing microsatellite loci from a first set of samples from subjects suffering from said disease with microsatellite loci from a second set of samples from subjects not suffering from said disease, thereby identifying an initial population (informative loci) of microsatellite loci.

[0054] In some cases, informative sites can be used directly as classifiers. In some cases, a classifier including informative sites can indicate the presence or absence of a subject's condition. In some cases, a classifier including informative sites can indicate an increased or decreased likelihood of a subject's condition developing. In some instances, a classifier including informative sites can indicate an increased or decreased likelihood of a subject benefiting from treatment, or an increased or decreased risk of adverse effects due to treatment. In some cases, a classifier including informative sites can indicate the responsiveness to treatment for a subject's condition. In some instances, a classifier including informative sites can indicate the prognosis of a subject's condition.

[0055] In some aspects, an initial population of microsatellite loci (informative loci) is used in a genetic algorithm executed as by a computer-implemented method. The method may include iteratively sorting a subset of the initial population of microsatellites by comparing microsatellite subsets from samples from subjects with the disease with microsatellites from samples from subjects without the disease. The method may include initialization, wherein an initial subset of the subset is randomly selected from the initial population of microsatellite loci. In some instances, a subset of the initial population of approximately 100 microsatellite loci is used throughout the genetic algorithm (optimization cycle), wherein the minimum number of microsatellites in the subset is 8, and the maximum number of microsatellites in the subset is 64. In some instances, the iterative sorting includes multiple optimization cycles, wherein the multiple optimization cycles include adding 10 new subsets of the initial population of microsatellites to subsets from previous optimization cycles. Seven of the 10 new subsets may be generated by randomly splitting and recombinating two subsets randomly selected from previous optimization cycles, and three of the 10 new subsets may be generated by randomly selecting microsatellites from the initial population of microsatellites. In some cases, the method includes sorting subsets during optimization cycles, where the 10 lowest-ranked subsets are discarded during the optimization cycle, thus maintaining a subset of 100 microsatellite populations throughout the optimization cycle. Genetic algorithms may include performing iterative sorting on all microsatellite combinations to identify the most informative microsatellite loci. Genetic algorithms can improve sensitivity and specificity by removing less informative microsatellite loci and selecting or weighting more informative microsatellite loci. In some cases, the condition identified by microsatellite loci through periodic optimization can indicate the presence or absence of a subject's health status, the increased or decreased likelihood of a subject's health status developing, the increased or decreased likelihood of a subject benefiting from treatment of a health status, the increased or decreased likelihood of a subject having an increased risk of adverse effects due to treatment of a health status, the subject's responsiveness to treatment of a health status, the prognosis of a subject's health status, or a combination thereof.

[0056] On the other hand, this disclosure provides an improved computer-implemented method comprising using multiple parameters to determine a classifier for a condition from samples of a subject, wherein each of the multiple parameters is a statistical measure of the correlation of each of multiple microsatellites from samples from subjects with the condition and / or from samples from subjects without said condition. In some cases, the multiple parameters include optimal weights, such as those determined by standard regression analysis and using genetic algorithms. In some cases, the classifier is determined by using minor allele frequency data. In some cases, the condition may indicate the presence or absence of a subject's health status, an increased or decreased likelihood of a subject's health status developing, an increased or decreased likelihood of a subject benefiting from treatment of the health status, an increased or decreased likelihood of a subject having an increased risk of adverse effects due to treatment of the health status, a subject's responsiveness to treatment of the health status, the prognosis of the subject's health status, or a combination thereof. In some cases, the health status is cancer, a neurological disorder, or a cardiovascular disease.

[0057] On the other hand, this disclosure provides a method for using a computer system to determine minor allele characteristics in a first sample from a subject, comparing the minor allele characteristics with a reference, and determining the subject's genomic age based on the comparison. A minor allele characteristic may be multiple minor alleles at a location, wherein the number of alleles is supported by at least one, at least two, at least three, or more next-generation sequencing reads. In some cases, a minor allele characteristic is a normalized total of minor allele reads to the total number of major allele reads at a location. The minor allele characteristics of a first sample from a subject can be compared with secondary minor allele characteristics of a second sample from the same subject to determine the rate of genomic aging.

[0058] This disclosure provides a pan-disease assay based on a classifier generated using microsatellite loci and optional minor allele information. In some cases, the pan-disease assay is a pan-cancer assay.

[0059] The terms “about” or “approximately” can mean within an acceptable range of error for a particular value as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, such as limitations of the measurement system. For example, “about” can mean within one or more standard deviations, depending on practice in a given value. “About” can mean + / - 10%, + / - 5%, + / - 2%, or + / - 1% of a value. Unless the context clearly specifies otherwise, the singular forms “a,” “an,” and “the” used in the specification and claims include plural references. For example, the term “nucleic acid” includes a plurality of nucleic acids, including mixtures thereof.

[0060] II. Microsatellite classifier method for identifying disease symptoms

[0061] This disclosure provides methods for microsatellite classifiers for identifying diseases, such as computer-implemented methods (e.g., see...). Figure 2 The method may include the presence or absence of a subject's health status, an increased or decreased likelihood of a subject's health status developing, an increased or decreased likelihood of a subject benefiting from treatment of a health status, an increased or decreased likelihood of a subject having an increased risk of adverse effects due to treatment of a health status, a subject's responsiveness to treatment of a health status, the prognosis of a subject's health status, or a combination thereof. The method may include identifying microsatellite loci (genotyping) in samples from subjects with and without the disease. The method may include identifying statistically informative microsatellite loci for the disease. The method may include developing a classification signature for the disease using statistically informative microsatellite loci. The classification signature can be validated and used to test samples from subjects.

[0062] A. Microsatellite locus genotyping

[0063] Methods for identifying microsatellite classifiers may include genotyping microsatellite loci in samples from subjects with and without the disease. In some cases, genotyping involves analyzing sequence information in a database. In others, genotyping involves obtaining samples and analyzing the nucleic acid molecules within them, for example, through next-generation sequencing.

[0064] 1. Sequence Information Database

[0065] In some cases, methods for identifying (e.g., genotyping) microsatellite loci may include analyzing sequence information from one or more databases. One or more databases may include sequence information (e.g., sequence reads) from subjects with a condition, such as those with cancer, or from nucleic acid samples from cancer cell lines. One or more databases may include reference sequences (e.g., the human genome or a portion thereof). One or more databases may include variant or polymorphic sequences from one or more subject populations.

[0066] One or more databases may include sequence information generated by high-throughput or next-generation sequencing. One or more databases may include data (e.g., sequence readout data) of sequences generated from whole-exome sequencing (WES), whole-genome sequencing (WGS), or a combination thereof, of samples from subjects. In some instances, one or more databases include sequence information (e.g., sequence readout information) generated from targeted sequencing. Targeted sequencing may include enrichment of targeted sequences from samples from subjects.

[0067] The database may include sequence information from the Cancer Genome Atlas (TCGA), such as exome data, for example, lung cancer exome data. The database may also originate from the 1000 Genomes Project.

[0068] 2. Sample

[0069] The sample can be a biological sample obtained from or derived from one or more subjects. The sample can be processed or graded to produce other samples, such as other biological samples. The samples described in this disclosure can include any material from which nucleic acid molecules can be obtained.

[0070] Samples can be obtained from subjects who have the disease. Samples can be obtained from subjects who have symptoms of the disease. Samples can be obtained from subjects who have the disease but do not have symptoms of the disease. Samples can be obtained from subjects who have never had the disease. Samples can be obtained from subjects who have cancer, are suspected of having cancer, or do not have or are not suspected of having cancer.

[0071] Samples may be obtained from or derived from human subjects. Samples may be stored under a variety of storage conditions prior to processing, such as different temperatures (e.g., at room temperature, under refrigerated or frozen conditions, at 25°C, at 4°C, at 18°C, at -20°C, or at -80°C) or different buffers (e.g., EDTA collection tubes, or cell-free DNA or RNA collection tubes).

[0072] Samples may be collected before and / or after treatment of subjects with cancer. Samples may be obtained from subjects during treatment or a treatment regimen. Multiple samples may be obtained from subjects to monitor the effects of treatment over time. Samples may be taken from subjects who are known to have or suspected of having cancer, where a final positive or negative diagnosis of cancer cannot be obtained through clinical trials. Samples may be taken from subjects suspected of having cancer. Samples may be taken from subjects experiencing unexplained symptoms, such as fatigue, nausea, weight loss, pain, weakness, or bleeding. Samples may be taken from subjects with explanatory symptoms. Samples may be taken from subjects at risk of developing cancer due to family history, age, high blood pressure or prehypertension, diabetes or prediabetes, overweight or obesity, environmental exposure, lifestyle risk factors (e.g., smoking, alcohol consumption, or drug use), or the presence of other risk factors.

[0073] The sample can be a biological sample from the subject. The sample can be whole blood, peripheral blood, plasma, serum, saliva, mucus, urine, semen, lymph, amniotic fluid, fecal extract, buccal swab, cells, or other bodily fluids or tissues, including tissues obtained through surgical biopsy or surgical excision. In some cases, the sample can be a cell line derived from the primary subject (e.g., a patient) or an archived subject (e.g., a patient) sample, such as a preserved sample, such as a formalin-fixed paraffin-embedded (FFPE) sample, or a fresh frozen sample. The sample (e.g., a biological sample) can be obtained from or derived from the subject using EDTA collection tubes, DNA or RNA collection tubes, or cell-free DNA or cell-free RNA collection tubes. The sample, such as a biological sample, can be obtained from a whole blood sample by grading. The sample, such as a biological sample or a derivative thereof, may include cells. The sample (e.g., a biological sample) can be a blood sample or a derivative thereof (e.g., blood or blood drops collected from a collection tube).

[0074] The sample may contain one or more analytes that can be measured. The sample may include one or more nucleic acid molecules. One or more nucleic acid molecules (or any nucleic acid molecules disclosed herein, including primers and probes) may be polymers of nucleotides of any length (e.g., deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs) or analogs thereof). Analogs may include non-naturally occurring bases, nucleotides attached to nucleotides other than naturally occurring phosphodiester bonds, or bases attached by bonds other than phosphodiester bonds. Nucleotide analogs include, for example, thiophosphates, dithiophosphates, triphosphates, aminophosphates, borate phosphates, methylphosphonates, chiral methylphosphonates, 2-O-methylribonucleotides, peptide-nucleic acids (PNAs), etc. The nucleic acid molecule may be deoxyribonucleic acid (DNA). DNA may be genomic DNA, viral DNA, mitochondrial DNA, plasmid DNA, amplified DNA, circular DNA, circulating DNA, cell-free DNA, or exosome DNA. In some instances, the DNA is single-stranded DNA (ssDNA), double-stranded DNA, denatured double-stranded DNA, synthetic DNA, and combinations thereof. Circular DNA may be cleaved or fragmented. The DNA may include coding or non-coding regions of a gene or gene segment of interest, loci defined by linkage analysis, exons, or introns. The DNA may be complementary DNA (cDNA). The nucleic acid molecule may be a recombinant nucleic acid, branched nucleic acid, plasmid, vector, or isolated DNA. The nucleic acid molecule may include one or more modified nucleotides, such as methylated nucleotides or nucleotide analogs. Modifications to the nucleotide structure may be made before or after the assembly of the nucleic acid molecule. The nucleotide sequence of the nucleic acid molecule may be interrupted by non-nucleotide components. The nucleic acid molecule may be further modified after polymerization, such as by conjugation or binding to a reporter agent.

[0075] Nucleic acid molecules may include loci, genetic loci, or genomic regions, which can be identified by their location within the genome or chromosome. In some examples, a locus may be referred to by a gene name and includes coding and non-coding regions associated with a physical region of the nucleic acid. A gene may include coding regions (exons), non-coding regions (introns), transcriptional control or other regulatory regions, and promoters. In another example, a genomic region may incorporate introns or exons or intron / exon boundaries within a named gene.

[0076] In some instances, nucleic acid molecules include ribonucleic acid (RNA). The RNA can be fragmented RNA. The RNA can be degraded RNA. The RNA can be microRNA or a portion thereof. The RNA can be selected from the following RNA molecules or fragmented RNA molecules (RNA fragments): microRNA (miRNA), premiRNA, pri-miRNA, messenger RNA (mRNA), premRNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), viral RNA, viroid RNA, circular RNA (circRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), pretRNA, long noncoding RNA (lncRNA), small nuclear RNA (snRNA), circulating RNA, cell-free RNA, exosome RNA, vector-expressed RNA, RNA transcripts, synthetic RNA, ribozymes, cell-free RNA, and combinations thereof.

[0077] In some cases, the sample includes cell-free nucleic acid molecules. Cell-free nucleic acid molecules can include, for example, all non-encapsulated nucleic acid molecules derived from the subject's bodily fluids. Cell-free nucleic acid (cfNA) molecules can be nucleic acids not contained in cells in a biological sample (e.g., cell-free RNA (cfRNA) molecules or cell-free DNA (cfDNA) molecules). cfDNA molecules can circulate freely in bodily fluids (such as in the bloodstream). Cell-free DNA molecules can be circulating tumor DNA, such as cfDNA derived from a tumor.

[0078] The sample can be a cell-free sample. A cell-free sample can be a biological sample that contains essentially no intact cells. A cell-free sample can be a biological sample that contains essentially no cells in itself, or it can be derived from a sample from which cells have been removed. Examples of cell-free samples include samples derived from blood, such as serum or plasma; urine; or samples derived from other sources, such as semen, sputum, feces, catheter exudate, lymph, or recovered lavage fluid.

[0079] Samples may include germline nucleic acid molecules (e.g., nucleic acids from non-disease cells or tissues, such as tumors). Samples may include nucleic acid molecules from tumors. In some cases, samples may include both germline nucleic acid molecules (e.g., from non-disease tissues) and nucleic acid molecules from diseased tissues (e.g., tumors).

[0080] The sample may include a target nucleic acid molecule. The target nucleic acid molecule may be a nucleic acid molecule having a nucleotide sequence, the presence, quantity and / or sequence of said nucleotide sequence, or variations thereof, need to be determined.

[0081] Nucleic acid molecules (e.g., RNA or DNA) can be extracted from a sample, for example using the Qiagen QIAmp DNA Blood Mini Kit, MP Biomedical's FastDNA Kit, or Norgen Biotek's Cell-Free Biological DNA Isolation Kit. These extraction methods can extract all RNA or DNA molecules from a sample. Alternatively, extraction methods can selectively extract a subset of RNA or DNA molecules from a sample. The RNA molecules extracted from the sample can be converted into DNA molecules via reverse transcription (RT). Reverse transcription can be performed by the action of reverse transcriptase to generate deoxyribonucleic acid (RNA) from a ribonucleic acid (RNA) template.

[0082] For example, the quality of extracted nucleic acids can be analyzed using the BIOANALYZER or NANODROP systems.

[0083] The subject can be a person or an individual. The subject can be a patient. The subject can be someone who has or is suspected of having cancer. The subject may exhibit symptoms indicative of a health or physiological state or condition. The subject may be asymptomatic in relation to a health or physiological state or condition. The subjects described herein can include mammals, including any member of the mammalian class: humans, non-human primates (such as chimpanzees, and other apes and monkeys); livestock, such as cattle, horses, sheep, goats, pigs; domesticated animals (such as rabbits, dogs, and cats); laboratory animals including rodents (such as rats, mice, and guinea pigs), etc. In one respect, the mammal is the human being.

[0084] Processing samples obtained from subjects may include placing the samples under conditions sufficient to isolate, enrich, or extract multiple nucleic acid molecules, and measuring multiple nucleic acid molecules to generate a dataset.

[0085] Samples from subjects can be analyzed to genotype one or more microsatellites. As described herein, a microsatellite, microsatellite locus, or microsatellite region can refer to a tandem repeat of 1 to 6 nucleotides in a nucleotide sequence. In some cases, microsatellites include tandem repeat sequences of more than 6 nucleotides. One or more microsatellites can be found upstream of an exon, downstream of an exon, within an exon, in an intergenic sequence, within an intron, in a region spanning exons and introns, in a 3' untranslated region (UTR), a 5' UTR, or any other region in the genome. In some instances, the pattern of microsatellites in a sample differs from that in a reference sample. Differences in microsatellite patterns can include single nucleotide polymorphisms (SNPs), the percentage of SNPs, insertions / deletions (the ratio of insertions, deletions, and combinations thereof), or the ratio of insertions / deletions to SNPs. In some instances, patterns of microsatellite difference include haplotypes, such as the percentage of homozygosity, heterozygosity, or minor alleles at a given locus. When the differential patterns of microsatellites are located in exon regions, the differentials can include non-synonymous SNPs, synonymous SNPs, frameshifted insertions / deletions, non-frameshifted insertions / deletions, stop gain, and stop loss. Samples may match, for example, age, sex, or race (e.g., Caucasian, African American, Hispanic American). In some cases, samples may not match. In some cases, samples may be accompanied by additional clinical metadata, including, for example, health status, cancer, cardiac or neurological condition, treatment status or response, or disease stage. Clinical metadata can be correlated with microsatellites to determine whether the microsatellites are informative relative to the clinical metadata.

[0086] The identity (e.g., genotype) of one or more microsatellites can be obtained by any available method or technology, including next-generation sequencing, high-throughput sequencing, synthetic sequencing, pyrosequencing, classical Sanger sequencing, ligation sequencing, synthetic sequencing, hybridization sequencing, RNA-Seq (Illumina), Illumina sequencing (using reversibly terminated nucleotides), paired-end sequencing, digital gene expression (Helicos), single-molecule sequencing (e.g., synthetic single-molecule sequencing (SMSS) (Helicos)), ion flooding (semiconductor) sequencing (Life Technologies / Thermo-Fisher), massively parallel sequencing, cloned single-molecule array (Solexa), nanopore sequencing, Pacific Biosciences SMRT sequencing, shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and any other sequencing method.

[0087] Next-generation sequencing can include sample multiplexing. Sample multiplexing can involve at least 12, 24, 48, 96, 192, 384, 768, or 1536 samples. Sequencing depth can range from approximately 1x to approximately 10x, approximately 10x to approximately 100x, approximately 100x to approximately 500x, or approximately 500x to approximately 1000x.

[0088] Sequencing depth can be at least, at most, or approximately 1x, 5x, 10x, 50x, 100x, 200x, 250x, 300x, 400x, or 500x. Base call consensus accuracy can be at least 95%, 96%, 97%, 98%, 99%, or more than approximately 99%. Quality score can be at least Q10 (e.g., error rate less than 1:10, inferred basic call accuracy greater than 90%), above Q20 (e.g., error rate less than 1:100, inferred basic call accuracy greater than 99%), above Q30 (e.g., error rate less than 1:1000, inferred basic call accuracy greater than 99.9%), above Q40 (e.g., error rate less than 1:10,000, inferred basic call accuracy greater than 99.99%), or above Q50 (e.g., error rate less than 1:100,000, inferred base detection accuracy greater than 99.999%). Assembly methods can generate at least 95%, 96%, 97%, 98%, or 99% accuracy for calling microsatellite genotypes in next-generation sequencing datasets.

[0089] After sequencing nucleic acid molecules, appropriate bioinformatics processing can be performed on the sequence reads. For example, the sequence reads can be aligned with one or more reference genomes (e.g., the genomes of one or more species, such as the human genome). Aligned sequence reads can be quantified at one or more sites (e.g., one or more microsatellite sites).

[0090] In some aspects, identifying (e.g., genotyping) one or more microsatellites involves amplifying the nucleotide sequences of one or more microsatellite sites, for example, by performing a polymerase chain reaction (PCR), for example, using primers, such as specific primers, flanking one or more microsatellite sites, and evaluating the amplified fragments, for example, by capillary electrophoresis or sequencing. PCR can be quantitative PCR (qPCR), digital PCR, or reverse transcriptase PCR. Amplification or amplification can increase the size or number of nucleic acid molecules. The amplified nucleic acid molecules can be single-stranded or double-stranded. Amplification can include generating one or more copies of the nucleic acid molecule or the amplification product. For example, amplification can be performed by extension (e.g., primer extension) or ligation. Amplification can include performing a primer extension reaction to generate a strand complementary to the single-stranded nucleic acid molecule, and in some cases, generating one or more copies of the single-stranded and / or single-stranded nucleic acid molecule.

[0091] Amplification of nucleic acid molecules (e.g., nucleic acid molecules including one or more microsatellite sites) can be performed, for example, by any of the following nucleic acid amplification methods: loop-mediated isothermal amplification (LAMP), sequence-based amplification (NASBA), self-sustaining sequence replication (3SR), rolling circle amplification (RCA), recombinase polymerase amplification (RPA), multiple displacement amplification (MDA), helicase-dependent amplification (HDA), strand displacement amplification (SDA), nicking enzyme amplification reaction (NEAR), exponential amplification reaction (EXPAR), polymerase helical reaction (PSR), isothermal multiple displacement amplification (IMDA), branching amplification method (RAM), single primer isothermal amplification (SPIA), signal-mediated RNA amplification technology (smart), beacon-assisted detection amplification (BADAMP), hinge-initiated primer-dependent nucleic acid amplification (HIP), SMART amplification process (SmartAmp), hybridization chain reaction (HCR), a bottom-mediated strand displacement (TMSD), ligase chain reaction, digital PCR (dPCR), droplet digital PCR (ddPCR), or transcription-mediated amplification. Amplification can involve multiplex amplification, for example, using AMPLISEQ. In some cases, RNA is converted to cDNA via reverse transcription before amplification. Assay readings can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or their normalized values. Other assays that can be used in the methods provided herein include immunoassays, electrochemical assays, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assays, molecular reverse probes, CRISPR / Cas-based assays (e.g., CRISPR-genotyping PCR (ctPCR), specific high-sensitivity enzyme reporter unlocking (SHERLOCK), DNA endonuclease-targeted CRISPR trans-reporter (DETECTR), CRISPR-mediated simulated multiple event recording device (CAMERA)), and laser transmission spectroscopy (LTS).

[0092] Multiplex amplification can include amplifying approximately 10 to 50 targets, approximately 50 to 100 targets, approximately 100 to 500 targets, or approximately 500 to 1000 targets. Adaptors (e.g., universal adaptors) can be added (e.g., ligated) to nucleic acid molecules to facilitate amplification and / or sequencing, for example, on the ILLUMINA sequencing platform. Universal primers can bind to universal adaptors for amplification.

[0093] Multiple samples can be analyzed, and each multiplexed sample can be barcoded. RNA or DNA molecules isolated or extracted from the samples can be labeled, for example, with identifiable markers, to allow multiplexing of multiple samples. Any number of RNA or DNA samples can be multiplexed. For example, a multiplex reaction can contain RNA or DNA from at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more initial samples. For example, multiple samples can be labeled with sample barcodes so that each DNA molecule can be traced back to the sample (and subject) from which the DNA molecule originated. These tags can be attached to RNA or DNA molecules by incorporation or by PCR amplification using primers.

[0094] In some cases, a decoy set (e.g., a hybridization probe, such as SURESELECT or SEQCAP) is used to acquire a target, such as a target nucleic acid molecule. The target may include RNA and / or DNA. The length of the hybridization probe may be at least 15, 25, 50, 75, 100, 120, or 150 bases. The length of the hybridization probe may be 15 to 50 bases, 50 to 100 bases, or 100 to 150 bases. The probe may be a nucleic acid molecule (e.g., RNA or DNA) that is sequence complementary to a nucleic acid sequence (e.g., RNA or DNA) at one or more sites (e.g., one or more microsatellites). Determining a sample using a probe selective for one or more sites (e.g., one or more microsatellites) may include using array hybridization (e.g., microarray-based), polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing).

[0095] In some aspects, analyzing nucleic acid molecules involves performing next-generation sequencing. In some cases, microsatellite sequencing can be performed directly, for example, without amplification. Next-generation sequencing methods can include whole-genome, whole-exome, and partial genome or exome sequencing. Next-generation sequencing methods can be used for targeted sequences, enriched sequences, or combinations thereof.

[0096] In some instances, enrichment is performed using an enrichment kit prior to sequencing and downstream analysis. In other cases, enrichment is performed using an enrichment kit to enrich microsatellite loci validated by genetic algorithms. Using an enrichment kit can increase the number of callable allelopathic types or genotypes in reads and can increase the ability to analyze a larger percentage of informative loci for a given sample. Enrichment kits may include enrichment arrays or probes that hybridize to the target sequence of the microsatellite and flanking sequences on either side or both sides of the microsatellite. In some cases, the use of enrichment increases the number of callable genotypes by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or more compared to the number of callable genotypes available without using an enrichment kit. In some instances, the use of an enrichment kit increases the number of callable genotypes by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times compared to the number of callable genotypes available without using an enrichment kit. In some respects, the enrichment kits disclosed herein include compositions that can be used to perform the methods described herein.

[0097] 3. Genotyping Algorithm

[0098] Algorithms can be used to genotype microsatellites. These algorithms can use, for example, Bayesian model selection guided by an empirically derived error model, or discretized Gaussian mixtures (e.g., GenoTan). For example, the algorithm could be Repeatseq. Dynamic programming-based or heuristic methods can be used for microsatellite genotyping. Other tools for microsatellite genotyping include PHOBOS, MISA, Tandem Repeats Finder, FullSSR, or bMSISEA.

[0099] B. Identifying informative microsatellites

[0100] Identifying informative microsatellites may include identifying a first set of microsatellite loci from samples from subjects with the disease and a second set of microsatellite loci from samples from subjects without the disease. In some cases, the second set of microsatellite loci may be obtained from a database of reference sequences.

[0101] 1. Statistical data

[0102] Differences between the first and second groups of microsatellite loci can be detected, and statistical comparisons can be made using one or more statistical tests (such as t-tests, Z-tests, ANOVA, regression analysis, Mann-Whitney-Wilkcockian test, chi-square test, correlation, Fisher's exact test, Bonferroni correction, and Benjamini-Hochberg test). In some cases, the generalized Fisher's exact test is used to quantify statistical differences. In other cases, Benjamini-Hochberg multiple test correction is applied to control for false discovery rates.

[0103] 2. Microsatellite filtering

[0104] If, for example, samples from subjects with the disease and samples from subjects without the disease do not match the factors, microsatellites can be filtered to control for any number of factors, such as age, ethnicity, sex, sequencing protocol (e.g., WSG, WES, or targeted sequencing). Microsatellites with potential bias can be excluded from subsequent analyses. Additional filters for microsatellite filtering may include the length of the microsatellite repetitive motif, the total length of the microsatellites (e.g., the copy number of the motif), the sequence of the motif (e.g., using only those with high GC content), and the purity of the microsatellite, for example, if it contains any bases that could disrupt a perfect set of copies of the motif. In some instances, microsatellites can be filtered by their location in the genome (e.g., exome, introns, intergenic regions, or untranslated regions). Filtering may include filtering by the genes or functional elements that are close to the microsatellite.

[0105] 3. Scoring the samples

[0106] Statistical tests can generate receiver operating characteristic (ROC) curves, where the area under the ROC curve is called the area under the curve (AUC). The AUC can be determined to assess the accuracy of comparisons of the groups of microsatellite loci. A larger AUC indicates higher accuracy in the correlation or association between the disease and the differences between the first and second groups of microsatellite loci. ROC curves can sensitively (e.g., true positives) determine the ratio of association or correlation between the disease and the differences between the first and second groups of microsatellite loci, and specifically (e.g., true negatives). Sensitivity, also known as the true positive rate, recall, or detection probability, measures the proportion of actual positives correctly identified as the presence or absence of a certain disease. Sensitivity can be quantified by dividing the number of true positives by the sum of the number of true positives and false negatives. Specificity, also known as the true negative rate, measures the proportion of actual negatives correctly identified as the presence or absence of a disease. Specificity can be quantified by dividing the number of true negatives by the sum of the number of true negatives and false positives.

[0107] In some instances, the statistically significant correlation or association between the condition and a first group of microsatellite loci (different from the second group) has a statistical accuracy of at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. In some cases, the statistically significant correlation or association between the condition and a first group of microsatellite loci (different from the second group) has a statistical specificity of at least 0.70, 0.80, 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99, and a statistical sensitivity of at least 0.70, 0.80, 0.85, or 0.99.

[0108] In some instances, identifying informative microsatellites involves identifying a first set of microsatellite loci from a database comprising nucleic acid sequences obtained from subjects with the disease, such as cancer type sequences from the Cancer Genome Atlas Project (TCGA), and a second set of microsatellite loci from a reference database (e.g., the HG19 or 1000 Genomes Project). Cancer types (e.g., breast cancer) can be subtypes based on factors such as stage, morphology, histology, gene expression, receptor profile, mutation profile, invasiveness, prognosis, malignancy, etc. Cancer types and subtypes can be applied at a more refined level, for example, to distinguish cancers or cancer subtypes of a specific tissue type, e.g., defined by mutation profiles or gene expression. Cancer stage can refer to the classification of cancer types based on histological and pathological features associated with disease progression. In some instances, the set of microsatellite loci is obtained from a database comprising nucleic acid sequences including nucleotide variations or polymorphisms. In some cases, the first set of microsatellite loci is obtained from samples of the disease and compared with a second set of microsatellite loci obtained from a database.

[0109] 4. Symptoms

[0110] In some cases, conditions associated with or related to differences in the aforementioned microsatellite loci can indicate the presence or absence of a subject's health status, the increased or decreased likelihood of a subject's health status developing, the increased or decreased likelihood of a subject benefiting from treatment of the health status, the increased or decreased likelihood of a subject having an increased risk of adverse effects due to treatment of the health status, the subject's responsiveness to treatment of the health status, the prognosis of the subject's health status, or a combination thereof. In some cases, the health status is cancer. In some cases, the cancer is solid or hematologic malignancy. In some cases, the cancer is metastatic, recurrent, or refractory. Cancers that may be associated with or linked to different groups of microsatellite loci include acute myeloid leukemia (LAML or AML), acute lymphoblastic leukemia (ALL), adrenocortical carcinoma (ACC), bladder urothelial carcinoma (BLCA), brainstem glioma, low-grade glioma (LGG), brain tumors, breast cancer (BRCA), bronchial tumors, Burkitt lymphoma, cancer of unknown primary site, carcinoid tumors, cancer of unknown primary site, atypical teratoid / rhabdoid tumors of the central nervous system, embryonal tumors of the central nervous system, and cervical squamous cell carcinoma.Cervical adenocarcinoma (CESC), childhood cancer, cholangiocarcinoma (CHOL), chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorders, colon (adenocarcinoma) cancer (COAD), colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, endocrine islet cell tumor, endometrial cancer, ependymoma, esophageal cancer (ESCA), sensory neuroblastoma, Ewing sarcoma, extracranial germ cell tumor, gonadal germ cell tumor, extrahepatic bile duct cancer, gallbladder cancer, gastric (stomach) cancer, gastrointestinal head and neck cancer (HNSD), cardia cancer, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma Pancreatic islet cell tumor, Kaposi's sarcoma, renal cell carcinoma, Langerhans cell histiocytosis, laryngeal cancer, lip cancer, liver cancer, lymphadenopathy, diffuse large B-cell lymphoma [DLBCL], malignant fibrous histiocytoma, bone cancer, medulloblastoma, medullary epithelioma, melanoma, Merkel cell carcinoma, Merkel cell skin cancer, mesothelioma (MESO), metastatic squamous neck cancer with occult primary cancer, oral cancer, multiple endocrine tumor syndrome, multiple myeloma, multiple myeloma nasal cavity carcinoma, nasopharyngeal carcinoma, neuroblastoma, non-Hodgkin lymphoma, non-melanoma skin cancer, non-small cell lung cancer, oral cancer, oral cavity cancer, oropharyngeal cancer. Osteosarcoma, other brain and spinal cord tumors, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumors, low-potency ovarian tumors, pancreatic cancer, papilloma, paranasal sinus cancer, parathyroid cancer, pelvic cancer, penile cancer, pharyngeal cancer, pheochromocytoma and paraganglioma (PCPG), intermediately differentiated pineal parenchymal tumors, pineal blastoma, pituitary tumors, plasmacytoma / tumor primary central nervous system (CNS) lymphoma, primary hepatocellular carcinoma, prostate cancer such as prostate adenocarcinoma (PRAD), rectal cancer, kidney cancer, renal cell (kidney) carcinoma, renal cell carcinoma, respiratory tract cancer, retinoblastoma, rhabdomyosarcoma, salivary glands Cancer, sarcoma (SARC), Sezary syndrome, cutaneous malignant melanoma (SKCM), small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, squamous neck cancer, gastric (stomach) cancer, supratentorial primitive neuroexon tumor, T-cell lymphoma, testicular cancer, testicular germ cell tumor (TGCT), laryngeal cancer thymoma (thymus), thyroid cancer (THCA), transitional cell carcinoma, renal pelvis and ureter transitional cell carcinoma, trophoblastic tumor, ureteral cancer, urethral cancer, uterine cancer, uveal melanoma (UVM), vaginal cancer, vulvar cancer, Waldenström macroglobulinemia or Wilm's tumor. In some respects,Types of cancer include acute lymphoblastic leukemia, acute myeloid leukemia, bladder cancer, breast cancer, brain cancer, cervical adenocarcinoma, bile duct cancer, colon cancer, colorectal cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, glioma, glioblastoma, head and neck cancer, kidney cancer, liver cancer, lung cancer, lymphoid tumor formation, melanoma, medullary tumor formation, ovarian cancer, pancreatic cancer, pheochromocytoma and paraganglioma, prostate cancer, rectal cancer, squamous cell carcinoma, testicular cancer, stomach cancer, or thyroid cancer.

[0111] In some cases, the healthy state is lung cancer or a subtype of lung cancer. Lung cancers that can be associated with or linked to different groups of microsatellite loci include non-small cell lung cancer (NSCLC) (e.g., lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), and large cell carcinoma), small cell lung cancer (SCLC), and lung carcinoid tumors.

[0112] In some cases, a health status is a neurological disorder. Examples of neurological disorders that may be associated with or related to differences in the microsatellite loci of the aforementioned groups include myotonic dystrophy, Fragile X-associated tremor / ataxia syndrome, spinocerebellar ataxia, Kennedy's disease, Huntington's disease, spinobulbar muscular atrophy, progressive myoclonic epilepsy 1 (Unverricht–Lundborg disease), Fragile X syndrome, Fragile XE syndrome, dentate nucleus-globus pallidus atrophy, Frederick's ataxia, oculopharyngeal muscular dystrophy, Fragile X-associated primary ovarian insufficiency, Huntington's disease-like 2, C9ORF72-associated frontotemporal dementia, and amyotrophic lateral sclerosis. A health status can also be autism.

[0113] In some cases, the health status is inflammatory bowel disease (IBD), which can include gastrointestinal disorders of the gastrointestinal tract. Non-limiting examples of IBD include Crohn's disease (CD), ulcerative colitis (UC), indeterminate colitis (IC), microscopic colitis, shunt colitis, Behçet's disease, and other indeterminate forms of IBD. In some instances, IBD includes fibrosis, fibrosis, strictures and / or penetrating disease, obstructive disease or refractory disease (e.g., mrUC, refractory CD), perianal CD, or other complex forms of IBD.

[0114] In some instances, health status is cardiovascular disease, which can include coronary artery disease (CAD), rheumatic heart disease, congenital heart disease, cardiomyopathy, cardiac tumors, vascular tumors, valvular heart disease, endocardial disease, stroke, aortic aneurysm, peripheral artery disease, deep vein thrombosis (DVT), or pulmonary embolism.

[0115] In some cases, a healthy state is a metabolic disease or disorder, which may include acid-base imbalance, metabolic brain disease, calcium metabolism disorder, DNA repair defect disorder, glucose metabolism disorder, hyperlactatemia, iron metabolism disorder, lipid metabolism disorder, malabsorption syndrome, metabolic syndrome X, congenital metabolic error, mitochondrial disease, phosphorus metabolism disorder, porphyria, protein deposition defect, metabolic skin disease, wasting syndrome, or electrolyte imbalance.

[0116] In some cases, health status is an autoimmune disease or disorder, which can include achalasia, Addison's disease, adult-onset Still's disease, agammaglobulinemia, alopecia areata, amyloidosis, ankylosing spondylitis, anti-GBM / anti-TBM nephritis, antiphospholipid syndrome, autoimmune angioedema, autoimmune autonomic disorders, autoimmune encephalomyelitis, autoimmune hepatitis, autoimmune inner ear disease (AIED), autoimmune myocarditis, autoimmune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune retinopathy, autoimmune urticaria, axonal and neuronal neuropathy (AMAN), Baló disease, and Behcet's chronic inflammatory demyelinating polyneuropathy (CI). DP), chronic relapsing multifocal osteomyelitis (CRMO), Chur-Strauss syndrome (CSS) or eosinophilic granulomatosis (EGPA), cicatricial pemphigoid, Cogan syndrome, cold agglutinin disease, congenital heart block, Coxsackie myocarditis, CREST syndrome, Crohn's disease, herpetic dermatitis, dermatomyositis, Devic disease (neuromyelitis optica), discoid lupus, Dressler syndrome, endometriosis, eosinophilic esophagitis (EoE), eosinophilic fasciitis, erythema nodosum allergic purpura (HSP), herpetic pregnancy or pemphigoid pregnancy (PG), hidradenitis suppurativa (HS) (acne inversion), hypoglobulinemia, IgA nephropathy, IgG4-associated sclerosis, immune Thrombocytopenic purpura (ITP), inclusion body myositis (IBM), interstitial cystitis (ic), juvenile arthritis, juvenile diabetes (type 1 diabetes), juvenile myositis (JM), Kawasaki disease, Lambert-Eton syndrome, leukocytic desquamative vasculitis, lichen planus, lichen sclerosus, woody conjunctivitis, linear multiple sclerosis, myasthenia gravis, myositis, narcolepsy, neonatal lupus, neuromyelitis optica, neutropenia, ocular cicatricial pemphigoid, PPT neuritis, palindromic rheumatism (PR), PANDAS, paraneoplastic cerebellar degeneration (PCD), paroxysmal nocturnal hemoglobinuria (PNH), Parry-Romberg syndrome, plaques (peripheral uveitis), Pars... Onage-Turner syndrome, pemphigus, peripheral neuropathy, perivenous encephalomyelitis, pernicious anemia (PA), POEMS syndrome, polyarteritis nodosa, type I and II polyglandular syndrome, Raynaud's phenomenon, reactive arthritis, reflex sympathetic dystrophy, relapsing polychondritis, restless legs syndrome (RLS), retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, sarcoidosis, Schmidt syndrome, scleritis, scleroderma, Sjögren's syndrome, sperm and testicular autoimmunity, stiff-person syndrome (SPS), subacute bacterial endocarditis (SBE), Sussac syndrome, sympathetic ophthalmia (SO), aortitis, temporal arteritis / giant cell arteritis, thrombocytopenic purpura (TTP).Tolosa-Hunter syndrome (THS), transverse myelitis, type 1 diabetes, ulcerative colitis (UC), undifferentiated connective tissue disease (UCTD), uveitis, vasculitis, vitiligo, or Vogt-Koyanagi-Harada disease.

[0117] C. Develop category signatures

[0118] This disclosure provides a computer-implemented method for generating a classifier for a condition from samples from subjects (see, for example, see...). Figure 2 and Figure 3 An informative list of microsatellite loci can be generated by statistically analyzing samples obtained or derived from a first group of subjects with the disease and / or samples obtained or derived from a second group of subjects who have never had the disease (e.g., cancers such as lung cancer). DNA sequencing from both groups of samples can be performed on multiple platforms. In some cases, targeted sequencing is performed with enrichment of certain targets. The quality of the sequencing results can then be analyzed and mapped to reveal differences between cancer samples and controls or references. This difference can then be analyzed using computer-implemented methods to generate a classifier. The classifier can be further optimized and validated using additional samples obtained or derived from subjects with the disease and / or samples obtained or derived from subjects who have never had the disease. In some respects, lists of informative genetic markers other than microsatellites can be generated using these methods for developing taxonomic signatures.

[0119] The symptom can indicate the presence or absence of a subject's health status. In some cases, the symptom indicates an increased or decreased likelihood of the subject's health status developing. In some instances, the symptom can indicate an increased or decreased likelihood of the subject benefiting from treatment, or an increased or decreased likelihood of the subject having an increased risk of adverse effects due to treatment (the classifier of the symptom can serve as a companion diagnostic to the treatment). In some cases, the symptom can indicate the subject's responsiveness to treatment of their health status. In some instances, the symptom indicates the prognosis of the subject's health status. In some cases, the classifier can be, for example, a value of quantity. For example, the value can indicate an increased or decreased likelihood (e.g., a probability value between 0 and 1). The value of the classifier (e.g., quantity) can be compared to a threshold (e.g., quantity). In some instances, the distance between the classifier value and the threshold can indicate an increased confidence or probability that the subject has or does not have the symptom. In some cases, a call is made when the standard deviation of the classifier value from the threshold is approximately 0.5, 1, 1.5, 2, 2.5, 3, or greater. Figure 24 ).

[0120] Computer-implemented methods for generating classifiers can perform processing, combination, statistical evaluation, or further analysis of the results, or any combination thereof. Computer-implemented methods can include supervised or unsupervised learning methods, including support vector machines (SVMs), neural networks, random forests, clustering algorithms (or software modules), gradient boosting, linear regression, logistic regression, and / or decision trees. Supervised learning algorithms can be algorithms that rely on using a set of labeled, paired training data examples to infer relationships between input and output data. Unsupervised learning algorithms can be algorithms used to derive inferences from training datasets to output data. Unsupervised learning algorithms can include clustering analysis, which can be used for exploratory data analysis to discover hidden patterns or groupings in process data. An example of an unsupervised learning method is principal component analysis. Principal component analysis can include reducing the dimensionality of a set of one or more variables. The dimension of a given set of variables can be at least 1, 5, 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or greater than 1800. The dimension of a given set of variables can be at most 1800, 1600, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less than 10.

[0121] Computer-implemented methods may include performing statistical techniques. In some instances, statistical techniques may include linear regression, classification, resampling methods, subset selection, shrinkage, dimensionality reduction, nonlinear models, tree-based methods, support vector machines, unsupervised learning, or any combination thereof.

[0122] Linear regression is a method of predicting a target variable by fitting the best linear relationship between the dependent and independent variables. The best fit can correspond to least squares, minimizing the sum of all distances between the shape at each point and the actual observed value. Linear regression can include simple linear regression and multiple linear regression. Simple linear regression can use a single independent variable to predict the dependent variable. Multiple linear regression can use more than one independent variable to predict the dependent variable by fitting the best linear relationship.

[0123] Classification can be a data mining technique that assigns classes to data collections in order to achieve accurate prediction and analysis. Classification techniques can include logistic regression and discriminant analysis. Logistic regression can be used when the dependent variable is dichotomous (binary). Logistic regression can be used to discover and describe the relationship between a binary variable and one or more independent variables at the nominal, ordinal, interval, or ratio level. Resampling can be a method that includes drawing repeated samples from an original data sample. In some cases, resampling may not involve using a common distribution table to calculate approximate probability values. Resampling can generate a unique sampling distribution based on actual data. In some cases, resampling can use experimental methods rather than analytical methods to generate a unique sample distribution. Resampling techniques can include bootstrapping and cross-validation. Bootstrapping can be performed by sampling with replacement from the original data and using the "unselected" data points as test cases. Cross-validation can be performed by dividing the training data into multiple parts.

[0124] Subset selection can identify subsets of predictor variables that are related to the response. Subset selection can include best subset selection, forward stepwise selection, backward stepwise selection, hybrid methods, or any combination thereof. In some instances, shrinkage fitting involves fitting a model that includes all the predictor variables, but the estimated coefficients are shrunk towards zero relative to the least squares estimates. This shrinkage can reduce variance. Shrinkage can include ridge regression and the lasso. Dimension reduction can simplify the problem of estimating n + 1 coefficients to a simpler problem of estimating m + 1 coefficients, where m < n. It can be obtained by computing n different linear combinations or projections of the variables. Then, these n projections can be used as predictor variables to fit a linear regression model, for example, by least squares. Dimension reduction can include principal component regression and partial least squares. Principal component regression can be used to derive a set of low-dimensional features from a large set of variables. The principal components used in principal component regression can capture the maximum variance in the data using linear combinations of the data in subsequent orthogonal directions. Partial least squares can be used as a supervised alternative to principal component regression because partial least squares can utilize the response variable to identify new features.

[0125] Nonlinear regression can be a form of regression analysis in which the observed data is modeled by a function that is a nonlinear combination of the model parameters and depends on one or more independent variables. Nonlinear regression can include step functions, piecewise functions, splines, generalized additive models, or any combination thereof.

[0126] Tree-based methods can be used for regression and classification problems. Regression and classification problems may involve hierarchically dividing or segmenting the space of predictor variables into many simple regions. Tree-based methods can include bagging, boosting, random forests, or any combination thereof. Bagging reduces the variance in predictions by generating multiple steps of the same size as the original data, using repeated combinations to generate additional training data from the original dataset. Boosting can compute the output using several different models and then average the results using a weighted averaging method. The random forest algorithm can draw random bootstrap samples from the training set. Support Vector Machines (SVMs) can be used for classification techniques. SVMs can involve finding a hyperplane that best separates two classes of points with the maximum margin. SVMs can constrain optimization problems so that maximizing the margin is constrained by its ability to perfectly classify data.

[0127] Unsupervised methods are methods that derive inferences from a dataset containing input data with unlabeled responses. Unsupervised methods can include clustering, principal component analysis, k-means clustering, hierarchical clustering, or any combination thereof.

[0128] 1. Genetic Algorithm

[0129] In some aspects, computer-implemented methods for generating classifiers include the use of genetic algorithms. The method may include generating an initial population (informative loci) of microsatellite loci that are disease-related or associated with the disease by identifying microsatellite loci that differ from those in samples without the disease from samples with the disease. The genetic algorithm may be used to determine a classification signature based on the informative loci. The genetic algorithm may select the subset of microsatellite loci with the highest informativeness to include in the final classifier. The genetic algorithm may assign weights to each subset. The weighting may be combined with other weighting schemes, such as proportionality to the relative risk of each microsatellite locus. Each subset of microsatellites may be iteratively ranked based on its relevance or association with the disease. The subset of the initial population of microsatellite loci may then be optimized by comparing the initial population with additional samples obtained or derived from subjects with and / or without the disease. In some cases, an initial population of approximately 100 subsets is used in the optimization. In some cases, an initial population of at least 100, 200, 300, 400, or 500 subsets is used in the optimization. In some instances, optimization includes at least one cycle comparing approximately 100 subsets with another sample. In some instances, optimization includes multiple cycles comparing approximately 100 subsets with another sample. Each subset may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 microsatellites.

[0130] Iterative ranking can be performed at the end of each cycle. In some cases, iterative ranking involves performing statistical analysis on subsets to perform receiver operating characteristic (ROC) analysis, thereby obtaining accuracy, sensitivity, and specificity in determining the presence or absence of the symptom in other samples. A predetermined number (e.g., 10) of subsets that perform the worst or rank lowest in indicating the presence or absence of the symptom can be identified and discarded. To maintain a constant number of subsets before the start of each optimization cycle, new subsets can be added to the subset group. In some cases, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more new subsets are generated by randomly splitting and recombining two randomly selected subsets from previous optimization cycles. In some instances, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more new subsets are randomly selected from previous optimization cycles. In some instances of adding 10 new subsets, 3 were generated by randomly splitting and rearranging 2 randomly selected subsets from previous optimization cycles, and 7 were randomly selected from subsets from previous optimization cycles. In some instances of adding 10 new subsets, 4 were generated by randomly splitting and rearranging 2 randomly selected subsets from previous optimization cycles, and 6 were randomly selected from subsets from previous optimization cycles. In some instances of adding 10 new subsets, 5 were generated by randomly splitting and rearranging 2 randomly selected subsets from previous optimization cycles, and 5 were randomly selected from subsets from previous optimization cycles. In some instances of adding 10 new subsets, 6 were generated by randomly splitting and rearranging 2 randomly selected subsets from previous optimization cycles, and 4 were randomly selected from subsets from previous optimization cycles. In some instances of adding 10 new subsets, 6 were generated by randomly splitting and rearranging 2 randomly selected subsets from previous optimization cycles, and 4 were randomly selected from subsets from previous optimization cycles. In some instances where 10 new subsets are added, 7 are generated by randomly splitting and recombining 2 randomly selected subsets from previous optimization cycles, and 3 are randomly selected from subsets from previous optimization cycles. Copies of the new subsets may be included in the optimization cycle. In some cases, copies of the new subsets are not included in the optimization cycle.

[0131] In some cases, the number of subsets discarded at the end of each optimization cycle is the same as the number of subsets added before each optimization cycle. In some cases, at the end of each optimization cycle, 5 of the lowest-ranked subsets are discarded, while 5 new subsets are added before each optimization cycle. In some cases, at the end of each optimization cycle, 10 of the lowest-ranked subsets are discarded, while 10 new subsets are added before each optimization cycle. In some cases, at the end of each optimization cycle, 20 of the lowest-ranked subsets are discarded, while 20 new subsets are added before each optimization cycle. In some cases, at the end of each optimization cycle, 50 of the lowest-ranked subsets are discarded, while 50 new subsets are added before each optimization cycle.

[0132] In some aspects, computer-implemented methods for generating classifiers include determining a statistically unweighted subset of microsatellites. In other aspects, computer-implemented methods for generating classifiers include determining a statistically weighted subset of microsatellites. In some cases, the weighted subsets are weighted by relative risk, hazard ratio, or odds ratio. The classifier can be unweighted or weighted. In some cases, the classifier generated by the aforementioned computer-implemented methods can be based on genetic markers other than microsatellites. In some cases, the classifier can be based on other genomic information, such as single nucleotide polymorphisms (SNPs) or genetic aberrations, such as copy number aberrations, insertions, deletions, etc. In some cases, the classifier can be based on the identity of the gene containing the microsatellite.

[0133] After the optimization cycle is completed, the computer-implemented method may include identifying microsatellites associated with or related to the disease with optimized accuracy, sensitivity, and specificity. In some aspects, the computer-implemented method can be validated using additional samples comprising samples with the disease, samples without the disease, or combinations thereof (e.g., see...). Figure 3 Validation may include using at least 10, 20, 30, 50, 100, or 1000 samples from subjects who have a condition (e.g., cancer) (the samples may be non-tumor (germ) samples or tumor samples) and at least 10, 20, 30, 50, 100, or 1000 samples from subjects who do not have the said condition (e.g., cancer, such as lung cancer).

[0134] When analyzing samples from subjects, optimized and validated computer-implemented methods can generate classifiers for symptoms. These symptoms can indicate the presence or absence of a subject's health status. In some cases, the symptoms indicate an increased or decreased likelihood of the subject's health status developing. In some instances, the symptoms can indicate an increased or decreased likelihood of the subject benefiting from treatment, or an increased or decreased risk of adverse effects due to treatment. In some cases, the symptoms can indicate the subject's responsiveness to treatment for their health status. In some instances, the symptoms indicate the prognosis of the subject's health status.

[0135] The symptoms may indicate the presence or absence of cancer. In some cases, the symptoms indicate an increased or decreased likelihood of cancer development. In some instances, the symptoms indicate an increased or decreased likelihood that a subject will benefit from treatment, or an increased or decreased risk of adverse effects due to treatment (the classifier may be a companion diagnostic to cancer treatment). In some cases, the symptoms may indicate responsiveness to cancer treatment. Treatment may be surgery, chemotherapy, radiation therapy, targeted drug therapy (e.g., afatinib, gefinib, bevacizumab, crizotinib, or celtinib), or immunotherapy (e.g., treatment with monoclonal antibodies, checkpoint inhibitors, therapeutic vaccines, or adoptive T-cell transfer). In some instances, the symptoms indicate the prognosis of the cancer. In some cases, the cancer is lung cancer, including non-small cell lung cancer (e.g., lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), and large cell carcinoma), small cell lung cancer (SCLC), or lung carcinoid.

[0136] The classifier may include microsatellite loci from any chromosome, such as chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X, or Y. In some cases, the classifier may not include microsatellite loci from the X chromosome and / or the Y chromosome.

[0137] III. Generating a weighted classifier for the disease.

[0138] This disclosure provides a method for weighting microsatellite loci identified as being associated with or related to a disease. Furthermore, this disclosure provides a method for weighting genetic markers other than microsatellite loci identified as being associated with or related to a disease. Weights or weightings can refer to the relative importance or prevalence of each individual microsatellite locus, which statistically contributes to the correlation or association with the disease. For example, high weights can be assigned to microsatellite loci that appear only in samples obtained from subjects with said disease and that appear at a higher frequency. In some cases, weights are assigned based on hazard ratios, odds ratios, or relative risks. Examples of numerical components that are part of the weight determination include sensitivity, specificity, negative predictive value, positive predictive value, odds ratio, hazard ratio, or any combination thereof. In some cases, a cutoff value (e.g., a threshold) is imposed on the numerical components used to calculate the weights. Samples with numerical classifiers below the cutoff value can be excluded from the weight calculation. Weights can be calculated based on a combination of linear, nonlinear, algebraic, trigonometric, statistical learning, Bayesian, regression, or correlation calculation methods. Classifiers can be generated using weighting schemes or regression methods that use values ​​(e.g., relative risks) associated with one or a group of microsatellites. Weighted classifiers can be evaluated to determine whether weighting improves the classifier's sensitivity or specificity. Regression analysis (e.g., standard regression analysis) can be used to calculate the optimal weights for each seat to maximize sensitivity and specificity (e.g., the sum of sensitivity and specificity).

[0139] In some cases, the weights assigned to each microsatellite are predetermined values, which determine the strength of the correlation or association between the sample size or the disease and the microsatellite locus. In some instances, the weights assigned to each microsatellite include relative risk, hazard ratio, or odds ratio. In some instances, the predetermined values ​​for the weights determine a numerical range (e.g., summation) of sensitivity, specificity, or a combination thereof. In some instances, the calculation and allocation of weights involve decision models implemented by computers via models such as support vector machines, decision trees, random forests, neural networks, or deep learning neural networks (e.g., artificial neural networks, recurrent neural networks, convolutional neural networks, perceptual networks, feedforward networks, radial basis function networks, deep feedforward networks, recurrent neural networks, long / short-term memory, gated recurrent units, autoencoders (AEs), mutated AEs, denoised AEs, sparse AEs, Markov chains, Hopfield networks, Boltzmann machines, restricted BMs, deep belief networks, deep convolutional networks, deconvolutional networks, deep convolutional inverse graph networks, generative adversarial networks, liquid state machines, extreme learning machines, per-state networks, deep residual networks, Kohonen networks, support vector machines, and neural Turing machines).

[0140] In some instances, the weights assigned to microsatellite sites are used as part of the computation of a classifier as described herein. In such instances, microsatellite sites with larger weights contribute more to the classifier's value than those with smaller weights. In some cases, the classifier computation involves using only the optimal weights. The optimal weights may include weights that are at least or greater than a predetermined threshold.

[0141] The symptoms identified by the weighted classifier can indicate the presence or absence of a subject's health status. In some cases, the symptoms identified by the weighted classifier indicate an increased or decreased likelihood of the subject's health status developing. In some instances, the symptoms identified by the weighted classifier indicate an increased or decreased likelihood of the subject benefiting from treatment, or an increased or decreased risk of adverse effects due to treatment. In some instances, the symptoms identified by the weighted classifier indicate the subject's responsiveness to treatment for their health status. In other instances, the symptoms identified by the weighted classifier can indicate the prognosis of the subject's health status. In some cases, the health status is cancer. In some cases, the cancer is lung cancer, such as non-small cell lung cancer (e.g., lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), and large cell carcinoma), small cell lung cancer (NSLC), or lung carcinoid tumors.

[0142] Classifiers can also be determined based on, for example, the distribution of minor alleles in microsatellites. In some cases, a classifier can be determined by calculating a weighted combination of informative microsatellite loci and minor allele distributions. Minor allele frequencies can be an additional weighting parameter for the classifier. Minor allele frequencies can serve as an indicator of overall genome stability. Classifiers based on minor allele frequencies can be statistically evaluated (e.g., through regression analysis) to determine whether adding minor allele frequencies to the classifier improves the classifier. IV. Pan-disease (e.g., cancer) risk assessment

[0143] This disclosure provides a computer-implemented method for generating pan-disease (e.g., cancer) classifiers (see, for example, see...). Figure 2 and Figure 4 Informative microsatellite loci lists can be generated through statistical analysis of samples from various disease (e.g., cancer) types and healthy reference sequences. DNA sequencing from both sets of samples can be performed on multiple platforms. In some cases, sequencing is targeted at additional enrichment, such as using decoy sets. The sequencing results are then subjected to quality analysis and mapping to reveal differences between disease (e.g., cancer) samples and reference samples. This difference can be analyzed using computational methods to generate a pan-disease (e.g., cancer) classifier. The pan-disease (e.g., cancer) classifier can be further optimized and validated with additional samples from various types of diseases (e.g., cancer).

[0144] A pan-disease (e.g., pan-cancer) classifier for one or more diseases can indicate the presence or absence of at least one of a plurality of health states in a subject, the increased or decreased likelihood of the subject developing at least one of a plurality of health states, the increased or decreased likelihood of the subject benefiting from treatment for at least one of a plurality of health states, the increased or decreased likelihood of the subject having an increased risk of adverse effects due to treatment for at least one of a plurality of health states, the subject's responsiveness to treatment for at least one of a plurality of health states, or a combination thereof. The plurality of health states can be any combination of the health states disclosed herein.

[0145] In some cases, a pancancer classifier can indicate the presence or absence of multiple types of cancer in a subject. In some instances, a pancancer classifier can indicate an increased or decreased likelihood of a subject developing multiple types of cancer. In some instances, multiple types of cancer are cancers that frequently develop together in the same subject. In other instances, multiple types of cancer are cancers that occur independently. In some instances, a pancancer classifier can indicate whether a subject may or may not benefit from treatment, or whether a subject may or may not be at increased risk of adverse effects due to treatment (the pancancer classifier can be a companion diagnostic to the treatment product). In some instances, a pancancer classifier can indicate a subject's responsiveness to cancer treatment. In other instances, a pancancer classifier can indicate the prognosis of a subject's cancer. The subjects described herein may have cancer symptoms or no cancer symptoms. In some cases, based on a subject's pancancer classifier, additional examinations may be used (e.g., physical examination, analysis of circulating or cell-free cancer biomarkers, imaging (e.g., computed tomography (CT), bone scan, magnetic resonance imaging (MRI), positron emission tomography (PET), ultrasound, and X-ray), biopsy, gene screening, gene or protein expression levels, etc.).

[0146] Computer-implemented methods for generating pan-disease (e.g., pan-cancer) classifiers may include performing processing, combination, statistical evaluation, or further analysis of results, or any combination thereof. In some aspects, computer-implemented methods for generating pan-disease (e.g., cancer) classifiers include first identifying microsatellite loci (different from microsatellite loci in samples obtained or derived from subjects who do not have multiple types of diseases (e.g., cancer)) to generate a population of subsets of microsatellite loci associated with or related to multiple types of diseases (e.g., cancer). The sequences of the microsatellites may be obtained first by any sequencing method.

[0147] Microsatellite loci associated with or related to a variety of conditions (e.g., cancer) can be identified using one or more statistical tests (such as t-test, Z-test, ANOVA, regression analysis, Mann-Whitney-Wilkcock test, chi-square test, correlation, Fisher exact test, Bonferroni correction, and Benjamini-Hochberg test).

[0148] Statistical tests can generate receiver operating characteristic (ROC) curves, where the area under the ROC curve is called the area under the curve (AUC). AUC determines the accuracy of identifying microsatellite loci associated with or related to multiple types of diseases (e.g., cancer). A larger AUC indicates higher accuracy in correlation or association. The ROC curve determines the ratio of sensitivity (e.g., true positive) and specificity (e.g., true negative) of the correlation or association of microsatellite loci with multiple types of diseases (e.g., cancer). Statistically significant correlation or association of microsatellite loci with multiple types of diseases (e.g., cancer) can have a statistical accuracy of at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. In some cases, the statistically significant correlation or association between microsatellite loci and multiple types of diseases (e.g., cancer) has a statistical specificity of at least 0.70, 0.80, 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99, and a statistical sensitivity of at least 0.70, 0.80, 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.96, 0.99, or 0.99.

[0149] In some instances, identifying microsatellite loci associated with or related to multiple types of diseases (e.g., cancer) involves identifying a first set of microsatellite loci from a database containing nucleic acid sequences of multiple types of diseases (e.g., cancer) and a second set of microsatellite loci from a reference database (e.g., hg19). In some cases, some microsatellites are identified as associated with or related to multiple types of diseases (e.g., cancer). In some cases, some microsatellites are identified as associated with or related to only one type of disease (e.g., cancer).

[0150] Multiple types of cancer can include solid or hematologic malignancies. In some cases, multiple types of cancer can be metastatic, recurrent, or refractory. Multiple types of cancer associated with or related to the identified microsatellite sites can include any number of cancers disclosed herein (e.g., about 4 to about 10, about 10 to about 15, about 15 to about 20, or about 4, about 10, about 15, about 20, about 25, about 30, or about 50).

[0151] The pancancer assay can detect or test for at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, or sixteen of the following cancers: breast cancer, ovarian cancer, prostate cancer, lung cancer, glioblastoma multiforme, endometrial cancer of the uterine corpus, colon adenocarcinoma, bladder cancer, urothelial carcinoma, squamous cell carcinoma of the head and neck, squamous cell carcinoma and adenocarcinoma of the cervix, gastric adenocarcinoma, thyroid cancer, low-grade glioma of the brain, papillary cell carcinoma of the kidney, and hepatocellular carcinoma.

[0152] In some cases, multiple types of cancer, including lung cancer, may be associated with or related to differences in the aforementioned microsatellite loci. Lung cancers that may be associated with or related to different sets of said microsatellite loci include non-small cell lung cancer (e.g., lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), and large cell carcinoma), small cell lung cancer (SCLC), and lung carcinoid.

[0153] Subclusters of microsatellite loci associated with or related to multiple types of diseases (e.g., cancer) may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 microsatellite loci in each subset. In some aspects, the subclusters are iteratively ordered based on their relevance or association with multiple types of diseases (e.g., cancer).

[0154] The subset of the microsatellite locus population can then be optimized by comparing the subset population with additional samples obtained or derived from subjects with multiple types of conditions (e.g., cancer) and / or subjects without multiple types of conditions (e.g., cancer). In some cases, a population of approximately 100 subsets is used in the optimization. In some cases, a population of at least 100, 200, 300, 400, 500, 1000, 2000, 3000, or 5000 subsets is used in the optimization. In some instances, the optimization includes at least one cycle of comparing approximately 100 identified subsets with additional samples. In some instances, the optimization includes multiple cycles of comparing approximately 100 identified subsets with additional samples.

[0155] Iterative ranking can be performed at the end of each cycle. In some cases, iterative ranking involves performing statistical analysis on subsets to perform receiver operating characteristic (ROC) analysis, thereby performing accuracy, sensitivity, and specificity in determining the presence or absence of multiple types of conditions (e.g., cancer) in additional samples. One or more subsets that perform the worst or rank lowest in indicating the presence or absence of multiple types of conditions (e.g., cancer) can be identified and discarded. To maintain a constant number of subsets before the start of each optimization cycle, new subsets can be added to the subset group. In some cases, new subsets are generated by randomly splitting and recombining two subsets randomly selected in previous optimization cycles. In some instances, new subsets are randomly selected from previous optimization cycles. In some cases, the number of subsets discarded at the end of each optimization cycle is the same as the number of subsets added to the subset group before each optimization cycle.

[0156] A computer-implemented method for generating a pan-disease (e.g., pan-cancer) classifier may include determining a statistically unweighted subset of microsatellite loci. In some aspects, the computer-implemented method for generating a pan-disease (e.g., pan-cancer) classifier includes determining a statistically weighted subset of microsatellite loci. The pan-disease (e.g., pan-cancer) classifier may be unweighted or weighted.

[0157] After the optimization cycle is completed, the computer-implemented method for generating a pan-disease (e.g., pan-cancer) classifier includes microsatellite loci associated with or related to the disease with optimized accuracy, sensitivity, and specificity. In some aspects, the computer-implemented method can be validated with additional samples, including samples obtained or derived from subjects with multiple types of diseases (e.g., cancer), samples obtained or derived from subjects without multiple types of diseases (e.g., cancer), or combinations thereof. When analyzing samples from subjects, the optimized and validated computer-implemented method can generate a pan-disease (e.g., pan-cancer) classifier. A pan-disease (e.g., pan-cancer) can indicate the presence or absence of a health condition (e.g., cancer) in a subject. In some cases, a pan-disease (e.g., pan-cancer) indicates an increased or decreased likelihood of a subject developing a health condition (e.g., cancer). In some cases, a pan-disease (e.g., pan-cancer) can indicate an increased or decreased likelihood of a subject benefiting from treatment, or an increased or decreased risk of adverse effects due to treatment (the pan-disease, e.g., pan-cancer, classifier may be a companion diagnostic to the treatment product). In some instances, a pan-disease (e.g., pan-cancer) indicates the responsiveness to treatment of a subject's health condition (e.g., cancer). In other instances, a pan-disease (e.g., pan-cancer) indicates the prognosis of a subject's health condition (e.g., cancer).

[0158] A classifier (e.g., a set of microsatellites) can be developed for each disease (e.g., cancer) in a pan-disease (e.g., pan-cancer) assay. In some cases, a single microsatellite locus can be a pan-disease (e.g., pan-cancer) microsatellite locus.

[0159] V. Assessing the sample of the subjects

[0160] The classifiers generated as described herein can be used to analyze subject (e.g., patient) samples. For example, samples from subjects can be analyzed in a Clinical Laboratory Improvement Amendment (CLIA) certified laboratory. In some cases, kits are prepared and subject samples are measured outside of a CLIA-certified laboratory. Figure 5The illustration shows an example of a workflow (500) for a subject (e.g., patient) sample analysis pipeline in, for example, a CLIA-certified laboratory; said workflow can be used to process samples for multiplex pan-cancer assays. Samples are obtained from multiple subjects (501), for example, from blood, urine, cerebrospinal fluid, semen, saliva, sputum, feces, lymph, tissues (e.g., thyroid, skin, heart, lung, kidney, breast, pancreas, liver, muscle, smooth muscle, bladder, gallbladder, colon, intestine, brain, esophagus, or prostate) or any combination thereof. Nucleic acid molecules, such as genomic DNA, are extracted from the samples. Targets, such as microsatellite targets, are enriched by multiplexing (e.g., using decoys, such as hybridization probes); the enriched targets can be barcoded and amplified (503). Next-generation sequencing assays are performed on the target-enriched samples, for example, in batches of approximately 4, 8, 12, 24, 96, 128, 384, or 1536 times (505). Sequencing data can be demultiplexed (e.g., using unique sequence tags added to each individual sample (e.g., barcodes)), quality control filters can be applied to raw sequence reads (e.g., Phred quality greater than Q30), and genotypes can be determined (e.g., using flanking sequences to align the reads at each locus with a reference sequence, and then calculating the two major alleles (genotypes)) and minor allele distributions (e.g., for each microsatellite locus of each sample (507), determining the number of minor alleles or the fraction of minor alleles relative to the major genotype (minor alleles may be supported by at least 1, at least 2, at least 3, or more than 3 sequence reads). Calculate (509) each A risk classifier for each sample of a cancer (e.g., based on at least 5, 10, 25, 50, or 100 microsatellite loci) (e.g., genotypes can be identified as modal or nonmodal relative to the most prominent genotype in the healthy population (e.g., GRCh38), and summed across all loci, and samples can be classified as at risk or not at risk for a certain condition, depending on their position relative to the cutoff point of loci with cancer or normal genotypes). Risk can be quantitative or indicated by classification assessment. A clinical laboratory report including the risk classifier is generated (511) and made available to healthcare providers, subjects, or insurance providers.

[0161] Figure 17 The illustration shows an example of a clinical laboratory report. A clinical laboratory report can include patient information, sample information, test summary, test results, notes, and result details. Result details can include the number of microsatellite loci for genotyping, one or more disease risk classifiers, one or more thresholds, and the relative risk of having or acquiring a disease (e.g., low risk, high risk, "at risk," "no risk").

[0162] The report may include the number of loci in samples from subjects with nonmodal (primarily cancer) genotypes. The sensitivity and specificity for detecting the presence of health conditions identified as high-risk can be greater than 90%, and these conditions are absent in control sample lines identified as "low-risk" for lung cancer. The accuracy of the assay can be greater than 99% by measuring highly conserved loci in reference controls.

[0163] In some instances, the condition can be verified or further examined through additional tests, such as physical examination, analysis of circulating or cell-free cancer biomarkers, imaging (e.g., computed tomography, bone scan, magnetic resonance imaging, positron emission tomography, ultrasound, and X-ray), biopsy, gene screening, gene expression, or protein expression, etc. VI. Minor alleles in microsatellites

[0164] This disclosure provides a computer-implemented method for determining a subject's genomic age and genomic aging rate. Genomic age can be given in terms of numbers calibrated to the nearest year. For example, if the genomic age is approximately equal to the subject's digital age, then overall genomic stability may be normal for that genomic age. In some instances, the genomic age may be younger, the same as, or older than the subject's chronological age. A genomic age older than the subject's chronological age, or a high genomic aging rate, may indicate genomic instability and a predisposition to age-related health conditions (e.g., diseases), such as cancer, cardiovascular disease, neurological disorders, etc. Genomic age and genomic aging rate may differ in samples obtained from different tissues (e.g., skin or blood) of the same subject. In some cases, genomic age and genomic aging rate can indicate a person's lifestyle (e.g., nutrition, physical or mental stress) or medical condition. Lifestyle modifications (e.g., quitting smoking, dietary changes, and exercise) can be recommended to a subject based on their genomic age.

[0165] Computer-implemented methods for determining genomic age and genomic aging rate may include identifying minor allelic traits in a first sample from a subject and comparing the minor allelic traits of the first sample with those of a reference to produce a first difference in the minor allelic traits. The reference may include the distribution of minor allelic content across a large population to determine the mean genomic age as a function of numerical age, ethnicity, sex, etc. The computer-implemented method can determine that the first difference in minor allelic traits between the first sample and the reference is the subject's genomic age. In some aspects, a second sample from the subject is compared with the reference at a time point after the comparison of the first sample and the reference to produce a second difference in the minor allelic traits. The change between the first and second differences can be determined by the computer-implemented method as the subject's genomic aging rate. In some cases, an additional genomic aging rate can be determined by obtaining and comparing later and earlier minor allelic traits.

[0166] As described herein, a minor allele can be the number of minor alleles at at least one locus. In some aspects, minor alleles include the percentage of SNPs, the percentage of amplified SNPs, the percentage of contracted SNPs, the ratio of amplified to contracted SNPs, the percentage of heterozygous SNPs, the percentage of homozygous SNPs, and the percentage of SNPs with minor alleles. In some cases, minor alleles include a combination of SNPs and insertion / deletion variants, microsatellite variants, synonymous SNPs, nonsynonymous SNPs, stop-gain SNPs, stop-loss SNPs, splicing variants (e.g., 2-bp within a splice junction), frameshift insertions / deletions, and non-frameshift insertions / deletions at at least one locus. In some cases, minor alleles are determined across multiple time points of the same subject.

[0167] Minor allele characteristics identified from a subject's sample may require at least one sequence read from any sequencing method. In some cases, minor allele characteristics may be identifiable from at least one, two, three, four, five, six, seven, eight, nine, ten, twenty, thirty, fifty, or one hundred sequence reads from any next-generation sequencing method. Minor allele characteristics identified from a subject's sample may require at least one, at least two, at least three, or more sequences read from any sequencing method.

[0168] In some instances, minor allelic traits determined from the sequence of a subject's sample are compared to a reference sequence. This comparison can yield differences in minor allelic traits from the reference sequence, which includes varying numbers of SNPs and combinations of insertion / deletion variants, microsatellite variants, synonymous SNPs, non-synonymous SNPs, stop-gain SNPs, stop-loss SNPs, splicing variants (e.g., 2-bp within a splice joint), frameshift insertions / deletions, and non-frameshift insertions / deletions at at least one locus. Differences in minor allelic traits between the sample and the reference can be determined using computer-implemented methods to generate genomic age.

[0169] In some cases, a first sequence from a first sample of a subject is compared with a reference sequence to generate a first major allele and a first genomic age. In some instances, a second sequence from a second sample of the same subject is compared with the same reference sequence to generate a second major allele and a second genomic age. The comparison between the first and second major alleles can determine the genomic aging rate. In some instances, multiple minor alleles can be obtained from samples of the same subject at later time points for comparison to generate multiple genomic aging rates at different ages of the subject.

[0170] This disclosure provides a computer-implemented method for determining the genomic age of a subject by identifying microsatellite minor allele characteristics in a first sample from the subject. Microsatellite minor allele characteristics may be minor alleles comprising microsatellites having different percentages of SNPs, amplification percentages, shrinkage percentages, amplification and shrinkage ratios to SNPs, heterozygous percentages, or homozygous percentages compared to a reference sequence. In some cases, microsatellite minor allele characteristics include minor alleles comprising microsatellites having different combinations of SNPs and insertion / deletion variants, microsatellite variants, synonymous SNPs, non-synonymous SNPs, stop-gain SNPs, stop-loss SNPs, splicing variants (e.g., 2-bp within the splice junction), frameshift insertion / deletion, or non-frameshift insertion / deletion, at least one locus when compared to a reference sequence. In some cases, microsatellite minor allele characteristics are determined across multiple time points of the same subject.

[0171] VI. Computer systems, processors, and memory

[0172] This disclosure provides a computer system configured to implement the methods described herein. In some instances, a system is disclosed comprising: a computer processing device optionally connected to a computer network; and software modules executed by the computer processing device. In some instances, the system includes a central processing unit (CPU), memory (e.g., random access memory, flash memory), electronic storage units, computer programs, communication interfaces for communicating with one or more other systems, and any combination thereof. In some instances, the system is coupled to a computer network, such as the Internet, an intranet, and / or an extranet communicating with the Internet, telecommunications, or data networks. In some aspects, this system includes storage units for storing data and information relating to any aspect of the methods described herein. Various aspects of the system are products, articles, or articles of manufacture.

[0173] A characteristic of a computer program includes a sequence of instructions written to perform a specific task, which can be executed in the CPU of a digital processing device. In some aspects, computer-readable instructions are implemented as program modules that perform specific tasks or implement specific abstract data types, such as functions, features, application programming interfaces (APIs), data structures, etc. In various implementations, computer programs can be written in various versions of various languages.

[0174] The functionality of computer-readable instructions can be combined or distributed in various environments as needed. In some instances, a computer program comprises one or more sequences of instructions. A computer program can be provided from one location. A computer program can be provided from multiple locations. In some aspects, a computer program comprises one or more software modules. In some aspects, a computer program may include, in whole or in part, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plugins, extensions, add-ons, or attachments, or combinations thereof.

[0175] Computer System

[0176] This disclosure provides computer systems programmed to implement the methods of this disclosure. Figure 18 A computer system (1801) is shown, which can be programmed or otherwise configured to perform the methods described herein. The computer system (1801) can regulate various aspects of this disclosure, including inputting nucleic acid location information, transferring inferred information to a dataset, and generating training algorithms from the dataset. The computer system (1801) can be a user electronic device or a remote computer system. The electronic device can be a mobile electronic device.

[0177] The computer system (1801) includes a central processing unit (CPU, also referred to herein as a “processor” and “computer processor”) (1805), which may be a single-core or multi-core processor that processes data sequentially or in parallel. The computer system (1801) also includes storage units or devices (1810) (e.g., random access memory, read-only memory, flash memory), storage units (1815) (e.g., hard disk), communication interfaces (1820) for communicating with one or more other systems (e.g., network adapters), and peripheral devices (1825), external or internal or both, such as printers, monitors, USB drives, and / or CD-ROM drives. The memory (1810), storage units (1815), interfaces (1820), and peripheral devices (1825) communicate with the CPU (1805) via a communication bus (solid line) such as a motherboard. The storage unit (1815) may be a data storage unit (or data repository) for storing data. The computer system (1801) can be operatively coupled to a computer network (“network”) (1830) via a communication interface (1820). The network (1830) can be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some cases, the network (1830) is a telecommunications network and / or a data network. The network (1830) can include one or more computer servers that can enable a peer-to-peer network supporting distributed computing. In some cases, the network (1830) can implement a client-server architecture by means of the computer system (1801), which allows devices coupled to the computer system (1801) to act as either clients or servers.

[0178] The CPU (1805) can execute a series of machine-readable instructions, which can be incorporated into a program or software. The instructions can be stored in memory (1810). The instructions can point to the CPU (1805), which can then be programmed or otherwise configured to implement the methods of this disclosure. Examples of operations performed by the CPU (1805) can include fetching, decoding, executing, and writing back.

[0179] The CPU (1805) may be part of a circuit (such as an integrated circuit). One or more other components of the system (1801) may be included in the circuit. In some implementations, the circuit is an application-specific integrated circuit (ASIC).

[0180] Storage unit (1815) may store files, such as drivers, libraries, and saved programs. Storage unit (1815) may store user data, such as user preferences and user programs. In some cases, computer system (1801) may include one or more additional data storage units outside of computer system (1801), such as those located on a remote server communicating with computer system (1801) via an intranet or the Internet.

[0181] The computer system (1801) can communicate with one or more remote computer systems via a network (1830). For example, the computer system (1801) can communicate with a remote computer system or a user. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets, or tablet PCs (e.g., tablet PCs). iPad GalaxyTab), telephone, smartphone (e.g., iPhone, Android-enabled devices (1830) or personal digital assistant. Users can access the computer system (1801) via a network (1830).

[0182] The method described herein can be implemented by machine-executable code (e.g., a computer processor) stored in an electronic storage location of a computer system (1801) (e.g., stored in memory (1810) or data storage unit (1815)). The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor (1805). In some cases, the code can be retrieved from the storage unit (1815) and stored in memory (1810) for access by the processor (1805) at any time. In some cases, the storage unit (1815) can be excluded, and the machine-executable instructions are stored in memory (1810).

[0183] The code may be pre-compiled and configured for use with a machine having a processor suitable for executing the code, or it may be compiled at runtime. The code may be supplied in a programming language, which may be selected to enable the code to be executed in a pre-compiled or compiled manner.

[0184] Aspects of the systems and methods provided herein, such as computer systems (1801), can be incorporated into programming. These aspects of the technology can be considered "products" or "manufactured goods," typically in the form of machine (or processor) executable code and / or associated data carried or contained in a type of machine-readable medium. Machine-executable code can be stored in storage units (such as hard disks) or memories (e.g., read-only memory, random access memory, flash memory). Storage-type media can include any or all tangible memory of a computer, processor, etc., or associated modules thereof, that can provide non-transitory storage for software programming at any time, including various semiconductor memories, tape drives, disk drives, etc. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another, such as from a management server or host computer to a computer platform for an application server. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices via wired and optical terrestrial networks and various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., can be considered as media carrying software. As used herein, unless limited to non-transitory, tangible "storage" media, the term "readable medium" for a computer or machine refers to any medium that participates in providing instructions to a processor for execution.

[0185] A. Electronic devices

[0186] In some aspects, the platforms, media, methods, and applications described herein include electronic devices, processors, or their use (also referred to as digital processing devices). In other aspects, the electronic device includes one or more hardware central processing units (CPUs) that perform device functions. In still other aspects, the electronic device also includes an operating system configured to execute executable instructions. In some aspects, the electronic device is optionally connected to a computer network. In other aspects, the electronic device is optionally connected to the Internet, enabling it to access the World Wide Web. In still other aspects, the electronic device is optionally connected to a cloud computing infrastructure. In some aspects, the electronic device is optionally connected to an intranet. In some aspects, the electronic device is optionally connected to a data storage device. As a non-limiting example based on the description herein, suitable electronic devices include server computers, desktop computers, laptop computers, notebook computers, netbook computers, network tablet computers, set-top box computers, handheld computers, Internet devices, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. In various embodiments, many smartphones are suitable for the systems described herein. In various embodiments, selected televisions, video players, and digital music players with optional computer network connections are suitable for the systems described herein. Suitable tablet computers include those with brochures, tablets, and convertible configurations.

[0187] In some aspects, electronic devices include an operating system configured to execute executable instructions. For example, an operating system is software, including programs and data, that manages the device's hardware and provides services for the execution of applications. In various implementations, as a non-limiting example, suitable server operating systems include FreeBSD, OpenBSD, etc. Linux, Ubuntu Linux Mac OS X Windows as well as In various implementations, as a non-limiting example, suitable personal computer operating systems include and UNIX-like operating systems, such as In some respects, the operating system is provided by cloud computing. In various implementations, as a non-limiting example, suitable mobile smartphone operating systems include, OS operating system, as well as

[0188] In some aspects, the apparatus includes storage and / or memory devices. Storage and / or memory devices are one or more physical devices for temporarily or permanently storing data or programs. In some aspects, the apparatus is volatile memory and requires power to maintain the stored information. In some aspects, the apparatus is non-volatile memory and retains the stored information when the electronic device is not powered. In other aspects, non-volatile memory includes flash memory. In some aspects, non-volatile memory includes dynamic random access memory (DRAM). In some aspects, non-volatile memory includes ferroelectric random access memory (FRAM). In some aspects, non-volatile memory includes phase-change random access memory (PRAM). In some aspects, non-volatile memory includes magnetoresistive random access memory (MRAM). In some aspects, the apparatus is a memory device, which, as a non-limiting example, includes CD-ROMs, DVDs, flash memory devices, disk drives, magnetic tape drives, optical disc drives, and cloud-based storage devices. In other aspects, storage and / or memory devices are combinations of devices such as those disclosed herein.

[0189] In some respects, the electronic device includes a display that transmits visual information to a subject. In some respects, the display is a cathode ray tube (CRT). In some respects, the display is a liquid crystal display (LCD). In other respects, the display is a thin-film transistor liquid crystal display (TFT-LCD). In some respects, the display is an organic light-emitting diode (OLED) display. In various other respects, the OLED display is a passive-matrix OLED (PMOLED) or an active-matrix OLED (AMOLED) display. In some respects, the display is a plasma display. In some respects, the display is electronic paper or electronic ink. In some respects, the display is a video projector. In yet another respect, the display is a combination of devices such as those disclosed herein.

[0190] In some aspects, the electronic device includes an input device for receiving information from a subject. In some aspects, the input device is a keyboard. In some aspects, the input device is a pointing device, including, as non-limiting examples, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some aspects, the input device is a touchscreen or multi-touchscreen. In some aspects, the input device is a microphone to capture speech or other sound input. In some aspects, the input device is a camera or other sensor to capture motion or visual input. In other aspects, the input device is Kinect, Leap Motion, etc. In yet another aspect, the input device is a combination of devices such as those disclosed herein.

[0191] B. Non-transitory computer-readable storage medium

[0192] In some aspects, the platforms, media, methods, and applications described herein include one or more non-transitory computer-readable storage media encoded with a program, said program comprising instructions executable by an operating system of an optionally networked digital processing device. In other aspects, the computer-readable storage medium is a tangible component of an electronic device. In still other aspects, the computer-readable storage medium is optionally removable from the electronic device. In some aspects, by way of non-limiting example, the computer-readable storage medium includes CD-ROMs, DVDs, flash memory devices, solid-state storage, disk drives, magnetic tape drives, optical disc drives, cloud computing systems and services, etc. In some cases, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitory encoded on the medium.

[0193] C. Computer programs

[0194] In some respects, the platforms, media, methods, and applications described herein include at least one computer program or its use. A computer program comprises a set of instructions written to perform a specific task, executable in the CPU of an electronic device. Computer-readable instructions can be implemented as program modules that perform a specific task or implement a specific abstract data type, such as functions, objects, application programming interfaces (APIs), data structures, etc. In various implementations, computer programs can be written in various versions of various languages.

[0195] The functionality of computer-readable instructions can be combined or distributed in various environments as needed. In some aspects, a computer program comprises a sequence of instructions. In some aspects, a computer program comprises multiple sequences of instructions. In some aspects, a computer program is provided from one location. In some aspects, a computer program is provided from multiple locations. In various aspects, a computer program comprises one or more software modules. In various aspects, a computer program partially or wholly comprises one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plugins, extensions, add-ons, or attachments, or combinations thereof.

[0196] D. Web Applications

[0197] In some respects, computer programs include web applications. In various implementations, web applications utilize one or more software frameworks and one or more database systems. In some respects, web applications are based on, for example... Web applications are built on the .NET or Ruby on Rails (RoR) software framework. In some respects, web applications utilize one or more database systems, including, as non-limiting examples, relational, non-relational, object-oriented, associational, and XML database systems. In other respects, suitable relational database systems include, as non-limiting examples, […]. SQL Server, MySQL TM ,and In various implementations and aspects, web applications are written in one or more versions of one or more languages. Web applications can be written in one or more markup languages, presentation-qualified languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some aspects, web applications are written to some extent in markup languages ​​such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or Extensible Markup Language (XML). In some aspects, web applications are written to some extent in presentation-qualified languages ​​such as Cascading Style Sheets (CSS). In some aspects, web applications are written to some extent in client-side scripting languages ​​such as Asynchronous JavaScript and XML (AJAX). Actionscript, Javascript or In some respects, web applications are written to some extent in server-side coding languages, such as Active Server Web Pages (ASP). Perl, Java TM JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python TM Ruby, Tcl, Smalltalk Or Groovy. In some respects, web applications are written to some extent in database query languages ​​such as Structured Query Language. In other respects, web applications integrate with enterprise server products, such as... In some respects, web applications include media player elements. In various other respects, media player elements utilize one or more of a number of suitable multimedia technologies, including, as non-limiting examples, HTML 5 Java TM and

[0198] E. Mobile Applications

[0199] In some aspects, computer programs include mobile applications provided to mobile electronic devices. In some aspects, mobile applications are provided to mobile electronic devices at the time of manufacture. In some aspects, mobile applications are provided to mobile electronic devices via the computer networks described herein.

[0200] In various implementations, mobile applications are created using a variety of technologies related to hardware, languages, and development environments. In various implementations, mobile applications are written in several languages. As non-limiting examples, suitable programming languages ​​include C, C++, C#, Objective-C, and Java. TM ,Javascript,Pascal,ObjectPascal,Python TM Ruby, VB.NET, WML, and XHTML / HTML or combinations thereof with or without CSS.

[0201] Suitable mobile application development environments can be obtained from several sources. As a non-limiting example, commercially available development environments include AirplaySDK, alcheMo, etc. Celsius, Bedrock, FlashLite, .NET CompactFramework, Rhomobile, and WorkLight Mobile Platform are all available free of charge. Other development environments, as non-restricted examples, include Lazarus, MobiFlex, MoSync, and Phonegap. Additionally, mobile device manufacturers distribute software development kits, as non-restricted examples, including the iPhone and iPad (iOS) SDKs and Android SDKs. TM SDK SDK, BREW SDK, OSSDK, Symbian SDK, webOSSDK, and Mobile SDK.

[0202] In various implementations, several business forums may be used to distribute mobile applications, as non-limiting examples, including... App Store, Android TM Market AppWorld, an app store for handheld devices; App Catalog for webOS. Mobile Market Ovi store for devices, Applications, and DSi Store.

[0203] F. Standalone Applications

[0204] In some respects, computer programs include standalone applications, which are programs that run as independent computer processes, rather than as add-ons to existing processes, e.g., not as plugins. In various implementations, standalone applications are often compiled. A compiler is a computer program that translates source code written in a programming language into binary object code (such as assembly language or machine code). As non-limiting examples, suitable compiled programming languages ​​include C, C++, Objective-C, COBOL, Delphi, Eiffel, and Java. TM Lisp, Python TM Visual Basic, VB.NET, or combinations thereof. Compilation is typically performed at least partially to create an executable program. In some respects, a computer program comprises one or more executable compiled applications.

[0205] G. Software Modules

[0206] In some aspects, the platforms, media, methods, and applications described herein include software, server, and / or database modules, or their use. In various implementations, software modules are created using various techniques involving machines, software, and languages. The software modules disclosed herein can be implemented in a variety of ways. In various aspects, a software module includes files, code segments, programming objects, programming structures, or combinations thereof. In other aspects, a software module includes multiple files, multiple code segments, multiple programming objects, multiple programming structures, or combinations thereof. In various aspects, by way of non-limiting example, one or more software modules include web applications, mobile applications, and standalone applications. In some aspects, the software module is within a computer program or application. In some aspects, the software module is within more than one computer program or application. In some aspects, the software module is hosted on a single machine. In some aspects, the software module is hosted on more than one machine. In other aspects, the software module is hosted on a cloud computing platform. In some aspects, the software module is hosted on one or more machines in one location. In some aspects, the software module is hosted on one or more machines in more than one location.

[0207] H. Database

[0208] In some aspects, the platforms, systems, media, and methods disclosed herein include one or more databases or their use. In various embodiments, many databases are suitable for storing and retrieving barcodes, routes, packages, subject information, or network information. In various aspects, by way of non-limiting examples, suitable databases include relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, association databases, and XML databases. In some aspects, the databases are Internet-based. In other aspects, the databases are network-based. In still other aspects, the databases are cloud-based. In some aspects, the databases are based on one or more local computer storage devices.

[0209] I. Data Transmission

[0210] The subjects described herein, including the methods and systems provided herein, can be configured to be performed in one or more facilities at one or more locations. Facility locations are not nationally restricted and include any country or region. In some instances, one or more steps are performed in a country different from another step of the methods described herein. In some instances, one or more steps for obtaining a sample are performed in a different country than one or more steps for detecting the presence or absence of a condition from the sample. In some aspects, one or more method steps involving a computer system are performed in a different country than another step of the methods provided herein. In some aspects, data processing and analysis are performed in a country or location different from one or more steps of the methods described herein. In some aspects, one or more articles, products, or data are transferred from one or more facilities to one or more different facilities for analysis or further analysis. Articles include, but are not limited to, one or more components obtained from a subject, such as processed cellular material. Processed cellular material includes, but is not limited to, cDNA reverse transcribed from RNA, amplified RNA, amplified cDNA, sequenced DNA, isolated and / or purified RNA, isolated and / or purified DNA, and isolated and / or purified polypeptides. Data includes, but is not limited to, information regarding subject stratification and any data generated by the methods disclosed herein. In some aspects of the methods and systems described herein, the analysis is performed and the subsequent data transmission step transmits or transfers the results of the analysis.

[0211] J.Web Browser Plugin

[0212] In some respects, computer programs include web browser plugins. In computing, a plugin is one or more software components that add specific functionality to a larger software application. Software application manufacturers support plugins, enabling third-party developers to create the ability to extend applications, easily add new features, and reduce application size. When supported, plugins enable the customization of a software application's functionality. For example, plugins are commonly used in web browsers to play videos, generate interactivity, scan for viruses, and display specific file types. In various implementations, several web browser plugins can be used, including... Player as well as In some respects, a toolbar includes one or more web browser extensions, add-ons, or add-ons. In other respects, a toolbar includes one or more browser bars, toolbars, or desktop bars.

[0213] In various implementations, several plug-in frameworks are available, enabling implementation in various programming languages ​​(including but not limited to C++, Delphi, Java). TM PHP, Python TM Develop plugins (using VB.NET or a combination thereof).

[0214] A web browser (also known as an internet browser) is a software application designed for use with network connectivity on electronic devices to retrieve, present, and traverse information resources on the World Wide Web. As a non-limiting example, suitable web browsers include... Internet Chrome Opera And KDE Konqueror. In some respects, a web browser is a mobile web browser. Mobile web browsers (also known as microbrowsers, mini-browsers, and wireless browsers) are designed for mobile electronic devices, including, as non-limiting examples, handheld computers, tablets, netbooks, laptops, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. As non-limiting examples, suitable mobile web browsers include... Browser Browser Blazer, Browser formobile Internet Mobile BasicWeb Browser Mobile, and PSP TM Browser.

[0215] K. Business methods utilizing computers

[0216] The methods described herein can utilize one or more computers. Computers can be used to manage customer and sample information, such as sample or customer tracking, database management, analysis of molecular profiling data, analysis of cytological data, data storage, billing, marketing, reporting results, storing results, or combinations thereof. Computers may include monitors or other graphical interfaces for displaying data, results, billing information, marketing information (e.g., demographic data), customer information, or sample information. Computers may also include means for data or information input. Computers may include processing units and fixed or removable media, or combinations thereof. Computers can be accessed by a user physically near the computer (e.g., via a keyboard and / or mouse), or by a user who does not need to access a physical computer via communication media (such as a modem, internet connection, telephone connection, or wired or wireless communication signal carrier). In some cases, computers may be connected to servers or other communication devices for relaying information from users to or from computers to users. In some cases, users may store data or information obtained from a computer via communication media on a medium (such as a removable medium). It is foreseeable that data related to these methods can be transmitted via such networks or connections for receipt and / or viewing by a party. The recipient may be, but is not limited to, an individual, a healthcare provider, or a healthcare administrator. In one example, a computer-readable medium includes a medium suitable for transmitting the results of biological sample analysis. The medium may include results from a subject, wherein such results were derived using the methods described herein.

[0217] Entities that obtain sample information may input it into a database for one or more of the following purposes: inventory tracking, test result tracking, order tracking, customer management, customer service, billing, and sales. Sample information may include, but is not limited to: customer name, unique customer identification, the customer's associated healthcare professional, instructed test, test result, adequacy status, adequacy of instructed tests, personal medical history, preliminary diagnosis, suspected diagnosis, sample history, insurance provider, healthcare provider, third-party testing center, or any information suitable for storage in the database. Sample history may include, but is not limited to: the sample's age, sample type, method of acquisition, storage method, or transportation method.

[0218] Customers, healthcare professionals, insurance providers, or other third parties may access the database. Database access may take the form of electronic communication, such as a computer or telephone. The database may be accessed through intermediaries (such as customer service representatives, business representatives, consultants, independent testing centers, or healthcare professionals). The availability or extent of database access or sample information (such as test results) may be altered upon payment for products and services already provided or to be provided. The extent of database access or sample information may be restricted to comply with generally accepted or legal requirements regarding patient or customer confidentiality.

[0219] Example

[0220] The examples provided below are for illustrative purposes only and are not intended to limit the scope of the claims provided herein.

[0221] Example 1: Germline microsatellite genotyping to differentiate childhood medulloblastoma (MB)

[0222] introduce

[0223] Medulloblastoma (MB) is a common malignant brain tumor in children. MB can be caused primarily by hereditary or spontaneous mutations because children with MB have not yet experienced lifelong environmental exposures and stresses. Extensive genomic characteristics have classified MB tumors into at least four common molecular subgroups: WNT, SHH, Group 3, and Group 4, each with distinct transcriptional profiles, copy number alterations, somatic mutations, and clinical outcomes. Typically, pediatric brain cancers, and especially MB, have 5–10 times fewer mutations than those commonly observed in adult solid tumors. Particularly uncommon are mutations in the most important tumor initiation genes, such as p53, PTEN, RB, and EGFR. Furthermore, the incidence of known hereditary tumor susceptibility mutations may be relatively low. A few known genetic variants, such as mutations in PTCH, SMO, and CTNNB1, and amplifications of MYC and MYCN, may not be sufficient to efficiently induce MB in animal models and may require an enhancing background, typically p53 inactivation, which is found in less than 5% of human tumors. Many genome-wide association studies (GWAS) in microsatellite transcription (MB) may focus on single nucleotide variants, neglecting non-coding regions and repetitive DNA. However, germline microsatellite (MS) insertions and deletions (insertions and deletions) have shown links to many neurological disorders such as Huntington's disease and Frederick's ataxia; the former is caused by microsatellite variants in coding sequences, and the latter by non-coding intron sequences. Furthermore, microsatellite variants may contribute to the genetic background of several cancers. In addition, many cancer-associated genes contain MS sites (e.g., PTEN and NF1), and in some cases, somatic MS insertions and deletions are causally linked to cancer. Based on these findings, a loosely structured genetic environment can be created by the cooperation of DNA microsatellite repetitive elements that influence the individual's transcriptional and translational landscape, making them more susceptible to tumorigenesis by modulating cellular base processes.

[0224] Microsatellite sites (MS) can comprise tandem repeats of 1–6 base pairs forming arrays. There are over 600,000 unique MSs in the human genome, which can embed within gene introns, exons, and regulatory regions. Due to chain slip replication and heterozygous instability, the length of microsatellite sites frequently changes, varying between alleles and individuals. These changes can affect gene expression by inducing Z-DNA and H-DNA folding; altering nucleosome positioning; and changing the spacing of DNA binding sites. Non-coding variations can alter the DNA secondary structure and protein / RNA binding of genes near their location, leading to changes in transcriptional and translational activity, as well as alternative splicing. For these reasons, MSs are referred to as the “regulatory knobs” of gene expression. Within exons, microsatellite sites containing 3 or 6 base pair repeats can lead to amino acid additions or losses by being held within the frame by codon triplets; other non-mot3 lengths result in frameshift mutations. Genes carrying MSs may disproportionately contribute to neurological disorders. The particular vulnerability of this pair of tandem repeat sequences (especially CAG motifs) to amplification indicates their importance in neural development. In fact, repetitive elements can play a role in neurological disorders; in particular, polyglutamate repeat sequences can play a role in Huntington's disease, spinocerebellar ataxia, and spinal dorsal horn muscle atrophy. Similarly, bioinformatics studies indicate that many genes containing tandem repeat sequences can have neural functions.

[0225] Advances in microsatellite genotyping algorithms and genome sequencing have enabled the identification of germline microsatellite genotypes, which can distinguish healthy individuals from those affected by different types of cancer (breast cancer, colon cancer, glioma, etc.). This example describes a set of microsatellite genotypes that can distinguish children with myxoma from healthy individuals based on germline DNA.

[0226] method

[0227] Patent Sample

[0228] Germline DNA WES and WGS from medulloblastoma (MB) patients were downloaded from the following datasets: phs000504, phs000409, EGAD00001000122, EGAD00001000275, EGAD00001000816, and Waszak, SM, et al. (Spectrum and prevalence of genetic susceptibility to medulloblastoma: retrospective genetic studies and prospective validation in clinical trial cohorts. The Lancet Oncology, Vol. 19, No. 6, pp. 785–798, the entire contents of which are incorporated herein by reference). Additionally, WES from blood DNA of 6 MB patients were newly generated using the TruSeq exome targeted enrichment kit and Illumina Sequencer HiSeq 2500. Germline DNA WES and WGS from healthy controls were downloaded from 1000 Genomes. Germ DNA WES from 100 healthy children was provided by the Hopp Children's Cancer Center (NCT Heidelberg) in Heidelberg, Germany.

[0229] Sequence mapping and covering

[0230] Bowtie2 was used to map WES and WGS readings to the human GRCh38 / hg38 reference genome. Overall, the coverage of the 120MB germline sample was 31-fold (31.0 ± 18.2). The coverage of the control group was 13-fold (13.4 ± 7.8).

[0231] Microsatellite list generation

[0232] The list of microsatellites in the human reference genome version GRCh38 / hg38 was generated using the self-qualified Perl script “searchTandemRepeats.pl” with default parameters. This script can be used for microsatellite research and is available free of charge online. In short, the “searchTandemRepeats.pl” script first searches for pure repeat stretches: impurities are not allowed. Then, imperfect repeats and compound repeats are handled using the “mergeGap” parameter, which has a default value of 10 base pairs. Essentially, impurities such as fragments that break pure repeat sequences are tolerated unless they exceed 10 base pairs. Similarly, repeats close to 10 base pairs are considered compound. The result is that the repeat sequences in the CAGm database are of high purity, and the components of compound repeat sequences are also of high purity. The initial list generated using this script includes 1,671,121 microsatellites. To mitigate the possibility of incorrect read mappings between microsatellites, a subset of all microsatellites with the same repeat motif between the 3' and 5' flanking regions of five base pairs in length were removed. For example, the microsatellite “GCTGC(A)”... 34"CTTAG" and "GCTGC(A)15CTTAG" were preemptively removed from the initial microsatellite list. Microsatellites can be embedded in larger repetitive motifs. The filtered list includes 625,195 unique microsatellites from the human genome.

[0233] Microsatellite genotyping

[0234] The Repeatseq program is used to determine the genotypes of microsatellites in next-generation sequencing reads. Repeatseq uses a Bayesian model selection guided by an empirically derived error model. The error model incorporates sequence and read attributes: units, length, and basic quality. Repeatseq operates on three input files: a reference genome, a file (.bam file) containing reads aligned to a human reference genome, and a list of known microsatellites (according to the methods and systems disclosed herein). The output is a variant call format (.vcf) file listing the genotypes of each microsatellite locus, which consists of two alleles with the most supported reads. A key advantage of Repeatseq over other microsatellite genotyping programs is that it realigns each read with the reference genome before array length detection. Repeatseq is available for microsatellite research and is freely available.

[0235] Repeatseq's capabilities have been extended to detect somatic microsatellite variations: for example, minor alleles. Minor alleles can differ from the major alleles of a genotype; they are acquired in normal tissues through somatic cells as one ages. Minor alleles are used as indicators of microsatellite mutations. In short, the detection of minor alleles is enabled through a two-step process built upon the Repeatseq output. First, a realigned read output is enabled in the call to Repeatseq. Second, the rearranged reads clear away all major alleles of the genotype. Of the remaining reads, those with an array length supported by at least three reads are counted as minor alleles. However, when comparing minor alleles in different samples, another method is used. Specifically, an array length supported by at least 20% of the total read depth is counted as a minor allele.

[0236] Statistical data

[0237] Power calculations were performed based on previous observations of microsatellite genotype distributions in other cancers and controls to select the size of the training set while ensuring sufficient samples in the test set for validation. A conservative Type I error probability associated with a null hypothesis test of 0.01 was selected as part of the validation. Responses within each subject group could be represented as a normal distribution with a standard deviation of 1. The null hypothesis that the population mean of the experimental and control groups was equal to the probability (power) greater than 0.99 was rejected for a true difference of 2 between the experimental and control means, with 120 experimental subjects and 426 control subjects. Therefore, the training set was predicted to have a sufficient number of usable samples.

[0238] For each microsatellite, the distribution of genotypes in the germline DNA of the two groups of samples in the training dataset was different: 120 MB and 425 healthy controls. In each case, the statistical difference was quantified using the generalized Fisher exact test. In short, for each microsatellite, the contingency table was populated with genotype counts for both groups: MB and normal (…). Figure 9 Then, the p-value for each contingency table is calculated using the Fisher test function in R. Benjamini-Hochberg multiple test correction (n = 43,457 tested microsatellites) is applied to control for the false detection rate.

[0239] Microsatellite filtering controls age, race, and sequencing scheme

[0240] This study was designed to identify germline microsatellite variations specific to disease conditions (MBs); specifically, statistically significant microsatellites were identified in 120 MB samples and 425 healthy controls. However, these samples were mismatched in age or sequencing protocol; furthermore, they were only partially matched in ethnicity. Therefore, this approach may risk identifying microsatellites with age, sequencing, and ethnic biases, rather than just disease conditions. To mitigate this risk, microsatellites were identified as potentially biased—age, sequencing, or ethnicity—and excluded from subsequent analyses.

[0241] Age control: To identify microsatellites whose genotypes do not change randomly with age, comparisons were made between 100 healthy European children and 501 European adults from the 1,000 Genomes Project. Fisher's precision test identified 738 (out of 29,061) statistically significant microsatellites: Benjamini-Hochberg correction (p-value < 0.05). Figure 10 ).

[0242] Controlled sequencing protocols: To identify microsatellites varying based on DNA sequencing protocols (WGS vs. WES), genotypes from paired WGS and WES experiments were compared in 16 individuals from the 1000 Genomes Project. Statistical differences in the genotype distribution of 37,511 microsatellites were tested (Fischer exact test); 157 differences were identified using Benjamini-Hochberg error discovery correction (p-value < 0.05). Figure 11 This could be due to the susceptibility of microsatellites to read mapping errors, especially when they carry a large number of insertions or deletions. Therefore, the 157 identified microsatellites were likely particularly prone to mapping errors or located in highly variable regions of the genome; they were excluded from subsequent analyses. Furthermore, 37,775 of the identified microsatellite calls were absent from the 134 WGS samples. Therefore, these 37,775 could not be used for microsatellite-based risk, diagnostic, or prognostic assays; they were excluded from subsequent analyses. Figure 11 ).

[0243] Controlling for race: To identify race-specific DNA microsatellites, the genotype distributions of 352 US samples and 502 European samples from the 1000 Genomes Project were compared and analyzed. A total of 184,981 statistical tests were performed, revealing significant differences in 1,037 microsatellites using Benjamini-Hochberg error-finding correction (p-value < 0.05). Furthermore, the distribution of microsatellite genotypes was examined in a cohort of 59 predominantly European MB samples and 55 predominantly US MB samples. Here, 13,899 tests were performed on 478 microsatellites, finding them to be different after Benjamini-Hochberg error-finding correction (p-value < 0.05). 71 microsatellites present in both lists were identified, and these 71 microsatellites were excluded from further analysis. Figure 12 ).

[0244] The number of unique microsatellites in the above three steps was 38,653; all of these were removed from further analysis.

[0245] Sample scoring metrics and ROC analysis

[0246] The scoring metric for the samples was designed based on the unique distribution of microsatellite genotypes. Essentially, the metric is a weighted sum of the genotypes belonging to each sample: the weights are derived from the frequency differences of each genotype in the MB and healthy groups. Figure 13 Provide a visual summary of the metrics.

[0247] ROC Analysis: Receiver operating characteristic (ROC) analysis is used to design classification schemes that can distinguish MB samples from healthy controls. In short, the area under the ROC curve (AUC) is used as a measure to differentiate between the two groups based on their scores. Then, a cutoff value is selected for all future classifications. Here, the cutoff value is a single score that minimizes sensitivity while maximizing specificity; it is identified using the Youden index. ROC analysis, AUC calculation, and Youden index optimization are performed using the freely available R package: ROCR.

[0248] Microsatellite subsets (genetic algorithm)

[0249] Genetic algorithms can be considered a class of biologically inspired algorithms. In short, this genetic algorithm was used to identify the most informative subset of labels from a set of 139 labels using a two-step iterative process. First, the algorithm was initialized with a random subset of 139 microsatellite labels; next, the top-ranked pre-formed subsets were continuously reorganized, re-evaluated, and re-ranked. Three hyperparameters (e.g., parameters set before the iterative algorithm begins) were used to control the maximum population size, the size of each subset, the performance of each subset, and the diversity of subsets within the population. Details of each step and hyperparameter are provided below.

[0250] Initialization: Each subset in the initial population consists of tags randomly selected from 139 perfect complements. Hyperparameters control the initial population size and the size of each subset. Once populated, the initial subsets are sorted based on the performance metrics described below.

[0251] Optimization: Each optimization cycle begins by placing 10 new subsets into the population; 7 of these are generated by recombining 2 members of the existing population (randomly selected), and 3 are randomly generated. Two subsets are recombined, and each subset is split; then, the two fragments (one from each subset) are rejoined. The split points and fragments are randomly selected. The 3 random subsets are generated during initialization to help maintain population diversity. Once the new subsets are generated, the population is reordered based on performance metrics. Finally, the 10 worst-performing subsets are discarded to maintain the population size.

[0252] Hyperparameters: The population size for 100 subsets is initialized and used throughout the algorithm. The minimum and maximum subset sizes are set to 8 and 64 labels respectively. Duplicate labels are not allowed in subsets. The performance of each subset is determined using ROC analysis with 120MB of samples and 425 healthy controls, e.g., using the same training samples throughout the study. The sum of sensitivity and specificity determines the performance of each subset and is used to perform ranking of the population in each generation of the genetic algorithm.

[0253] Robustness: The parameters of a genetic algorithm are chosen for computational feasibility. However, the results of a genetic algorithm are not sensitive to the choice of hyperparameters. Furthermore, details of the optimization cycle (such as the number of new subsets in each cycle) do not affect the results of the genetic algorithm.

[0254] verify

[0255] Sample used: To ensure sufficient power of the study, 102 experimental subjects and 428 control subjects were selected in the validation study. The subjects (MB) and control distributions identified during the analysis of the training set were used. Figure 7A The responses within each subject group were normally distributed with a standard deviation of 1.1. A null hypothesis was used to reject the study based on a true difference of 4.4 between the experimental and control means; that is, for this sample size and control validation set, the population means of the experimental and control groups equal a probability (power) greater than 0.99 for a Type I error probability of 0.01. All control samples used in training and validation underwent whole-exome sequencing. For MB, the collection included both whole-exome and whole-genome samples. The whole-genome sequencing samples were used exclusively for validation.

[0256] Procedure: Each validation sample was scored using the same metrics as the training samples. A cutoff value (identified during training) was used to predict which of the 530 validation samples had MB and which were healthy controls. MB were predicted as validation samples exceeding the cutoff value. The predictions were compared to the known identities of 102 MB samples and 428 healthy controls. The sensitivity and specificity of these predictions were comparable to the training values.

[0257] microsatellite mutation

[0258] To test whether individuals with MB are more prone to microsatellite variants, the total number of alleles for each microsatellite genotype (allele load) was used as a measure of its mutations, and this measure was compared between disease and control cohorts. Allele constraints made the counting robust to two sources of error: (a) by requiring at least two reads to support each allele, the potential impact of PCR products was mitigated; and (b) to normalize for differences in read coverage between samples, each allele needed to be supported by at least 20% of the total number of reads mapped to the microsatellite. Alleles were counted only for microsatellites with mapped reads present in at least 20% of the samples. Fisher's exact test was then performed to establish statistical significance between MB patients and healthy individuals. This procedure was repeated 50 times, with a mean p-value of 0.077.

[0259] Two additional lines of evidence were used to assess the integrity of the germline mismatch repair mechanism in medulloblastoma: (a) recording homozygous and heterozygous genotypes on all microsatellites (71,192 in total) in both MB and control samples; and (b) comparing the median microsatellite array length on all microsatellites (71,192 in total) in both MB and control samples. For the former analysis, aberrant mismatch repair was expected to increase the number of heterozygous genotypes; however, the difference between case and control samples was not statistically significant. The medulloblastoma samples contained 299,802 heterozygous genotypes and 2,596,324 homozygous genotypes; the control samples contained 283,037 heterozygous genotypes and 2,449,046 homozygous genotypes. For the latter analysis, aberrant mismatch repair was expected to result in an accumulation of longer or shorter median microsatellite array lengths in medulloblastoma samples compared to controls; again, the results were not statistically significant. In the medulloblastoma sample, the median array length of 1,031 microsatellites was relatively short, while that of 907 microsatellites was relatively long; the median array length of the remaining 69,254 microsatellites did not differ.

[0260] Downstream Analysis

[0261] Functional analysis was performed using genes associated with 139 microsatellite loci, the genotypes of which differed significantly between MB subjects and controls. A total of 124 genes were included in the analysis, excluding microsatellites located in intergenic regions. Pathway analysis was performed using Ingenuity Pathway Analysis (QIAGENInc.). Mutations and co-occurrences were analyzed using PedcBioPortal. Protein-protein interaction (PPI) network construction was performed using STRING with a minimum interaction score of 0.7 (high confidence) and no more than five molecules in the first shell. This setup generated a hub with 129 nodes and 49 edges, resulting in a PPI-enriched network with a p-value of 0.0007.

[0262] result

[0263] Identification of microsatellite informative sites in medulloblastoma

[0264] Single nucleotide mutations can be characterized using whole-genome (MB) analysis. Here, the impact of microsatellite variations on susceptibility to medulloblastoma was investigated. To this end, a computational workflow was developed to identify genotypic germline microsatellites between children with medulloblastoma and control subjects, while correcting for variations with age, ethnicity, and DNA sequencing protocol. Figure 6A metric was also developed to score microsatellite genotypes based on a unique collection for each sample. This method was applied to germline DNA sequencing data from 222 children with medulloblastoma and 853 healthy controls. The data were divided into two groups, each containing affected and healthy subjects: the first group was used for training, containing 120 medulloblastoma patients and 425 controls, and the second group was used for validation, containing 102 medulloblastoma patients and 428 controls. In the first phase of the analysis, using the training set, 43,457 distinct microsatellites present in the 120 medulloblastoma samples and 425 healthy controls were genotyped. For each of these microsatellites, the statistical difference in genotype distribution between the two groups was assessed using the generalized Fisher exact test. 2,094 microsatellites were identified with p-values ​​<0.05. After Benjamini-Hochberg multiple testing correction (α = .05), 422 passed the false discovery test. Three additional steps are performed to remove microsatellites that vary with age, ethnicity, and DNA sequencing protocol. Figure 6 , Figure 10 , Figure 11 and Figure 12 A total of 283 microsatellites were removed from the list of 422 satellites, reducing the list to 139. Figure 19 In summary, this method identified 139 microsatellites from germline DNA, whose genotypes differed significantly between medulloblastoma subjects and healthy controls.

[0265] Microsatellite classifier set for medulloblastoma

[0266] To identify the microsatellite subset with optimal performance in distinguishing medulloblastoma samples from healthy controls, a medulloblastoma classifier was trained using a set of 139 microsatellites. First, an index was designed to score each medulloblastoma and control sample based on the genotypes of the 139 microsatellites (see Methods and...). Figure 13 Next, receiver operating characteristics (ROCs) are generated and used to determine the sample scores as a binary classifier for medulloblastoma. A subset optimization strategy based on a genetic algorithm is used to identify the best subset of distinguishing labels using a two-step iterative process. First, subsets are randomly generated from the complete list and sorted by their F-metric. Second, the best-performing subsets are continuously mixed, re-evaluated, and re-sorted. The algorithm converges within 87 cycles to reveal a subset of 43 microsatellites with an F-metric of 0.90 and an area under the curve (AUC) of 0.962 (Figure 7). Figure 20 The Youden index was determined, indicating that the optimal cutoff score for distinguishing medulloblastoma samples from healthy controls was 0.155. Figure 14When applied to the training set, the sensitivity was 0.88 and the specificity was 0.92. Figure 7B ).exist Figure 15 The chromosomal locations of these 43 markers in the human genome are shown. Therefore, a set of 43 microsatellites was identified, and the genotype distribution of this set of 43 microsatellites was able to distinguish between medulloblastoma patients and healthy controls with 88% sensitivity and 92% specificity.

[0267] Independent germline DNA cohorts from medulloblastoma patients and healthy controls were used to validate previous results. For the validation study, which included 102 experimental subjects and 428 control subjects, the subject (medulloblastoma) and control distributions identified during the analysis of the training set (Figure 7) were used to ensure adequate power of the study. In the training set, responses within each subject group were normally distributed with a standard deviation of 1.1. A true difference of 4.4 between the experimental and control means was found to allow rejection of the null hypothesis that the population means of the experimental and control groups are equal with a probability (power) greater than 0.99, resulting in a Type I error probability of 0.01 for this sample size and the control group. Applying the optimal cutoff (0.155) to the independent validation sample set, the classifier was found to distinguish between cases and controls with a sensitivity of 0.95 and a specificity of 0.90. Figure 7C and Figure 7D In summary, the genotype distribution of a group of 43 MS was identified and validated, enabling the differentiation of MB patients and healthy controls using germline DNA with high sensitivity and specificity.

[0268] Mutations in informative microsatellite sites in medulloblastoma

[0269] In germline phylogenetic studies, the rate of insertions and deletions in MS was significantly higher than the rate of single nucleotide substitutions elsewhere in the genome, correspondingly 10-1. -4 Up to 10 -3 In contrast, the ratio of each generation to each seat is 10. -8 However, the mutation rates of different MSs vary based on repeat length, their repeat motifs, and their impact on DNA folding. Assuming that among 139 MSs whose genotypes are not randomly associated with MB ( Figure 20The differences found in the study could be a result of increased microsatellite genotypic variation inherent in MB individuals. To test whether individuals with MB are more prone to microsatellite variation, the total number of alleles for each microsatellite genotype (allele load) was used as a measure of its mutations, and this measure was compared between disease and control cohorts. There was no significant difference in the number of genotypic alleles between healthy and MB individuals, supporting the conclusion that there is no generalized microsatellite instability in MB patients. The predictive power associated with informative microsatellite characteristics was investigated by sorting all MS by allele load to determine whether 139 markers were located at the most mutated sites analyzed. It was found that although they belonged to more mutated MS, they did not include the most mutated sites. Furthermore, the number of homozygous and heterozygous genotypes and the length of the microsatellite array were compared as potential sources of MB variation. In both cases, there were no statistically significant differences between MB and control germline DNA. These results and data suggest that the association of those 139 microsatellites with MB is a result of these individual microsatellite genotypes, and not just structural hypervariation.

[0270] The role of informative MST-associated genes

[0271] Of the 139 MS loci with genotypic differences between MB and control samples, 114 were located in intronic regions, 15 in intergenic regions, 6 in the 3' UTR, 3 in exon regions, and 1 in the 5' UTR. Figure 8A To understand the potential mechanistic roles of these genes, Ingenuity was investigated. The study analyzed 124 genes (excluding MS located in intergenic regions) associated with informative metastatic sites (MS). The analysis revealed statistically significant associations with cancer and molecular cellular functions such as cell cycle, DNA replication, recombination and repair, and cell growth and proliferation, indicating a relationship with cancer biology. Figure 8B and Figure 21 The occurrence of mutations in these 124 genes associated with informative MS was examined in a 4MB cohort available in cBioportal. Although the mutation rate is known to be low in MB tumors, an average of 17% of MB cancer samples contained mutations in at least one of these 124 genes. Figure 22 The mutation rate in neuroblastoma tumors was 4.5%. Mutation co-occurrence analysis using the 2016 dataset of affected children from cBioportal indicated that 135 pairs (9,591 = 139*(139-1) / 2) of all possible microsatellite pairs were found to co-occur significantly (p-value <0.05). Two patients were found to have simultaneous mutations at the 20MB and 10MB informative MS sites, respectively. Figure 23 ).

[0272] A protein-protein interaction (PPI) network composed of 124 genes associated with informative MS sites was discovered. Figure 8C The network comprises 129 nodes and 49 edges, resulting in a PPI-enriched p-value of 0.0007. Despite the small number of proteins used as inputs, it is a crucial hub associated with mTOR, an important pathway in macrophage tumors (PI3K / AKT / mTOR).

[0273] Three informative microsatellite loci are located in the protein-coding sequence ( Figure 8A These are all trinucleotide repeat sequences (RAI1, BCL6B, TNS1). Variations in trinucleotide repeat sequences are considered a cause of neurological and neuromuscular diseases such as Huntington's disease, spinocerebellar ataxia, and Fragile X syndrome. Two of these genes (RAI1, BCL6B) are transcription factors located on the short arm of chromosome 17, and their deletion is a recurrent alteration in the most common subgroup of MB tumors. The BCL6B gene is associated with colon cancer, gastric cancer, and liver cancer. The predominant genotype of MB tumors is 33 / 33, while the control group is 30 / 33. Figure 16 In this read frame, the codon CAG is translated to serine. RAI1 (retinoic acid-induced protein) encodes a nuclear protein of unknown function, and haploinadequacy of this protein leads to Smith-Marges syndrome. The two main genotypes of RAI1 in MB tumors are 38 / 41 and 41 / 41, while in controls they are 38 / 38 and 38 / 41. Figure 16 In addition to inducing changes in the structure of the RAI1 protein, short polyglutamine amplification is also thought to regulate transcription factor activity. The RAI1 protein is highly expressed in the cerebellum, a region where MB tumors develop.

[0274] In this study, a subset of 139 MS were identified as having distinct genotypes between MB patients and healthy controls. A subset of 43 MS were able to distinguish MB individuals from controls based on their germline DNA, with sensitivity and specificity of 0.95 and 0.90, respectively.

[0275] This study identified three groups of microsatellites: (a) 43 microsatellites that collectively distinguished medulloblastoma samples from healthy controls; (b) 139 microsatellites whose genotypes showed statistically significant differences between medulloblastoma samples and healthy controls; and (c) 422 microsatellites identified in the initial screening. All microsatellites in all three groups passed the false discovery stage. The group of microsatellites identified in the initial screening (c) comprised 283 microsatellites sensitive to age, race, and / or DNA sequencing; therefore, they were not used in subsequent analyses. Some of the racially biased microsatellites may also play a role in medulloblastoma. The prevalence of many diseases, including medulloblastoma, can show racial differences. Therefore, a re-examination of the 283 microsatellites could be feasible once more understanding of the genetic mechanisms leading to medulloblastoma is gained.

[0276] Furthermore, the relationship between the 139 microsatellite set (b) and its 43 microsatellite subset (a) was investigated: the latter distinguished medulloblastoma samples from healthy controls, while the former did not. Mutations in the 43 microsatellite set could have a greater impact on gene expression; or genes carrying these microsatellites could have a greater impact on disease onset. This is supported by the presence of two coding microsatellites in the 43 coding microsatellite sets; in both cases, mutations directly affect the primary structure of the protein and have potential effects on secondary structure and function. Additionally, a larger proportion of the 43 microsatellites in this set contained 5' and 3' UTR regions; it is possible that the MS in these regions has a stronger impact on gene expression / translation. These indications can be determined by studying the expression of these genes containing informative microsatellites in tumor tissue.

[0277] These results indicate that polyglutamine microsatellites embedded in the BCL6B and RAI1 genes may play a role in medulloblastoma. Only 181 polyglutamine microsatellites were found in the complete list of screened microsatellites (out of 627,174). Therefore, mere chance cannot explain the presence of two in the final list of 43 informative microsatellites; using computer simulations, the probability of this occurring randomly is estimated to be approximately 1 in 1,000,000. Secondly, polyglutamine microsatellites may play a role in diseases such as spinal and bulbar muscular atrophy, Huntington's disease, and various spinocerebellar ataxias. Furthermore, both the BCL6B and RAI1 genes may be associated with diseases; the former in lymphoma and the latter in Smith-Marges syndrome. Polyglutamine diseases are characterized by the aggregation of insoluble proteins: this is not seen in some cancers. On the other hand, polyglutamine amplification can confer or lose function depending on the affected proteins.

[0278] This study yielded two overall conclusions. First, the identified microsatellites—specifically the set of 139 and a subset of 43—may play a role in the etiology of medulloblastoma. The effects of variations in microsatellite array length included influences on DNA secondary structure, nucleosome localization, and DNA binding sites. Three identified microsatellites affected the protein primary sequence. The microsatellites could aid in differentiating medulloblastoma patients from healthy controls; the classification scheme showed high sensitivity and specificity of 0.95 and 0.90, respectively.

[0279] Treatment of medulloblastoma can leave survivors with a lifelong burden, including hearing loss, cognitive impairment, endocrine disorders, and a high risk of stroke and secondary malignancies. Identifying at-risk groups for medulloblastoma development can enable early detection strategies, leading to less invasive and more localized tumor control. However, the most effective way to improve the lives of these children is to prevent their tumors from forming in the first place. Recent advances in immunotherapy, including cancer vaccines, have created the potential for immunizing individuals with tumor-specific antigens. Such strategies may require selecting individuals suitable for such interventions.

[0280] Example 2: Informative microsatellite tag recognition

[0281] Nucleic acid sequence samples were obtained from public domain databases from subjects with the disease (Group 1) and healthy controls (Group 2). Microsatellite loci were identified in both groups. Microsatellites were compared to reveal differences in microsatellite loci found only in Group 1 and their specific association with the disease. Statistical analysis and modeling were applied to these different microsatellites for their correlation or association with the disease. In some instances, the microsatellites were statistically weighted. After a set of microsatellites was identified as strongly associated with the disease, these microsatellites were assembled into a training algorithm to further optimize the accuracy, sensitivity, and specificity of the association between these microsatellites and the disease. During training, microsatellites could be randomly recombined to generate additional microsatellite combinations. After training, the algorithm was validated using additional independent sample sets.

[0282] For example, nucleic acid sequences from cancer patients and corresponding healthy controls were downloaded from The Cancer Genome Atlas (TCGA) and the Thousand Genomes Project, respectively. Microsatellite loci were identified in both groups. Comparison of microsatellites between the two groups revealed a population of microsatellite loci found only in the cancer patient group and specifically associated with or related to a single cancer type. An algorithm was then trained on these cancer-type-associated microsatellites to improve the accuracy, sensitivity, and specificity of these microsatellites in relation to cancer. After training, the algorithm was validated using additional samples from the same groups, either containing cancer or from healthy controls. Once validated, the algorithm could be applied to the patient samples.

[0283] Example 3: Patient risk assessment

[0284] Serum samples were isolated from subjects during routine health checkups. DNA was extracted from the serum samples and sequenced. The sequencing data was processed and analyzed to generate a set of microsatellites unique to each subject. This set of microsatellites was then analyzed using a computer-implemented method designed to determine the risk of developing cancer based on comparisons between the subject's microsatellites and microsatellites from a pan-cancer database. Each of the identified informative microsatellites was assigned a weight ranging from 0 to 1. The weights were generated based on the accuracy, sensitivity, and specificity of the identified microsatellites. The sum of the weights was then determined and used to create a classifier to determine the likelihood of developing a type of cancer. A pan-cancer classifier was then compiled and reported for risk assessment of the subject, targeting multiple likelihoods of developing various cancers. The pan-cancer classifier provides a risk assessment of the likelihood of a subject developing cancer (such as breast cancer, lung cancer, prostate cancer, cervical adenocarcinoma, glioblastoma multiforme, endometrial cancer, colon adenocarcinoma, bladder and urothelial carcinoma, head and neck squamous cell carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, gastric adenocarcinoma, thyroid cancer, low-grade glioma of the brain, renal papillary cell carcinoma, and hepatocellular carcinoma).

[0285] Risk assessment was communicated to subjects via laboratory reports. Figure 5 and Figure 17 Information about the patient, healthcare professional, and serum sample is listed along with the test summary. The summary reveals that although the subject currently does not have cancer, several identified microsatellites in the subject's genome increase the likelihood of developing lung cancer. The classifier for the likelihood of developing lung cancer includes a numerical output and is compared to a threshold for the likelihood of developing lung cancer. The threshold for the likelihood of developing lung cancer is 0.3, with a standard deviation range of 0.1 and 0.5. Figure 24 The classifier for the likelihood of a subject developing lung cancer was 2.3, indicating a high probability of future cancer development. Therefore, additional clinical attention was given to the subject's lungs and respiratory system. Regular, more routine lung imaging was recommended. Subjects were also advised not to start smoking and to avoid prolonged exposure to environments containing known aerosol carcinogens. Furthermore, the summary provides an overview of the risk assessment parameters, such as the statistical methods used, the type of thresholds, and the number of microsatellite loci analyzed.

[0286] Example 4: Measuring genomic age using minor alleles

[0287] DNA samples from primary skin fibroblasts were obtained from subjects aged 17 to 30 years. DNA-seq libraries were constructed, followed by sequencing using a next-generation sequencing platform and mapping to hg19. Enrichment was performed to identify hotspots of minor alleles predisposing to clusters. Minor alleles with at least 5 reads were independently identified by Sanger sequencing. True positive minor alleles were analyzed and weighted. Examples of locations where minor alleles were found included upstream or downstream of genes, exon regions, intergenic regions, regions spanning introns and exons, 3'UTR, and 5'UTR. Minor alleles could be nonsynonymous variants, synonymous variants, frameshift insertions / deletions, non-frameshift insertions / deletions, stop gain, stop loss, or combinations thereof.

[0288] Secondary alleles obtained by comparing samples from age 17 with the hg19 reference sequence were analyzed using computer-based methods to reveal genomic age. An increase in the number of secondary alleles or loci indicates a genomic age that is older than the subject's actual age and physical health. Samples from the same subject at ages 17 and 30 can be compared to reveal additional accumulation or shifts in secondary allele patterns within the same subject. Comparison of secondary alleles between ages 17 and 30 revealed a slight increase in the total number of secondary alleles in the subject. This increase was analyzed using computer-based methods to reveal an accelerated rate of genomic aging in the subject. Therefore, a lifestyle emphasizing nutritional balance and stress reduction is recommended for the subject.

[0289] This invention provides, but is not limited to, the following embodiments:

[0290] 1. A computer-implemented method for constructing an optimized classifier for a disease, the method comprising sorting a subset of a plurality of microsatellites for a classifier for the disease in a plurality of optimization cycles, wherein the subset of the plurality of microsatellites includes microsatellites in an initial group of microsatellites associated with the disease, thereby identifying an optimized subset of the subset of microsatellites as the optimized classifier for the disease.

[0291] 2. The method according to embodiment 1 further includes comparing microsatellites from a first group of samples from subjects suffering from the disease and microsatellites from a second group of samples from subjects not suffering from the disease, thereby identifying the initial population of microsatellites.

[0292] 3. The method according to embodiment 1, wherein the sorting includes comparing a subset of the microsatellites with microsatellites from samples of subjects with the disease and microsatellites from samples of subjects without the disease.

[0293] 4. The method according to embodiment 1 further includes initialization, wherein the initialization includes randomly selecting an initial subset of microsatellites from the initial population of microsatellites for sorting in the optimization cycles of the plurality of optimization cycles.

[0294] 5. The method according to embodiment 1, wherein at least 100 subsets of the initial population of the microsatellites are used in the plurality of optimization cycles.

[0295] 6. The method according to embodiment 1, wherein the minimum number of microsatellites in the subset of the microsatellites is 8.

[0296] 7. The method according to embodiment 1, wherein the maximum number of microsatellites in the subset of the subset of microsatellites is 64.

[0297] 8. The method according to embodiment 1, wherein duplicate microsatellites are not allowed in a subset of the subset of microsatellites.

[0298] 9. The method according to embodiment 1, wherein the sorting includes performing receiver operating characteristic (ROC) analysis using (i) the subset of microsatellites, (ii) microsatellites from samples of subjects with the disease, and (iii) microsatellites from samples of subjects without the disease.

[0299] 10. The method according to embodiment 9, wherein the sorting in the optimization loop of the plurality of optimization loops includes: determining the sum of the sensitivity and specificity of microsatellites in each subset of the subset that serves as the classifier for the disease.

[0300] 11. The method according to embodiment 10, wherein the optimization cycle of the plurality of optimization cycles includes: adding 10 new subsets of the initial population of microsatellites to subsets from previous optimization cycles of the plurality of optimization cycles.

[0301] 12. According to the method of embodiment 11, 7 of the 10 new subsets are generated by randomly splitting and recombining 2 randomly selected subsets from the previous optimization cycle, and 3 of the 10 new subsets are generated by randomly selecting microsatellites from the initial population of the microsatellites.

[0302] 13. The method according to embodiment 12 further includes, at least in part, discarding 10 subsets of a subset during the optimization period based on having the lowest ranking during the optimization period.

[0303] 14. The method according to embodiment 1, wherein the condition includes the presence or absence of the subject's health status.

[0304] 15. The method according to embodiment 1, wherein the condition includes an increased or decreased likelihood of the subject developing a healthy state.

[0305] 16. The method according to embodiment 1, wherein the condition includes an increased or decreased likelihood that the subject benefits from treatment of a healthy state.

[0306] 17. The method according to embodiment 1, wherein the condition includes an increased or decreased likelihood that the subject has an increased risk of adverse effects due to treatment of a health condition.

[0307] 18. The method according to embodiment 1, wherein the condition includes the subject's responsiveness to treatment of a health state.

[0308] 19. The method according to embodiment 1, wherein the condition includes the prognosis of the subject's health status.

[0309] 20. The method according to any one of embodiments 14 to 19, wherein the health condition is cancer.

[0310] 21. The method according to embodiment 20, wherein the cancer is lung cancer.

[0311] 22. The method according to any one of embodiments 14 to 19, wherein the health state is a neurological disease.

[0312] 23. The method according to any one of embodiments 14 to 19, wherein the health condition is cardiovascular disease.

[0313] 24. A computer-implemented method comprising using a plurality of parameters to determine the value of a classifier for a condition from samples of a subject, wherein each of the plurality of parameters is a statistical measure of the correlation of each of a plurality of microsatellites from samples of a subject having the condition or from samples of a subject not having the condition.

[0314] 25. The method according to embodiment 24, wherein the plurality of parameters includes a plurality of weights.

[0315] 26. The method according to embodiment 25, wherein the plurality of weights includes a plurality of optimal weights.

[0316] 27. The method according to embodiment 26 further includes determining the plurality of optimal weights.

[0317] 28. The method according to embodiment 27, wherein determining the plurality of optimal weights includes applying standard regression analysis to the plurality of weights.

[0318] 29. The method according to embodiment 24, wherein determining the plurality of optimal weights includes using a genetic algorithm.

[0319] 30. The method according to embodiment 24, wherein determining the value of the classifier includes using minor allele frequency data.

[0320] 31. The method according to embodiment 24, wherein the plurality of microsatellites includes at least 10 microsatellites.

[0321] 32. The method according to embodiment 24, wherein each of the plurality of microsatellites is associated with the disease.

[0322] 33. The method according to embodiment 24 further includes comparing the value of the classifier with a threshold.

[0323] 34. The method according to embodiment 24, wherein the condition includes the presence or absence of the subject's health status.

[0324] 35. The method according to embodiment 24, wherein the condition includes an increased or decreased likelihood of the subject developing a healthy state.

[0325] 36. The method according to embodiment 24, wherein the condition includes an increased or decreased likelihood that the subject benefits from treatment of a healthy state.

[0326] 37. The method according to embodiment 24, wherein the condition includes an increased or decreased likelihood that the subject has an increased risk of adverse effects due to treatment of a health condition.

[0327] 38. The method according to embodiment 24, wherein the condition includes the subject's responsiveness to treatment of a health state.

[0328] 39. The method according to any one of embodiments 34 to 38, wherein the disease is cancer, cardiovascular disease or nervous system disease.

[0329] 40. The method according to embodiment 39, wherein the cancer is lung cancer.

[0330] 41. A computer-implemented method for determining the genomic age of a subject, the method comprising:

[0331] a) Identify microsatellite minor allele signatures in the first sample from the subjects;

[0332] b) Using a reference to process the microsatellite minor allele characteristics; and

[0333] c) Determine the subject's genomic age based on the treatment.

[0334] 42. The method according to embodiment 41, wherein the processing includes comparing the microsatellite minor allele characteristics with the reference.

[0335] 43. The method according to embodiment 41, wherein the minor allele feature is a plurality of minor alleles at a genotype locus.

[0336] 44. The method according to embodiment 43, wherein the number of minor alleles is supported by at least three next-generation sequencing reads.

[0337] 45. The method according to embodiment 41, wherein the minor allele feature is the total number of minor allele readings, which is normalized to the total number of major allele readings at the locus.

[0338] 46. ​​The method according to embodiment 41 further includes performing next-generation sequencing on the first sample from the subject to generate sequence reads of the subject's microsatellites.

[0339] 47. The method according to embodiment 46, wherein the first sample comprises blood, saliva, or a tumor.

[0340] 48. The method according to embodiment 45 further includes: after operation c), determining minor allele characteristics in a second sample from the subject.

[0341] 49. The method according to embodiment 47 further includes: assessing the minor allele characteristics in the first sample from the subject and the minor allele characteristics in the second sample from the subject, and determining the genomic aging rate of the subject based on the assessment.

[0342] 50. A computer-implemented method, comprising:

[0343] a) Use microsatellites from the samples from the subjects to determine multiple classifiers for the samples from the subjects;

[0344] b) Process the multiple classifiers using multiple reference classifiers for various diseases; and

[0345] c) Based on the treatment, identify at least one condition for the subject from among the multiple conditions.

[0346] 51. The method according to embodiment 50, wherein the processing includes comparing the plurality of classifiers with the plurality of reference classifiers for the plurality of diseases.

[0347] 52. The method according to embodiment 50, wherein the at least one of the multiple conditions includes the presence or absence of at least one of the multiple health states of the subject.

[0348] 53. The method according to embodiment 50, wherein the at least one of the multiple conditions includes an increased or decreased probability of developing at least one health condition from the multiple health conditions of the subject.

[0349] 54. The method according to embodiment 50, wherein the at least one of the multiple conditions includes an increased or decreased likelihood that the subject benefits from treatment of at least one of the multiple health conditions of the subject.

[0350] 55. The method according to embodiment 50, wherein the at least one of the multiple conditions includes the possibility that treatment of at least one of the multiple health conditions of the subject increases or decreases the risk of adverse effects on the subject.

[0351] 56. The method according to embodiment 50, wherein the at least one of the multiple conditions includes the subject's responsiveness to treatment for at least one of the multiple health conditions of the subject.

[0352] 57. The method according to any one of embodiments 51 to 56, wherein the multiple health conditions include multiple cancers.

[0353] 58. The method according to embodiment 57, wherein the plurality of cancers includes ovarian cancer, breast cancer, low-grade glioma, glioblastoma, lung cancer, prostate cancer, or melanoma.

[0354] 59. The method according to embodiment 50, wherein the multiple health conditions include multiple neurological diseases or multiple cardiovascular diseases.

[0355] 60. A non-transitory computer-readable medium comprising executable instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of embodiments 1 to 59.

[0356] 61. A computer system including a hardware processor configured to execute the instructions on the non-transitory computer-readable medium of claim 60.

[0357] Although preferred aspects of this example have been shown and described herein, it will be apparent to those skilled in the art that such aspects are provided by way of example only. Many variations, alterations, and substitutions will now occur to those skilled in the art without departing from this disclosure. It should be understood that various alternatives to the aspects of this disclosure described herein may be employed in practice. The following claims are intended to define the scope of this disclosure and thereby cover the methods and structures within the scope of those claims and their equivalents.

Claims

1. A computer-implemented method for constructing an optimized classifier for a condition, the method comprising ranking subsets of a plurality of microsatellites as classifiers for the condition in a plurality of optimization cycles, wherein the subsets of the plurality of microsatellites comprise microsatellites in an initial population of microsatellites associated with the condition, thereby identifying an optimized subset of subsets of the microsatellites as the optimized classifier for the condition.

2. The method of claim 1, further comprising comparing microsatellites in a first set of samples from subjects having the condition and microsatellites in a second set of samples from subjects not having the condition, thereby identifying the initial population of microsatellites.

3. The method of claim 1, wherein the ranking comprises comparing the subsets of the microsatellites to microsatellites in samples from subjects having the condition and microsatellites in samples from subjects not having the condition.

4. The method of claim 1, further comprising initialization, wherein the initialization comprises randomly selecting a population of initial subsets of microsatellites from the initial population of microsatellites for ranking in an optimization cycle of the plurality of optimization cycles.

5. The method of claim 1, wherein a population of at least 100 subsets of the initial population of microsatellites is used in the plurality of optimization cycles.

6. A computer-implemented method comprising determining a value of a classifier of a condition for a sample from a subject using a plurality of parameters, wherein each parameter of the plurality of parameters is a statistical measure of correlation of each of a plurality of microsatellites in a sample from a subject having the condition or a sample from a subject not having the condition.

7. A computer-implemented method of determining a genomic age of a subject, the method comprising: a) determining a microsatellite minor allele signature in a first sample from a subject; b) processing the microsatellite minor allele signature with a reference; and c) determining the genomic age of the subject based on the processing.

8. A computer-implemented method comprising: a) determining a plurality of classifiers for a sample from a subject using microsatellites in the sample from the subject; b) processing the plurality of classifiers with a plurality of reference classifiers for a plurality of conditions; and c) determining at least one condition for the subject from the plurality of conditions based on the processing.

9. A non-transitory computer-readable medium comprising executable instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any of the preceding claims.

10. A computer system comprising a hardware processor configured to execute the instructions on the non-transitory computer-readable medium of claim 9.