Methods and systems for microsatellite analysis

By constructing an optimized microsatellite classifier, using genetic algorithms and standard regression analysis, the problem that microsatellite analysis in the prior art is difficult to accurately predict and detect complex health status in the early stages, and the effectiveness of early diagnosis and treatment choices is achieved.

CN114026253BActive Publication Date: 2025-08-12ORBIT GENOMICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080045548.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-22
Filing Date
2020-04-21
Publication Date
2025-08-12
Estimated Expiration
2040-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict, detect and characterize complex multigene health statuses such as cancer, neurological diseases or cardiovascular diseases in microsatellite analysis, resulting in unreliability and difficulty in detection and diagnosis.

Method used

Through computer-implemented methods, an optimized microsatellite classifier is constructed, a classifier is generated using microsatellite sites, combined with genetic algorithms and standard regression analysis, to identify microsatellite populations related to the disease, and to determine the health status of the subjects by comparing the samples of subjects with and without the disease.

Benefits of technology

Accurately predict and detect complex health status early, improve the reliability of diagnosis and effectiveness of treatment choices, and can identify the development risks of health status and the treatment responsiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114026253B_ABST
    Figure CN114026253B_ABST
Patent Text Reader

Abstract

The present disclosure provides methods and systems for classifying microsatellites and minor alleles in a sample. Furthermore, the present disclosure provides methods and systems for generating classifiers for conditions based on microsatellite loci and for performing pan-cancer assays. The methods and systems may involve next-generation sequencing of nucleic acid samples from subjects and genotyping of microsatellite loci in the samples.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 837,109, filed April 22, 2019, the entire contents of which are incorporated herein by reference. Background Art

[0003] Microsatellites (MS) and their changes and instabilities can be the genetic driving force behind many complex, polygenic health conditions, including cancer, neurological diseases, or cardiovascular disease. Currently, predicting, detecting, diagnosing, and characterizing these health conditions through microsatellites can involve matching a patient's microsatellite profile with a database of microsatellites associated with these health conditions. Such methods are only applicable to the later stages of a health condition's progression, which can lead to unreliable and difficult detection, prognosis, diagnosis, treatment selection, and treatment outcomes. Therefore, there remains a need for improved methods for predicting, detecting, and characterizing these health conditions in both the early and late stages through analysis of microsatellite loci. Summary of the Invention

[0004] In one aspect, the present disclosure provides a computer-implemented method for constructing an optimized classifier for a condition, the method comprising ranking a subset of a plurality of microsatellites into a classifier for the condition in a plurality of optimization cycles, wherein the subset of the plurality of microsatellites includes microsatellites from an initial population of microsatellites associated with the condition, thereby identifying an optimized subset of the subset of the plurality of microsatellites as the optimized classifier for the condition. In some aspects, the computer-implemented method further comprises comparing the microsatellites in a first set of samples from subjects having the condition with the microsatellites in a second set of samples from subjects not having the condition, thereby identifying the initial population of microsatellites.

[0005] Sorting can include comparing microsatellites in a first set of samples from subjects with the condition with microsatellites in a second set of samples from subjects without the condition, thereby identifying an initial population of microsatellites. The computer-implemented method can include initialization, wherein initialization includes randomly selecting an initial subset of microsatellites from the initial population of microsatellites for sorting in optimization cycles of a plurality of optimization cycles. A population of at least about 100 subsets of the initial population of microsatellites can be used in the plurality of optimization cycles. The minimum number of microsatellites in a subset of the subset of microsatellites can be 8. The maximum number of microsatellites in a subset of the subset of microsatellites can be 64. In some cases, duplicate microsatellites are not allowed in the subset of the subset of microsatellites. Sorting can include performing a receiver operating characteristic (ROC) analysis using (i) the subset of microsatellites, (ii) the microsatellites in samples from subjects with the condition, and (iii) the microsatellites in samples from subjects without the condition. Sorting in the optimization cycles of the plurality of optimization cycles can include determining the sum of the sensitivity and specificity of the microsatellites in each subset of the subset as a classifier for the condition. An optimization cycle of the plurality of optimization cycles may include adding 10 new subsets of the initial population of microsatellites to subsets from a previous optimization cycle of the plurality of optimization cycles. Seven of the 10 new subsets may be generated by randomly splitting and recombining two randomly selected subsets from the previous optimization cycle, and three of the 10 new subsets may be generated by randomly selecting microsatellites from the initial population of microsatellites. The method may also include discarding 10 subsets of the subsets in the optimization cycle based, at least in part, on having the lowest ranking in the optimization cycle. In some cases, the condition may be the presence or absence of a health state in the subject. The condition may be an increased or decreased likelihood that the subject will develop the health state. The condition may be an increased or decreased likelihood that the subject will benefit from treatment for the health state. In some cases, the condition may be an increased or decreased likelihood that the subject will have an increased risk of adverse effects from treatment for the health state. The condition may be the subject's responsiveness to treatment for the health state. In some cases, the condition may be a prognosis of the subject's health state. In some cases, the health state may be cancer. The cancer may be lung cancer. In other cases, the health state may be a neurological disease or a cardiovascular disease.

[0006] In another aspect, the present disclosure provides a computer-implemented method comprising determining a value of a classifier for a condition in a sample from a subject using a plurality of parameters, wherein each of the plurality of parameters is a statistical measure of the association of each of a plurality of microsatellites in samples from subjects having the condition and / or in samples from subjects not having the condition.

[0007] Multiple weights may include multiple optimal weights. In some aspects, the computer-implemented method may include determining multiple optimal weights. Determining multiple optimal weights may include applying standard regression analysis to the multiple weights. Determining multiple optimal weights may include using a genetic algorithm. Determining a classifier may include using minor allele frequency data. Multiple microsatellites may include at least 10 microsatellites. In some instances, each of the multiple microsatellites is associated with the presence of a condition. The value of the classifier may also include comparing the classifier with a threshold value. In some aspects, the condition may be the presence or absence of a subject's health state, an increase or decrease in the likelihood that the subject will benefit from treatment for the health state, an increase or decrease in the likelihood that the subject will have an increased risk of adverse effects due to treatment for the health state, a subject's responsiveness to treatment for the health state, or a combination thereof. In some cases, the health state is cancer, cardiovascular disease, or a neurological disease. When the health state is cancer, the cancer may be lung cancer.

[0008] In another aspect, the present disclosure provides a computer-implemented method of determining a genomic age of a subject, the method comprising: determining a microsatellite minor allele signature in a first sample from the subject; processing the microsatellite minor allele signature with a reference; and determining the genomic age of the subject based on the processing.

[0009] In some cases, processing includes comparing the microsatellite minor allele signature with a reference. The minor allele signature can be many important alleles on the genetic locus. The number of minor alleles can be supported by at least three next generation sequencing sequence reads. The minor allele signature can be that the total number of readings of the minor allele is standardized to the total number of readings of the major allele at the genetic locus. The method can also include performing next generation sequencing on the first sample from the experimenter to generate the sequence readings of the microsatellite of the experimenter. The first sample can include blood, saliva or tumor. The method can also include, after determining the first genome age, determining the minor allele signature in the second sample from the experimenter. The method can include assessing the minor allele signature in the first sample from the experimenter and the minor allele signature in the second sample from the experimenter, and determining the genome aging rate of the experimenter based on the assessment.

[0010] In another aspect, the present disclosure provides a computer-implemented method comprising: determining a plurality of classifiers for a sample from a subject using microsatellites in the sample from the subject; processing the plurality of classifiers with a plurality of reference classifiers for a plurality of conditions; and determining at least one condition for the subject from the plurality of conditions based on the processing.

[0011] The processing may include comparing a plurality of classifiers to a plurality of reference classifiers for a plurality of conditions. In some cases, at least one of the plurality of conditions includes the presence or absence of at least one of the plurality of health states of the subject. In some cases, at least one of the plurality of conditions includes an increased or decreased likelihood of developing at least one of the plurality of health states of the subject. At least one of the plurality of conditions may include an increased or decreased likelihood that the subject will benefit from treatment of at least one of the plurality of health states of the subject. At least one of the plurality of conditions may include an increased or decreased likelihood that treatment of at least one of the plurality of health states of the subject will result in an increased or decreased risk of an adverse effect in the subject. At least one of the plurality of conditions may include a subject's responsiveness to treatment of at least one of the plurality of health states of the subject. The plurality of health states may include a plurality of cancers, wherein the plurality of cancers includes ovarian cancer, breast cancer, low-grade glioma, glioblastoma, lung cancer, prostate cancer, or melanoma. In some cases, the plurality of health states may include a plurality of neurological diseases or a plurality of cardiovascular diseases.

[0012] In one aspect, the present disclosure provides a non-transitory computer-readable medium comprising executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method for constructing an optimized classifier for a condition, the method comprising ranking a subset of a plurality of microsatellites as a classifier for the condition in a plurality of optimization cycles, wherein the subset of the plurality of microsatellites comprises microsatellites from an initial population of microsatellites associated with the condition, thereby identifying an optimized subset of the subset of the plurality of microsatellites as the optimized classifier for the condition. The computer-implemented method may further comprise comparing microsatellites from a first set of samples from subjects suffering from the condition with microsatellites from a second set of samples from subjects not suffering from the condition, thereby identifying an initial population of microsatellites.

[0013] Sorting can include comparing microsatellites in a first set of samples from subjects with the condition with microsatellites in a second set of samples from subjects without the condition, thereby identifying an initial population of microsatellites. The computer-implemented method can include initialization, wherein initialization includes randomly selecting an initial subset of microsatellites from the initial population of microsatellites for sorting in optimization cycles of a plurality of optimization cycles. A population of at least about 100 subsets of the initial population of microsatellites can be used in the plurality of optimization cycles. The minimum number of microsatellites in a subset of the subset of microsatellites can be 8. The maximum number of microsatellites in a subset of the subset of microsatellites can be 64. In some embodiments, duplicate microsatellites are not allowed in the subset of the subset of microsatellites. Sorting can include performing a receiver operating characteristic (ROC) analysis using (i) the subset of microsatellites, (ii) the microsatellites in samples from subjects with the condition, and (iii) the microsatellites in samples from subjects without the condition. Sorting in the optimization cycles of the plurality of optimization cycles can include determining the sum of the sensitivity and specificity of the microsatellites in each subset of the subset as a classifier for the condition. An optimization cycle of the plurality of optimization cycles may include adding 10 new subsets of the initial population of microsatellites to subsets from a previous optimization cycle of the plurality of optimization cycles. Seven of the 10 new subsets may be generated by randomly splitting and recombining two randomly selected subsets from the previous optimization cycle, and three of the 10 new subsets may be generated by randomly selecting microsatellites from the initial population of microsatellites. The method may also include discarding 10 subsets of the subsets in the optimization cycle based, at least in part, on having the lowest ranking in the optimization cycle. The condition may be the presence or absence of a health state in the subject. The condition may be an increased or decreased likelihood that the subject will develop the health state. The condition may be an increased or decreased likelihood that the subject will benefit from treatment for the health state. The condition may be an increased or decreased likelihood that the subject will have an increased risk of adverse effects from treatment for the health state. The condition may be the responsiveness of the subject to treatment for the health state. The condition may be a prognosis of the subject's health state. The health state may be cancer. The cancer may be lung cancer. The health state may be a neurological disease or a cardiovascular disease.

[0014] In another aspect, the present disclosure provides a non-transitory computer-readable medium comprising executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising determining a value of a classifier for a condition in a sample from a subject using a plurality of parameters, wherein each of the plurality of parameters is a statistical measure of the association of each of a plurality of microsatellites in samples from subjects having the condition and / or in samples from subjects not having the condition.

[0015] The plurality of weights may include a plurality of optimal weights. The computer-implemented method may include determining the plurality of optimal weights. Determining the plurality of optimal weights may include applying a standard regression analysis to the plurality of weights. Determining the plurality of optimal weights may include using a genetic algorithm. Determining the classifier may include using minor allele frequency data. The plurality of microsatellites may include at least 10 microsatellites. Each of the plurality of microsatellites may be associated with the presence of a condition. The value of the classifier may also include comparing the classifier to a threshold value. The condition may be the presence or absence of a subject's health state, an increased or decreased likelihood that the subject will benefit from treatment for the health state, an increased or decreased likelihood that the subject will have an increased risk of adverse effects due to treatment for the health state, a subject's responsiveness to treatment for the health state, or a combination thereof. The health state may be cancer, cardiovascular disease, or a neurological disease. The cancer may be lung cancer.

[0016] In another aspect, the present disclosure provides a non-transitory computer-readable medium comprising executable instructions, which, when executed by one or more processors, cause the one or more processors to perform a method for determining a genomic age of a subject, the method comprising: determining a microsatellite minor allele signature in a first sample from the subject; processing the microsatellite minor allele signature with a reference; and determining the genomic age of the subject based on the processing.

[0017] The processing can include comparing the microsatellite minor allele signature with a reference. The minor allele signature can be many important alleles on the genetic locus. The number of minor alleles can be supported by at least three next generation sequencing sequence reads. The minor allele signature can be that the total number of readings of the minor allele is standardized to the total number of readings of the major allele at the genetic locus. The method can also include performing next generation sequencing on the first sample from the experimenter to generate the sequence readings of the microsatellite of the experimenter. The first sample can include blood, saliva or tumor. The method can also include, after determining the first genome age, determining the minor allele signature in the second sample from the experimenter. The method can include assessing the minor allele signature in the first sample from the experimenter and the minor allele signature in the second sample from the experimenter, and determining the genome aging rate of the experimenter based on the assessment.

[0018] In another aspect, the present disclosure provides a non-transitory computer-readable medium comprising executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method comprising: determining a plurality of classifiers for a sample from a subject using microsatellites in the sample from the subject; processing the plurality of classifiers with a plurality of reference classifiers for a plurality of conditions; and determining at least one condition for the subject from the plurality of conditions based on the processing.

[0019] The processing may include comparing a plurality of classifiers to a plurality of reference classifiers for a plurality of conditions. At least one of the plurality of conditions may include the presence or absence of at least one of the plurality of health states of the subject. At least one of the plurality of conditions may include an increased or decreased likelihood of developing at least one of the plurality of health states of the subject. At least one of the plurality of conditions may include an increased or decreased likelihood that the subject will benefit from treatment of at least one of the plurality of health states of the subject. At least one of the plurality of conditions may include an increased or decreased likelihood that treatment of at least one of the plurality of health states of the subject will result in an increased risk of an adverse effect in the subject. At least one of the plurality of conditions may include a subject's responsiveness to treatment of at least one of the plurality of health states of the subject. The plurality of health states may include a plurality of cancers, wherein the plurality of cancers may include ovarian cancer, breast cancer, low-grade glioma, glioblastoma, lung cancer, prostate cancer, or melanoma. The plurality of health states may include a plurality of neurological diseases or a plurality of cardiovascular diseases.

[0020] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code, which, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0021] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, wherein the computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.

[0022] Other aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are illustrated and described. As will be appreciated, the present disclosure is capable of other and different embodiments, and its several details can be modified in various obvious aspects without departing from the present disclosure. Therefore, the drawings and description are to be considered illustrative in nature, and not restrictive.

[0023] Incorporation by reference

[0024] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated as incorporated by reference. To the extent that publications and patents or patent applications incorporated by reference conflict with the disclosure contained in this specification, the specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained by reference to the following detailed description, which sets forth illustrative embodiments in which the principles of the invention are utilized, and in the accompanying drawings:

[0026] Figure 1 An example of a workflow of a computer-implemented method for generating a microsatellite classifier is illustrated.

[0027] Figure 2 An example of a development process using a computer-implemented method to identify informative microsatellite loci and generate a classifier for a condition is illustrated.

[0028] Figure 3 An example of the validation process for a lung cancer assay is illustrated.

[0029] Figure 4 An example of validating a pan-cancer assay is shown.

[0030] Figure 5 An example of a workflow for analyzing patient samples is illustrated.

[0031] Figure 6 A schematic diagram of the method used to identify and validate medulloblastoma (MB)-associated MS is presented. The method consists of three stages: computational identification of informative MS loci using a training set, validation of microsatellite markers in an independent validation cohort, and downstream analysis of genes associated with these MS. The first stage includes a filter to eliminate MS that varies with age, ethnicity, and sequencing technology.

[0032] 7A to 7D An example of validation and training data is shown. Figure 7A Figure 2. Distribution of metric scores across the training cohort. Figure 7C The distribution of index scores in the validation cohort is shown. Figure 7B ) and validation (102 MB subjects and 428 control subjects) cohorts ( Figure 7D ) performed ROC analysis.

[0033] Figure 8A A pie chart showing the genomic locations of 139 MS informative sites for MB is shown. Figure 8B Gene ontology analysis of informative medulloblastoma MS loci is shown. Figure 8C The protein-protein interaction (PPI) network of 124 genes associated with informative MS sites is shown. The PPI contains 129 nodes and 49 edges, resulting in an enriched network with a p-value of 0.0007.

[0034] Figure 9 Figure 1 shows an example of the genotype distribution and contingency table used in the studies described herein. Genotype distribution of microsatellite marker 242626 at base pair 153645035 on chromosome 1. The p-value for this example is 3.5e -4 The table on the right is a contingency table for the same microsatellite marker.

[0035] Figure 10 An overview of the workflow for identifying age-sensitive MS is shown.

[0036] Figure 11 A schematic overview of the workflow for identifying MS that is sensitive to sequencing technologies is shown.

[0037] Figure 12 An overview of the workflow for identifying ethnically sensitive MS is shown.

[0038] Figure 13 The diagram shows an example of an indicator used to assign scores to samples. Consider a hypothetical sample with the genotypes 22|22, 12|12, and 13|13, respectively, for the aforementioned markers. To apply the indicator to this sample, the difference in frequency of each genotype in the MB and healthy groups is summed, resulting in a score of 0.95. In other words, for each genotype, its frequency in the healthy group is subtracted from its frequency in the MB group; the difference is then summed. Therefore, healthy controls predominantly have negative scores, while affected individuals have positive scores.

[0039] Figure 14 The Youden Index, used to determine the criterion for distinguishing MB from healthy samples, is shown. The Youden Index was used to determine the cutoff value for the receiver operating characteristic (ROC) curve in the training set. The optimal criterion for the 43-marker list was 0.155. The same criterion was used to calculate specificity and sensitivity in the validation cohort.

[0040] Figure 15 A circos diagram showing the chromosomal locations of 43 informative loci indicating MB is shown.

[0041] Figure 16The diagram shows the genotype distribution of microsatellite markers 166663 (an exonic microsatellite located in the RAI gene) and 164048 (an exonic microsatellite located in the BLC6B gene). Adding a CAG triplet can alter protein structure and impair its function, similar to a missense mutation.

[0042] Figure 17 Illustrated is an example of output from a computer-implemented method for reporting the results of a microsatellite analysis to assess a subject's risk of developing cancer.

[0043] Figure 18 A computer system is illustrated that is programmed or otherwise configured to implement the methods provided herein.

[0044] Figure 19 A list of 139 informative germline MSs associated with MBs is shown.

[0045] Figure 20 A list of 43 microsatellite loci in the MB signature set is shown.

[0046] Figure 21 Ingenuity Pathway analysis of informative MBMS sites is shown.

[0047] Figure 22 Illustration of mutations in genes associated with informative MBMS loci in the cBioportalMB cohort.

[0048] Figure 23 Graph depicting analysis of the cBioportal MB cancer study, which revealed 135 gene pairs whose mutations tended to co-occur significantly in the MB cancer risk classifier.

[0049] Figure 24 The thresholds with a confidence interval of 1 standard deviation are shown. Classifiers outside the interval indicate that the subject has the condition (above 0.5) or does not have the condition (below 0.1). Classifier values further away from the threshold carry a stronger indication. DETAILED DESCRIPTION

[0050] I. Overview

[0051] The present disclosure provides computer-implemented methods for generating classifiers for disorders using, for example, microsatellites. Figure 1An example of a workflow for generating a classifier using a computer-implemented method is illustrated. Deoxyribonucleic acid (DNA) sequences are obtained from a database of sequence information (101) from samples of subjects with a condition and from a database of sequence information (102) from reference subjects without the condition. Microsatellite loci from 101 and 102 are identified (genotyped) and compared to each other to reveal only microsatellite populations that are associated or correlated with the condition (103). The microsatellite populations are then further analyzed and weighted to obtain an initial set of microsatellite loci (104) for optimization of the classifier (105). The optimization iteratively ranks how the microsatellites are associated or correlated with the condition. The optimization can be repeated with additional microsatellite groups for additional optimization cycles. In some cases, the microsatellite groups are randomly split and recombined to produce new initial microsatellite groups for additional optimization cycles (106). After the optimization is complete, the computer-implemented method identifies the microsatellite group that is most informative for generating the classifier (107). Additional validation or optimization steps can be performed by analyzing additional samples (e.g., from a database) of subjects known to have or not have the condition (108).After 108, the computer-implemented method can be used to generate a final classifier (109).

[0052] In one aspect, the present disclosure provides an improved computer-implemented method for identifying a set of microsatellites as markers (classifiers) for a disease. The method may also include comparing the microsatellite loci of a first set of samples from a subject suffering from the disease and the microsatellite loci of a second set of samples from a subject not suffering from the disease, thereby identifying an initial population of microsatellite loci (informative loci).

[0053] In some cases, informative sites can be used directly as classifiers. In some cases, a classifier comprising informative sites can indicate the presence or absence of a subject's illness. In some cases, a classifier comprising informative sites can indicate an increase or decrease in the likelihood that a subject's illness will develop. In some instances, a classifier comprising informative sites can indicate an increase or decrease in the likelihood that a subject will benefit from treatment, or an increase or decrease in the likelihood that a subject will have an increased risk of adverse effects due to treatment. In some cases, a classifier comprising informative sites can indicate the responsiveness of a treatment for a subject's illness. In some instances, a classifier comprising informative sites can indicate the prognosis of a subject's illness.

[0054] In some aspects, the initial population of microsatellite loci (informative sites) is used for the genetic algorithm as performed by the method implemented by a computer. Said method can include by comparing the microsatellite subset in the sample from the experimenter suffering from the disease and the microsatellite in the sample from the experimenter not suffering from the disease, and iteratively sorting the subset of the initial population of microsatellite. Said method can include initialization, wherein the initial subset of the random selection subset from the initial population of microsatellite loci. In some instances, the subset of the initial population of about 100 microsatellite loci is used in the whole genetic algorithm (optimization cycle), wherein the microsatellite of the minimum number of subsets is 8, and the microsatellite of the maximum number of subsets is 64. In some instances, iterative sorting includes multiple optimization cycles, wherein multiple optimization cycles include adding 10 new subsets of the initial population of microsatellite to the subset from the previous optimization cycle. 7 of the 10 new subsets can be generated by randomly splitting and recombining 2 subsets randomly selected from the previous optimization cycle, and 3 of the 10 new subsets are generated by randomly selecting microsatellites from the initial population of said microsatellite. In some cases, the method is included in the optimization cycle and the subset is sorted, wherein 10 of the subsets with the lowest sorting in the optimization cycle are discarded, so that 100 subsets of the microsatellite population are maintained in the entire optimization cycle. The genetic algorithm can include performing iterative sorting on all microsatellite combinations to identify the microsatellite loci with the greatest information. The genetic algorithm can improve sensitivity and specificity by removing the microsatellite loci with less information, and selecting or weighting the microsatellite loci with the larger information. In some cases, the illness optimized by the cycle identified by the microsatellite loci can indicate the presence or absence of the subject's health state, the possibility of the subject's health state development increases or decreases, the possibility of the subject benefiting from the treatment of the health state increases or decreases, the possibility of the subject having an increased adverse effect risk due to the treatment of the health state increases or decreases, the responsiveness of the subject to the treatment of the health state, the prognosis of the subject's health state, or a combination thereof.

[0055] On the other hand, the present disclosure provides improved computer-implemented methods, which include using multiple parameters to determine a classifier for a disorder from a sample of an experimenter, wherein each of the multiple parameters is a statistical measure of the correlation of each of a plurality of microsatellites from a sample of an experimenter suffering from the disorder and / or from a sample of an experimenter not suffering from the disorder. In some cases, multiple parameters include optimal weights, such as those determined by standard regression analysis and using a genetic algorithm. In some cases, the classifier is determined by using minor allele frequency data. In some cases, the disorder can indicate the presence or absence of a subject's health state, the likelihood of the subject's health state developing increases or decreases, the likelihood of the subject benefiting from the treatment of the health state increases or decreases, the likelihood of the subject having an increased adverse effect risk due to the treatment of the health state increases or decreases, the responsiveness of the subject to the treatment of the health state, the prognosis of the subject's health state, or a combination thereof. In some cases, the health state is cancer, a disease of the nervous system, or a cardiovascular disease.

[0056] On the other hand, the present disclosure provides the method for using a computer system to determine the minor allele signature in the first sample from the experimenter, comparing the minor allele signature with a reference and determining the genome age of the experimenter based on the comparison. The minor allele signature can be a plurality of minor alleles at a seat, wherein the quantity of the allele is supported by at least one, at least two, at least three or more next generation sequencing sequence readings. In some cases, the minor allele signature is to standardize the total number of readings of the minor allele to the total number of readings of the major allele at the seat. The minor allele signature from the first sample of the experimenter can be compared with the second minor allele signature from the second sample of the same experimenter to determine the genome aging rate.

[0057] The present disclosure provides pan-disorder assays based on classifiers generated using microsatellite loci and optional minor allele information. In some cases, the pan-disorder assay is a pan-cancer assay.

[0058] The term "about" or "approximately" can mean within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, such as limitations of the measurement system. For example, depending on the practice in a given value, "about" can mean within 1 or more standard deviations. Approximately can mean + / - 10%, + / - 5%, + / - 2%, or + / - 1% of a value. Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" as used in the specification and in the claims include plural references. For example, the term "nucleic acid" includes a plurality of nucleic acids, including mixtures thereof.

[0059] II. Methods for Determining Microsatellite Classifiers for Conditions

[0060] The present disclosure provides methods, e.g., computer-implemented methods (e.g., see Figure 2 ) and system. The disease can be the presence or absence of the subject's health state, the possibility of the subject's health state development increases or decreases, the possibility of the subject benefiting from the treatment of the health state increases or decreases, the possibility of the subject having an increased adverse effect risk due to the treatment of the health state increases or decreases, the responsiveness of the subject to the treatment of the health state, the prognosis of the subject's health state, or a combination thereof. The method can include identifying microsatellite loci (genotyping) in samples from subjects with and without the disease. The method can include identifying statistically informative microsatellite loci for the disease. The method can include developing a classification signature for the disease using statistically informative microsatellite loci. The classification signature can be verified and used to test samples from subjects.

[0061] A. Microsatellite genotyping

[0062] The method for identifying microsatellite sorter can comprise carrying out genotyping to the microsatellite loci in the sample from the experimenter suffering from illness and not suffering from illness.In some cases, genotyping comprises the sequence information in analytical database.In some cases, genotyping comprises obtaining sample and analyzing the nucleic acid molecules in sample, for example, by next generation sequencing.

[0063] 1. Sequence Information Database

[0064] In some cases, the method for identifying (for example, genotyping) microsatellite loci can include analyzing the sequence information from one or more databases.One or more databases can include sequence information (for example, sequence reading) from a subject suffering from illness, for example, a subject suffering from cancer or a nucleic acid sample from a cancer cell line.One or more databases can include reference sequences (for example, human genome or a part thereof).One or more databases can include variation or polymorphic sequences of one or more subject colonies.

[0065] One or more databases may include sequence information generated by high throughput or next generation sequencing. One or more databases may include data (e.g., sequence read data) of sequences generated by whole exome sequencing (WES), whole genome sequencing (WGS), or a combination thereof from a sample of a subject. In certain instances, one or more databases include sequence information (e.g., sequence read information) generated from targeted sequencing. Targeted sequencing may include enrichment of targeted sequences from a sample of a subject.

[0066] The database may include sequence information from The Cancer Genome Atlas (TCGA), such as exome data, such as lung cancer exome data. The database may be from the 1000 Genomes Project.

[0067] 2. Sample

[0068] The sample can be a biological sample obtained or derived from one or more subjects. The sample can be processed or fractionated to produce other samples, such as other biological samples. The samples described in this disclosure can include any material from which nucleic acid molecules can be obtained.

[0069] The sample can be obtained from a subject suffering from a condition. The sample can be obtained from a subject suffering from symptoms of a condition. The sample can be obtained from a subject suffering from a condition, but the subject does not have symptoms of the condition. The sample can be obtained from a subject not suffering from the condition. The sample can be obtained from a subject suffering from cancer, a subject suspected of having cancer, or a subject not suffering from or not suspected of having cancer.

[0070] The sample can be obtained or derived from a human subject. The sample can be stored under a variety of storage conditions before processing, such as different temperatures (e.g., at room temperature, under refrigerated or frozen conditions, at 25°C, at 4°C, at 18°C, at -20°C, or at -80°C) or different buffer devices (e.g., EDTA collection tubes, or cell-free DNA or RNA collection tubes).

[0071] Samples can be collected before and / or after treating a subject suffering from cancer. Samples can be obtained from a subject during treatment or a treatment regimen. Multiple samples can be obtained from a subject to monitor the effect of treatment over time. Samples can be taken from a subject known or suspected of having cancer, and the final positive or negative diagnosis of the cancer cannot be obtained via clinical trials. Samples can be taken from a subject suspected of having cancer. Samples can be taken from a subject with unexplained symptoms, such as fatigue, nausea, weight loss, pain, weakness or bleeding. Samples can be taken from a subject with an explanation for the symptoms. Samples can be taken from a subject who develops a risk of cancer due to the presence of family history, age, hypertension or pre-hypertension, diabetes or pre-diabetes, overweight or obesity, environmental exposure, lifestyle risk factors (e.g., smoking, drinking or drug abuse) or other risk factors.

[0072] Sample can be a biological sample from an experimenter.Sample can be whole blood, peripheral blood, plasma, serum, saliva, mucus, urine, semen, lymph fluid, amniotic fluid, fecal extract, cheek swab, cell or other body fluids or tissue, including tissue obtained by surgical biopsy or surgical resection.In some cases, sample can be a cell line derived from a main experimenter (for example, a patient) or an archived experimenter (for example, a patient) sample, such as a preserved sample, such as a paraffin-embedded (FFPE) sample fixed with formalin, or a fresh frozen sample.Sample (for example a biological sample) can be obtained or derived from an experimenter using an ethylenediaminetetraacetic acid collection tube, a DNA or RNA collection tube or a cell-free DNA or cell-free RNA collection tube.Sample, for example a biological sample, can be obtained from a whole blood sample by classification.Sample, for example a biological sample or its derivatives can include cells.Sample (for example a biological sample) can be a blood sample or its derivatives (for example, blood or blood drops collected from a collection tube).

[0073] The sample may comprise one or more assays that can be measured. The sample may comprise one or more nucleic acid molecules. One or more nucleic acid molecules (or any nucleic acid molecules disclosed herein, including primers and probes) may be a polymeric form of nucleotides (e.g., deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs) or analogs thereof) of any length. Analogs may include non-naturally occurring bases, nucleotides attached to other nucleotides other than naturally occurring phosphodiester bonds, or bases attached by bonds other than phosphodiester bonds. Nucleotide analogs include, for example, phosphorothioates, phosphorodithioates, triphosphates, phosphoramidates, boronate phosphates, methylphosphonates, chiral methylphosphonates, 2-O-methyl ribonucleotides, peptide-nucleic acids (PNAs), and the like. The nucleic acid molecule may be deoxyribonucleic acid (DNA). The DNA may be genomic DNA, viral DNA, mitochondrial DNA, plasmid DNA, amplified DNA, circular DNA, circulating DNA, cell-free DNA, or exosomal DNA. In some examples, the DNA is single-stranded DNA (ssDNA), double-stranded DNA, denatured double-stranded DNA, synthetic DNA, and combinations thereof. Circular DNA can be cut or fragmented. The DNA can include coding or non-coding regions of a gene or gene fragment of interest, sites (locus) defined by linkage analysis, exons, or introns. The DNA can be complementary DNA (cDNA). The nucleic acid molecule can be a recombinant nucleic acid, a branched nucleic acid, a plasmid, a vector, or isolated DNA. The nucleic acid molecule can include one or more modified nucleotides, such as methylated nucleotides or nucleotide analogs. Modification of the nucleotide structure can be performed before or after the assembly of the nucleic acid molecule. The nucleotide sequence of the nucleic acid molecule can be interrupted by non-nucleotide components. The nucleic acid molecule can be further modified after polymerization, such as by conjugating or binding to a reporter.

[0074] Nucleic acid molecule can comprise seat, genetic seat or genomic region, and described seat, genetic seat or genomic region can be identified by its position in genome or chromosome.In some examples, seat can be referred to by gene name, and comprises the coding and non-coding region that are associated with the physical region of nucleic acid.Gene can comprise coding region (exon), non-coding region (intron), transcription control or other regulatory region and promoter.In another example, genomic region can be incorporated into intron or exon or intron / exon boundary in naming gene.

[0075] In some instances, the nucleic acid molecule comprises ribonucleic acid (RNA). The RNA may be fragmented RNA. The RNA may be degraded RNA. The RNA may be a microRNA or a portion thereof. The RNA may be an RNA molecule or a fragmented RNA molecule (RNA fragment) selected from the group consisting of microRNA (miRNA), pre-miRNA, pri-miRNA, messenger RNA (mRNA), pre-mRNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), viral RNA, viroid RNA, circular RNA (circRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), pre-tRNA, long noncoding RNA (lncRNA), small nuclear RNA (snRNA), circulating RNA, cell-free RNA, exosomal RNA, vector-expressed RNA, RNA transcripts, synthetic RNA, ribozymes, cell-free RNA, and combinations thereof.

[0076] In some cases, the sample includes cell-free nucleic acid molecules. Cell-free nucleic acid molecules can include, for example, all non-encapsulated nucleic acid molecules derived from a subject's body fluids. Cell-free nucleic acid (cfNA) molecules can be nucleic acids in a biological sample that are not contained in cells (e.g., cell-free RNA (cfRNA) molecules or cell-free DNA (cfDNA) molecules). cfDNA molecules can circulate freely in body fluids (such as in the bloodstream). Cell-free DNA molecules can be circulating tumor DNA, such as cfDNA derived from a tumor.

[0077] The sample can be a cell-free sample. A cell-free sample can be a biological sample that is substantially free of intact cells. A cell-free sample can be a biological sample that is substantially free of cells itself, or can be derived from a sample from which cells have been removed. Examples of cell-free samples include samples derived from blood, such as serum or plasma; urine; or samples derived from other sources, such as semen, sputum, stool, ductal exudate, lymph fluid, or recovered lavage fluid.

[0078] The sample can include germline nucleic acid molecules (e.g., nucleic acids from non-diseased cells or tissues, such as tumors). The sample can include nucleic acid molecules from tumors. In some cases, the sample can include germline nucleic acid molecules (e.g., from non-diseased tissues) and nucleic acid molecules from diseased tissues (e.g., tumors).

[0079] The sample may include a target nucleic acid molecule. The target nucleic acid molecule may be a nucleic acid molecule having a nucleotide sequence, the presence, amount and / or sequence of which, or a change in one or more thereof, is to be determined.

[0080] Nucleic acid molecules (e.g., RNA or DNA) can be extracted from a sample using, for example, a Qiagen QIAmp DNA blood mini kit, MP Biomedical's FastDNA kit protocol, or Norgen Biotek's cell-free biotechnology DNA isolation kit protocol. The extraction method can extract all RNA or DNA molecules from a sample. The extraction method can selectively extract a portion of RNA or DNA molecules from a sample. The RNA molecules extracted from a sample can be converted into DNA molecules by reverse transcription (RT). Reverse transcription can be the generation of deoxy RNA (DNA) from a ribonucleic acid (RNA) template via the action of a reverse transcriptase.

[0081] For example, the quality of the extracted nucleic acids can be analyzed using the BIOANALYZER or NANODROP systems.

[0082] The subject may be a person or an individual. The subject may be a patient. The subject may be a person suffering from or suspected of having cancer. The subject may show symptoms indicating a health or physiological state or condition. The subject may be asymptomatic in terms of health or physiological state or condition. The subjects described herein may include mammals, including any member of the mammalian class: humans, non-human primates (such as chimpanzees, and other apes and monkeys); livestock, such as cattle, horses, sheep, goats, pigs; domestic animals (such as rabbits, dogs and cats); laboratory animals including rodents (such as rats, mice and guinea pigs), etc. In one aspect, the mammal is a human.

[0083] Processing a sample obtained from a subject can include subjecting the sample to a condition sufficient to isolate, enrich, or extract the plurality of nucleic acid molecules, and assaying the plurality of nucleic acid molecules to generate a dataset.

[0084] The sample of experimenter can be analyzed to carry out genotyping to one or more microsatellites.Microsatellite as described herein, microsatellite site or microsatellite region can refer to the tandem repeat of 1 to 6 nucleotide in nucleotide sequence.In some cases, microsatellite comprises the tandem repeat sequence that exceeds 6 nucleotide.One or more microsatellites can be in the upstream of exon, the downstream of exon, in exon, in intergenic sequence, intron, in the region across exon and intron, in 3 ' untranslated region (UTR), 5 ' UTR or any other region in genome and find.In some instances, the pattern of the microsatellite in sample is different from the pattern of the microsatellite in reference sample.The difference of microsatellite pattern can comprise the ratio of single nucleotide polymorphism (SNP), SNP, indel (insert, disappearance, insertion and deletion ratio and combination thereof) or indel and SNP.In some instances, the pattern of microsatellite difference comprises haplotype, for example, the ratio of homozygosity, heterozygosity or minor allele of given site. In the case that the difference pattern of microsatellite is positioned at exon region, difference can comprise non-synonymous SNP, synonymous SNP, frameshift insertion and deletion, non-frameshift insertion and deletion, stop gain and stop loss.Sample can be matched, for example, age, sex or race (for example, Caucasian, African American, Hispanic American).In some cases, sample does not match.In some cases, sample can be accompanied by other clinical metadata, comprise for example health status, cancer, heart or nervous system condition, treatment situation or response or disease stage.Clinical metadata can be associated with microsatellite, to determine whether microsatellite has informativeness with respect to clinical metadata.

[0085] The identity (e.g., genotype) of one or more microsatellites can be obtained by any available method or technology, including next-generation sequencing, high-throughput sequencing, sequencing by synthesis, pyrosequencing, classical Sanger sequencing methods, sequencing by ligation, sequencing by synthesis, sequencing by hybridization, RNA-Seq (Illumina), Illumina sequencing (using reversibly terminated nucleotides), paired-end sequencing, digital gene expression (Helicos), single-molecule sequencing (e.g., single-molecule sequencing by synthesis (SMSS) (Helicos)), Ion Torrent (semiconductor) sequencing (Life Technologies / Thermo-Fisher), massively parallel sequencing, clonal single molecule arrays (Solexa), nanopore sequencing, Pacific Biosciences SMRT sequencing, shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and any other sequencing method.

[0086] Next generation sequencing can include sample multiplexing. Sample multiplexing can be at least or up to or about 12 samples, 24 samples, 48 samples, 96 samples, 192 samples, 384 samples, 768 samples, or 1536 samples. Sequencing depth can be from about 1-fold to about 10-fold, about 10-fold to about 100-fold, about 100-fold to about 500-fold, or about 500-fold to about 1000-fold.

[0087] The sequencing depth can be at least, at most, or about 1-fold, 5-fold, 10-fold, 50-fold, 100-fold, 200-fold, 250-fold, 300-fold, 400-fold, or 500-fold. The base calling consensus accuracy can be at least 95%, 96%, 97%, 98%, 99%, or greater than about 99%. The quality score can be at least Q10 (e.g., error rate less than 1:10, inferred base calling accuracy greater than 90%), Q20 or greater (e.g., error rate less than 1:100, inferred base calling accuracy greater than 99%), Q30 or greater (e.g., error rate less than 1:1000, inferred base calling accuracy greater than 99.9%), Q40 or greater (e.g., error rate less than 1:10,000, inferred base calling accuracy greater than 99.99%), or Q50 or greater (e.g., error rate less than 1:100,000, inferred base calling accuracy greater than 99.999%). The assembly method can generate at least 95%, 96%, 97%, 98%, or 99% accuracy for calling microsatellite genotypes in next-generation sequencing datasets.

[0088] After sequencing the nucleic acid molecules, suitable bioinformatics processing can be performed on the sequence reads. For example, the sequence reads can be aligned with one or more reference genomes (e.g., genomes of one or more species, such as the human genome). The aligned sequence reads can be quantified at one or more sites (e.g., one or more microsatellite sites).

[0089] In some aspects, identification (e.g., genotyping) of one or more microsatellites includes amplifying the nucleotide sequence of one or more microsatellite sites, e.g., by performing polymerase chain reaction (PCR), e.g., using primers, e.g., specific primers, flanking one or more microsatellite sites, and e.g., by capillary electrophoresis or sequencing to assess the amplified fragment. PCR can be quantitative PCR (qPCR), digital PCR, or reverse transcriptase PCR. Amplification or amplification can increase the size or quantity of nucleic acid molecules. Amplified nucleic acid molecules can be single-stranded or double-stranded. Amplification can include generating one or more copies or amplification products of nucleic acid molecules. For example, amplification can be performed by extension (e.g., primer extension) or connection. Amplification can include performing a primer extension reaction to generate a chain complementary to a single-stranded nucleic acid molecule, and generating one or more copies of a chain and / or a single-stranded nucleic acid molecule in some cases.

[0090] Amplification of nucleic acid molecules (e.g., nucleic acid molecules comprising one or more microsatellite loci) can be performed, for example, using any of the following nucleic acid amplification methods: loop-mediated isothermal amplification (LAMP), nucleic acid sequence-based amplification (NASBA), self-sustained sequence replication (3SR), rolling circle amplification (RCA), recombinase polymerase amplification (RPA), multiple displacement amplification (MDA), helicase-dependent amplification (HDA), strand displacement amplification (SDA), nicking enzyme amplification reaction (NEAR), exponential amplification reaction (EXPAR), polymerase spiral reaction (PSR), isothermal multiple displacement amplification (IMDA), branched amplification method (RAM), single primer isothermal amplification (SPIA), signal-mediated RNA amplification technology (smart), beacon-assisted detection amplification (BADAMP), hinge-initiated primer-dependent nucleic acid amplification (HIP), SMART amplification process (SmartAmp), hybridization chain reaction (HCR), a substrate-mediated strand displacement (TMSD), ligase chain reaction, digital PCR (dPCR), droplet digital PCR (ddPCR), or transcription-mediated amplification. Amplification can relate to multiple amplification, for example, using AMPLISEQ.In some cases, RNA is converted into cDNA by reverse transcription before amplification.Determination reading can include quantitative PCR (qPCR) value, digital PCR (dPCR) value, digital droplet PCR (ddPCR) value, fluorescence value etc., or its normalized value.Other determinations that can be used for the method provided herein include immunoassay, electrochemical determination, surface enhanced Raman spectroscopy (SERS), determination based on quantum dots (QD), molecular inversion probe, determination based on CRISPR / Cas (for example, CRISPR- typing PCR (ctPCR), specific high sensitivity enzyme reporter unlocking (SHERLOCK), DNA endonuclease targeting CRISPR trans reporter (DETECTR), CRISPR-mediated analog multi-event recording equipment (CAMERA)) and laser transmission spectroscopy (LTS).

[0091] Multiplex amplification can include amplifying about 10 to about 50 targets, about 50 to about 100 targets, about 100 to about 500 targets, or about 500 to about 1000 targets. Adapters (e.g., universal adapters) can be added (e.g., ligated) to nucleic acid molecules to facilitate amplification and / or sequencing, e.g., on an ILLUMINA sequencing platform. Universal primers can be combined with universal adapters for amplification.

[0092] In some embodiments, the present invention relates to a method for the analysis of RNA or DNA molecules. A plurality of samples can be analyzed, and each multiplexed sample can have a barcode. The RNA or DNA molecules separated or extracted from the sample can be marked, for example, with an identifiable label to allow the multiplexing of a plurality of samples. Any number of RNA or DNA samples can be multiplexed. For example, the multiplex reaction can include RNA or DNA from at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more than 100 initial samples. For example, a plurality of samples can be labeled with a sample barcode so that each DNA molecule can be traced back to the sample (and experimenter) in which the DNA molecule originated. Such tags can be attached to RNA or DNA molecules by attachment or PCR amplification via primers.

[0093] In some cases, a bait set (e.g., a hybridization probe, such as SURESELECT or SEQCAP) is used to obtain a target, such as a target nucleic acid molecule. The target may include RNA and / or DNA. The length of the hybridization probe may be at least 15, 25, 50, 75, 100, 120, or 150 bases. The length of the hybridization probe may be 15 to 50 bases, 50 to 100 bases, or 100 to 150 bases. The probe may be a nucleic acid molecule (e.g., RNA or DNA) having sequence complementarity with a nucleic acid sequence (e.g., RNA or DNA) of one or more sites (e.g., one or more microsatellites). Using a probe that is selective for one or more sites (e.g., one or more microsatellites) to assay a sample may include using array hybridization (e.g., microarray-based), polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing).

[0094] In some aspects, analyzing nucleic acid molecules includes performing next generation sequencing. In some cases, microsatellite sequencing can be performed directly, for example, without the need to perform amplification. Next generation sequencing methods can include full genomes, full exomes, and partial genomes or exomes. Next generation sequencing methods can be used for targeted sequences, enrichment sequences, or combinations thereof.

[0095] In some instances, before sequencing and downstream analysis, enrichment is performed with an enrichment kit. In some cases, enrichment is performed with an enrichment kit to enrich for microsatellite sites verified by a genetic algorithm. Using an enrichment kit can increase the number of allelopathic types or genotypes that can be called in readings, and can increase the ability to analyze a larger percentage of informative sites for a given sample. The enrichment kit can include an enrichment array or probe that hybridizes with the targeting sequence of the microsatellite and the flanking sequence on either or both sides of the microsatellite. In some cases, compared to the number of callable genotypes that can be obtained without the use of an enrichment kit, the use of enrichment increases the number of callable genotypes by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or more. In some instances, compared to the number of callable genotypes that can be obtained without the use of an enrichment kit, the use of an enrichment kit increases the number of callable genotypes by at least 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times or more. In some aspects, the enrichment kits disclosed herein include compositions that can be used to perform the methods described herein.

[0096] 3. Genotyping Algorithm

[0097] Algorithms can be used to carry out genotyping of microsatellites. Said algorithm can use the Bayesian model selection that for example is guided by the error model derived from experience, or discretization Gaussian mixture (for example, GenoTan). For example, said algorithm can be Repeatseq. Method based on dynamic programming or heuristic method can be used to carry out genotyping of microsatellites. Other tools for microsatellite genotyping include PHOBOS, MISA, Tandem Repeats Finder, FullSSR or bMSISEA.

[0098] B. Identification of informative microsatellites

[0099] Identifying informative microsatellites can include identifying a first set of microsatellite loci from a sample of a subject with a disorder and identifying a second set of microsatellite loci from a sample of a subject without a disorder. In some cases, the second set of microsatellite loci can be obtained from a database of reference sequences.

[0100] 1. Statistics

[0101] The difference between the first group of microsatellite loci and the second group of microsatellite loci can be detected, and statistically compared with one or more statistical tests (such as t test, Z test, analysis of variance, regression analysis, Mann-Whitney-Wilcoxon test, chi-square test, correlation, Fisher's exact test, Bonferroni correction and Benjamini-Hochberg test). In some cases, generalized Fisher's exact test is used to quantify statistical differences. In some cases, Benjamini-Hochberg multiple testing correction is used to control false discovery rate.

[0102] 2. Microsatellite Filtering

[0103] If for example from the sample of the experimenter suffering from illness and from the sample and factor mismatch of the experimenter not suffering from described illness, then can filter microsatellite to control any number of factors, for example age, race, sex, order-checking scheme (for example, WSG, WES or targeted sequencing).Microsatellite with potential deviation can be excluded from analyzing subsequently.Other filter for filtering microsatellite can comprise the length of microsatellite repeat motif, the total length of microsatellite (for example, the copy number of motif), the sequence of motif (for example, only using those with high GC content) and the purity of microsatellite, for example, if it has any base that can interrupt the perfect copy set of motif.In some instances, microsatellite can be filtered by their position in genome (for example, exon group, intron, intergenic region or untranslated region).Filtering can comprise filtering by the gene near microsatellite or functional element.

[0104] 3. Score the samples

[0105] Statistical tests can produce receiver operating characteristic (ROC) curves, wherein the area under the ROC curve is referred to as area under the curve (AUC). AUC can be determined to assess the accuracy of the group of microsatellite loci comparisons. A larger AUC can indicate the higher accuracy of the correlation or correlation of the difference between a disease and the first group of microsatellite loci and the second group of microsatellite loci. The ROC curve can sensitively (e.g., true positive) determine the ratio and specificity (e.g., true negative) of the correlation or correlation of the difference between a disease and the first group of microsatellite loci and the second group of microsatellite loci. Sensitivity, also referred to as true positive rate, recall rate or probability of detection, can measure the ratio of the actual positive that is correctly identified as the presence or absence of a certain disease. Sensitivity can be quantified to avoid false negatives by calculating the true positive number divided by the sum of the true positive number and the false negative number. Specificity, also referred to as true negative rate, can measure the ratio of the actual negative that is correctly identified as the presence or absence of a disease. Specificity can be quantified to avoid false positives by calculating the true negative number divided by the sum of the true negative number and the false positive number.

[0106] In some instances, the condition has a statistically significant correlation or association with a first set of microsatellite loci that is different from a second set of microsatellite loci with a statistical accuracy of at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. In some cases, the condition has a statistically significant correlation or association with a first set of microsatellite loci that is different from a second set of microsatellite loci with a statistical specificity of at least 0.70, 0.80, 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99, and a statistical sensitivity of at least 0.70, 0.80, 0.85, or 0.99.

[0107] In some instances, identifying informative microsatellites includes identifying a first set of microsatellite sites from a database, the database including nucleic acid sequences obtained from a subject with a disease, such as cancer type sequences from the Cancer Genome Atlas Project (TCGA), and identifying a second set of microsatellite sites from a reference database (e.g., hg19 or 1000 Genome Project). The type of cancer (e.g., breast cancer) can be based on a subtype such as staging, morphology, histology, gene expression, receptor spectrum, mutation spectrum, aggressiveness, prognosis, malignant characteristics, etc. Cancer types and cancer subtypes can be applied at a more refined level, for example, to distinguish cancer or cancer subtypes of a tissue type, for example, defined according to mutation spectrum or gene expression. Cancer stage can refer to classification of cancer types based on histological and pathological features associated with disease progression. In some instances, the group of microsatellite sites is obtained from a database including nucleic acid sequences, including nucleotide variations or polymorphisms. In some cases, the first group of microsatellite sites is obtained from a sample suffering from the disease and is compared with the second group of microsatellite sites obtained from the database.

[0108] 4. Symptoms

[0109] In some cases, the condition associated or related to the difference in the set of microsatellite loci can indicate the presence or absence of a subject's health state, an increase or decrease in the likelihood of the subject's health state developing, an increase or decrease in the likelihood that the subject will benefit from the treatment of the health state, an increase or decrease in the likelihood that the subject will have an increased risk of adverse effects due to the treatment of the health state, an increase or decrease in the likelihood that the subject will have an increased risk of adverse effects due to the treatment of the health state, the responsiveness of the subject to the treatment of the health state, the prognosis of the subject's health state, or a combination thereof. In some cases, the health state is cancer. In some cases, the cancer is solid or hematological malignancy. In some cases, the cancer is metastatic, recurrent, or refractory. Cancers that may be associated or correlated with different sets of microsatellite loci include acute myeloid leukemia (LAML or AML), acute lymphoblastic leukemia (ALL), adrenocortical carcinoma (ACC), bladder urothelial carcinoma (BLCA), brainstem glioma, low-grade glioma (LGG), brain tumor, breast cancer (BRCA), bronchogenic tumor, Burkitt lymphoma, cancer of unknown primary site, carcinoid, cancer of unknown primary site, atypical teratoid / rhabdoid tumor of the central nervous system, embryonal tumor of the central nervous system, squamous cell carcinoma of the cervix,Endocervical adenocarcinoma (CESC) cancer, childhood cancer, bile duct cancer (CHOL), chordoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myeloproliferative disorders, colon (adenocarcinoma) cancer (COAD), colorectal cancer, craniopharyngioma, cutaneous T-cell lymphoma, endocrine islet cell tumor, endometrial cancer, ependymoma, ependymoma, esophageal cancer (ESCA), esthesioneuroblastoma, Ewing sarcoma, extracranial germ cell tumor, extragonadal germ cell tumor, extrahepatic bile duct cancer, gallbladder cancer, stomach (gastric) cancer, gastrointestinal head and neck cancer (HNSD), cardia cancer, Hodgkin lymphoma, hypopharyngeal cancer, intraocular melanoma , pancreatic islet cell tumor, Kaposi's sarcoma, kidney cancer, Langerhans cell histiocytosis, laryngeal cancer, lip cancer, liver cancer, lymphoid neoplasms diffuse large B-cell lymphoma [DLBCL], malignant fibrous histiocytoma bone cancer, medulloblastoma, medullary epithelioma, melanoma, Merkel cell carcinoma, Merkel cell skin cancer, mesothelioma (MESO), metastatic squamous neck cancer with occult primary, oral cancer, multiple endocrine neoplasia syndrome, multiple myeloma, multiple myeloma nasal cancer, nasopharyngeal cancer, neuroblastoma, non-Hodgkin lymphoma, non-melanoma skin cancer, non-small cell lung cancer, oral cancer, oral cancer, oropharyngeal cancer, Osteosarcoma, other brain and spinal cord tumors, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, ovarian low malignant potential tumor, pancreatic cancer, papillomatosis, paranasal sinus cancer, parathyroid cancer, pelvic cancer, penile cancer, pharyngeal cancer, pheochromocytoma and paraganglioma (PCPG), intermediately differentiated pineal parenchymal tumor, pineoblastoma, pituitary tumor, plasmacytoma / tumor, primary central nervous system (CNS) lymphoma, primary hepatocellular carcinoma, prostate cancer such as prostatic adenocarcinoma (PRAD), rectal cancer, kidney cancer, renal cell (kidney) cancer, renal cell carcinoma, respiratory tract cancer, retinoblastoma, rhabdomyosarcoma, salivary gland cancer In some aspects,Cancer types include acute lymphoblastic leukemia, acute myeloid leukemia, bladder cancer, breast cancer, brain cancer, cervical adenocarcinoma, bile duct cancer, colon cancer, colorectal cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, glioma, glioblastoma, head and neck cancer, kidney cancer, liver cancer, lung cancer, lymphoid neoplasia, melanoma, myeloid neoplasia, ovarian cancer, pancreatic cancer, pheochromocytoma and paraganglioma, prostate cancer, rectal cancer, squamous cell carcinoma, testicular cancer, stomach cancer, or thyroid cancer.

[0110] In some cases, the health condition is lung cancer or a subtype of lung cancer. Lung cancers that may be associated or correlated with different sets of microsatellite loci include non-small cell lung cancer (NSCLC) (e.g., lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), and large cell carcinoma), small cell lung cancer (SCLC), and lung carcinoid tumors.

[0111] In some cases, the health state is a neurological disease. Examples of neurological diseases that may be associated or correlated with the differences in the set of microsatellite loci include myotonic dystrophy, fragile X-associated tremor / ataxia syndrome, spinocerebellar ataxia, Kennedy disease, Huntington's disease, spinobulbar muscular atrophy, progressive myoclonic epilepsy 1 (Unverricht–Lundborg disease), fragile X syndrome, fragile XE syndrome, dentate-pallidular atrophy, Friedrich's ataxia, oculopharyngeal muscular dystrophy, fragile X-associated primary ovarian insufficiency, Huntington's disease-like 2, C9ORF72-related frontotemporal dementia, and amyotrophic lateral sclerosis. The health state may be autism.

[0112] In some cases, the health state is inflammatory bowel disease (IBD), which can include gastrointestinal diseases of the gastrointestinal tract. Non-limiting examples of IBD include Crohn's disease (CD), ulcerative colitis (UC), indeterminate colitis (IC), microscopic colitis, diversion colitis, Behcet's disease and other indeterminate forms of IBD. In some instances, IBD includes fibrosis, fibrostenosis, stricture and / or penetrating disease, obstructive disease or refractory disease (e.g., mrUC, refractory CD), perianal CD or other complex forms of IBD.

[0113] In some instances, the health condition is cardiovascular disease, which can include coronary artery disease (CAD), rheumatic heart disease, congenital heart disease, cardiomyopathy, cardiac tumors, vascular tumors, heart valve disease, disease of the inner layer of the heart, stroke, aortic aneurysm, peripheral arterial disease, deep vein thrombosis (DVT), or pulmonary embolism.

[0114] In some cases, the health condition is a metabolic disease or disorder and can include acid-base imbalance, metabolic brain disease, calcium metabolism disorder, DNA repair defect disorder, glucose metabolism disorder, hyperlactatemia, iron metabolism disorder, lipid metabolism disorder, malabsorption syndrome, metabolic syndrome X, inborn errors of metabolism, mitochondrial disease, phosphorus metabolism disorder, porphyria, protein deposition defect, metabolic skin disease, wasting syndrome, or water and electrolyte imbalance.

[0115] In some cases, the health condition is an autoimmune disease or disorder, which can include achalasia, Addison's disease, adult-onset Still's disease, agammaglobulinemia, alopecia areata, amyloidosis, ankylosing spondylitis, anti-GBM / anti-TBM nephritis, antiphospholipid syndrome, autoimmune angioedema, autoimmune autonomic dysfunction, autoimmune encephalomyelitis, autoimmune hepatitis, autoimmune inner ear disease (AIED), autoimmune myocarditis, autoimmune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune retinopathy, autoimmune urticaria, axonal and neuronal neuropathy (AMAN), Baló disease, Behçet's chronic inflammatory demyelinating polyneuropathy (CIP), and other autoimmune diseases. DP), chronic recurrent multifocal osteomyelitis (CRMO), Churg-Strauss syndrome (CSS) or eosinophilic granulomatosis (EGPA), cicatricial pemphigoid, Cogan syndrome, cold agglutinin disease, congenital heart block, Coxsackie myocarditis, CREST syndrome, Crohn's disease, dermatitis herpetiformis, dermatomyositis, Devic's disease (neuromyelitis optica), discoid lupus, Dressler syndrome, endometriosis, eosinophilic esophagitis (EoE), eosinophilic fasciitis, erythema nodosum purpura (HSP), herpes gestationis or pemphigoid gestationis (PG), hidradenitis suppurativa (HS) (acne inversa), hypoglobulinemia, IgA nephropathy, IgG4-related sclerosing disease, immune ITP, IBM, cystitis interstitialis, juvenile arthritis, juvenile diabetes mellitus (type 1 diabetes), juvenile myositis (JM), Kawasaki disease, Lambert-Eaton syndrome, leukocytoclastic vasculitis, lichen planus, lichen sclerosus, conjunctivitis leucoderma, linear multiple sclerosis, myasthenia gravis, myositis, narcolepsy, neonatal lupus, neuromyelitis optica, neutropenia, ocular cicatricial pemphigoid, PPT neuritis, palindromic rheumatism (PR), PANDAS, paraneoplastic cerebellar degeneration (PCD), paroxysmal nocturnal hemoglobinuria (PNH), Parry-Romberg syndrome, pars planitis (peripheral uveitis), Pars Positive-Turner syndrome, pemphigus, peripheral neuropathy, perivenous encephalomyelitis, pernicious anemia (PA), POEMS syndrome, polyarteritis nodosa, polyglandular syndrome type I and II, Raynaud's phenomenon, reactive arthritis, reflex sympathetic dystrophy, relapsing polychondritis, restless legs syndrome (RLS), retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, sarcoidosis, Schmidt syndrome, scleritis, scleroderma, Sjögren syndrome, sperm and testicular autoimmunity, stiff-man syndrome (SPS), subacute bacterial endocarditis (SBE), Susac syndrome, sympathetic ophthalmia (SO), Takayasu arteritis, temporal arteritis / giant cell arteritis, thrombocytopenic purpura (TTP),Tolosa-Hunt syndrome (THS), transverse myelitis, type 1 diabetes, ulcerative colitis (UC), undifferentiated connective tissue disease (UCTD), uveitis, vasculitis, vitiligo, or Vogt-Koyanagi-Harada disease.

[0116] C. Developing Classification Signatures

[0117] The present disclosure provides computer-implemented methods for generating a classifier for a condition from a sample from a subject (e.g., see Figure 2 and Figure 3 ). An informative microsatellite locus list can be generated by statistically analyzing a sample obtained from a first group of subjects suffering from an illness or derived and / or a sample obtained or derived from a second group of subjects who have never suffered from an illness (for example, a cancer such as lung cancer). Can be on a multiplex platform to DNA sequencing from two groups of samples. In some cases, targeted sequencing is performed when some targets of enrichment are present. The quality of sequencing results can then be analyzed and mapped to reveal the difference between cancer sample and control or reference. This difference can then be analyzed using a computer-implemented method to generate a sorter. The sorter can be further optimized and verified with the sample obtained or derived from the subject suffering from the illness and / or the sample obtained or derived from the subject who has never suffered from the illness. In some respects, the list of informative genetic markers except microsatellites can be generated by these methods for developing classification signatures.

[0118] The condition may indicate the presence or absence of a subject's health state. In some cases, the condition indicates an increase or decrease in the likelihood of the subject's health state developing. In some instances, the condition may indicate an increase or decrease in the likelihood that the subject will benefit from treatment, or an increase or decrease in the likelihood that the subject will have an increased risk of adverse effects due to treatment (the classifier for the condition may be used as a companion diagnosis for a therapeutic agent). In some cases, the condition may indicate the responsiveness of the subject to the treatment of a health state. In some instances, the condition indicates the prognosis of the subject's health state. In some cases, the classifier may be a value such as a quantity. For example, the value may indicate an increase or decrease in likelihood (e.g., a probability value between 0 and 1). The value of the classifier (e.g., quantity) may be compared to a threshold value (e.g., quantity). In some instances, the distance between the classifier value and the threshold value may indicate an increase in confidence or probability that the condition is true or not. In some cases, a call is made when the standard deviation of the classifier value from the threshold value is approximately 0.5, 1, 1.5, 2, 2.5, 3, or more than 3. Figure 24 ).

[0119] The computer-implemented method for generating a classifier can perform processing, combining, statistical evaluation or further analysis of the results, or any combination thereof. The computer-implemented method can include supervised or unsupervised learning methods, including support vector machines (SVMs), neural networks, random forests, clustering algorithms (or software modules), gradient boosting, linear regression, logistic regression and / or decision trees. A supervised learning algorithm can be an algorithm that relies on using a set of labeled, paired training data examples to infer the relationship between input data and output data. An unsupervised learning algorithm can be an algorithm for drawing inferences from a training data set to output data. An unsupervised learning algorithm can include cluster analysis, which can be used for exploratory data analysis to discover hidden patterns or groupings in process data. An example of an unsupervised learning method is principal component analysis. Principal component analysis can include reducing the dimensionality of a set of one or more variables. The number of dimensions of a given set of variables may be at least 1, 5, 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or greater than 1800. The number of dimensions of a given set of variables may be at most 1800, 1600, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less than 10.

[0120] The computer-implemented method may include performing statistical techniques. In some instances, the statistical techniques may include linear regression, classification, resampling methods, subset selection, shrinkage, dimensionality reduction, nonlinear models, tree-based methods, support vector machines, unsupervised learning, or any combination thereof.

[0121] Linear regression can be a method for predicting a target variable by fitting the best linear relationship between the dependent variable and the independent variable. The best fit can correspond to least squares, minimizing the sum of all distances between the shape at each point and the actual observed value. Linear regression can include simple linear regression and multiple linear regression. Simple linear regression can use a single independent variable to predict the dependent variable. Multiple linear regression can use more than one independent variable to predict the dependent variable by fitting the best linear relationship.

[0122] Classification can be a data mining technique that assigns categories to data collections in order to achieve accurate prediction and analysis. Classification techniques can include logistic regression and discriminant analysis. Logistic regression can be used when the dependent variable is dichotomous (binary). Logistic regression can be used to discover and describe the relationship between a binary variable and one or more independent variables at the nominal, ordinal, interval, or ratio level. Resampling can be a method that includes drawing repeated samples from an original data sample. In some cases, resampling may not involve using a common distribution table to calculate approximate probability values. Resampling can generate a unique sampling distribution based on actual data. In some cases, resampling can use experimental methods rather than analytical methods to generate a unique sample distribution. Resampling techniques can include bootstrapping and cross-validation. Bootstrapping can be performed by sampling with replacement from the original data and using the "unselected" data points as test cases. Cross-validation can be performed by dividing the training data into multiple parts.

[0123] Subset selection can identify a subset of the predictor variables that are related to the response. Subset selection can include best subset selection, forward stepwise selection, backward stepwise selection, hybrid methods, or any combination thereof. In some instances, shrinkage fitting involves fitting a model that includes all the predictor variables, but the estimated coefficients are shrunk towards zero relative to the least squares estimates. This shrinkage can reduce variance. Shrinkage can include ridge regression and the lasso. Dimension reduction can simplify the problem of estimating n + 1 coefficients to a simpler problem of estimating m + 1 coefficients, where m < n. It can be obtained by computing n different linear combinations or projections of the variables. Then, these n projections can be used as predictor variables to fit a linear regression model, for example, by least squares. Dimension reduction can include principal component regression and partial least squares. Principal component regression can be used to derive a set of low-dimensional features from a large set of variables. The principal components used in principal component regression can capture the maximum variance in the data using linear combinations of the data in subsequent orthogonal directions. Partial least squares can be used as a supervised alternative to principal component regression because partial least squares can utilize the response variable to identify new features.

[0124] Nonlinear regression can be a form of regression analysis in which the observed data are modeled by a function that is a nonlinear combination of the model parameters and depends on one or more independent variables. Nonlinear regression can include step functions, piecewise functions, splines, generalized additive models, or any combination thereof.

[0125] Tree-based methods can be used for both regression and classification problems. Regression and classification problems can involve stratifying or partitioning the space of predictor variables into many simple regions. Tree-based methods can include bagging, boosting, random forests, or any combination thereof. Bagging can reduce the variance of predictions by generating additional training data from the original dataset using repeated combinations to generate multiple steps of the same size / volume as the original data. Boosting can calculate the output using several different models and then average the results using a weighted average method. The random forest algorithm can draw random bootstrap samples of the training set. Support vector machines can be used for classification techniques. Support vector machines can involve finding a hyperplane that best separates two classes of points with the largest margin. Support vector machines can constrain the optimization problem so that maximizing the margin is constrained by its ability to perfectly classify the data.

[0126] Unsupervised methods can be methods that draw inferences from a dataset that includes input data without labeled responses. Unsupervised methods can include clustering, principal component analysis, k-means clustering, hierarchical clustering, or any combination thereof.

[0127] 1. Genetic Algorithm

[0128] In some aspects, the computer-implemented method for generating a sorter includes the use of a genetic algorithm. The method can include identifying a microsatellite locus different from the microsatellite locus in a sample that does not suffer from the disease from a sample that suffers from the disease, generating an initial population (informative site) of a subset of microsatellite loci that is relevant or associated with the disease. The genetic algorithm can be used to determine a classification signature based on the informative site. The genetic algorithm can select the microsatellite locus subset with the largest informativeness to be included in the final sorter. The genetic algorithm can assign a weight to each subset. Weighting can be combined with other weighting schemes, such as being proportional to the relative risk of each microsatellite locus. Each subset of microsatellites can be iteratively sorted based on the correlation or association of the subset and the disease. The subset of the initial population of microsatellite loci can then be optimized by comparing the initial population with another sample obtained or derived from a subject suffering from the disease and / or a subject not suffering from the disease. In some cases, an initial population of about 100 subsets is used in the optimization. In some cases, an initial population of at least 100, 200, 300, 400, or 500 subsets is used in the optimization. In some instances, optimization comprises at least one cycle of comparing about 100 subsets to the additional samples. In some instances, optimization comprises multiple cycles of comparing about 100 subsets to the additional samples. Each subset can include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 microsatellites.

[0129] In some cases, the iterative sorting comprises performing statistical analysis on the subset to perform receiver operating characteristic (ROC) analysis, thereby obtaining accuracy, sensitivity and specificity when determining the presence or absence of the disease in other samples. The predetermined number (e.g., 10) of the subsets with the worst performance or the lowest ranking in terms of the presence or absence of the disease can be identified and discarded. In order to maintain a constant number of subsets before each optimization cycle begins, a new subset can be added to the subset population. In some cases, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more new subsets are generated by randomly splitting and reorganizing 2 randomly selected subsets from the previous optimization cycle. In some instances, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more new subsets are randomly selected from the previous optimization cycle. In some instances where 10 new subsets were added, 3 were generated by randomly splitting and recombining 2 randomly selected subsets from the previous optimization cycle, and 7 were randomly selected from the subsets of the previous optimization cycle. In some instances where 10 new subsets were added, 4 were generated by randomly splitting and recombining 2 randomly selected subsets from the previous optimization cycle, and 6 were randomly selected from the subsets of the previous optimization cycle. In some instances where 10 new subsets were added, 5 were generated by randomly splitting and recombining 2 randomly selected subsets from the previous optimization cycle, and 5 were randomly selected from the subsets of the previous optimization cycle. In some instances where 10 new subsets were added, 6 were generated by randomly splitting and recombining 2 randomly selected subsets from the previous optimization cycle, and 4 were randomly selected from the subsets of the previous optimization cycle. In some instances where 10 new subsets were added, 6 were generated by randomly splitting and recombining 2 randomly selected subsets from the previous optimization cycle, and 4 were randomly selected from the subsets of the previous optimization cycle. In some instances where 10 new subsets are added, 7 are generated by randomly splitting and recombining 2 randomly selected subsets from the previous optimization cycle, and 3 are randomly selected from subsets from the previous optimization cycle. Copies of the new subsets may be included in the optimization cycle. In some cases, copies of the new subsets are not included in the optimization cycle.

[0130] In some cases, the number of subsets discarded at the end of each optimization cycle is the same as the number of subsets added to the subsets before each optimization cycle. In some cases, the five lowest-ranked subsets are discarded at the end of each optimization cycle, while five new subsets are added before each optimization cycle. In some cases, the ten lowest-ranked subsets are discarded at the end of each optimization cycle, while ten new subsets are added before each optimization cycle. In some cases, the twenty lowest-ranked subsets are discarded at the end of each optimization cycle, while twenty new subsets are added before each optimization cycle. In some cases, the fifty lowest-ranked subsets are discarded at the end of each optimization cycle, while fifty new subsets are added before each optimization cycle.

[0131] In some aspects, the computer-implemented method for generating a sorter comprises determining a statistically unweighted subset of microsatellites. In some aspects, the computer-implemented method for generating a sorter comprises determining a statistically weighted subset of microsatellites. In some cases, the weight subset is weighted by relative risk, hazard ratio or odds ratio. The sorter can be unweighted or weighted. In some cases, the sorter generated by the above-mentioned computer-implemented method can be based on genetic markers except microsatellites. In some cases, the sorter can be based on other genomic information, such as single nucleotide polymorphisms (SNPs) or genetic aberrations, such as copy number aberrations, insertions and deletions etc. In some cases, the sorter can be based on the identity of the gene at microsatellite place.

[0132] After the optimization cycle is complete, the computer-implemented method can include determining microsatellites associated or correlated with the condition with optimized accuracy, sensitivity, and specificity. In some aspects, the computer-implemented method can be validated with the set of additional samples comprising samples having the condition, samples not having the condition, or a combination thereof (e.g., see Figure 3 Validation can include using at least 10, 20, 30, 50, 100, or 1000 samples from subjects with a disorder (e.g., cancer) (the samples can be non-tumor (germline) samples or tumor samples) and at least 10, 20, 30, 50, 100, or 1000 samples from subjects without the disorder (e.g., cancer, such as lung cancer).

[0133] When analyzing a sample from a subject, the computer-implemented method of optimization and verification can generate a classifier for a condition. The condition can indicate the presence or absence of a subject's health state. In some cases, the condition indicates an increase or decrease in the likelihood that the subject's health state will develop. In some instances, the condition can indicate an increase or decrease in the likelihood that the subject will benefit from treatment, or an increase or decrease in the likelihood that the subject will have an increased risk of adverse effects due to treatment. In some cases, the condition can indicate the responsiveness of the subject to the treatment of the health state. In some instances, the condition indicates a prognosis for the subject's health state.

[0134] In some cases, the disease may indicate the presence or absence of cancer. In some cases, the disease indicates that the possibility of cancer development increases or decreases. In some instances, the disease indicates that the possibility of the subject benefiting from treatment increases or decreases, or the subject has the possibility of increased adverse effect risk due to treatment increases or decreases (the classifier can be a companion diagnosis for cancer treatment). In some cases, the disease may indicate the responsiveness to cancer treatment. Treatment can be surgery, chemotherapy, radiotherapy, drug-targeted therapy (for example, afatinib, gefenib, bevacizumab, crizotinib or serotoninib) or immunotherapy (for example, with monoclonal antibodies, checkpoint inhibitors, therapeutic vaccines or adoptive T cell transfer therapy). In some instances, the disease indicates the prognosis of cancer. In some cases, cancer is lung cancer, including non-small cell lung cancer (for example, lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC) and large cell carcinoma), small cell lung cancer (SCLC) or lung carcinoid.

[0135] The classifier can include microsatellite loci from any chromosome, e.g., chromosome 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X, or Y. In some cases, the classifier does not include microsatellite loci from chromosome X and / or chromosome Y.

[0136] III. Generating weighted classifiers for disease states

[0137] The present disclosure provides a method for weighting microsatellite loci that have been identified as being associated or related to a disease. In addition, the present disclosure provides a method for weighting genetic markers other than microsatellite loci that have been identified as being associated or related to a disease. Weight or weights can refer to the relative importance or prevalence of each microsatellite locus, which statistically contributes to the correlation or association with the disease. For example, a high weight can be assigned to a microsatellite locus that only appears in samples obtained from subjects with the disease and that appears with higher frequency. In some cases, weights are assigned based on a hazard ratio, an odds ratio, or a relative risk. Examples of numerical components that are part of the weight determination include sensitivity, specificity, negative predictive value, positive predictive value, odds ratio, hazard ratio, or any combination thereof. In some cases, a cutoff value (e.g., a threshold) is imposed on the numerical components used to calculate the weights. Samples whose numerical classifiers are below the cutoff value can be excluded from the weight calculation. Weights can be calculated based on a combination of linear, nonlinear, algebraic, trigonometric, statistical learning, Bayesian, regression, or related computational methods. A weighting scheme or regression method for values (e.g., relative risk) associated with one or a group of microsatellites can be used to generate a classifier. The weighted classifier can be evaluated to determine whether the weighting improves the sensitivity or specificity of the classifier. Regression analysis (e.g., standard regression analysis) can be used to calculate the optimal weight for each seat to maximize sensitivity and specificity (e.g., the sum of sensitivity and specificity).

[0138] In some cases, the weight assigned to each microsatellite is a predetermined value, wherein the predetermined value determines the correlation or the strength of association between sample size or illness and the microsatellite site. In some instances, the weight assigned to each microsatellite comprises relative risk, hazard ratio or odds ratio. In some instances, the predetermined value of weight determines the numerical range (for example, summation) of sensitivity, specificity or its combination. In some instances, the calculation and assignment of weights include a decision model implemented by a computer via a model, such as a support vector machine, a decision tree, a random forest, a neural network, or a deep learning neural network (e.g., an artificial neural network, a recurrent neural network, a convolutional neural network, perception, feedforward, a radial basis network, a deep feedforward, a recurrent neural network, a long / short-term memory, a gated recurrent unit, an autoencoder (AE), a mutation AE, a denoising AE, a sparse AE, a Markov chain, a Hopfield network, a Boltzmann machine, a restricted BM, a deep belief network, a deep convolutional network, a deconvolutional network, a deep convolutional inverse graph network, a generative adversarial network, a liquid state machine, an extreme learning machine, a per-state network, a deep residual network, a Kohonen network, a support vector machine, and a neural Turing machine).

[0139] In some instances, the weight assigned to microsatellite loci is used as a part for the calculation of a sorter as described herein. In this type of instance, the microsatellite loci with larger weights contribute more to the value of the sorter than the microsatellite loci with less weights. In some cases, the calculation of the sorter comprises and only uses optimal weights. Optimal weights can include a weight at least or greater than a predetermined threshold.

[0140] The illness determined by the weighted classifier can indicate the presence or absence of the subject's health state. In some cases, the possibility of the illness determined by the weighted classifier indicating the development of the subject's health state increases or decreases. In some instances, the possibility of the illness determined by the weighted classifier indicating the subject benefiting from treatment increases or decreases, or the possibility of the subject having an increased risk of adverse effects due to treatment increases or decreases. In some instances, the responsiveness of the illness determined by the weighted classifier indicating the subject to the treatment of health state. In other examples, the illness determined by the weighted classifier can indicate the prognosis of the subject's health state. In some cases, health state is cancer. In some cases, cancer is lung cancer, for example, non-small cell lung cancer (for example, lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC) and large cell carcinoma), small cell lung cancer (NSLC) or lung carcinoid tumor.

[0141] Classifiers can also be determined based on, for example, the minor allele distribution of microsatellites. In some cases, a classifier can be determined by calculating a weighted combination of informative microsatellite loci and minor allele distributions. Minor allele frequency can be another weighted parameter of a classifier. Minor allele frequency can be used as an indicator of overall genome stability. Classifiers based on minor allele frequency can be statistically evaluated (e.g., by regression analysis) to determine whether adding a minor allele frequency to a classifier improves the classifier. IV. Pan-disease (e.g., cancer) risk determination

[0142] The present disclosure provides computer-implemented methods for generating a pan-disease (e.g., cancer) classifier (e.g., see Figure 2 and Figure 4 ). An informative microsatellite locus list can be generated by statistically analyzing samples of various disease (e.g., cancer) types and healthy reference sequences. DNA sequencing from two groups of samples can be performed on a multiplex platform. In some cases, the targeting of sequencing is another enrichment, such as using a bait set. The sequencing result is then quality analyzed and mapped to reveal the difference between the disease (e.g., cancer) sample and the reference sample. This difference can be analyzed by a computer-implemented method to generate a general disease (e.g., cancer) classifier. The general disease (e.g., cancer) classifier can be further optimized and verified with other samples of various types of disease (e.g., cancer).

[0143] A pan-disorder (e.g., pan-cancer) classifier for one or more conditions can indicate the presence or absence of at least one of a plurality of health states in a subject, an increased or decreased likelihood that at least one of a plurality of health states will develop in a subject, an increased or decreased likelihood that a subject will benefit from treatment for at least one of a plurality of health states, an increased or decreased likelihood that a subject will have an increased risk of adverse effects due to treatment for at least one of a plurality of health states, the responsiveness of a subject to treatment for at least one of a plurality of health states, or a combination thereof. The plurality of health states can be any combination of the health states disclosed herein.

[0144] In some cases, a pan-cancer condition can indicate the presence or absence of multiple types of cancer in a subject. In some instances, a pan-cancer condition can indicate an increased or decreased likelihood of developing multiple types of cancer in a subject. In some instances, multiple types of cancer are cancers that often develop together in the same subject. In other instances, multiple types of cancer are cancers that arise independently. In some instances, a pan-cancer condition can indicate that a subject may or may not benefit from treatment, or that a subject may or may not be at increased risk of adverse effects due to treatment (a pan-cancer classifier can be a companion diagnostic for a therapeutic product). In some instances, a pan-cancer condition can indicate a subject's responsiveness to cancer treatment. In other instances, a pan-cancer condition can indicate a subject's prognosis for cancer. The subjects described herein may have or may not have cancer symptoms. In some cases, additional tests (e.g., physical examination, analysis of circulating or cell-free cancer biomarkers, imaging (e.g., computed tomography (CT), bone scan, magnetic resonance imaging (MRI), positron emission tomography (PET), ultrasound, and X-ray), biopsy, genetic screening, gene or protein expression levels, etc.) can be used based on the subject's pan-cancer classifier.

[0145] The computer-implemented method for generating a pan-disease (e.g., pan-cancer) classifier may include further analysis of execution processing, combination, statistical evaluation or results, or any combination thereof. In some aspects, the computer-implemented method for generating a pan-disease (e.g., cancer) classifier includes first identifying microsatellite loci (different from the microsatellite loci in the sample obtained or derived from a subject suffering from a variety of diseases (e.g., cancer)) to generate a colony of a subset of microsatellite loci associated with or associated with a variety of diseases (e.g., cancer) by obtaining or deriving from a subject suffering from a variety of diseases (e.g., cancer). The sequence of microsatellite can first be obtained by any sequencing method.

[0146] One or more statistical tests (such as t-test, Z-test, analysis of variance, regression analysis, Mann-Whitney-Wilcoxon test, chi-square test, association, Fisher's exact test, Bonferroni correction, and Benjamini-Hochberg test) can be used to identify microsatellite loci that are associated or correlated with various types of disorders (e.g., cancer).

[0147] Statistical tests can produce receiver operating characteristic (ROC) curves, wherein the area under the ROC curve is referred to as area under the curve (AUC). AUC can determine the accuracy of identifying microsatellite loci that are relevant or associated with various types of illnesses (e.g., cancer). Larger AUC can indicate the higher accuracy of correlation or association. The ROC curve can determine the sensitivity (e.g., true positive) and specificity (e.g., true negative) ratio of the correlation or association of microsatellite loci and various types of illnesses (e.g., cancer). The statistically significant correlation or association of microsatellite loci and various types of illnesses (e.g., cancer) can have a statistical accuracy of at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%. In some cases, the statistically significant association or correlation of a microsatellite locus with multiple types of disorders (e.g., cancer) has a statistical specificity of at least 0.70, 0.80, 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, or 0.99, and a statistical sensitivity of at least 0.70, 0.80, 0.85, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.96, 0.99, or 0.99.

[0148] In some instances, identification of microsatellite loci associated or related to various types of illness (e.g., cancer) includes identifying a first group of microsatellite loci from a database comprising nucleic acid sequences of various types of illness (e.g., cancer) and identifying a second group of microsatellite loci from a reference database (e.g., hg19). In some cases, some microsatellites are identified as being associated or related to various types of illness (e.g., cancer). In some cases, some microsatellites are identified as being associated or related to a type of illness (e.g., cancer).

[0149] Various types of cancer can include solid or hematological malignant cancers. In some cases, various types of cancer can be metastatic, recurrent or refractory. Various types of cancer associated or associated with the identified microsatellite loci can include any number (e.g., about 4 to about 10, about 10 to about 15, about 15 to about 20, or about 4, about 10, about 15, about 20, about 25, about 30, or about 50) of cancers disclosed herein.

[0150] The pan-cancer assay can assay for or can test for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 of the following cancers: breast cancer, ovarian cancer, prostate cancer, lung cancer, glioblastoma multiforme, uterine corpus endometrial cancer, colon adenocarcinoma, bladder cancer, urothelial carcinoma, head and neck squamous cell carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, gastric adenocarcinoma, thyroid cancer, brain low-grade glioma, kidney renal papillary cell carcinoma, and hepatocellular carcinoma.

[0151] In some cases, the various types of cancers associated or correlated with differences in the set of microsatellite loci include lung cancer. Lung cancers that may be associated or correlated with different sets of microsatellite loci include non-small cell lung cancer (e.g., lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), and large cell carcinoma), small cell lung cancer (SCLC), and lung carcinoid.

[0152] The population of microsatellite loci subsets comprising those associated or correlated with various types of disorders (e.g., cancer) can include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 microsatellite loci of each subset. In some aspects, the population of subsets is iteratively ranked based on the relevance or association of the subsets with various types of disorders (e.g., cancer).

[0153] In some cases, the colony of about 100 subsets is used in the optimization. In some cases, the colony of at least 100,200,300,400,500,1000,2000,3000 or 5000 subsets is used in the optimization. In some instances, optimization comprises at least one cycle that the subset of about 100 identifications is compared with other samples. In some instances, optimization comprises a plurality of cycles that the subset of about 100 identifications is compared with other samples.

[0154] Can carry out iterative sorting when each cycle is completed.In some cases, iterative sorting comprises that subset is performed statistical analysis, to perform receiver operating characteristic (ROC) analysis, thereby when determining the presence or absence of multiple types of illness (for example, cancer) in other sample, perform accuracy, sensitivity and specificity.Can identify and discard one or more in the subset that performs the worst or ranks lowest in indicating the presence or absence of multiple types of illness (for example, cancer).In order to maintain a constant number of subsets before each optimization cycle begins, new subsets can be added to the subset population.In some cases, new subsets are generated by randomly splitting and reorganizing 2 subsets randomly selected in the previous round optimization cycle.In some instances, new subsets are randomly selected from the previous optimization cycle.In some cases, the number of the subsets discarded at the end of each optimization cycle is identical with the number of the subsets added to the subset before each optimization cycle.

[0155] A computer-implemented method for generating a pan-disease (e.g., pan-cancer) classifier can include determining a statistically unweighted subset of microsatellite loci. In some aspects, a computer-implemented method for generating a pan-disease (e.g., pan-cancer) classifier includes determining a statistically weighted subset of microsatellite loci. The pan-disease (e.g., pan-cancer) classifier can be unweighted or weighted.

[0156] After the optimization cycle is completed, the computer-implemented method for generating a pan-disease (e.g., pan-cancer) classifier includes microsatellite loci associated or associated with the disease with the accuracy, sensitivity and specificity of optimization. In some aspects, the computer-implemented method can be verified with the other samples of the group, and the other samples of the group include samples obtained or derived from a subject suffering from multiple types of diseases (e.g., cancer), samples obtained or derived from a subject never suffering from multiple types of diseases (e.g., cancer), or a combination thereof. When analyzing the sample from the subject, the computer-implemented method for optimization and verification can generate a pan-disease (e.g., pan-cancer classifier). Pan-disease (e.g., pan-cancer) can indicate the presence or absence of a health state (e.g., cancer) of a subject. In some cases, the possibility of a health state (e.g., cancer) development of a pan-disease (e.g., pan-cancer) indication subject increases or decreases. In some cases, pan-disease (e.g., pan-cancer) can indicate that the possibility of a subject benefiting from treatment increases or decreases, or the possibility of a subject having an increased risk of adverse effects due to treatment increases or decreases (pan-disease, e.g., pan-cancer, classifier can be a companion diagnosis for therapeutic products). In some instances, a pan-disorder (e.g., pan-cancer) indicates responsiveness to treatment of a subject's health condition (e.g., cancer). In other instances, a pan-disorder (e.g., pan-cancer) indicates a prognosis of a subject's health condition (e.g., cancer).

[0157] A classifier (e.g., a set of microsatellites) can be developed for each disorder (e.g., cancer) in a pan-disease (e.g., pan-cancer) assay. In some cases, a single microsatellite locus can be a pan-disease (e.g., pan-cancer) microsatellite locus.

[0158] V. Sample of Evaluation Subjects

[0159] The classifier generated as described herein can be used to analyze subject (e.g., patient) samples. For example, samples from a subject can be analyzed in a laboratory certified by the Clinical Laboratory Improvement Amendments (CLIA). In some cases, a test kit is prepared and the subject's sample is measured outside a CLIA laboratory. Figure 5An example of a workflow (500) for a subject (e.g., patient) sample analysis pipeline, such as in a CLIA-certified laboratory, is illustrated; the workflow can be used to process samples for a multiplexed pan-cancer assay. Samples are obtained from multiple subjects (501), such as blood, urine, cerebrospinal fluid, semen, saliva, sputum, stool, lymph, tissue (e.g., thyroid, skin, heart, lung, kidney, breast, pancreas, liver, muscle, smooth muscle, bladder, gallbladder, colon, intestine, brain, esophagus, or prostate), or any combination thereof. Nucleic acid molecules, such as genomic DNA, are extracted from the samples. Targets, such as microsatellite targets, are enriched by multiplexing (e.g., using baits, such as hybridization probes); the enriched targets can be barcoded and amplified (503). Next generation sequencing assays are performed on the target-enriched samples, such as in batches of about 4, 8, 12, 24, 96, 128, 384, or 1536 times (505). Sequencing data can be demultiplexed (e.g., using unique sequence tags (e.g., barcodes) added to each individual sample), quality control filters can be applied to the raw sequence reads (e.g., Phred quality greater than Q30), and genotypes determined (e.g., aligning reads for each locus to a reference sequence using flanking sequences and then calculating two major alleles (genotypes)) and minor allele distributions (e.g., determining the number of minor alleles or the fraction of minor alleles relative to the major genotype for each microsatellite locus for each sample (507)). Calculation (509) is performed for each microsatellite locus. A risk classifier (e.g., based on at least 5, 10, 25, 50, or 100 microsatellite loci) is generated for each sample for a type of cancer (e.g., genotypes can be determined to be modal or non-modal relative to the most prominent genotype in a healthy population (e.g., GRCh38) and summed across all loci, and samples can be classified as at risk or not at risk for a condition depending on their position relative to a cutoff point for the fraction of loci having a cancer or normal genotype). Risk can be quantitative or can be indicated by a categorical assessment. A clinical laboratory report including the risk classifier is generated (511) and provided to a healthcare provider, subject, or insurance provider.

[0160] Figure 17 An example of a clinical laboratory report is shown. A clinical laboratory report may include patient information, sample information, a test summary, test results, annotations, and result details. Result details may include the number of microsatellite loci genotyped, one or more condition risk classifiers, one or more thresholds, and the relative risk of having or acquiring a condition (e.g., lung cancer) (e.g., low risk, high risk, "at risk," "no risk").

[0161] The report can include the number of sites in samples from subjects with non-modal (primarily cancer) genotypes. The sensitivity and specificity of detecting the presence of healthy states determined to be high risk can be greater than 90%, and their absence in control sample lines determined to be "low risk" for lung cancer. The accuracy of the assay can be greater than 99% by measuring highly conserved sites in reference controls.

[0162] In some examples, the condition can be confirmed or further investigated by additional tests, such as physical examination, analysis of circulating or cell-free cancer biomarkers, imaging (e.g., computed tomography, bone scan, magnetic resonance imaging, positron emission tomography, ultrasound, and X-ray), biopsy, genetic screening, gene expression, or protein expression, etc. VI. Minor Alleles in Microsatellites

[0163] The present disclosure provides a computer-implemented method for determining the genome age and genome aging rate of a subject. The genome age can be given with a number calibrated to years. For example, if the genome age is approximately equal to the digital age of the subject, then for the genome age, overall genome stability can be normal. In some instances, the genome age may be smaller, the same, or larger than the actual age of the subject. A genome age older than the actual age of the subject, or a high genome aging rate, may suggest that the genome is unstable and prone to developing health conditions (e.g., diseases) associated with aging, such as cancer, cardiovascular disease, neurological diseases, etc. From samples obtained from different tissues (e.g., skin or blood) of the same subject, the genome age and genome aging rate may be different. In some cases, the genome age and genome aging rate can indicate a person's lifestyle (e.g., nutrition, physical or mental stress) or medical condition. A change in lifestyle (e.g., quitting smoking, changing diet and exercising) can be recommended to the subject based on the genome age of the subject.

[0164] The computer-implemented method for determining genome age and genome aging rate can comprise determining the minor allele signature from the first sample of experimenter, and the minor allele signature of the first sample is compared with the minor allele signature of reference, to produce the first difference of minor allele signature. Said reference can comprise the distribution of minor allele content across large population, to determine the average genome age as a function of digital age, race, sex etc. By computer-implemented method, it can be determined that the first difference of the minor allele signature between the first sample and the reference is the genome age of experimenter. In some aspects, the time point after the first sample is compared with the reference will be compared with the reference to produce the second difference of minor allele signature. The variation between the first difference and the second difference can be determined as the genome aging rate of experimenter by computer-implemented method. In some cases, other genome aging rate can be determined by obtaining and comparing later minor allele signature and earlier minor allele signature.

[0165] In some cases, the minor allele signature comprises the combination of SNP and insertion / deletion variation, microsatellite variation, synonymous SNP, non-synonymous SNP, stop gain SNP, stop losing SNP, splice variation (for example, the 2-bp in the splice joint), frameshift insertion / deletion and the non-frameshift insertion / deletion at at least one seat. In some cases, the minor allele signature comprises a plurality of time points across same experimenter to determine.

[0166] The minor allele signature determined from a sample of an experimenter may need to read at least 1 sequence from any sequencing method. In some cases, the minor allele signature can be identified at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, or 100 sequence reads from any next generation sequencing method. The minor allele signature determined from a sample of an experimenter may need to read at least 1, at least 2, at least 3, or more than 3 sequences from any sequencing method.

[0167] In some instances, the minor allele signature determined from the sequence of the sample of the experimenter is compared with the reference sequence.Described comparison can produce the difference of minor allele signature from reference sequence, and described reference sequence comprises the SNP of different numbers and insertion and deletion variation, microsatellite variation, synonymous SNP, non-synonymous SNP, stop gain SNP, stop loss SNP, splice variation (for example, 2-bp in splice junction), frameshift insertion and deletion and the combination of the non-frameshift insertion and deletion of at least one seat.The difference of minor allele signature between sample and reference can be determined by computer-implemented method, to produce genome age.

[0168] In some cases, a first sequence from a first sample of a subject is compared to a reference sequence to generate a first minor allele signature and a first genomic age. In some instances, a second sequence from a second sample of the same subject is compared to the same reference sequence to generate a second minor allele signature and a second genomic age. The comparison between the first minor allele signature and the second minor allele signature can determine a genomic aging rate. In certain instances, multiple minor allele signatures can be obtained from samples of the same subject at a later time point for comparison to generate multiple genomic aging rates for different ages of the subject.

[0169] The disclosure provides the computer-implemented method for determining the genome age of an experimenter by determining the microsatellite minor allele signature from the first sample of the experimenter.The microsatellite minor allele signature can be a minor allele, and it comprises the microsatellite with different SNP percentages, amplification percentage, contraction percentage, amplification and contraction and the ratio of SNP, heterozygous site percentage or homozygous site percentage compared with reference sequences.In some cases, the microsatellite minor allele signature comprises a minor allele, and the minor allele comprises the microsatellite with the different combinations of SNP and insertion / deletion variation, microsatellite variation, synonymous SNP, non-synonymous SNP, stop gain SNP, stop losing SNP, splice variation (for example, 2-bp in splice joint), frameshift insertion / deletion or non-frameshift insertion / deletion, when compared with reference sequences, at least one seat.In some cases, the microsatellite minor allele signature is determined across multiple time points of same experimenter.

[0170] VI. Computer Systems, Processors, and Memory

[0171] The present disclosure provides a computer system configured to implement the method described in the present disclosure. In some instances, a system is disclosed herein, comprising: a computer processing device, which is optionally connected to a computer network; and a software module, which is executed by the computer processing device. In some instances, the system includes a central processing unit (CPU), a memory (e.g., random access memory, flash memory), an electronic storage unit, a computer program, a communication interface for communicating with one or more other systems, and any combination thereof. In some instances, the system is coupled to a computer network, such as the Internet, an intranet, and / or an extranet that communicates with the Internet, a telecommunications or data network. In some aspects, this system includes a storage unit for storing data and information about any aspect of the method described in the present disclosure. Each aspect of the system is a product or article or manufactured product.

[0172] One feature of a computer program is a sequence of instructions written to perform specific tasks that can be executed in a CPU of a digital processing device. In some aspects, the computer-readable instructions are implemented as program modules that perform specific tasks or implement specific abstract data types, such as functions, features, application programming interfaces (APIs), data structures, etc. In various embodiments, computer programs can be written in various versions of various languages.

[0173] The functionality of the computer-readable instructions can be combined or distributed in various environments as needed. In some instances, a computer program comprises one or more instruction sequences. The computer program can be provided from one location. The computer program can be provided from multiple locations. In some aspects, the computer program comprises one or more software modules. In some aspects, the computer program comprises, in part or in whole, one or more web applications, one or more mobile applications, one or more stand-alone applications, one or more web browser plug-ins, extensions, add-ons, or combinations thereof.

[0174] Computer system

[0175] The present disclosure provides computer systems programmed to implement the methods of the present disclosure. Figure 18 A computer system (1801) is shown that can be programmed or otherwise configured to perform the methods described herein. The computer system (1801) can adjust various aspects of the present disclosure, including inputting nucleic acid position information, transferring the inferred information to a data set, and generating a training algorithm with the data set. The computer system (1801) can be a user electronic device or a remote computer system. The electronic device can be a mobile electronic device.

[0176] The computer system (1801) includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") (1805), which can be a single-core or multi-core processor that processes data sequentially or in parallel. The computer system (1801) also includes a storage unit or device (1810) (e.g., random access memory, read-only memory, flash memory), a storage unit (1815) (e.g., a hard disk), a communication interface (1820) (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices (1825), either external or internal or both, such as a printer, a monitor, a USB drive, and / or a CD-ROM drive. The memory (1810), storage unit (1815), interface (1820), and peripheral devices (1825) communicate with the CPU (1805) via a communication bus (solid line) such as a motherboard. The storage unit (1815) can be a data storage unit (or data repository) for storing data. The computer system (1801) can be operatively coupled to a computer network ("network") (1830) via a communication interface (1820). The network (1830) can be the Internet, an intranet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The network (1830) is, in some cases, a telecommunications network and / or a data network. The network (1830) can include one or more computer servers that can enable a peer-to-peer network supporting distributed computing. In some cases, the network (1830) can implement a client-server architecture with the computer system (1801), which can enable a device coupled to the computer system (1801) to behave as a client or a server.

[0177] The CPU (1805) can execute a series of machine-readable instructions, which can be incorporated into a program or software. The instructions can be stored in a memory (1810). The instructions can be directed to the CPU (1805), which can then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU (1805) can include fetching, decoding, executing, and writing back.

[0178] The CPU (1805) may be part of a circuit, such as an integrated circuit. One or more other components of the system (1801) may be included in the circuit. In some embodiments, the circuit is an application specific integrated circuit (ASIC).

[0179] The storage unit (1815) can store files, such as drivers, libraries, and saved programs. The storage unit (1815) can store user data, such as user preferences and user programs. In some cases, the computer system (1801) can include one or more additional data storage units external to the computer system (1801), such as located on a remote server that communicates with the computer system (1801) via an intranet or the Internet.

[0180] The computer system (1801) can communicate with one or more remote computer systems via a network (1830). For example, the computer system (1801) can communicate with a remote computer system or user. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets or tablet PCs (e.g., iPad, GalaxyTab), phones, smartphones (e.g. iPhone, Android-supported devices, ) or a personal digital assistant. A user can access the computer system (1801) via a network (1830).

[0181] The methods described herein may be implemented by machine (e.g., computer processor) executable code stored in an electronic storage location of a computer system (1801), such as in a memory (1810) or a data storage unit (1815). The machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by the processor (1805). In some cases, the code may be retrieved from the storage unit (1815) and stored in the memory (1810) for ready access by the processor (1805). In some cases, the storage unit (1815) may be eliminated, and the machine executable instructions may be stored in the memory (1810).

[0182] The code may be precompiled and configured for use with a machine having a processor suitable for executing the code, or it may be compiled at runtime. The code may be supplied in a programming language that may be selected to enable the code to be executed in a precompiled or compiled manner.

[0183] Aspects of the systems and methods provided herein, such as the computer system (1801), can be incorporated into programming. Various aspects of the technology can be considered "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried or contained in a type of machine-readable medium. The machine executable code can be stored in a storage unit (such as a hard disk) or a memory (e.g., read-only memory, random access memory, flash memory). Storage-type media can include any or all tangible memories of a computer, processor, etc., or their associated modules, that can provide non-transitory storage for software programming at any time, including various semiconductor memories, tape drives, disk drives, etc. All or part of the software can sometimes be communicated over the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another, such as from a management server or host computer to a computer platform of an application server. Thus, another type of medium that can carry software elements includes optical waves, radio waves, and electromagnetic waves, such as those used across physical interfaces between local devices via wired and optical landline networks and various air links. The physical elements that carry such waves, such as wired or wireless links, optical links, etc., may be considered the medium that carries the software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0184] A.Electronic devices

[0185] In some aspects, the platforms, media, methods, and applications described herein include electronic devices, processors, or their use (also referred to as digital processing devices). In other aspects, the electronic device includes one or more hardware central processing units (CPUs) that perform device functions. In yet other aspects, the electronic device also includes an operating system configured to execute executable instructions. In some aspects, the electronic device is optionally connected to a computer network. In other aspects, the electronic device is optionally connected to the Internet so that it accesses the World Wide Web. In yet other aspects, the electronic device is optionally connected to a cloud computing infrastructure. In some aspects, the electronic device is optionally connected to an intranet. In some aspects, the electronic device is optionally connected to a data storage device. According to the description herein, as non-limiting examples, suitable electronic devices include server computers, desktop computers, laptop computers, notebook computers, small notebook computers, netbook computers, network tablet computers, set-top box computers, handheld computers, Internet devices, mobile smart phones, tablet computers, personal digital assistants, video game consoles, and vehicles. In various embodiments, many smart phones are suitable for the systems described herein. In various embodiments, selected televisions, video players, and digital music players with optional computer network connections are suitable for the systems described herein. Suitable tablet computers include those available in booklet, tablet, and convertible configurations.

[0186] In some aspects, the electronic device includes an operating system configured to execute executable instructions. For example, an operating system is software, including programs and data, that manages the hardware of the device and provides services for the execution of application programs. In various embodiments, suitable server operating systems include, by way of non-limiting example, FreeBSD, OpenBSD, Linux, Ubuntu Linux, Mac OS X Windows as well as In various embodiments, suitable personal computer operating systems include, by way of non-limiting example, and UNIX-like operating systems such as In some aspects, the operating system is provided by cloud computing. In various embodiments, suitable mobile smartphone operating systems include, by way of non-limiting example, OS, OS, operating system, as well as

[0187] In some aspects, the device includes a storage and / or memory device. A storage device and / or memory device is one or more physical devices for temporarily or permanently storing data or programs. In some aspects, the device is a volatile memory and requires power to maintain the stored information. In some aspects, the device is a non-volatile memory and retains the stored information when the electronic device is not powered. In other aspects, the non-volatile memory includes flash memory. In some aspects, the non-volatile memory includes dynamic random access memory (DRAM). In some aspects, the non-volatile memory includes ferroelectric random access memory (FRAM). In some aspects, the non-volatile memory includes phase change random access memory (PRAM). In some aspects, the non-volatile memory includes magnetoresistive random access memory (MRAM). In some aspects, the device is a memory device, which, as non-limiting examples, includes a CD-ROM, a DVD, a flash memory device, a disk drive, a tape drive, an optical drive, and a cloud computing-based storage device. In other aspects, the storage and / or memory device is a combination of devices such as those disclosed herein.

[0188] In some aspects, the electronic device includes a display that sends visual information to the subject. In some aspects, the display is a cathode ray tube (CRT). In some aspects, the display is a liquid crystal display (LCD). In other aspects, the display is a thin film transistor liquid crystal display (TFT-LCD). In some aspects, the display is an organic light emitting diode (OLED) display. In various further aspects, the OLED display is a passive matrix OLED (PMOLED) or active matrix OLED (AMOLED) display. In some aspects, the display is a plasma display. In some aspects, the display is electronic paper or electronic ink. In some aspects, the display is a video projector. In yet other aspects, the display is a combination of devices such as those disclosed herein.

[0189] In some aspects, the electronic device includes an input device for receiving information from the subject. In some aspects, the input device is a keyboard. In some aspects, the input device is a pointing device, including, as non-limiting examples, a mouse, trackball, trackpad, joystick, game controller, or stylus. In some aspects, the input device is a touch screen or multi-touch screen. In some aspects, the input device is a microphone to capture voice or other sound input. In some aspects, the input device is a camera or other sensor to capture motion or visual input. In other aspects, the input device is a Kinect, Leap Motion, etc. In still other aspects, the input device is a combination of devices such as those disclosed herein.

[0190] B. Non-transitory computer-readable storage media

[0191] In some aspects, the platforms, media, methods, and applications described herein include one or more non-transitory computer-readable storage media encoded with a program comprising instructions executable by an operating system of an optionally networked digital processing device. In other aspects, the computer-readable storage medium is a tangible component of the electronic device. In still other aspects, the computer-readable storage medium is optionally removable from the electronic device. In some aspects, by way of non-limiting example, computer-readable storage media include CD-ROMs, DVDs, flash memory devices, solid-state memories, magnetic disk drives, tape drives, optical disk drives, cloud computing systems and services, and the like. In some cases, the program and instructions are encoded on the medium permanently, substantially permanently, semi-permanently, or non-transitorily.

[0192] C.Computer Program

[0193] In some aspects, the platforms, media, methods, and applications described herein include at least one computer program or use thereof. A computer program comprises a series of instructions that are written to perform specific tasks and that can be executed in a CPU of an electronic device. Computer-readable instructions can be implemented as program modules that perform specific tasks or implement specific abstract data types, such as functions, objects, application programming interfaces (APIs), data structures, and the like. In various embodiments, computer programs can be written in various versions of various languages.

[0194] The functionality of the computer-readable instructions can be combined or distributed as needed in various environments. In some aspects, the computer program comprises a single sequence of instructions. In some aspects, the computer program comprises multiple sequences of instructions. In some aspects, the computer program is provided from a single location. In some aspects, the computer program is provided from multiple locations. In various aspects, the computer program comprises one or more software modules. In various aspects, the computer program comprises, in part or in whole, one or more web applications, one or more mobile applications, one or more stand-alone applications, one or more web browser plug-ins, extensions, add-ons, or combinations thereof.

[0195] D. Web Application

[0196] In some aspects, the computer program comprises a web application. In various embodiments, in various aspects, the web application utilizes one or more software frameworks and one or more database systems. In some aspects, the web application is a web application that is hosted on a server such as a server. .NET or Ruby on Rails (RoR) software framework. In some aspects, the web application utilizes one or more database systems, including, by way of non-limiting example, relational, non-relational, object-oriented, associative, and XML database systems. In other aspects, suitable relational database systems include, by way of non-limiting example, SQL Server, mySQL TM ,and In various embodiments, in various aspects, a web application is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation-defining languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some aspects, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or Extensible Markup Language (XML). In some aspects, a web application is written to some extent in a presentation-defining language such as Cascading Style Sheets (CSS). In some aspects, a web application is written to some extent in a client-side scripting language such as Asynchronous JavaScript and XML (AJAX), Actionscript, Javascript, or In some aspects, web applications are written to some extent in server-side coding languages such as Active Server Pages (ASP), Perl, Java TM , JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python TM 、Ruby、Tcl、Smalltalk、 or Groovy. In some respects, web applications are written to some extent in database query languages such as Structured Query Language. In some respects, web applications integrate with enterprise server products such as In some aspects, the web application includes a media player element. In various other aspects, the media player element utilizes one or more of a number of suitable multimedia technologies, including, by way of non-limiting example, HTML 5, Java TM and

[0197] E. Mobile Application

[0198] In some aspects, the computer program includes a mobile application provided to the mobile electronic device. In some aspects, the mobile application is provided to the mobile electronic device during manufacture. In some aspects, the mobile application is provided to the mobile electronic device via a computer network as described herein.

[0199] In various embodiments, the mobile application is created using various techniques using hardware, languages, and development environments. In various embodiments, the mobile application is written in several languages. Suitable programming languages include, by way of non-limiting example, C, C++, C#, Objective-C, Java, and C++. TM ,Javascript,Pascal,ObjectPascal,Python TM , Ruby, VB.NET, WML and XHTML / HTML with or without CSS or a combination thereof.

[0200] Suitable mobile application development environments are available from several sources. As non-limiting examples, commercially available development environments include AirplaySDK, alcheMo, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available free of charge and include, as non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. In addition, mobile device manufacturers distribute software development kits including, as non-limiting examples, the iPhone and iPad (iOS) SDK, SDK, BREW SDK, OS SDK, SymbianSDK, webOS SDK, and Mobile SDK.

[0201] In various embodiments, several commercial forums may be used to distribute mobile applications, including, by way of non-limiting example, App Store, AppWorld, an application store for handheld devices, App Catalog for webOS, Mobile market, Ovi Store on the device, applications, and DSi Store.

[0202] F. Standalone Application

[0203] In some aspects, a computer program includes a stand-alone application, which is a program that runs as an independent computer process, rather than as an add-on to an existing process, e.g., not a plug-in. In various embodiments, a stand-alone application is often compiled. A compiler is a computer program that converts source code written in a programming language into binary object code (such as assembly language or machine code). By way of non-limiting example, suitable compiled programming languages include C, C++, Objective-C, COBOL, Delphi, Eiffel, Java, and many more. TM , Lisp, Python TM , Visual Basic, and VB.NET or a combination thereof. Compilation is typically performed at least in part to create an executable program. In some aspects, a computer program includes one or more executable compiled applications.

[0204] G. Software Modules

[0205] In some aspects, the platforms, media, methods, and applications described herein include software, server, and / or database modules, or their use. In various embodiments, software modules are created using various techniques for machines, software, and languages. The software modules disclosed herein can be implemented in a variety of ways. In various aspects, a software module includes a file, a code segment, a programming object, a programming structure, or a combination thereof. In other aspects, a software module includes multiple files, multiple code segments, multiple programming objects, multiple programming structures, or a combination thereof. In various aspects, as non-limiting examples, one or more software modules include a web application, a mobile application, and a standalone application. In some aspects, a software module is in a computer program or application. In some aspects, a software module is in more than one computer program or application. In some aspects, a software module is hosted on a single machine. In some aspects, a software module is hosted on more than one machine. In other aspects, a software module is hosted on a cloud computing platform. In some aspects, a software module is hosted on one or more machines in a single location. In some aspects, a software module is hosted on one or more machines in more than one location.

[0206] H. Database

[0207] In some aspects, the platforms, systems, media, and methods disclosed herein include one or more databases or their use. In various embodiments, many databases are suitable for storing and retrieving barcodes, routes, packages, subjects, or network information. In various aspects, as non-limiting examples, suitable databases include relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. In some aspects, the database is based on the Internet. In other aspects, the database is based on a network. In yet other aspects, the database is based on cloud computing. In some aspects, the database is based on one or more local computer storage devices.

[0208] I. Data Transfer

[0209] The subject matter described herein, including the methods and systems provided herein, can be configured to be performed in one or more facilities in one or more locations. The facility location is not limited to a country and includes any country or region. In some instances, one or more steps are performed in a country different from another step of the method. In some instances, one or more steps for obtaining a sample are performed in a country different from one or more steps for detecting the presence or absence of a disease from a sample. In some aspects, one or more method steps involving a computer system are performed in a country different from another step of the method provided herein. In some aspects, data processing and analysis are performed in a country or location different from one or more steps of the method described herein. In some aspects, one or more items, products, or data are transferred from one or more facilities to one or more different facilities for analysis or further analysis. Articles include, but are not limited to, one or more components obtained from a subject, such as processed cell material. Processed cell material includes, but is not limited to, cDNA reverse transcribed from RNA, amplified RNA, amplified cDNA, sequenced DNA, isolated and / or purified RNA, isolated and / or purified DNA, and isolated and / or purified polypeptides. Data includes, but is not limited to, information about subject stratification, as well as any data generated by the methods disclosed herein. In some aspects of the methods and systems described herein, an analysis is performed and a subsequent data transmission step communicates or transmits the results of the analysis.

[0210] J.Web browser plug-in

[0211] In some aspects, a computer program includes a web browser plug-in. In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Manufacturers of software applications support plug-ins to enable third-party developers to create applications that extend the capabilities of the application, enable easy addition of new features, and reduce the size of the application. When supported, plug-ins enable customization of the functionality of the software application. For example, plug-ins are commonly used in web browsers to play videos, generate interactivity, scan for viruses, and display specific file types. In various embodiments, several web browser plug-ins may be used, including Player, as well as In some aspects, the toolbar includes one or more web browser extensions, add-ons, or add-ons. In some aspects, the toolbar includes one or more browser bars, toolbars, or desktop bars.

[0212] In various embodiments, several plug-in frameworks are available that enable plugins to be written in various programming languages including, but not limited to, C++, Delphi, Java, TM , PHP, Python TM and VB.NET or their combination) to develop plug-ins.

[0213] A web browser (also known as an Internet browser) is a software application designed for use with a network-connected electronic device for retrieving, presenting, and traversing information resources on the World Wide Web. Suitable web browsers include, by way of non-limiting example, Internet Chrome, Opera and KDE Konqueror. In some aspects, the web browser is a mobile web browser. Mobile web browsers (also known as microbrowsers, minibrowsers, and wireless browsers) are designed for use with mobile electronic devices, including, by way of non-limiting example, handheld computers, tablet computers, netbook computers, notebook computers, smartphones, music players, personal digital assistants (PDAs), and handheld video game systems. Suitable mobile web browsers include, by way of non-limiting example, Browser, Browser, Blazer, Browser, formobile, Internet Mobile, BasicWeb, Browser, Mobile, and PSP TM Browser.

[0214] K. Business methods utilizing computers

[0215] Methods described herein can utilize one or more computers.Computers can be used to manage client and sample information, such as sample or client tracking, database management, analysis of molecular spectrum data, analysis of cytology data, storage data, billing, marketing, reporting results, storage results or its combination.Computers can include monitors or other graphical interfaces for displaying data, results, billing information, marketing information (e.g., demographic data), client information or sample information.Computers can also include devices for data or information input.Computers can include processing units and fixed or removable media or their combination.Computers can be accessed by users physically approaching the computer (e.g., via keyboard and / or mouse), or by users who do not have to access physical computers through communication media (e.g., modem, internet connection, telephone connection or wired or wireless communication signal carrier).In some cases, computers can be connected to servers or other communication devices for relaying information from users to computers or from computers to users.In some cases, users can store data or information obtained from computers through communication media on media (e.g., removable media).It is foreseeable that the data relevant to these methods can be transmitted through such networks or connections for one party to receive and / or view. The recipient can be, but is not limited to, an individual, a healthcare provider, or a healthcare administrator. In one example, a computer-readable medium includes a medium suitable for transmitting biological sample analysis results. The medium can include the results of the subject, wherein such results are derived using the methods described herein.

[0216] Entities that obtain sample information may enter it into a database for one or more of the following purposes: inventory tracking, assay result tracking, order tracking, customer management, customer service, billing, and sales. Sample information may include, but is not limited to, the customer's name, unique customer identifier, the customer's associated healthcare professional, the assay indicated, the assay result, the adequacy status, the adequacy of the test indicated, personal medical history, preliminary diagnosis, suspected diagnosis, sample history, insurance provider, medical provider, third-party testing center, or any other information suitable for storage in a database. Sample history may include, but is not limited to, the age of the sample, the type of sample, acquisition method, storage method, or shipping method.

[0217] The database may be accessed by clients, medical professionals, insurance providers, or other third parties. Database access may take the form of electronic communication, such as a computer or telephone. The database may be accessed by an intermediary, such as a customer service representative, a business representative, a consultant, an independent testing center, or a medical professional. The availability or extent of database access or sample information (such as assay results) may be changed upon payment of the fees for products and services provided or to be provided. The extent of database access or sample information may be limited to comply with generally accepted or legal requirements for patient or client confidentiality.

[0218] Example

[0219] The examples provided below are for illustrative purposes only and are not intended to limit the scope of the claims provided herein.

[0220] Example 1 Germline microsatellite genotypes differentiate childhood medulloblastoma ( MB )

[0221] introduce

[0222] Medulloblastoma (MB) is a common childhood malignant brain tumor. MB likely results primarily from inherited or spontaneous mutations, as children with MB have not yet experienced a lifetime of environmental exposures and stressors. Extensive genomic characterization has divided MB tumors into at least four shared molecular subgroups: WNT, SHH, Group 3, and Group 4, each with distinct transcriptional profiles, copy number alterations, somatic mutations, and clinical outcomes. Pediatric brain cancers in general, and MB in particular, harbor 5-10-fold fewer mutations than typically observed in adult solid tumors. Particularly less common are mutations in the most important tumor-initiating genes, such as p53, PTEN, RB, and EGFR. Furthermore, the incidence of known inherited tumor-susceptibility mutations may be relatively low. The few known genetic variants, such as mutations in PTCH, SMO, and CTNNB1, as well as amplifications of MYC and MYCN, may not be sufficient to efficiently cause MB in animal models and may require an enhanced background, typically p53 inactivation, which is found in less than 5% of human tumors. Many genome-wide association studies (GWAS) in MB may focus on single nucleotide variants, neglecting noncoding regions and repetitive DNA. However, germline microsatellite (MS) insertions and deletions (indels) have been shown to be associated with many neurological diseases such as Huntington's disease and Friedreich's ataxia; the former is caused by microsatellite variants in coding sequences, and the latter by noncoding intronic sequences. Furthermore, microsatellite variants may contribute to the genetic background of several cancers. Furthermore, many cancer-associated genes contain MS loci (e.g., PTEN and NF1), and in some cases, somatic MS indels have been causally linked to cancer. Based on these findings, it is possible that the cooperation of DNA microsatellite repeat elements, which influence the transcriptional and translational landscapes of individuals, creates a permissive structural genetic environment that predisposes them to tumor formation by regulating basic cellular processes.

[0223] Microsatellite sequences (MSs) can consist of tandemly repeated sequences of 1-6 base pairs forming arrays. Over 600,000 unique MSs exist in the human genome, embedded within gene introns, exons, and regulatory regions. Due to strand-slipping replication and heterozygosity instability, the length of microsatellite loci often varies, both between alleles and between individuals. These changes can affect gene expression by inducing Z-DNA and H-DNA folding, altering nucleosome positioning, and shifting the spacing of DNA binding sites. Noncoding variants can alter DNA secondary structure and protein / RNA binding in genes near their location, leading to changes in transcriptional and translational activity and alternative splicing. For these reasons, MSs have been called "regulatory knobs" of gene expression. Within exons, microsatellite loci containing 3- or 6-base pair repeat elements can result in amino acid additions or deletions by remaining in frame with codon triplets; other non-modulo 3 lengths can lead to frameshift mutations. Genes carrying MSs may disproportionately contribute to neurological disorders. This particular vulnerability to expansion of tandem repeat sequences, particularly the CAG motif, suggests its importance in neural development. Indeed, repetitive elements have been shown to play a role in neurological diseases; in particular, polyglutamate repeats have been shown to play a role in Huntington's disease, spinocerebellar ataxia, and spinal cord muscular atrophy. Similarly, bioinformatics studies have indicated that many genes containing tandem repeats may have neurological functions.

[0224] The development of microsatellite genotyping algorithms and advances in genome sequencing have allowed the identification of germline microsatellite genotypes that can distinguish healthy individuals from affected individuals with different types of cancer (breast cancer, colon cancer, glioma, etc.) Described in this example is a panel of microsatellite genotypes that can distinguish children with MB from healthy individuals based on germline DNA.

[0225] method

[0226] Patent Sample

[0227] Germline DNA WES and WGS of medulloblastoma (MB) patients were downloaded from the following datasets: phs000504, phs000409, EGAD00001000122, EGAD00001000275, EGAD00001000816, and Waszak, SM, et. Al (Spectrum and prevalence of genetic susceptibility to medulloblastoma: a retrospective genetic study and prospective validation in a clinical trial cohort. The Lancet Oncology, Vol. 19, No. 6, pp. 785–798, the entire contents of which are incorporated herein by reference). In addition, WES from blood DNA of 6 MB patients was newly generated using the TruSeq Exome Targeted Enrichment Kit and the Illumina Sequencer HiSeq 2500. Germline DNA WES and WGS from healthy controls were downloaded from 1000 Genomes. Germline DNA WES from 100 healthy children was provided by the Hopp Center for Children's Cancer, NCT Heidelberg, Heidelberg, Germany.

[0228] Sequence mapping and coverage

[0229] Bowtie2 was used to map the WES and WGS reads to the human GRCh38 / hg38 reference genome. Overall, the coverage of the 120MB germline samples was 31-fold (31.0 ± 18.2), while the coverage of the control samples was 13-fold (13.4 ± 7.8).

[0230] Microsatellite list generation

[0231] The list of microsatellites in the human reference genome version GRCh38 / hg38 was generated using the self-defined Perl script "searchTandemRepeats.pl" using default parameters. This script can be used for microsatellite research and is freely available online. In brief, the "searchTandemRepeats.pl" script first searches for pure repeat stretches: no impurities are allowed. Imperfect repeats and composite repeats are then processed using the "mergeGap" parameter with a default value of 10 base pairs. Essentially, impurities that interrupt pure repeat sequence stretches are tolerated unless they exceed 10 base pairs. Similarly, repeats close to 10 base pairs are considered composite. The result is that the repeat sequences in the CAGm database are highly pure, and the components of the composite repeat sequences are also highly pure. The initial list generated using this script included 1,671,121 microsatellites. To mitigate the possibility of incorrect read mapping between microsatellites, a subset of microsatellites with the same repeat motif between the five base pair long 3' and 5' flanking regions were removed. For example, the microsatellite "GCTGC(A) 34"CTTAG" and "GCTGC(A)15CTTAG" were preemptively removed from the initial microsatellite list. Microsatellites can be embedded in larger repeat motifs. The filtered list included 625,195 unique microsatellites in the human genome.

[0232] Microsatellite genotyping

[0233] The program Repeatseq is used to determine the genotype of microsatellites in next-generation sequencing reads. Repeatseq uses Bayesian model selection guided by an empirically derived error model. The error model incorporates sequence and read attributes: unit, length, and basic quality. Repeatseq operates on three input files: a reference genome, a file (.bam file) containing reads aligned with the human reference genome, and a list of known microsatellites (according to the methods and systems disclosed herein). The output is a variation call format (.vcf) file, which lists the genotype of each microsatellite seat, which is composed of two alleles with the most support reads. The advantage of Repeatseq over other microsatellite genotyping programs is that it realigns each read with the reference genome before array length detection. Repeatseq can be used for the study of microsatellites and is freely available.

[0234] Repeatseq's capabilities extend to detecting somatic microsatellite variants: for example, minor alleles. Minor alleles can differ from the genotyped major allele and can be acquired somatically in normal tissues with aging. Minor alleles are used as indicators of microsatellite mutations. Briefly, detection of minor alleles is enabled through two steps built on the Repeatseq output. First, realigned read output is enabled in the Repeatseq call. Second, the realigned reads are cleared of all major alleles of the genotype. Of the remaining reads, those array lengths supported by at least three reads are counted as minor alleles. However, when comparing minor alleles across different samples, a different approach is used. Specifically, array lengths supported by at least 20% of the total read depth are counted as minor alleles.

[0235] Statistics

[0236] Based on the observation of the microsatellite genotype distribution of other cancers and controls previously, power calculation was carried out to select the size of the training set while ensuring that there were enough samples in the test set for verification. The conservative I type error probability associated with the null hypothesis test of 0.01 was selected as a part for verification. The response in each subject group can be shown as a normal distribution with a standard deviation of 1. For the real difference of the experimental and control mean being 2, the null hypothesis that the population mean of the experimental group and the control group was equal to probability (power) greater than 0.99 was rejected, and the research had 120 experimental subjects and 426 control subjects. Therefore, the training set was predicted to have enough available sample quantities.

[0237] For each microsatellite, the distribution of genotypes in the germline DNA of the two groups of samples in the training dataset was different: 120 MB and 425 healthy controls. In each case, the statistical difference was quantified using the generalized Fisher's exact test. Briefly, for each microsatellite, a contingency table was populated with the genotype counts of the two groups: MB and normal ( Figure 9 ). The p-value for each contingency table was then calculated using the Fisher test function in R. Benjamini-Hochberg multiple testing correction (n = 43,457 tested microsatellites) was applied to control the false discovery rate.

[0238] Microsatellite filtering controls for age, ethnicity, and sequencing protocol

[0239] This study was designed to identify germline microsatellite variants specific for medullary leukemia (MB); specifically, statistically significant microsatellites were identified in 120 MB samples and 425 healthy controls. However, these samples were not matched for age or sequencing protocols; furthermore, they were only partially matched for ethnicity. Therefore, this approach carries the risk of identifying microsatellites that are biased by age, sequencing, and ethnicity, rather than disease status alone. To mitigate this risk, microsatellites identified as potentially biased—by age, sequencing, or ethnicity—were excluded from subsequent analyses.

[0240] Controlling for age: To identify microsatellites whose genotypes vary non-randomly with age, 100 healthy European children and 501 European adults from the 1,000 Genomes Project were compared. Fisher's exact test identified 738 (of a total of 29,061) statistically significant microsatellites: Benjamini-Hochberg correction (p-value < 0.05) ( Figure 10 ).

[0241] Controlling for sequencing protocols: To identify microsatellites that vary based on DNA sequencing protocols (WGS vs. WES), genotypes from paired WGS and WES experiments were compared for 16 individuals in the 1000 Genomes Project. The genotype distributions of 37,511 microsatellites were tested for statistical differences (Fisher's exact test); 157 were found to be different (p-value < 0.05) using Benjamini-Hochberg false discovery correction ( Figure 11 ). This may be due to microsatellites being prone to read mapping errors, especially when they carry large insertions or deletions. Therefore, the 157 microsatellites identified may be particularly prone to mapping errors or are located in highly variable regions of the genome; they were excluded from subsequent analyses. In addition, 37,775 identified microsatellite calls were not present in 134 WGS samples. Therefore, these 37,775 cannot be used in microsatellite-based risk, diagnostic, or prognostic assays; they were excluded from subsequent analyses ( Figure 11 ).

[0242] Controlling for ethnicity: To identify DNA microsatellites that differ by ethnicity, the genotype distributions of 352 American and 502 European samples from the 1000 Genomes Project were compared and analyzed. A total of 184,981 statistical tests were performed, of which 1,037 microsatellites were revealed to be significantly different using Benjamini-Hochberg false discovery correction (p-value < 0.05). In addition, the distribution of microsatellite genotypes was examined in a set of 59 predominantly European MB samples and 55 predominantly American MB samples. Here, 13,899 tests were performed on 478 microsatellites that were found to be different after Benjamini-Hochberg false discovery correction (p-value < 0.05). 71 microsatellites that were present in both lists were identified and were excluded from further analysis ( Figure 12 ).

[0243] The number of unique microsatellites from the three steps above was 38,653; all of these were removed from further analysis.

[0244] Sample scoring indicators and ROC analysis

[0245] A metric was designed to score the samples based on their unique distribution of microsatellite genotypes. Essentially, the metric is a weighted sum of the genotypes belonging to each sample: the weights are derived from the difference in frequency of each genotype between the MB and healthy groups. Figure 13 Provides a visual summary of the metrics.

[0246] ROC Analysis: Receiver Operating Characteristic (ROC) analysis is used to design a classification scheme capable of distinguishing MB samples from healthy controls. Briefly, the area under the ROC curve (AUC) is used as a measure of the ability of the scores between the two groups to distinguish between them. A cutoff value is then selected for all future classifications. Here, the cutoff value is a single score that minimizes sensitivity while maximizing specificity; it is determined using the Youden index. ROC analysis, AUC calculation, and Youden index optimization are performed using the freely available R package: ROCR.

[0247] Microsatellite subset (genetic algorithm)

[0248] Genetic algorithms can be a class of algorithms inspired by biology. Briefly, a genetic algorithm was used to identify the most informative subset of markers from a set of 139 markers using a two-step iterative process. First, the algorithm was initialized with a random subset of 139 microsatellite markers; next, the top-ranked preformed subsets were continuously reorganized, reevaluated, and reranked. Three hyperparameters (i.e., parameters set before the iterative algorithm begins) were used to control the maximum population size, the size of each subset, the performance of each subset, and the diversity of subsets within the population. Details of each step and hyperparameter are provided below.

[0249] Initialization: Each subset of the initial population consists of randomly selected markers from the 139 perfect complements. Hyperparameters control the initial population size and the size of each subset. Once populated, the initial subsets are ranked based on the performance metrics described below.

[0250] Optimization: Each optimization cycle begins by placing 10 new subsets in the population; 7 of these are generated by recombining 2 members of the existing population (selected randomly), and 3 are randomly generated. The 2 subsets are recombined, and each subset is split; then, the two fragments (one from each subset) are rejoined. The split point and the fragments are selected randomly. The 3 random subsets are generated at initialization to help maintain the diversity of the population. Once the new subsets are generated, the population is re-ranked based on the performance metric. Finally, the 10 worst-performing subsets are discarded to maintain the population size.

[0251] Hyperparameters: A population size of 100 subsets was initialized and used throughout the algorithm. The minimum and maximum subset sizes were set to 8 and 64 markers, respectively. Duplicate markers were not allowed within the subsets. The performance of each subset was determined through receiver operating characteristic (ROC) analysis using 120MB of samples and 425 healthy controls, i.e., the same training sample was used throughout the study. The sum of sensitivity and specificity determined the performance of each subset and was used to rank the population within each generation of the genetic algorithm.

[0252] Robustness: The parameters of the genetic algorithm are chosen for computational feasibility. However, the results of the genetic algorithm are not sensitive to the choice of hyperparameters. In addition, the details of the optimization cycle (such as the number of new subsets in each cycle) do not affect the results of the genetic algorithm.

[0253] verify

[0254] Samples used: To ensure adequate power for the study, 102 experimental subjects and 428 control subjects were selected in the validation study. The subject (MB) and control distributions found when analyzing the training set ( Figure 7A ), the responses within each subject group were normally distributed with a standard deviation of 1.1. A true difference of 4.4 between the experimental and control means was rejected based on the null hypothesis that, for a sample and control validation set of this size, the probability (power) that the population means of the experimental and control groups are equal is greater than 0.99 for a Type I error probability of 0.01. All control samples used in training and validation were whole-exome sequenced. For MB, the collection included both whole-exome and whole-genome samples. Whole-genome sequenced samples were used exclusively for validation.

[0255] Procedure: Each validation sample was scored using the same metrics as the training samples. A cutoff value (identified during training) was used to predict which of the 530 validation samples had medullary leukemia (MB) and which were healthy controls. MB was predicted for validation samples exceeding the cutoff value. The predictions were compared with the known identities of 102 MB samples and 428 healthy controls. The sensitivity and specificity of these predictions were comparable to those of the training.

[0256] microsatellite mutations

[0257] To test whether individuals with MB are more prone to microsatellite variation, the total number of genotyped alleles for each microsatellite (allelic burden) was used as a measure of its mutation, and this measure was compared between disease and control cohorts. The allele restriction made the counts robust to two sources of error: (a) by requiring each allele to be supported by at least two reads, the potential influence of PCR products was mitigated; and (b) to normalize for differences in read coverage between samples, each allele needed to be supported by at least 20% of the total reads mapping to the microsatellite. Alleles were counted only for microsatellites that had mapped reads in at least 20% of the samples. Fisher's exact test was then performed to establish statistical significance between MB patients and healthy individuals. This process was repeated 50 times, with an average p-value of 0.077.

[0258] Two additional lines of evidence were used to assess the integrity of the germline mismatch repair mechanism in medulloblastoma: (a) recording homozygous and heterozygous genotypes across all microsatellites (71,192 total) in MB and control samples; and (b) comparing median microsatellite array length across all microsatellites (71,192 total) in MB and control samples. For the former analysis, aberrant mismatch repair would be expected to increase the number of heterozygous genotypes; however, the difference between case and control samples was not statistically significant. Medulloblastoma samples had a total of 299,802 heterozygous and 2,596,324 homozygous genotypes; control samples had 283,037 heterozygous and 2,449,046 homozygous genotypes. For the latter analysis, aberrant mismatch repair would be expected to lead to the accumulation of longer or shorter median microsatellite array lengths in medulloblastoma samples compared with controls; again, the results were not statistically significant. In medulloblastoma samples, 1,031 microsatellites had a shorter median array length, and 907 had a longer median array length; the remaining 69,254 microsatellites had no difference in median array length.

[0259] Downstream analysis

[0260] Functional analysis was performed using genes associated with 139 microsatellite loci whose genotypes were significantly different between MB subjects and controls. A total of 124 genes were included in the analysis, excluding microsatellites located in intergenic regions. Pathway analysis was performed using Ingenuity Pathway Analysis (QIAGEN Inc.). Mutations and co-occurrences were analyzed using PedcBioPortal. Protein-protein interaction (PPI) network construction was performed using STRING with a minimum interaction score of 0.7 (high confidence) and no more than five molecules in the first shell. This setting generated a hub with 129 nodes and 49 edges, resulting in a network with a PPI enrichment p-value of 0.0007.

[0261] result

[0262] Identification of informative microsatellite loci in medulloblastoma

[0263] Single nucleotide mutations can be characterized using whole-genome analysis of MB. Here, we investigated the impact of microsatellite variation on susceptibility to medulloblastoma. To this end, we developed a computational workflow to identify germline microsatellites that differ in genotype between children with medulloblastoma and control subjects, while correcting for those that vary with age, ethnicity, and DNA sequencing protocols. Figure 6). A metric was also developed to score each sample based on its unique collection of microsatellite genotypes. This method was applied to germline DNA sequencing data from 222 children with medulloblastoma and 853 healthy control subjects. The data were divided into two groups, both containing affected and healthy subjects, the first group for training, containing 120 medulloblastoma patients and 425 control individuals, and the second group for validation, containing 102 medulloblastoma patients and 428 control individuals. In the first stage of the analysis, using the training set, 43,457 different microsatellites present in the 120 medulloblastoma samples and 425 healthy controls were genotyped. For each of these microsatellites, a generalized Fisher's exact test was used to assess the statistical difference in the genotype distribution of each microsatellite between the two groups. 2,094 microsatellites were identified with a p-value < 0.05. After Benjamini-Hochberg multiple testing correction (α = .05), 422 passed false discovery. Three additional steps were performed to remove microsatellites that varied with age, ethnicity, and DNA sequencing protocol ( Figure 6 、 Figure 10 、 Figure 11 and Figure 12 A total of 283 microsatellites were removed from the list of 422 satellites, reducing the list to 139 ( Figure 19 In total, this approach identified 139 microsatellites from germline DNA whose genotypes were significantly different between medulloblastoma subjects and healthy controls.

[0264] Medulloblastoma Microsatellite Classifier Set

[0265] To identify the microsatellite subset with the best performance in distinguishing medulloblastoma samples from healthy controls, a 139-microsatellite set was used to train a medulloblastoma classifier. First, a metric was designed to score each medulloblastoma and control sample based on the genotype of the 139 microsatellites (see Methods and Figure 13 ). Next, a receiver operating characteristic (ROC) was generated and used to determine the ability of the sample score to act as a binary classifier for medulloblastoma. A subset optimization strategy based on a genetic algorithm approach was used to identify the best subset of distinguishing markers using a 2-step iterative process. First, subsets were randomly generated from the complete list and sorted by their F-measure. Second, the best performing subsets were continuously mixed, re-evaluated, and re-ranked. The algorithm converged within 87 cycles to reveal a subset of 43 microsatellites with an F-measure of 0.90 and an area under the curve (AUC) of 0.962 (Figure 7, Figure 20 The Youden index was determined, indicating that the optimal cutoff score for distinguishing medulloblastoma samples from healthy controls was 0.155 ( Figure 14When applied to the training set, the sensitivity was 0.88 and the specificity was 0.92 ( Figure 7B ).exist Figure 15 The chromosomal locations of these 43 markers in the human genome are shown. Thus, a set of 43 microsatellites was identified, and the genotype distribution of this set of 43 microsatellites was able to distinguish medulloblastoma patients from healthy controls with 88% sensitivity and 92% specificity.

[0266] An independent germline DNA cohort from medulloblastoma patients and healthy controls was used to validate the previous results. For the validation study, 102 experimental subjects and 428 control subjects were included, and the subject (medulloblastoma) and control distributions found when analyzing the training set (Figure 7) were used to ensure that the study had adequate power. In the training set, the responses within each subject group were normally distributed with a standard deviation of 1.1. For a true difference in the experimental and control means of 4.4, it was found that the null hypothesis that the population means of the experimental and control groups were equal could be rejected with a probability (power) greater than 0.99 and a type I error probability of 0.01 for a sample and control group of this size. The optimal cutoff value (0.155) was applied to the independent validation sample set and the classifier was found to be able to distinguish between cases and controls with a sensitivity of 0.95 and a specificity of 0.90 ( Figure 7C and Figure 7D In conclusion, a panel of 43 MS genotype profiles was identified and validated, enabling the discrimination of MB patients from healthy controls using germline DNA with high sensitivity and specificity.

[0267] Mutations of informative microsatellite loci in medulloblastoma

[0268] In the germline, the rate of indels in MS is significantly higher than the rate of single nucleotide substitutions elsewhere in the genome, corresponding to 10 -4 to 10 -3 , compared to a ratio of 10 per seat per generation. -8 However, different MSs have different mutation rates based on their repeat length, their repeat motifs, and their effects on DNA folding. Figure 20) may be the result of increased microsatellite genotypic variation inherent in individuals with MB. To test whether individuals with MB are more susceptible to microsatellite variation, the total number of genotyped alleles for each microsatellite (allelic burden) was used as a measure of its mutation, and this measure was compared between the disease and control cohorts. There was no significant difference in the number of genotyped alleles between healthy individuals and individuals with MB, supporting the conclusion that there is no widespread microsatellite instability in patients with MB. The predictive power of the informative microsatellite signatures themselves was investigated by sorting all MS by allelic burden to determine whether the 139 markers were located at the most mutable sites analyzed. It was found that although they belonged to the more mutable MS, they did not include the most mutable sites. In addition, the number of homozygous and heterozygous genotypes and the length of the microsatellite array were compared as potential sources of MB variation. In both cases, there were no statistically significant differences between MB and control germline DNA. These results and data indicate that the association of these 139 microsatellites with MB is a result of these individual microsatellite genotypes, rather than solely due to structural hypermutation.

[0269] Role of informative MST-associated genes

[0270] Of the 139 MS sites that differed in genotype between MB and control samples, 114 were located in intronic regions, 15 in intergenic regions, 6 in 3'UTR, 3 in exonic regions, and 1 in 5'UTR ( Figure 8A To understand the potential mechanistic roles of these genes, Ingenuity 124 genes associated with informative MS sites were analyzed (excluding MS located in intergenic regions). The analysis revealed statistically significant associations with cancer and molecular cellular functions such as cell cycle, DNA replication, recombination and repair, and cell growth and proliferation, indicating a relationship with cancer biology ( Figure 8B and Figure 21 The occurrence of mutations in these 124 genes associated with informative MS was examined in the 4MB cohort available in the cBioportal. Although MB tumors are known to have a low mutation rate, an average of 17% of MB cancer samples contained mutations in at least one of these 124 genes ( Figure 22 ), while the mutation rate in neuroblastoma tumors was 4.5%. Mutation co-occurrence analysis using the Sick Children 2016 dataset within the cBioportal indicated that 135 of all possible microsatellite pairs (9,591 = 139*(139-1) / 2) were found to be significantly co-occurring (p-value < 0.05). Two patients were found to have mutations in both the 20MB and 10MB informative MS sites, respectively ( Figure 23 ).

[0271] A protein-protein interaction (PPI) network consisting of 124 genes associated with informative MS sites was found ( Figure 8C ) included 129 nodes and 49 edges, resulting in a network with a PPI enrichment p-value of 0.0007. Although the number of proteins used as input was small, it was a significant hub associated with mTOR, a pathway important in macrophage tumors (PI3K / AKT / mTOR).

[0272] Three informative microsatellite loci are located in protein-coding sequences ( Figure 8A ); they are all trinucleotide repeat sequences (RAI1, BCL6B, TNS1). Variations in trinucleotide repeat sequences are believed to be the cause of neurological and neuromuscular diseases such as Huntington's disease, spinocerebellar ataxia and fragile X syndrome. Two of these genes (RAI1, BCL6B) are transcription factors located on the short arm of chromosome 17, and their deletion is a recurrent alteration in the most common subgroup of MB tumors. The BCL6B gene is associated with colon cancer, gastric cancer and liver cancer. The predominant genotype in MB tumors is 33 / 33, while the control group is 30 / 33 ( Figure 16 ); in this reading frame, the codon CAG is translated into serine. RAI1 (retinoic acid-induced protein) encodes a nuclear protein of unknown function, and its haploinsufficiency causes Smith-Magis syndrome. The two major genotypes of RAI1 in MB tumors are 38 / 41 and 41 / 41, while in controls they are 38 / 38 and 38 / 41 ( Figure 16 In addition to inducing changes in the RAI1 protein structure, short polyglutamine expansions are also thought to regulate transcription factor activity. RAI1 protein is highly expressed in the cerebellum, a region where MB tumors arise.

[0273] In this study, a panel of 139 MSs were identified as having genotypes that differed between MB patients and healthy controls. A subset of 43 MSs was able to distinguish MB individuals from controls based on their germline DNA with a sensitivity and specificity of 0.95 and 0.90, respectively.

[0274] This study identified three groups of microsatellites: (a) 43 microsatellites that collectively distinguished medulloblastoma samples from healthy controls; (b) 139 microsatellites whose genotypes were statistically different between medulloblastoma samples and healthy controls; and (c) 422 microsatellites identified in the initial screen. Microsatellites in all three groups were falsely discovered. The group of microsatellites identified in the initial screen (c) included 283 microsatellites that were sensitive to age, race, and / or DNA sequencing; therefore, they were not used in subsequent analyses. Some of the microsatellites with racial biases may also play a role in medulloblastoma. The prevalence of many diseases, including medulloblastoma, can show racial differences. Therefore, once more is understood about the genetic mechanisms that lead to medulloblastoma, re-examination of the 283 microsatellites may be feasible.

[0275] Furthermore, the relationship between the set of 139 microsatellites (b) and its subset of 43 microsatellites (a) was investigated: the latter distinguished medulloblastoma samples from healthy controls, while the former did not. Mutations within the 43 microsatellites could have a greater impact on gene expression, or genes harboring these microsatellites could have a greater influence on disease onset. This is supported by the presence of two coding microsatellites within the set of 43; in both cases, mutations directly affect the primary structure of the protein and potentially impact secondary structure and function. Furthermore, a greater proportion of the 43 microsatellites in this set are embedded in the 5' and 3' UTR regions; it is possible that MS in these regions have a stronger impact on gene expression / translation. These indications could be confirmed by studying the expression of genes harboring informative microsatellites in tumor tissue.

[0276] These results suggest that polyglutamine microsatellites embedded in the BCL6B and RAI1 genes may play a role in medulloblastoma. In the complete list of microsatellites screened, only 181 polyglutamine microsatellites (out of 627,174) were identified. Therefore, chance alone cannot explain the presence of two of the 43 informative microsatellites in the final list; using computer simulations, the probability of this occurring randomly was estimated to be approximately 1 in 1,000,000. Furthermore, polyglutamine microsatellites may play a role in diseases such as spinal and bulbar muscular atrophy, Huntington's disease, and various spinocerebellar ataxias. Furthermore, both the BCL6B and RAI1 genes may be disease-associated; the former with lymphoma and the latter with Smith-Magis syndrome. Polyglutamine disorders are characterized by insoluble protein aggregates, which are not seen in some cancers. On the other hand, polyglutamine expansions can confer both gain- and loss-of-function, depending on the affected protein.

[0277] This study revealed two overall conclusions. First, the identified microsatellites—particularly the set of 139 and a subset of 43—may play a role in the etiology of medulloblastoma. The effects of variations in microsatellite array length include effects on DNA secondary structure, nucleosome positioning, and DNA binding sites. Three of the identified microsatellites affect protein primary sequence. Microsatellites can help distinguish medulloblastoma patients from healthy controls; the classification scheme demonstrated high sensitivity and specificity of 0.95 and 0.90, respectively.

[0278] Treatment of medulloblastoma can leave survivors with lifelong burdens, including hearing loss, cognitive deficits, endocrine disorders, and a higher risk of stroke and secondary malignancies. Identifying individuals at risk for developing medulloblastoma could enable early detection strategies, leading to less invasive and more localized means of tumor control. However, an effective approach to improving the lives of these children is to prevent their tumors from developing. Recent advances in immunotherapy, including cancer vaccines, have created the potential to immunize individuals against tumor-specific antigens. Such strategies may require the selection of individuals suitable for such interventions.

[0279] Example 2 : Identification of informative microsatellite markers

[0280] From public domain database, obtain the nucleic acid sequence sample of experimenter (first group) and healthy control person (second group) suffering from disease.In two groups, all identify microsatellite locus.Compare microsatellite to reveal the difference of microsatellite locus that is only found in the first group, and specifically relevant or associated with described disease.Statistical analysis and modeling are applied to these different microsatellites, are used for their relevance or correlation with disease.In some instances, microsatellite is statistically weighted.After one group of microsatellites has been identified as being strongly associated with disease, these microsatellites are assembled in the training algorithm, to further optimize the accuracy, sensitivity and specificity that these microsatellites are associated with disease.Microsatellites during training can be randomly reorganized to generate other microsatellite combination.After training is complete, can verify algorithm with other independent sample set.

[0281] For example, the nucleic acid sequences of cancer patients and corresponding healthy controls are downloaded from The Cancer Genome Atlas (TCGA) and the Thousand Genomes Project accordingly. Microsatellite loci are all identified in both groups. The comparison of the microsatellites between the two groups reveals a microsatellite locus colony that is only found in the cancer patient group and is specifically related or associated with a type of cancer. These microsatellites associated with the cancer type are then trained on an algorithm to improve the accuracy, sensitivity, and specificity of these microsatellites associated with cancer. After training is complete, the algorithm is verified by using another sample from the group, and the other sample set either contains cancer or is from a healthy control. After verification, the algorithm can be applied to patient samples.

[0282] Example 3 : Patient Risk Assessment

[0283] During a routine health checkup, a serum sample is collected from a subject. DNA is extracted from the serum sample and sequenced. The sequencing data is processed and analyzed to generate a panel of microsatellites unique to the subject. This panel of microsatellites is then analyzed using a computer-implemented method designed to determine the risk of developing cancer based on a comparison between the subject's microsatellites and microsatellites from the Pan-Cancer Database. Each of the identified informative microsatellites is assigned a weight ranging from 0 to 1. The weights are generated based on the accuracy, sensitivity, and specificity of the identified microsatellites. The sum of the weights is then determined and used to create a classifier to determine the likelihood of developing a particular cancer. The Pan-Cancer Classifier then compiles and reports multiple classifiers for the various probabilities of developing multiple cancers for use in risk assessment of the subject. The pan-cancer classifier provides a risk assessment of a subject's likelihood of developing cancer (e.g., breast cancer, lung cancer, prostate cancer, cervical adenocarcinoma, glioblastoma multiforme, endometrial cancer, colon adenocarcinoma, bladder, urothelial carcinoma, head and neck squamous cell carcinoma, cervical squamous cell carcinoma and cervical adenocarcinoma, gastric adenocarcinoma, thyroid cancer, brain low-grade glioma, kidney papillary cell carcinoma, and hepatocellular carcinoma).

[0284] Inform subject risk assessment via laboratory reports ( Figure 5 and Figure 17 ). Information about the patient, healthcare professional, and serum sample is listed along with a test summary. The summary reveals that, although the subject does not currently have cancer, there are several identified microsatellites in the subject's genome that increase the likelihood of the subject developing lung cancer. The classifier for the likelihood of developing lung cancer includes a numerical output and is compared to a threshold for the likelihood of developing lung cancer. The threshold for the likelihood of developing lung cancer is 0.3, with a 1 standard deviation range of 0.1 and 0.5 ( Figure 24 ). The classifier for the subject's likelihood of developing lung cancer is 2.3, indicating a high probability that the subject will develop cancer in the future. Therefore, additional clinical attention is given to the subject's lungs and respiratory system. Regular and more routine lung imaging is recommended. The subject is also advised not to start smoking and to avoid prolonged exposure to certain environments containing known aerosolized carcinogens. In addition, the summary provides an overview of the risk assessment parameters, such as the type of statistical methods and thresholds used and the number of microsatellite loci analyzed.

[0285] Example 4 : Measuring genomic age using minor alleles

[0286] DNA samples from primary skin fibroblasts were obtained from subjects aged 17 and 30 years. DNA-seq libraries were constructed, subsequently sequenced using a next-generation sequencing platform, and mapped to hg19. Enrichment can be performed to enrich for hotspots in the population where minor alleles tend to appear. Minor alleles with a minimum of 5 reads were independently confirmed by Sanger sequencing. True positive minor alleles were analyzed and weighted. Examples of locations where minor alleles appear include upstream or downstream of a gene, exonic regions, intergenic regions, regions spanning introns and exons, 3'UTRs, and 5'UTRs. Minor alleles can be non-synonymous variants, synonymous variants, frameshift indels, non-frameshift indels, stop gains, stop losses, or a combination thereof.

[0287] The minor alleles obtained from the comparison between the sample obtained at age 17 and the hg19 reference sequence were analyzed by computer-implemented methods to reveal genomic age. An increase in the number of minor alleles or loci of minor alleles would result in a genomic age that is older than the subject's actual age and physical health. Samples obtained from the same subject at age 17 and age 30 can be compared to each other to reveal additional accumulation or shift in minor allele patterns within the same subject. Comparison of the minor alleles between the 17 and 30 years of age revealed a slight increase in the total number of minor alleles in the subject. This increase was analyzed by computer-implemented methods to reveal an accelerated rate of genomic aging in the subject. Therefore, it is recommended that the subject adopt a lifestyle that emphasizes nutritional balance and reduced mental stress.

[0288] Although the preferred aspects of this example have been shown and described herein, it will be apparent to those skilled in the art that such aspects are provided only as examples. Without departing from the present disclosure, those skilled in the art will now appreciate that many variations, changes, and replacements may be employed. It should be understood that in practicing the present disclosure, various alternatives to the various aspects of the present disclosure described herein may be employed. The following claims are intended to define the scope of the present disclosure, and are therefore encompassed within the scope of these claims and their equivalents.

Claims

1. A computer-implemented method for constructing an optimized classifier for a condition, the method comprising ranking a subset of a plurality of microsatellites as a classifier for the condition in a plurality of optimization cycles, wherein the subset of the plurality of microsatellites comprises microsatellites in an initial population of microsatellites associated with the condition, thereby identifying an optimized subset of the subset of microsatellites as the optimized classifier for the condition, The sorting includes: performing a receiver operating characteristic (ROC) analysis using (1) the subset of the plurality of microsatellites and (2) microsatellites in reference samples from subjects having the condition and subjects not having the condition; and adding the sensitivity and specificity of the classifier for the condition for each subset of the plurality of microsatellites, An optimization cycle in the plurality of optimization cycles comprises adding a new subset of the initial population of microsatellites to a subset from a previous optimization cycle of the plurality of optimization cycles, wherein a first portion of the new subset is generated by randomly splitting microsatellites of the randomly selected subset from the previous optimization cycle into a plurality of microsatellite fragments and recombining the plurality of microsatellite fragments of the randomly selected subset from the previous optimization cycle, and a second portion of the new subset is generated by selecting microsatellites from the initial population of microsatellites.

2. The method of claim 1 , further comprising comparing microsatellites in a first set of samples from subjects having the condition with microsatellites in a second set of samples from subjects not having the condition, thereby identifying the initial population of microsatellites.

3. The method of claim 1, wherein the ranking comprises comparing a subset of the plurality of microsatellites to microsatellites in a sample from a subject having the condition and to microsatellites in a sample from a subject not having the condition.

4. The method of claim 1 , further comprising initializing the ranking by randomly selecting a population of an initial subset of microsatellites from the initial population of microsatellites for ranking in an optimization cycle of the plurality of optimization cycles.

5. The method of claim 1, wherein a population of at least 100 subsets of the initial population of microsatellites is used in the plurality of optimization cycles.

6. The method of claim 1, wherein the minimum number of microsatellites in the subset of the subset of microsatellites is 8.

7. The method of claim 1, wherein the maximum number of microsatellites in the subset of the subset of microsatellites is 64.

8. The method of claim 1, wherein no duplicate microsatellites are allowed in the subset of the subset of microsatellites.

9. The method of claim 1, wherein the new subset comprises 10 new subsets.

10. The method of claim 9, wherein the first portion of the new subsets comprises 7 of 10 new subsets, and the second portion of the new subsets comprises 3 of 10 new subsets.

11. The method of claim 10, further comprising discarding 10 subsets of the subsets in the optimization cycle based at least in part on having the lowest ranking in the optimization cycle.

12. The method of claim 1, wherein the condition comprises the presence or absence of a health state in the subject.

13. The method of claim 1, wherein the condition comprises an increased or decreased likelihood that the subject will develop a health condition.

14. The method of claim 1, wherein the condition comprises an increased or decreased likelihood that the subject will benefit from treatment for a health condition.

15. The method of claim 1, wherein the condition comprises an increased or decreased likelihood that the subject will have an increased risk of adverse effects from treatment of a health condition.

16. The method of claim 1, wherein the condition comprises the subject's responsiveness to treatment for a health condition.

17. The method of claim 1, wherein the condition comprises a prognosis of the subject's health status.

18. The method of any one of claims 12 to 17, wherein the health condition is cancer.

19. The method of claim 18, wherein the cancer is lung cancer.

20. The method of any one of claims 12 to 17, wherein the health condition is a neurological disease.

21. The method of any one of claims 12 to 17, wherein the health condition is cardiovascular disease.

22. The method of claim 1 , wherein splitting and recombining the subset selected from the previous optimization cycle comprises randomly splitting and recombining the subset randomly selected from the previous optimization cycle, and selecting a microsatellite from the initial population of microsatellites comprises randomly selecting a microsatellite from the initial population of microsatellites.

23. A computer system configured to construct an optimized classifier for a condition, the computer system comprising: a memory comprising computer-executable instructions; and One or more processors that are individually or collectively programmed to execute the computer-executable instructions, wherein the computer-executable instructions include: ranking a subset of a plurality of microsatellites as a classifier for the condition in a plurality of optimization cycles, wherein the subset of the plurality of microsatellites includes microsatellites in an initial population of microsatellites associated with the condition, thereby identifying an optimized subset of the subset of microsatellites as the optimized classifier for the condition, The sorting includes: performing a receiver operating characteristic (ROC) analysis using (1) the subset of the plurality of microsatellites and (2) microsatellites in reference samples from subjects having the condition and subjects not having the condition; and adding the sensitivity and specificity of the classifier for the condition for each subset of the plurality of microsatellites, wherein an optimization cycle in the plurality of optimization cycles comprises adding a new subset of the initial population of microsatellites to a subset from a previous optimization cycle of the plurality of optimization cycles, and wherein a first portion of the new subset is generated by randomly splitting microsatellites of the randomly selected subset from the previous optimization cycle into a plurality of microsatellite fragments and recombining the plurality of microsatellite fragments of the randomly selected subset from the previous optimization cycle, and a second portion of the new subset is generated by selecting microsatellites from the initial population of microsatellites.

24. The computer system of claim 23, wherein the computer-executable instructions further comprise: The initial population of microsatellites is identified by comparing microsatellites in a first set of samples from subjects having the disorder to microsatellites in a second set of samples from subjects not having the disorder.

25. The computer system of claim 23, wherein the computer-executable instructions further comprise: The subset of microsatellites is compared to microsatellites in a sample from a subject having the disorder and to microsatellites in a sample from a subject not having the disorder.

26. The computer system of claim 23, wherein the computer-executable instructions further comprise: The ranking is initialized by randomly selecting a population of an initial subset of microsatellites from the initial population of microsatellites for ranking in an optimization cycle of the plurality of optimization cycles.

27. The computer system of claim 23, wherein a population of at least 100 subsets of the initial population of microsatellites is used in the plurality of optimization cycles.

28. The computer system of claim 23, wherein no duplicate microsatellites are allowed in the subset of the subset of microsatellites.

29. A non-transitory computer-readable medium comprising instructions that, when executed by a processor of a processing system, cause the processing system to perform a method for constructing an optimized classifier for a condition, the method comprising ranking a subset of a plurality of microsatellites as a classifier for the condition in a plurality of optimization cycles, wherein the subset of the plurality of microsatellites comprises microsatellites from an initial population of microsatellites associated with the condition, thereby identifying an optimized subset of the subset of microsatellites as the optimized classifier for the condition, The sorting includes: performing a receiver operating characteristic (ROC) analysis using (1) the subset of the plurality of microsatellites and (2) microsatellites in reference samples from subjects having the condition and subjects not having the condition; and adding the sensitivity and specificity of the classifier for the condition for each subset of the plurality of microsatellites, An optimization cycle in the plurality of optimization cycles comprises adding a new subset of the initial population of microsatellites to a subset from a previous optimization cycle of the plurality of optimization cycles, wherein a first portion of the new subset is generated by randomly splitting microsatellites of the randomly selected subset from the previous optimization cycle into a plurality of microsatellite fragments and recombining the plurality of microsatellite fragments of the randomly selected subset from the previous optimization cycle, and a second portion of the new subset is generated by selecting microsatellites from the initial population of microsatellites.

30. The non-transitory computer-readable medium of claim 29, wherein splitting and recombining the subset selected from the previous optimization cycle comprises randomly splitting and recombining the subset randomly selected from the previous optimization cycle, and selecting a microsatellite from the initial population of microsatellites comprises randomly selecting a microsatellite from the initial population of microsatellites.

Citation Information

Patent Citations

  • Computer system and methods for constructing biological classifiers and uses thereof

    US20070269804A1

  • Methods, systems, and software for identifying functional bio-molecules

    US20150065357A1