Microsatellite instability detection in cell-free DNA
By quantifying the repeat sequence length of microsatellite loci in cfDNA samples and using computer-implemented method scoring and threshold comparison, the challenge of detecting microsatellite instability in cell-free DNA samples was solved, achieving highly sensitive and specific assessment results to support personalized treatment decisions.
Patent Information
- Application Number
- CN201980071631.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-04
- Filing Date
- 2019-08-30
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2039-08-30
AI Technical Summary
The existing technology lacks effective methods to assess microsatellite instability status in cell-free DNA (cfDNA) samples, especially in plasma-based next-generation DNA sequencing (NGS) tests, which cannot accurately detect microsatellite instability (MSI) status and do not consider the impact of variable tumor shedding.
By quantifying the different repeat sequence lengths at multiple microsatellite loci, a computer-implemented method is used to perform site scoring and threshold comparison to identify unstable microsatellite loci. Combined with population-trained thresholds, the microsatellite instability status is determined and used to guide disease prognosis and treatment decisions.
The accurate assessment of microsatellite instability status in cfDNA samples was achieved, and the results were consistent with traditional PCR methods, with high sensitivity and strong specificity, which can guide the formulation of personalized treatment plans.
Smart Images

Figure CN112930569B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 726,182, filed on August 31, 2018, U.S. Provisional Patent Application No. 62 / 823,578, filed on March 25, 2019, and U.S. Provisional Patent Application No. 62 / 857,048, filed on June 4, 2019, and relies on the filing dates of these U.S. Provisional Patent Applications, the entire disclosures of which are incorporated herein by reference.
[0003] background
[0004] Repeat nucleic acid elements are patterns of nucleotides (DNA or RNA) that occur in multiple copies throughout the genomes of eukaryotes and prokaryotes. Examples of such repeat elements include microsatellites, short tandem repeats (STRs), and minisatellites, among others. Microsatellites typically include repeating units of less than 10 base pairs. STRs typically include repeating units of 2 to 13 nucleotides that are typically repeated hundreds of times in a given segment of nuclear DNA. STR analysis is a commonly used tool in forensic analysis. Minisatellites are repeating elements that typically have repeating units of about 10 to 60 base pairs.
[0005] In particular, microsatellites are highly polymorphic repeat regions of DNA. Microsatellite instability (MSI) is a guideline-recommended biomarker used to assess prognosis and treatment selection, including the recently approved checkpoint inhibitors for the treatment of cancers with MSI-high (MSI-H) status. Plasma-based next-generation DNA sequencing (NGS) tests are increasingly used for comprehensive genomic profiling of cancers, however, methods for detecting MSI status based on cell-free DNA (cfDNA) data are underdeveloped. Furthermore, the impact of variable tumor shedding on MSI detection has not been previously evaluated.
[0006] There remains a need for methods and related aspects that can be used to assess the status of repetitive element instability, including MSI, in various samples, especially cfDNA samples.
[0007] Overview
[0008] The present application discloses methods, computer-readable media, and systems that can be used to determine the microsatellite and / or other repetitive DNA instability status of a cell-free DNA (cfDNA) sample from a patient and help guide disease prognosis and treatment decisions. Typically, at least a portion of the methods disclosed herein are computer-implemented, and the results obtained are highly consistent with those obtained using more conventional polymerase chain reaction (PCR)-based MSI assessment methods.
[0009] From the following detailed description, other aspects and advantages of the present disclosure will become apparent to those skilled in the art, in which only illustrative embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure can have other and different embodiments, and its several details can be modified in various obvious aspects, all of which do not depart from the present disclosure. Accordingly, the drawings and description are considered to be illustrative in nature, rather than restrictive.
[0010] On the one hand, present disclosure provides a method for determining the repetitive nucleic acid instability state of a nucleic acid sample. The method includes (a) quantifying the number of different repeat sequence lengths present at each of more than one repetitive nucleic acid locus according to sequence information, to generate a site score for each of more than one repetitive nucleic acid locus. The sequence information is from a colony of repetitive nucleic acid loci in a nucleic acid sample. The method also includes (b) when the site score of a given repetitive nucleic acid locus exceeds the site-specific training threshold value of a given repetitive nucleic acid locus, a given repetitive nucleic acid locus is identified (call) as unstable, to generate a repetitive nucleic acid instability score, wherein the repetitive nucleic acid instability score includes the number of unstable repetitive nucleic acid loci from more than one repetitive nucleic acid locus. In addition, the method also includes (c) when the repetitive nucleic acid instability score exceeds the colony training threshold value of the colony of the repetitive nucleic acid locus in a nucleic acid sample, the repetitive nucleic acid instability state of a nucleic acid sample is classified as unstable, thereby determining the repetitive nucleic acid instability state of a nucleic acid sample.
[0011] On the other hand, the present disclosure provides a method for determining the instability state of the repetitive DNA of a sample (e.g., a cell-free DNA (cfDNA) sample). The method includes (a) quantifying the number of different repeat sequence lengths present at each of more than one repetitive DNA loci according to sequence information, to generate a site score for each of more than one repetitive DNA loci. The sequence information is from a population of repetitive DNA loci in the sample. The method also includes (b) for each of more than one repetitive DNA loci, comparing the site score of a given repetitive DNA locus with the site-specific training threshold of a given repetitive DNA locus. The method also includes (c) when the site score of a given repetitive DNA locus exceeds the site-specific training threshold of a given repetitive DNA locus, identifying a given repetitive DNA locus as unstable, to generate a repetitive DNA instability score, wherein the repetitive DNA instability score includes the number of unstable repetitive DNA loci from more than one repetitive DNA locus. In addition, the method further comprises (d) classifying the repetitive DNA instability state of the sample as unstable when the repetitive DNA instability score exceeds a population training threshold for the population of repetitive DNA loci in the sample, thereby determining the repetitive DNA instability state of the sample. The methods disclosed herein are generally at least partially computer-implemented.
[0012] On the other hand, the present disclosure provides a method for determining the microsatellite instability (MSI) status of a sample. The method includes (a) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on sequence information to generate a site score for each of more than one microsatellite loci, wherein the sequence information is from a population of microsatellite loci in the sample. The method also includes (b) for each of more than one microsatellite loci, comparing the site score of a given microsatellite locus with a site-specific training threshold for a given microsatellite locus. The method also includes (c) when the site score of a given microsatellite locus exceeds the site-specific training threshold for a given microsatellite locus, identifying the given microsatellite locus as unstable to generate a microsatellite instability score, the microsatellite instability score including the number of unstable microsatellite loci from more than one microsatellite locus. In addition, the method also includes (d) when the microsatellite instability score exceeds the population training threshold for a population of microsatellite loci in the sample, classifying the MSI status of the sample as unstable, thereby determining the MSI status of the sample.
[0013] On the other hand, the present disclosure provides a method for determining the microsatellite instability (MSI) status of a sample. The method includes (a) receiving sequence information from a population of microsatellite loci in the sample, and (b) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on the sequence information to generate a site score for each of more than one microsatellite loci. The method also includes (c) for each of more than one microsatellite loci, comparing the site score of a given microsatellite locus with a site-specific training threshold for a given microsatellite locus. The method also includes (d) when the site score of a given microsatellite locus exceeds the site-specific training threshold for a given microsatellite locus, identifying the given microsatellite locus as unstable to generate a microsatellite instability score, the microsatellite instability score including the number of unstable microsatellite loci from more than one microsatellite locus. In addition, the method also includes (e) when the microsatellite instability score exceeds the population training threshold for the population of microsatellite loci in the sample, classifying the MSI status of the sample as unstable, thereby determining the MSI status of the sample.
[0014] On the other hand, present disclosure provides a method for identifying one or more customized therapies for treating a disease of a subject. The method includes (a) quantifying the number of different repeat lengths present at each of more than one microsatellite locus according to sequence information, to generate a site score for each of more than one microsatellite locus, wherein the sequence information is from a population of microsatellite loci in a sample. The method also includes (b) for each of more than one microsatellite locus, the site score of a given microsatellite locus is compared with a site-specific training threshold value of a given microsatellite locus. The method also includes (c) when the site score of a given microsatellite locus exceeds the site-specific training threshold value of a given microsatellite locus, a given microsatellite locus is identified as unstable, to generate a microsatellite instability score, wherein the microsatellite instability score includes the number of unstable microsatellite loci from more than one microsatellite locus. The method also includes (d) when the microsatellite instability score exceeds the population training threshold value of a population of microsatellite loci in a sample, the MSI status of the sample is classified as unstable, to identify an unstable sample. Additionally, the method includes (e) comparing the microsatellite instability status of the sample to one or more comparator results indexed with one or more therapies to identify one or more customized therapies for treating the subject's disease.
[0015] On the other hand, the present disclosure provides a method for treating a disease of a subject. The method includes (a) quantifying the number of different repeat lengths present at each of more than one microsatellite locus according to sequence information, to generate a site score for each of more than one microsatellite locus, wherein the sequence information is from a population of microsatellite loci in a sample. The method also includes (b) for each of more than one microsatellite locus, comparing the site score of a given microsatellite locus with a site-specific training threshold of a given microsatellite locus. The method also includes (c) when the site score of a given microsatellite locus exceeds the site-specific training threshold of a given microsatellite locus, identifying a given microsatellite locus as unstable, to generate a microsatellite instability score, wherein the microsatellite instability score includes the number of unstable microsatellite loci from more than one microsatellite locus. The method also includes (d) when the microsatellite instability score exceeds the population training threshold of a population of microsatellite loci in a sample, classifying the MSI status of the sample as unstable, to identify an unstable sample. The method further includes (e) comparing the microsatellite instability status of the sample with one or more comparative results indexed by the one or more therapies to identify one or more customized therapies for treating the subject's disease. Furthermore, the method further includes (f) when there is a substantial match between the microsatellite instability status of the sample and the comparative results, administering at least one identified customized therapy to the subject, thereby treating the subject's disease.
[0016] On the other hand, the present disclosure provides a method for treating a disease in a subject. The method includes administering one or more customized therapies to the subject, thereby treating the disease in the subject, wherein the customized therapy has been identified by: (a) quantifying the number of different repeat lengths present at each of more than one microsatellite loci according to sequence information, to generate a site score for each of more than one microsatellite loci, wherein the sequence information is from a population of microsatellite loci in a sample. The method also includes (b) for each of more than one microsatellite loci, comparing the site score of a given microsatellite locus with a site-specific training threshold value of a given microsatellite locus. The method also includes (c) when the site score of a given microsatellite locus exceeds the site-specific training threshold value of a given microsatellite locus, identifying a given microsatellite locus as unstable, to generate a microsatellite instability score, wherein the microsatellite instability score includes the number of unstable microsatellite loci from more than one microsatellite locus. The method further includes (d) classifying the MSI status of the sample as unstable when the microsatellite instability score exceeds a population training threshold for the population of microsatellite loci in the sample to identify an unstable sample. The method further includes (e) comparing the microsatellite instability status of the sample with one or more comparison results indexed by one or more therapies. In addition, the method further includes (f) identifying one or more customized therapies for treating the disease of the subject when there is a substantial match between the microsatellite instability status of the sample and the comparison results.
[0017] In some embodiments, the site score of more than one microsatellite locus includes a likelihood score (likelihood score). In certain embodiments of these embodiments, the likelihood score includes a score based on probability log likelihood, and the score based on probability log likelihood distinguishes the biological signal obtained from the nucleic acid fragments (in some embodiments -cfDNA fragments) of many somatic cell sources in the sample from the noise generated by the artifacts after the sample is collected in the sample. In some embodiments, the method includes using at least two parameters to determine the score based on probability log likelihood of the individual microsatellite locus in the sequence information from the sample, wherein at least the first parameter includes allele frequency, and at least the second parameter includes at least one error pattern. Typically, the allele frequency includes the frequency of nucleic acids including different repeat lengths in the sequence information from the sample. In some embodiments, at least one error pattern includes a random error pattern and a chain-specific error pattern. In certain embodiments, the site scores for more than one microsatellite loci include the difference or ratio between (a) a score that measures the observed sequence support for the null hypothesis that a given microsatellite locus is stable, and (b) a score that measures the observed sequence support for the alternative hypothesis that a given microsatellite locus is unstable. In some embodiments, the site scores for more than one microsatellite loci are generated using one or more of the following: likelihood criterion, log-likelihood criterion, posterior probability criterion, Akaike information criterion (AIC), Bayesian information criterion, and / or the like.
[0018] In some embodiments, the site score for more than one microsatellite locus includes a site score based on the Akaike Information Criterion (AIC), which tests for the presence of somatic indels at more than one microsatellite locus. In certain embodiments of these embodiments, a given AIC-based site score is calculated using the following formula:
[0019] AIC = k-log likelihood,
[0020] Wherein k is the number of parameters used in the model. Optionally, the method comprises estimating the parameters of the model using maximum likelihood estimation (MLE). In some of these embodiments, the method comprises determining the MLE using the Nelder-Mead algorithm. In certain embodiments, the method comprises calculating the null hypothesis score of the model (e.g., a score that measures whether the observed sequence supports the null hypothesis that a given microsatellite locus is stable) using the following formula:
[0021] AIC0=k-log(Pr(obs|β,γ)),
[0022] Wherein AIC0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the length of the observed repeat sequence of the sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter. In certain embodiments, obs is the number of sequencing reads observed to cover the given microsatellite locus. In some of these embodiments, the method includes calculating the alternative hypothesis score of the model (e.g., a score that measures whether the observed sequence supports the alternative hypothesis that the given microsatellite locus is unstable) using the following formula:
[0023] AIC min =min α (k-log(Pr(obs|β,γ,α)),
[0024] Among them, AIC min is the alternative hypothesis, min α is the effect of minimizing all values of α, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1. In some embodiments, obs is the number of sequencing reads observed to cover a given microsatellite locus. In certain of these embodiments, the method includes detecting a change in the model to determine a site score (i.e., ΔAIC) using the following formula:
[0025] ΔAIC=AIC0-AIC min .
[0026] In some of these embodiments, γ comprises: (a) a read-level error rate at which the length of a microsatellite observed in a sequencing read is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule strand; and / or (b) a read-level error rate at which the length of a microsatellite observed in a sequencing read is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule strand. In certain of these embodiments, β comprises: (a) a strand-level error rate at which the expected microsatellite length of the sense strand is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (b) a strand-level error rate at which the expected microsatellite length of the antisense strand is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (c) a strand-level error rate at which the expected microsatellite length of the sense strand is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule; and / or (d) a strand-level error rate at which the expected microsatellite length of the antisense strand is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule. Generally, the method comprises identifying a given microsatellite locus as unstable when the site score for the given microsatellite locus statistically exceeds a site-specific training threshold for the given microsatellite locus.
[0027] In some embodiments, the AIC-based site score is calculated using the following formula:
[0028] AIC = 2(k-log likelihood),
[0029] Where k is the number of parameters used in the model. In these embodiments, AIC0 and AIC min Calculate using the above formula.
[0030] For purposes of clarity, in an embodiment in which the AIC-based score is determined using the formula AIC=2 (k-log likelihood), the site-specific threshold for classifying a site as unstable will be twice the site-specific threshold used in the previous embodiment in which the AIC-based score is determined using the following formula: AIC=k-log likelihood.
[0031] Typically, the mutant allele fraction (MAF) of a sample (e.g., a cfDNA sample) is estimated. In some embodiments of these embodiments, the tumor fraction of a sample (e.g., a cfDNA sample) is estimated. In certain embodiments, the tumor fraction includes the maximum mutant allele fraction (MAF) of all somatic mutations identified in the nucleic acid in the sample (e.g., a cfDNA sample). In some embodiments, the tumor fraction is lower than about 0.05%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14% or about 15% of all nucleic acids in the sample (e.g., a cfDNA sample). In some embodiments, more than one microsatellite locus includes all of the colony of the microsatellite locus, and in other embodiments, more than one microsatellite locus includes a subset of the colony of the microsatellite locus. In certain embodiments, the method includes determining a site-specific training threshold and / or a population training threshold based on sequence information from a population of microsatellite loci in one or more training DNA samples. In some of these embodiments, the training DNA samples include non-tumor cfDNA training samples and / or DNA from one or more tumor types.
[0032] In some embodiments, the method comprises a sensitivity of at least about 94% at the limit of detection (LOD) of about 0.1%-0.4% tumor fraction of nucleic acid in the sample. In some embodiments, the method comprises an analytical specificity of at least about 99% for non-tumor DNA in the sample. In certain embodiments, across a tumor fraction range of about 1% to about 15%, the determined MSI status of the sample comprises at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% consistency with the corresponding MSI status of the sample determined using a PCR-based MSI assessment technique. In some embodiments of these embodiments, the consistency is 100%. In some embodiments, the method includes classifying the MSI status of the sample as MSI-high (MSI-H) when the microsatellite instability score is greater than about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 30, about 40, about 50 or more unstable microsatellite loci from more than one microsatellite locus. In certain embodiments, the method includes classifying the MSI status of the sample as MSI-high (MSI-H) when the number of unstable microsatellite loci constitutes about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20% or about 25% of more than one microsatellite loci. In some embodiments, the number of different repeat sequence lengths includes the frequency of each different repeat sequence length present at each of more than one microsatellite loci.
[0033] In various embodiments, the present disclosure includes methods of selecting a customized therapy for treating a disease in a subject, and / or methods of treating a disease in a subject. In some of these embodiments, the disease comprises cancer, the cancer comprising at least one tumor type selected from the group consisting of, but not limited to, biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma carcinoma), transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatocellular carcinoma (liver carcinoma), hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma.
[0034] In some embodiments, the therapy includes at least one immunotherapy (e.g., checkpoint inhibitor antibodies, autologous cytotoxic T cells, personalized cancer vaccines, etc.). In certain embodiments, for example, immunotherapy includes antibodies against: PD-1, PD-2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40, or CD47. In some embodiments, immunotherapy includes administering proinflammatory cytokines against at least one tumor type. Optionally, immunotherapy includes administering T cells against at least one tumor type.
[0035] In some embodiments, the method includes obtaining a sample from a subject. Essentially any sample type is optionally used. In certain embodiments, for example, the sample is tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, feces, synovial fluid, spinal fluid, saliva and / or the like. Typically, the subject is a mammalian subject (e.g., a human subject). In some embodiments, the sample is blood. In some embodiments, the sample is plasma. In some embodiments, the sample is serum. In some embodiments, the sample includes cell-free DNA (i.e., cfDNA sample). In some embodiments, the cfDNA sample includes circulating tumor nucleic acid.
[0036] In certain embodiments, the method includes receiving sequence information generated from a sample, wherein the sequence information includes sequencing reads from a population of microsatellite loci in the sample. In some embodiments, the method includes amplifying one or more segments of nucleic acid in the sample to produce at least one amplified nucleic acid. In certain embodiments, the method includes sequencing nucleic acid from the sample to generate sequence information. In some embodiments, the sample can be a cfDNA sample. In these embodiments, the sequence information includes cfDNA sequencing reads from a population of microsatellite loci in the cfDNA sample. In some embodiments, the sequence information is obtained from a targeted segment of nucleic acid in the sample, wherein the targeted segment is obtained by selectively enriching one or more regions of nucleic acid in the sample before sequencing. In some of these embodiments, the method includes amplifying the obtained targeted segment before sequencing. In these embodiments, the method typically includes attaching one or more adapters comprising molecular barcodes to the nucleic acid before amplification. In some embodiments, the method includes attaching one or more sample indexes via amplification before sequencing. Essentially any nucleic acid sequencing technology is optionally used or applicable to perform the methods disclosed herein. For example, sequencing is optionally selected from targeted sequencing, intron sequencing, exome sequencing, whole genome sequencing, and / or the like. In some embodiments, sequencing is targeted sequencing. In some embodiments, the method comprises sequencing at least about 50, about 100, about 150, about 200, about 250, about 500, about 750, about 1,000, about 1,500, about 2,000 or more targeted genomic regions in the nucleic acid of the sample to generate sequence information.
[0037] In another aspect, the present disclosure provides a system comprising a controller comprising or having access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) receiving sequence information from a population of microsatellite loci in a sample; (b) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on the sequence information to generate a site score for each of the more than one microsatellite loci; (c) for each of the more than one microsatellite loci, assigning a given microsatellite locus a given repeat length; and (d) assigning a given repeat length to a given microsatellite locus. (d) when the site score of the given microsatellite locus exceeds the site-specific training threshold for the given microsatellite locus, identifying the given microsatellite locus as unstable to generate a microsatellite instability score, wherein the microsatellite instability score includes the number of unstable microsatellite loci from more than one microsatellite locus; and (e) when the microsatellite instability score exceeds the population training threshold for the population of microsatellite loci in the sample, classifying the MSI status of the sample as unstable, thereby determining the MSI status of the sample.
[0038] In some embodiments, the system includes a nucleic acid sequencer operably connected to a controller, which is configured to provide sequence information from a population of microsatellite loci in a sample. In some embodiments in these embodiments, the nucleic acid sequencer is configured to perform pyrophosphate sequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, synthetic sequencing (sequencing-by-synthesis), connection sequencing (sequencing-by-ligation) or hybridization sequencing (sequencing-by-hybridization) to nucleic acid to produce sequencing reads. In certain embodiments, the system includes a sample preparation component operably connected to a controller, which is configured to prepare a sample (in some cases, a cfDNA sample) to be sequenced by the nucleic acid sequencer. In some embodiments in these embodiments, the sample preparation component is configured to selectively enrich the region from the nucleic acid in the sample. In certain embodiments, the sample preparation component is configured to attach one or more adapters comprising molecular barcodes to nucleic acid. In some embodiments, the system includes a nucleic acid amplification component operably connected to a controller, which is configured to amplify DNA (in some cases, cfDNA). In certain of these embodiments, the nucleic acid amplification component is configured to amplify a selectively enriched region of nucleic acid from the sample.
[0039] In certain embodiments, the system includes a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between the nucleic acid sequencer and the sample preparation component. In some embodiments, the system includes a database operably connected to the controller, the database including one or more comparative results indexed by one or more therapies, and wherein the electronic processor further performs at least the following: (f) comparing the microsatellite instability status of the sample to the one or more comparative results, wherein a substantial match between the microsatellite instability score and the comparative result indicates a predicted response of the subject to the therapy.
[0040] In yet another aspect, the present disclosure provides a computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) receiving sequence information from a population of microsatellite loci in a sample; (b) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on the sequence information to generate a site score for each of the more than one microsatellite loci; (c) for each of the more than one microsatellite loci, comparing the site score for a given microsatellite locus to a given site score; (d) when the site score for a given microsatellite locus exceeds the site-specific training threshold for the given microsatellite locus, identifying the given microsatellite locus as unstable to generate a microsatellite instability score, the microsatellite instability score including the number of unstable microsatellite loci from more than one microsatellite locus; and (e) when the microsatellite instability score exceeds the population training threshold for the population of microsatellite loci in the sample, classifying the MSI status of the sample as unstable, thereby determining the MSI status of the sample.
[0041] The system disclosed herein and computer readable medium include various embodiments. In some embodiments, for example, the site score of more than one microsatellite locus includes a likelihood score. In certain embodiments of these embodiments, the likelihood score includes a score based on probability log likelihood, and the score based on probability log likelihood will distinguish the biological signal obtained from the nucleic acid fragments (in some embodiments -cfDNA fragments) of many somatic cell sources in the sample from the noise generated by the artifacts after the sample is collected in the sample. Typically, at least two parameters are used to determine the score based on probability log likelihood of the individual microsatellite locus in the sequence information from the sample, wherein at least the first parameter includes allele frequency, and at least the second parameter includes at least one error pattern. The allele frequency includes the frequency of nucleic acids containing different repeat lengths in the sequence information from the sample. At least one error pattern typically includes a random error pattern and a chain-specific error pattern. In some embodiments, the site scores for more than one microsatellite loci include the difference or ratio between (a) a score that measures the observed sequence support for the null hypothesis that the given microsatellite locus is stable, and (b) a score that measures the observed sequence support for the alternative hypothesis that the given microsatellite locus is unstable. In some embodiments, the site scores for more than one microsatellite loci are generated using one or more statistical model selection criteria, such as likelihood criteria, log-likelihood criteria, posterior probability criteria, Akaike Information Criterion (AIC), Bayesian Information Criterion, and / or the like.
[0042] In some embodiments of the system or computer-readable medium, the site scores for more than one microsatellite loci include site scores based on the Akaike Information Criterion (AIC), which tests for the presence of somatic gain / loss at more than one microsatellite loci. In certain embodiments, for example, a given AIC-based site score is calculated using the following formula:
[0043] AIC = k-log likelihood,
[0044] Wherein k is the number of parameters used in the model. Optionally, the parameters of the model are estimated using maximum likelihood estimation (MLE). In some of these embodiments, the MLE is determined using the Nelder-Mead algorithm. In certain embodiments, the null hypothesis score of the model is calculated using the following formula:
[0045] AIC0=k-log(Pr(obs|β,γ)),
[0046] Where AIC0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of the sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter. In some embodiments, the alternative hypothesis score of the model is calculated using the following formula:
[0047] AIC min =min α (k-log(Pr(obs|β,γ,α)),
[0048] Among them, AIC min is the alternative hypothesis, min α is the effect of minimizing all values of α, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1. In these embodiments, changes in the model are typically detected using the following formula to determine the site score:
[0049] ΔAIC=AIC0-AIC min .
[0050] In some embodiments, γ comprises: (a) the read-level error rate for a microsatellite length observed in a sequencing read that is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule strand; and / or (b) the read-level error rate for a microsatellite length observed in a sequencing read that is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule strand. In certain embodiments, β comprises: (a) the strand-level error rate for the expected microsatellite length of the sense strand that is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (b) the strand-level error rate for the expected microsatellite length of the antisense strand that is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (c) the strand-level error rate for the expected microsatellite length of the sense strand that is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule; and / or (d) the strand-level error rate for the expected microsatellite length of the antisense strand that is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule.
[0051] In certain embodiments of a system or computer-readable medium, when the site score of a given microsatellite locus statistically exceeds the site-specific training threshold value of a given microsatellite locus, a given microsatellite locus is identified as unstable. Typically, tumor score is estimated, and the tumor score is included in the maximum mutant allele fraction (MAF) of all somatic mutations identified in the nucleic acid in the sample. In certain embodiments, tumor score is lower than about 0.05%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14% or about 15% of all nucleic acids in the sample. In some embodiments, more than one microsatellite locus includes all of the colony of the microsatellite locus, and in other embodiments, more than one microsatellite locus includes the subset of the colony of the microsatellite locus. In certain embodiments, site-specific training thresholds and / or population training thresholds are determined based on sequence information from a population of microsatellite loci in one or more training DNA samples. Optionally, the MSI status of a sample is classified as MSI-high (MSI-H) when the microsatellite instability score is greater than about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50 or more unstable microsatellite loci from more than one microsatellite locus. In some embodiments, the MSI status of a sample is classified as MSI-high (MSI-H) when the number of unstable microsatellite loci constitutes about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20%, or about 25% of more than one microsatellite loci.
[0052] In yet another aspect, the present disclosure provides a system comprising a communication interface that obtains sequencing information of one or more nucleic acids in a sample from a subject via a communication network; and a computer in communication with the communication interface, wherein the computer comprises at least one computer processor and a computer-readable medium containing machine-executable code that, when executed by the at least one computer processor, implements a method comprising: (a) receiving sequence information from a population of microsatellite loci in the sample; (b) quantifying, based on the sequence information, the number of different repeat lengths present at each of more than one microsatellite loci to generate a locus for each of the more than one microsatellite loci; (c) for each of more than one microsatellite loci, comparing the site score for the given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (d) identifying the given microsatellite locus as unstable when the site score for the given microsatellite locus exceeds the site-specific training threshold for the given microsatellite locus to generate a microsatellite instability score, the microsatellite instability score comprising the number of unstable microsatellite loci from the more than one microsatellite loci; and (e) classifying the MSI status of the sample as unstable when the microsatellite instability score exceeds the population training threshold for the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.
[0053] In some embodiments, sequence information is provided by a nucleic acid sequencer. Typically, a nucleic acid sequencer performs pyrophosphate sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, synthesis sequencing, ligation sequencing, hybridization sequencing, and / or another sequencing technology on nucleic acids to generate sequencing reads. In some embodiments, a nucleic acid sequencer uses a clonal single-molecule array derived from a sequencing library to generate sequencing reads. In certain embodiments, a nucleic acid sequencer includes a chip with a micropore array to sequence the sequencing library to generate sequencing reads.
[0054] The computer-readable medium of the system disclosed herein generally includes a memory, a hard drive or a computer server. In some embodiments, the communication network includes one or more computer servers capable of distributed computing. In some embodiments, distributed computing is cloud computing. In some embodiments, the computer is located on a computer server located away from the nucleic acid sequencer. In some embodiments, the system disclosed herein includes an electronic display that communicates with the computer through a network, wherein the electronic display includes a user interface for displaying the result after implementing the method disclosed herein. In some embodiments of these embodiments, the user interface is a graphical user interface (GUI) or a network-based user interface. In some embodiments, the electronic display is in a personal computer. In certain embodiments, the electronic display is in an internet-enabled computer. In some embodiments of these embodiments, the internet-enabled computer is located away from the computer. Typically, the computer-readable medium includes a memory, a hard drive or a computer server. In some embodiments, the communication network includes a telecommunications network, the internet, an extranet or an intranet.
[0055] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate a report. The report can be in paper format or electronic format. For example, the MSI score and / or MSI status obtained by the methods and systems disclosed herein can be directly displayed in such a report. Alternatively or additionally, diagnostic information or one or more customized therapies based on the MSI status can be included in the report. In some embodiments, the report is transmitted to the subject (e.g., patient) or a healthcare provider.
[0056] In some embodiments, the method, system, or computer-readable medium further comprises classifying the repetitive nucleic acid instability state of the nucleic acid sample as stable if the repetitive nucleic acid instability score is below or at a population training threshold for a population of repetitive nucleic acid loci in the nucleic acid sample.
[0057] In some embodiments, the method, system, or computer-readable medium further comprises classifying the repetitive DNA instability state of the sample as stable if the repetitive DNA instability score is below or at a population training threshold for the population of repetitive DNA loci in the sample.
[0058] In some embodiments, the method, system, or computer-readable medium further comprises classifying the microsatellite instability status of the sample as stable if the microsatellite instability score is below or at a population training threshold for the population of microsatellite loci in the sample.
[0059] The various steps of the methods disclosed herein, or steps performed by the systems disclosed herein, can be performed at the same or different times, in the same or different geographic locations (eg, countries), and by the same or different persons. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The accompanying drawings illustrate certain embodiments and, together with the written description, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein, and are incorporated into and constitute a part of this specification. The description provided herein is better understood when read in conjunction with the accompanying drawings, which are included by way of example and not limitation. It will be understood that, unless the context indicates otherwise, like reference numerals identify similar components throughout the drawings. It will also be understood that some or all of the drawings may be schematic representations for illustrative purposes and do not necessarily depict the actual relative sizes or positions of the elements shown.
[0062] Figure 1 is a flow chart schematically depicting exemplary method steps for determining microsatellite instability (MSI) status according to some embodiments of the present invention.
[0063] Figure 2 is a schematic diagram of an exemplary system suitable for use with certain embodiments of the present invention.
[0064] Figure 3 is a graph of the limit of detection (LoD) for simulated samples (probability of detection (y-axis); mutant allele fraction (MAF) (x-axis)).
[0065] Figure 4A (MSI score (y-axis); flow cell (x-axis)) and Figure 4B (MSI score (y-axis); reference sample (x-axis)) is a graph of data from repeatability and reproducibility analysis.
[0066] Figure 5 A (maximum MAF of somatic cells (y-axis); MSI score (x-axis)) and Figure 5 B (maximum somatic MAF (y-axis); MSI score (x-axis)) is a graph showing that tumor fraction does not correlate with MSI score.
[0067] Figure 6A and Figure 6B It is a diagram showing the technical features of microsatellite detection. Figure 6AIt is a diagram showing the hierarchical clustering of the Akaike information criterion scores for 99 candidate microsatellite loci from cfDNA sequencing results of 84 healthy donors. Loci with poor unique molecular coverage are shown in black, while loci with too many technical artifacts are shown in dark gray. Robust but consistent measurements of microsatellite repeat lengths defining informative sites are shown in light gray. Arrows indicate the three Bethesda loci included in this study. Figure 6B is a graph showing the observed reduction in error rate associated with each component of digital sequencing.
[0068] Figures 7A-7C is a graph showing analytical validation of ctDNA MSI detection. The observed MSI detection rate is plotted by titration level (grey dots), and probit regression is used to determine the MSI detection rate for 5 ng ( Figure 7A ) and 30ng( Figure 7B )95% detection limit of cfDNA input. Figure 7C is a graph showing sample-level MSI scores for 499 independent replicates of two microsatellite stable (MSS) and two MSI-H artifacts run in 499 separate sequencing runs. The dashed line indicates the sample-level threshold for MSI detection.
[0069] Figure 8 is a graph showing a precision study using artificial samples processed in triplicate at three input levels (5 ng, 10 ng and 30 ng) within and between runs. Each greyscale shade represents a different run.
[0070] Figure 9 is a graph showing representative tumor types in the clinical validation cohort. In particular, only tumor types with at least 5 representative samples are shown. All other tumor types (n=25 different tumor types) are grouped in the "Other" category.
[0071] Figures 10A-10C Concordance data of ctDNA MSI status with tissue testing are shown. Figure 10A Graph showing sample-level MSI scores for 1137 cfDNA samples categorized by tissue test results and observed tumor fraction. The dashed line indicates the sample-level threshold for MSI detection. Figure 10B is a graph showing the consistency results categorized by organizational testing method. Figure 10C is a table showing descriptive statistics for the unique cohort of patients that could be evaluated.
[0072] Figures 11A-11C is a graph showing the ctDNA MSI landscape of 28,459 clinical samples. Figure 11A is a graph showing that the positive axis reports the prevalence of ctDNA MSI in the 16 most prevalent tumor types in the sample set. The negative axis reports the prevalence of tissue MSI in the 16 most prevalent tumor types in the sample set based on Hause et al. (52). The total number of samples is reported, each with the number of MSI-H samples in parentheses. Figure 11B is a graph showing sample-level MSI scores by tumor type for tumor types with ≥5 MSI-H samples. The dotted line indicates the sample-level threshold for MSI detection. Figure 11C Figure 2 is a graph showing the frequency of individual microsatellite loci contributing to MSI-H samples by tumor type for tumor types with ≥5 MSI-H samples. UCEC, uterine corpus endometrial carcinoma; STAD, gastric adenocarcinoma; COAD, colon adenocarcinoma; PRAD, prostate adenocarcinoma; COUP, carcinoma of unknown primary; BLCA, bladder cancer; CHCA, bile duct carcinoma; HNSC, head and neck squamous cell carcinoma; LUSC, lung squamous cell carcinoma; BRST, breast cancer; PANC, pancreatic adenocarcinoma; LUNG, lung cancer, not otherwise specified; LIHC, hepatocellular carcinoma; KIRC, renal cell carcinoma; OV, ovarian cancer; LUAD, lung adenocarcinoma.
[0073] Figure 12A and Figure 12B is a graph showing tumor mutation burden by MSI status. Across 278 MSI-H and 28,181 MSS samples, single nucleotide variants (SNVs) detected per sample categorized by MSI status ( Figure 12A ) and gain and loss position ( Figure 12B ) number.
[0074] Figures 13A-13E Clinical outcome data of immune checkpoint blockade (ICB) therapy in ctDNA MSI-H patients are shown. Figure 13A is a swimmer plot of the duration of pembrolizumab therapy by week. Patient 2's baseline ( Figure 13B and Figure 13C ) and after therapy ( Figure 13D and Figure 13E ) of CT( Figure 13B and Figure 13D ) and gastroscopy ( Figure 13C and Figure 13E ).
[0075] definition
[0076] In order to more easily understand the present disclosure, certain terms are first defined below. The following terms and other definitions of other terms can be set forth through the specification. If the definition of a term set forth below is inconsistent with the definition in the application or patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.
[0077] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods, and / or steps of the type described herein and / or which will become apparent to one of ordinary skill in the art upon reading this disclosure, and so forth.
[0078] It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. When describing and claiming methods, computer-readable media, and systems, the following terms and their grammatical variations will be used according to the definitions set forth below.
[0079] About: As used herein, "about" or "approximately" when applied to one or more values or elements of interest refers to a value or element similar to a stated reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that falls within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less of the stated reference value or element in either direction (greater than or less than), unless otherwise stated or otherwise apparent from the context (except where such a number would exceed 100% of the possible value or element).
[0080] Adaptor: As used herein, "adaptor" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and is used to connect either end or both ends of a given sample nucleic acid molecule. An adaptor can include a nucleic acid primer binding site and / or a sequencing primer binding site that allows amplification of nucleic acid molecules flanked by adaptors at both ends, and / or the sequencing primer binding site includes primer binding sites for sequencing applications such as various next generation sequencing (NGS) applications. An adaptor can also include a binding site for capture probes such as oligonucleotides attached to a flow cell support. An adaptor can also include a nucleic acid tag as described herein. Nucleic acid tags are typically positioned relative to the binding sites of amplification primers and sequencing primers so that the nucleic acid tag is included in the amplicon and sequencing reads of a given nucleic acid molecule. The same or different adaptors can be connected to the corresponding ends of a nucleic acid molecule. In some embodiments, adaptors of the same sequence that are different except for the nucleic acid tag are connected to the corresponding ends of a nucleic acid molecule. In some embodiments, the adapter is a Y-shaped adapter, one end of which is blunt-ended or tailed as described herein, for connecting nucleic acid molecules that are also blunt-ended or tailed with one or more complementary nucleotides. In still other exemplary embodiments, the adapter is a bell-shaped adapter, which comprises a blunt-ended or tailed end for connecting to the nucleic acid molecule to be analyzed. Other examples of adapters include T-tailed and C-tailed adapters.
[0081] Administration: As used herein, "administering" or "administering" a therapeutic agent (e.g., an immunotherapeutic agent) to a subject means giving a composition to a subject, applying a composition to a subject, or contacting a composition with a subject. Administration can be accomplished by any of a number of routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.
[0082] Akaike Information Criterion: As used herein, "Akaike Information Criterion" or "AIC" refers to a criterion for selecting a statistical model from a finite set of models and includes a penalty term for the number of parameters in the model. In some embodiments, the model with the lowest AIC is selected.
[0083] Allele frequency: As used herein, "allele frequency" refers to the relative frequency of an allele at a particular locus in a population or a given subject. Allele frequency is typically expressed as a fraction or percentage.
[0084] Amplification: As used herein, "amplify" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide or a portion of a polynucleotide, typically starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), wherein the amplification product or amplicon is typically detectable. Amplification of polynucleotides encompasses various chemical and enzymatic processes.
[0085] Barcode: As used herein, "barcode" or "molecular barcode" in the context of nucleic acids refers to a nucleic acid molecule that contains a sequence that can be used as a molecular identifier. For example, during next-generation sequencing (NGS) library preparation, a separate "barcode" sequence is typically added to each DNA fragment so that each sequencing read can be identified and sorted before final data analysis.
[0086] Cancer type: As used herein, "cancer," "cancer type," or "tumor type" refers to a type or subtype of cancer, for example, as defined by histopathology. Cancer type can be defined by any conventional criteria, such as based on occurrence in a given tissue (e.g., blood cancers, central nervous system (CNS) cancers, brain cancers, lung cancers (small cell and non-small cell), skin cancers, nose cancers, throat cancers, liver cancers, bone cancers, lymphomas, pancreatic cancers, intestinal cancers, rectal cancers, thyroid cancers, bladder cancers, kidney cancers, oral cancers, stomach cancers, breast cancers, prostate cancers, ovarian cancers, lung cancers, small intestine cancers, soft tissue cancers, neuroendocrine cancers, gastroesophageal cancers, head and neck cancers, gynecological cancers, colorectal cancers, urothelial cancers, solid state cancers, heterogeneous cancers, homogeneous cancers). Cancer can also be classified by stage (e.g., 1, 2, 3, or 4) and whether it is a primary or secondary origin.
[0087] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acids that are not contained within cells or otherwise associated with cells, or in some embodiments, refers to nucleic acids that are naturally retained in a sample after the removal of intact cells. Cell-free nucleic acids may include, for example, all unencapsulated nucleic acids derived from body fluids from a subject (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids may be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids may be released into body fluids by secretion or cell death processes, such as cell necrosis, apoptosis, etc. Some cell-free nucleic acids are released from cancer cells into body fluids, for example, circulating tumor DNA (ctDNA). Others are released from healthy cells. ctDNA can be fragmented DNA of tumor origin that is not encapsulated. Another example of cell-free nucleic acids is fetal DNA that circulates freely in the maternal bloodstream, also referred to as cell-free fetal DNA (cffDNA). Cell-free nucleic acids can have one or more epigenetic modifications, for example, cell-free nucleic acids can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated and / or citrullinated.
[0088] Comparison result: As used herein, "comparison result" means a result or a set of results to which a given test sample or test result can be compared to one or more possible properties of the test sample or result, and / or one or more possible prognostic results, and / or one or more customized therapies for a subject, the test sample being obtained from the subject or otherwise obtained. Comparison results are typically obtained from a set of reference samples (e.g., from a subject having the same disease or cancer type as the test subject and / or from a subject who receives or has received the same therapy as the test subject). In certain embodiments, for example, the microsatellite instability status of a sample (e.g., an unstable cfDNA sample) is compared with a comparison result to identify a substantial match between the microsatellite instability status of the cfDNA test sample and the microsatellite instability status determined for a set of reference samples. The microsatellite instability score determined for the set of reference samples is typically indexed with one or more customized therapies. Therefore, when a substantial match is identified, the corresponding customized therapy is also identified as a potential treatment approach for the subject from whom the test sample was obtained.
[0089] Control sample: As used herein, a "control sample" or "control DNA sample" refers to a sample of known composition and / or with known properties and / or known parameters (e.g., known tumor fraction, known coverage, known microsatellite instability score, and / or the like) that is analyzed along with or compared to a test sample to assess the accuracy of an analytical procedure.
[0090] Coverage: As used herein, "coverage" refers to the number of nucleic acid molecules that represent a specific base position.
[0091] Tailored therapy: As used herein, "tailored therapy" refers to a therapy that is associated with a desired treatment outcome for a subject or population of subjects selected based on given criteria, such as having a given microsatellite instability status or being within a defined range of a microsatellite instability score.
[0092] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to a natural or modified nucleotide with a hydrogen group at the 2'-position of the sugar moiety. DNA generally includes a nucleotide chain containing the following four types of nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to a natural or modified nucleotide with a hydroxyl group at the 2'-position of the sugar moiety. RNA generally includes a nucleotide chain containing the following four types of nucleotides: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to a natural or modified nucleotide. Certain nucleotide pairs specifically bind to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When the first nucleic acid chain is combined with a second nucleic acid chain consisting of nucleotides complementary to those in the first chain, the two chains combine to form a double strand. As used herein, "nucleic acid sequencing data", "nucleic acid sequencing information", "sequence information", "nucleic acid sequence", "nucleotide sequence", "genomic sequence", "gene sequence", or "fragment sequence", or "nucleic acid sequencing read" refers to any information or data indicating the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, or a fragment) of a nucleic acid such as DNA or RNA. It should be understood that the present teachings contemplate the use of all available various techniques, platforms or technologies, including but not limited to sequence information obtained by capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrophosphate sequencing, ion-based or pH-based detection systems, and electronic signature-based systems.
[0093] Immunotherapy: As used herein, “immunotherapy” refers to treatment with one or more agents that stimulate the immune system to kill or at least inhibit the growth of cancer cells, and is preferably used to reduce further growth of cancer, reduce the size of cancer, and / or eliminate cancer. Some such agents bind to targets present on cancer cells; some bind to targets present on immune cells but not on cancer cells; some bind to targets present on both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of immune system pathways that maintain self-tolerance and regulate the duration and amplitude of physiological immune responses in peripheral tissues to minimize collateral tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Exemplary agents include antibodies to any of the following: PD-1, PD-2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40, or CD47. Other exemplary agents include proinflammatory cytokines such as IL-1β, IL-6, and TNF-α. Other exemplary agents are T cells activated against tumors, such as T cells activated by expressing chimeric antigens that target tumor antigens recognized by the T cells.
[0094] Indel: As used herein, "Indel" refers to a mutation involving the insertion or deletion of one or more nucleotides in the genome of a subject.
[0095] Indexed: As used herein, "indexed" means that a first factor (eg, a microsatellite instability score) is correlated with a second factor (eg, a given therapy).
[0096] Instability state: As used herein, "instability state" or "instability score" in the context of repetitive nucleic acids (e.g., repetitive nucleic acid / repetitive DNA instability state or score, microsatellite instability state or score) refers to a measurement or determination of whether a given repetitive nucleic acid locus or a population of repetitive nucleic acid loci in one or more nucleic acid samples exhibits a mutation level or degree (e.g., variable repeat sequence length, etc.) that is above, at, or below a threshold level determined for the locus or population of loci. For the sake of clarity, instability state and instability score are not interchangeable, but are related concepts. Instability state is based on instability score. For example, if the instability score of a sample is below a population training threshold or at a population training threshold, the sample is classified as a stable sample (e.g., MSI-MSS or MSI-low), and if the instability score of a sample is above a population training threshold, the sample is classified as an unstable sample (e.g., MSI-MSI-high).
[0097] Limit of Detection (LoD): As used herein, "limit of detection" or "LoD" means the minimum amount of a substance (eg, nucleic acid) in a sample that can be measured by a given assay or analytical method.
[0098] Maximum MAF: As used herein, "maximum MAF" or "max MAF" refers to the maximum MAF of all somatic variants in a sample.
[0099] Microsatellite: As used herein, "microsatellite" refers to a repetitive nucleic acid having repeating units less than about 10 base pairs or nucleotides in length.
[0100] Minisatellite: As used herein, "minisatellite" refers to a repetitive nucleic acid having repeating units of about 10 to about 60 base pairs or nucleotides in length.
[0101] Mutant allele fraction: As used herein, "mutant allele fraction," "mutation dose," or "MAF" refers to the fraction of nucleic acid molecules that carry an allelic alteration or mutation at a given genomic location. MAF is typically expressed as a fraction or percentage. For example, MAF is typically less than about 0.5, 0.1, 0.05, or 0.01 (i.e., less than about 50%, 10%, 5%, or 1%) of all somatic variants or alleles present at a given locus.
[0102] Mutation: As used herein, "mutation" refers to a variation from a known reference sequence and includes mutations such as, for example, single nucleotide variations (SNVs), copy number variants or variations (CNVs) / aberrations, insertions or deletions (gains and losses), gene fusions, transversions, translocations, frameshifts, duplications, repeat sequence expansions, and epigenetic variations. The mutation can be a germline mutation or a somatic mutation. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the species of the subject providing the test sample, typically the human genome.
[0103] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to an abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor refers to a cancerous or cancerous tumor.
[0104] Next Generation Sequencing: As used herein, "next generation sequencing" or "NGS" refers to sequencing technologies with increased throughput compared to traditional Sanger and capillary electrophoresis-based methods, e.g., sequencing technologies with the ability to generate hundreds of thousands of relatively small sequence reads at a time. Some examples of next generation sequencing technologies include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.
[0105] Nucleic acid tag: As used herein, a "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides, about 100 nucleotides, about 50 nucleotides, or about 10 nucleotides in length) that is used to distinguish nucleic acids from different samples (e.g., representing a sample index), or different nucleic acid molecules of different types or that have undergone different treatments in the same sample (e.g., representing a molecular barcode). Nucleic acid tags comprise predetermined, fixed, non-random, random, or semi-random oligonucleotide sequences. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags optionally have the same length or different lengths. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, including 5' or 3' single-stranded regions (e.g., overhangs), and / or include one or more other single-stranded regions at other positions within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tag can be decoded to reveal the information of sample source, form or the processing carried out to given nucleic acid such as given nucleic acid.For example, nucleic acid tag can also be used for realizing gathering and / or parallel processing and comprise multiple samples of nucleic acid with different molecular bar codes and / or sample index, and wherein nucleic acid is deconvolved (deconvolved) by detecting (for example, reading) nucleic acid tag subsequently. Nucleic acid tag can also be referred to as identifier (for example molecular identifier, sample identifier).Additionally or selectively, nucleic acid tag can be used as molecular bar code (for example, to distinguish the amplicon of different molecules or different parent molecules in same sample or subsample).This comprises, for example, uniquely labeling the different nucleic acid molecules in given sample, or non-uniquely labeling such molecule.In the case of non-unique labeling application, a limited number of labels (for example, molecular bar code) can be used to label nucleic acid molecules so that different molecules can be distinguished based on the combination of its endogenous sequence information (for example, its mapping to the starting position and / or termination position of selected reference genome, one end of sequence or two end subsequences, and / or the length of sequence) and at least one molecular bar code. Typically, a sufficient number of different molecular barcodes are used so that the probability that any two molecules may have the same endogenous sequence information (e.g., start and / or end position, subsequence at one or both ends of the sequence, and / or length) and also have the same molecular barcode is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% probability).
[0106] Polynucleotide: As used herein, "polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogs) linked by internucleoside linkages. Typically, a polynucleotide comprises at least three nucleosides. Oligonucleotides typically range in size from a few monomeric units (e.g., 3-4) to several hundred monomeric units. Whenever a polynucleotide is represented by a string of letters such as "ATGCCTG," it will be understood that the nucleotides are in 5'→3' order from left to right, and in the case of DNA, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine, unless otherwise indicated. As is standard in the art, the letters A, C, G, and T may be used to refer to the bases themselves, nucleosides, or nucleotides comprising these bases.
[0107] Population trained threshold: As used herein, "population trained threshold" in the context of repetitive nucleic acids refers to the individually determined aggregate maximum number of unstable repetitive nucleic acid loci (e.g., the number of unstable microsatellite loci) expected to be observed in training DNA samples (e.g., non-tumor samples, tumor samples, etc.) that include the unstable repetitive nucleic acid loci. The population trained threshold is typically used to characterize the experimentally determined repetitive nucleic acid instability score for a particular sample.
[0108] Processing: As used herein, the terms "processing," "calculating," and "comparing" can be used interchangeably. In certain applications, the term refers to determining differences, such as differences in number or sequence. For example, repeat DNA instability scores (e.g., microsatellite instability scores), gene expression, copy number variation (CNV), gain / loss, and / or single nucleotide variation (SNV) values or sequences can be processed.
[0109] Reference sequence: As used herein, a "reference sequence" refers to a known sequence for the purpose of comparison with an experimentally determined sequence. For example, the known sequence can be an entire genome, a chromosome, or any segment thereof. A reference sequence typically comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more nucleotides. A reference sequence can be aligned to a single continuous sequence of a genome or chromosome, or can include non-contiguous segments aligned to different regions of a genome or chromosome. Exemplary reference sequences include, for example, human genomes, such as hG19 and hG38.
[0110] Repeat length: As used herein, "repeat length" in the context of a repetitive nucleic acid refers to the number of repeat units present at a given repetitive nucleic acid locus. For illustration, the following single-stranded nucleic acid chain has a repeat length of 8:
[0111]
[0112] Repeat unit: As used herein, a "repeat unit" in the context of a repetitive nucleic acid refers to an individual nucleotide pattern or motif (e.g., a homopolymer or heteropolymer) that is repeated at a given repetitive nucleic acid genome. For illustration, the following single-stranded nucleic acid chain has a repeat unit of "ATT":
[0113]
[0114] Repetitive nucleic acids: As used herein, "repetitive nucleic acids" or "repetitive elements" refer to a pattern of nucleotides that occur repeatedly in multiple copies throughout a given genome and / or a population of genomes. Repetitive nucleic acids include repetitive DNA and repetitive RNA. Non-limiting examples of repetitive nucleic acids include microsatellites, terminal repeats, tandem repeats, minisatellites, satellite DNA, interspersed repeats, transposable elements (e.g., DNA transposons, retrotransposons (e.g., LTR-retrotransposons (HERVs) and LTR-retrotransposons (HERVs)), etc.), clustered regularly interspaced short palindromic repeats (CRISPRs), direct repeats, inverted repeats, mirror repeats, and everted repeats.
[0115] Repetitive nucleic acid instability score: As used herein, a "repetitive nucleic acid instability score" in the context of repetitive nucleic acids (e.g., repetitive DNA instability score, microsatellite instability score, etc.) refers to the total number of repetitive nucleic acid loci from a population of repetitive nucleic acid loci identified or otherwise determined to be unstable in a given sample. This repetitive nucleic acid instability score is a sample-level score (or sample score) and is distinct from a site score, which is specific to a locus.
[0116] Sample: As used herein, "sample" means anything capable of being analyzed by the methods and / or systems disclosed herein.
[0117] Sensitivity: As used herein, "sensitivity" means the probability of detecting the presence of a mutation at a given MAF and coverage.
[0118] Sequencing: As used herein, "sequencing" refers to any of a number of techniques used to determine the sequence (eg, the identity and order of monomeric units) of a biomolecule, eg, a nucleic acid such as DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy terminator sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid phase sequencing, high throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturing temperature co-amplification PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD TM Sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as, for example, commercially available genetic analyzers from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many other companies.
[0119] Sequence information: As used herein, "sequence information" in the context of a nucleic acid polymer means the order and identity of the monomeric units (eg, nucleotides, etc.) in the polymer.
[0120] Site score: As used herein, "site score" refers to a measure of the likelihood that, in addition to the germline repeat length, additional repeat lengths are present at a given repetitive nucleic acid locus in a sample. In certain embodiments, the site score for a given locus is determined by calculating the Δ Akaike Information Criterion (AIC) for the locus.
[0121] Site specific trained threshold: As used herein, "site specific trained threshold" refers to the maximum value of a site score individually determined for a given repetitive nucleic acid locus (e.g., a given microsatellite locus) that makes the locus stable.
[0122] Somatic mutation: As used herein, "somatic mutation" refers to a genomic mutation that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0123] Specificity: As used herein, "specificity" in the context of a diagnostic analysis or assay refers to the degree to which the assay or assay detects the intended target analyte to the exclusion of other components of a given sample.
[0124] Substantial match: As used herein, "substantial match" means that at least a first value or element is at least approximately equal to at least a second value or element. For example, in certain embodiments, a customized therapy is identified when there is at least a substantial match or a near match between the microsatellite instability score and the comparator outcome.
[0125] Subject: As used herein, "subject" refers to an animal, such as a mammalian species (e.g., a human), or an avian (e.g., bird) species, or other organisms, such as a plant. More specifically, a subject can be a vertebrate, e.g., a mammal, such as a mouse, a primate, an ape, or a human. Animals include farm animals (e.g., production cattle, dairy cows, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual having or suspected of having a disease or a predisposition to the disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended to be interchangeable with "subject."
[0126] For example, the subject can be an individual who has been diagnosed with cancer, is about to receive cancer treatment, and / or has received at least one cancer treatment. The subject can be in cancer remission. As another example, the subject can be an individual who has been diagnosed with an autoimmune disease. As another example, the subject can be a female individual who is pregnant or planning a pregnancy and who may have been diagnosed with or is suspected of having a disease, such as cancer or an autoimmune disease.
[0127] Threshold: As used herein, "threshold" refers to an individually determined value used to characterize or classify an experimentally determined value.
[0128] Training DNA Samples: As used herein, "training DNA samples" refers to DNA samples used to estimate site-specific training thresholds and population training thresholds. A training DNA sample dataset includes one or more training DNA samples. The training DNA samples include one or more normal DNA samples and / or tumor DNA samples. In some embodiments, the training DNA samples include one or more samples with MSI-high and / or MSI-low / MSS status.
[0129] Tumor fraction: As used herein, "tumor fraction" refers to an estimate of the fraction of nucleic acid molecules in a given sample that are derived from a tumor. For example, the tumor fraction of a sample can be a measure derived from the maximum MAF of the sample, or the coverage of the sample, or the length of cfDNA fragments in the sample, or any other selected characteristic of the sample. In some embodiments, the tumor fraction of a sample is equal to the maximum MAF of the sample.
[0130] Instability: As used herein, "instability" or "instability" in the context of repetitive nucleic acids refers to the level of mutations (e.g., gains and losses, or the like) observed at a given repetitive nucleic acid locus or in a given population of repetitive nucleic acid loci in a nucleic acid sample (e.g., a cfDNA sample) that exceeds a threshold value (e.g., a site-specific training threshold-locus level; a population training threshold-sample level; or the like).
[0131] Details
[0132] introduction
[0133] Cancer comprises a large group of genetic diseases characterized by abnormal cell growth and the potential to metastasize beyond the site of origin within the body. The underlying molecular basis of these diseases is mutational and / or epigenetic changes that lead to phenotypic transformations of cells, whether these deleterious changes are inherited or have a somatic basis. Further complicating matters, these molecular changes often vary not only among patients with the same type of cancer but even within a given patient's own tumor.
[0134] Given the mutational variability observed in most cancers, one of the challenges of cancer care is identifying the therapy to which a patient will most likely respond given their individual cancer type. Various biomarkers are used to match cancer patients with appropriate therapies, including cancer immunotherapy. One biomarker of response is microsatellite instability (MSI), a condition or predisposition for gene hypermutation caused by impaired DNA mismatch repair (MMR) machinery. Cancer patients with cancer classified as high microsatellite instability (MSI-H or MSI-high) often exhibit accumulation of somatic mutations in tumor cells, which results in a series of molecular and biological changes, including high tumor mutation burden, increased expression of neoantigens, and abundant tumor-infiltrating lymphocytes. Chang et al. “Microsatellite Instability: A Predictive Biomarker for Cancer Immunotherapy,” Appl Immunohistochem Mol Morphol, 26(2):e15-e21 (2018). These changes are associated with increased sensitivity to checkpoint inhibitor drugs such as pembrolizumab. Pembrolizumab For the treatment of advanced melanoma, head and neck squamous cell carcinoma, non-small cell lung cancer (NSCLC) and classical Hodgkin lymphoma.To date, the application of this response biomarker has been largely limited to the assessment of MSI status in solid tumor samples using standard PCR-based techniques.
[0135] The present disclosure provides methods, computer-readable media, and systems that can be used to determine and analyze MSI in patient samples, particularly cell-free DNA (cfDNA) samples. MSI status determined using these methods and related aspects helps guide disease prognosis and treatment decisions. Results achieved using the methods and related aspects disclosed herein are generally highly consistent with those obtained using, for example, more conventional PCR-based MSI assessment methods.
[0136] Methods for determining microsatellite instability status
[0137] The present application discloses various methods for accurately determining the microsatellite instability (MSI) status and / or other repetitive DNA instability status of a sample, especially a cell-free DNA (cfDNA) sample. In certain embodiments, the method for assessing MSI status includes targeted sequencing of cfDNA, for example, using a digital sequencing platform from Guardant Health, Inc. (Redwood City, CA, USA), allowing broad coverage of simple repeat sequences, where microsatellite instability can occur across a wide range of cancer types. The digital sequencing platform is an NGS panel of cancer-related genes that utilizes high-quality sequencing of cell-free DNA (which may include circulating tumor DNA) isolated from a simple, non-invasive blood draw. Digital sequencing uses a digital library of individually tagged cfDNA molecules prepared before sequencing, combined with post-sequencing bioinformatics reconstruction to eliminate almost all false positives. To illustrate, Figure 1 A flow chart schematically depicting exemplary method steps for determining MSI status according to some embodiments of the present invention is provided. As shown, method 100 includes, at step 110, quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on sequence information to generate a site score for each of more than one microsatellite loci. Sequence information is typically obtained from a population of microsatellite loci in a cfDNA sample. As further described herein, in some embodiments, a site score based on probability log-likelihood is used to quantify the number of different repeat lengths present at a given microsatellite locus. As further described herein, other quantitative methods are optionally used as long as they are equally accurate in distinguishing biological signals obtained from relatively small amounts of somatic cell-derived cfDNA fragments from noise generated by, for example, technical or sample collection artifacts (e.g., amplification artifacts, sequencing artifacts, and the like) in the sample.
[0138] Method 100 also includes, at step 112, comparing the site score of a given microsatellite locus with the site-specific training threshold of the specific microsatellite locus. For each of more than one microsatellite locus, the experimentally determined site score of the specific locus is typically compared with its corresponding site-specific training threshold. The site-specific training threshold of a given locus is typically a predetermined value of a specific locus derived from a population of training DNA samples (such as a normal cfDNA sample or a cohort of non-tumor cfDNA samples). As shown, method 100 also includes, at step 114, identifying the given microsatellite locus as unstable when the site score (e.g., likelihood score, etc.) of the given microsatellite locus exceeds (e.g., is statistically greater than) the site-specific training threshold of the given microsatellite locus. Based on these comparisons, a microsatellite instability score is generated, which includes the number of microsatellite loci identified as unstable from more than one microsatellite locus (e.g., the overall MSI score or the total MSI score of the sample). In addition, method 100 also includes, at step 116, classifying the MSI status of the cfDNA sample as unstable when the microsatellite instability score exceeds a population training threshold for the population of microsatellite loci in the cfDNA sample, thereby identifying an unstable cfDNA sample (e.g., scoring or predicting the sample as MSI-high). In other words, in certain embodiments, the MSI status of the sample is determined by the presence of a minimum number of unstable microsatellite loci. The population training threshold is typically a predetermined value obtained from a population of training DNA samples, such as a predetermined value obtained from a cohort of normal cfDNA samples or non-tumor cfDNA samples.
[0139] In some embodiments, a threshold value (e.g., a site-specific training threshold value, a population training threshold value, etc.) is determined or otherwise obtained from at least one training DNA sample data set. The training DNA sample data set typically includes at least about 25 to at least about 30,000 or more training samples. In some embodiments, the training DNA sample data set includes about 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000 or more training DNA samples.
[0140] In certain embodiments, method 100 includes additional upstream and / or downstream steps. In some embodiments, for example, method 100 starts at step 102 and provides a sample from a subject at step 104 (e.g., providing a blood sample obtained from a subject). In these embodiments, the workflow of method 100 generally further includes, at step 106, amplifying the nucleic acid in the sample to produce amplified nucleic acid, and at step 108, sequencing the amplified nucleic acid to produce sequence information, and then at step 110, quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on the sequence information. Nucleic acid amplification (including related sample preparation), nucleic acid sequencing, and related data analysis are also described herein.
[0141] In some embodiments, method 100 includes various steps downstream of identifying an unstable cfDNA sample at step 116. Some examples of these steps include comparing the microsatellite instability status of the cfDNA sample with a comparison result indexed by therapy to identify a customized therapy for treating the subject's disease (e.g., cancer or another genetically based disease, disorder, or condition) at step 118. In other exemplary embodiments, method 100 further includes administering at least one identified customized therapy to the subject (e.g., to treat the subject's cancer or another disease, disorder, or condition) at step 120 when there is a substantial match between the microsatellite instability status of the sample and the comparison result, before ending at step 122.
[0142] Method described herein includes various selectable embodiments.For example, the site score of microsatellite locus optionally includes likelihood score.In some embodiments, likelihood score includes the score based on probability log likelihood.In some embodiments in these embodiments, method includes using each parameter, such as allele frequency and one or more error patterns (for example, random error pattern, chain-specific error pattern and / or similar error pattern), determines the score based on probability log likelihood of the individual microsatellite locus in the sequence information obtained from the sample.Allele frequency is generally included in the observed frequency of the nucleic acid with different repeat lengths at the given microsatellite locus in the sequence information obtained from the sample.In some embodiments, the site score of specific microsatellite locus includes the difference or ratio between following (a) and (b): (a) metric observed nucleotide sequence supports the scoring of stable null hypothesis given microsatellite locus, and (b) metric observed nucleotide sequence supports the scoring of unstable alternative hypothesis given microsatellite locus. The null hypothesis is the hypothesis with the smallest AIC score among all hypotheses assuming that the site is stable, and the alternative hypothesis is the hypothesis with the smallest AIC score among all hypotheses assuming that the site is unstable. Typically, site scores are generated using various measures of model accuracy, such as likelihood criterion, log-likelihood criterion, posterior probability criterion, Akaike Information Criterion (AIC), Bayesian Information Criterion, and / or the like. Additional details regarding statistical modeling, including measures of statistical model accuracy, optionally applicable to performing the methods disclosed herein, are provided in, for example, Bruce, Practical Statistics for Data Scientists: 50 Essential Concepts, 1st ed., O'Reilly Media (2017); Freedman et al., Statistics, 4th ed., WW Norton & Company (2007); James et al., An Introduction to Statistical Learning: with Applications in R, 1st ed., Springer (2013), and Hastie et al., The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed., Springer (2016), each of which is incorporated herein by reference in its entirety.
[0143] For further illustration, the site score for a given microsatellite locus optionally includes an AIC-based site score that tests for the presence of somatic gain / loss at the microsatellite locus. In some of these embodiments, the given AIC-based site score is calculated using the following formula:
[0144] AIC = k-log likelihood,
[0145] Where k is the number of parameters used in the model. In some embodiments, the method includes estimating the parameters of the model using maximum likelihood estimation (MLE) (e.g., using the Nelder-Mead algorithm or another simplex search algorithm). The method optionally includes calculating the null hypothesis of the model using the following formula:
[0146] AIC0=k-log(Pr(obs|β,γ)).
[0147] Wherein AIC0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs is the number of sequencing reads observed to cover a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter. In some of these embodiments, the method comprises calculating the alternative hypothesis of the model using the following formula:
[0148] AIC min =min α (k-log(Pr(obs|β,γ,α)),
[0149] Among them, AIC min is the alternative hypothesis, min α is the effect of minimizing all values of α, k is the number of parameters used in the model, Pr is the probability, obs is the observed number of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of the values equals 1. Changes in the model used to determine the site score (ΔAIC) are typically detected using the following formula:
[0150] ΔAIC=AIC0-AIC min .
[0151] In certain embodiments, the parameter γ comprises: (a) the read-level error rate at which the microsatellite length observed in the sequencing read is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule strand; and / or (b) the read-level error rate at which the microsatellite length observed in the sequencing read is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule strand. In some embodiments, the parameter β comprises: (a) the strand-level error rate at which the expected microsatellite length of the sense strand is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (b) the strand-level error rate at which the expected microsatellite length of the antisense strand is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (c) the strand-level error rate at which the expected microsatellite length of the sense strand is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule; and / or (d) the strand-level error rate at which the expected microsatellite length of the antisense strand is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule.
[0152] In some embodiments, the AIC-based site score is calculated using the following formula:
[0153] AIC = 2(k-log likelihood),
[0154] Where k is the number of parameters used in the model. In these embodiments, AIC0 and AIC min Calculate using the above formula.
[0155] For purposes of clarity, in an embodiment in which the AIC-based score is determined using the formula AIC=2 (k-log likelihood), the site-specific threshold for classifying a site as unstable will be twice the site-specific threshold used in the previous embodiment in which the AIC-based score is determined using the following formula: AIC=k-log likelihood.
[0156] The sample using the method for analyzing described herein generally includes various mutant allele fractions (MAF) (for example, the sample fraction that shows the different repeat lengths of specific microsatellite locus or other allele changes). In addition, in some embodiments, the sample includes a tumor fraction. In certain embodiments, maximum MAF (max MAF) is used as the approximate value of the tumor fraction in a given sample. The tumor fraction is generally lower than about 0.05%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14% or about 15% of all nucleic acids in the sample.
[0157] In some embodiments, the methods disclosed herein typically comprise a sensitivity of at least about 94% at a limit of detection (LOD) of about 0.2% tumor fraction of nucleic acid in a given sample. The methods also typically have a specificity of at least about 99% for non-tumor DNA in the sample. Across a tumor fraction range of about 1.4% to about 15%, the determined MSI status of the sample typically has at least about 95%, 96%, 97%, 98%, or 99% concordance with the corresponding MSI status of the sample determined using a standard PCR-based MSI assessment technique. In some embodiments, the concordance is 100%.
[0158] In certain embodiments, the MSI status of a particular sample is classified as MSI-high (MSI-H) when the microsatellite instability score of the sample is greater than about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, or more than 100 unstable microsatellite loci in the sample. In certain embodiments, the population training threshold for determining the instability status (e.g., MSI status) of a sample is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, or more than 100 unstable repetitive nucleic acid (e.g., microsatellite) loci. In some embodiments, the population training threshold for a sample is about 5 unstable microsatellite loci. In some embodiments, the population training threshold for a sample is about 6 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold for a sample is about 10 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 15 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 16 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 20 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 25 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 26 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 30 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 35 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 36 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 40 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 45 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold of the sample is about 46 unstable repetitive nucleic acid loci. In some embodiments, the population training threshold for a sample is about 50 unstable repetitive nucleic acid loci.In some embodiments, the repetitive nucleic acid loci can be microsatellite loci.In some embodiments, the MSI status of a given sample is classified as MSI-H when the number of unstable microsatellite loci constitutes about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20%, or about 25% of all microsatellite loci estimated in the sample. In some embodiments, about 50, about 60, about 70, about 80, about 90, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000 or more repetitive nucleic acid (e.g., microsatellite) loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 50 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 60 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 70 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 80 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 90 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 100 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 200 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 300 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 400 repetition nucleic acid loci are used to determine the repetition nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 500 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1000 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1100 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1200 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample.In some embodiments, about 1300 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 1400 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, at least 1500 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, about 1600 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, at least 1700 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, at least 1800 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, at least 1900 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) state of a given sample. In some embodiments, at least 2000 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, the repetitive nucleic acid loci can be microsatellite loci. In some embodiments, the repetitive nucleic acid instability status can be an MSI status.
[0159] In some embodiments, the method includes obtaining a sample from a subject. Substantially any sample type is optionally used. In certain embodiments, for example, the sample is tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, feces, synovial fluid, spinal fluid, saliva and / or the like. Other exemplary sample types optionally used are also described herein. Typically, the subject is a mammalian subject (e.g., a human subject). Substantially any type of nucleic acid (e.g., DNA and / or RNA) can be assessed according to the methods disclosed in this application. Some examples include cell-free nucleic acids (e.g., cfDNA from tumor origin, fetal origin, maternal origin, and / or similar sources), cellular nucleic acids, including those from circulating tumor cells (e.g., obtained by lysing intact cells in a sample), circulating tumor nucleic acids, etc. In some embodiments, the sample includes cell-free DNA (cfDNA sample). In some embodiments, the cfDNA sample includes circulating tumor nucleic acids.
[0160] The methods disclosed herein generally comprise obtaining sequence information from a nucleic acid in a sample taken from a subject. In certain embodiments, the sequence information is obtained from a targeted segment of the nucleic acid. Essentially any number of genomic regions are optionally targeted. The targeted segment may comprise at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000, or at least 50,000 (e.g., 25, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100, 3200, 3300, 3400, 3500, 3600, 3700, 3800, 3900, 4000, 4100, 4200, 4300, 4400, 4500, 4600, 4700, 4800, 4900, 5000 In some embodiments, the targeted segments include selected regions of at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, or at least 700 genes. In some embodiments, the targeted segments include selected regions of at least 70 genes. In some embodiments, the targeted segment includes a region of at least 500 genes.
[0161] In these embodiments, the method generally also includes each sample or library preparation step, to prepare the nucleic acid for sequencing. Many different sample preparation techniques are well known to those skilled in the art. Substantially any technology in these technologies is used for or is applicable to perform methods described herein. For example, in addition to each purification step for separating nucleic acid from other components in a given sample, the typical steps for preparing the nucleic acid for sequencing include labeling the nucleic acid with a molecular identifier or a barcode, adding an adapter (for example, which may include a barcode), amplifying nucleic acid once or more times, enriching the targeted section of nucleic acid (for example, using various targeted capture strategies, etc.), and / or similar steps. Exemplary library preparation processes are also described herein. Additional details regarding nucleic acid sample / library preparation are also described in, for example, van Dijk et al., Library preparation methods for next-generation sequencing: Tone down the bias, Experimental Cell Research, 322(1):12-20 (2014), Micic (Ed.), Sample Preparation Techniques for Soil, Plant, and Animal Samples (Springer Protocols Handbooks), 1st ed., Humana Press (2016), and Chiu, Next-Generation Sequencing and Sequence Data Analysis, Bentham Science Publishers (2018), each of which is herein incorporated by reference in its entirety.
[0162] The instability state of microsatellites and / or other repetitive nucleic acids determined by the methods disclosed herein is optionally used to diagnose a disease or condition in a subject, particularly the presence of cancer, to characterize such a disease or condition (e.g., staging a given cancer, determining the heterogeneity of the cancer, etc.), monitor the response to treatment, assess the potential risk of developing a given disease or condition, and / or assess the prognosis of a disease or condition. The instability state of microsatellites and / or other repetitive nucleic acids is also optionally used to characterize a specific form of cancer. Since cancer is typically heterogeneous in both composition and staging, the instability state data of microsatellites and / or other repetitive nucleic acids can allow characterization of a specific subtype of cancer, thereby helping diagnosis and treatment selection. This information can also provide clues about the prognosis of a specific type of cancer for a subject or a health care practitioner, and enables a subject and / or a health care practitioner to adjust treatment options according to the progression of the disease. As cancer progresses, some cancers become more aggressive and genetically unstable. Other tumors remain benign, inactive, or dormant.
[0163] The instability state of microsatellite and / or other repetition nucleic acid can also be used to determine disease progression and / or monitor recurrence.In some cases, for example, successful treatment may initially increase the instability of observed microsatellite and / or other repetition nucleic acid along with the number of cancer cell death and shedding nucleic acid.In these cases, along with therapy progress, the instability of microsatellite and / or other repetition nucleic acid will usually reduce along with the continuation of tumor size subsequently.In other cases, successful treatment can also reduce the instability of microsatellite and / or other repetition nucleic acid, and does not have the initial increase of such instability.In addition, if observe that cancer is alleviated after treatment, the instability state of microsatellite and / or other repetition nucleic acid can be used to monitor the disease remaining in the patient or the recurrence of disease.
[0164] sample
[0165] The sample can be any biological sample separated from the subject. The sample can include body tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (white blood cell) or white blood cells (leucocyte), endothelial cells, tissue biopsy (for example, from a biopsy of a known or suspected solid tumor), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid (for example, from fluid between cells), gingival fluid, gingival sulcus fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. The sample is preferably a body fluid, particularly blood and its fraction, and urine. Such a sample includes nucleic acids shed from a tumor. Nucleic acids can include DNA and RNA, and can be in double-stranded form and single-stranded form. The sample can be in the form initially separated from the subject, or can be subjected to further processing to remove or add components, such as cells, enrich a component relative to another component, or convert a form of nucleic acid into another, such as RNA into DNA or single-stranded nucleic acid into double-stranded. Thus, for example, the bodily fluid sample used for analysis is plasma or serum containing cell-free nucleic acids, such as cell-free DNA (cfDNA).
[0166] In some embodiments, the volume of the sample taken from the subject's body fluid depends on the desired read depth of the sequencing region. Exemplary volumes are about 0.4 ml to 40 ml, about 5 ml to 20 ml, or about 10 ml to 20 ml. For example, the volume can be about 0.5 ml, about 1 ml, about 5 ml, about 10 ml, about 20 ml, about 30 ml, about 40 ml, or more milliliters. The volume of the sampled plasma is typically between about 5 ml and about 20 ml.
[0167] Samples can contain various amounts of nucleic acid. Typically, the amount of nucleic acid in a given sample is equivalent to multiple genome equivalents. For example, a sample of about 30 ng of DNA may contain about 10,000 (10 4 ) haploid human genome equivalents, and in the case of cfDNA, can contain approximately 200 billion (2 × 10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, can contain about 600 billion individual molecules.
[0168] In some embodiments, the sample includes nucleic acids from different sources, for example, from cell sources and from cell-free sources (e.g., blood samples, etc.). Typically, the sample includes nucleic acids carrying mutations. For example, the sample optionally includes DNA carrying germline mutations and / or somatic mutations. Typically, the sample includes DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations).
[0169] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), e.g., about 1 picogram (pg) to about 200 nanograms (ng), about 1 ng to about 100 ng, about 10 ng to about 1000 ng. In some embodiments, the sample includes up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the amount is up to about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the method comprises obtaining between about 5 ng and about 30 ng of cell-free nucleic acid molecules from a sample. In certain embodiments, the method comprises obtaining between about 5 ng and about 100 ng of cell-free nucleic acid molecules from a sample. In certain embodiments, the method comprises obtaining between about 5 ng and about 150 ng of cell-free nucleic acid molecules from a sample. In certain embodiments, the method comprises obtaining between about 5 ng and about 200 ng of cell-free nucleic acid molecules from a sample. In some embodiments, the amount is up to about 100 ng of cell-free nucleic acid molecules from a sample. In some embodiments, the amount is up to about 150 ng of cell-free nucleic acid molecules from a sample. In some embodiments, the amount is up to about 200 ng of cell-free nucleic acid molecules from a sample. In some embodiments, the amount is up to about 250 ng of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is up to about 300 ng of cell-free nucleic acid molecules from the sample. In some embodiments, the method comprises obtaining between about 1 fg and about 200 ng of cell-free nucleic acid molecules from the sample.
[0170] The cell-free nucleic acids typically have a size distribution between about 100 nucleotides in length and about 500 nucleotides in length, with molecules between about 110 nucleotides in length and about 230 nucleotides in length representing about 90% of the molecules in the sample, with a mode of about 168 nucleotides in length and a secondary minor peak ranging from about 240 nucleotides in length to about 440 nucleotides in length. In certain embodiments, the cell-free nucleic acids are about 160 nucleotides in length to about 180 nucleotides in length, or about 320 nucleotides in length to about 360 nucleotides in length, or about 440 nucleotides in length to about 480 nucleotides in length.
[0171] In some embodiments, cell-free nucleic acid is separated from body fluid by a partitioning step, in which the cell-free nucleic acid found in the solution is separated from intact cells and other insoluble components in the body fluid. In some embodiments of these embodiments, distribution includes techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid are lysed, and the cell-free nucleic acid and cellular nucleic acid are processed together. Typically, after adding buffer and washing steps, the cell-free nucleic acid is precipitated with, for example, alcohol. In certain embodiments, other clean-up steps such as silica-based columns are used to remove contaminants or salts. For example, non-specific bulk carrier nucleic acids are optionally added throughout the reaction to optimize certain aspects of the exemplary procedure, such as yield. After such processing, the sample typically comprises various forms of nucleic acid, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. Optionally, single-stranded DNA and / or single-stranded RNA are converted into double-stranded forms so that they are included in subsequent processing and analysis steps.
[0172] Nucleic acid tags
[0173] In some embodiments, nucleic acid molecules (from samples of polynucleotides) can be tagged with sample indexes and / or molecular barcodes (commonly referred to as "tags"). Tags can be incorporated into adapters or otherwise connected to adapters by chemical synthesis, connection (e.g., blunt end connection or sticky end connection) or overlap extension polymerase chain reaction (PCR) and other methods. Such adapters can ultimately be connected to target nucleic acid molecules. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are typically applied to introduce sample indexes into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification can be carried out in one or more reaction mixtures (e.g., more than one microwell in an array). Molecular barcodes and / or sample indexes can be introduced simultaneously or in any order. In some embodiments, molecular barcodes and / or sample indexes are introduced before and / or after performing a sequence capture step. In some embodiments, only molecular barcodes are introduced before probe capture, and sample indexes are introduced after performing a sequence capture step. In some embodiments, both molecular barcodes and sample indexes are introduced before performing a probe-based capture step. In some embodiments, the sample index is introduced after the sequence capture step is performed. In some embodiments, the molecular barcode is incorporated into the nucleic acid molecules (e.g., cfDNA molecules) in the sample via connection (e.g., blunt end connection or sticky end connection) through an adapter. In some embodiments, the sample index is incorporated into the nucleic acid molecules (e.g., cfDNA molecules) in the sample by overlap extension polymerase chain reaction (PCR). Typically, the sequence capture scheme involves the introduction of a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence, such as a coding sequence of a genomic region, and mutations in such a region are associated with cancer types.
[0174] In some embodiments, the label can be located at one end or both ends of the sample nucleic acid molecule. In some embodiments, the label is a predetermined or random or semi-random sequence oligonucleotide. In some embodiments, the length of the label can be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 nucleotides. The label can be randomly or non-randomly connected to the sample nucleic acid.
[0175] In some embodiments, each sample is uniquely labeled by a combination of sample indexes or sample indexes. In some embodiments, each nucleic acid molecule of a sample or subsample is uniquely labeled by a molecular barcode or a combination of molecular barcodes. In other embodiments, more than one molecular barcode can be used so that the molecular barcodes are not necessarily unique to each other in the more than one molecular barcode (e.g., non-unique molecular barcodes). In these embodiments, molecular barcodes are typically attached (e.g., by connection) to individual molecules so that the combination of molecular barcodes and the sequences to which they can be attached produces a unique sequence that can be tracked individually. The detection of a combination of non-uniquely labeled molecular barcodes and endogenous sequence information (e.g., corresponding to the beginning (start) and / or end (termination) portion of the original nucleic acid molecule sequence in the sample, a subsequence of the sequence read at one end or both ends, the length of the sequence read and / or the length of the original nucleic acid molecule in the sample) generally allows for the assignment of a unique identity to a specific molecule. The length or base pair number of the individual sequence reads are also optionally used to assign a unique identity to a given molecule. As described herein, fragments from a single strand of nucleic acid that have been assigned a unique identity can thereby allow subsequent identification of fragments from the parental strand and / or the complementary strand.
[0176] In some embodiments, molecular barcodes are introduced with a ratio of a set of expected identifiers (e.g., a combination of unique molecular barcodes or non-unique molecular barcodes) to the molecules in the sample. An exemplary form uses about 2 to about 1,000,000 different molecular barcodes, or about 5 to about 150 different molecular barcodes, or about 20 to about 50 different molecular barcodes. Alternatively, about 25 to about 1,000,000 different molecular barcodes can be used. The molecular barcode can be connected to both ends of the target molecule. For example, 20-50 × 20-50 molecular barcodes can be used. In some embodiments, 20-50 different molecular barcodes can be used. In some embodiments, 5-100 different molecular barcodes can be used. In some embodiments, 5-150 molecular barcodes can be used. In some embodiments, 5-200 different molecular barcodes can be used. Such a number of identifiers is generally sufficient to provide a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of different molecules having the same start and end points receiving different identifier combinations. In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same molecular barcode combination.
[0177] In some embodiments, the assignment of unique or non-unique molecular barcodes in a reaction is performed using methods and systems such as those described in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, only endogenous sequence information (e.g., start and / or end position, subsequences at one or both ends of the sequence, and / or length) can be used to identify different nucleic acid molecules of a sample.
[0178] Nucleic acid amplification
[0179] Sample nucleic acids flanked by adapters are typically amplified by PCR and other amplification methods using nucleic acid primers that bind to primer binding sites in the adapters flanking the DNA molecule to be amplified. In some embodiments, the amplification method involves cycles of extension, denaturation, and annealing generated by thermal cycling, or can be isothermal, for example, in transcription-mediated amplification. Other exemplary amplification methods that are optionally used include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and autonomous continuous sequence-based replication, among other methods.
[0180] Conventionally, one or more rounds of amplification cycles are used to introduce sample index into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification is typically carried out in one or more reaction mixtures. Molecular tags and sample index / labels are optionally introduced simultaneously or in any order. In some embodiments, molecular tags and sample index / labels are introduced before and / or after performing a nucleic acid molecule capture step (i.e., nucleic acid enrichment). In some embodiments, only molecular tags are introduced before probe capture, and sample index / labels are introduced after performing a sequence capture step. In certain embodiments, both molecular tags and sample index / labels are introduced before performing a probe-based capture step. In some embodiments, sample index / labels are introduced after performing a sequence capture step. Typically, sequence capture schemes involve introducing single-stranded nucleic acid molecules complementary to targeted nucleic acid sequences, such as coding sequences in genomic regions, and mutations in such regions are associated with cancer types. Typically, the amplification reaction produces more than one non-uniquely or uniquely tagged nucleic acid amplicon having a molecular tag and a sample index / tag having a size range of about 200 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicon has a size of about 300 nt. In some embodiments, the amplicon has a size of about 500 nt.
[0181] Nucleic acid enrichment
[0182] In some embodiments, before nucleic acid sequencing, enrichment sequence is carried out. Specific target region (" target sequence ") is optionally enriched. In some embodiments, the target region of interest can be enriched with the nucleic acid capture probe (" bait ") selected for one or more bait set groups (baitset panels) using differential tiling (differential tiling) and capture scheme. Differential tiling and capture scheme typically use the bait group of different relative concentrations to cross differential tiling (for example, with different " resolutions ") in the genomic region relevant to the bait, subject to one group of restrictions (for example, sequencer restrictions, such as sequencing capacity, the effectiveness of every kind of bait etc.), and capture the targeted nucleic acid with the desired level of downstream sequencing. These targeted genomic regions of interest optionally include the natural nucleotide sequence or the synthetic nucleotide sequence of the nucleic acid construct. In some embodiments, the biotin-labeled beads with the probe for one or more regions of interest can be used to capture the target sequence, and optionally subsequently amplify these regions with the region of interest of enrichment.
[0183] Sequence capture is generally related to the use of oligonucleotide probes that hybridize with the target nucleic acid sequence.In certain embodiments, the probe set strategy is related to the probe being tiled on the region of interest.The length of such probe can be, for example, about 60 to about 120 nucleotides.The group can have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 30x, 40x, 50x or more than 50x.The effectiveness of sequence capture depends in part on the length of the sequence in the target molecule with the sequence complementarity (or almost complementary) of the probe usually.
[0184] Nucleic acid sequencing
[0185] The sample nucleic acid of optional flank is generally subjected to sequencing with or without pre-amplification.The sequencing method optionally used or commercially available form include, for example, Sanger sequencing, high throughput sequencing, pyrophosphate sequencing, synthesis sequencing, single molecule sequencing, sequencing based on nanopore, semiconductor sequencing, connection sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next generation sequencing (NGS), single molecule synthesis sequencing (SMSS) (Helicos), large-scale parallel sequencing, cloned single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, use PacBio, SOLiD, Ion Torrent or nanopore platform sequencing.Sequencing reaction can be carried out in many kinds of sample processing units, and the sample processing unit can include multiple lanes (multiple lanes), multichannel, multiporous or other devices that substantially process multiple sample groups simultaneously.The sample processing unit can also include multiple sample chambers, to realize processing multiple operations simultaneously.
[0186] Sequencing reaction can be carried out to one or more nucleic acid fragment types or the region of the marker (for example, microsatellite and / or other repeat nucleic acid elements) known to comprise cancer or other diseases. Sequencing reaction can also be carried out to any nucleic acid fragment present in the sample. Sequence reaction can provide at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% genomic sequence coverage. In other cases, genomic sequence coverage can be less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of genome. In some embodiments, the sequence coverage of a genome can be less than about 0.01%, 0.02%, 0.05%, 0.1%, 0.2%, 0.5%, 1%, 2% or 5% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 0.01% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 0.02% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 0.05% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 0.1% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 0.2% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 0.5% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 1% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 2% of the genome. In some embodiments, the sequence coverage of a genome can be less than about 5% of the genome. In some embodiments, the sequence coverage of a genome can be at least about 5% of the genome. In some embodiments, the sequence coverage of a genome can be at least about 10% of the genome. In some embodiments, the sequence coverage of a genome can be at least about 20% of the genome.
[0187] Multiple sequencing techniques can be used to perform simultaneous sequencing reactions. In some embodiments, the cell-free polynucleotides are sequenced using at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, the cell-free polynucleotides are sequenced using less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. The sequencing reactions are typically performed sequentially or simultaneously. Subsequent data analysis is typically performed on all or a portion of the sequencing reactions. In some embodiments, data analysis is performed on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, data analysis can be performed on less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. Exemplary read depths are about 1000 to about 50,000 reads / locus (base positions) or >50,000 reads / locus.
[0188] In some embodiments, the nucleic acid colony for sequencing is prepared by enzymatically forming a flat end on the double-stranded nucleic acid with a single-stranded overhang at one end or both ends. In these embodiments, in the presence of the nucleotides (for example, A, C, G and T or U) in the form of dNTPs, the colony is usually treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity. The exemplary enzyme optionally used or its catalytic fragment includes Klenow large fragment and T4 polymerase. At the 5' overhang, the enzyme usually extends the 3' end of the depression on the relative chain until it is flush with the 5' end to produce a flat end. At the 3' overhang, the enzyme usually digests from the 3' end, reaches the 5' end of the relative chain and sometimes exceeds the 5' end of the relative chain. If this digestion advances past the 5' end of the relative chain, the room can be filled by having and filling with the enzyme with identical polymerase activity used for the 5' overhang. The formation of the flat end on the double-stranded nucleic acid is conducive to the attachment of, for example, adapters and the subsequent amplification.
[0189] In some embodiments, the nucleic acid population is subjected to additional processing, such as converting single-stranded nucleic acids into double-stranded nucleic acids and / or converting RNA into DNA. These forms of nucleic acids are also optionally ligated with adapters and amplified.
[0190] The nucleic acids subjected to the above-described blunt-end treatment, with or without prior amplification, and optionally other nucleic acids in the sample, can be sequenced to produce sequenced nucleic acids. Sequenced nucleic acids can refer to the sequence (i.e., sequence information) of a nucleic acid or a nucleic acid whose sequence has been determined. Sequencing can be performed to provide sequence data for individual nucleic acid molecules in the sample, directly or indirectly, from the consensus sequence of the amplification products of the individual nucleic acid molecules in the sample.
[0191] In some embodiments, after the double-stranded nucleic acid with single-stranded overhang in the sample is formed into a flat end, it is connected to an adapter containing a barcode at both ends, and sequencing determines the nucleic acid sequence and the inline barcode introduced by the adapter. The flat-ended DNA molecule is optionally flat-ended with an adapter that is at least partially double-stranded (e.g., a Y-shaped adapter or a bell-shaped adapter). Alternatively, the flat ends of the sample nucleic acid and the adapter can be tailed with complementary nucleotides to facilitate connection (e.g., sticky end connection).
[0192] Typically, the nucleic acid sample is contacted with a sufficient number of adapters so that the probability of any two identical nucleic acids receiving the same adapter barcode combination from the adapters connected at the two ends is low (e.g., less than <1% or <0.1%). Using adapters in this way allows the identification of a family of nucleic acid sequences that have the same start and end points on a reference nucleic acid and are connected to the same barcode combination. Such a family represents the sequence of the amplified product of the nucleic acid in the sample before amplification. The sequences of the family members can be compiled to obtain one or more common nucleotides or a complete common sequence of the nucleic acid molecules in the original sample, which are modified by blunt end formation and adapter attachment. In other words, the nucleotides occupying a specific position of the nucleic acid in the sample are determined to be the common nucleotides of the nucleotides occupying the corresponding position in the family member sequence. The family can include the sequence of one or two chains of a double-stranded nucleic acid. If a member of the family includes sequences from two chains of a double-stranded nucleic acid, the sequence of one chain can be converted into their complementary sequences for the purpose of compiling all sequences to obtain one or more common nucleotides or sequences. Some families contain only single member sequences. In this case, the sequence can be regarded as the sequence of the nucleic acid in the sample before amplification. Alternatively, families with only a single member sequence can be excluded from subsequent analysis.
[0193] By comparing the nucleic acid through sequencing with the reference sequence, it is possible to determine the nucleotide variation in the nucleic acid through sequencing. A reference sequence is typically a known sequence, for example, a known whole or part of a genome sequence from a subject (for example, a complete genome sequence of a human subject). A reference sequence can be, for example, hG19 or hG38. As described above, the nucleic acid through sequencing can represent the consensus sequence of the sequence of the nucleic acid in the sample directly determined or the amplification product of such nucleic acid. Comparisons can be made at one or more designated positions on the reference sequence. When the corresponding sequence is aligned to the greatest extent, a subset of the nucleic acid through sequencing can be identified, the subset comprising the position corresponding to the designated position of the reference sequence. In such a subset, it is possible to determine which (if any) nucleic acid through sequencing comprises nucleotide variation at the designated position, and optionally which (if any) comprises reference nucleotides (that is, the same as in the reference sequence). If the number of the nucleic acid through sequencing comprising nucleotide variation in the subset exceeds a selected threshold value, variant nucleotides can be identified at the designated position. The threshold value can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 9 or 10 sequenced nucleic acids in the subset that include the nucleotide variation, or the threshold value can be a ratio of sequenced nucleic acids in the subset that include the nucleotide variation, such as at least 0.5%, 1%, 2%, 3%, 4%, 5%, 10%, 15% or 20%, among other possibilities. Repetitive comparisons can be performed for any specified position of interest in the reference sequence. Sometimes, comparisons can be performed for specified positions that occupy at least about 20, 100, 200 or 300 consecutive positions on the reference sequence, for example, about 20-500 or about 50-300 consecutive positions.
[0194] Additional details about nucleic acid sequencing, including the formats and applications described herein, are also provided in the following literature: for example, Levy et al., Annual Review of Genomics and Human Genetics, 17:95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, 2012, Article ID 251364:1-11 (2012); Voelkerding et al., Clinical Chem., 55:641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7:287-296 (2009); Astier et al., J Am Chem Soc., 128(5):1705-10(2006); U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. No. 7,482,120, U.S. Patent No. 7,501,245, U.S. Patent No. 6,818,395, U.S. Patent No. 6,911,345, U.S. Patent No. 7,501,245, U.S. Patent No. 7,329,492, U.S. Patent No. 7,170,050, U.S. Patent No. 7,302,146, U.S. Patent No. 7,313,308, and U.S. Patent No. 7,476,503, each of which is incorporated herein by reference in its entirety.
[0195] Comparison results
[0196] The instability state of the microsatellite and / or other repetitive nucleic acids of the given subject determined according to the method disclosed in the application is usually compared with the database of the comparison result from the reference population, to identify the customized therapy or targeted therapy for the subject.In some embodiments, the instability state of the microsatellite and / or other repetitive nucleic acids of the test subject and the comparison result are measured across such as whole genome or whole exon group, and in other embodiments, these markers are measured based on a subset or targeted region of such as genome or exon group, which are optionally extrapolated to determine the microsatellite instability of such as whole genome or whole exon group.Usually, reference population includes the patient with the same cancer type as the test subject and / or is receiving or has received the patient of the therapy identical with the test subject.In some embodiments, the instability state of the microsatellite and / or other repetitive nucleic acids of the test subject and the microsatellite and / or other repetitive nucleic acids of the comparison are measured by determining the mutation count or load in the set of predetermined or selected gene or genomic region.Substantially any gene (e.g., oncogene) is optionally selected for such analysis. In some of these embodiments, the selected genes or genomic regions include at least about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1,500, 2,000 or more selected genes or genomic regions. In some of these embodiments, the selected genes or genomic regions optionally include one or more genes listed in Table 1.
[0197] Table 1
[0198]
[0199] Cancer and other diseases
[0200] In certain embodiments, the methods and systems disclosed herein are used to identify customized therapies to treat a given disease, disorder, or condition in a patient. Typically, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, cancer), liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.
[0201] Non-limiting examples of other genetically based diseases, disorders, or conditions that are optionally assessed using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), Cry-a-cat syndrome, Crohn's disease, cystic fibrosis, Dercum disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, Factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson disease, etc.
[0202] Customized therapy and related administration
[0203] In some embodiments, methods disclosed herein relate to identifying customized therapy and administering customized therapy to patients with an unstable state of a given microsatellite and / or other repetitive nucleic acids.Substantially any cancer therapy (e.g., surgical therapy, radiotherapy, chemotherapy and / or similar therapy) is included as a part of these methods. Typically, customized therapy includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method for enhancing the immune response for a given cancer type. In certain embodiments, immunotherapy refers to a method for enhancing the T cell response for a tumor or cancer.
[0204] In some embodiments, immunotherapy or immunotherapeutic agents target immune checkpoint molecules. Certain tumors can evade the immune system by co-opting immune checkpoint pathways. Therefore, targeting immune checkpoints has become an effective method for combating the ability of tumors to evade the immune system and activating anti-tumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012, 12: 252-264.
[0205] In certain embodiments, immune checkpoint molecules are inhibitory molecules that reduce the signals involved in the response of T cells to antigens. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 limits the activity of T cells in peripheral tissues during inflammatory responses. In addition, the ligands of PD-1 (PD-L1 or PD-L2) are usually upregulated on the surface of many different tumors, resulting in downregulation of anti-tumor immune responses in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a ligand of PD-1, such as PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a ligand of CTLA4, such as CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin-like receptor (KIR), T-cell membrane protein 3 (TIM3), galectin 9 (GAL9), or adenosine A2a receptor (A2aR).
[0206] Antagonists targeting these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Therefore, in certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the antagonist of the inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti-CTLA4 antibody, an anti-PD-1 antibody, an anti-PD-L1 antibody, or an anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal anti-PD-1 antibody. In some embodiments, the antibody is a monoclonal anti-PD-L1 antibody. In certain embodiments, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, a combination of an anti-CTLA4 antibody and an anti-PD-L1 antibody, or a combination of an anti-PD-L1 antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is pembrolizumab or nivolumab In certain embodiments, the anti-CTLA4 antibody is ipilimumab. In certain embodiments, the anti-PD-L1 antibody is atezolizumab Avelumab or durvalumab One or more of .
[0207] In certain embodiments, immunotherapy or immunotherapeutic agent is an antagonist (eg, antibody) for CD80, CD86, LAG3, KIR, TIM3, GAL9 or A2aR. In other embodiments, the antagonist is a soluble form of an inhibitory immune checkpoint molecule, such as a soluble fusion protein comprising the extracellular domain of an inhibitory immune checkpoint molecule and the Fc domain of an antibody. In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1 or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9 or A2aR. In one embodiment, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.
[0208] In certain embodiments, immune checkpoint molecules are co-stimulatory molecules that amplify the signals involved in the response of T cells to antigens. For example, CD28 is a co-stimulatory receptor expressed on T cells. When T cells bind to antigens through their T cell receptors, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells to amplify T cell receptor signaling and promote T cell activation. Because CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 can offset or regulate the co-stimulatory signaling mediated by CD28. In certain embodiments, immune checkpoint molecules are co-stimulatory molecules selected from CD28, inducible T cell co-stimulatory factor (ICOS), CD137, OX40 or CD27. In other embodiments, immune checkpoint molecules are ligands of co-stimulatory molecules, including, for example, CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L or CD70.
[0209] Agonists targeting these co-stimulatory checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Therefore, in certain embodiments, immunotherapy or immunotherapeutic agents are agonists of co-stimulatory checkpoint molecules. In certain embodiments, the agonist of co-stimulatory checkpoint molecules is an agonist antibody, and preferably a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS antibody, an anti-CD137 antibody, an anti-OX40 antibody or an anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80 antibody, an anti-CD86 antibody, an anti-B7RP1 antibody, an anti-B7-H3 antibody, an anti-B7-H4 antibody, an anti-CD137L antibody, an anti-OX40L antibody or an anti-CD70 antibody.
[0210] Therapeutic options for treating specific genetically based diseases, disorders or conditions other than cancer are generally well known to those of ordinary skill in the art and will be apparent in view of the particular disease, disorder or condition under consideration.
[0211] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions comprising immunotherapeutics are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutics, etc.) can also be administered by any method known in the art, including, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraaural, which can include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, ointments, ointments, or the like.
[0212] System and computer-readable medium
[0213] The present disclosure also provides various systems and computer program products or machine-readable media. In some embodiments, for example, the methods described herein are optionally performed or facilitated at least in part using systems, distributed computing hardware and applications (e.g., cloud computing servers), electronic communication networks, communication interfaces, computer program products, machine-readable media, electronic storage media, software (e.g., machine-executable code or logic instructions), etc. For illustration, Figure 2A schematic diagram of an exemplary system suitable for implementing at least aspects of the methods disclosed herein is provided. As shown, system 200 includes at least one controller or computer, such as a server 202 (e.g., a search engine server), the server 202 including a processor 204 and a memory, storage device, or memory component 206; and one or more other communication devices 214 and 216 (e.g., a client computer terminal, a phone, a tablet, a laptop, other mobile devices, etc.), the one or more other communication devices 214 and 216 being located remotely from the remote server 202 and communicating with the remote server 202 via an electronic communication network 212 (such as the Internet or other interconnected network). The communication devices 214 and 216 typically include an electronic display (e.g., an Internet-enabled computer or the like) that communicates with the server 202 computer via the network 212, wherein the electronic display includes a user interface (e.g., a graphical user interface (GUI), a web-based user interface, and / or the like) for displaying results after implementing the methods described herein. In certain embodiments, the communication network also includes physical transfer of data from one location to another, for example using a hard drive, thumb drive, or other data storage mechanism. The system 200 also includes a program product 208 stored on a computer or machine-readable medium, such as, for example, one or more memories of various types, such as memory 206 of server 202, which can be read by server 202 to facilitate, for example, directing a search application or other application executed by one or more other communication devices, such as 214 (schematically shown as a desktop computer or personal computer) and 216 (schematically shown as a tablet computer). In some embodiments, the system 200 optionally also includes at least one database server, such as, for example, server 210 associated with an online website, which has data stored thereon that can be searched directly or through the search engine server 202 (e.g., control samples or comparative result data, indexed customized therapies, etc.). The system 200 optionally also includes one or more other servers located remotely from the server 202, each of the other servers optionally associated with one or more database servers 210, which are located remotely from each of the other servers or are located locally with each of the other servers. The other servers can advantageously provide services to geographically remote users and enhance geographically distributed operations.
[0214] As will be understood by those skilled in the art, the memory 206 of the server 202 may optionally include volatile memory and / or non-volatile memory, including, for example, RAM, ROM, and magnetic or optical disks, among others. Those skilled in the art will also appreciate that, although illustrated as a single server, the illustrated configuration of the server 202 is provided by way of example only, and other types of servers or computers configured according to various other methods or architectures may also be used. Figure 2 Schematically shown in the server 202 represents a server or server cluster or server farm, and is not limited to any individual physical server. The server site can be deployed as a server farm or server cluster managed by a server hosting provider. The number of servers and their architecture and configuration can increase based on the use, demand and capacity requirements of the system 200. As those of ordinary skill in the art also understand, other user communication devices 214 and 216 in these embodiments can be, for example, laptop computers, desktop computers, tablet computers, personal digital assistants (PDAs), mobile phones, servers or other types of computers. As known and understood by those of ordinary skill in the art, network 212 can include the Internet, an intranet, a telecommunications network, an extranet or a world wide web of more than one computer / server communicating with one or more other computers through a communication network, and / or a part of a local area network or other regional network.
[0215] As further understood by those skilled in the art, the exemplary program product or machine-readable medium 208 is optionally in the form of microcode, a program, a cloud computing format, a routine, and / or a symbolic language that provides one or more sets of ordered operations that control the functionality of the hardware and direct its operation. According to exemplary embodiments, the program product 208 need not reside entirely in volatile memory, but rather can be selectively loaded as needed according to various methods known and understood by those skilled in the art.
[0216] As further understood by one of ordinary skill in the art, the term "computer-readable medium" or "machine-readable medium" refers to any medium that participates in providing instructions to a processor for execution. For illustration, the term "computer-readable medium" or "machine-readable medium" includes distributed media, cloud computing formats, intermediate storage media, the execution memory of a computer, and any other medium or device capable of storing a program product 208 that implements the functions or methods of the various embodiments of the present disclosure, such as for reading by a computer. "Computer-readable medium" or "machine-readable medium" can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks. Volatile media include dynamic memory, such as the main memory of a given system. Transmission media include coaxial cables, copper wire, and optical fibers, including the wires that make up a bus. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communications. Exemplary forms of computer readable media include floppy disks, flexible disks, hard disks, magnetic tape, flash drives, or any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tape, any other physical media with patterns of holes, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chip or cartridge, a carrier wave, or any other medium from which a computer can read.
[0217] Program product 208 is optionally copied from a computer-readable medium to a hard disk or similar intermediate storage medium. When program product 208 or a portion thereof is to be executed, it is optionally loaded from its distribution medium, its intermediate storage medium, etc., into the execution memory of one or more computers, configuring one or more computers to function according to the functions or methods of various embodiments. All such operations are well known to those skilled in the art of computer systems, for example.
[0218] For further explanation, in certain embodiments, the application provides a system comprising one or more processors and one or more memory components in communication with the processor. The memory component typically includes one or more instructions that, when executed, cause the processor to provide information that causes sequence information, microsatellite and / or other repetitive nucleic acid instability status, comparison results, customized therapy, and / or the like to be displayed (e.g., via communication devices 214, 216, or similar devices) and / or receive information from other system components and / or from a system user (e.g., via communication devices 214, 216, or similar devices).
[0219] In some embodiments, the program product 208 includes non-transitory computer-executable instructions that, when executed by the electronic processor 204, perform at least the following: (i) receiving sequence information from a population of microsatellite loci in a sample; (ii) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on the sequence information to generate a site score for each of the more than one microsatellite loci; (iii) for each of the more than one microsatellite loci, comparing the site score for the given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (iv) generating a site score for the given microsatellite locus when the site score for the given microsatellite locus exceeds a given threshold; (v) classifying the MSI status of the sample as unstable when the microsatellite instability score exceeds a population training threshold for a population of microsatellite loci in the sample to identify an unstable sample; and optionally (vi) comparing the microsatellite instability score for the unstable sample to one or more comparative outcomes, wherein a substantial match between the microsatellite instability score for the unstable sample and the comparative outcome indicates a predicted response of the subject to a therapy.
[0220] The system 200 typically also includes additional system components configured to perform various aspects of the methods described herein. In some of these embodiments, one or more of these additional system components are located remotely from the remote server 202 and communicate with the remote server 202 via an electronic communications network 212, while in other embodiments, one or more of these additional system components are located locally and communicate with the server 202 (i.e., in the absence of an electronic communications network 212) or directly with, for example, a desktop computer 214.
[0221] In some embodiments, for example, additional system components including a sample preparation component 218 are operably connected (directly or indirectly (e.g., via the electronic communication network 212)) to the controller 202. The sample preparation component 218 is configured to prepare nucleic acids in the sample (e.g., prepare a library of nucleic acids) for amplification and / or sequencing by a nucleic acid amplification component (e.g., a thermal cycler, etc.) and / or a nucleic acid sequencer. In certain of these embodiments, the sample preparation component 218 is configured to separate nucleic acids from other components in the sample, attach one or more adapters including barcodes to the nucleic acids as described herein, selectively enrich one or more regions of the genome or transcriptome prior to sequencing, and / or the like.
[0222] In certain embodiments, the system 200 further includes a nucleic acid amplification component 220 (e.g., a thermal cycler, etc.) operably connected (directly or indirectly (e.g., via an electronic communication network 212)) to the controller 202. The nucleic acid amplification component 220 is configured to amplify nucleic acids in a sample from a subject. For example, the nucleic acid amplification component 220 is optionally configured to amplify a selectively enriched region of a genome or transcriptome from a sample as described herein.
[0223] The system 200 typically also includes at least one nucleic acid sequencer 222 operably connected (directly or indirectly (e.g., via an electronic communication network 212)) to the controller 202. The nucleic acid sequencer 222 is configured to provide sequence information of nucleic acids (e.g., amplified nucleic acids) in a sample from a subject. Essentially any type of nucleic acid sequencer can be suitable for use in these systems. For example, the nucleic acid sequencer 222 is optionally configured to perform pyrophosphate sequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, synthesis sequencing, ligation sequencing, hybridization sequencing, or other techniques on nucleic acids to generate sequencing reads. Optionally, the nucleic acid sequencer 222 is configured to group sequence reads into families of sequence reads, each family including sequence reads generated from nucleic acids in a given sample. In some embodiments, the nucleic acid sequencer 222 uses a clonal single molecule array derived from a sequencing library to generate sequencing reads. In certain embodiments, the nucleic acid sequencer 222 includes at least one chip having a micropore array for sequencing the sequencing library to generate sequencing reads.
[0224] To facilitate full or partial system automation, the system 200 typically also includes a material transfer component 224 that is operably connected (directly or indirectly (e.g., via the electronic communication network 212)) to the controller 202. The material transfer component 224 is configured to transfer one or more materials (e.g., nucleic acid samples, amplicons, reagents, and / or the like) to and / or from the nucleic acid sequencer 222, the sample preparation component 218, and the nucleic acid amplification component 220.
[0225] Additional details about computer systems and networks, databases, and computer program products are also provided in, for example, Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Edition (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Edition (2016); Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Edition (2010); Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Edition (2014); Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Edition (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is incorporated herein by reference in its entirety. Example
[0226] Example 1
[0227] Using non-tumor samples as background, MSI high (MSI-H) samples were computationally simulated with variable tumor scores and the number of unstable sites. The distribution observed in a cohort of samples of 3000 different cancer types was used as a priori for the number of unstable sites. The analysis demonstrated a sensitivity of 94% at the limit of detection (LoD) of 0.2% tumor content. The method for determining MSI status according to the embodiments described herein had an expected specificity of 99.999% for non-tumor donor samples. Comparison of these results with standard or conventional PCR-based MSI assessments showed 100% consistency across a tumor content range of 1.4%-15%. In addition, the performance of the analysis was tested on 155 clinical samples from three cancer types for which standard PCR-based MSI status assessments were available (10 MSI-H, 145 microsatellite stable (MSS)). The MSI identification generated according to the embodiments described herein showed 100% consistency with standard PCR-based MSI assessments.
[0228] Example 2
[0229] The MSI status of 82 samples was assessed using a digital sequencing clinical platform (Guardant Health, Inc., Redwood City, CA, USA). The digital sequencing platform is an NGS panel test for cancer-related genes that utilizes high-quality sequencing of cell-free DNA (which may include circulating tumor DNA) isolated from simple, non-invasive blood draws. Digital sequencing uses a digital library of individually tagged cfDNA molecules prepared before sequencing, combined with bioinformatics reconstruction after sequencing to eliminate almost all false positives. Sequence information was obtained using targeted sequencing of cfDNA in the sample. The site scores (ΔAIC) of the 61 most informative microsatellite loci in each sample were determined. The tumor scores of the samples ranged from 0.5% to 15%. The site score for each sample was compared with the corresponding site-specific training threshold to identify the number of unstable microsatellite loci in each sample. The number of unstable microsatellite loci identified in a given sample was used as the microsatellite instability score (i.e., MSI sample score) for that particular sample. The population training thresholds for 61 microsatellite loci in the sample were determined, wherein a microsatellite instability score greater than or equal to 5 was predicted to classify the sample as MSI-high (MSI-H), while a microsatellite instability score less than or equal to 4 was predicted to classify the sample as microsatellite stable (MSS). 9 samples in 82 samples were classified as MSI-H. The remaining 73 samples were classified as MSS. In all 82 samples, the predicted stability status matched the expected stability status. The MSI status was confirmed based on orthogonal validation.
[0230] Example 3
[0231] introduction
[0232] Microsatellite instability (MSI) is a guideline-recommended biomarker with prognostic significance for many tumor types and predictive significance for treatment with immune checkpoint inhibitors. Traditionally, microsatellite instability detection has relied on testing tumor tissue by PCR or immunohistochemistry. Recently, next-generation sequencing (NGS) methods have been developed that also rely on the availability of tumor tissue. In contrast, plasma-based MSI detection methods can provide non-invasive real-time assessments of MSI status. Guardant Health's large panel of cell-free DNA (cfDNA) NGS assays assesses 500 cancer-related genes to identify genomic alterations and tumor mutation burden (TMB). In addition to single nucleotide variants (SNVs), gains and losses, copy number amplification (CNA), fusions, and TMB, this panel can detect microsatellite instability-high (MSI-high) status based on somatic changes in >1,000 MSI sites. The analytical validation presented in this example has four main components used to determine the performance of a 500 cancer-related gene cfDNA NGS assay for MSI-high detection: Accuracy, Limit of Detection (LoD), Precision, and Limit of Blank (LoB).
[0233] method
[0234] Accuracy analysis used 258 samples from 3 sources whose MSI status was predicted by 500 cancer-related gene cfDNA NGS assays and compared with the truth based on tissue MSI status determined by orthogonal methods. 36 collaborator samples with tissue MSI status (true: tissue MSI status), 121 healthy donors (true: microsatellite stable, MSS), and 101 samples sequenced by 500 cancer-related gene cfDNA NGS assays (large panel detection assays) and 73 cancer-related gene cfDNA NGS assays (small panel detection assays, MSI status is true) were used. Reproducibility and repeatability analysis used 2 sets of replicates (56 replicates in total). MSI status and MSI score were compared within and between runs. The LoD of both was obtained by simulation. The LoB was calculated using healthy donor samples and known MSS samples.
[0235] result
[0236] 1. Accuracy Analysis
[0237] Of the 13 MSI-high samples based on tissue MSI status, 12 were identified as MSI-high by the large panel assay. All MSS / MSI-low (MSI-L) samples were correctly detected (Table 2). When restricted to the microsatellite region covered by the small panel assay (~90 loci), the 12 samples detected as MSI-high by the large panel assay also met the small panel assay threshold for identification as MSI-high.
[0238] Table 2
[0239]
[0240] 2.LoD Analysis
[0241] Simulations on >1,000 sites for detecting MSI at 5 tumor fractions spanning the range of 0.05% to 1% indicated a LoD of 0.1% ( Figure 3 ).
[0242] 3. MSI-high detection using cfDNA NGS of 500 cancer-related genes is reproducible and reproducible
[0243] Twenty-four MSI-high replicates were tested in two runs. All replicates were detected as MSI-high using the 500 cancer-related gene cfDNA NGS assay. Within-run and between-run MSI values were scored as ±4 ( Figure 4A ). 10 MSS / MSI low samples had 2-3 replicates tested in the same flow cell (32 replicates in total). All replicates were tested as MSS / MSI low using the 500 cancer-related gene cfDNA NGS assay. The MSI score within each sample was ±3 ( Figure 4B ).
[0244] 4.LoB Analysis
[0245] 121 healthy donors and 25 collaborator samples with known MSS status were used in the LoB analysis. All 146 samples were correctly classified as MSS / MSI-low using the large panel detection assay, showing a 0% false positive rate.
[0246] 5. Tumor score and MSI score
[0247] The MSI scores in >2,000 large panel assay samples were plotted against the maximum mutant allele fraction (MAF) of somatic calls ( Figure 5 A and Figure 5 B), shows that tumor fraction (as measured by MAF) does not correlate with MSI status.
[0248] in conclusion
[0249] MSI high detection using 500 cancer-related gene cfDNA NGS assay showed high sensitivity (>90%) and specificity (100%). The repeatability and reproducibility within and between runs were high. The LoD for MSI high detection was 0.1% MAF. The LoB study showed a false positive rate of 0%. The 500 cancer-related gene cfDNA NGS assay provides a reliable prediction of MSI high status using cfDNA without the need for tissue samples, which will give physicians therapeutic value.
[0250] Example 4
[0251] introduction
[0252] Because of the importance of microsatellite instability (MSI) as a predictive biomarker for response to immune checkpoint blockade (ICB), as exemplified by the pan-cancer approval of pembrolizumab (10, 11), microsatellite instability (MSI) is a biomarker recommended by the National Comprehensive Cancer Network (NCCN) clinical practice guidelines in at least nine cancer types: cervical cancer, bile duct cancer, colorectal cancer, endometrial cancer, esophageal and esophagogastric cancer, gastric cancer, ovarian cancer, pancreatic cancer, and prostate cancer (1-9). Detection of MSI in patients with advanced cancer may also alert clinicians to assess the hereditary cancer risk of the patient's asymptomatic family members.
[0253] MSI is the prototypical manifestation of defective DNA mismatch repair (dMMR), which results in a marked increase in the mutation rate throughout the genome, including gains and / or losses of nucleotides in repetitive motifs called microsatellite tracts, from which the entity derives its name. MSI is most prevalent in endometrial, colorectal, and gastroesophageal cancers, where it can be a sequelae of sporadic mutations in MMR-related genes or a manifestation of Lynch syndrome, an inherited cancer susceptibility syndrome most commonly caused by germline mutations in MLH1, MSH2, MSH6, PMS2, or EPCAM (12). However, despite its increased prevalence in these cancer types, landscape analyses have shown that MSI also occurs at non-negligible rates in most other solid tumors, including common tumor types such as lung, prostate, and breast cancer (13).
[0254] Recent studies have shown that MSI predicts clinical benefit from ICB with PD-1 / PD-L1 inhibitors, which has led to the approval of these agents in several indications (when MSI is present), including nivolumab ± ipilimumab for MSI-high (MSI-H, MSI-positive) metastatic colorectal cancer and pembrolizumab for unresectable or metastatic MSI-H solid tumors after progression on previously approved therapies. In addition to its value as a predictive biomarker for ICB benefit, MSI also has prognostic significance, most notably in colorectal cancer (CRC), where testing is recommended for all patients in clinical practice guidelines (3,15).
[0255] Currently, MSI testing is most commonly performed via polymerase chain reaction (PCR) and / or immunohistochemistry (IHC) analysis of tumor tissue samples. The former assesses five canonical microsatellite loci, originally recommended by the Bethesda panel (16,17), and compares their length in tumor DNA relative to the germline genotype assessed in matched non-tumor DNA; instability in the length of each microsatellite region is used as direct evidence of MSI. However, this limited microsatellite panel was developed primarily for CRC and has more limited sensitivity in other cancer types (18). In contrast, IHC methods assess the levels of four MMR proteins, of which the lack of expression of one or more (defective MMR, dMMR) is strongly associated with MSI status. However, approximately 5% to 11% of MSI-H cases exhibit intact MMR staining and localization (proficient MMR, pMMR) because the antigenicity and intracellular trafficking of otherwise non-functional proteins are retained (19). Recent publications (20,21) have demonstrated that next-generation sequencing (NGS) can also accurately characterize the MSI status of tumors, allowing for comprehensive profiling of targetable genomic biomarkers as well as MSI status via a single NGS test.
[0256] Despite recommendations for many cancer types in the NCCN guidelines and associated FDA-approved treatment options, current rates of MSI testing outside of CRC and gastroesophageal cancer remain very low (22). Even in CRC, where MSI testing recommendations have been available since 2005 (17,23), less than 50% of patients undergo testing (24), leading to missed opportunities for ICB treatment and failure to identify patients whose family members may be at increased risk for cancer. Although multifactorial, such undergenotyping of MSI is often due to barriers related to tissue acquisition and complex testing recommendations / algorithms. For example, testing of archival diagnostic specimens can result in significant delays associated with locating and obtaining this material and inaccurate assessment of MSI status due to tumor evolution and / or heterogeneity. Similarly, testing of newly obtained tissue specimens can also result in significant delays and failures related to the timing of biopsies and, in addition, are associated with the risk and cost of surgical complications. Consequently, invasive tissue acquisition procedures are contraindicated in many heavily pretreated and / or frail patients. Furthermore, the rapidly growing number of biomarkers and the diversity of testing options present daunting complexity for already overburdened physicians.
[0257] Cell-free circulating tumor DNA (ctDNA) assays (“liquid biopsies”) have successfully addressed this barrier in many genotyping indications by enabling minimally invasive profiling of contemporaneous tumor DNA. Thus, by identifying patients whose tumors carry biomarkers of interest that would otherwise remain unidentified due to tissue sampling limitations, liquid biopsies expand patient access to standard-of-care targeted therapies, including ICB, and are more rapid than typical tissue testing (25). Furthermore, comprehensive liquid biopsies can provide information on all guideline-recommended somatic genomic biomarkers for all adult solid tumors in a single test. In this study, we sought to enhance the utility of a previously validated ctDNA-based genotyping test by adding MSI detection. Here, we describe the design and validation of MSI assessment on this platform, report its performance in the largest ctDNA-tissue MSI validation cohort to date (n=1145), and evaluate response prediction in 16 patients with advanced gastric cancer treated with ICB. Also reported here is the MSI-H landscape of more than 28,000 consecutive patients with solid tumors tested in a Clinical Laboratory Improvement Amendments (CLIA)-certified, College of American Pathologists (CAP)-accredited, New York State Department of Health-approved laboratory.
[0258] Materials and methods
[0259] 1. Microsatellite locus selection
[0260] Guardant Health's small panel cell-free DNA (cfDNA) NGS assay is a 74-gene panel that has been previously validated for the detection of SNVs, gains and losses, CNAs, and fusions in all guideline-recommended indications for advanced solid tumors (26,27). The assay initially included 99 putative microsatellite loci consisting of short tandem repeats (STRs) of length 7 or longer, which were selected to include sites susceptible to instability across multiple tumor types, including three of the five Bethesda panel loci (BAT-25, BAT-26, and NR-21). The remaining two Bethesda loci (NR-24 and MONO-27) were not included because the mappability of the regions was extremely low. The coverage and noise profile of these loci were evaluated using sequencing data from a pool of 84 healthy donor samples to exclude uninformative loci from the final MSI detection algorithm.
[0261] 2. Model Description
[0262] MSI detection is based on integrating the observed read sequences with molecular barcode information into a single probabilistic model that compares the likelihood of the observed data under the assumption of PCR and sequencing noise with the likelihood of the data under the assumption of somatic MSI instability. Each individual locus is independently scored using the Akaike Information Criterion (AIC) (28). The AIC model produces a locus score (ranging from 0 to infinity) that reflects the likelihood that the variability observed at any given microsatellite locus is due to biological instability relative to noise, and if the locus score (i.e., site score) is above a site-specific training threshold, the locus is considered unstable. The number of affected loci is calculated across the last 90 loci, and a sample is identified as positive if the number of unstable loci ("MSI score") is above a population training threshold (n=6). The threshold value of individual locus and the total MSI score of each sample are used to establish with the error parameter of the data and individual locus from healthy donor sample and the total number of unstable loci in the simulation sample based on arrangement (permutation), and the frequency of molecules with different repeat lengths in the healthy donor sample is different. By this method, simulation is used here to inquire about 100,000 combinations of microsatellite length and unstable locus number, which allows the different landscapes of evaluation scenario, some of which may not be shown in non-simulated data sets. This algorithm does not distinguish between microsatellite stable (MSS) and MSI-low (a classification defined by observing a single unstable Bethesda locus using a PCR method), and they are grouped into a single classification. This is based on previous reports that the MSI-L state is not an obvious phenotype but an artifact of testing a small amount of microsatellites, so that when testing a large amount of microsatellite loci, the MSI-L sample previously characterized simulates an MSS phenotype in the overall MSI load.
[0263] 3. Sample
[0264] MSI algorithm development and training were performed using simulated data and a collection of 84 healthy donor samples. The clinical validation study included 1145 archived samples (residual plasma and / or cell-free DNA) collected and processed as part of routine care clinical testing standards in the Guardant Health CLIA laboratory as previously described (26), or archived patient plasma samples collected in EDTA blood collection tubes. 20 healthy donor samples were also used for analytical specificity studies. Artificial samples used in the analytical validation study included cfDNA pools extracted from cell line supernatants and healthy donor plasma. Cell-free DNA (ATCC, Inc.) prepared from culture supernatants from the following cell lines was used: KM12, NCI-H660, HCC1419, NCI-H2228, NCI-H1650, NCI-H1648, NCI-H1975, NCI-H1993, NCIH596, HCC78, GM12878, MCF-7. cfDNA isolated from cell line culture supernatants mimics the fragment size and mechanisms of extracellular release (29), library conversion, and sequencing properties of patient-derived cfDNA, while also providing a renewable source of well-defined material in sufficient quantities to support the high material demands of studies, such as detection limits and precision.
[0265] 4. Sample Processing and Bioinformatics Analysis
[0266] Cell-free DNA was extracted from plasma samples or cell line supernatants (QIAmp Circulating Nucleic Acid Kit, Qiagen, Inc.), and up to 30 ng of extracted cfDNA was labeled with nonrandom oligonucleotide barcodes (IDT, Inc.) and then subjected to library preparation, hybridization capture enrichment (Agilent Technologies, Inc.), and sequencing by paired-end synthesis (NextSeq 500 / 550 or HiSeq 2500, Illumina, Inc.) as previously described (26). Bioinformatics analysis and variant detection were performed as previously described (26).
[0267] 5. Analytical Verification Methods
[0268] The studies conducted for analytical validation were based on established CLIA, Nex-StoCT working groups, and the Association of Molecular Pathologists / CAP guidelines for performance characteristics and validation principles. To determine the sensitivity of the assay for MSI status, cfDNA from cell line supernatants of an MSI-H cell line (KM12) (29) was diluted with cfDNA from a microsatellite stable (MSS) cell line (NCI-H660) (30, 31) and tested with both standard (30 ng) cfDNA input and low (5 ng) cfDNA input. For 5 ng input, the dilution series targeted the maximum mutant allele fraction (maximum MAF) of 0.03%-2%, and for 30 ng input, the dilution series targeted the maximum mutant allele fraction (maximum MAF) of 0.01%-1%. Targeted tumor fractions were validated using known germline variations unique to the titrant and diluent materials. Repeatability (intra-run precision) and reproducibility (inter-run precision) were assessed based on clinical and artificial model samples. Six clinical samples (three MSI-H and three MSS) of certain precision were selected with maximum MAF values of 1%-2%, representing -2-3x the predicted LoD at 5 ng.MSI assay specificity was determined by analyzing 20 healthy donor samples and 245 known MSS artificial samples.
[0269] 6. Clinical Validation Methods
[0270] Archived plasma or cfDNA from clinical samples of patients with available results from standard-of-care tissue-based MSI testing (n=1145) were tested using the ctDNA MSI algorithm. Tissue-based MSI status was derived from IHC, PCR, or less commonly, NGS. Clinical outcome data were extracted from patient medical records and deidentified by the treating physician.
[0271] 7. Landscape Analysis of Plasma MSI Status from 28,459 Advanced Cancer Patient Samples
[0272] The cohort included 28,459 consecutive samples from patients with advanced cancer who were tested during their clinical care using a cfDNA NGS assay (a small panel assay) covering 73 cancer-related genes. All analyses were performed using de-identified data and according to an IRB-approved protocol. The prevalence of MSI-H in this cohort was assessed across 16 primary tumor types: bladder cancer, breast cancer, bile duct cancer, colon adenocarcinoma, cancer of unknown primary, head and neck squamous cell carcinoma, hepatocellular carcinoma, lung adenocarcinoma, lung no special type, lung squamous cell carcinoma, "other" cancer diagnosis, pancreatic adenocarcinoma, prostate cancer, gastric adenocarcinoma, and endometrial cancer.
[0273] 8. Statistics
[0274] Statistical analysis was performed using Student's t-test to analyze the number of variances per sample and Fisher's exact test for comparisons of proportions. The lower and upper limits of the 95% confidence intervals (CI) for binomial proportions were calculated using Wilson score intervals with continuity correction.
[0275] 9. Ethics
[0276] This study was conducted under a protocol approved by the Quorum Institutional Review Board using deidentified data.
[0277] result
[0278] 1. MSI algorithm development
[0279] Traditional challenges for ctDNA genotyping using NGS include efficient molecular capture due to low input and low tumor fractions in circulation (26,27) and correction of sequencing and other technical artifacts. MSI detection presents additional challenges because of the need for 1) efficient molecular capture, sequencing, and mapping of repetitive genomic regions that accurately reflect MSI status; 2) error correction and variant detection within repetitive regions; and 3) differentiation of signals due to MSI from those due to non-MSI somatic variants and strong PCR slippage artifacts at loci commonly affected by somatic instability. Indeed, technical PCR errors are often at least an order of magnitude higher than typical sequencing error rates in homopolymer loci, necessitating iterative site selection and optimized use of molecular barcodes to achieve relevant signal-to-noise detection ratios across a large number of candidate microsatellite loci.
[0280] While tissue sequencing panels typically contain sufficient informative microsatellite loci simply due to the large panel size and long DNA fragment length (13,32), the medium-sized ctDNA panel and short cell-free DNA (cfDNA) fragment length used here necessitated purposeful microsatellite selection and inclusion. To achieve this, an iterative approach informed by the literature and tissue sequencing compendia was used to evaluate candidate loci to provide pan-cancer MSI detection with minimal background noise. The list of candidate loci was further refined based on the performance criteria described above using a healthy donor cfDNA reference.
[0281] Based on performance evaluation of training healthy donor samples, informative loci were defined as those that were efficiently captured, sequenced, and mapped and were associated with minor variants in MSS samples ( Figure 6A Uninformative loci were not captured, sequenced, or mapped, leading to inadequate molecular representation (in Figure 6A ), or exhibit significant variation in MSS samples, leading to excessive artifact signals (shown in black in Figure 6A Interestingly, the BAT-25, BAT-26, and NR-21 Bethesda loci, used in traditional MSI tissue testing (16,17) and some ctDNA panels (33), performed poorly relative to other candidate loci and were excluded from the final marker set (in Figure 6A indicated by arrows).
[0282] Using this approach, 90 microsatellite loci were selected for inclusion in the final test format: 89 mononucleotide repeats and a single trinucleotide repeat, all of which included a repeat length of 7 or greater. Evaluation of the unique molecular coverage distribution showed that 65% of these loci had coverage greater than 0.5X the median sample coverage.
[0283] In addition to efficient molecular capture and mapping, MSI detection requires highly accurate discrimination between cancer-associated signals and background noise due to sequencing and polymerase errors at the very low allelic fractions at which ctDNA is typically found (26,27,34). Importantly, the same repetitive genomic background that makes microsatellite candidates informative for MSI detection due to polymerase slippage during cell replication in vivo also makes them particularly susceptible to the same polymerase slippage during in vitro library preparation and sequencing, resulting in high levels of technical noise. To address this issue, digital sequencing error correction is used to define true biological indel events at microsatellite loci with high fidelity, as previously described (26,27). The digital sequencing platform is an NGS panel of cancer-associated genes that utilizes high-quality sequencing of cell-free DNA (which may include circulating tumor DNA) isolated from a simple, non-invasive blood draw. Digital sequencing employs pre-sequencing preparation of digital libraries of individually tagged cfDNA molecules, combined with post-sequencing bioinformatics reconstruction to eliminate virtually all false positives.
[0284] In these high-background error repeats, digital sequencing was associated with a 100-fold reduction in sequencing errors per molecule relative to standard sequencing methods ( Figure 6B ), allowing efficient and accurate reconstruction of the microsatellite sequences of individual unique molecules present in the original patient blood sample. Threshold simulations based on permuted healthy donor samples were then used to establish site-specific MSI status determination thresholds and aggregate sample-level MSI status determination thresholds. When these per-site thresholds and per-sample thresholds were combined with the effect of digital sequencing correction, the false positive rate per sample was estimated to be ~10-7.3. In addition, titration simulations adjusted for the distribution of clinical input predicted robust MSI detection to a tumor fraction of ~0.2%, after which detection efficiency dropped significantly. Therefore, samples with a circulating tumor fraction of <0.2% (as defined by the maximum somatic variant allele fraction) were considered unevaluable for MSI status.
[0285] 2. Analytical Validation Studies
[0286] To assess the analytical sensitivity of MSI detection, cfDNA derived from the supernatant of the MSI-H cell line KM12 was diluted into MSS cfDNA, targeting five titration points, including 15 independently treated replicates, encompassing the limits of detection (LoD) predicted by the computer simulation described above. Each titration series was analyzed with 5 ng (the minimum acceptable cfDNA input) and 30 ng (the maximum and most common cfDNA input). Using probit analysis, the 95% LoD (LOD95) was calculated to be 0.4% ( Figure 7A ), and was calculated to be 0.1% ( Figure 7B).
[0287] To assess analytical intermediate precision, replicates of four different anthropogenic materials (two MSS and two MSI-H) were analyzed ( Figure 7C Across 499 replicates, the classification agreement for MSI status was 100% (499 / 499, 95% CI 99-100%), with the coefficient of variation for the quantitative MSI score for MSI-H samples ranging from 6.3% to 7.2% (Table 3). Repeatability and input robustness were also assessed by repeated testing of MSS and MSI-H artifacts at 5 ng, 10 ng, and 30 ng cfDNA input, which similarly demonstrated 100% agreement (27 / 27, 95% CI 85-100%, Figure 8 Clinical accuracy was confirmed in 72 independent patient sample replicates, representing a range of MSI scores and tumor scores across three independent batches, days, operators, and reagent batches, demonstrating 100% qualitative agreement (72 / 72, 95% CI 94-100%), with coefficients of variation for the basic quantitative MSI score ranging from 2.0 to 15.2% (Table 5).
[0288] Table 3
[0289]
[0290] Table 4
[0291] Input average value Standard Deviation %cv 5ng 28.1 2.1 7.6 10ng 30.6 5.1 16.5 30ng 32.0 3.0 9.2
[0292] Table 5
[0293]
[0294] To assess analytical specificity, healthy donor plasma samples (different from those used for training), MSS artifacts, and MSS patient samples were analyzed for pseudo-MSI-H identification. Analytical specificity was 100% across healthy donor samples (20 / 20, 95% CI 83-100%), artifacts (245 / 245, 95% CI 98-100%), and patient samples (48 / 48, 95% CI 92-100%).
[0295] 3. Clinical validation studies
[0296] Because no orthogonal cfDNA-based methods were available as comparators, clinical accuracy was determined by comparing ctDNA MSI assessments with MSI status from medical records determined using standard of care tissue testing (a mix of IHC, PCR, and NGS methods) on 1,145 samples encompassing 40 different cancer types, with at least five representative samples for 15 cancer types ( Figure 9 Among 949 unique evaluable patients, ctDNA detected 87% of patients as MSI-H (71 / 82, 95% CI 77-93%) and 99.5% of patients as MSS / MSI-L (863 / 867, 95% CI 98.7-99.8%), with an overall accuracy of 98.4% (934 / 949, 95% CI 97.3-99.1%) and a positive predictive value (PPV) of 95% (71 / 75, 95% CI 86-98%) ( Figure 10C , Table 8). Consistent with the in silico modeling studies, MSI-H detection was rare (0 / 19) in samples classified as unevaluable due to low tumor fraction ( Figure 10A ), which explains 57% (16 / 28) of the observed ctDNA-tissue discordance in the total unique patient sample set (Table 6-Table 9). For samples with a tumor fraction higher than 1%, ctDNA PPA rose to 93% (54 / 58, 95% CI 82-98%, Table 9).
[0297] Table 6
[0298] A. All samples, regardless of maximum VAF
[0299]
[0300] Table 7
[0301] B. Excluding undetected tumors
[0302]
[0303] Table 8
[0304] C. Maximum VAF ≥ 0.2%
[0305]
[0306] Table 9
[0307] D. Maximum VAF ≥ 1%
[0308]
[0309] Interestingly, despite the high correlation between IHC and PCR tissue testing reported in the literature (23, 35), here, the concordance between ctDNA and tissue MSI status varied by tissue testing method (97.4% (450 / 462) for PCR, 98.0% (239 / 244) for NGS, and 83.0% (93 / 112) for IHC, Figure 10B And Table 10 and Table 11). In further investigation, it was noted that this discordance was due to an increased tissue IHC-positive ctDNA-negative population (2.4% for PCR, 2.0% for NGS, and 12.5% for IHC, Fisher's exact test p < 0.001 for IHC-PCR and IHC-NGS) and an increased tissue IHC-negative ctDNA-positive population (0.2% for PCR, 0% for NGS, and 4.5% for IHC, Fisher's exact test p < 0.01 for both comparisons). These differences prompted us to investigate whether IHC limitations might lead to the observed IHC-ctDNA discordance rather than impaired ctDNA accuracy. Of the 25 samples for which IHC and another tissue test result were available, 12 samples exhibited IHC-ctDNA discordance. Importantly, in 5 of the 12 discordances, PCR and / or NGS tissue testing supported the ctDNA NGS results, rather than tissue IHC. Together, these data support previous reports that IHC may be less reliable than the PCR diagnostic prototype for MSI determination (36).
[0310] Table 10
[0311] A. All samples
[0312] organize cfDNA PCR NGS IHC Negative Negative 408 226 78 Positive Positive 42 13 15 Positive Negative 11 5 14 Negative Positive 1 0 5 total 462 244 112
[0313] Table 11
[0314] B. Evaluable
[0315] organize cfDNA PCR NGS IHC Negative Negative 353 179 64 Positive Positive 42 13 15 Positive Negative 5 2 5 Negative Positive 1 0 5 total 401 194 89
[0316] 4. ctDNA MSI status in 28,459 consecutive patients with advanced cancer
[0317] While many studies have assessed the prevalence of MSI in tissues across different tumor types (13, 32, 37), to date, there has been no published landscape analysis of ctDNA MSI status across cancer types. To this end, the MSI algorithm described above was applied to clinical samples from 28,459 consecutive patients with advanced cancer tested in Guardant Health clinical laboratories. In this cohort, 278 samples (median tumor score 6.55%, range 0.09%-89%) across 16 different tumor types were identified as MSI-H by ctDNA, which corresponds to an overall pan-cancer prevalence of ~1%, similar to that previously reported for tissues (13, 32, 37). Similarly, the prevalence of MSI-H across tumor types also closely mirrored that observed in tissue-based analyses ( Figure 11A ); As expected, MSI-H was most prevalent in endometrial, colorectal, and gastric cancers, while other tumors such as lung, bladder, and head and neck cancers showed lower prevalence. Specific exceptions to previous MSI-H prevalence estimates include slightly lower prevalence in endometrial, colorectal, and gastric cancers, and slightly higher prevalence in prostate cancer.
[0318] Given the pan-solid tumor nature of the ctDNA intended use population and the availability of immunotherapies approved for MSI-H tumors, this panel of microsatellite loci was intentionally selected to be informative about MSI status across all solid tumor types. Figure 1 Consistent with this, we analyzed the distribution of MSI scores at the sample level and locus level—that is, MSI scores and locus scores, respectively ( Figure 11B and Figure 11C ) showed consistent performance across tumor types, with MSI-H samples showing a signal significantly above the threshold. Furthermore, the diagnostic yield of MSI assessment was substantial beyond tumor types commonly tested for MSI; more than half of the identified cases (143 / 278) occurred in tumor types for which MSI testing was very rare, and therefore the identified patients may have never been tested.
[0319] Consistent with what has been reported in tissues (38), the number of gain / loss sites and SNVs (including both nonsynonymous and synonymous variants) was significantly increased in MSI-H samples relative to those characterized as having MSS status (Figure 12). Specifically, the median number of SNVs was 6.3 in MSI-H samples versus 1.4 in MSS (chi-square p < 0.0001), and the median number of gain / loss sites was 2.6 in MSI-H samples versus 0.4 in MSS (chi-square p < 0.0001).
[0320] 5. ctDNA MSI status predicts immunotherapy response
[0321] The most significant utility of MSI status today is its ability to select patients for immunotherapy. Despite this and the barriers to obtaining tissue in many patients, the ability of ctDNA MSI status to predict immunotherapy response has not been reported. To establish the clinical validity of this biomarker, we present the clinical outcomes of 16 patients with ctDNA MSI-H metastatic gastric cancer treated with pembrolizumab (n=15) or nivolumab (n=1) after failure of standard of care chemotherapy in a Phase II pembrolizumab trial in gastric cancer (NCT#02589496). cfDNA and tissue PCR MSI assessments performed in pre-treatment samples were 100% concordant for MSI-H (16 / 16, 95% CI 76-100%). Ten of the 16 patients achieved an investigator-assessed complete (n=3) objective response or an investigator-assessed partial (n=7) objective response per RECIST 1.1 criteria, with an additional 3 patients having stable disease ( Figure 13A ), the objective response rate was 63% (10 / 16, 95% CI 36-84%), and the disease control rate was 81% (13 / 16, 95% CI 54-95%), similar to responses previously reported for patients with MSI-H defined by tissue testing (39). Importantly, even in this pre-treated population, these responses were durable, with a mean treatment duration of 39 weeks. In fact, for example, patient 21 experienced complete regression of disease after treatment with pembrolizumab after failure of fluoropyrimidine / platinum chemotherapy and remains disease-free more than 6 months after completing 35 cycles of therapy ( Figures 13B-13E ).
[0322] discuss
[0323] A novel targeted NGS method based on cfDNA was validated for MSI detection—by using a large microsatellite panel, this method achieved high sensitivity relative to tissue-based methods while maintaining very high specificity. The prevalence of MSI-H across 16 common solid tumors detected by plasma was similar to that of published tissue-based compilations, demonstrating the pan-tumor performance expected from the design of the MSI detection algorithm. Furthermore, clinical utility was demonstrated by showing that patients with MSI-H as detected by cfDNA benefited from ICB therapy in a manner similar to that reported for tissue-defined cohorts (39), extending the availability of MSI testing to all patients regardless of tissue availability or the need to undergo invasive tissue acquisition procedures.
[0324] This example demonstrates the robust analytical performance of MSI detection on a ctDNA panel that was previously validated for detection of four other variant types across all guideline-recommended indications (26). In particular, the analytical sensitivity of MSI detection in human samples demonstrated reproducible detection up to 0.1%, consistent with previous reports of similar sensitivity for gain / loss and SNVs (26). Importantly, this example evaluated the performance of the ctDNA MSI test in 1145 samples with orthogonal tissue MSI, which constitutes the largest ctDNA-tissue MSI concordance cohort yet described. Relative to standard of care tissue MSI testing for the same patients, ctDNA MSI assessment demonstrated a high PPV (95%), which is comparable to the reported PPV of 90%-92% reported for local tissue-based versus central tissue-based MSI assessment (36), and demonstrated a high PPA (87%) in the evaluable population, consistent with previous studies examining the concordance of plasma and tissue genotyping for other variant types (25, 26, 45, 46). Factors that may contribute to incomplete concordance may include tumor heterogeneity, differential shedding of primary versus metastatic lesions, inconsistencies in the timing of tissue and plasma collection, and low tumor shedding in some tumors (40,44,47-49). Interestingly, the gastric cancer patients identified as MSI-H by plasma and quintuplex PCR in this report were previously reported to represent a discrete tumor population that included both MSS and MSI-H disease, as assessed by both IHC and PCR performed on tissue (40). The same study found 9% discordance for MSI-H between paired tissue biopsies from the same patient (40), highlighting the potential contribution of intratumor heterogeneity to discordance in MSI status. Furthermore, the meaningful discordance between PCR and IHC tissue methods observed in this report highlights the importance of accurate MSI testing, which has been reported as a major source of ICB failure (36). Consistent with the challenges presented by tissue genotyping in advanced solid tumors, a study of metastatic non-small cell lung cancer (NSCLC) has shown that, relative to tissue, plasma-based testing increased the number of patients with successful tumor genotyping results and the frequency of detecting targetable mutations, while returning results at least one week faster than typical tissue genotyping results (25,45).
[0325] This example presents the first ctDNA-based landscape analysis of MSI in a large, advanced pan-cancer cohort. Overall, in a collection of >28,000 consecutive clinical samples, the relative prevalence across tumor types was consistent with that reported for tissue (13, 32, 37), with only minor differences. For example, the prevalence in CRC and endometrial cancer was lower than that reported for tissue (13), which likely reflects the fact that the tissue-based landscape analysis included a large number of early-stage MSI-H tumors, which have a better prognosis (15) and are unlikely to be part of the advanced cancer population tested with ctDNA. On the other hand, the higher-than-expected prevalence of MSI-H prostate cancer may be attributable to the increased expression of MSI-H disease in advanced patients; two recent studies focusing on MSI status in advanced prostate cancer have shown MSI-H prevalence of 3.1% and 3.8% in this patient population (50, 51), which is similar to the 2.6% observed in this study. Not surprisingly, given the design intent of the pan-cancer MSI assay, landscape analysis did not reveal tumor type-specific patterns of microsatellite instability. However, this does not exclude the possibility that in plasma, similar to what has been shown in tissues (37,52), tumor type-specific patterns may emerge following evaluation of a large number of microsatellite loci and a large number of representative MSI-H samples.
[0326] The clinical results reported here are limited to gastric cancer; however, the observed objective response rates are consistent with expectations from tissue-based studies, suggesting that ICB therapy based on cfDNA MSI results should achieve the expected results across solid tumor types. In addition, the lack of germline dMMR data hinders conclusions regarding the familial significance of MSI detected by cfDNA. Finally, treatment data for the majority of patients with cfDNA- / tissue+ discordance are unavailable, but it is expected that at least some patients have received ICB therapy based on tumor results, which may have suppressed MSI-H disease and contributed to the lack of MSI detection by cfDNA. Therefore, the 87% sensitivity of MSI-H detection may be higher in patients who are untreated or have not received therapy, relative to tissue. Future studies should be directed to address these issues.
[0327] In summary, a cfDNA-based targeted NGS panel has been developed and validated that accurately assesses MSI status while also providing comprehensive tumor genotyping, allowing pan-solid tumor guide-testing to be completed from a single peripheral blood draw with high sensitivity, specificity, and accuracy. Clinical validation using both comparisons with tissue testing, population-level prevalence analyses, and the first reported results of cfDNA MSI-H patients treated with ICB therapy support the clinical accuracy and relevance of this approach. Such simultaneous characterization of MSI status and tumor genotype from a simple peripheral blood draw has the potential to expand both targeted therapies and immunotherapies to all patients with advanced cancer, including those for whom current tissue-based testing modalities are inadequate.
[0328] References
[0329] 1. Koh WJ, Abu-Rustum NR, Bean S, Bradley K, Campos SM, Cho KR, et al. Cervical Cancer, Version 3.2019, NCCN Clinical Practice Guidelines inOncology. J Natl Compr Canc Netw. 2019; 17:64–84.
[0330] 2. Benson, Al B., D'Angelica, Michael I., Abbott, Daniel E., Abrams, Thomas A., Alberts, Steven R., Anaya, Daniel A., et al. Hepatobiliary Cancers, Version 4. 2018: Featured Updates to the NCCN Guidelines. National Comprehensive Cancer Network Clinical Practice Guidelines in Oncology[Internet].2018;Availablefrom:https: / / www.nccn.org / professionals / physician_gls / pdf / hepatobiliary.pdf.
[0331] 3.Benson,Al B.,Venook,Alan P.,Bekaii-Saab,Tanios,Chan,Emily,Chen,Yi-Jen,Cooper,Harry S.,et al.Colon Cancer Version 1.2016,NCCN PracticeGuidelines in Oncology.2015;Available from:https: / / www.nccn.org / professionals / physician_gls / pdf / colorectal.pdf.
[0332] 4.Koh W-J,Abu-Rustum NR,Bean S,Bradley K,Campos SM,Cho KR,etal.Uterine Neoplasms,Version 1.2018,NCCN Clinical Practice Guidelines inOncology.J Natl Compr Canc Netw.2018;16:170–99.
[0333] 5.Ajani,Jaffer A.,D’Amico,Thomas A.,Baggstrom,Maria,Bentrem,David J.,Chao,Joseph,Das,Prajnan,et al.Esophageal and Esophagogastric Junction CancersVersion 4.2017.2017;Available from:https: / / www.nccn.org / professionals / physician_gls / pdf / esophageal.pdf.
[0334] 6.Ajani,Jaffer A.,D’Amico,Thomas A.,Baggstrom,Maria,Bentrem,David J.,Chao,Joseph,Das,Prajnan,et al.Gastric Cancer,Version 5.2017,NCCN ClinicalPractice Guidelines in Oncology.2017;Available from:https: / / www.nccn.org / professionals / physician_gls / pdf / gastric.pdf.
[0335] 7.Armstrong,Deborah K.,Alvarez,Ronald D.,Bakkum-Gamez,Jamie N.,Barroilhet,Lisa,Behbakht,Kian,Berchuck,Andrew,et al.NCCN Guidelines Version1.2019 Ovarian Cancer.2019;Available from:https: / / www.nccn.org / professionals / physician_gls / pdf / ovarian.pdf.
[0336] 8.Tempero,Margaret A.,Malafa,Mokenge P.,Al-Hawary,Mahmoud,Asbun,Horacio,Bain,Andrew,Behrman,Stephen W.,et al.Pancreatic AdenocarcinomaVersion 2.2018,NCCN Clinical Practice Guidelines in Oncology.2018;Availablefrom:httpshttps: / / www.nccn.org / professionals / physician_gls / pdf / pancreatic.pdf.
[0337] 9.Mohler JL,Lee RJ,Antonarakis ES,Armstrong AJ,D’Amico AV,Davis BJ,etal.NCCN Guidelines Version 1.2018 Prostate Cancer[Internet].NCCN ClinicalPractice Guidelines in Oncology(NCCN Guidelines).2018.Available from:https: / / www.nccn.org.
[0338] 10.Diaz LA,Le DT.PD-1 Blockade in Tumors with Mismatch-RepairDeficiency.N Engl J Med.2015;373:1979.
[0339] 11.Le DT,Durham JN,Smith KN,Wang H,Bartlett BR,Aulakh LK,etal.Mismatch-repair deficiency predicts response of solid tumors to PD-1blockade.Science.2017.
[0340] 12.Buza N,Ziai J,Hui P.Mismatch repair deficiency testing in clinicalpractice.Expert Rev Mol Diagn.2016;16:591–604.
[0341] 13.Bonneville R,Krook MA,Kautto EA,Miya J,Wing MR,Chen H-Z,etal.Landscape of Microsatellite Instability Across 39 Cancer Types.JCO PrecisOncol.2017.
[0342] 14.Marcus L,Lemery SJ,Keegan P,Pazdur R.FDA Approval Summary:Pembrolizumab for the treatment of microsatellite instability-high solidtumors.Clin Cancer Res.2019.
[0343] 15.Benson AB,Arnoletti JP,Bekaii-Saab T,Chan E,Chen Y-J,Choti MA,etal.Colon cancer.J Natl Compr Canc Netw.2011;9:1238–90.
[0344] 16.Boland CR,Thibodeau SN,Hamilton SR,Sidransky D,Eshleman JR,BurtRW,et al.A National Cancer Institute Workshop on Microsatellite Instabilityfor cancer detection and familial predisposition:development of internationalcriteria for the determination of microsatellite instability in colorectalcancer.Cancer Res.1998;58:5248–57.
[0345] 17.Umar A,Boland CR,Terdiman JP,Syngal S,de la Chapelle A,Rüschoff J,et al.Revised Bethesda Guidelines for hereditary nonpolyposis colorectalcancer(Lynch syndrome)and microsatellite instability.J Natl Cancer Inst.2004;96:261–8.
[0346] 18.Lu Y,Soong TD,Elemento O.A novel approach for characterizingmicrosatellite instability in cancer cells.PLoS ONE.2013;8:e63056.
[0347] 19.Dudley JC,Lin M-T,Le DT,Eshleman JR.Microsatellite Instability asa Biomarker for PD-1 Blockade.Clin Cancer Res.2016;22:813–20.
[0348] 20.Salipante SJ,Scroggins SM,Hampel HL,Turner EH,PritchardCC.Microsatellite instability detection by next generation sequencing.ClinChem.2014;60:1192–9.
[0349] 21.Latham A,Srinivasan P,Kemel Y,Shia J,Bandlamudi C,Mandelker D,etal.Microsatellite Instability Is Associated With the Presence of LynchSyndrome Pan-Cancer.J Clin Oncol.2019;37:286–95.
[0350] 22.Guardant Health,Inc.Tissue Findings Submitted with Guardant360Test Requisitions–Data on File.Redwood City,California;2019.
[0351] 23.Hampel H,Frankel WL,Martin E,Arnold M,Khanduja K,Kuebler P,etal.Screening for the Lynch syndrome(hereditary nonpolyposis colorectalcancer).N Engl J Med.2005;352:1851–60.
[0352] 24.Shaikh T,Handorf EA,Meyer JE,Hall MJ,Esnaola NF.Mismatch RepairDeficiency Testing in Patients With Colorectal Cancer and Nonadherence toTesting Guidelines in Young Adults.JAMA Oncol.2017;e173580.
[0353] 25.Leighl NB,Page RD,Raymond VM,Daniel DB,Divers SG,Reckamp KL,etal.Clinical Utility of Comprehensive Cell-Free DNA Analysis to IdentifyGenomic Biomarkers in Patients with Newly Diagnosed Metastatic Non-Small CellLung Cancer.Clinical Cancer Research.2019;clincanres.0624.2019.
[0354] 26.Odegaard JI,Vincent JJ,Mortimer S,Vowles JV,Ulrich BC,Banks KC,etal.Validation of a Plasma-Based Comprehensive Cancer Genotyping AssayUtilizing Orthogonal Tissue-and Plasma-Based Methodologies.Clin CancerRes.2018;24:3539–49.
[0355] 27.Lanman RB,Mortimer SA,Zill OA,Sebisanovic D,Lopez R,Blau S,etal.Analytical and Clinical Validation of a Digital Sequencing Panel forQuantitative,Highly Accurate Evaluation of Cell-Free Circulating TumorDNA.PLoS ONE.2015;10:e0140712.
[0356] 28.Akaike H.Information Theory and an Extension of the MaximumLikelihood Principle.In:Petrov BN,Csaki F,editors.Proceedings of the 2ndInternational Symposium on Information Theory(pp 267-281)Budapest:AkademiaiKiado.1973.
[0357] 29.Berg KCG,Eide PW,Eilertsen IA,Johannessen B,Bruun J,Danielsen SA,et al.Multi-omics of 34 colorectal cancer cell lines-a resource forbiomedical studies.Mol Cancer.2017;16:116.
[0358] 30.Cosmic.COSMIC-Catalogue of Somatic Mutations in Cancer[Internet].[cited 2019 Apr 17].Available from:https: / / cancer.sanger.ac.uk / cosmic.
[0359] 31.Forbes SA,Beare D,Boutselakis H,Bamford S,Bindal N,Tate J,etal.COSMIC:somatic cancer genetics at high-resolution.Nucleic AcidsResearch.2017;45:D777–83.
[0360] 32.Hause RJ,Pritchard CC,Shendure J,Salipante SJ.Classification andcharacterization of microsatellite instability across 18 cancer types.NatMed.2016;22:1342–50.
[0361] 33.Georgiadis A,Wood D,Murphy D,Parpart-Li S,Riley D,Sengamalay N,etal.Abstract 1286:Analytical validation of an integrated next-generationsequencing pan-cancer liquid biopsy approach for detection of microsatelliteinstability.Cancer Res.2018;78:1286.
[0362] 34.Zill OA,Banks KC,Fairclough SR,Mortimer SA,Vowles JV,Mokhtari R,etal.The Landscape of Actionable Genomic Alterations in Cell-Free CirculatingTumor DNA from 21,807 Advanced Cancer Patients.Clin Cancer Res.2018;24:3528–38.
[0363] 35.Lindor NM,Burgart LJ,Leontovich O,Goldberg RM,Cunningham JM,Sargent DJ,et al.Immunohistochemistry versus microsatellite instabilitytesting in phenotyping colorectal tumors.J Clin Oncol.2002;20:1043–8.
[0364] 36.Cohen R,Hain E,Buhard O,Guilloux A,Bardier A,Kaci R,etal.Association of Primary Resistance to Immune Checkpoint Inhibitors inMetastatic Colorectal Cancer With Misdiagnosis of Microsatellite Instabilityor Mismatch Repair Deficiency Status.JAMA Oncology.2019;5:551.
[0365] 37.Vanderwalde A,Spetzler D,Xiao N,Gatalica Z,MarshallJ.Microsatellite instability status determined by next-generation sequencingand compared with PD-L1 and tumor mutational burden in 11,348 patients.CancerMedicine.2018;7:746–56.
[0366] 38.Bonneville R,Krook MA,Kautto EA,Miya J,Wing MR,Chen H-Z,etal.Landscape of Microsatellite Instability Across 39 Cancer Types.JCO PrecisOncol.2017;2017.
[0367] 39.Le DT,Uram JN,Wang H,Bartlett BR,Kemberling H,Eyring AD,et al.PD-1Blockade in Tumors with Mismatch-Repair Deficiency.New England Journal ofMedicine.2015;372:2509–20.
[0368] 40.Kim ST,Cristescu R,Bass AJ,Kim K-M,Odegaard JI,Kim K,etal.Comprehensive molecular characterization of clinical responses to PD-1inhibition in metastatic gastric cancer.Nat Med.2018;24:1449–58.
[0369] 41.Accordino MK,Wright JD,Buono D,Neugut AI,Hershman DL.Trends in useand safety of image-guided transthoracic needle biopsies in patients withcancer.J Oncol Pract.2015;11:e351-359.
[0370] 42.Lokhandwala T,Bittoni MA,Dann RA,D’Souza AO,Johnson M,Nagy RJ,etal.Costs of Diagnostic Assessment for Lung Cancer:A Medicare ClaimsAnalysis.Clin Lung Cancer.2017;18:e27–34.
[0371] 43.De Mattos-Arruda L,Weigelt B,Cortes J,Won HH,Ng CKY,Nuciforo P,etal.Capturing intratumor genetic heterogeneity by de novo mutation profilingof circulating cell-free tumor DNA:a proof-of-principle.Ann Oncol.2014;25:1729–35.
[0372] 44.Goyal L,Saha SK,Liu LY,Siravegna G,Leshchiner I,Ahronian LG,etal.Polyclonal Secondary FGFR2 Mutations Drive Acquired Resistance to FGFRInhibition in Patients with FGFR2 Fusion-Positive Cholangiocarcinoma.CancerDiscov.2017;7:1–12.
[0373] 45.Aggarwal C,Thompson JC,Black TA,Katz SI,Fan R,Yee SS,etal.Clinical Implications of Plasma-Based Genotyping With the Delivery ofPersonalized Therapy in Metastatic Non-Small Cell Lung Cancer.JAMAOncol.2018.
[0374] 46.Siravegna G,Lazzari L,Crisafulli G,Sartore-Bianchi A,Mussolin B,Cassingena A,et al.Radiologic and Genomic Evolution of Individual Metastasesduring HER2 Blockade in Colorectal Cancer.Cancer Cell.2018;34:148-162.e7.
[0375] 47.Pectasides E,Stachler MD,Derks S,Liu Y,Maron S,Islam M,etal.Genomic Heterogeneity as a Barrier to Precision Medicine inGastroesophageal Adenocarcinoma.Cancer Discov.2018;8:37–48.
[0376] 48.Thompson JC,Yee SS,Troxel AB,Savitch SL,Fan R,Balli D,etal.Detection of Therapeutically Targetable Driver and Resistance Mutations inLung Cancer Patients by Next-Generation Sequencing of Cell-Free CirculatingTumor DNA.Clin Cancer Res.2016;22:5772–82.
[0377] 49.Sacher AG,Komatsubara KM,Oxnard GR.Application of plasmagenotyping technologies in non-small cell lung cancer:a practical review.JThorac Oncol.2017.
[0378] 50.Mayrhofer M,De Laere B,Whitington T,Van Oyen P,Ghysel C,Ampe J,etal.Cell-free DNA profiling of metastatic prostate cancer revealsmicrosatellite instability,structural rearrangements and clonalhematopoiesis.Genome Med.2018;10:85.
[0379] 51.Abida W,Armenia J,Gopalan A,Brennan R,Walsh M,Barron D,etal.Prospective Genomic Profiling of Prostate Cancer Across Disease StatesReveals Germline and Somatic Alterations That May Affect Clinical DecisionMaking.JCO Precision Oncology.2017;1–16.
[0380] 52.Hause RJ,Pritchard CC,Shendure J,Salipante SJ.Classification and characterization of microsatellite instability across 18cancer types.NatureMedicine.2016;22:1342–50.
[0381] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent to those skilled in the art upon reading this disclosure that various changes in form and detail may be made without departing from the true scope of the disclosure and that the present invention may be practiced within the scope of the appended claims. For example, all methods, systems, computer-readable media, and / or component features, steps, elements, or other aspects may be used in various combinations.
[0382] All patents, patent applications, websites, other publications or documents, accession numbers, etc. cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be so incorporated. If different versions of a sequence are associated with an accession number at different times, this refers to the version associated with the accession number on the effective filing date of the present application. If applicable, the effective filing date refers to the earlier of the actual filing date or the filing date of the priority application that mentions the accession number. Similarly, if different versions of a publication, website, etc. are published at different times, this refers to the version that was most recently published on the effective filing date of the present application, unless otherwise indicated.
Claims
1. A computer-implemented method for determining a repetitive nucleic acid instability state of a nucleic acid sample, the method comprising: (a) quantifying the number of different repeat sequence lengths present at each of more than one repetitive nucleic acid loci based on sequence information to generate a site score for each of the more than one repetitive nucleic acid loci, wherein the sequence information is from a population of repetitive nucleic acid loci in the nucleic acid sample, wherein the site scores for the more than one repetitive nucleic acid loci include a likelihood score, wherein the site scores for the more than one repetitive nucleic acid loci include a site score based on the Akaike information criterion (AIC), wherein the site score based on the Akaike information criterion (AIC) tests for the presence of somatic gain / loss positions at the more than one repetitive nucleic acid loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, where the method includes: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given repetitive nucleic acid locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs comprising the observed repeat sequence length of sequencing reads covering said given repetitive nucleic acid locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (b) identifying a given repetitive nucleic acid locus as unstable when the site score of the given repetitive nucleic acid locus exceeds a site-specific training threshold for the given repetitive nucleic acid locus to generate a repetitive nucleic acid instability score comprising the number of unstable repetitive nucleic acid loci from the more than one repetitive nucleic acid loci; and (c) classifying the repetitive nucleic acid instability state of the nucleic acid sample as unstable when the repetitive nucleic acid instability score exceeds a population training threshold for the population of the repetitive nucleic acid loci in the nucleic acid sample, thereby determining the repetitive nucleic acid instability state of the nucleic acid sample.
2. A computer-implemented method for determining the repetitive DNA instability state of a sample, the method comprising: (a) quantifying the number of different repeat sequence lengths present at each of more than one repetitive DNA loci based on sequence information to generate a site score for each of the more than one repetitive DNA loci, wherein the sequence information is from a population of repetitive DNA loci in the sample, wherein the site scores for the more than one repetitive DNA loci include a likelihood score, wherein the site scores for the more than one repetitive DNA loci include a site score based on the Akaike information criterion (AIC), wherein the site score based on the Akaike information criterion (AIC) tests for the presence of somatic gain / loss positions at the more than one repetitive DNA loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, where the method includes: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given repetitive DNA locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs comprising the observed repeat length of sequencing reads covering said given repetitive DNA locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (b) for each of the more than one repetitive DNA loci, comparing a site score for a given repetitive DNA locus to a site-specific training threshold for the given repetitive DNA locus; (c) identifying the given repetitive DNA locus as unstable when the site score of the given repetitive DNA locus exceeds a site-specific training threshold for the given repetitive DNA locus to generate a repetitive DNA instability score, the repetitive DNA instability score comprising the number of unstable repetitive DNA loci from the more than one repetitive DNA loci; and (d) classifying the repetitive DNA instability state of the sample as unstable when the repetitive DNA instability score exceeds a population training threshold for the population of the repetitive DNA loci in the sample, thereby determining the repetitive DNA instability state of the sample.
3. A computer-implemented method for determining the microsatellite instability (MSI) status of a sample, the method comprising: (a) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on sequence information to generate a site score for each of the more than one microsatellite loci, wherein the sequence information is from a population of microsatellite loci in the sample, wherein the site scores for the more than one microsatellite loci include a likelihood score, wherein the site scores for the more than one microsatellite loci include an Akaike information criterion (AIC)-based site score, wherein the Akaike information criterion (AIC)-based site score tests for the presence of somatic gain / loss at the more than one microsatellite loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, where the method includes: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs including the observed repeat length of sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (b) for each of the more than one microsatellite loci, comparing the site score for a given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (c) identifying the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds a site-specific training threshold for the given microsatellite locus to generate a microsatellite instability score, the microsatellite instability score comprising the number of unstable microsatellite loci from the more than one microsatellite loci; and (d) classifying the MSI status of the sample as unstable when the microsatellite instability score exceeds a population training threshold for the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.
4. A computer-implemented method for determining the microsatellite instability (MSI) status of a sample, the method comprising: (a) receiving sequence information from a population of microsatellite loci in the sample; (b) quantifying the number of different repeat lengths present at each of the more than one microsatellite loci based on the sequence information to generate a site score for each of the more than one microsatellite loci, wherein the site scores for the more than one microsatellite loci include a likelihood score, wherein the site scores for the more than one microsatellite loci include a site score based on the Akaike information criterion (AIC), and the site score based on the Akaike information criterion (AIC) tests for the presence of somatic gain / loss at the more than one microsatellite loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, and the method includes: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs including the observed repeat length of sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (c) for each of the more than one microsatellite loci, comparing the site score for a given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (d) identifying the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds a site-specific training threshold for the given microsatellite locus to generate a microsatellite instability score, the microsatellite instability score comprising the number of unstable microsatellite loci from the more than one microsatellite loci; and (e) classifying the MSI status of the sample as unstable when the microsatellite instability score exceeds a population training threshold for the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.
5. A system comprising a controller comprising or having access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method of identifying one or more customized therapies for treating a disease in a subject, the method comprising: (a) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on sequence information to generate a site score for each of the more than one microsatellite loci, wherein the sequence information is from a population of microsatellite loci in the sample, wherein the site scores for the more than one microsatellite loci include a likelihood score, wherein the site scores for the more than one microsatellite loci include an Akaike information criterion (AIC)-based site score, wherein the Akaike information criterion (AIC)-based site score tests for the presence of somatic gain / loss at the more than one microsatellite loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, and the method includes: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs including the observed repeat length of sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (b) for each of the more than one microsatellite loci, comparing the site score for a given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (c) identifying the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds a site-specific training threshold for the given microsatellite locus to generate a microsatellite instability score, the microsatellite instability score comprising the number of unstable microsatellite loci from the more than one microsatellite loci; (d) classifying the microsatellite instability (MSI) status of the sample as unstable when the microsatellite instability score exceeds a population-trained threshold for the population of microsatellite loci in the sample to identify an unstable sample; and (e) comparing the microsatellite instability status of the sample to one or more comparative outcomes indexed by one or more therapies to identify one or more customized therapies for treating the subject's disease.
6. A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method of identifying one or more customized therapies for treating a disease in a subject, the method comprising: (a) quantifying the number of different repeat lengths present at each of more than one microsatellite loci based on sequence information to generate a site score for each of the more than one microsatellite loci, wherein the sequence information is from a population of microsatellite loci in the sample, wherein the site scores for the more than one microsatellite loci include a likelihood score, wherein the site scores for the more than one microsatellite loci include an Akaike information criterion (AIC)-based site score, wherein the Akaike information criterion (AIC)-based site score tests for the presence of somatic gain / loss at the more than one microsatellite loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, and the method includes: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs including the observed repeat length of sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (b) for each of the more than one microsatellite loci, comparing the site score for a given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (c) identifying the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds a site-specific training threshold for the given microsatellite locus to generate a microsatellite instability score, the microsatellite instability score comprising the number of unstable microsatellite loci from the more than one microsatellite loci; (d) classifying the microsatellite instability (MSI) status of the sample as unstable when the microsatellite instability score exceeds a population-trained threshold for the population of microsatellite loci in the sample to identify an unstable sample; and (e) comparing the microsatellite instability status of the sample to one or more comparative outcomes indexed by one or more therapies to identify one or more customized therapies for treating the subject's disease.
7. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, comprising estimating parameters of the model using maximum likelihood estimation (MLE).
8. The method, system, or computer-readable medium of claim 7, comprising determining the MLE using a Nelder-Mead algorithm.
9. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, wherein γ include: (a) The read-level error rate for microsatellite lengths observed in sequencing reads that are one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule strand; and / or (b) Read-level error rate for microsatellite lengths observed in sequenced reads that are one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule strand.
10. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, wherein β include: (a) strand-level error rate where the expected microsatellite length of the sense strand is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (b) the strand-level error rate where the expected microsatellite length of the antisense strand is one repeat unit longer than the expected microsatellite length of the starting nucleic acid molecule; (c) a strand-level error rate where the expected microsatellite length of the sense strand is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule; and / or (d) Strand-level error rate where the expected microsatellite length of the antisense strand is one repeat unit shorter than the expected microsatellite length of the starting nucleic acid molecule.
11. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, comprising identifying the given microsatellite locus as unstable when the site score for the given microsatellite locus statistically exceeds a site-specific training threshold for the given microsatellite locus.
12. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the mutant allele fraction of the sample is estimated.
13. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the tumor fraction of the sample is estimated.
14. The method, system, or computer-readable medium of claim 13, wherein the tumor score comprises the maximum mutant allele fraction (MAF) of all somatic mutations identified in nucleic acids in the sample.
15. The method, system, or computer-readable medium of claim 14, wherein the tumor fraction is less than 15% of all nucleic acids in the sample.
16. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the more than one microsatellite locus comprises the entirety of the population of microsatellite loci.
17. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the more than one microsatellite loci comprise a subset of the population of microsatellite loci.
18. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, comprising determining the site-specific training threshold and / or the population training threshold based on sequence information from a population of microsatellite loci in one or more training DNA samples.
19. The method, system, or computer-readable medium of claim 18, wherein the training DNA sample comprises a non-tumor cfDNA sample.
20. The method, system, or computer-readable medium of claim 18, wherein the training DNA samples comprise DNA from one or more tumor types.
21. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the method comprises having a sensitivity of at least 94% at a limit of detection (LOD) of 0.2% tumor fraction for nucleic acid in the sample.
22. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the method comprises having at least 99% specificity for non-tumor DNA in the sample.
23. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the determined MSI status of the sample comprises at least 95% concordance with a corresponding MSI status of the sample determined using a PCR-based MSI assessment technique across a tumor fraction range of 1.4% to 15%.
24. The method, system, or computer-readable medium of claim 23, wherein the determined MSI status of the sample comprises at least 96% concordance with a corresponding MSI status of the sample determined using a PCR-based MSI assessment technique across a tumor fraction range of 1.4% to 15%.
25. The method, system, or computer-readable medium of claim 23, wherein the determined MSI status of the sample comprises at least 97% concordance with a corresponding MSI status of the sample determined using a PCR-based MSI assessment technique across a tumor fraction range of 1.4% to 15%.
26. The method, system, or computer-readable medium of claim 23, wherein the determined MSI status of the sample comprises at least 98% concordance with a corresponding MSI status of the sample determined using a PCR-based MSI assessment technique across a tumor fraction range of 1.4% to 15%.
27. The method, system, or computer-readable medium of claim 23, wherein the determined MSI status of the sample comprises at least 99% concordance with a corresponding MSI status of the sample determined using a PCR-based MSI assessment technique across a tumor fraction range of 1.4% to 15%.
28. The method, system, or computer-readable medium of claim 23, wherein the consistency is 100%.
29. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, comprising classifying the MSI status of the sample as MSI-high MSI-H when the microsatellite instability score is greater than 1 unstable microsatellite locus from the more than one microsatellite loci.
30. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, comprising classifying the MSI status of the sample as MSI-high MSI-H when the number of unstable microsatellite loci constitutes 0.1%, 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, or 25% of the more than one microsatellite loci.
31. The system of claim 5 or the computer-readable medium of claim 6, wherein the disease comprises cancer, the cancer comprising at least one tumor type selected from the group consisting of biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, colorectal cancer, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, gallbladder cancer, renal cell carcinoma, Wilms' tumor, leukemia, liver cancer, bile duct cancer, lung cancer, mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, mantle cell lymphoma, T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pseudopapillary tumor, acinar cell carcinoma, prostate cancer, skin cancer, melanoma, small intestine cancer, gastric cancer, uterine cancer, and uterine sarcoma.
32. The system of claim 5 or the computer-readable medium of claim 6, wherein the disease comprises a cancer comprising at least one tumor type selected from the group consisting of cervical squamous cell carcinoma, rectal cancer, colon cancer, esophageal adenocarcinoma, ocular melanoma, cutaneous melanoma, gallbladder adenocarcinoma, clear cell renal cell carcinoma, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), hepatoblastoma, non-small cell lung cancer (NSCLC), diffuse large B-cell lymphoma, peripheral T-cell lymphoma, oral squamous cell carcinoma, pancreatic ductal adenocarcinoma, prostate adenocarcinoma, gastric epithelial carcinoma, hepatoepithelial carcinoma, and hepatocellular carcinoma.
33. The system of claim 5 or the computer-readable medium of claim 6, wherein the disease comprises cancer comprising at least one tumor type selected from the group consisting of: hereditary nonpolyposis colorectal cancer, esophageal squamous cell carcinoma, uveal melanoma, hepatocellular carcinoma, and precursor T-lymphoblastic lymphoma / leukemia.
34. The system of claim 5 or the computer-readable medium of claim 6, wherein the disease is malignant melanoma or colorectal adenocarcinoma.
35. The system of claim 5 or the computer-readable medium of claim 6, wherein the therapy comprises at least one immunotherapy.
36. The system or computer-readable medium of claim 35, wherein the immunotherapy comprises at least one checkpoint inhibitor antibody.
37. The system or computer-readable medium of claim 35, wherein the immunotherapy comprises an antibody to PD-1, PD-2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40, or CD47.
38. The system or computer-readable medium of claim 35, wherein the immunotherapy comprises administering a pro-inflammatory cytokine directed against at least one tumor type.
39. The system or computer-readable medium of claim 35, wherein the immunotherapy comprises administering T cells directed against at least one tumor type.
40. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, further comprising obtaining the sample from a subject.
41. The method, system, or computer-readable medium of claim 40, wherein the sample is selected from the group consisting of tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, stool, synovial fluid, spinal fluid, and saliva.
42. The method, system, or computer-readable medium of claim 40, wherein the subject is a mammalian subject.
43. The method, system, or computer-readable medium of claim 42, wherein the mammalian subject is a human subject.
44. The method, system, or computer-readable medium of claim 40, wherein the sample comprises cell-free nucleic acid.
45. The method, system, or computer-readable medium of claim 44, wherein the sample comprises circulating tumor nucleic acids.
46. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, further comprising receiving the sequence information generated from the sample, wherein the sequence information comprises cfDNA sequencing reads from the population of microsatellite loci in the sample.
47. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, comprising amplifying one or more segments of nucleic acid in the sample to produce at least one amplified nucleic acid.
48. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, further comprising sequencing nucleic acid from the sample to generate the sequence information.
49. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the sequence information is obtained from a targeted segment of nucleic acid in the sample, wherein the targeted segment is obtained by selectively enriching one or more regions of nucleic acid in the sample prior to sequencing.
50. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, further comprising amplifying the obtained targeted segment prior to sequencing.
51. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, further comprising attaching one or more adaptors comprising a barcode to the nucleic acid prior to sequencing.
52. The method, system, or computer readable medium of claim 48, wherein the sequencing is selected from the group consisting of targeted sequencing, intron sequencing, exome sequencing, and whole genome sequencing.
53. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, comprising sequencing at least 50 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
54. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 100 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
55. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 150 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
56. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 200 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
57. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 250 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
58. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 500 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
59. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 750 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
60. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 1,000 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
61. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 1,500 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
62. The method, system, or computer-readable medium of claim 53, comprising sequencing at least 2,000 targeted genomic regions in the nucleic acid of the sample to generate the sequence information.
63. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, wherein the number of different repeat lengths comprises the frequency of each different repeat length present at each of the more than one microsatellite loci.
64. The method of any one of claims 1-4, wherein at least a portion of the method is computer-implemented.
65. The method of any one of claims 1-4, the system of claim 5, or the computer-readable medium of claim 6, further comprising generating a report comprising information about the instability status of the sample, information about the instability score of the sample, and / or information about one or more customized therapies for treating a disease in a subject.
66. The method, system, or computer-readable medium of claim 65, further comprising transmitting the report to a third party.
67. The method, system, or computer-readable medium of claim 66, further comprising transmitting the report to the subject or a health care provider.
68. The method of claim 1, further comprising classifying the repetitive nucleic acid instability state of the nucleic acid sample as stable if the repetitive nucleic acid instability score is below or at a population training threshold for the population of repetitive nucleic acid loci in the nucleic acid sample.
69. The method of claim 2, further comprising classifying the repetitive DNA instability state of the sample as stable if the repetitive DNA instability score is below or at a population training threshold for the population of the repetitive DNA loci in the sample.
70. The method of any one of claims 3-4, the system of claim 5, or the computer-readable medium of claim 6, further comprising classifying the microsatellite instability state of the sample as stable if the microsatellite instability score is below or at a population training threshold for the population of microsatellite loci in the sample.
71. A system comprising a controller comprising or having access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) receiving sequence information from a population of microsatellite loci in a sample; (b) quantifying the number of different repeat lengths present at each of the more than one microsatellite loci based on the sequence information to generate a site score for each of the more than one microsatellite loci, wherein the site scores for the more than one microsatellite loci include a likelihood score, wherein the site scores for the more than one microsatellite loci include a site score based on the Akaike information criterion (AIC), and the site score based on the Akaike information criterion (AIC) tests for the presence of somatic gain / loss at the more than one microsatellite loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, where: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs including the observed repeat length of sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (c) for each of the more than one microsatellite loci, comparing the site score for a given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (d) identifying the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds a site-specific training threshold for the given microsatellite locus to generate a microsatellite instability score, the microsatellite instability score comprising the number of unstable microsatellite loci from the more than one microsatellite loci; and (e) classifying the microsatellite instability (MSI) status of the sample as unstable when the microsatellite instability score exceeds a population training threshold for the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.
72. The system of claim 71, comprising a nucleic acid sequencer operably connected to the controller, the nucleic acid sequencer configured to provide the sequence information from the population of microsatellite loci in the sample.
73. The system of claim 72, wherein the nucleic acid sequencer is configured to perform pyrosequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, synthesis sequencing, ligation sequencing, or hybridization sequencing on the nucleic acid to generate sequencing reads.
74. The system of claim 71, comprising a sample preparation component operably connected to the controller, the sample preparation component configured to prepare the sample to be sequenced by a nucleic acid sequencer.
75. The system of claim 74, wherein the sample preparation component is configured to selectively enrich a region of nucleic acid from the sample.
76. The system of claim 74, wherein the sample preparation component is configured to attach one or more adaptors comprising a molecular barcode to the nucleic acid.
77. The system of claim 71, comprising a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component configured to amplify nucleic acid.
78. The system of claim 77, wherein the nucleic acid amplification component is configured to amplify a selectively enriched region of nucleic acid from the sample.
79. The system of claim 71, comprising a material transfer component operably connected to the controller, the material transfer component configured to transfer one or more materials between the nucleic acid sequencer and the sample preparation component.
80. The system of claim 71 , comprising a database operably connected to the controller, the database comprising one or more comparative results indexed with one or more therapies, and wherein the electronic processor further performs at least the following: (f) comparing the microsatellite instability status of the sample to one or more comparative results, wherein a substantial match between the microsatellite instability score and the comparative result indicates a predicted subject response to a therapy.
81. A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform at least the following: (a) receiving sequence information from a population of microsatellite loci in a sample; (b) quantifying the number of different repeat lengths present at each of the more than one microsatellite loci based on the sequence information to generate a site score for each of the more than one microsatellite loci, wherein the site scores for the more than one microsatellite loci include a likelihood score, wherein the site scores for the more than one microsatellite loci include a site score based on the Akaike information criterion (AIC), and the site score based on the Akaike information criterion (AIC) tests for the presence of somatic gain / loss at the more than one microsatellite loci, wherein a given AIC-based site score is calculated using the following formula: AIC = k -log likelihood , in k is the number of parameters used in the model, where: (i) Calculate the null hypothesis score of the model using the following formula: in AIC 0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the observed repeat length of sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter; (ii) Calculate the alternative hypothesis score for the model using the following formula: in AIC min is the alternative hypothesis, min α It will α The effect of minimizing all values of k is the number of parameters used in the model, Pr is the probability, obs including the observed repeat length of sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that one or more α i The sum of is equal to 1; and (iii) Detecting the site score using the following formula: ; (c) for each of the more than one microsatellite loci, comparing the site score for a given microsatellite locus to a site-specific training threshold for the given microsatellite locus; (d) identifying the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds a site-specific training threshold for the given microsatellite locus to generate a microsatellite instability score, the microsatellite instability score comprising the number of unstable microsatellite loci from the more than one microsatellite loci; and (e) classifying the microsatellite instability (MSI) status of the sample as unstable when the microsatellite instability score exceeds a population training threshold for the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.
82. A system comprising a controller comprising or having access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform the method of any one of claims 1-3.
83. A computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform the method of any one of claims 1-3.
Citation Information
Patent Citations
Improvement in lanterns
US170510A
Oligonucleotides
US20010053519A1
Method and apparatus for imaging a sample on a device
US20030152490A1
Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels
US20110160078A1
Coupled amplification and ligation method
US5912148A