Detection of microsatellite instability in cell-free DNA

The method for determining microsatellite instability in cfDNA samples using site and population-trained thresholds addresses the need for accurate MSI evaluation, achieving high agreement with conventional methods and guiding cancer treatment decisions.

JP7699261B2Active Publication Date: 2025-06-26GUARDANT HEALTH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024060146
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-04
Filing Date
2024-04-03
Publication Date
2025-06-26
Estimated Expiration
2039-08-30

AI Technical Summary

Technical Problem

There is a need for methods to evaluate repetitive element instability, including microsatellite instability (MSI), in various samples, particularly cell-free DNA (cfDNA) samples, as existing methods are inadequate for comprehensive genomic profiling and prognosis in cancer treatment.

Method used

A method is disclosed for determining the microsatellite and/or other repetitive DNA instability status of a cfDNA sample by quantifying different repeat lengths at repetitive nucleic acid loci, comparing site scores to locus-specific thresholds, and classifying instability based on population-trained thresholds, using computer-executed processes.

Benefits of technology

The method achieves high agreement with conventional PCR-based MSI evaluation techniques, providing accurate MSI status determination for guiding disease prognosis and treatment decisions, with sensitivity of at least 94% and analytical specificity of at least 99%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699261000019
    Figure 0007699261000019
  • Figure 0007699261000020
    Figure 0007699261000020
  • Figure 0007699261000021
    Figure 0007699261000021
Patent Text Reader

Abstract

To provide a method for determining the repetitive nucleic acid instability status of a nucleic acid sample.SOLUTION: A method for determining a microsatellite instability status of a sample includes the steps of: quantifying many different repeat lengths present at each of a plurality of repetitive nucleic acid genetic loci from sequence information to generate a site score of each of the plurality of repetitive nucleic acid genetic loci; calling a given repetitive nucleic acid genetic locus as being unstable in the case that the site score of the given repetitive nucleic acid genetic locus exceeds a site specific trained threshold for the given repetitive nucleic acid genetic locus to generate a repetitive nucleic acid instability score including many unstable repetitive nucleic acid genetic loci from the plurality of repetitive nucleic acid genetic loci; and classifying the repetitive nucleic acid instability status of a nucleic acid sample as being unstable in the case that the repetitive nucleic acid unstable score exceeds a collectively trained threshold of a group of the repetitive nucleic acid genetic loci in the nucleic acid sample to determine the repetitive nucleic acid unstable status of the nucleic acid sample by it.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 726,182, filed Aug. 31, 2018, U.S. Provisional Patent Application No. 62 / 823,578, filed Mar. 25, 2019, and U.S. Provisional Patent Application No. 62 / 857,048, filed Jun. 4, 2019, and relies on the filing dates of the foregoing provisional applications, the entire disclosures of which are incorporated herein by reference.

Background Art

[0002] Repetitive nucleic acid elements are patterns of nucleotides (DNA or RNA) that exist in multiple copies throughout eukaryotic and prokaryotic genomes. Examples of such repetitive elements include, among others, microsatellites, short tandem repeats (STRs), and minisatellites. Microsatellites typically contain repeat units of less than 10 base pairs. STRs generally contain repeat units of 2 - 13 nucleotides that are often repeated hundreds of times within a given stretch of nuclear DNA. STR analysis is a common tool used in forensic analysis. Minisatellites are repetitive elements that typically have repeat units of about 10 - 60 base pairs.

[0003] Microsatellites are, in particular, highly polymorphic DNA repeat regions. Microsatellite instability (MSI) is a guideline - recommended biomarker used in the evaluation of prognosis and treatment selection, including checkpoint inhibitors recently approved for the treatment of cancers with high MSI (MSI - H) status. Plasma - based next - generation DNA sequencing (NGS) assays are increasingly being used for comprehensive genomic profiling of cancers, but methods for detecting MSI status from cell - free DNA (cfDNA) data are under development. In addition, the impact of variable tumor dropout on MSI detection has not been evaluated in the past.

Summary of the Invention

Means for Solving the Problems

[0004] There is still a need for methods and related aspects useful for the evaluation of repetitive element instability, including MSI, in various samples, particularly cfDNA samples.

[0005] The present application discloses a method, a computer-readable medium, and a system useful for determining the microsatellite and / or other repetitive DNA instability status of a cell-free DNA (cfDNA) sample from a patient and useful for guiding disease prognosis and treatment decisions. Typically, at least a portion of the methods disclosed herein are computer-executed and achieve results that have a high degree of agreement with the results obtained using more conventional polymerase chain reaction (PCR)-based MSI evaluation techniques.

[0006] Further aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, in which merely exemplary embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different embodiments and some of the details thereof are capable of variation in various obvious respects without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0007] In one aspect, the present disclosure provides a method for determining the repetitive nucleic acid instability status of a nucleic acid sample. The method includes (a) quantifying a number of different repeat lengths present at each of a plurality of repetitive nucleic acid loci from sequence information to generate a site score for each of the plurality of repetitive nucleic acid loci. The sequence information is from a population of repetitive nucleic acid loci in the nucleic acid sample. The method also includes (b) calling a given repetitive nucleic acid locus as unstable when the site score of the given repetitive nucleic acid locus exceeds a site-specific trained threshold for the given repetitive nucleic acid locus, thereby generating a repetitive nucleic acid instability score that includes a number of unstable repetitive nucleic acid loci from the plurality of repetitive nucleic acid loci. Additionally, the method includes (c) classifying the repetitive nucleic acid instability status of the nucleic acid sample as unstable when the repetitive nucleic acid instability score exceeds a population-trained threshold for the population of repetitive nucleic acid loci in the nucleic acid sample, thereby determining the repetitive nucleic acid instability status of the nucleic acid sample.

[0008] In another aspect, the present disclosure provides a method for determining the repetitive DNA instability status of a sample (e.g., a cell-free DNA (cfDNA) sample). The method includes (a) quantifying the number of different repeat lengths present at each of a plurality of repetitive DNA loci from sequence information to generate a site score for each of the plurality of repetitive DNA loci. The sequence information is from a population of repetitive DNA loci in the sample. The method also includes (b) comparing the site score of a given repetitive DNA locus for each of the plurality of repetitive DNA loci to a locus-specific trained threshold of the given repetitive DNA locus. The method further includes (c) calling the given repetitive DNA locus as unstable if the site score of the given repetitive DNA locus exceeds the locus-specific trained threshold of the given repetitive DNA locus, thereby generating a repetitive DNA instability score that includes a number of unstable repetitive DNA loci from the plurality of repetitive DNA loci. In addition, the method includes (d) classifying the repetitive DNA instability status of the sample as unstable if the repetitive DNA instability score exceeds a population-trained threshold of the population of repetitive DNA loci in the sample, thereby determining the repetitive DNA instability status of the sample. The methods disclosed herein are typically at least partially computer-executed.

[0009] In another aspect, the present disclosure provides a method for determining the microsatellite instability (MSI) status of a sample. The method includes (a) quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in the sample. The method also includes (b) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci to a locus-specific trained threshold of the given microsatellite locus. The method further includes (c) calling the given microsatellite locus as unstable if the site score of the given microsatellite locus exceeds the locus-specific trained threshold of the given microsatellite locus, thereby generating a microsatellite instability score that includes a number of unstable microsatellite loci from the plurality of microsatellite loci. Additionally, the method includes (d) classifying the MSI status of the sample as unstable if the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.

[0010] In another aspect, the present disclosure provides a method for determining the microsatellite instability (MSI) status of a sample. The method includes (a) receiving sequence information from a population of microsatellite loci in the sample, and (b) quantifying the number of different repeat lengths present at each of a plurality of microsatellite loci from the sequence information to generate a site score for each of the plurality of microsatellite loci. The method further includes (c) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci to a locus-specific trained threshold of the given microsatellite locus. The method also includes (d) calling the given microsatellite locus as unstable if the site score of the given microsatellite locus exceeds the locus-specific trained threshold of the given microsatellite locus, thereby generating a microsatellite instability score that includes a number of unstable microsatellite loci from the plurality of microsatellite loci. Additionally, the method includes (e) classifying the MSI status of the sample as unstable if the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.

[0011] In another aspect, the present disclosure provides a method for identifying one or more customized therapies for treating a disease in a subject. The method includes (a) quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in a sample. The method also includes (b) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci to a site-specific trained threshold of the given microsatellite locus. The method further includes (c) calling the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds the site-specific trained threshold of the given microsatellite locus, and generating a microsatellite instability score that includes a number of unstable microsatellite loci from the plurality of microsatellite loci. The method also includes (d) classifying the MSI status of the sample as unstable and identifying an unstable sample when the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample. Additionally, the method includes (e) comparing the microsatellite instability status of the sample to one or more comparator results indexed by one or more therapies to identify one or more customized therapies for treating the disease in the subject.

[0012] In another aspect, the present disclosure provides a method for treating a disease in a subject. The method includes: (a) quantifying the number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in a sample; (b) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci with a locus-specific trained threshold of the given microsatellite locus; (c) when the site score of a given microsatellite locus exceeds the locus-specific trained threshold of the given microsatellite locus, calling the given microsatellite locus as unstable and generating a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; (d) when the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample, classifying the MSI status of the sample as unstable and identifying an unstable sample; (e) comparing the microsatellite instability status of the sample with one or more comparator results indexed by one or more therapies to identify one or more customized therapies for treating the disease in the subject; and (f) when the microsatellite instability status of the sample substantially matches the comparator result, administering at least one of the identified customized therapies to the subject, thereby treating the disease in the subject.

[0013] In another aspect, the present disclosure provides a method of treating a disease in a subject. The method includes administering to the subject one or more customized therapies, thereby treating the disease in the subject, wherein the customized therapy in this case is (a) quantifying the number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in a sample, as identified by the step. The method also includes (b) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci to a site-specific trained threshold of the given microsatellite locus. The method further includes (c) calling the given microsatellite locus as unstable when the site score of the given microsatellite locus exceeds the site-specific trained threshold of the given microsatellite locus, and generating a microsatellite instability score that includes a number of unstable microsatellite loci from the plurality of microsatellite loci. The method also includes (d) classifying the MSI status of the sample as unstable and identifying an unstable sample when the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample. The method further includes (e) comparing the microsatellite instability status of the sample to one or more comparator results indexed by one or more therapies. Additionally, the method includes (f) identifying one or more customized therapies for treating the disease in the subject when the microsatellite instability status of the sample and the comparator results substantially match.

[0014] In some embodiments, the locus scores for a plurality of microsatellite loci include likelihood scores. In certain embodiments of these embodiments, the likelihood score includes a probabilistically-based log-likelihood score that discriminates biological signals derived from a large number of nucleic acid fragments of somatic origin (in some embodiments, -cfDNA fragments) in a sample from noise arising from post-sample collection artifacts. In some embodiments, the method uses at least two parameters, where at least a first parameter includes allele frequencies and at least a second parameter includes at least one error mode, to determine a probabilistically-based log-likelihood score for individual microsatellite loci in sequence information from a sample. Typically, the allele frequencies include the frequencies of nucleic acids with different repeat lengths in the sequence information from the sample. In some embodiments, the at least one error mode includes a random error mode and a strand-specific error mode. In certain embodiments, the locus scores for a plurality of microsatellite loci include the difference or ratio between (a) a score that measures the support of the observed sequence for the null hypothesis that a given microsatellite locus is stable and (b) a score that measures the support of the observed sequence for the alternative hypothesis that a given microsatellite locus is unstable. In some embodiments, the locus scores for a plurality of microsatellite loci are generated using one or more of a likelihood criterion, a log-likelihood criterion, a posterior probability criterion, an Akaike information criterion (AIC), a Bayesian information criterion, and / or the like.

[0015] In some embodiments, the locus scores for a plurality of microsatellite loci include locus scores based on the Akaike information criterion (AIC) that test for the presence of somatic indels at the plurality of microsatellite loci. In certain embodiments of these embodiments, the locus score based on a given AIC is given by the formula: AIC = k - log-likelihood (where k is the number of parameters used in the model) It is calculated using. Optionally, the method includes the step of estimating the parameters of the model using maximum likelihood estimation (MLE). In some of these embodiments, the method includes the step of determining the MLE using the Nelder-Mead algorithm. In certain embodiments, the method is given by the following equation: AIC0 = k - log(Pr(obs|β,γ)) (where AIC0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the repeat lengths of the observed sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter) including calculating a null hypothesis score for the model (e.g., a score measuring the support of the observed sequence for the null hypothesis that a given microsatellite locus is stable) using. In certain embodiments, obs is a number of observed sequencing reads covering a given microsatellite locus. In some of these embodiments, the method is given by the following equation: AIC min = min α (k - log(Pr(obs|β,γ,α)) (where AIC min is the alternative hypothesis, min α is the minimization effect over all values of α, k is the number of parameters used in the model, Pr is the probability, obs includes the repeat lengths of the observed sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, α is at least one allele frequency, where α is a vector of allele frequencies such that the sum of one or more α i is equal to 1) Including the step of calculating the model's alternative hypothesis score (e.g., a score that measures the support of the observed sequence for the alternative hypothesis that a given microsatellite locus is unstable) using. In some embodiments, obs is a number of observed sequencing reads covering a given microsatellite locus. In certain embodiments of these embodiments, the method determines a model change (i.e., ΔAIC) to determine a site score using the following formula: ΔAIC = AIC0 - AIC min Including the step of detecting using. In some of these embodiments, γ includes (a) the read-level error rate where the microsatellite length observed within a sequencing read is one repeat unit longer than the expected microsatellite length for the strand of the originating nucleic acid molecule; and / or (b) the read-level error rate where the microsatellite length observed within a sequencing read is one repeat unit shorter than the expected microsatellite length for the strand of the originating nucleic acid molecule. In certain embodiments of these embodiments, β includes (a) the strand-level error rate where the expected microsatellite length of the sense strand is one repeat unit longer than the expected microsatellite length of the nucleic acid-derived molecule; (b) the strand-level error rate where the expected microsatellite length of the antisense strand is one repeat unit longer than the expected microsatellite length of the nucleic acid-derived molecule; (c) the strand-level error rate where the expected microsatellite length of the sense strand is one repeat unit shorter than the expected microsatellite length of the nucleic acid-derived molecule; and / or (d) the strand-level error rate where the expected microsatellite length of the antisense strand is one repeat unit shorter than the expected microsatellite length of the nucleic acid-derived molecule. Typically, the method includes the step of calling a given microsatellite locus as unstable when the site score of the given microsatellite locus statistically exceeds the site-specific trained threshold of the given microsatellite locus.

[0016] In some embodiments, the site score based on AIC is given by the following formula: AIC = 2(k - log likelihood) (where k is the number of parameters used in the model) is calculated using these. In these embodiments, AIC0 and AIC min are calculated using the above formula.

[0017] For clarity, in embodiments where a score based on AIC is determined using the formula AIC = 2(k - log likelihood), the site-specific threshold used to classify a site as unstable is twice the site-specific threshold used in the previous embodiments where a score based on AIC is determined using the formula AIC = k - log likelihood.

[0018] Typically, the mutant allele fraction (MAF) of a sample (e.g., a cfDNA sample) is estimated. In some of these embodiments, the tumor fraction of a sample (e.g., a cfDNA sample) is estimated. In certain embodiments, the tumor fraction includes the maximum mutant allele fraction (MAF) of all somatic mutations identified in the nucleic acids in the sample (e.g., a cfDNA sample). In some embodiments, the tumor fraction is less than about 0.05%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14% or about 15% of all nucleic acids in the sample (e.g., a cfDNA sample). In some embodiments, the plurality of microsatellite loci includes all of a population of microsatellite loci, while in other embodiments, the plurality of microsatellite loci includes a subset of a population of microsatellite loci. In certain embodiments, the method includes determining a site-specific trained threshold and / or a population-trained threshold from sequence information from a population of microsatellite loci in one or more training DNA samples. In some of these embodiments, the training DNA samples include non-tumor cfDNA training samples and / or DNA from one or more tumor types.

[0019] In some embodiments, the method has a sensitivity of at least about 94% at a limit of detection (LOD) of about 0.1% to 0.4% tumor fraction of nucleic acids in a sample. In some embodiments, the method has an analytical specificity of at least about 99% for non-tumor DNA in the sample. In certain embodiments, the determined MSI status of the sample includes at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% agreement with the corresponding MSI status of the sample determined using a PCR-based MSI assessment technique over a tumor fraction range of about 1% to about 15%. In some of these embodiments, the agreement is 100%. In some embodiments, the method includes classifying the MSI status of the sample as high MSI (MSI-H) if the microsatellite instability score is greater than an unstable microsatellite locus number of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 30, about 40, about 50 or more from a plurality of microsatellite loci. In certain embodiments, the method includes classifying the MSI status of the sample as high MSI (MSI-H) if the number of unstable microsatellite loci constitutes about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20% or about 25% of the plurality of microsatellite loci. In some embodiments, the number of different repeat lengths includes the frequency of each different repeat length present at each of the plurality of microsatellite loci.

[0020] In various embodiments, the present disclosure includes a method of selecting a customized therapy for treating a disease in a subject and / or a method of treating a disease in a subject. In some embodiments of these embodiments, the disease is biliary cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, uveal melanoma, choroidal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myelogenous (CML), chronic myelomonocytic (CMML), liver cancer, hepatocarcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma, including but not limited to at least one tumor type selected from the group consisting of cancer.

[0021] In some embodiments, the therapy includes at least one immunotherapy (e.g., checkpoint inhibitory antibodies, autologous cytotoxic T cells, personalized cancer vaccines, etc.). In certain embodiments, for example, the immunotherapy includes antibodies against PD-1, PD-2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40, or CD47. In some embodiments, the immunotherapy includes administration of a pro-inflammatory cytokine against at least one tumor type. Optionally, the immunotherapy includes administration of T cells against at least one tumor type.

[0022] In some embodiments, the method includes obtaining a sample from a subject. Essentially, any sample type can be utilized as needed. In certain embodiments, for example, the sample is tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, feces, synovial fluid, cerebrospinal fluid, saliva, and / or the like. Typically, the subject is a mammalian subject (e.g., a human subject). In some embodiments, the sample is blood. In some embodiments, the sample is plasma. In some embodiments, the sample is serum. In some embodiments, the sample includes cell-free DNA (i.e., a cfDNA sample). In some embodiments, the cfDNA sample includes circulating tumor nucleic acids.

[0023] In certain embodiments, the method includes receiving array information generated from a sample, the array information including sequencing reads from a population of microsatellite loci in the sample. In some embodiments, the method includes amplifying one or more segments of nucleic acids in the sample to produce at least one amplified nucleic acid. In certain embodiments, the method includes sequencing nucleic acids from the sample to generate the array information. In some embodiments, the sample can be a cfDNA sample. In these embodiments, the array information includes cfDNA sequencing reads from a population of microsatellite loci in the cfDNA sample. In some embodiments, the array information is obtained from targeted segments of nucleic acids in the sample, the targeted segments being obtained by selectively enriching one or more regions from the nucleic acids in the sample prior to sequencing. In some of these embodiments, the method includes amplifying the obtained targeted segments prior to sequencing. In these embodiments, the method typically includes binding one or more adapters including molecular barcodes to the nucleic acids prior to amplification. In some embodiments, the method includes binding one or more sample indices by amplification prior to sequencing. Essentially any nucleic acid sequencing technique can be used as needed or adapted for use in performing the methods disclosed herein. For example, the sequencing is optionally selected from targeted sequencing, intron sequencing, exome sequencing, whole genome sequencing, and / or the like. In some embodiments, the sequencing is targeted sequencing. In some embodiments, the method includes sequencing at least about 50, about 100, about 150, about 200, about 250, about 500, about 750, about 1,000, about 1,500, about 2,000, or more targeted genomic regions within the nucleic acids of the sample to generate the array information.

[0024] In another aspect, the present disclosure, when executed by at least one electronic processor, includes at least (a) receiving sequence information from a population of microsatellite loci in a sample; (b) quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from the sequence information to generate a site score for each of the plurality of microsatellite loci; (c) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci to a locus-specific trained threshold of the given microsatellite locus; (d) calling the given microsatellite locus as unstable if the site score of the given microsatellite locus exceeds the locus-specific trained threshold of the given microsatellite locus, to generate a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; and (e) classifying the MSI status of the sample as unstable if the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample, thereby determining the MSI status of the sample. The present disclosure provides a system including a computer-readable medium including non-transitory computer-executable instructions for performing the above steps, or a controller capable of accessing such a computer-readable medium.

[0025] In some embodiments, the system includes a nucleic acid sequencer operably connected to a controller, the nucleic acid sequencer being configured to provide sequence information from a population of microsatellite loci in a sample. In some of these embodiments, the nucleic acid sequencer is configured to perform pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, sequencing by ligation, or sequencing by hybridization using nucleic acids to generate sequencing reads. In certain embodiments, the system includes a sample preparation component operably connected to the controller, the sample preparation component being configured to prepare a sample (in some cases, a cfDNA sample) to be sequenced by the nucleic acid sequencer. In some of these embodiments, the sample preparation component is configured to selectively enrich regions from the nucleic acids in the sample. In certain embodiments, the sample preparation component is configured to bind one or more adapters, including molecular barcodes, to the nucleic acids. In some embodiments, the system includes a nucleic acid amplification component operably connected to the controller, the nucleic acid amplification component being configured to amplify DNA (in some cases, cfDNA). In certain of these embodiments, the nucleic acid amplification component is configured to selectively amplify regions from the nucleic acids in the sample.

[0026] In certain embodiments, the system includes a material transfer component operably connected to a controller, the material transfer component being configured to transfer one or more materials between a nucleic acid sequencer and a sample preparation component. In some embodiments, the system includes a database operably connected to a controller, the database including one or more comparator results indexed by one or more therapies, and the electronic processor further performs at least (f) comparing the microsatellite locus status of a sample with one or more comparator results, and a substantial match between the microsatellite instability score and the comparator results indicates a predicted response to the therapy of interest.

[0027] In yet another aspect, the present disclosure provides a non-transitory computer-readable medium including computer-executable instructions that, when executed by at least one electronic processor, perform at least: (a) receiving sequence information from a population of microsatellite loci in a sample; (b) quantifying a number of different repeat lengths present at each of the plurality of microsatellite loci from the sequence information to generate a site score for each of the plurality of microsatellite loci; (c) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci with a locus-specific trained threshold of the given microsatellite locus; (d) calling the given microsatellite locus as unstable if the site score of the given microsatellite locus exceeds the locus-specific trained threshold of the given microsatellite locus, thereby generating a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; and (e) classifying the MSI status of the sample as unstable if the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample, thereby determining the MSI status of the sample.

[0028] The systems and computer-readable media disclosed herein include various embodiments. In some embodiments, for example, the locus scores of a plurality of microsatellite loci include likelihood scores. In certain specific embodiments of these embodiments, the likelihood score includes a probabilistically based log-likelihood score that distinguishes the biological signals derived from a large number of nucleic acid fragments of somatic origin (in some embodiments, cfDNA fragments) in a sample from the noise resulting from artifacts after sample collection. The probabilistically based log-likelihood score for an individual microsatellite locus in the sequence information from a sample is typically determined using at least two parameters, where at least the first parameter includes allele frequencies and at least the second parameter includes at least one error mode. The allele frequencies include the frequencies of nucleic acids with different repeat lengths in the sequence information from the sample. The at least one error mode typically includes a random error mode and a strand-specific error mode. In some embodiments, the locus scores of a plurality of microsatellite loci include the difference or ratio between (a) a score that measures the support of the observed sequence for the null hypothesis that a given microsatellite locus is stable and (b) a score that measures the support of the observed sequence for the alternative hypothesis that a given microsatellite locus is unstable. In some embodiments, the locus scores of a plurality of microsatellite loci are generated using one or more statistical model selection criteria such as likelihood criteria, log-likelihood criteria, posterior probability criteria, Akaike information criterion (AIC), Bayesian information criterion, and / or the like.

[0029] In some embodiments of the system or computer-readable media, the locus scores of a plurality of microsatellite loci include locus scores based on the Akaike information criterion (AIC) that test for the presence of somatic indels at the plurality of microsatellite loci. In certain specific embodiments, for example, a locus score based on a given AIC is given by the following formula: AIC = k - log-likelihood (where k is the number of parameters used in the model) It is calculated using. Optionally, the model parameters are estimated using maximum likelihood estimation (MLE). In some embodiments of these embodiments, the MLE is determined using the Nelder-Mead algorithm. In certain embodiments, the null hypothesis score of the model is given by the following equation: AIC0=k-log(Pr(obs|β,γ)) (where AIC0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the repeat length of the observed sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter) It is calculated using. In some embodiments, the alternative hypothesis score of the model is given by the following equation: AIC min =min α (k-log(Pr(obs|β,γ,α)) (where AIC min is the alternative hypothesis, min α is the minimization effect for all values of α, k is the number of parameters used in the model, Pr is the probability, obs includes the repeat length of the observed sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, and α is at least one allele frequency, where α is a vector of allele frequencies such that the sum of one or more α i is equal to 1) It is calculated using. In these embodiments, to determine the site score, the following equation: ΔAIC=AIC0 - AIC min is used to typically detect changes in the model.

[0030] In some embodiments, γ includes: (a) a read-level error rate where the microsatellite length observed in the sequencing read is one repeat unit longer than the expected microsatellite length for the strand of the originating nucleic acid molecule; and / or (b) a read-level error rate where the microsatellite length observed in the sequencing read is one repeat unit shorter than the expected microsatellite length for the strand of the originating nucleic acid molecule. In certain embodiments, β includes: (a) a strand-level error rate where the expected microsatellite length of the sense strand is one repeat unit longer than the expected microsatellite length of the nucleic acid-derived molecule; (b) a strand-level error rate where the expected microsatellite length of the antisense strand is one repeat unit longer than the expected microsatellite length of the nucleic acid-derived molecule; (c) a strand-level error rate where the expected microsatellite length of the sense strand is one repeat unit shorter than the expected microsatellite length of the nucleic acid-derived molecule; and / or (d) a strand-level error rate where the expected microsatellite length of the antisense strand is one repeat unit shorter than the expected microsatellite length of the nucleic acid-derived molecule.

[0031] In certain embodiments of the system or computer-readable medium, a given microsatellite locus is called unstable if the site score of the given microsatellite locus statistically exceeds the site-specific trained threshold of the given microsatellite locus. Typically, the tumor fraction is estimated, including the maximum mutant allele fraction (MAF) of all somatic mutations identified in the nucleic acids in the sample. In certain embodiments, the tumor fraction is less than about 0.05%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14% or about 15% of all nucleic acids in the sample. In some embodiments, the plurality of microsatellite loci includes all of the population of microsatellite loci, while in other embodiments, the plurality of microsatellite loci includes a subset of the population of microsatellite loci. In certain embodiments, the site-specific trained threshold and / or the population-trained threshold are determined from sequence information from a population of microsatellite loci in one or more training DNA samples. Optionally, the MSI status of the sample is classified as high MSI (MSI-H) if the microsatellite instability score exceeds a value of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50 or more from a plurality of microsatellite loci. In some embodiments, the MSI status of the sample is classified as high MSI (MSI-H) if the number of unstable microsatellite loci constitutes about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20% or about 25% of the plurality of microsatellite loci.

[0032] In yet another aspect, the present disclosure provides a system comprising a communication interface for obtaining sequencing information from one or more nucleic acids in a sample from a subject via a communication network; and a computer in communication with the communication interface, the computer comprising at least one computer processor and, when executed by the at least one computer processor, (a) receiving sequence information from a population of microsatellite loci in the sample; (b) quantifying a number of different repeat lengths present at each of the plurality of microsatellite loci from the sequence information to generate a locus score for each of the plurality of microsatellite loci; (c) comparing the locus score of a given microsatellite locus for each of the plurality of microsatellite loci to a locus-specific trained threshold of the given microsatellite locus; (d) calling the given microsatellite locus as unstable if the locus score of the given microsatellite locus exceeds the locus-specific trained threshold of the given microsatellite locus to generate a microsatellite instability score comprising a number of unstable microsatellite loci from the plurality of microsatellite loci; and (e) classifying the MSI status of the sample as unstable if the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample, thereby determining the MSI status of the sample, the computer-readable medium comprising machine-executable code for executing a method comprising these steps.

[0033] In some embodiments, the array information is provided by a nucleic acid sequencer. Typically, a nucleic acid sequencer performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, and / or another sequencing technique using nucleic acids to generate sequencing reads. In some embodiments, the nucleic acid sequence sequencer generates sequencing reads using a clonal single molecule array derived from a sequencing library. In certain embodiments, the nucleic acid sequence sequencer includes a chip having an array of microwells to sequence a sequencing library to generate sequencing reads.

[0034] The computer-readable media of the systems disclosed herein typically include a memory, a hard drive, or a computer server. In some embodiments, the communication network includes one or more servers capable of distributed computing. In some embodiments, the distributed computing is cloud computing. In some embodiments, the computer is installed on a computer server that is remotely located from the nucleic acid sequencer. In some embodiments, the systems disclosed herein include an electronic display that communicates with a computer via a network, the electronic display including a user interface for displaying results upon execution of (i)-(iv). In some of these embodiments, the user interface is a graphical user interface (GUI) or a web-based user interface. In some embodiments, the electronic display is installed on a personal computer. In certain embodiments, the electronic display is installed on a computer capable of connecting to the Internet. In some of these embodiments, the computer capable of connecting to the Internet is installed remotely from the computer. Typically, the computer-readable media includes a memory, a hard drive, or a computer server. In some embodiments, the communication network includes a telecommunications network, the Internet, an extranet, or an intranet.

[0035] In some embodiments, the results of the systems and methods disclosed herein are used as input data for generating a report. The report may be in paper or electronic form. For example, the MSI score and / or MSI status obtained by the methods and systems disclosed herein can be directly displayed in such a report. Alternatively or in addition, diagnostic information or one or more customized therapies based on the MSI status can be included in the report. In some embodiments, the report is communicated to a subject (e.g., a patient) or a healthcare provider.

[0036] In some embodiments, the method, system, or computer-readable medium further includes classifying the repeat nucleic acid instability status of the nucleic acid sample as stable if the repeat nucleic acid instability score is at or below the population-trained threshold of the population of repeat nucleic acid loci in the nucleic acid sample.

[0037] In some embodiments, the method, system, or computer-readable medium further includes classifying the repeat DNA instability status of the sample as stable if the repeat DNA instability score is at or below the population-trained threshold of the population of repeat DNA loci in the sample.

[0038] In some embodiments, the method, system, or computer-readable medium further includes classifying the microsatellite instability status of the sample as stable if the microsatellite instability score is at or below the population-trained threshold of the population of microsatellite loci in the sample.

[0039] The various steps of the methods disclosed herein, or the steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographical locations, e.g., countries, and / or by the same or different people.

[0040] The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate certain embodiments and, together with the written description herein, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein, including the accompanying drawings which are included by way of example and not limitation, will be better understood when read in conjunction with the accompanying drawings. It should be understood that like reference numerals identify like components throughout the drawings, unless the context dictates otherwise. It should also be understood that some or all of the figures may be schematic illustrations for purposes of illustration and do not necessarily depict the actual relative sizes and positions of the elements shown.

Brief Description of the Drawings

[0041]

Figure 1

[0042]

Figure 2

[0043]

Figure 3

[0044]

Figure 4

[0045]

Figure 5

[0046]

Figure 6

[0047]

Figure 7

[0048]

Figure 8

[0049]

Figure 9

[0050]

Figure 10

[0051]

Figure 11

[0052]

Figure 12

[0053]

Figure 13

[0054] definition In order that this disclosure may be more readily understood, certain terms are first defined below. Additional definitions for the following terms, as well as other terms, may be set forth throughout this specification. In the event that a definition of a term set forth below conflicts with a definition in an application or patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.

[0055] As used in this specification and the appended claims, the singular forms "a," "an," "an "An" and "the" are used interchangeably unless the context clearly dictates otherwise. A reference to "a method" includes one or more The present invention includes methods, methods, and / or steps of the type described herein and / or that will become apparent to those of skill in the art upon reading this disclosure or otherwise.

[0056] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Further, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In the descriptions and claims of methods, computer-readable media, and systems, the following terminology, and their grammatical variants, will be used according to the definitions set forth below.

[0057] About: As used herein, "about" or "approximately" when applied to one or more values or elements of interest refers to a value or element that is similar to the reference value or element being described. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater than or less than) of the reference value or element being described, unless otherwise stated or apparent from the context (except where such numbers exceed 100% of the possible values or elements).

[0058] Adapter: As used herein, an "adapter" is typically at least partially double-stranded and is a short nucleic acid (e.g., less than about 500 nucleotides in length, less than about 100 nucleotides in length, or less than about 50 nucleotides in length) used to ligate to one or both of a given sample nucleic acid molecule. An adapter can include a nucleic acid primer binding site that enables amplification of a nucleic acid molecule with an adapter adjacent to both ends, and / or a sequencing primer binding site for sequencing applications such as various next-generation sequencing (NGS) applications. An adapter can also include a binding site for a capture probe, such as an oligonucleotide bound to a flow cell support or the like. An adapter can also include a nucleic acid tag as described herein. The nucleic acid tag is typically positioned such that the nucleic acid tag is included in the amplicon and sequence read of a given nucleic acid molecule, relative to the amplification primer and the sequencing primer binding site. The same or different adapters can be ligated to each end of the nucleic acid molecule. In some embodiments, adapters of the same sequence except for different nucleic acid tags are ligated to each end of the nucleic acid molecule. In some embodiments, the adapter is a Y-shaped adapter with one end blunt-ended or having a tail as described herein, and is a Y-shaped adapter for binding to a nucleic acid molecule that is also blunt-ended or has a tail with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter with a blunt end or a tail end for binding to the nucleic acid molecule to be analyzed. Other examples of adapters include adapters having a T tail and adapters having a C tail.

[0059] Administering: As used herein, “administering” or “administration” of a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means giving, applying, or contacting the composition to the subject. Administration can be accomplished by any of a number of routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.

[0060] Akaike Information Criterion: As used herein, “Akaike Information Criterion” or “AIC” refers to a criterion for selecting a statistical model from a finite set of models, and includes a penalty term for the number of parameters in the model. In some embodiments, the model having the lowest AIC is selected.

[0061] Allele frequency: As used herein, “allele frequency” refers to the relative frequency of an allele at a particular locus within a population or in a given subject. Allele frequencies are typically expressed as a ratio or percentage.

[0062] Amplify: As used herein, “amplify” or “amplification” in the context of a nucleic acid refers to the production of multiple copies of a polynucleotide or a portion of a polynucleotide, typically starting from a small amount of polynucleotide (e.g., a single polynucleotide molecule) where the amplification product or amplicon is generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes.

[0063] Barcode: As used herein, “barcode” or “molecular barcode” in the context of a nucleic acid refers to a nucleic acid molecule that includes a sequence that can serve as a molecular identifier. For example, individual “barcode” sequences are typically added to each DNA fragment during next-generation sequencing (NGS) library preparation, and can thus be used to identify each sequencing read and sort prior to final data analysis.

[0064] Cancer type: As used herein, "cancer", "cancer type" or "tumor type" refers to a type or subtype of cancer, defined, for example, by histopathology. Cancer types can be defined according to any conventional criteria, for example, based on their occurrence in a given tissue (e.g., blood cancer, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nasal cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, colorectal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer, solid state cancer, heterogeneous cancer, homogeneous cancer), unknown primary and the like, and / or based on the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma or glioblastoma) and / or cancer indicated by cancer markers such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors and NMP-22. Cancer can also be classified by stage (e.g., stage 1, 2, 3 or 4) and whether it is primary or secondary.

[0065] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acids that are not contained within cells or otherwise bound, or in some embodiments, nucleic acids that naturally remain in a sample after removal of intact cells. Cell-free nucleic acids can include, for example, all unencapsulated nucleic acids sourced from a subject's body fluid (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or any fragments thereof. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids by secretion or cell death processes, such as necrosis, apoptosis, or the like. There is also cell-free nucleic acid released from cancer cells into body fluids, such as circulating tumor DNA (ctDNA). There is also that released from healthy cells. CtDNA can be unencapsulated tumor-derived fragmented DNA. Another example of cell-free nucleic acid is fetal DNA freely circulating in maternal blood, which is also called cell-free fetal DNA (cffDNA). Cell-free nucleic acids may have one or more epigenetic modifications, for example, cell-free nucleic acids may be acetylated, 5-methylated, ubiquitinated, phosphorylated, SUMOylated, ribosylated, and / or citrullinated.

[0066] Comparator result: As used herein, "comparator result" means a result or set of results that can be compared to identify one or more likely characteristics of a given test sample or test result, and / or one or more possible prognostic prediction results, and / or one or more customized therapies for a subject from whom the test sample was taken or otherwise derived. Comparator results are typically obtained from a set of reference samples (e.g., from subjects having the same disease or cancer type as the test subject, and / or from subjects who have received or have previously received the same therapy as the test subject). In certain embodiments, for example, the microsatellite instability status of a sample (e.g., an unstable cfDNA sample) is compared to a comparator result to identify a substantial match with the microsatellite instability status determined for the cfDNA test sample and the set of reference samples. The microsatellite instability scores determined for the set of reference samples are typically indexed with one or more customized therapies. Thus, when a substantial match is identified, the corresponding customized therapy is thereby also identified as a possible treatment pathway for the subject from whom the test sample was taken.

[0067] Control sample: As used herein, "control sample" or "control DNA sample" refers to a sample of known composition and / or having known characteristics and / or known parameters (e.g., known tumor fraction, known coverage, known microsatellite instability score, and / or the like) that is analyzed with or compared to a test sample to evaluate the accuracy of an analytical procedure.

[0068] Coverage: As used herein, "coverage" refers to the number of nucleic acid molecules occupying a particular base position.

[0069] Customized therapy: As used herein, "customized therapy" refers to a therapy with a desired treatment outcome for a subject or population of subjects selected based on a given criterion, such as having a given microsatellite instability status or being within a defined range of microsatellite instability scores.

[0070] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to natural or modified nucleotides having a hydrogen group at the 2'-position of the sugar moiety. DNA typically comprises a chain of nucleotides containing four nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to natural or modified nucleotides having a hydroxyl group at the 2'-position of the sugar moiety. RNA typically comprises a chain of nucleotides containing four nucleotides: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to natural or modified nucleotides. Certain pairs of nucleotides specifically bind to each other in a complementary fashion (this is called complementary base pairing). In the case of DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In the case of RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides that are complementary to the nucleotides in this first strand, these two strands bind to form a double strand. As used herein, "nucleic acid sequencing data", "nucleic acid sequencing information", "sequence information", "nucleic acid sequence", "nucleotide sequence", "genome sequence", "gene sequence", or "fragment sequence", or "nucleic acid sequencing read" means any information or data indicating the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a nucleic acid molecule such as DNA or RNA (e.g., whole genome, whole transcriptome, exosome, oligonucleotide, polynucleotide, or fragment).It should be understood that the present teachings contemplate sequence information obtained using all available types of techniques, platforms or technologies including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion or pH-based detection systems, and electronic signature-based systems.

[0071] Immunotherapy: As used herein, "immunotherapy" refers to treatment with one or more agents that stimulate the immune system to kill cancer cells or at least inhibit the growth of cancer cells, preferably to reduce further growth of cancer, shrink the size of cancer and / or eliminate cancer, such that it acts to reduce the size of cancer and / or eliminate cancer. Some such agents bind to targets present on cancer cells; some bind to targets present on immune cells but not to targets on cancer cells; some bind to targets present on both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of pathways of the immune system that maintain self-tolerance and modulate the duration and amplitude of physiological immune responses in peripheral tissues to minimize attendant tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Exemplary agents include antibodies against any of PD-1, PD -2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40 or CD47. Other exemplary agents include pro-inflammatory cytokines such as IL-1β, IL-6, and TNF-α. Other exemplary agents are T cells activated against a tumor, such as T cells activated by expressing a chimeric antigen that targets a tumor antigen recognized by the T cells.

[0072] Indel: As used herein, "Indel" refers to a mutation that includes an insertion or deletion of one or more nucleotides in a target genome.

[0073] Indexed: As used herein, "indexed" refers to a first element (e.g., a microsatellite instability score) that is associated with a second element (e.g., a given therapy).

[0074] Instability status: As used herein, "instability status" or "instability score" in the context of repetitive nucleic acids (e.g., repetitive nucleic acid / repetitive DNA instability status or score, microsatellite instability status or score) refers to a measure or determination as to whether a given repetitive nucleic acid locus or population of repetitive nucleic acid loci in one or more nucleic acid samples exhibits a level or degree of mutation (e.g., variable repeat length) that is higher than a threshold level determined for that locus or population of loci, is at that threshold level of mutation (e.g., variable repeat length), or is lower than that threshold level of mutation (e.g., variable repeat length). For clarity, instability status and instability score are not interchangeable, but rather related concepts. Instability status is based on the instability score. For example, if the instability score of a sample is at or below a population-trained threshold, the sample is classified as a stable sample (e.g., for MSI - MSS or low MSI), and if the instability score of the sample is higher than the population-trained threshold, the sample is classified as an unstable sample (e.g., for MSI - high MSI).

[0075] Limit of detection (LoD): As used herein, "limit of detection" or "LoD" means the minimum amount of a substance (e.g., nucleic acid) in a sample that can be measured by a given assay or analytical technique.

[0076] Maximum MAF: As used herein, "maximum MAF" or "max MAF" refers to the maximum MAF of all somatic variants in a sample.

[0077] Microsatellite: As used herein, "microsatellite" refers to a repetitive nucleic acid having repeat units of about 10 base pairs or less in length.

[0078] Minisatellite: As used herein, "minisatellite" refers to a repetitive nucleic acid having repeat units of about 10 to about 60 base pairs or nucleotides in length.

[0079] Mutant allele fraction: As used herein, "mutant allele fraction", "mutation dose", or "MAF" refers to the fraction of nucleic acid molecules carrying an allelic change or mutation at a given genomic position. MAF is generally expressed as a fraction or percentage. For example, MAF is typically less than about 0.5, 0.1, 0.05 or 0.01 (i.e., less than about 50%, 10%, 5% or 1%) of all somatic variants or alleles present at a given locus.

[0080] Mutation: As used herein, "mutation" refers to a variation from a known reference sequence, and includes, for example, single nucleotide variants (SNVs), copy number variants or mutations (CNVs) / abnormalities, insertions or deletions (indels), gene fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants. Mutations may be germline mutations or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the species providing the test sample, typically the human genome.

[0081] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used synonymously. They refer to the abnormal growth of cells in a subject. A neoplasm or tumor may be benign, potentially malignant, or malignant. Malignant tumors are referred to as cancers or cancerous tumors.

[0082] Next-generation sequencing: As used herein, "next-generation sequencing" or "NGS" refers to sequencing technologies that have increased throughput compared to traditional Sanger-based and capillary electrophoresis-based methods, e.g., the ability to simultaneously generate hundreds of thousands of relatively short sequence reads. Some examples of next-generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.

[0083] Nucleic acid tag: As used herein, a "nucleic acid tag" is a short nucleic acid (e.g., about 500 nucleotides in length, about 100 nucleotides, about 50 nucleotides, or less than about 10 nucleotides), a short nucleic acid used to distinguish nucleic acids from different samples, of different types or that have undergone different treatments (e.g., representing a sample index), or a short nucleic acid used to distinguish different nucleic acids of different types or that have undergone different treatments within the same sample (e.g., representing a molecular barcode). Nucleic acid tags include a predetermined, fixed, non-random, random, or semi-random oligonucleotide sequence. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or secondary samples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or a diverse length as needed. Nucleic acid tags may include double-stranded molecules having one or more blunt ends, may include 5' or 3' single-stranded regions (e.g., overhangs), and / or may include one or more other single-stranded regions in other regions of a given molecule. A nucleic acid tag can be attached to one or both ends of another nucleic acid (e.g., the sample nucleic acid to be amplified and / or sequenced). The nucleic acid tag can be decoded to reveal information such as the origin sample, form, or treatment of a given nucleic acid. For example, nucleic acid tags can be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indexes, and these nucleic acids are then deconvolved by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags may also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally, or alternatively, nucleic acid tags can be used as molecular barcodes (e.g., to distinguish different molecules or amplicons of different parental molecules within the same sample or secondary sample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample or non-uniquely tagging such molecules.In the case of applying non-unique tagging, a limited number of tags (i.e., molecular barcodes) can be used to tag nucleic acid molecules, such that different molecules can be distinguished based on their endogenous sequence information (e.g., their start and / or end positions within a selected reference genome, partial sequences at one or both ends of the sequence, and / or the length of the sequence), in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used such that the probability that any two molecules can have the same endogenous sequence information (e.g., start and / or end positions, partial sequences at one or both ends of the sequence, and / or length) and also have the same molecular barcode is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1%).

[0084] Polynucleotide: As used herein, "polynucleotide", "nucleic acid", "nucleic acid molecule", or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) linked by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomer units, e.g., 3 - 4, to several hundred monomer units. When a polynucleotide is represented by a sequence of letters such as "ATGCCTG", it is always understood, unless otherwise specified, that the nucleotides are in the 5'→3' order from left to right, and in the case of DNA, that "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine. The letters A, C, G, and T are used, as is common in the art, sometimes to refer to the bases themselves, sometimes to refer to the nucleosides, and sometimes to refer to the nucleotides containing the bases.

[0085] Population-trained threshold: As used herein, the "population-trained threshold" in the context of repetitive nucleic acids refers to the separately determined maximum total number of unstable repetitive nucleic acid loci (e.g., a number of unstable microsatellite loci) expected to be observed in training DNA samples (e.g., non-tumor samples, tumor samples, etc.) that contain those loci. The population-trained threshold is typically used to characterize the experimentally determined repetitive nucleic acid instability score for a particular sample.

[0086] Processing: As used herein, the terms "processing", "calculating", and "comparing" are used synonymously. In certain embodiments, these terms refer to determining a difference, e.g., a difference in numbers or sequences. For example, one can process the values or sequences of repetitive DNA instability scores (e.g., microsatellite instability scores), gene expression, copy number variations (CNVs), indels, and / or single nucleotide variants (SNVs).

[0087] Reference sequence: As used herein, a "reference sequence" refers to a known sequence used for the purpose of comparison with an experimentally determined sequence. For example, the known sequence may be a whole genome, a chromosome, or any segment thereof. A reference sequence typically includes at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more nucleotides. A reference sequence can align with a single continuous sequence of a genome or chromosome, or can include discontinuous segments that align with different regions of a genome or chromosome. Exemplary reference sequences include, for example, the human genomes such as hG19 and hG38.

[0088] Repeat length: As used herein, in the context of repetitive nucleic acids, "repeat length" refers to the number of repeat units present at a given repetitive nucleic acid locus. By way of example, the following single-stranded nucleic acid strand has a repeat length of 8:

Chem.

[0089] Repeat unit: As used herein, in the context of repetitive nucleic acids, "repeat unit" refers to an individual nucleotide pattern or motif (e.g., homopolymer or heteropolymer) that is repeated at a given repetitive nucleic acid locus. By way of example, the following single-stranded nucleic acid strand has a repeat unit of "ATT":

Chem.

[0090] Repetitive nucleic acid: As used herein, "repetitive nucleic acid" or "repetitive element" refers to a pattern in which nucleotides repeatedly occur in multiple copies across a given genome and / or across an entire population of genomes. Repetitive nucleic acids include repetitive DNA and repetitive RNA. Non-limiting examples of repetitive nucleic acids include microsatellites, terminal repeats, tandem repeats, minisatellites, satellite DNA, interspersed repeats, transposable elements (e.g., DNA transposons, retrotransposons (e.g., LTR-type retrotransposons (HERV) and LTR-type retrotransposons (HERV)), etc.), palindromic repeats (CRISPR) that form clusters and have regular spacing, direct repeats, inverted repeats, mirror repeats, and inverted repeats.

[0091] Repeat nucleic acid instability score: As used herein, the "repeat nucleic acid instability score" (e.g., repeat DNA instability score, microsatellite instability score, etc.) in the context of repeat nucleic acids refers to the total number of repeat nucleic acid loci from a population of repeat nucleic acid loci in a given sample that are called unstable or otherwise determined. This repeat nucleic acid instability score is a sample-level score (or sample score) and is different from a site score that is locus-specific.

[0092] Sample: As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.

[0093] Sensitivity: As used herein, "sensitivity" means the probability of detecting the presence of a mutation at a given MAF and coverage.

[0094] Sequencing: As used herein, "sequencing" refers to any of a number of techniques used to determine the sequence (e.g., the identity and order of monomer units) of a biomolecule, such as a nucleic acid, e.g., DNA or RNA. Exemplary sequencing methods include targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, sequencing based on electron microscopy, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxynucleotide termination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single nucleotide extension sequencing, solid phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR at lower denaturation temperature (COLD-PCR), multiplex PCR, sequencing by reversible dye terminators, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof, but are not limited thereto. In some embodiments, sequencing can be performed by a genetic analysis apparatus, such as, for example, a genetic analysis apparatus commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among others.

[0095] Array information: As used herein, "array information" in the context of a nucleic acid polymer refers to the order and identity of monomer units (e.g., nucleotides, etc.) in that polymer.

[0096] Site score: As used herein, "site score" refers to a measure of the likelihood of the presence of additional repeat lengths other than the germline repeat length at a given repetitive nucleic acid locus in a sample. In certain embodiments, the site score is determined for a given locus by calculating the delta Akaike information criterion (ΔAIC) for that locus.

[0097] Site-specific trained threshold: As used herein, "site-specific trained threshold" refers to a separately determined maximum value of the site score for a given repetitive nucleic acid locus (e.g., a given microsatellite locus) such that the locus is stable.

[0098] Somatic mutation: As used herein, "somatic mutation" means a mutation of the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.

[0099] Specificity: As used herein, "specificity" in the context of a diagnostic assay or analysis refers to the degree to which the assay or analysis detects the intended target analyte rather than other components of a given sample.

[0100] Substantial match: As used herein, "substantial match" means that at least a first value or element is at least approximately equal to at least a second value or element. In certain embodiments, for example, a customized therapy is identified when the microsatellite instability score and the comparator result are at least substantially or approximately matched.

[0101] Subject: As used herein, "subject" refers to an animal, such as a mammalian species (e.g., human) or an avian (e.g., bird) species, or other organisms, such as plants. More specifically, the subject can be a vertebrate, such as a mammal, such as a mouse, primate, monkey or human. Animals include domestic animals (e.g., beef cattle, dairy cows, poultry, horses, pigs and the like), sport animals, and companion animals (e.g., pets or service animals). The subject may be a healthy individual, or an individual having a disease or predisposed to a disease, or suspected of having a disease or being predisposed to a disease, or an individual in need of treatment or suspected of being in need of treatment. The terms "individual" or "patient" are intended to be interchangeable with "subject".

[0102] For example, the subject may be an individual diagnosed with having cancer, scheduled to receive cancer therapy, and / or having received at least one type of cancer therapy. The subject may be in remission from cancer. As another example, the subject may be an individual diagnosed with having an autoimmune disease. As another example, the subject may be a female individual who is pregnant or planning to become pregnant and who has been diagnosed with or suspected of having a disease, such as cancer, autoimmune disease.

[0103] Threshold: As used herein, "threshold" refers to a separately determined value used to characterize or classify an experimentally determined value.

[0104] Training DNA sample: As used herein, "training DNA sample" refers to a DNA sample used for the estimation of site-specific trained thresholds and population-trained thresholds. A training DNA sample dataset includes one or more training DNA samples. A training DNA sample includes one or more normal DNA samples and / or tumor DNA samples. In some embodiments, the training DNA sample includes one or more samples having high MSI and / or low MSI / MSS status.

[0105] Tumor fraction: As used herein, "tumor fraction" refers to an estimate of the proportion of nucleic acid molecules derived from tumor in a given sample. For example, the tumor fraction of a sample may be a measure derived from the max MAF of the sample or the coverage of the sample or the length of cfDNA fragments in the sample or any other selected feature of the sample. In some embodiments, the tumor fraction of a sample is equal to the max MAF of the sample.

[0106] Unstable: As used herein, "unstable" or "instability" in the context of repetitive nucleic acids refers to the level of mutations (e.g., indels or the like) observed at a given repetitive nucleic acid locus in a nucleic acid sample (e.g., a cfDNA sample) or in a given population of repetitive nucleic acid loci that exceeds a threshold (e.g., a site-specific trained threshold - locus level; a population-trained threshold - sample level; or the like).

[0107] Detailed Description Introduction

[0108] Cancer encompasses a large group of genetic diseases that share the common feature of abnormal cell growth and have the potential to metastasize beyond the site of origin of cells in the body. The molecular basis underlying this disease consists of mutations and / or epigenetic changes that lead to the transformed cell phenotype, whether these deleterious changes are genetically acquired or have a somatic basis. Unfortunately, these molecular changes are typically different not only between patients with the same cancer type but even within the tumors of a given patient themselves.

[0109] Given the variability of mutations observed in most cancers, one challenge in cancer treatment is to identify the therapy to which a patient is most likely to respond, based on the patient's individual cancer type. A variety of biomarkers are used to select the appropriate treatment for a cancer patient, including cancer immunotherapy. One biomarker of response is microsatellite instability (MSI), which is a state of genomic hypermutability caused by a reduction in the DNA mismatch repair (MMR) mechanism or a predisposition to such genomic hypermutability. Cancer patients with microsatellite instability classified as high (MSI-H or high MSI) often exhibit the accumulation of somatic mutations in tumor cells, resulting in a range of molecular and biological changes, including high tumor mutation burden, increased neoantigen expression, and a large number of tumor-infiltrating lymphocytes. Chang et al., "Microsatellite Instability: A Predictive Biomarker for Cancer Immunotherapy," Appl Immunohistochem Mol Morphol, 26(2):e15-e21 (2018). These changes have been associated with increased sensitivity to checkpoint inhibitors such as pembrolizumab (Keytruda®) used to treat advanced melanoma, head and neck squamous cell carcinoma, non-small cell lung cancer (NSCLC), and classical Hodgkin lymphoma. To date, the application of this response biomarker has been essentially limited to the assessment of MSI status in solid tumor samples using standard PCR-based techniques.

[0110] The present disclosure provides methods, computer-readable media, and systems that are useful for the determination and analysis of MSI in patient samples, particularly cell-free DNA (cfDNA) samples. The MSI status determined using these methods and related aspects can help guide disease prognosis and treatment decisions. The results achieved with the methods and related aspects disclosed herein generally have a high degree of agreement with the results obtained using more conventional PCR-based MSI assessment techniques, for example.

[0111] Method for determining microsatellite instability status

[0112] This application discloses various methods for accurately determining the microsatellite instability (MSI) status and / or other repetitive DNA instability status of a sample (particularly, a cell-free DNA (cfDNA) sample). In certain embodiments, the method for assessing the MSI status enables broad coverage of simple repeats where microsatellite instability can exist across a wide range of cancer types and includes, for example, targeted sequencing of cfDNA using a digital sequencing platform from Guardant Health, Inc. (Redwood City, CA, USA). The digital sequencing platform is an NGS panel of cancer-related genes that utilizes high-quality sequencing of cell-free DNA isolated by simple non-invasive blood sampling (which may include circulating tumor DNA). Digital sequencing utilizes pre-sequencing preparation of a digital library of individually tagged cfDNA molecules, combined with post-sequencing bioinformatics reconstruction, to eliminate almost all false positives. By way of example, FIG. 1 provides a flowchart schematically depicting exemplary method steps for determining the MSI status according to some embodiments of the present invention. As shown, method 100, in step 110, quantifies the number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci. The sequence information is typically obtained from a population of microsatellite loci in the cfDNA sample. As further described herein, the number of different repeat lengths present at a given microsatellite locus is quantified in some embodiments using a site score based on probabilistic log-likelihood. Also as further described herein, other quantification techniques may be utilized as needed if they can very accurately distinguish a relatively small number of cfDNA fragments of somatic origin in the sample from noise resulting from artifacts (e.g., amplification artifacts, sequencing artifacts, and the like) after sample collection.

[0113] Method 100 also includes, at step 112, comparing the site score of a given microsatellite locus for each of the particular microsatellite loci to the site-specific trained threshold for that particular microsatellite locus. The experimentally determined site score for a particular locus and its corresponding site-specific trained threshold are typically compared for each of the plurality of microsatellite loci. The site-specific trained threshold for a given locus is generally a predetermined value for that particular locus derived from a population of training DNA samples such as a cohort of normal or non-tumor cfDNA samples. As shown, method 100 includes, at step 114, calling a given microsatellite locus as unstable if the site score (e.g., likelihood score or the like) of the given microsatellite locus exceeds (e.g., is statistically greater than) the site-specific trained threshold of the given microsatellite locus. Based on these comparisons, a microsatellite instability score is generated, which score includes the number of microsatellite loci called as unstable from the plurality of microsatellite loci (e.g., the overall or total MSI score of the sample). Additionally, method 100 includes, at step 116, classifying the MSI status of the cfDNA sample as unstable if the microsatellite instability score exceeds the population-trained threshold of the population of microsatellite loci in the cfDNA sample, thereby identifying an unstable cfDNA sample (e.g., scoring or predicting the sample as being high MSI). In other words, the MSI status of the sample is determined, in certain embodiments, by the presence of a minimum number of unstable microsatellite loci. The population-trained threshold is generally a predetermined value derived from a population of training DNA samples such as a cohort of normal or non-tumor cfDNA samples.

[0114] In some embodiments, a threshold (e.g., a site-specific trained threshold, a population-trained threshold, and the like) is determined from or otherwise derived from at least one training DNA sample dataset. A training DNA sample dataset typically includes at least about 25 to at least about 30,000 or more training samples. In some embodiments, a training DNA sample dataset includes about 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000, or more training DNA samples.

[0115] In certain embodiments, method 100 includes additional upstream and / or downstream steps. In some embodiments, for example, method 100 begins, at step 102, by providing a sample from the subject at step 104 (e.g., providing a blood sample taken from the subject). In these embodiments, the workflow of method 100 typically also includes, at step 106, amplifying the nucleic acids in the sample to produce amplified nucleic acids, and, at step 108, sequencing the amplified nucleic acids to generate sequence information, and then, at step 110, quantifying the number of different repeat lengths present at each of a plurality of microsatellite loci from the sequence information. Nucleic acid amplification (including related sample preparation), nucleic acid sequencing, and related data analysis are further described herein.

[0116] In some embodiments, method 100 includes various steps downstream from the identification of an unstable cfDNA sample at step 116. Some examples of these include comparing the microsatellite instability status of the cfDNA sample at step 118 to comparator results indexed by therapy to identify a customized therapy for treating a disease (e.g., cancer or another genetic disease, disorder, or condition) in a subject. In other exemplary embodiments, method 100 includes administering to the subject at least one of the identified customized therapies if, prior to ending at step 122 (e.g., for treating the subject's cancer or another disease, disorder, or condition), the microsatellite instability status of the sample and the comparator results substantially match at step 120.

[0117] The methods described herein include various alternative embodiments. For example, the locus score of a microsatellite locus optionally includes a likelihood score. In some embodiments, the likelihood score includes a score based on a probabilistic log-likelihood. In some of these embodiments, the method uses various parameters such as allele frequencies or one or more error modes (e.g., random error mode, strand-specific error mode, and / or the like) to determine a score based on the probabilistic log-likelihood for an individual microsatellite locus in the sequence information obtained from a sample. Allele frequencies generally include the observed frequencies of nucleic acids with different repeat lengths at a given microsatellite locus in the sequence information obtained from the sample. In some embodiments, the locus score for a particular microsatellite locus includes the difference or ratio between (a) a score measuring the support of the observed nucleic acid sequence for the null hypothesis that a given microsatellite locus is stable and (b) a score measuring the support of the observed nucleic acid sequence for the alternative hypothesis that a given microsatellite locus is unstable. The null hypothesis is the hypothesis with the minimum AIC score among all hypotheses under the assumption that the locus is stable, and the alternative hypothesis is the hypothesis with the minimum AIC score among all hypotheses under the assumption that the locus is unstable. Typically, the locus score is generated using various measures of model accuracy such as likelihood criteria, log-likelihood criteria, posterior probability criteria, Akaike information criterion (AIC), Bayesian information criterion, and / or the like. Further details regarding statistical modeling, including measures of statistical model accuracy adapted as needed for use in performing the methods disclosed herein, can be found, for example, in Bruce, Practical Statistics for Data Scientists: 50 Essential Concepts, 1 st Ed., O'Reilly Media (2017), Freedman et al., Statistics, 4 thEd., W.W. Norton & Company (2007), James et al., An Introduction to Statistical Learning: with Applications in R, 1 st Ed., Springer (2013), and Hastie et al., The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2 nd are provided in Ed., Springer (2016), each of which is hereby incorporated by reference in its entirety.

[0118] To further illustrate with examples, the site score of a given microsatellite locus optionally includes the AIC-based site score for testing somatic indels at that microsatellite locus. In some embodiments of these embodiments, the AIC-based site score for a given one is given by the following formula: AIC = k - log likelihood (where k is the number of parameters used in the model) is calculated using. In some embodiments, the method includes the step of estimating the parameters of the model using maximum likelihood estimation (MLE) (e.g., using the Nelder-Mead algorithm or another simplex search algorithm). The method is given by the following formula: AIC0 = k - log(Pr(obs|β,γ)) (where AIC0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs is a number of observed sequencing reads covering a given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter) optionally includes the step of calculating the null hypothesis of the model using. In some embodiments of these embodiments, the method is given by the following formula: AIC min= min α (k - log(Pr(obs|β, γ, α)) (where AIC min is the alternative hypothesis, min α is the minimization effect over all values of α, k is the number of parameters used in the model, Pr is the probability, obs is the number of observed sequencing reads covering a given microsatellite locus, β is at least one strand - specific error parameter, γ is at least one random error parameter, α is at least one allele frequency, where α is a vector of allele frequencies such that the sum of one or more α i equals 1) includes the step of calculating the alternative hypothesis of the model using. The change in the model (ΔAIC) used to determine the site score is given by the following equation: ΔAIC = AIC0 - AIC min is typically detected using.

[0119] In certain embodiments, the parameter γ includes (a) the read - level error rate where the microsatellite length observed within a sequencing read is 1 repeat unit longer than the expected microsatellite length for the strand of the originating nucleic acid molecule; and / or (b) the read - level error rate where the microsatellite length observed within a sequencing read is 1 repeat unit shorter than the expected microsatellite length for the strand of the originating nucleic acid molecule. In some embodiments, the parameter β includes (a) the strand - level error rate where the expected microsatellite length of the sense strand is 1 repeat unit longer than the expected microsatellite length of the nucleic acid - derived molecule, (b) the strand - level error rate where the expected microsatellite length of the antisense strand is 1 repeat unit longer than the expected microsatellite length of the nucleic acid - derived molecule, (c) the strand - level error rate where the expected microsatellite length of the sense strand is 1 repeat unit shorter than the expected microsatellite length of the nucleic acid - derived molecule, and / or (d) the strand - level error rate where the expected microsatellite length of the antisense strand is 1 repeat unit shorter than the expected microsatellite length of the nucleic acid - derived molecule.

[0120] In some embodiments, the site score based on AIC is given by the following formula: AIC = 2(k - log likelihood) (where k is the number of parameters used in the model) is used for the calculation. In these embodiments, AIC0 and AIC min are calculated using the above formula.

[0121] For clarity, in embodiments where the score based on AIC is determined using the formula AIC = 2(k - log likelihood), the site-specific threshold used to classify a site as unstable is twice the site-specific threshold used in the previous embodiments where the score based on AIC is determined using the formula AIC = k - log likelihood.

[0122] Samples analyzed using the methods described herein typically include various mutant allele frequencies (MAFs) (e.g., sample fractions showing different repeat lengths, specific microsatellite loci, or other allelic variations). Additionally, the samples include tumor fractions in some embodiments. In certain embodiments, the maximum MAF (max MAF) serves as an approximation of the tumor fraction in a given sample. The tumor fraction is typically less than about 0.05%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, or about 15% of all nucleic acids in the sample.

[0123] In some embodiments, the methods disclosed herein typically include at least about 94% sensitivity at a limit of detection (LOD) of about 0.2% tumor fraction of nucleic acids in a given sample. The methods generally also have at least about 99% specificity for non-tumor DNA in the sample. The MSI status determined for a sample typically also has at least about 95%, 96%, 97%, 98% or 99% concordance with the corresponding MSI status of the sample determined using standard PCR-based MSI assessment techniques over a tumor fraction range of about 1.4% to about 15%. In some embodiments, this concordance is 100%.

[0124] In certain embodiments, the MSI status of a particular sample is classified as high MSI (MSI-H) if the microsatellite instability score for the sample is the number of unstable microsatellite loci of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, or more than 100 in that sample. In certain embodiments, the population-trained threshold used to determine the instability status (e.g., MSI status) of a sample is the number of unstable repetitive nucleic acid (e.g., microsatellite) loci of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, or more than 100. In some embodiments, the population-trained threshold of a sample is about 5 unstable microsatellite loci. In some embodiments, the population-trained threshold of a sample is about 6 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 10 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 15 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 16 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 20 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 25 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 26 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 30 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 35 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 36 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 40 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold of a sample is about 45 unstable repetitive nucleic acid loci.In some embodiments, the population-trained threshold for the sample is about 46 unstable repetitive nucleic acid loci. In some embodiments, the population-trained threshold for the sample is about 50 unstable repetitive nucleic acid loci. In some embodiments, the repetitive nucleic acid locus can be a microsatellite locus. In some embodiments, the MSI status of a given sample is classified as MSI-H if the number of unstable microsatellite loci constitutes about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20%, or about 25% of all the microsatellite loci evaluated in that sample. In some embodiments, about 50, about 60, about 70, about 80, about 90, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1100, about 1200, about 1300, about 1400, about 1500, about 1600, about 1700, about 1800, about 1900, about 2000, or more than 2000 repetitive nucleic acid (e.g., microsatellite) loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 50 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 60 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 70 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 80 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 90 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 100 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 200 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 300 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample.In some embodiments, about 400 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 500 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1000 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1100 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1200 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1300 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1400 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, at least 1500 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, about 1600 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, at least 1700 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, at least 1800 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, at least 1900 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, at least 2000 repetitive nucleic acid loci are used to determine the repetitive nucleic acid instability (e.g., MSI) status of a given sample. In some embodiments, the repetitive nucleic acid loci can be microsatellite loci. In some embodiments, the repetitive nucleic acid instability status can be an MSI status.

[0125] In some embodiments, the method includes obtaining a sample from a subject. Essentially, any sample type can be utilized as needed. In certain embodiments, for example, the sample is tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, feces, synovial fluid, cerebrospinal fluid, saliva, and / or the like. Additional exemplary sample types that can be utilized as needed are further described herein. Typically, the subject is a mammalian subject (e.g., a human subject). Essentially any type of nucleic acid (e.g., DNA and / or RNA) can be evaluated according to the methods disclosed in the present application. Some examples include cell-free nucleic acids (e.g., tumor-derived, fetal-derived, maternal-fetal-derived, and / or the like, such as cfDNA), nucleic acids of cells including circulating tumor cells (e.g., those obtained by lysing intact cells in the sample), circulating tumor nucleic acids, and the like. In some embodiments, the sample includes cell-free DNA (a cfDNA sample). In some embodiments, the cfDNA sample includes circulating tumor nucleic acids.

[0126] The methods disclosed herein generally include obtaining sequence information from nucleic acids in a sample collected from a subject. In certain embodiments, the sequence information is obtained from targeted segments of nucleic acids. Essentially any number of genomic regions can be targeted as needed. The targeted segments can include at least 10, at least 50, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 20,000 or at least 50,000 (e.g., 25, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 25,000, 30,000, 35,000, 40,000, 45,000) different or overlapping genomic regions. In some embodiments, the targeted segments include a selected region of at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, or at least 700 genes. In some embodiments, the targeted segments include a selected region of at least 70 genes. In some embodiments, the targeted segments include a region of at least 500 genes.

[0127] In these embodiments, the method typically includes various sample or library preparation steps for preparing nucleic acids for sequencing. Many different sample preparation techniques are well known to those of ordinary skill in the art. Essentially any of these techniques can be used or adapted for use in performing the methods described herein. For example, in addition to various purification steps for isolating nucleic acids in a given sample from other components, typical steps for preparing nucleic acids for sequencing include tagging the nucleic acids with molecular identifiers or barcodes, adding adapters (which may, for example, include barcodes), amplifying the nucleic acids one or more times, enriching targeted segments of the nucleic acids (e.g., using various target capture strategies, etc.), and / or the like. Exemplary library preparation processes are further described herein. Further details regarding nucleic acid sample / library preparation are described, for example, in vanDijk et al., Library preparation methods for next-generationsequencing: Tonedown the bias, Experimental Cell Research, 322(1):12-20 (2014), Micic (Ed.),Sample Preparation Techniques for Soil, Plant, and Animal Samples (SpringerProtocols Handbooks), 1 st Ed., Humana Press (2016), and Chiu,Next-Generation Sequencing and Sequence DataAnalysis, Bentham SciencePublishers (2018), each of which is hereby incorporated by reference in its entirety.

[0128] The microsatellite and / or other repetitive nucleic acid instability status determined by the methods disclosed herein is used as needed to diagnose the presence of a disease or condition, particularly cancer, in a subject, to characterize such a disease or condition (e.g., to stage a given cancer, to determine cancer heterogeneity, and for similar purposes), to monitor response to treatment, to assess the potential risk of developing a given disease or condition, and / or to assess the prognosis of a disease or condition. The microsatellite and / or other repetitive nucleic acid instability status is also used as needed to characterize specific forms of cancer. Since cancers are often heterogeneous in both composition and staging, the microsatellite and / or other repetitive nucleic acid instability status enables the characterization of specific subtypes of cancer, thereby assisting in the selection of diagnosis and treatment. This information can also provide clues regarding the prognosis of a specific type of cancer to the subject or healthcare provider, enabling the subject and / or healthcare provider to adapt treatment options as the disease progresses. Some cancers become more aggressive and genetically unstable as they progress. Other tumors remain benign, inactive, or in a quiescent state.

[0129] Microsatellites and / or other repetitive nucleic acid instability status may also be useful in determining disease progression and / or monitoring recurrence. In certain cases, for example, a successful treatment may initially increase the observed microsatellite and / or other repetitive nucleic acid instability as an increase in the number of cancer cell deaths and shed nucleic acids. In these cases, subsequently, as the treatment progresses, the microsatellite and / or other repetitive nucleic acid instability will typically decrease as the tumor size continues to decrease. In other cases, a successful treatment may also be able to decrease the microsatellite and / or other repetitive nucleic acid instability without such an initial increase in instability. Additionally, if it is recognized that the cancer is in remission after treatment, the microsatellite and / or other repetitive nucleic acid instability status can be used to monitor for residual disease or recurrence of the disease in the patient.

[0130] Sample

[0131] A sample can be any biological sample isolated from a subject. Samples include body tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsy samples (e.g., biopsy samples from known or suspected solid tumors ) may include cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid (e.g., fluid from the intercellular space), gingival exudate, crevicular fluid, bone marrow, pleural exudate, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions, as well as urine. Such samples contain nucleic acids shed from tumors. The nucleic acids can include DNA and RNA and can be in double-stranded and single-stranded forms. The sample may be in the form initially isolated from the subject, or may have been subjected to further processing to remove or add components such as cells, to concentrate one component relative to another component, or to convert one form of nucleic acid to another form, e.g., RNA to DNA or single-stranded nucleic acid to double-stranded nucleic acid. Thus, for example, a body fluid sample for analysis is serum or plasma containing cell-free nucleic acids, e.g., cell-free DNA (cfDNA).

[0132] In some embodiments, the sample volume of the body fluid collected from the subject depends on the desired read depth for the region to be sequenced. Exemplary volumes are about 0.4 - 40 ml, about 5 - 20 ml, about 10 - 20 ml. For example, the volume can be about 0.5 ml, about 1 ml, about 5 ml, about 10 ml, about 20 ml, about 30 ml, about 40 ml, or more milliliters. The volume of plasma from which the sample is taken is typically between about 5 ml and about 20 ml.

[0133] Samples can contain various amounts of nucleic acids. Typically, the amount of nucleic acids in a given sample is considered equivalent to a plurality of genome equivalents. For example, a sample of about 30 ng of DNA is about 10,000 (10 4 ) diploid human genome equivalents, and in the case of cfDNA, about 200 billion (2×10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA is about 30,000 diploid human genome equivalents, and in the case of cfDNA, about 600 billion individual molecules.

[0134] In some embodiments, the sample comprises nucleic acids from different sources, e.g., from cells and from cell-free sources (e.g., blood samples, etc.). Typically, the sample comprises nucleic acids having mutations. For example, the sample optionally comprises DNA having germline mutations and / or somatic mutations. Typically, the sample comprises DNA having cancer-related mutations (e.g., cancer-related somatic mutations).

[0135] Exemplary amounts of cell-free nucleic acids in a sample prior to amplification are typically in the range of about 1 femtogram (fg) to about 1 microgram (μg), such as, for example, about 1 picogram (pg) to about 200 nanograms (ng), about 1 ng to about 100 ng, about 10 ng to about 1000 ng. In some embodiments, the sample contains cell-free nucleic acid molecules of about 600 ng or less, about 500 ng or less, about 400 ng or less, about 300 ng or less, about 200 ng or less, about 100 ng or less, about 50 ng or less, or about 20 ng or less. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the amount is about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng or less of cell-free nucleic acid molecules. In certain embodiments, the method includes obtaining cell-free nucleic acid molecules between about 5 ng and about 30 ng from the sample. In certain embodiments, the method includes obtaining cell-free nucleic acid molecules between about 5 ng and about 100 ng from the sample. In certain embodiments, the method includes obtaining cell-free nucleic acid molecules between about 5 ng and about 150 ng from the sample. In certain embodiments, the method includes obtaining cell-free nucleic acid molecules between about 5 ng and about 200 ng from the sample. In some embodiments, the amount is about 100 ng or less of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is about 150 ng or less of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is about 200 ng or less of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is about 250 ng or less of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is about 300 ng or less of cell-free nucleic acid molecules from the sample. In certain embodiments, the method includes obtaining cell-free nucleic acid molecules between about 1 fg and about 200 ng from the sample.

[0136] Cell-free nucleic acid molecules typically have a size distribution between about 100 nucleotides and about 500 nucleotides in length, with molecules between about 110 nucleotides and about 230 nucleotides in length corresponding to about 90% of the molecules in the sample, a mode of about 168 nucleotides in length, and a second minor peak within the range between about 240 and about 440 nucleotides in length. In certain embodiments, the cell-free nucleic acid is about 160 to about 180 nucleotides in length, or about 320 to about 360 nucleotides in length, or about 440 to about 480 nucleotides in length.

[0137] In some embodiments, cell-free nucleic acid is isolated from a body fluid by a partitioning step in which cell-free nucleic acid as found in solution is separated from intact cells and other insoluble components of the body fluid. In some of these embodiments, partitioning includes techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid are lysed and the cell-free nucleic acid and cellular nucleic acid are processed together. Generally, after addition of a buffer and a washing step, the cell-free nucleic acid is precipitated, for example, with alcohol. In certain embodiments, additional purification steps, such as a silica-based column to remove contaminants or salts, are used. For example, non-specific bulk carrier nucleic acid is added as needed throughout the reaction to optimize certain aspects of exemplary procedures, such as yield. After such processing, the sample typically contains nucleic acids in various forms, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. Optionally, single-stranded DNA and / or single-stranded RNA are converted to double-stranded form for inclusion in subsequent processing and analysis steps.

[0138] Nucleic acid tag

[0139] In some embodiments, nucleic acid molecules (from a sample of polynucleotides) can be tagged with a sample index and / or a molecular barcode (commonly referred to as a "tag"). Tags can be incorporated into adapters or otherwise attached by, among other ways, chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR). Ultimately, such adapters can be attached to target nucleic acid molecules. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are applied to introduce a sample index into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., multiple microwells in the form of an array). Molecular barcodes and / or sample indexes can be introduced simultaneously or in any order. In some embodiments, molecular barcodes and / or sample indexes are introduced before and / or after a sequence capture step is performed. In some embodiments, only the molecular barcode is introduced prior to probe capture and the sample index is introduced after the sequence capture step is performed. In some embodiments, both the molecular barcode and the sample index are introduced prior to the performance of the probe-based capture step. In some embodiments, the sample index is introduced after the sequence capture step is performed. In some embodiments, the molecular barcode is incorporated into nucleic acid molecules (e.g., cfDNA molecules) in a sample via an adapter by ligation (e.g., blunt-end ligation or sticky-end ligation). In some embodiments, the sample index is incorporated into nucleic acid molecules (e.g., cfDNA molecules) in a sample by overlap extension polymerase chain reaction (PCR). Typically, a sequence capture protocol involves introducing single-stranded nucleic acid molecules complementary to a targeted nucleic acid sequence, e.g., the coding sequence of a genomic region, where mutations in such regions are associated with cancer types.

[0140] In some embodiments, the tag may be located at one end of the sample nucleic acid molecule or at both ends. In some embodiments, the tag is a predetermined or random or semi-random sequence oligonucleotide. In some embodiments, the tag is less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 nucleotide in length. The tag can be ligated to the sample nucleic acid either randomly or deliberately.

[0141] In some embodiments, each sample is uniquely tagged with a sample index or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or a secondary sample is uniquely tagged with a molecular barcode or a combination of molecular barcodes. In other embodiments, multiple molecular barcodes (e.g., non-unique molecular barcodes) may be used such that the molecular barcodes among the multiple molecular barcodes are not necessarily unique to each other. In these embodiments, the molecular barcode is generally attached to an individual molecule (e.g., by ligation), resulting in a unique sequence that can be individually tracked by the combination of the molecular barcode and the sequence to which it can be attached. Detection of non-uniquely tagged molecular barcodes in combination with endogenous sequence information (e.g., the beginning (start) and / or end (termination) portions corresponding to the sequence of the original nucleic acid molecule in the sample, partial sequences at one or both ends of the sequence read, the length of the sequence read, and / or the length of the original nucleic acid molecule in the sample) enables the identification of the unique identity of a particular molecule. The length of an individual sequence read, or the number of base pairs of an individual sequence read, may also be used as needed to provide a unique identity to a given molecule. As described herein, a fragment from a single strand of a nucleic acid with a unique identity can thereby enable subsequent identification of fragments from the parental strand and / or complementary strand.

[0142] In some embodiments, molecular barcodes are introduced into the molecules in a sample at an expected ratio of a set of identifiers (e.g., a combination of unique and non-unique molecular barcodes). One example format uses from about 2 to about 1,000,000 different molecular barcodes, or from about 5 to about 150 different molecular barcodes, or from about 20 to about 50 different molecular barcodes. Alternatively, from about 25 to about 1,000,000 different molecular barcodes may be used. The molecular barcodes can be ligated to both ends of the target molecule. For example, 20 - 50×20 - 50 molecular barcodes can be used. In some embodiments, 20 - 50 different molecular barcodes can be used. In some embodiments, 5 - 100 different molecular barcodes can be used. In some embodiments, 5 - 150 molecular barcodes can be used. In some embodiments, 5 - 200 different molecular barcodes can be used. Such numbers of identifiers are typically sufficient for different molecules having the same start and end points to have a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of receiving different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95% or about 99% of the molecules have the same combination of molecular barcodes.

[0143] In some embodiments, the assignment of unique or non-unique molecular barcodes during the reaction can be performed using, for example, the methods and systems described in U.S. Patent Application Publication Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patents Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, only endogenous sequence information (e.g., start and / or end positions, partial sequences at one or both ends of the sequence, and / or length) can be used to identify different nucleic acid molecules in a sample.

[0144] Nucleic acid amplification

[0145] Sample nucleic acids adjacent to the adapter are typically amplified by PCR and other amplification methods using nucleic acid primers that bind to primer binding sites within the adapter that are adjacent to the DNA molecule to be amplified. In some embodiments, the amplification method includes cycles of extension, denaturation, and annealing that result from thermal cycling, or may be isothermal amplification, such as in the case of transcription-mediated amplification. Other exemplary amplification methods that may be utilized as needed include, among a number of techniques, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence replication.

[0146] One or more rounds of amplification cycles are generally applied to introduce a sample index into a nucleic acid molecule using conventional nucleic acid amplification methods. Amplification is typically performed in one or more reaction mixtures. The molecular tag and the sample index / tag can be introduced simultaneously or in any order as needed. In some embodiments, the molecular tag and the sample index / tag are introduced before and / or after a nucleic acid molecule capture step (i.e., nucleic acid concentration) is performed. In some embodiments, only the molecular tag is introduced before probe capture, and the sample index / tag is introduced after the sequence capture step is performed. In certain embodiments, both the molecular tag and the sample index / tag are introduced before the performance of the probe-based capture step. In some embodiments, the sample index / tag is introduced after the sequence capture step is performed. Typically, the sequence capture protocol involves introducing a single-stranded molecule complementary to a targeted nucleic acid sequence, e.g., the coding sequence of a genomic region, and a mutation in such a region associated with a cancer type. Typically, the amplification reaction generates a plurality of nucleic acid amplicons that are non-uniquely or uniquely tagged with a molecular tag and a sample index / tag in a size range of about 200 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicon has a size of about 300 nt. In some embodiments, the amplicon has a size of about 500 nt.

[0147] Nucleic acid concentration

[0148] In some embodiments, the array is enriched prior to nucleic acid sequencing. Enrichment is performed as needed for specific target regions (“target sequences”). In some embodiments, differential tiling and capture mechanisms can be used to enrich target regions of interest using nucleic acid capture probes (“baits”) selected for one or more bait set panels. Differential tiling and capture mechanisms typically use bait sets of different relative concentrations to tile differentially (e.g., at different “resolutions”) across genomic regions associated with baits that will be subject to a set of constraints (e.g., sequencer constraints such as sequencing load, utility of each bait, etc.) and to capture a desired level of target nucleic acids for downstream sequencing. These targeted genomic regions of interest optionally include the natural or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, biotinylated beads with probes for one or more regions of interest can be used to capture target sequences and, optionally, then amplify those regions to enrich the regions of interest.

[0149] Array capture typically involves the use of oligonucleotide probes that hybridize to target nucleic acid sequences. In certain embodiments, probe set strategies involve tiling probes across the region of interest. Such probes can be, for example, about 60 to about 120 nucleotides in length. The set can have a depth of about 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 8-fold, 9-fold, 10-fold, 15-fold, 20-fold, 30-fold, 40-fold, 50-fold or greater than 50-fold. The effectiveness of array capture generally depends in part on the length of the sequence within the target molecule that is complementary (or nearly complementary) to the sequence of the probe.

[0150] Nucleic Acid Sequencing

[0151] Sample nucleic acids, which may or may not have been pre-amplified and which may have an adapter adjacent thereto as needed, will generally be subjected to sequencing. Sequencing methods or commercially available formats that may be used as needed include, for example, Sanger sequencing; high-throughput sequencing; pyrosequencing; sequencing by synthesis; single molecule sequencing; nanopore-based sequencing; semiconductor sequencing; sequencing by ligation; sequencing by hybridization; RNA-Seq (Illumina); Digital Gene Expression (Helicos); next generation sequencing (NGS); Single Molecule Sequencing by Synthesis (SMSS) (Helicos); massively parallel sequencing; Clonal Single Molecule Array (Solexa); shotgun sequencing; Ion Torrent; Oxford Nanopore; Roche Genia; Maxim-Gilbert sequencing; primer walking; sequencing using PacBio, SOLiD, Ion Torrent or nanopore platforms. The sequencing reaction can be carried out in various sample processing units, which may include other means of substantially simultaneously processing multiple lanes, multiple channels, multiple wells, or multiple sets of samples. The sample processing unit can also include multiple sample chambers that allow multiple runs to be processed simultaneously.

[0152] The sequencing reaction can be carried out using one or more nucleic acid fragment types or regions known to contain markers for cancer or other diseases (e.g., microsatellites and / or other repetitive nucleic acid elements). The sequencing reaction can also be carried out using any nucleic acid fragments present in the sample. The sequencing reaction can result in a genomic sequence coverage of at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome. In other cases, the genomic sequence coverage can be less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome. In some embodiments, the genomic sequence coverage can be less than about 0.01%, 0.02%, 0.05%, 0.1%, 0.2%, 0.5%, 1%, 2% or 5% of the genome. In some embodiments, the genomic sequence coverage can be less than about 0.01% of the genome. In some embodiments, the genomic sequence coverage can be less than about 0.02% of the genome. In some embodiments, the genomic sequence coverage can be less than about 0.05% of the genome. In some embodiments, the genomic sequence coverage can be less than about 0.1% of the genome. In some embodiments, the genomic sequence coverage can be less than about 0.2% of the genome. In some embodiments, the genomic sequence coverage can be less than about 0.5% of the genome. In some embodiments, the genomic sequence coverage can be less than about 1% of the genome. In some embodiments, the genomic sequence coverage can be less than about 2% of the genome. In some embodiments, the genomic sequence coverage can be less than about 5% of the genome. In some embodiments, the genomic sequence coverage can be at least about 5% of the genome. In some embodiments, the genomic sequence coverage can be at least about 10% of the genome. In some embodiments, the genomic sequence coverage can be at least about 20% of the genome.

[0153] Using multiplex sequencing techniques, simultaneous sequencing reactions can be performed. In some embodiments, the cell-free polynucleotide is sequenced using at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, the cell-free polynucleotide is sequenced using less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. The sequencing reactions are typically performed sequentially or simultaneously. Subsequent data analysis is generally performed on all or a portion of the sequencing reactions. In some embodiments, the data analysis is performed on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, the data analysis may be performed on less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. Exemplary read depths are about 1000 to about 50000 reads per locus (base position), or >50,000 reads per locus.

[0154] In some embodiments, the nucleic acid population is prepared for sequencing by enzymatically forming blunt ends on double-stranded nucleic acids having single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides in dNTP form (e.g., A, C, G, and T or U). Exemplary enzymes or catalytic fragments thereof that may be used as needed include the large Klenow fragment, and T4 polymerase. In the 5' overhang, the enzyme typically extends the recessed 3' end of the opposing strand until it precisely overlaps the 5' end to generate a blunt end. In the 3' overhang, the enzyme generally digests from the 3' end to the 5' end of the opposing strand and may also digest beyond the 5' end. If this digestion proceeds beyond the 5' end of the opposing strand, an enzyme having the same polymerase activity used for the 5' overhang can fill the gap. The formation of blunt ends on double-stranded nucleic acids can, for example, facilitate adapter ligation and subsequent amplification.

[0155] In some embodiments, the nucleic acid population will undergo further processing such as conversion of single-stranded nucleic acids to double-stranded and / or conversion of RNA to DNA. These forms of nucleic acids may also be ligated to adapters and amplified as needed.

[0156] Pre-amplified or unamplified nucleic acids will undergo the process of forming the blunt ends described above and, if needed, other nucleic acids in the sample can be sequenced to generate sequenced nucleic acids. Sequenced nucleic acids may refer to the sequence of the nucleic acid (i.e., sequence information) or to the nucleic acid whose sequence has been determined. Sequencing can be performed to obtain sequence data of individual nucleic acid molecules directly or indirectly from the consensus sequence of the amplification products of the individual nucleic acid molecules in the sample.

[0157] In some embodiments, after blunt-end formation, both ends of the double-stranded nucleic acid having single-stranded overhangs in the sample are ligated to an adapter containing a barcode, and by sequencing, not only the nucleic acid sequence but also the barcodes arranged in series introduced by the adapter are determined. The blunt-end DNA molecules are ligated as needed to the blunt ends of at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters). Alternatively, to facilitate ligation (e.g., sticky-end ligation), tails having complementary nucleotides can be added to the blunt ends of the sample nucleic acid and the adapter.

[0158] The nucleic acid sample is typically contacted with a sufficient number of adapters such that the probability that any two identical nucleic acids receive the same combination of adapter barcodes linked to both ends is low (e.g., <1 or <0.1%). The use of adapters in this way enables the identification of families of nucleic acid sequences that have the same starting and ending points as those on the reference nucleic acid and are linked to the same combination of barcodes. Such families represent the pre-amplification sequences of the amplification products of the nucleic acids in the sample. The sequences of the family members can be compiled to derive the consensus nucleotide or complete consensus sequence of the nucleic acids in the original sample modified by blunt-end formation and adapter ligation. In other words, the nucleotide occupying a particular position in the nucleic acids in the sample is determined to be the consensus nucleotide of the nucleotides occupying that corresponding position within the family member sequences. A family may include the sequence of one strand of the double-stranded nucleic acid or may include the sequences of both strands. When the members of the family include the sequences of both strands of the double-stranded nucleic acid, the sequence of one strand is converted to its complementary sequence in order to compile all the sequences and derive the consensus nucleotide or sequence. Some families contain only a single-member sequence. In this case, this sequence can be considered the sequence of the nucleic acid in the sample before amplification. Alternatively, families having only a single-member sequence can be excluded from subsequent analysis.

[0159] Nucleotide variations in the sequenced nucleic acid can be determined by comparing the sequenced nucleic acid with a reference sequence. The reference sequence is often known, for example, the entire or partial genomic sequence from a subject (e.g., the entire genomic sequence of a human subject) is known. The reference sequence can be, for example, hG19 or hG38. The sequenced nucleic acid may represent the sequence directly determined for the nucleic acid in the sample as described above, or may represent the consensus sequence of the sequences of the amplification products of such nucleic acids. Comparisons can be performed at one or more designated positions on the reference sequence. A subset of the sequenced nucleic acids that includes positions that match the designated positions of the reference sequence when the respective sequences are maximally aligned can be identified. Within such a subset, if any, it can be determined which of the sequenced nucleic acids contain a nucleotide variation at the designated position, and if any, which contain the reference nucleotide (i.e., the same as that in the reference sequence) can be determined as needed. If the number of sequenced nucleic acids in the subset containing the nucleotide variant exceeds a selected threshold, the variant nucleotide can be called at the designated position. The threshold can be, among other possibilities, a simple number such as at least 1, 2, 3, 4, 5, 6, 7, 9, or 10 of the sequenced nucleic acids in the subset containing the nucleotide variant, or a ratio such as at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20 with respect to the sequenced nucleic acids in the subset containing the nucleotide variant. The comparison for any designated position of interest in the reference sequence can be repeated. Sometimes, comparisons can be performed for designated positions that occupy at least about 20, 100, 200, or 300 consecutive positions on the reference sequence, for example, at least about 20 - 500 or about 50 - 300 consecutive positions.

[0160] Further details regarding nucleic acid sequencing, including the formats and applications described herein, can be found, for example, in Levy et al., Annual Review of Genomics and Human Genetics, 17:95-115 (2016), Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012), Voelkerding et al., Clinical Chem., 55: 641-658 (2009), MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009), Astier et al., J Am Chem Soc., 128(5):1705-10 (2006), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Patent No. 7,501,245, U.S. Patent No. 6,818,395, U.S. Patent No. 6,911,345, U.S. Patent No. 7,501,245, U.S. Patent No. 7,329,492, U.S. Patent No. 7,170,050, U.S. Patent No. 7,302,146, U.S. Patent No. 7,313,308, and U.S. Patent No. 7,476,503, each of which is hereby incorporated by reference in its entirety.

[0161] Comparator results

[0162] The microsatellite and / or other repetitive nucleic acid instability status of a given subject determined according to the methods disclosed herein is typically compared to a database of comparator results from a reference population to identify a customized or targeted therapy for that subject. In some embodiments, the microsatellite and / or other repetitive nucleic acid instability status of the subject being tested and the comparator results are measured, for example, across the entire genome or exome, while in other embodiments, these markers are measured based on, for example, a subset or targeted region of the genome or exome and extrapolated as needed, for example, to determine microsatellite instability across the entire genome or exome. Typically, the reference population includes patients having the same cancer type as the subject being tested and / or patients who are or have been treated with the same therapy as the subject being tested. In some embodiments, the subject microsatellite and / or other repetitive nucleic acid instability status, as well as the comparator microsatellite and / or other repetitive nucleic acid instability status, are measured by determining the number or burden of mutations in a predetermined or selected set of genes or genomic regions. Essentially any gene (e.g., a cancer gene) can be selected for such analysis as needed. In certain embodiments of these embodiments, the genes or genomic regions selected include at least about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1,500, 2,000 or more genes or genomic regions selected. In some embodiments of these embodiments, the genes or genomic regions selected optionally include one or more genes listed in Table 1.

Table 1

[0163] Cancer and Other Diseases

[0164] In certain embodiments, the methods and systems disclosed herein are used to identify customized therapies for treating a given disease, disorder or condition in a patient. Typically, the disease of interest is a type of cancer. Non-limiting examples of such cancers include biliary cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, eye melanoma, choroidal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myeloid (CML), chronic myelomonocytic (CMML), liver cancer, hepatocarcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B cell lymphoma, non-Hodgkin lymphoma, diffuse large B cell lymphoma, mantle cell lymphoma, T cell lymphoma, non-Hodgkin lymphoma, precursor T lymphoblastic lymphoma / leukemia, peripheral T cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostatic adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma.

[0165] Non-limiting examples of other genetic diseases, disorders or conditions that may be evaluated as needed using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cat cry, Crohn's disease, cystic fibrosis, Dercum's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (scid), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson's disease, or the like.

[0166] Customized Therapies and Related Administration

[0167] In some embodiments, the methods disclosed herein relate to identifying and administering customized therapies to patients having a given microsatellite and / or other repeat nucleic acid instability status. Essentially any cancer therapy (e.g., surgical therapy, radiation therapy, chemotherapy, and / or the like) is included as part of these methods. Typically, the customized therapy includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to methods of enhancing the immune response against a given cancer type. In certain embodiments, immunotherapy refers to methods of enhancing the T cell response against a tumor or cancer.

[0168] In some embodiments, the immunotherapy or immunotherapeutic agent targets immune checkpoint molecules. Certain tumors can evade the immune system by exploiting the immune checkpoint pathway. Thus, targeting immune checkpoints has emerged as an effective approach to counter the ability of tumors to evade the immune system and to activate antitumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012, 12:252-264.

[0169] In certain embodiments, the immune checkpoint molecule is an inhibitory molecule that reduces signals involved in the T cell response to an antigen. For example, CTLA4 is expressed on T cells and is involved in downregulating T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 restricts the activity of T cells in peripheral tissues during an inflammatory response. In addition, the ligands of PD-1 (PD-L1 or PD-L2) are generally upregulated on the surface of many different tumors, resulting in downregulation of antitumor immunity in the tumor microenvironment. In certain embodiments, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other embodiments, the inhibitory immune checkpoint molecule is a ligand of PD-1, such as PD-L1 or PD-L2. In other embodiments, the inhibitory immune checkpoint molecule is a ligand of CTLA4, such as CD80 or CD86. In other embodiments, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin-like receptor (KIR), T cell membrane protein 3 (TIM3), galectin 9 (GAL9), or adenosine A2a receptor (A2aR).

[0170] Antagonists targeting these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain embodiments, the inhibitory immune checkpoint molecule is PD-1. In certain embodiments, the inhibitory immune checkpoint molecule is PD-L1. In certain embodiments, the antagonist of the inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain embodiments, the antibody or monoclonal antibody is an anti-CTLA4, anti-PD-1, anti-PD-L1, or anti-PD-L2 antibody. In certain embodiments, the antibody is a monoclonal PD-1 antibody. In some embodiments, the antibody is a monoclonal PD-L1 antibody. In certain embodiments, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, an anti-CTLA4 antibody and an anti-PD-L1 antibody, or an anti-PD-L1 antibody and an anti-PD-1 antibody. In certain embodiments, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In certain embodiments, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In certain embodiments, the anti-PD-L1 antibody is one or more of atezolizumab (Tecentriq®), avelumab (Bavencio®) or durvalumab (Imfinzi®).

[0171] In certain embodiments, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., an antibody) to CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other embodiments, the antagonist is a soluble version of an inhibitory immune checkpoint molecule, e.g., a soluble fusion protein comprising the extracellular domain of an inhibitory immune checkpoint molecule and the Fc domain of an antibody. In certain embodiments, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1, or PD-L2. In some embodiments, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one embodiment, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.

[0172] In certain embodiments, the immune checkpoint molecule is a costimulatory molecule that amplifies signals involved in the T cell response to an antigen. For example, CD28 is a costimulatory receptor expressed on T cells. When a T cell binds to an antigen via its T cell receptor, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on the antigen-presenting cell to amplify T cell receptor signaling and promote T cell activation. Since CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 can interfere with or regulate costimulatory signaling mediated by CD28. In certain embodiments, the immune checkpoint molecule is a costimulatory molecule selected from CD28, inducible T cell costimulatory molecule (ICOS), CD137, OX40, or CD27. In other embodiments, the immune checkpoint molecule is a ligand of a costimulatory molecule, including, for example, CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.

[0173] Using agonists that target these co-stimulatory checkpoint molecules can enhance antigen-specific T cell responses against certain cancers. Thus, in certain embodiments, the immunotherapy or immunotherapeutic agent is an agonist of a co-stimulatory checkpoint molecule. In certain embodiments, the agonist of the co-stimulatory checkpoint molecule is an agonist antibody, preferably a monoclonal antibody. In certain embodiments, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti-OX40, or anti-CD27 antibody. In other embodiments, the agonist antibody or monoclonal antibody is an anti-CD80, anti-CD86, anti-B7RP1, anti-B7-H3, anti-B7-H4, anti-CD137L, anti-OX40L or anti-CD70 antibody.

[0174] Treatment options for specific genetic diseases, disorders or conditions other than cancer are generally well known to those skilled in the art and will be apparent considering the particular disease, disorder or condition being considered.

[0175] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) can also be administered by any method known in the art, including, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal and / or intratympanic, and such administration can include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, plasters, ointments, or the like.

[0176] Systems and Computer-Readable Media

[0177] The present disclosure also provides various systems and computer program products or machine-readable media. In some embodiments, for example, the methods described herein are at least partially implemented or facilitated using a system, distributed computing hardware and applications (such as cloud computing services), an electronic communication network, a communication interface, a computer program product, a machine-readable media, an electronic storage medium, software (such as machine-executable code or logical instructions), and / or the like, as necessary. By way of example, FIG. 2 provides a schematic diagram of an exemplary system suitable for use in connection with the implementation of at least aspects of the methods disclosed herein. As shown, system 200 includes at least one controller or computer, such as server 202 (such as a search engine server), including a processor 204 and a memory, storage device, or memory component 206, and one or more other communication devices 214 and 216 (such as client-side computer terminals, telephones, tablets, laptops, other mobile devices, etc.) located remotely from remote server 202 and communicating with remote server 202 via an electronic communication network 212, such as the Internet or other Internet network. Communication devices 214 and 216 typically include, for example, an electronic display (such as a computer capable of connecting to the Internet or the like) that communicates with server 202 computer via network 212, and the electronic display includes a user interface (such as a graphical user interface (GUI), a web-based user interface, and / or the like) for displaying results during the execution of the methods described herein. In certain embodiments, the communication network also includes the physical transfer of data from one location to another, for example, using a hard drive, a thumb drive, or other data storage mechanism.System 200 also includes a program product 208 stored on a computer or machine-readable medium, such as one or more of various types of memory, such as memory 206 of server 202, which is readable by server 202, for example, to facilitate an inductive search application or the like executable by one or more other communication devices such as 214 (schematically shown as a desktop or personal computer) and 216 (schematically shown as a tablet computer). In some embodiments, system 200 also optionally includes at least one database server, such as server 210, associated with an online website storing data (such as control sample or comparator result data, indexed customized therapies, etc.), which is searchable, for example, directly or by search engine server 202. System 200 also optionally includes one or more other servers located remotely from server 202, each of these other servers optionally associated with one or more database servers 210, which are either remotely located or installed within the same area as each of the other servers. These other servers can beneficially provide services to remote users and can enhance geographically distributed operations.

[0178] As will be understood by those skilled in the art, the memory 206 of the server 202 includes volatile and / or non-volatile memory as needed, including, for example, among others, RAM, ROM, and magnetic or optical disks. Although illustrated as a single server, it will also be understood by those skilled in the art that the illustrated configuration of the server 202 is provided merely as an example, and that other types of servers or computers configured according to various other methodologies or configurations may also be used. The server 202 schematically shown in FIG. 2 represents a server or a server cluster or a server farm and is not limited to any particular physical server. The server site can be deployed as a server farm or a server cluster managed by a server hosting provider. The number of servers and their structures and configurations can be increased based on the usage, demand, and capacity requirements for the system 200. As will also be understood by those skilled in the art, for example, the other user communication devices 214 and 216 in these embodiments can be laptops, desktops, tablets, personal digital assistants (PDAs), mobile phones, servers, or other types of computers. As is known and understood by those skilled in the art, the network 212 can include the Internet, an intranet, a telecommunications network, an extranet, or the World Wide Web of multiple computers / servers communicating with one or more other computers via a communication network and / or via a local or other area network as part of it.

[0179] Furthermore, as will be understood by those skilled in the art, if desired, the exemplary program product or machine-readable medium 208 may be in the form of microcode, programs, cloud computing formats, routines and / or symbolic languages that cause a set or sets of ordered operations to control the functions of the hardware and direct its operation. According to an exemplary embodiment, the program product 208 also need not be entirely within volatile memory and may be selectively loaded as desired, according to various methodologies known and understood by those skilled in the art.

[0180] Furthermore, as will be understood by those skilled in the art, the terms "computer-readable medium" or "machine-readable medium" refer to any medium involved in providing instructions for execution to a processor. By way of example, the terms "computer-readable medium" or "machine-readable medium" include, for example, a distribution medium, a cloud computing format, an intermediate storage medium, a computer's execution memory, and any other medium or device capable of storing a program product 208 that executes the functions and processes of various embodiments of the present disclosure for reading by a computer. The "computer-readable medium" or "machine-readable medium" can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, for example, optical or magnetic disks. Examples of volatile media include, for example, dynamic memory such as the main memory of a given system. Examples of transmission media include coaxial cables, copper wire, and fiber optics, such as wires including a bus. Transmission media can also take the form of acoustic or light waves, among other things, such as those generated during radio wave and infrared data communications. Exemplary forms of computer-readable media include floppy (registered trademark) disks, flexible disks, hard disks, magnetic tapes, flash drives, or any other magnetic medium, CD-ROMs, any other optical medium, punch cards, paper tapes, any other physical medium having a pattern of holes, RAM, PROM, and EPROM, FLASH (registered trademark)-EPROM, any other memory chip or cartridge, a carrier wave, or any other medium readable by a computer.

[0181] The program product 208 is optionally copied from a computer-readable medium to a hard disk or similar intermediate storage medium. When the program product 208 or a portion thereof is tried out, it is optionally loaded from those distribution media, those intermediate storage media, or the like into one or more computer-execution memories that make up a computer, such that the computer is configured to operate in accordance with the functions or methods of the various embodiments. All such operations are well known to those of ordinary skill in the art of computer systems, for example.

[0182] To further illustrate by way of example, in certain embodiments, the present application provides a system that includes one or more processors and one or more memory components in communication with the processors. The memory components typically include one or more instructions that, when executed, cause the processors to display information such as array information, microsatellite and / or other repetitive nucleic acid instability status, comparator results, customized therapies, and / or the like (e.g., by communication devices 214, 216 or the like), and / or to receive information from other system components and / or from system users (e.g., by communication devices 214, 216 or the like).

[0183] In some embodiments, the program product 208, when executed by the electronic processor 204, at least (i) receives sequence information from a population of microsatellite loci in a sample; (ii) quantifies the number of different repeat lengths present at each of the plurality of microsatellite loci from the sequence information to generate a locus score for each of the plurality of microsatellite loci; (iii) compares the locus score of a given microsatellite locus for each of the plurality of microsatellite loci to a locus-specific trained threshold of the given microsatellite locus; (iv) if the locus score of the given microsatellite locus exceeds the locus-specific trained threshold of the given microsatellite locus, calls the given microsatellite locus as unstable and generates a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; (v) if the microsatellite instability score exceeds a population-trained threshold of the population of microsatellite loci in the sample, classifies the MSI status of the sample as unstable and identifies an unstable sample; and optionally, (vi) compares the microsatellite instability score of the unstable sample to one or more comparator results, and a substantial match between the microsatellite instability score of the unstable sample and the comparator result indicates a predicted response to a therapy of interest.

[0184] System 200 also typically includes additional system components configured to execute various aspects of the methods described herein. In some embodiments of these, one or more of these additional system components are located remotely from remote server 202 and communicate with remote server 202 via electronic communication network 212, while in other embodiments, one or more of these additional system components are located within the same area and communicate with server 202 (i.e., in the absence of electronic communication network 212) or communicate directly with, for example, desktop computer 214.

[0185] In some embodiments, for example, the additional system components include a sample preparation component 218 operably (either directly or indirectly (e.g., via electronic communication network 212)) connected to controller 202. Sample preparation component 218 is configured to prepare nucleic acids in a sample (e.g., prepare a library of nucleic acids) that are to be amplified and / or sequenced by a nucleic acid amplification component (such as a thermal cycler, etc.) and / or a nucleic acid sequencer. In certain embodiments of these, sample preparation component 218 is configured to separate nucleic acids in the sample from other components in order to, for example, ligate one or more adapters containing barcodes to the nucleic acids as described herein, selectively enrich one or more regions from a genome or transcriptome prior to sequencing, and / or for similar purposes.

[0186] In certain embodiments, system 200 also includes a nucleic acid amplification component 220 (such as a thermal cycler, etc.) operably connected (either directly or indirectly, e.g., via electronic communication network 212) to controller 202. The nucleic acid amplification component 220 is configured to amplify nucleic acids in a sample from a subject. For example, the nucleic acid amplification component 220 is configured as needed to amplify regions selectively enriched from the genome or transcriptome in a sample as described herein.

[0187] System 200 typically also includes at least one nucleic acid sequencer 222 operably connected (either directly or indirectly, e.g., via electronic communication network 212) to controller 202. The nucleic acid sequencer 222 is configured to provide sequence information from nucleic acids (e.g., amplified nucleic acids) in a sample from a subject. Essentially any type of nucleic acid sequencer can be adapted for use in these systems. For example, the nucleic acid sequencer 222 is configured as needed to perform pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, or other sequencing techniques using nucleic acids to generate sequencing reads. Optionally, the nucleic acid sequencer 222 is configured to group sequencing reads into families of sequencing reads, each family containing sequencing reads generated from nucleic acids in a given sample. In some embodiments, the nucleic acid sequence sequencer 222 generates sequencing reads using a clonal single molecule array derived from a sequencing library. In certain embodiments, the nucleic acid sequence sequencer 222 includes at least one chip having an array of microwells for sequencing a sequencing library to generate sequencing reads.

[0188] To facilitate full or partial system automation, system 200 typically also includes a material transfer component 224 operably (either directly or indirectly (e.g., via electronic communication network 212)) connected to controller 202. The material transfer component 224 is configured to transfer one or more materials (e.g., nucleic acid samples, amplicons, reagents, and / or the like) to and / or from the nucleic acid sequencer 222, the sample preparation component 218, and the nucleic acid amplification component 220.

[0189] Further details regarding computer systems and networks, databases, and computer program products are provided, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), and these references are hereby incorporated by reference in their entirety.

Example

[0190] (Example 1) Non-tumor samples were used as background to computationally simulate highly microsatellite unstable (MSI-H) samples using variable tumor fractions and number of unstable sites. The distribution observed in a cohort of 3000 samples of different cancer types was used as prior for the number of unstable sites. This analysis demonstrated 94% sensitivity at a limit of detection (LoD ) of 0.2% tumor content. The predicted specificity of the method for determining MSI status according to the embodiments described herein for non-tumor donor samples was 99.999%. Comparison of these results with standard or conventional PCR-based MSI evaluations showed 100% concordance over a tumor content range of 1.4 - 15%. Additionally, the performance of this analysis was tested on 155 clinical samples from three cancer types where standard PCR-based evaluations of MSI status were available (10 MSI-H, 145 microsatellite stable (MSS)). The MSI calls generated according to the embodiments described herein showed 100% concordance with standard PCR-based MSI evaluations.

[0191] (Example 2) The MSI status of 82 samples was evaluated using the Digital Sequencing Clinical Platform (Guardant Health, Inc., Redwood City, CA, USA). The Digital Sequencing Platform is an NGS panel of cancer-related genes that utilizes high-quality sequencing of cell-free DNA (which may include circulating tumor DNA) isolated by simple non-invasive blood sampling. Digital Sequencing uses pre-sequencing preparation of a digital library of individually tagged cfDNA molecules, combined with post-sequencing bioinformatics reconstruction, to eliminate nearly all false positives. Targeted sequencing of cfDNA in the samples was used to obtain sequence information. A site score (ΔAIC) was determined for 61 of the most informative microsatellite loci in each sample. The tumor fraction of the samples ranged from 0.5% to 15%. Unstable microsatellite loci in each sample were identified by comparing the site score for each sample to the corresponding site-specific trained threshold. The number of unstable microsatellite loci identified in a given sample was used as the microsatellite instability score (i.e., MSI sample score) for that particular sample. A population-trained threshold was determined for the 61 microsatellite loci in the samples, where a microsatellite instability score greater than or equal to 5 was predicted to classify the sample as highly MSI (MSI-H), while a microsatellite instability score less than 4 was predicted to classify the sample as microsatellite stable (MSS). Nine of the 82 samples were classified as MSI-H. The remaining 73 samples were classified as MSS. In all 82 samples, the predicted stability status matched the expected stability status. The MSI status was confirmed based on orthogonal validation.

[0192] (Example 3) Introduction

[0193] Microsatellite instability (MSI) is a biomarker for guideline recommendations regarding the prognostic significance of various tumor types as well as the predictive significance for treatment with immune checkpoint inhibitors. Historically, microsatellite instability detection has relied on examination of tumor tissue by PCR or immunohistochemistry. More recently, next-generation sequencing (NGS) methods have been developed, but these methods also rely on the availability of tumor tissue. In contrast, plasma-based MSI detection methods have been able to provide a non-invasive real-time assessment of MSI status. Guardant Health's large panel cell-free DNA (cfDNA) NGS assay evaluates 500 cancer-related genes to identify genomic changes and tumor mutational burden (TMB). In addition to single nucleotide variants (SNVs), indels, copy number amplifications (CNAs), fusions, and TMB, this panel can detect high microsatellite instability (high MSI) status based on somatic changes at >1,000 MSI sites. The analytical validation presented in this example serves to determine the performance of a 500-cancer-related gene cfDNA NGS assay for the detection of high MSI, with four main components: accuracy, limit of detection (LoD), precision, and limit of blank (LoB).

[0194] Method

[0195] Accuracy analysis was performed by comparing the MSI status predicted by cfDNA NGS assays of 500 cancer-related genes with the truth based on the tissue MSI status determined by an orthogonal method, using 258 samples from 3 sources. Thirty-six collaborator samples with tissue MSI status (truth: tissue MSI status), 121 healthy donors (truth: microsatellite stable, MSS), and 101 samples sequenced by cfDNA NGS assays of 500 cancer-related genes (large panel assay) and cfDNA GNS assays of 73 cancer-related genes (small panel, MSI status as truth) were used. Reproducibility and repeatability analyses were performed using 2 sets of replicates (56 replicates in total). MSI status and MSI score were compared within and between trials. LoD was obtained by simulation for both. LoB was calculated using samples from healthy donors and known MSS samples.

[0196] Results

[0197] 1. Accuracy analysis Twelve out of 13 high MSI samples based on tissue MSI status were called high MSI by the large panel assay. All MSS / low MSI (MSI-L) samples were accurately detected (Table 2). When limited to the microsatellite regions (about 90 sites) covered by the small panel assay, the 12 samples detected as high MSI by the large panel assay also met the small panel assay threshold for being called high MSI. [Table 2]

[0198] 2. LoD analysis

[0199] Simulations at >1,000 sites used to detect MSI across five tumor proportions in the range of 0.05% - 1% showed a LoD of 0.1% (Figure 3).

[0200] Detection of high MSI in cfDNA NGS assays of 3,500 cancer-related genes is reproducible and repeatable.

[0201] Twenty-four high MSI replicates were assayed in two trials. All replicates were detected as high MSI in the cfDNA NGS assay of 500 cancer-related genes. The MSI numerical score was ±4 within and between trials (Figure 4A). Ten MSS / low MSI samples had 2 - 3 replicates (32 replicates in total) examined on the same flow cell. All replicates were detected as MSS / low MSI in the cfDNA NGS assay of 500 cancer-related genes. The MSI score was ±3 within each sample (Figure 4B).

[0202] 4. LoB analysis

[0203] One hundred and twenty-one samples from healthy donors and 25 samples from collaborators with known MSS status were used for LoB analysis. All 146 samples were accurately classified as MSS / low MSI in the large panel assay, showing a false positive rate of 0%.

[0204] 5. Tumor fraction and MSI score

[0205] The MSI score was plotted against the maximum somatic call mutant allele fraction (MAF) in large panel assay samples over 2,000 (Figure 5A and B), revealing that the tumor fraction (measured by MAF) does not correlate with the MSI status.

[0206] Conclusion

[0207] Detection of high microsatellite instability (MSI) by cfDNA NGS assay of 500 cancer-related genes showed high sensitivity (>90%) and specificity (100%). Repeatability and reproducibility were high within and across trials. The LoD for detection of high MSI was 0.1% MAF. The LoB study showed a false positive rate of 0%. The cfDNA NGS assay of 500 cancer-related genes provides a reliable prediction of high MSI status in cfDNA, giving physicians treatment values without the need for tissue samples.

[0208] (Example 4) Introduction

[0209] Microsatellite instability (MSI) is an important predictive biomarker for response to immune checkpoint blockade (ICB), as exemplified by the broad approval of pembrolizumab, for at least nine cancer types - cervical, cholangiocarcinoma, colorectal, endometrial, esophageal and esophagogastric, gastric, ovarian, pancreatic, and prostate cancer (1-9) - and is a biomarker recommended by the National Comprehensive Cancer Network (NCCN) clinical practice guidelines. Detection of MSI in patients with advanced cancer can also alert clinicians to evaluate asymptomatic family members of the patient for risk of hereditary cancer.

[0210] MSI is a typical manifestation of DNA mismatch repair deficiency (dMMR) that results in a dramatic increase in the mutation rate across the genome, including the acquisition and / or loss of nucleotides within repetitive motifs known as microsatellite tracts, from which the name of this entity is derived. MSI is most prevalent in endometrial, colorectal, and gastroesophageal cancers, and it can be a sequela of sporadic mutations in MMR-related genes or the manifestation of Lynch syndrome, the most common hereditary cancer predisposition syndrome caused by germline mutations in MLH1, MSH2, MSH6, PMS2, or EPCAM (12). However, despite the increased prevalence in these cancer types, landscape analysis has shown that MSI is present at non-negligible rates in most other solid tumors, including frequently seen tumor types such as lung, prostate, and breast cancers (13).

[0211] Recent studies have shown that MSI predicts clinical benefit from ICB with PD-1 / PD-L1 inhibitors, leading to the approval of these agents, such as nivolumab ± ipilimumab for highly microsatellite-instable (MSI-H, positive for MSI) metastatic colorectal cancer, and pembrolizumab for unresectable or metastatic MSI-H solid tumors after progression on prior approved therapy, in several indications where MSI is present (14). In addition to its value as a predictive biomarker for the benefit of ICB, MSI also has prognostic significance, which is most prominent in colorectal cancer (CRC) and is recommended to be tested in clinical practice guidelines for all patients (3, 15).

[0212] Recently, MSI testing is most commonly performed by polymerase chain reaction (PCR) and / or immunohistochemistry (IHC) analysis of tumor tissue specimens. The former evaluates five standard microsatellite loci (16, 17) originally recommended by the Bethesda panel, compares their lengths in tumor DNA with those in germline genotypes evaluated with matched non-tumor DNA, and uses the instability of the length of each microsatellite tract as direct evidence of MSI. However, this limited microsatellite panel was developed mainly for CRC and has more limited sensitivity in other cancer types (18). In contrast, the IHC approach evaluates the levels of four MMR proteins, and the absence of expression of one or more (MMR deficiency, dMMR) strongly correlates with the MSI status. However, about 5-11% of MSI-H cases show intact MMR staining and localization (functional MMR, pMMR) due to the retention of antigenicity and intracellular transport of otherwise non-functional proteins (19). Recent published literature (20, 21) has demonstrated that next-generation sequencing (NGS) can also accurately characterize the MSI status in tumors, such that a single NGS assay can comprehensively profile not only the MSI status but also targetable genomic biomarkers.

[0213] Despite the recommendations across many cancer types in the NCCN Guidelines and the accompanying FDA-approved treatment options, current MSI testing rates for cancers other than CRC and gastroesophageal cancers remain very low (22). Even in CRC, where MSI testing recommendations have been in place since 2005 (17, 23), fewer than 50% of patients are tested (24), resulting in missed opportunities for ICB treatment and the inability to identify patients who may be at high risk of developing cancer in their families. Multifactorial, such under-genotyping of MSI is often due to barriers related to tissue acquisition and complex testing recommendations / algorithms. For example, testing of archival diagnostic specimens can lead to significant delays related to the location and acquisition of this material and may also result in inaccurate assessment of MSI status due to tumor progression and / or heterogeneity. Similarly, testing of newly obtained tissue specimens can also result in significant delays related to biopsy scheduling and failure, with additional risks and costs associated with the complexity of the procedure. Therefore, invasive tissue acquisition procedures are contraindicated for many patients who have undergone intensive pretreatment and / or are frail. In addition, the increasing number of biomarkers and the diversification of testing options create mind-boggling complexity for physicians who are already overburdened.

[0214] Cell-free circulating tumor DNA (ctDNA) assays (“liquid biopsies”) have successfully addressed such hurdles in the genotyping of many indications by enabling minimally invasive profiling of contemporaneous tumor DNA. Thus, liquid biopsies identify patients harboring biomarkers of interest that cannot be identified by other means due to tissue sampling limitations and do so more rapidly than typical tissue testing, expanding patients' access to standard of care targeted therapies, including ICB (25). Furthermore, comprehensive liquid biopsies can provide all guideline-recommended somatic genomic biomarker information for all adult solid tumors in a single test. In this study, we sought to enhance the utility of a previously validated ctDNA-based genotyping assay by the addition of MSI detection. We describe the design and validation of MSI assessment on this platform, report its performance in the largest ctDNA-tissue MSI validation cohort (n = 1145) not yet described, and evaluate response prediction in 16 patients with advanced gastric cancer treated with ICB. We also report the MSI-H landscape of more than 28,000 consecutive solid tumor patients tested in a laboratory certified by the US Clinical Laboratory Improvement Amendment (CLIA), accredited by the College of American Pathologists (CAP), and approved by the New York State Department of Health.

[0215] Materials and Methods

[0216] 1. Microsatellite Locus Selection

[0217] Guardant Health's small-panel cell-free DNA (cfDNA) NGS assay is a 74-gene panel (26, 27) previously validated for the detection of SNVs, indels, CNAs, and fusions in all guideline-recommended indications for advanced solid tumors. This assay initially incorporated 99 putative microsatellite loci consisting of short tandem repeats (STRs) of length 7 or greater, selected to include sites prone to instability across multiple tumor types, including 3 of the 5 Bethesda panel sites (BAT-25, BAT-26, and NR-21). The remaining 2 Bethesda sites (NR-24 and MONO-27) were not included because of extremely low regional mappability. Sequencing data from a set of 84 samples from healthy donors were used to evaluate coverage and noise profiles at these sites, excluding sites of no informational value from the final MSI detection algorithm.

[0218] 2. Description of the Model

[0219] MSI detection is based on incorporating observed read sequences with molecular barcoding information into a single probability model, thereby comparing the likelihood of the observed data assuming PCR and sequencing noise with the likelihood of the observed data assuming somatic MSI instability. Each individual locus is scored independently using the Akaike information criterion (AIC) (28). The AIC model generates a locus score (ranging from 0 to infinity) that reflects the likelihood that the variability observed at any given microsatellite locus is due to biological instability relative to noise, and a locus is considered unstable if its score (i.e., the site score) is above a site-specific trained threshold. The number of affected loci is calculated across the final 90 sites, and if the number of unstable loci ("MSI score") is above a population-trained threshold (n = 6), the sample is called positive. Thresholds for individual loci and for the total MSI score per sample were established using simulations based on permutations of data from samples of healthy donors, varying the frequencies of molecules with different repeat lengths and error parameters at individual loci and the total number of unstable loci in the simulated samples. Simulations by this method were used here to examine 100,000 combinations of microsatellite length and number of unstable loci. This enabled the evaluation of the diverse landscape of scenarios, some of which may not be represented in the unsimulated dataset. The algorithm groups microsatellite stable (MSS) and low-level MSI (a category defined by the observation of a single unstable Bethesda locus using the PCR method) into a single category without distinguishing between them. This is based on previous reports that the MSI-L status is not a distinct phenotype but an artifact of the examination of a small number of microsatellites, and thus, when examining a large number of microsatellite loci, previously characterized MSI-L samples mimic the MSS phenotype among the total MSI burden.

[0220] 3. Sample

[0221] In addition to simulation data, a set of 84 samples from healthy donors was also used to develop and train the MSI algorithm. Clinical validation studies were conducted on 1145 archived samples (residual plasma and / or cell-free DNA) collected and processed as part of a routine standard care clinical trial (26) at the previously described Guardant Health CLIA laboratory, or archived patient plasma samples collected in EDTA tubes. Samples from 20 healthy donors were also used in the analytical specificity study. The devised samples used in the analytical validation study included cell line supernatants and cfDNA pools extracted from the plasma of healthy donors. Cell-free DNA prepared from the culture supernatants of the following cell lines was used (ATCC, Inc.): KM12, NCI-H660, HCC1419, NCI-H2228, NCI-H1650, NCI-H1648, NCI-H1975, NCI-H1993, NCIH596, HCC78, GM12878, MCF-7. CfDNA isolated from cell line culture supernatants mimics the fragment size and mechanism of extracellular release of patient-derived cfDNA (29), library conversion, and sequencing characteristics, and moreover, provides a renewable source of well-defined material in sufficient quantity to support advanced research material requirements such as detection limit and accuracy.

[0222] 4. Sample Processing and Bioinformatics Analysis

[0223] Cell-free DNA was extracted from plasma samples or cell line supernatants (QIAmp Circulating Nucleic Acid Kit, Qiagen, Inc.), labeled with non-random oligonucleotide barcodes (IDT, Inc.) for extracted cfDNA of 30 ng or less, and then library preparation, hybrid capture enrichment (Agilent Technologies, Inc.), and sequencing by paired-end synthesis (NextSeq 500 / 550 or HiSeq 2500, Illumina, Inc.) were performed as previously described (26). Bioinformatics analysis and variant detection were performed as previously described (26).

[0224] 5. Analytical Verification Methods

[0225] The analytical validation studies conducted were based on established CLIA, Nex-StoCT Working Group, and Association published guidelines for performance characteristics and validation principles. of Molecular Pathologist / CAP guidance. To determine the sensitivity of the assay for MSI status, cfDNA from cell line supernatants from an MSI-H cell line (KM12) (29) was diluted with cfDNA from a microsatellite stable (MSS) cell line (NCI-H660) (30, 31) and tested at both standard (30 ng) and low (5 ng) cfDNA inputs. The dilution series targeted maximum mutant allele fractions (max MAFs) of 0.03–2% for the 5 ng input and 0.01–1% for the 30 ng input. Target tumor fractions were validated using known germline variants specific to the titration and dilution materials. Evaluation of repeatability (precision within a run) and reproducibility (match between runs) was based on clinical and design model samples. Six of the clinical samples for accuracy (three MSI-H, and three MSS) were selected, which had max MAF values ​​of 1-2%, corresponding to approximately 2-3 times the expected LoD at 5 ng. MSI assay specificity was determined by analyzing 20 healthy donor samples and 245 conjugate samples of known MSS.

[0226] 6. Clinical validation methods

[0227] Archived plasma or cfDNA from clinical samples from patients for whom standard-of-care tissue-based MSI test results were available were assayed using the ctDNA MSI algorithm (n=1145). Tissue-based MSI status was derived from IHC, PCR, or less frequently NGS. Clinical outcome data were extracted from patient medical records by the primary physician and de-identified.

[0228] 7. Landscape analysis of plasma MSI status from 28,459 advanced cancer patient samples

[0229] The cohort included 28,459 consecutive advanced cancer patient samples that were tested using 73 cancer-related gene cfDNA NGS assays (small panels) during their clinical care. All analyses were performed using anonymized data in accordance with an IRB-approved protocol. The prevalence of MSI-H in this cohort was evaluated across the following 16 primary tumor types: bladder cancer, breast cancer, cholangiocarcinoma, colon adenocarcinoma, cancer of unknown primary, head and neck squamous cell carcinoma, hepatocellular carcinoma, lung adenocarcinoma, non-specific lung cancer, lung squamous cell carcinoma, "other" cancer diagnosis, pancreatic adenocarcinoma, prostate adenocarcinoma, gastric adenocarcinoma, and uterine endometrial cancer.

[0230] 8. Statistics

[0231] Statistical analysis used the Student's t-test for analysis of the number of variants per sample and Fisher's exact test for comparison of ratios. The lower and upper bounds of the 95% confidence interval (CI) for binomial ratios were calculated using Wilson's score interval with continuity correction.

[0232] 9. Ethical Standards

[0233] This study was conducted using anonymized data in accordance with a protocol approved by the Quorum Institutional Review Board.

[0234] Results

[0235] 1. MSI Algorithm Development

[0236] Traditional challenges in ctDNA genotyping using NGS include efficient molecular capture due to low input amounts and low tumor fractions in circulation (26, 27), as well as correction of sequencing and other technical artifacts. MSI detection presents further challenges as it requires: 1) efficient molecular capture, sequencing, and mapping of repetitive genomic regions that accurately represent the MSI status; 2) error correction and variant detection within repetitive regions; and 3) discrimination between signals from MSI and strong PCR slip artifacts resulting from somatic mutations in non-MSI somatic cells typically affected by somatic instability. In fact, technical PCR errors are typically at least one order of magnitude greater than typical sequencing error rates at homopolymer sites, and thus optimal use of repetitive site selection and molecular barcoding is required to achieve an appropriate signal-to-noise detection ratio across a large number of candidate microsatellite sites.

[0237] Tissue sequencing panels often contain informative microsatellite loci simply due to their large panel size and longer DNA fragment lengths (13, 32), but the medium-sized ctDNA panels and short cell-free DNA (cfDNA) fragment lengths utilized here require intentional microsatellite selection and incorporation. To accomplish this, candidate sites were evaluated using iterative methods that obtained information from the literature and tissue sequencing summaries to perform pan-cancer MSI detection with minimal background noise. The list of candidate loci was further refined using cfDNA from healthy donors based on the performance criteria described above.

[0238] Based on the performance evaluation in training healthy donor samples, informative loci were defined as those that were effectively captured, sequenced, and mapped, and were associated with the near absence of mutations in the MSS samples (shown in light gray in Figure 6A). Uninformative loci either could not be captured, sequenced, or mapped, resulting in the display of inappropriate molecules (shown in black in Figure 6A), or demonstrated substantial mutations in the MSS samples, resulting in the generation of excessive artifact signals (shown in dark gray in Figure 6A). Interestingly, the BAT-25, BAT-26, and NR-21 Bethesda loci used in traditional MSI tissue assays (16, 17) and some ctDNA panels (33) were inferior in performance compared to other candidates and were excluded from the final marker set (shown by arrows in Figure 6A).

[0239] Using this approach, 90 microsatellite loci, which are 89 mononucleotide repeats and a single trinucleotide repeat, all having a repeat length of 7 or more, were selected for incorporation into the final test version. Evaluation of the unique molecular coverage distribution demonstrated that 65% of these loci had coverage exceeding 0.5 times the median sample coverage.

[0240] In addition to effective molecular capture and mapping, MSI detection necessarily involves highly accurate discrimination between cancer-related signals and background noise resulting from sequencing and polymerase errors at the very low allelic fractions at which ctDNA is typically found (26, 27, 34). Importantly, the same repetitive genomic context that makes microsatellite candidates informative for MSI detection due to polymerase slippage during in vivo cell replication also makes them particularly susceptible to the same polymerase slippage during in vitro library preparation and sequencing, resulting in elevated technical noise levels. To address this, error correction for digital sequencing was used to define true biological insertion-deletion events with high fidelity at microsatellite loci, as previously described (26, 27). The digital sequencing platform is an NGS panel of cancer-related genes that utilizes high-quality sequencing of cell-free DNA isolated by simple non-invasive blood draws, which may include circulating tumor DNA. Digital sequencing utilizes pre-sequencing preparation of a digital library of individually tagged cfDNA molecules, combined with post-sequencing bioinformatics reconstruction, to eliminate nearly all false positives.

[0241] Among these advanced background error repeats, digital sequencing was associated with a 100-fold reduction in sequencing errors per molecule (Figure 6B) compared to standard sequencing methods, enabling the efficient and accurate reconstruction of the microsatellite sequences of individual unique molecules present in the original patient blood sample. Threshold simulations based on permutations of samples from healthy donors were then used to establish site-specific and total sample level MSI status determination thresholds. Combining these per-site and per-sample thresholds with the effect of digital sequencing correction, the false positive rate per sample was estimated to be approximately 10−7.3. In addition, titration simulations adjusted for the distribution of clinical input amounts predicted robust MSI detection up to approximately 0.2% tumor fraction, but subsequent detection efficiency decreased significantly. Therefore, samples with a <0.2% circulating tumor fraction (defined by the maximum somatic variant allele fraction) were considered unevaluable for MSI status.

[0242] 2. Analytical validation study

[0243] To evaluate the analytical sensitivity of MSI detection, cfDNA collected from the supernatant of the MSI-H cell line KM12 was diluted into MSS cfDNA targeting five dilution points, which included 15 independently processed replicates straddling the detection limit (LoD) predicted by the above computer simulations. Each titration series was analyzed at both the minimum allowable cfDNA input amount of 5 ng and the maximum most common cfDNA input amount of 30 ng. Probit analysis was used to calculate the 95% LoD (LOD95) to be 0.4% at the 5 ng input amount (Figure 7A) and 0.1% at the 30 ng input amount (Figure 7B).

[0244] To evaluate the intermediate accuracy of the analysis, replicates of four different devised materials, two MSS and two MSI-H, were analyzed (Figure 7C). Over 499 replicates, the category agreement of the MSI status was 100% (499 / 499, 95% CI 99–100%), and the coefficient of variation for the quantitative MSI score was 6.3–7.2% for the MSI-H samples (Table 3). Repeatability and input amount robustness were also evaluated by replicate testing of MSS and MSI-H devised materials at cfDNA input amounts of 5 ng, 10 ng, and 30 ng, and similarly demonstrated 100% agreement (27 / 27, 95% CI 85–100%, Figure 8 and Table 4). Clinical accuracy was confirmed in replicates of 72 independent patient samples representing various MSI scores and tumor fractions processed across three independent batches, dates, operators, and reagent lots, demonstrating 100% qualitative agreement (72 / 72, 95% CI 94–100%) with a coefficient of variation of 2.0–15.2% for the underlying quantitative MSI score (Table 5).

Table 3

Table 4

Table 5

[0245] To evaluate the analytical specificity, plasma samples from healthy donors (distinct from those used in training), MSS devised materials, and MSS patient samples were analyzed for false MSI-H calls. The analytical specificity was 100% across healthy donor samples (20 / 20, 95% CI 83–100%), devised samples (245 / 245, 95% CI 98–100%), and patient samples (48 / 48, 95% CI 92–100%).

[0246] 3. Clinical Validation Study

[0247] Since there was no orthogonal cfDNA-based method that could be used for use as a comparator, the MSI status from the medical records and the ctDNA MSI evaluation were compared for 1145 samples including 40 different cancer types, 15 of which had at least 5 representative specimens, to determine clinical accuracy. In 949 unique evaluable patients, ctDNA detected 87% (71 / 82, 95% CI 77–93%) of patients reported as MSI-H and 99.5% (863 / 867, 95% CI 98.7–99.8%) of patients reported as MSS / MSI-L, with an overall accuracy of 98.4% (934 / 949, 95% CI 97.3–99.1%) and a positive predictive value (PPV) of 95% (71 / 75, 95% CI 86–98%) (Figure 10C, Table 8). Consistent with the computer modeling study, MSI-H detection was rare (0 / 19) in samples classified as non-evaluable due to a low tumor fraction (Figure 10A) that explained 57% (16 / 28) of the ctDNA-tissue discordance observed across the entire unique patient sample set (Tables 6–9). For samples with a tumor fraction higher than 1%, the ctDNA PPA increased to 93% (54 / 58, 95% CI 82–98%, Table 9).

Table 6

Table 7

Table 8

Table 9

[0248] Interestingly, despite the high correlation between IHC tissue testing and PCR tissue testing reported in the literature (23, 35), the concordance between ctDNA and tissue MSI status here differed by tissue testing methodology (97.4% (450 / 462) by PCR, 98.0% (239 / 244) by NGS, and 83.0% (93 / 112) by IHC, Figure 10B and Tables 10 and 11). In further investigation, this discordance was due to both an increase in the tissue IHC positive, ctDNA negative population (2.4% by PCR, 2.0% by NGS, and 12.5% by IHC, Fisher's exact probability test p < 0.001 for IHC-PCR and IHC-NGS) and an increase in the tissue IHC negative, ctDNA positive population (0.2% by PCR, 0% by NGS, and 4.5% by IHC, Fisher's exact probability test p < 0.01 for both comparisons). Due to these discrepancies, the inventors sought to investigate whether IHC limitations could contribute to the observed IHC-ctDNA discordance rather than a decrease in ctDNA accuracy. Of the 25 samples for which IHC and another tissue test result were available, 12 samples demonstrated IHC-ctDNA discordance. Importantly, PCR and / or NGS tissue testing supported the ctDNA NGS result rather than the tissue IHC in 5 of the 12 discordances. Collectively, these data support previous reports (36) that IHC may not be as reliable as the PCR diagnostic gold standard in MSI determination.

Table 10

Table 11

[0249] 4.28,459 Consecutive Advanced Cancer Patients' ctDNA MSI Status

[0250] Multiple studies have evaluated the prevalence of MSI across different tumor types in tissue (13, 32, 37), but to date, a landscape analysis of ctDNA MSI status across cancer types has not been published. To achieve this, the above MSI algorithm was applied to 28,459 consecutive advanced cancer patient clinical samples tested at Guardant Health Clinical Laboratory. In this cohort, 278 samples (median tumor fraction 6.55%, range 0.09 - 89%) including 16 different tumor types were identified as MSI-H by ctDNA. This corresponds to an overall pan-cancer prevalence of approximately 1%, similar to that previously reported for tissue (13, 32, 37). Similarly, the prevalence of MSI-H across tumor types also closely reflected that observed in tissue-based analyses (Figure 11A); as expected, MSI-H was most prevalent in endometrial, colorectal, and gastric cancers, while other tumors such as lung, bladder, and head and neck cancers demonstrated lower prevalence. Specific deviations from previous MSI-H prevalence estimates included slightly lower prevalence in endometrial, colorectal, and gastric cancers, as well as slightly higher prevalence in prostate cancer.

[0251] Considering the nature of pan-solid tumors in the intended ctDNA use population and the immunotherapies approved for MSI-H tumors, microsatellite loci for this panel were intentionally selected such that MSI status information could be obtained across all solid tumor types. Consistent with this design intent, analysis of the sample-level and locus-level MSI scores - i.e., the MSI score and site score, respectively - (Figures 11B and 11C) demonstrated consistent performance across tumor types, with MSI-H samples demonstrating significantly higher signals than the threshold. Furthermore, the diagnostic rate of MSI assessment for tumor types other than those where MSI is commonly tested was substantial, with more than half of the identified cases (143 / 278) present in tumor types where MSI testing is very rare, thus identifying patients who otherwise would never have been tested.

[0252] Consistent with what has been reported in the literature (38), the number of indels and SNVs (including non-synonymous and synonymous variants) is significantly increased in MSI-H samples compared to those characterized as having an MSS status (Figure 12). Specifically, the median number of SNVs in MSI-H samples was 6.3 as compared to 1.4 in MSS (chi-square p < 0.0001), and the median number of indels in MSI-H samples was 2.6 as compared to 0.4 in MSS (chi-square p < 0.0001).

[0253] 5. ctDNA MSI status predicts response to immunotherapy

[0254] Currently, the most prominent utility of MSI status is its ability to select patients for immunotherapy. Despite this and the obstacles to obtaining tissue from many patients, the ability of ctDNA MSI status to predict response to immunotherapy has not been reported. To establish the clinical validity of this biomarker, we present the clinical outcomes of 16 patients with ctDNA MSI-H metastatic gastric cancer treated with pembrolizumab (n = 15) or nivolumab (n = 1) after failure of standard care chemotherapy in a phase II pembrolizumab trial (NCT#02589496) in gastric cancer. cfDNA and tissue PCR MSI assessments in pretreatment samples were 100% concordant for MSI-H (16 / 16, 95% CI 76–100%). Ten of the 16 patients achieved a complete response (n = 3) or partial response (n = 7) as objectively evaluated by the study physician according to RECIST 1.1 criteria, and an additional 3 patients had stable disease (Figure 13A), with an objective response rate of 63% (10 / 16, 95% CI 36–84%) and a disease control rate of 81% (13 / 16, 95% CI (54 - 95%), which was similar to the previously reported response (39) for MSI-H patients defined by tissue testing. Importantly, even in this pretreatment population, these responses were durable, with a median treatment duration of 39 weeks. Indeed, for example, Patient 21 experienced a complete regression of disease after pembrolizumab treatment following failure of fluoropyrimidine / platinum chemotherapy and remained disease-free for more than 6 months after completion of 35 cycles of treatment (Figures 13B - 13E).

[0255] Discussion

[0256] We validated a novel cfDNA-based targeted NGS approach for MSI detection - by using a large panel of microsatellites, this approach achieved high sensitivity compared to tissue-based methods while maintaining very high specificity. The MSI-H prevalence detected in plasma across 16 common solid tumors was similar to published tissue-based summaries, thus demonstrating the pan-tumor performance intended in MSI detection algorithm design. Furthermore, by showing that MSI-H patients detected by cfDNA benefited from ICB therapy in the same way as those reported for the tissue-defined population (39), we demonstrated the clinical utility of MSI detection being extended to all patients regardless of tissue availability or the need to undergo invasive tissue acquisition procedures.

[0257] This example demonstrates the robust analytical performance of MSI detection with a ctDNA panel (26) previously validated for the detection of the other four variant types in all guideline-recommended indications. In particular, the analytical sensitivity for MSI detection in the devised samples demonstrated reproducible detection down to 0.1%, consistent with previous reports of similar sensitivities for indels and SNVs (26). Importantly, this example evaluated the performance of ctDNA MSI testing in 1145 samples including orthogonal tissue MSI, which constitute the largest ctDNA-tissue MSI concordant cohort not yet described. Compared to standard care tissue MSI testing for the same patients, ctDNA MSI evaluation had a high PPV (95%) that was not inferior to the reported PPV of 90 - 92% reported for MSI evaluation based on local vs central tissue (36), as well as a high PPA (87%) in the evaluable population consistent with previous studies (25, 26, 45, 46) that investigated concordance of plasma and tissue genotype determination for other variant types. Factors that could contribute to incomplete concordance include tumor heterogeneity, primary lesion Examples include different shedding due to metastatic lesions, temporal inconsistencies in tissue and plasma collection, and low tumor shedding by some tumors (40, 44, 47-49). Interestingly, gastric cancer patients identified as MSI-H by plasma and pentaplex PCR in this report constituted distinct tumor populations of MSS and MSI-H diseases as previously reported when evaluated by both IHC and PCR performed using tissue (40). The same study found a 9% discordance in MSI-H between paired tissue biopsy samples in the same patients (40), highlighting the potential contribution of tumor heterogeneity to MSI status discordance. Additionally, the observation of a non-trivial discordance between the PCR tissue method and the IHC tissue method in this report emphasizes the importance of accurate MSI testing and has been reported as a major cause of ICB failure (36). Consistent with the challenges presented by tissue genotyping in advanced solid tumors, studies in metastatic non-small cell lung cancer (NSCLC) have shown that plasma-based tests not only increase the likelihood of successful tumor genotyping results compared to tissue, but also produce results at least one week faster than typical tissue genotyping results and increase the frequency of detecting targetable mutations (25, 45).

[0258] This example presents the first ctDNA-based landscape analysis of MSI in a large advanced cancer cohort. Overall, the relative prevalence across tumor types in a set of >28,000 consecutive clinical samples was consistent with that reported for tissue (13, 32, 37), with little difference. For example, the prevalence in CRC and endometrial cancer was lower than that reported for tissue (13), which most likely reflects the fact that tissue-based landscape analysis includes a number of early MSI-H tumors with a better prognosis (15) and a lower likelihood of being part of the advanced cancer population examined by ctDNA. On the other hand, the higher-than-expected prevalence of MSI-H prostate cancer is due to the increased indication of MSI-H disease in advanced patients, and two recent studies focusing on the MSI status in advanced prostate cancer have shown MSI-H prevalence rates of 3.1% and 3.8% in their patient populations, similar to the 2.6% observed by the inventors in this study (50, 51). Naturally, by design of the pan-cancer MSI detection, the landscape analysis did not reveal tumor type-specific patterns of microsatellite instability. However, this does not rule out the possibility that tumor type-specific patterns can emerge in plasma, as shown in tissue (37, 52), by evaluating a larger number of microsatellite loci and a larger number of representative MSI-H samples.

[0259] The clinical outcomes reported here were limited to gastric cancer; nevertheless, the observed objective response rate was consistent with predictions from tissue-based studies, suggesting that ICB treatment based on cfDNA MSI results should achieve expected outcomes across solid tumor types. In addition, the lack of germline dMMR data precluded drawing conclusions about the familial impact of cfDNA-detected MSI. Finally, treatment data for the majority of patients with cfDNA- / tissue+ discrepancies were unavailable, but at least some patients likely received ICB therapy based on tumor results, which may suppress MSI-H disease and contribute to the lack of cfDNA-detected MSI. Therefore, the 87% sensitivity for MSI-H detection compared to tissue may be higher in treatment-naïve or untreated patients. Further studies should be pursued to address these issues.

[0260] In summary, we developed and validated a cfDNA-based targeted NGS panel that accurately assesses MSI status and also provides comprehensive tumor genotyping, thus enabling a single peripheral blood draw to perform a test that meets pan-solid tumor guidelines with high sensitivity, specificity, and accuracy. Clinical validation using both tissue testing, population-level prevalence analysis, and comparison to the initially reported outcomes of ICB-treated cfDNA MSI-H patients supported the clinical accuracy and relevance of this approach. Such simultaneous characterization of MSI status and tumor genotype by simple peripheral blood draw may expand opportunities to utilize both targeted and immunotherapies in all advanced cancer patients, including those for whom the current tissue-based testing paradigm is inadequate. Clinical validation using both tissue testing, population-level prevalence analysis, and comparison to the initially reported outcomes of ICB-treated cfDNA MSI-H patients supported the clinical accuracy and relevance of this approach. Such simultaneous characterization of MSI status and tumor genotype by simple peripheral blood draw may expand opportunities to utilize both targeted and immunotherapies in all advanced cancer patients, including those for whom the current tissue-based testing paradigm is inadequate.

[0261] References

Chem.

Chem.

Chem.

Chem.

Chem.

[0262] For the purpose of clarifying and understanding the above disclosure, it has been described in some detail by way of explanation and examples. However, various changes in form and detail can be made without departing from the true scope of the present disclosure, and it will be apparent to those skilled in the art upon reading this disclosure that it can be implemented within the scope recited in the appended claims. For example, methods, systems, computer-readable media, and / or all of these component features, steps, elements or other aspects can be used in various combinations.

[0263] All patents, patent applications, websites, other published documents or literature, accession numbers and the like cited in this specification are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference in this manner. Even if different versions of a sequence are associated with an accession number at different times, the version associated with that accession number as of the effective filing date of the present application is intended. The effective filing date means, where available, the earlier of the actual filing date or the filing date of the priority application, with reference to the accession number. Similarly, even if different versions of a published document, website or the like were published at different times, unless otherwise indicated, the version published closest to the effective filing date of the present application is intended. The present invention provides, for example, the following items. (Item 1) A method for determining the repetitive nucleic acid instability status of a nucleic acid sample, comprising (a) Quantifying the many different repeat lengths present at each of a plurality of repetitive nucleic acid loci from sequence information to generate a site score for each of the plurality of repetitive nucleic acid loci, wherein the sequence information is from a population of repetitive nucleic acid loci in the nucleic acid sample; (b) When the site score of a given repetitive nucleic acid locus exceeds the site-specific trained threshold of the given repetitive nucleic acid locus, calling the given repetitive nucleic acid locus as unstable and generating a repetitive nucleic acid instability score including a number of unstable repetitive nucleic acid loci from the plurality of repetitive nucleic acid loci; and (c) When the repetitive nucleic acid instability score exceeds the population-trained threshold of the population of repetitive nucleic acid loci in the nucleic acid sample, classifying the repetitive nucleic acid instability status of the nucleic acid sample as unstable, thereby determining the repetitive nucleic acid instability status of the nucleic acid sample A method comprising. (Item 2) A method for determining the repetitive DNA instability status of a sample, comprising (a) Quantifying the many different repeat lengths present at each of a plurality of repetitive DNA loci from sequence information to generate a site score for each of the plurality of repetitive DNA loci, wherein the sequence information is from a population of repetitive DNA loci in the sample; (b) Comparing the site score of a given repetitive DNA locus for each of the plurality of repetitive DNA loci with the site-specific trained threshold of the given repetitive DNA locus; (c) When the site score of the given repetitive DNA locus exceeds the site-specific trained threshold of the given repetitive DNA locus, calling the given repetitive nuclear DNA locus as unstable and generating a repetitive DNA instability score including a number of unstable repetitive DNA loci from the plurality of repetitive DNA loci; and (d) If the repetitive DNA instability score exceeds the population-trained threshold of the population of the repetitive DNA loci in the sample, classifying the repetitive DNA instability status of the sample as unstable, thereby determining the repetitive DNA instability status of the sample A method comprising. (Item 3) A method for determining the microsatellite instability (MSI) status of a sample, comprising: (a) Quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in the sample; (b) Comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci with the site-specific trained threshold of the given microsatellite locus; (c) If the site score of the given microsatellite locus exceeds the site-specific trained threshold of the given microsatellite locus, calling the given microsatellite locus as unstable and generating a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; and (d) If the microsatellite instability score exceeds the population-trained threshold of the population of the microsatellite loci in the sample, classifying the MSI status of the sample as unstable, thereby determining the MSI status of the sample A method comprising. (Item 4) A method for determining the microsatellite instability (MSI) status of a sample, comprising: (a) Receiving sequence information from a population of microsatellite loci in the sample; (b) Quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from the array information to generate a site score for each of the plurality of microsatellite loci; (c) Comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci with the site-specific trained threshold of the given microsatellite locus; (d) When the site score of the given microsatellite locus exceeds the site-specific trained threshold of the given microsatellite locus, calling the given microsatellite locus as unstable and generating a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; and (e) When the microsatellite instability score exceeds the population-trained threshold of the population of the microsatellite loci in the sample, classifying the MSI status of the sample as unstable, thereby determining the MSI status of the sample A method comprising. (Item 5) A method for identifying one or more customized therapies for treating a disease in a subject, comprising: (a) Quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from the array information to generate a site score for each of the plurality of microsatellite loci, wherein the array information is from a population of microsatellite loci in a sample; (b) Comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci with the site-specific trained threshold of the given microsatellite locus; (c) If the site score of the given microsatellite locus exceeds the site-specific trained threshold of the given microsatellite locus, calling the given microsatellite locus as unstable and generating a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; (d) If the microsatellite instability score exceeds the population-trained threshold of the population of microsatellite loci in the sample, classifying the MSI status of the sample as unstable and identifying an unstable sample; and (e) Comparing the microsatellite instability status of the sample with one or more comparator results indexed by one or more therapies to identify one or more customized therapies for treating the disease in the subject A method comprising. (Item 6) A method for treating a disease in a subject, comprising: (a) Quantifying a number of different repeat lengths present in each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in a sample; (b) Comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci with the site-specific trained threshold of the given microsatellite locus; (c) If the site score of the given microsatellite locus exceeds the site-specific trained threshold of the given microsatellite locus, calling the given microsatellite locus as unstable and generating a microsatellite instability score including a number of unstable microsatellite loci from the plurality of microsatellite loci; (d) classifying the MSI status of the sample as unstable if the microsatellite instability score exceeds a population-trained threshold for the population of microsatellite loci in the sample, thereby identifying an unstable sample; (e) comparing the microsatellite instability status of the sample with one or more comparator results indexed with one or more therapies to identify one or more customized therapies for treating the disease in the subject; and (f) administering at least one of the identified customized therapies to the subject if the microsatellite instability status of the sample and the comparator result substantially match, thereby treating the disease in the subject. A method comprising: (Item 7) 1. A method of treating a disease in a subject, comprising administering to the subject one or more customized therapies, thereby treating the disease in the subject, wherein the customized therapies comprise: (a) quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate a site score for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in a sample; (b) comparing the site score of a given microsatellite locus for each of the plurality of microsatellite loci to a site-specific trained threshold for the given microsatellite locus; (c) calling the given microsatellite locus as unstable if the site score for the given microsatellite locus exceeds the site-specific trained threshold for the given microsatellite locus, generating a microsatellite instability score that includes a number of unstable microsatellite loci from the plurality of microsatellite loci; (d) If the microsatellite instability score exceeds the population-trained threshold of the population of microsatellite loci in the sample, classifying the MSI status of the sample as unstable and identifying unstable samples; (e) comparing the microsatellite instability status of the sample with one or more comparator results indexed by one or more therapies; and (f) identifying one or more customized therapies for treating the disease in the subject if the microsatellite instability status of the sample substantially matches the comparator results The method identified by. (Item 8) The method according to any one of the preceding items, wherein the site score of the plurality of microsatellite loci includes a likelihood score. (Item 9) The method according to item 8, wherein the likelihood score includes a probability log-likelihood based score that distinguishes a biological signal derived from a large number of cfDNA fragments of somatic origin in the sample from noise generated from artifacts after sample collection. (Item 10) Determining the probability log-likelihood based score for individual microsatellite loci in the sequence information from the sample using at least two parameters, wherein at least a first parameter includes an allele frequency and at least a second parameter includes at least one error mode, the method according to item 9. (Item 11) The method according to item 10, wherein the allele frequency includes the frequency of nucleic acids including different repeat lengths in the sequence information from the sample. (Item 12) The method according to item 10, wherein the at least one error mode includes a random error mode and a strand-specific error mode. (Item 13) The site score of the plurality of microsatellite loci is (a) A score for measuring the support of the observed sequence for the null hypothesis that the given microsatellite locus is stable, and (b) A score for measuring the support of the observed sequence for the alternative hypothesis that the given microsatellite locus is unstable The method according to any one of the preceding items, comprising a difference or ratio of. (Item 14) The method according to any one of the preceding items, wherein the site scores of the plurality of microsatellite loci are generated using one or more of a likelihood criterion, a log-likelihood criterion, a posterior probability criterion, an Akaike information criterion (AIC), and a Bayesian information criterion. (Item 15) The method according to any one of the preceding items, wherein the site scores of the plurality of microsatellite loci include site scores based on the Akaike information criterion (AIC) for testing for the presence of somatic indels in the plurality of microsatellite loci. (Item 16) A site score based on a given AIC is given by the following formula: AIC = k - log-likelihood The method according to item 15, wherein k is the number of parameters used in the model and is calculated using. (Item 17) The method according to item 10 or 16, comprising the step of estimating the parameters of the model using maximum likelihood estimation (MLE). (Item 18) The method according to item 17, comprising the step of determining the MLE using the Nelder-Mead algorithm. (Item 19) The following formula: AIC0 = k - log(Pr(obs|β,γ)) including the step of calculating the null hypothesis score of the model using the formula, where AIC0 is the null hypothesis, k is the number of parameters used in the model, Pr is the probability, obs includes the repeat length of the observed sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, and γ is at least one random error parameter, the method according to item 16. (Item 20) The following formula: AIC min =min α (k - log(Pr(obs|β,γ,α)) including the step of calculating the alternative hypothesis score of the model using the formula, where AIC min is the alternative hypothesis, min α is the minimization effect for all values of α, k is the number of parameters used in the model, Pr is the probability, obs includes the repeat length of the observed sequencing reads covering the given microsatellite locus, β is at least one strand-specific error parameter, γ is at least one random error parameter, α is at least one allele frequency, where α is a vector of allele frequencies such that the sum of one or more α i is equal to 1, the method according to item 16. (Item 21) The following formula: ΔAIC = AIC0 - AIC min including the step of detecting the site score using the formula, the method according to any one of the preceding items. (Item 22) γ is (a) the error rate at the read level where the microsatellite length observed within the sequencing read is 1 repeat unit longer than the expected microsatellite length for the strand of the originating nucleic acid molecule; and / or (b) The read-level error rate where the microsatellite length observed within the sequencing read is one repeat unit shorter than the predicted microsatellite length for the strand of the originating nucleic acid molecule The method according to any one of the preceding items, comprising (Item 23) β is (a) The strand-level error rate where the predicted microsatellite length of the sense strand is one repeat unit longer than the predicted microsatellite length of the nucleic acid-derived molecule; (b) The strand-level error rate where the predicted microsatellite length of the antisense strand is one repeat unit longer than the predicted microsatellite length of the nucleic acid-derived molecule; (c) The strand-level error rate where the predicted microsatellite length of the sense strand is one repeat unit shorter than the predicted microsatellite length of the nucleic acid-derived molecule; and / or (d) The strand-level error rate where the predicted microsatellite length of the antisense strand is one repeat unit shorter than the predicted microsatellite length of the nucleic acid-derived molecule The method according to any one of the preceding items, comprising (Item 24) The method according to any one of the preceding items, comprising the step of calling the given microsatellite locus as unstable when the site score of the given microsatellite locus statistically exceeds the site-specific trained threshold of the given microsatellite locus. (Item 25) The method according to any one of the preceding items, wherein the sample comprises a mutant allele ratio. (Item 26) The method according to any one of the preceding items, wherein the sample comprises a tumor ratio. (Item 27) The method according to Item 26, wherein the tumor ratio comprises the maximum mutant allele frequency (MAF) of all somatic mutations identified in the nucleic acid in the sample. (Item 28) The method according to item 27, wherein the tumor proportion is less than about 0.05%, about 0.1%, about 0.2%, about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14% or about 15% of all nucleic acids in the sample. (Item 29) The method according to any one of the preceding items, wherein the plurality of microsatellite loci includes all of the population of the microsatellite loci. (Item 30) The method according to any one of items 1 to 28, wherein the plurality of microsatellite loci includes a subset of the population of the microsatellite loci. (Item 31) The method according to any one of the preceding items, comprising determining the site-specific trained threshold and / or the population-trained threshold from sequence information from a population of microsatellite loci in one or more training DNA samples. (Item 32) The method according to item 31, wherein the training DNA sample includes a non-tumor cfDNA sample. (Item 33) The method according to item 31, wherein the training DNA sample includes DNA from one or more tumor types. (Item 34) The method according to any one of the preceding items, having at least about 94% sensitivity at a detection limit (LOD) of about 0.2% tumor proportion of nucleic acids in the sample. (Item 35) The method according to any one of the preceding items, having at least about 99% specificity for non-tumor DNA in the sample. (Item 36) The method according to any one of the preceding items, wherein the determined MSI status of the sample includes at least about 95%, 96%, 97%, 98% or 99% agreement with the corresponding MSI status of the sample determined using a PCR-based MSI evaluation technique over a tumor proportion range of about 1.4% to about 15%. (Item 37) The method according to item 36, wherein the match is 100%. (Item 38) The method according to any one of the preceding items, comprising classifying the MSI status of the sample as high MSI (MSI-H) when the microsatellite instability score is the number of unstable microsatellite loci exceeding a value of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 30, about 40, about 50 or more from the plurality of microsatellite loci. (Item 39) The method according to any one of the preceding items, comprising classifying the MSI status of the sample as high MSI (MSI-H) when the number of unstable microsatellite loci constitutes about 0.1%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 15%, about 20% or about 25% of the plurality of microsatellite loci. (Item 40) The cancer according to any one of items 5 to 7, including cancer comprising at least one tumor type selected from the group consisting of biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, uveal melanoma, choroidal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myelogenous (CML), chronic myelomonocytic (CMML), liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostatic adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma. (Item 41) The method according to any one of items 5 to 7, wherein the therapy comprises at least one immunotherapy. (Item 42) The method according to item 41, wherein the immunotherapy comprises at least one checkpoint inhibitory antibody. (Item 43) The method according to item 41, wherein the immunotherapy comprises an antibody against PD-1, PD-2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40 or CD47. (Item 44) The method according to item 41, wherein the immunotherapy comprises administration of a pro-inflammatory cytokine against at least one tumor type. (Item 45) The method according to item 41, wherein the immunotherapy comprises administration of T cells against at least one tumor type. (Item 46) The method according to any one of the preceding items, further comprising the step of obtaining the sample from the subject. (Item 47) The method according to item 46, wherein the sample is selected from the group consisting of tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, feces, synovial fluid, cerebrospinal fluid, and saliva. (Item 48) The method according to item 46, wherein the subject is a mammalian subject. (Item 49) The method according to item 48, wherein the mammalian subject is a human subject. (Item 50) The method according to item 46, wherein the sample contains cell-free nucleic acid. (Item 51) The method according to item 50, wherein the sample contains circulating tumor nucleic acid. (Item 52) The method according to any one of the preceding items, further comprising the step of receiving the sequence information generated from the sample, wherein the sequence information includes cfDNA sequencing reads from a population of the microsatellite loci in the sample. (Item 53) The method according to any one of items 46 to 52, comprising the step of amplifying one or more segments of the nucleic acid in the sample to generate at least one amplified nucleic acid. (Item 54) The method according to any one of items 46 to 53, further comprising the step of sequencing the nucleic acid from the sample to generate the sequence information. (Item 55) The method according to any one of items 46 to 54, wherein the sequence information is obtained from a targeted segment of the nucleic acid in the sample, and the targeted segment is obtained by selectively enriching one or more regions from the nucleic acid in the sample prior to sequencing. (Item 56) The method according to item 55, further comprising the step of amplifying the obtained targeted segment before sequencing. (Item 57) The method according to any one of items 46 to 56, further comprising the step of binding one or more adapters containing barcodes to the nucleic acid before sequencing. (Item 58) The method according to any one of items 54 to 57, wherein the sequencing is selected from the group consisting of targeted sequencing, intron sequencing, exome sequencing, and whole genome sequencing. (Item 59) The method according to any one of items 54 to 58, comprising the step of sequencing at least about 50, about 100, about 150, about 200, about 250, about 500, about 750, about 1,000, about 1,500, about 2,000, or more targeted genomic regions in the nucleic acid of the sample to generate the sequence information. (Item 60) The method according to any one of the preceding items, wherein the number of different repeat lengths includes the frequency of each different repeat length present in each of the plurality of microsatellite loci. (Item 61) The method according to any one of the preceding items, wherein at least a part of the method is executed by a computer. (Item 62) The method according to any one of the preceding items, further comprising the step of generating a report that optionally includes information regarding the instability status of the sample, information regarding the instability score of the sample, and / or information regarding the one or more customized therapies for the treatment of the disease in the subject. (Item 63) The method or system according to item 62, further comprising the step of communicating the report to a third party such as the subject or a healthcare provider. (Item 64) The method according to any one of the preceding items, further comprising classifying the repetitive nucleic acid instability status of the nucleic acid sample as stable if the repetitive nucleic acid instability score is at or below the population-trained threshold of the population of the repetitive nucleic acid locus in the nucleic acid sample. (Item 65) The method according to any one of the preceding items, further comprising classifying the repetitive DNA instability status of the sample as stable if the repetitive DNA instability score is at or below the population-trained threshold of the population of the repetitive DNA locus in the sample. (Item 66) The method according to any one of the preceding items, further comprising classifying the microsatellite instability status of the sample as stable if the microsatellite instability score is at or below the population-trained threshold of the population of the microsatellite locus in the sample.

Claims

1. A method for obtaining a site score (SS) generated from sequence information of a number of different repeat lengths present at each of a number of microsatellite loci as an index for determining the microsatellite instability (MSI) status of a sample, the sample comprising cell-free nucleic acid, the method comprising: (a) quantifying a number of different repeat lengths present at each of a plurality of microsatellite loci from sequence information to generate the SS for each of the plurality of microsatellite loci, wherein the sequence information is from a population of microsatellite loci in the sample; The SS of the plurality of microsatellite loci is (i) (A) A score that measures the support of the observed sequence for the null hypothesis that a given microsatellite locus is stable; and (B) A score that measures the support of the observed sequence for the alternative hypothesis that a given microsatellite locus is unstable; and (ii) generated using one or more of a likelihood criterion, a log likelihood criterion, a posterior probability criterion, and a Bayesian information criterion; The method includes determining a probabilistic log likelihood-based score for each microsatellite locus in the sequence information from the sample using at least two parameters, the at least two parameters being: (iii) an allele frequency (AF) comprising the frequency of nucleic acids comprising different repeat lengths in the sequence information from the sample; and (iv) at least one error mode, including a random error mode and a strand-specific error mode; and (b) comparing the SS of a given microsatellite locus to a site-specific trained threshold for the given microsatellite locus, for each of the plurality of microsatellite loci; (c) calling the given microsatellite locus as unstable if the SS of the given microsatellite locus exceeds the site-specific trained threshold for the given microsatellite locus, and generating a microsatellite instability (MI) score that includes a number of unstable microsatellite loci from the plurality of microsatellite loci; and (d) classifying the MSI status of the sample as unstable if the MI score exceeds a population-trained threshold for the population of microsatellite loci in the sample, thereby determining the MSI status of the sample; A method comprising: (e) comparing the MSI status of the sample to one or more comparator results indexed with one or more therapies; 13. The method of claim 1, wherein if the MSI status of the sample and the comparator result substantially match, the subject is a candidate for one or more customized therapies to treat disease in the subject.

3. At least one error mode including a random error mode and a strand-specific error mode, (1) the read-level error rate where the observed microsatellite length within a sequencing read is one repeat unit longer than the expected microsatellite length for the strand of the original nucleic acid molecule; (2) the read-level error rate where the observed microsatellite length within a sequencing read is one repeat unit shorter than the expected microsatellite length for the strand of the original nucleic acid molecule; (3) a strand-level error rate in which the predicted microsatellite length of the sense strand is one repeat unit longer than the predicted microsatellite length of the nucleic acid-derived molecule; (4) a strand-level error rate in which the predicted microsatellite length of the antisense strand is one repeat unit longer than the predicted microsatellite length of the nucleic acid-derived molecule; (5) a strand-level error rate in which the predicted microsatellite length of the sense strand is one repeat unit shorter than the predicted microsatellite length of the nucleic acid-derived molecule; and / or (6) a strand-level error rate in which the predicted microsatellite length of the antisense strand is one repeat unit shorter than the predicted microsatellite length of the nucleic acid-derived molecule; The method according to any one of claims 1 to 2, comprising:

4. determining the site-specific trained threshold and / or the population trained threshold from sequence information from a population of microsatellite loci in one or more training DNA samples; 4. The method of any one of claims 1-3, wherein the training DNA samples comprise non-tumor cfDNA samples, DNA from one or more tumor types, or both.

5. A method according to any one of claims 1 to 4, having a sensitivity of at least 94% at a limit of detection (LOD) of 0.2% tumor fraction of nucleic acid in the sample.

6. A method described in any one of claims 1 to 5, having a specificity of at least 99% for non-tumor DNA in the sample.

7. A method according to any one of claims 1 to 6, comprising a step of determining the MSI status of the sample having at least 95%, 96%, 97%, 98% or 99% concordance with the corresponding MSI status of the sample determined using a PCR-based MSI assessment technique across a tumor fraction range of 1.4% to 15%, e.g. the concordance is 100%.

8. The method of claim 1, further comprising classifying the MSI status of the sample as MSI-H based on an MI score exceeding 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50 or more unstable microsatellite loci from the plurality of microsatellite loci.

9. The method of claim 1, further comprising classifying the MSI status of the sample as MSI-H based on a number of unstable microsatellite loci comprising 0.1%, 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20% or 25% of the plurality of microsatellite loci.

10. (a) The disease is selected from the group consisting of biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, renal clear cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic (CLL), chronic myelogenous (CML), chronic myelomonocytic (CMML), liver cancer, liver cancer, hepatoma, cell carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B cell lymphoma, non-Hodgkin's lymphoma, diffuse large B cell lymphoma, mantle cell lymphoma, T cell lymphoma, non-Hodgkin's lymphoma, precursor T lymphoblastic lymphoma / leukemia, peripheral T cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal carcinoma, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, and uterine sarcoma; (b) the therapy comprises at least one immunotherapy, e.g., the immunotherapy comprises: at least one checkpoint inhibitor antibody; antibodies against PD-1, PD-2, PD-L1, PD-L2, CTLA-4, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, CD40 or CD47; Administration of a pro-inflammatory cytokine to at least one tumor type; or Administration of T Cells Against at Least One Tumor Type The method of claim 2 , comprising:

11. The method of claim 1, wherein the sample is obtained from a subject.

12. The method of any one of claims 1 to 11, wherein the sample is selected from the group consisting of tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, feces, synovial fluid, spinal fluid and saliva.

13. The method of any one of claims 1 to 12, wherein the subject is a mammalian subject.

14. The method of any one of claims 1 to 13, wherein the sample contains circulating tumor nucleic acid.

15. A method according to any one of claims 1 to 14, wherein the method comprises a step of receiving the sequence information generated from the sample, the sequence information comprising cfDNA sequencing reads from the population of microsatellite loci in the sample.

16. A method according to any one of claims 1 to 15, wherein the sequence information is obtained from a targeted segment of nucleic acid in the sample.

17. The method of claim 1, wherein the sample comprises a mutant allele proportion and / or the sample comprises a tumor proportion.

18. The method of claim 17, wherein the tumor fraction comprises a maximum mutated allele fraction (MAF) of all somatic mutations identified in the nucleic acid in the sample, and the tumor fraction is less than 0.05%, 0.1%, 0.2%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14% or 15% of all nucleic acid in the sample.

19. A method according to any one of claims 1 to 18, wherein the number of different repeat lengths comprises the frequency of each different repeat length present in each of the plurality of microsatellite loci.

Citation Information

Patent Citations

  • Variant based disease diagnostics and tracking

    US20170213008A1

  • Methods for multi-resolution analysis of cell-free nucleic acids

    WO2018064629A1