Techniques for detecting minimal residual disease
By using trinucleotide context error rates and statistical hypothesis testing on the same biological sample, the method enhances MRD detection accuracy, reducing false positives and negatives in cancer recurrence identification.
Patent Information
- Application Number
- JP2024571845
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-06
- Filing Date
- 2023-06-06
- Publication Date
- 2025-07-15
AI Technical Summary
Current methods for detecting minimal residual disease (MRD) in cancer patients are inaccurate due to sequencing errors, leading to false positives and negatives, as they assume similar error rates between healthy and cancerous samples, which is not always true.
A method that uses sequencing data from the same biological sample to identify sequencing errors by monitoring trinucleotide context (TNC) error rates, grouping them into clusters, and applying a statistical hypothesis test to distinguish between expected and actual mutations, thereby improving the accuracy of MRD detection.
This approach significantly reduces false positives and negatives in MRD detection, allowing for earlier and more reliable identification of cancer recurrence.
Smart Images

Figure 2025522347000001_ABST
Abstract
Description
Technical Field
[0001] Related Patent Application This patent application claims the benefit of UK (GB) Patent Application No. 2208273.9, entitled "Techniques For Detecting Minimum Residual Disease", filed on Jun. 6, 2022, as indicated by Mewburn Ellis Attorney Docket No. 008275836. The entire content of the above patent application, including all texts, tables, and drawings, is incorporated herein by reference.
Background Art
[0002] Background A central challenge in cancer treatment is the early detection of minimal residual disease following cancer treatment. Minimal residual disease (MRD) can generally be an indicator of cancer recurrence that occurs before standard surveillance imaging. One strategy for identifying MRD is to monitor biological samples from patients for circulating tumor DNA (ctDNA) that may be of cancer origin.
Summary of the Invention
Means for Solving the Problems
[0003] Summary Some embodiments of the present disclosure provide a method for determining whether sequencing data of a biological sample of a subject provides an indicator that the subject has minimal residual disease. The method uses at least one computer hardware processor to: (A) obtain sequencing data, where the sequencing data has been previously generated by sequencing a biological sample of the subject, and the sequencing data includes sequence reads that cover positions monitored for mutations; (B) use at least a first subset of the sequence reads to determine a first value indicating the expected number of mutations present in the sequencing data due to sequencing errors, where determining includes using the first subset of the sequence reads to identify a plurality of trinucleotide context (TNC) error rates for a plurality of TNC error types, grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, using at least some of the plurality of TNC error rates to identify a TNC group error rate for the plurality of TNC error rate groups, and using the TNC group error rate to determine a first value indicating the expected number of mutations present in the sequencing data; (C) use at least a second subset of the sequence reads to identify a second value indicating the actual number of mutations present at positions monitored for mutations; and (D) use the first value indicating the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicating the actual number of mutations present in the sequencing data at positions monitored for mutations to determine whether the sequencing data provides an indicator that the subject has minimal residual disease.
[0004] In some embodiments, the sequence reads cover at least 10 positions that are being monitored for mutations. In some embodiments, the sequence reads cover from 10 to 200 positions that are being monitored for mutations. In some embodiments, the sequence reads cover from 50 to 200 positions that are being monitored for mutations.
[0005] In some embodiments, the method further comprises obtaining the sequencing data by sequencing a biological sample. In some embodiments, the method further comprises obtaining a biological sample from a body fluid of a subject. In some embodiments, the biological sample comprises circulating tumor DNA (ctDNA). In some embodiments, each of the sequence reads covers at least one of the positions that are being monitored for mutations. In some embodiments, the sequence reads are obtained using whole exome sequencing. In some embodiments, the sequence reads are obtained using a targeted gene sequencing panel. In some embodiments, the targeted gene sequencing panel targets sequences that cover the positions that are being monitored for mutations. In some embodiments, primers are used to amplify sequences that include the positions that are being monitored for mutations. In some embodiments, the sequences targeted by the targeted gene sequencing panel are determined using sequence data from the subject's primary tumor.
[0006] In some embodiments, the first subset of array reads and the second subset of array reads are the same. In some embodiments, (B) is performed using at least the first subset of array reads and one or more array reads in the sequencing data that do not cover the positions being monitored for mutations. In some embodiments, performing (B) further includes generating a consensus array read using at least the first subset of array reads, each of the consensus array reads being generated from array reads within at least the first subset of array reads that are associated with their respective common unique molecular identifier (UMI), and identifying each of the plurality of trinucleotide context (TNC) error rates for the plurality of TNC error types is performed using the generated consensus array reads.
[0007] In some embodiments, each of the consensus array reads is generated from at least a threshold number of array reads associated with their respective common UMI. In some embodiments, the threshold number of array reads is between 2 and 20.
[0008] In some embodiments, the method further includes selecting a subset of the consensus array reads, and identifying each of the plurality of TNC error rates for the plurality of TNC error types is performed using only the selected subset of the consensus array reads. In some embodiments, the consensus array reads include a plus-strand consensus array read and a minus-strand consensus array read, and selecting the subset is performed using criteria that apply a similarity measure between the corresponding plus-strand consensus array read and the minus-strand consensus array read. In some embodiments, the consensus array reads include a plus-strand consensus array read and a minus-strand consensus array read, and selecting a subset of the consensus reads uses one or more criteria applied to the plus-strand consensus array read and the minus-strand consensus array read.
[0009] In some embodiments, the method further includes identifying a similarity measure between a corresponding plus-strand consensus sequence read and a minus-strand consensus sequence read. In some embodiments, the consensus sequence reads include a plus-strand consensus sequence read and a minus-strand consensus sequence read, and selecting a subset is performed using criteria applied to the relative numbers of the plus-strand consensus sequence reads and the corresponding minus-strand consensus sequence reads.
[0010] In some embodiments, identifying, using the consensus sequence reads, a plurality of TNC error rates for respective ones of a plurality of trinucleotide context (TNC) error types includes identifying the plurality of TNC error rates using a background region of the consensus sequence reads, the positions being monitored for mutations include a first position, the consensus sequence reads include a first consensus sequence read covering the first position, the background region includes a first background region of the first consensus sequence read, and the first background region includes nucleotides within the first consensus sequence read that are at least a first threshold distance from the first position. In some embodiments, the background region does not include the positions being monitored for mutations.
[0011] In some embodiments, the consensus sequence reads include a first group of plus-strand consensus sequence reads associated with a plus-strand primer binding sequence at the 3' end of each of the plus-strand consensus sequence reads for a first position of the positions being monitored for mutations, and a second group of minus-strand consensus sequence reads associated with a minus-strand primer binding sequence at the 3' end of the minus-strand consensus sequence reads in a second group. Identifying each of a plurality of trinucleotide context (TNC) error rates using the consensus sequence reads includes identifying the plurality of TNC error rates by using nucleotides in any sequence read within the first group of plus-strand consensus sequence reads located within a second threshold distance of the plus-strand primer binding sequence and nucleotides in any sequence read within the second group of minus-strand consensus sequence reads located within a third threshold distance of the minus-strand primer binding sequence. In some embodiments, identifying each of a plurality of trinucleotide context (TNC) error rates using the consensus sequence reads includes identifying the frequency of occurrence of each of the TNC error types within the consensus sequence reads.
[0012] In some embodiments, identifying each of a plurality of trinucleotide context (TNC) error rates using the consensus sequence reads includes identifying the plurality of TNC error rates from a background region of the consensus sequence reads, the consensus sequence reads including a first consensus sequence read, the background region including a first background region of the first consensus sequence read, and the TNC error rates being identified based on the frequency with which each of the TNC error types occurs in the first background region of the first consensus sequence read. In some embodiments, the TNC error type corresponds to a mutation at any position of a given TNC.
[0013] In some embodiments, each of the TNC error types corresponds to a specific mutation of the central nucleotide within a given TNC.
[0014] In some embodiments, the method further includes determining a confidence interval for the TNC error rate before identifying a plurality of trinucleotide context (TNC) error rates and before grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, and selecting at least some of the plurality of TNC error rates for grouping using a criterion applied to the confidence interval of the TNC error rate. In some embodiments, grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups includes clustering the plurality of TNC error rates. In some embodiments, grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups includes grouping using partitioning around medoids (PAM) clustering. In some embodiments, grouping at least some of the plurality of TNC error rates includes grouping into four TNC error rate groups.
[0015] In some embodiments, determining a first value indicative of the expected number of mutations present in the sequencing data is performed using at least some of the TNC group error rates and the number of times at least some of the positions being monitored for mutations are covered by sequence reads within a first subset of the sequence reads. In some embodiments, determining a first value indicative of the expected number of mutations present in the sequencing data includes determining the first value as a weighted linear combination of the TNC error group rates, each one of a particular of the TNC error group rates weighted by the number of times the position being monitored corresponds to a TNC error type belonging to the TNC error group and covered by sequence reads within a first subset of the sequence reads.
[0016] Performing (C) further includes generating a second consensus array read using at least a second subset of the array reads, each of the second consensus array reads being generated from the array reads within at least the second subset of the array reads associated with respective common unique molecular identifiers (UMIs), and identifying a second value indicative of the actual number of mutations present at positions being monitored for mutations is performed using the second consensus array read.
[0017] In some embodiments, (D) is performed using a statistical hypothesis test with a null hypothesis by comparing the second value to a distribution associated with the null hypothesis, the distribution having one or more parameters that depend on the first value. In some embodiments, the distribution is a Poisson distribution having a mean value (λ) set to the first value. In some embodiments, using the statistical hypothesis test includes determining a measure of the likelihood of observing the actual number of mutations indicated by the second value under the null hypothesis. In some embodiments, (D) is performed using a one-sided Poisson hypothesis test. In some embodiments, using the one-sided Poisson hypothesis test includes setting the mean value (λ) of the Poisson distribution to the first value and determining a measure of the likelihood of observing the actual number of mutations indicated by the second value under the Poisson distribution.
[0018] In some embodiments, determining whether the sequencing data provides an indication that the subject has minimal residual disease uses a measure of likelihood. In some embodiments, if the second value indicates that the null hypothesis can be rejected, the subject is likely to have minimal residual disease. In some embodiments, (D) further includes providing an indication that the subject has minimal residual disease.
[0019] In some embodiments, the method comprises obtaining, using at least one computer hardware processor, one or more of additional sequencing data previously generated by sequencing one or more additional biological samples of a subject, wherein each of the one or more additional sequencing data comprises additional sequence reads covering positions being monitored for mutations; determining, for each of the additional sequence reads of the one or more additional sequencing data, a further first value indicative of the expected number of mutations present in each of the sequencing data due to sequencing errors, the determining comprising identifying a further plurality of trinucleotide context (TNC) error rates for each of a plurality of TNC error types, grouping at least some of the further plurality of TNC error rates into further plurality of TNC error rate groups, identifying a further TNC group error rate for the further plurality of TNC error rate groups using at least some of the further plurality of TNC error rates, and determining a further first value indicative of the expected number of mutations present in each of the sequencing data using the further TNC group error rate; identifying, using at least a second subset of the additional sequence reads, a further second value indicative of the actual number of mutations present at the positions; and determining whether each of the sequencing data provides an indicator that the subject has minimal residual disease using the further first value indicative of the expected number of mutations present in each of the sequencing data due to sequencing errors and the further second value indicative of the actual number of mutations present at the positions being monitored for mutations in each of the sequencing data.
[0020] The various aspects described above can be used in place of or in addition to aspects in any of the systems, methods, and / or processes described herein. Further, the system can be configured to operate according to a method using one or more of the above aspects. Such a system can include at least one computer hardware processor and at least one non-transitory computer-readable storage medium storing processor-executable instructions, which, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to execute such a method. Further, the non-transitory computer-readable medium can include processor-executable instructions that, when executed by at least one computer hardware processor of a data processing system, cause the at least one computer hardware processor to execute according to a method using one or more of the above aspects. Accordingly, the above is a non-limiting summary of the present invention, and the scope of the present invention is defined by the appended claims.
Brief Description of the Drawings
[0021]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10A
Figure 10B
Figure 10C
Figure 10D
Figure 11A
Figure 11B
Figure 11C
Figure 11D
Figure 11E
Figure 12
DETAILED DESCRIPTION OF THE INVENTION
[0022] Detailed Description Early detection of cancer relapse / recurrence is an important aspect of effective cancer treatment. One strategy for detecting cancer recurrence is to look for minimal residual disease (MRD) in biological samples collected from subjects with a history of cancer treatment. MRD can be identified by sequencing a biological sample from a subject to identify circulating tumor DNA (ctDNA), which is an indicator of cancer recurrence. CtDNA generally contains the same sequence as the subject's wild-type DNA, except for cancer-related mutations.
[0023] Identifying the presence of cancer-related mutations (indicators of MRD) in cfDNA or ctDNA is difficult for many reasons, including sequencing errors in the sequencing experiment. Sequencing errors can be introduced at multiple points throughout the process of sample acquisition, sample preparation, sequencing, and post-sequencing analysis of the data. Sequencing errors can, for example, lead to false-positive identification of MRD due to the false-positive appearance of cancer-related mutations. Therefore, when identifying MRD, a method for correcting the false-positive appearance of cancer-related mutations is needed.
[0024] Conventional methods correct sequencing errors by sequencing negative subject samples from healthy subjects (see, for example, Abbosh, Christopher et al., Nature 545.7655 (2017): 446-451). These methods assume that sequencing errors associated with healthy subjects and those associated with cancer patients are generally the same, and thus, if the MRD signal of a cancer subject significantly exceeds the MRD signal of a healthy subject, a positive MRD call is made. However, the assumption that sequencing errors are similar between healthy and cancer subjects is not true in all cases. Sequencing errors between sequencing data collected from different biological samples can depend on the collection of the biological sample, the preparation of the biological sample for sequencing, the sequencing instrument (both the type of instrument and the maintenance of the instrument), and the analysis of the sequencing data. As a result, conventional techniques for identifying how much sequencing error is present in the sequencing data of a sample based on the errors found in the sequencing data of other samples can lead to inaccurate estimates of the errors. For example, in a recent experiment, genomic sequencing data from the same cell line sequenced in different scenarios were compared, and false negatives (no mutations when mutations were expected) were found at a rate of 42-51%, and false positives (mutations present when mutations were not expected) were found at a rate of 5-8% (see Kim, Young-Ho et al., PloS one 14.9 (2019): e0222535). In contrast to conventional methods, the inventors have developed techniques for identifying sequencing errors present in a sample using sequencing data from substantially the same sample (e.g., the indicators of sequencing errors and MRD can be identified using the same sequencing data from the same biological sample).
[0025] This technique involves determining a sequencing error rate by monitoring the error rate in a nucleotide or group of nucleotides (i.e., nucleotide context (NC)) (e.g., a value representing the rate of inaccurate nucleotides identified at a position; inaccurate nucleotides are identifiable at a position due to events occurring during sample collection, preparation, sequencing, post-sequencing analysis, or any other scenario in which the sample or data is manipulated). Nucleotide context (NC) refers to a series of consecutive nucleic acids having a specific base within a nucleic acid sequence or sequence read. In some embodiments, the error rate in a single nucleotide (single nucleotide context) is monitored. In some embodiments, the error rate in a group of two nucleotides (dinucleotide context), three nucleotides (trinucleotide context), four nucleotides (tetranucleotide context), five nucleotides (pentanucleotide context), six nucleotides (hexanucleotide context), or more nucleotides is monitored. In some embodiments, the error rate within a group of trinucleotide contexts is monitored as described herein. And the estimated sequencing error rate can be compared to the actual number of mutations observed at the positions being monitored for mutations to identify an indicator of MRD. In some embodiments, this technique involves estimating sequencing errors from sequencing results at positions where cancer-related mutations are not being monitored (such a collection of sequence read positions can be referred to herein as the "background region").
[0026] Accordingly, some embodiments provide a computer-implemented method for determining whether sequencing data of a biological sample (e.g., plasma) of a subject (e.g., a human subject) provides an indicator that the subject has minimal residual disease. In some embodiments, the method comprises: (A) obtaining sequencing data, wherein the sequencing data has been previously generated by sequencing a biological sample of the subject, and the sequencing data comprises sequence reads that cover positions to be monitored for mutations (e.g., the positions can be determined by analyzing the results of sequencing the subject's primary tumor to identify positions useful for subsequent monitoring for MRD); and (B) using at least a first subset of the sequence reads (e.g., a first subset that can be selected based on data quality) to determine a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors, wherein determining comprises using the first subset of sequence reads to identify, for each of a plurality of NC error types (e.g., different types of mutations observable in a nucleotide context), a plurality of nucleotide context (NC) error rates (e.g., the rate at which nucleotide repeats can mutate due to sequencing errors) selected from single nucleotide context, dinucleotide context, trinucleotide context, 4-nucleotide context, 5-nucleotide context, and 6-nucleotide context, grouping at least some of the plurality of NC error rates into a plurality of NC error rate groups (e.g., using PAM clustering, k-nearest neighbor clustering, or agglomerative hierarchical clustering);Clustering can reduce statistical noise by increasing the number of array reads used to identify the NC group error rate used to determine the first value, identifying the NC group error rate of multiple NC error rate groups (e.g., the population weighted average of the NC error rates can be the NC group error rate) using at least some of the multiple NC error rates, and determining a first value indicating the expected number of mutations present in the sequencing data using the NC group error rate, and (C) using at least a second subset of the array reads to identify a second value indicating the actual number of mutations present at positions being monitored for mutations (e.g., the number of times a mutation is present at a position being monitored for mutations), and (D) using the first value indicating the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicating the actual number of mutations present in the sequencing data at positions being monitored for mutations (e.g., using a one-sided Poisson test where the first value is λ) to determine whether the sequencing data provides an indicator (e.g., the probability that the subject has MRD) that the subject has minimal residual disease.;
[0027] Embodiments, examples, claims, and figures herein are often referred to as trinucleotide context (TNC), but it is understood that the methods described in the embodiments, examples, claims, and figures herein can be performed using any suitable nucleotide context (e.g., single nucleotide context (SNC), dinucleotide context (DNC), trinucleotide context (TNC), 4-nucleotide context (4NC), 5-nucleotide context (5NC), 6-nucleotide context (6NC), etc.).
[0028] Accordingly, some embodiments provide a computer-implemented method of determining whether sequencing data of a biological sample (e.g., plasma) of a subject (e.g., a human subject) provides an indication that the subject has minimal residual disease. In some embodiments, the method comprises: (A) obtaining sequencing data that has been previously generated by sequencing a biological sample of the subject, the sequencing data including sequence reads that cover positions to be monitored for mutations (e.g., the positions can be determined by analyzing the results of sequencing the subject's primary tumor to identify positions useful for subsequent monitoring for MRD), and (B) using at least a first subset of the sequence reads (e.g., a first subset that can be selected based on data quality) to determine a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors, the determining comprising using the first subset of sequence reads to identify a plurality of nucleotide context (NC) error rates (e.g., the rate at which nucleotide repeats can mutate due to sequencing errors) for a plurality of different types of mutations observable in a nucleotide context (e.g., different types of mutations in the nucleotide context), and grouping at least some of the plurality of NC error rates into a plurality of NC error rate groups (e.g., using PAM clustering;Pooling can reduce statistical noise by increasing the number of sequence reads used to identify the NC group error rate used to determine the first value), identifying the NC group error rate of a plurality of NC error rate groups (e.g., the population weighted average of the NC error rates can be the NC group error rate) using at least some of the plurality of NC error rates, and determining a first value indicative of the expected number of mutations present in the sequencing data using the NC group error rate, and (C) using at least a second subset of the sequence reads to identify a second value indicative of the actual number of mutations present at positions being monitored for mutations (e.g., the number of times a mutation is present at a position being monitored for mutations), and (D) using the first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicative of the actual number of mutations present in the sequencing data at positions being monitored for mutations (e.g., using a one-sided Poisson test where the first value is λ) to determine whether the sequencing data provides an indicator (e.g., the probability or likelihood that the subject has MRD) that the subject has minimal residual disease.;
[0029] Accordingly, some embodiments provide a computer-implemented method for determining whether sequencing data of a biological sample (e.g., plasma) of a subject (e.g., a human subject) provides an indicator that the subject has minimal residual disease. In some embodiments, the method comprises: (A) obtaining sequencing data, the sequencing data having been previously generated by sequencing a biological sample of the subject, the sequencing data including sequence reads that cover positions to be monitored for mutations (e.g., the positions can be determined by analyzing the results of sequencing of the subject's primary tumor to identify positions useful for subsequent monitoring for MRD), and (B) using at least a first subset of the sequence reads (e.g., a first subset that can be selected based on data quality) to determine a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors, the determining comprising using the first subset of the sequence reads to identify a plurality of trinucleotide context (TNC) error rates (e.g., the rate at which a 3-nucleotide repeat can mutate due to sequencing error) for each of a plurality of different types of mutations observable in the TNC, and grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups (e.g., using PAM clustering;Pooling can reduce statistical noise by increasing the number of sequence reads used to identify the TNC group error rate used to determine the first value, identifying the TNC group error rate of a plurality of TNC error rate groups (e.g., the population weighted average of the TNC error rates can be the TNC group error rate) using at least some of the plurality of TNC error rates, and determining a first value indicative of the expected number of mutations present in the sequencing data using the TNC group error rate, and (C) using at least a second subset of the sequence reads to identify a second value (e.g., the number of times a mutation is present at a position being monitored for mutations) indicative of the actual number of mutations present at positions being monitored for mutations, and (D) using the first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicative of the actual number of mutations present in the sequencing data at positions being monitored for mutations (e.g., using a one-sided Poisson test where the first value is λ) to determine whether the sequencing data provides an indicator (e.g., the probability or likelihood that the subject has MRD) that the subject has minimal residual disease.;
[0030] In some embodiments, coverage and / or resolution play a major role in determining the optimal context size (e.g., NC) for error rate identification, e.g., where coverage refers to the average maximum number of observations of an error rate context given the depth of sequencing of a sample, and resolution refers to the total number of error rate contexts of a given size. Sometimes, the formula (N = 3 *According to 4^k), the larger the context size, the more context is provided, where, for example, "k" is the context size. The more context (i.e., the higher the resolution), the more accurate the error rate estimation caused by the bias of the array around the variant can be in many cases. This sometimes leads to a cost that is directly proportional to the potential coverage (depth / N). For example, in an example where the minimum depth of the sample is 10,000 reads, the trinucleotide context has the theoretical potential to detect the error rate up to 1 / 52 (1.9%) on average while still increasing the overall resolution compared to the dinucleotide context or the mononucleotide context. The inventors have found herein that the trinucleotide context is often the largest context size that provides an acceptable detectable error rate over many sequencing depths, and thus often provides technical advantages over larger or smaller nucleotide contexts.
[0031] In some embodiments, the biological sample can be plasma and can contain cell-free DNA and / or ctDNA. Aspects of biological samples are described herein, including the following section called "biological sample".
[0032] In some embodiments, the subject can be a human subject previously treated for cancer (e.g., lung cancer). Various subjects and cancer types are described herein, including the following section called "subject".
[0033] In some embodiments, the minimal residual disease indicator can be identified using a statistical test (e.g., a statistical hypothesis test, which can be one-sided or two-sided and can be, for example, a Poisson test). Aspects of statistical tests are described herein. In some embodiments, minimal residual disease (MRD) can generally be an indicator of cancer recurrence that occurs before standard surveillance imaging detects cancer recurrence. Aspects of minimal residual disease (MRD) are described, including the following section called "minimal residual disease (MRD)".
[0034] In some embodiments, obtaining sequencing data can include sequencing nucleic acids in a biological sample to obtain sequencing data. Aspects of sequencing data are described herein, including the following section referred to as "sequencing data". In some embodiments, sequence reads covering the sequences being monitored for mutations can further cover regions not being monitored for mutations (e.g., background regions). Thus, sequencing data from a biological sample of a subject can include sequence reads that monitor positions being monitored for mutations and can be used to identify sequencing errors. In some embodiments, the positions being monitored for mutations may have been previously determined by sequencing the subject's primary tumor. Aspects of the positions being monitored for mutations are described herein, including the following section referred to as "positions being monitored for mutations". In some embodiments, at least 10, 10 - 200, or 50 - 200 positions are monitored for mutations. Sequencing data can be obtained using a suitable method. In some embodiments, sequencing data can be obtained using whole genome sequencing. In some embodiments, sequencing data can be obtained using whole exome sequencing. In some embodiments, sequencing data can be obtained using a targeted gene sequencing panel or method. Aspects of the targeted gene sequencing panel are described herein, including the following section referred to as "sequencing data".
[0035] As described above, in some embodiments, the method includes using at least a first subset of the sequence reads to obtain a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors. In some embodiments, the first subset of sequence reads includes any number or combination of sequence reads in the sequencing data. For example, the first subset may include a consensus sequence read. In some embodiments, the consensus sequence read can be identified or specified using a suitable alignment method (e.g., pileup, Bowtie, BarraCUDA, BFAST, CUSHAW, ELAND, FASTA, SOAP, etc., variations or combinations thereof). The consensus sequence read can be determined using multiple sequence reads having the same unique molecular identifier (e.g., barcode). In particular, the first subset may include a deep consensus sequence read. In some embodiments, the deep consensus read can be a consensus read determined using at least 2 sequence reads, at least 3 sequence reads, at least 4 sequence reads, at least 5 sequence reads, at least 6 sequence reads, at least 7 sequence reads, at least 8 sequence reads, at least 9 sequence reads, at least 10 sequence reads, at least 15 sequence reads, or at least 20 sequence reads having the same unique molecular identifier.
[0036] In some embodiments, the first value can be determined using sequence reads that cover one or more positions being monitored for a plurality of mutations (cancer-related mutations). However, in other embodiments, the first value can be determined using one or more sequence reads that each cover one or more positions being monitored for a plurality of mutations (cancer-related mutations) and one or more sequence reads that do not cover any of the positions being monitored for the plurality of mutations. In still other embodiments, the first value can be determined using only sequence reads that do not cover any of the positions being monitored for cancer-related mutations.
[0037] In some embodiments, performing (B) includes generating consensus sequence reads using at least a first subset of the array reads, each of the consensus sequence reads being generated from the array reads within at least the first subset of the array reads associated with respective common unique molecular identifiers (UMIs). In some embodiments, generating consensus reads using UMIs can reduce polymerase chain reaction amplification biases that can occur during sample preparation for sequencing. Aspects of the consensus sequence reads are described herein. In some embodiments, each of the consensus sequence reads can be generated from at least a threshold number (e.g., 2-20) of array reads associated with respective common UMIs. In some embodiments, the method includes selecting a subset of the consensus sequence reads, and determining a plurality of NC error rates (e.g., trinucleotide context (TNC) error rates) for each of a plurality of NC or TNC error types is performed using only the selected subset of the consensus sequence reads. In some embodiments, the consensus sequence reads include a plus-strand consensus sequence read and a minus-strand consensus sequence read, and selecting the subset can be performed based on a similarity measure between the corresponding plus-strand consensus sequence read and the minus-strand consensus sequence read. In some embodiments, the consensus sequence reads include a plus-strand consensus sequence read and a minus-strand consensus sequence read, and selecting a subset of the consensus reads uses one or more criteria applied to the plus-strand consensus sequence read and the minus-strand consensus sequence read. In some embodiments, the method includes determining a similarity measure between the corresponding plus-strand consensus sequence read and the minus-strand consensus sequence read using a suitable statistical test. In some embodiments, the consensus sequence reads include a plus-strand consensus sequence read and a minus-strand consensus sequence read, and selecting the subset can be performed based on the relative numbers of the plus-strand consensus sequence read and the corresponding minus-strand consensus sequence read.Accordingly, in some embodiments, by including thresholds and making these selections, the consensus read may be more likely to represent the actual DNA sequence from the subject.
[0038] In some embodiments, nucleotide context (NC) refers to a particular base within a nucleic acid sequence or sequence read. In some embodiments, nucleotide context (NC) refers to a series of consecutive nucleic acids (e.g., 2, 3, 4, 5, 6, 7, 8, or more consecutive nucleic acids) having a particular base within a nucleic acid sequence or sequence read. An NC error type refers to a particular mutation (from the wild type and / or reference genome) within any given NC. For example, there are three error types that can occur in any single NC (e.g., for an A nucleotide, there is A>T, A>C, or A>G). In yet another example, for a dinucleotide context having the wild type sequence AT, the NC error types can include, but are not limited to, AA, AG, AC, GT, GA, GC, GG, CT, CA, CC, CG, TT, TA, TC, and TG. An NC error type can refer to the frequency with which each NC error type occurs within an NC. In some embodiments, only selected NC error types or selected pluralities of NC error types may be considered.
[0039] A trinucleotide context (TNC) refers to a series of three consecutive nucleic acids having a specific base within a nucleic acid sequence or sequence read (e.g., TAC). Aspects of the trinucleotide context are described herein, including the following section referred to as "trinucleotide context (TNC)". A TNC error type refers to a specific mutation (from a wild-type and / or reference genome) in any given TNC (e.g., A>T, A>C, or A>G). For example, if the expected (wild-type) TNC is TAC, the TNC error types can include, but are not limited to, TTC, TCC, AAC, TAG, CAC, GAC, TAT, TAA, and TGC. The TNC error types are further described herein, including the description with reference to Figure 3. The TNC error type can refer to the frequency at which each TNC error type occurs within the context of the TNC. In some embodiments, only a plurality of each TNC error type may be considered. For example, in some embodiments, a plurality of each TNC error type can refer to a TNC error type in which the central portion of the TNC is mutated and the first (3') and third (5') positions may not be mutated.
[0040] In some embodiments, using array reads (e.g., consensus array reads) to identify the plurality of trinucleotide context (TNC) error rates for each of the plurality of TNC error types includes identifying the plurality of TNC error rates using the background region of the consensus array reads, the positions being monitored for mutations include a first position, the consensus array reads include a first consensus array covering the first position, the background region includes a first background region of the first consensus array reads, and the first background region includes nucleotides within the first consensus array reads that are at least a first threshold distance away from the first position. Thus, the TNC error rate can be identified using TNCs that are a threshold distance away from the positions being monitored for mutations. In some embodiments, the threshold distance is used to exclude TNCs having high errors (e.g., it has been found that nucleotides at the end of an array read may be less reliable than nucleotides at the beginning of the array read). The TNC error rate and aspects of the background region are described in the following sections with reference to this specification and FIG. 3. In some embodiments, the background region does not include the positions being monitored for mutations.
[0041] In some embodiments, the consensus sequence reads include a first group of plus-strand consensus sequence reads associated with a plus-strand primer binding sequence at the 3’ end of each of the plus-strand consensus sequence reads at a first position among the positions being monitored for mutations, and a second group of minus-strand consensus sequence reads associated with a minus-strand primer binding sequence at the 3’ end of the minus-strand consensus sequence reads in a second group. Identifying a plurality of trinucleotide context (TNC) error rates of respective TNC error types using the consensus sequence reads includes using a nucleotide within any sequence read within the first group of plus-strand consensus sequence reads that is within a second threshold distance of the plus-strand primer binding sequence and a nucleotide within any sequence read within the second group of minus-strand consensus sequence reads that is within a third threshold distance of the minus-strand primer binding sequence to identify the plurality of TNC error rates. Thus, the TNC error rate can be identified using TNCs that are a threshold distance away from the start of the sequence reads determined by the location of the plus-strand primer binding sequence or the minus-strand primer binding sequence. Aspects of the second threshold and the third threshold are described, including the following section with reference to Figure 3.
[0042] In some embodiments, identifying the plurality of TNC error rates for each of the plurality of trinucleotide context (TNC) error types using consensus sequence reads includes identifying the plurality of TNC error rates from the background region of the consensus sequence reads, the consensus sequence reads include a first consensus sequence read, the background region includes a first background region of the first consensus sequence read, and the TNC error rate is identified based on the frequency at which each of the TNC error types occurs in the first background region of the first consensus sequence read. Thus, the TNC error rate can, in some embodiments, be calculated using only the background region of the sequence reads. In other embodiments, the TNC error rate can be calculated using both the background region and the positions being monitored for mutations, such embodiments relying on the understanding that the majority of TNCs in the sequence reads are not being monitored for mutations and thus the positions being monitored for mutations can be few in number such that including them does not likely change the sequencing errors much when identifying the sequencing errors.
[0043] In some embodiments, the method includes identifying and / or removing positions in the background region having an error rate of 0.5% or greater, 1% or greater, 1.5% or greater, 2% or greater, or 3% or greater. This can remove positions that may bias the estimation of background error (e.g., positions where the genomic sequence of a biological sample truly differs from the reference genomic sequence). In some embodiments, the reference genomic sequence is a patient-specific genomic sequence (e.g., wild-type sequence). In some embodiments, the reference genomic sequence is a reference genome (e.g., hg19).
[0044] In some embodiments, the method further includes, after identifying a plurality of trinucleotide context (TNC) error rates and before clustering at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, identifying a confidence interval (e.g., a 99% binomial confidence interval for each TNC error rate) of the TNC error rates; and based on the confidence intervals of the TNC error rates, selecting at least some of the plurality of TNC error rates to be clustered when the confidence intervals of all of the TNC error rates exceed a threshold (99% binomial confidence interval). By this step, before determining the first value, the highest TNC error rate (which may be abnormally high and may reduce the accuracy of the algorithm) can be removed. Aspects of selecting TNC error rates to be clustered are discussed herein, including the discussion with reference to FIG. 4. Thus, in some embodiments, high TNC error rates can be removed before determining the first value.
[0045] Aspects of clustering TNC error rates are described herein, including the description with reference to FIG. 5. In some embodiments, each TNC error rate group includes at least one TNC error rate. In some embodiments, clustering at least some of the plurality of TNC error rates into a plurality of TNC error rate groups includes clustering using a suitable clustering method, non-limiting examples of which include k-means clustering, agglomerative hierarchical clustering, partitioning around medoids (PAM) clustering, and the like. In some embodiments, clustering at least some of the plurality of TNC error rates into a plurality of TNC error rate groups includes clustering using partitioning around medoids (PAM) clustering. In some embodiments, clustering at least some of the plurality of TNC error rates includes clustering into four TNC error rate groups. Clustering the TNC error rates into TNC error rate groups can be performed in some embodiments to have sufficient sequence read depth for identifying the TNC group error rates.
[0046] In some embodiments, determining a first value indicative of the expected number of mutations present in the sequencing data can be performed using at least some of the TNC group error rates and the number of times at least some of the positions being monitored for mutations are covered by sequence reads within a first subset of the sequence reads. In some embodiments, the TNC group error rate for a given TNC error rate group can be identified by calculating the population weighted average of the TNC derived error rate groups. In some embodiments, the number of times at least some of the positions being monitored for mutations are covered by sequence reads within a first subset of the sequence reads can be identified by counting the number of times each position being monitored is covered by sequence reads within a subset of the sequence reads of the sequencing data. Aspects of the TNC group error rate are described herein, including the description with reference to FIG. 5.
[0047] In some embodiments, determining a first value indicative of the expected number of mutations present in the sequencing data includes determining the first value as a weighted linear combination of each one of the identified TNC error group rates weighted by the number of times the position being monitored is covered by sequence reads within a first subset of the sequence reads corresponding to the TNC error type belonging to the TNC error group. Thus, in some embodiments, the first value can be calculated based on information from all of the TNC error rate groups, thereby increasing the data used for the estimation of the TNC error rate and providing an improvement in the estimation of the TNC error rate.
[0048] The methods described herein for identifying and using TNC error group rates can be applied to any suitable NC (e.g., a single NC, dinucleotide context, TNC, etc.).
[0049] As described above, in some embodiments, the method includes identifying a second value indicative of the actual number of mutations present at positions being monitored for mutations using at least a second subset of the sequence reads. In some embodiments, the second subset of the sequence reads includes sequence reads that include the positions being monitored for mutations. In some embodiments, the first subset of the sequence reads and the second subset of the sequence reads can be the same subset of the sequence reads. In some embodiments, the second value can be identified based on counting the number of times each position being monitored for mutations mutates in the second subset of the reads. In some embodiments, performing (C) further includes generating a second consensus sequence read using at least a second subset of the sequence reads, each of the second consensus sequence reads can be generated from sequence reads within at least the second subset of the sequence reads associated with respective common unique molecular identifiers (UMIs), and identifying a second value indicative of the actual number of mutations present at positions being monitored for mutations can be performed using the second consensus sequence read. In some embodiments, the second subset of the consensus sequence reads can include consensus reads constructed using 2 to 20 reads. Thus, similar to the calculation of the first value, in some embodiments, the second value can be calculated using a consensus read for the same reason (e.g., to control for PCR amplification bias).
[0050] As described above, in some embodiments, the method includes determining whether sequencing data provides an indicator that a subject has minimal residual disease using a first value indicative of an expected number of mutations present in the sequencing data due to sequencing errors and a second value indicative of an actual number of mutations present in the sequencing data at a position being monitored for mutations. In some embodiments, (D) can be performed using a suitable statistical test having a null hypothesis (e.g., a one-sided Poisson test or a t-test), and the distribution associated with the null hypothesis has one or more parameters that depend on the first value. In some embodiments, the distribution can be a Poisson distribution having a mean value (λ) that can be set to the first value. In some embodiments, using the statistical test includes identifying a measure of the likelihood of observing the actual number of mutations indicated by the second value under the null hypothesis. In some embodiments, if the second value indicates that the null hypothesis can be rejected, the subject is likely to have minimal residual disease. In some embodiments, performing a one-sided Poisson hypothesis test includes setting the mean value (λ) of the Poisson distribution to the first value and identifying a measure of the likelihood of observing the actual number of mutations indicated by the second value under the Poisson distribution. In some embodiments, the sequencing data can provide an indicator that the subject has minimal residual disease (MRD) using the measure of likelihood from the Poisson test. In some embodiments, the indicator of MRD can be based on a p-value from a statistical test less than a predetermined alpha (e.g., p-value ≤ 0.01). In some embodiments, rejection of the null hypothesis of the Poisson test can be an indicator of MRD. In some embodiments, failure to reject the null hypothesis of the Poisson test may not be an indicator of MRD. In some embodiments, (D) further includes providing an indicator that the subject has minimal residual disease. Aspects of the indicator of MRD are described herein, including the following section called "Indicator of Minimal Residual MRD".
[0051] In some embodiments, the presence of MRD is determined using a chi-square test over each monitored position that compares observations to predictions (e.g., as in (D)). In some embodiments, positions with zero deep alternate observations (DAO) are filtered out.
[0052] In some embodiments, the presence of MRD is determined by subtracting a first value indicating the expected number of mutations in the sequencing data from a second value indicating the actual number of mutations present at positions being monitored for mutations, and determining whether the difference is greater than a threshold (e.g., as in (D)). In some embodiments, the threshold can be zero.
[0053] In some embodiments, any of the embodiments of this method can be repeated using additional biological samples collected from the subject, as described above. In some embodiments, the methods described herein include obtaining one or more additional sequencing data previously generated by sequencing one or more additional biological samples from the subject over time and analyzing these samples according to the embodiments described above, for monitoring for MRD. Thus, in some embodiments, a subject can be incrementally monitored for MRD over a period of weeks, months, years, or the entire remaining lifespan of the subject. In some embodiments, the method further includes recommending treatment to the subject if MRD is identified and not recommending treatment to the subject if MRD is not identified.
[0054] In some embodiments, the method is to obtain second sequencing data previously generated by sequencing a second biological sample (e.g., an additional biological sample) of a subject, the second sequencing data including second sequence reads that cover positions being monitored for mutations, and obtaining, using at least a first subset of the second sequence reads, a third value indicative of the expected number of mutations present in the second sequencing data due to sequencing errors (e.g., identified in the same manner as the first value), determining including identifying second multiple trinucleotide context (TNC) error rates for each of the multiple TNC error types, grouping at least some of the second multiple TNC error rates into second multiple TNC error rate groups, identifying a second TNC group error rate of the second multiple TNC error rates using at least some of the second TNC error rates of the second multiple TNC error rates, and using the second TNC group error rate (e.g., determined in the same manner as the first value but determined using the second sequencing data) to determine a third value indicative of the expected number of mutations present in the second sequencing data, determining, using at least a second subset of the second sequence reads, identifying a fourth value indicative of the actual number of mutations present at the positions, and using a third value indicative of the expected number of mutations present in the second sequencing data due to sequencing errors and a fourth value indicative of the actual number of mutations present in the second sequencing data at positions being monitored for mutations to determine whether the second sequencing data provides an indicator that the subject has minimal residual disease.
[0055] In some embodiments, the method uses at least one computer hardware processor to obtain one or more of additional sequencing data (e.g., second sequence data, third sequence data, etc.) previously generated by sequencing one or more additional biological samples of a subject (e.g., a second biological sample collected over time, a third biological sample, etc.), wherein each of the one or more of the additional sequencing data includes additional sequence reads that cover positions being monitored for mutations, and obtaining; for each of the one or more additional sequence reads of the additional sequencing data, using at least a first subset of the additional sequence reads to determine an additional first value indicative of the expected number of mutations present in each of the sequencing data due to sequencing errors, the determining including identifying a plurality of additional trinucleotide context (TNC) error rates for each of a plurality of TNC error types, grouping at least some of the plurality of additional TNC error rates into a plurality of additional TNC error rate groups, identifying an additional TNC group error rate for the plurality of additional TNC error rate groups using at least some of the plurality of additional TNC error rates, and determining an additional first value indicative of the expected number of mutations present in each of the sequencing data using the additional TNC group error rate, the determining; using at least a second subset of the additional sequence reads to identify an additional second value indicative of the actual number of mutations present at the positions; and using an additional first value indicative of the expected number of mutations present in each of the sequencing data due to sequencing errors and an additional second value indicative of the actual number of mutations present at the positions being monitored for mutations in each of the sequencing data to determine whether each of the sequencing data provides an indicator that the subject has minimal residual disease.
[0056] Some embodiments of the present disclosure provide a method for determining whether sequencing data of a biological sample of a subject has a mutation (e.g., a cancer-related mutation). In these embodiments, the method uses at least one computer hardware processor to: (A) obtain sequencing data, where the sequencing data has been previously generated by sequencing a biological sample of the subject, and the sequencing data includes sequence reads that cover positions monitored for mutations; (B) use at least a first subset of the sequence reads to determine a first value indicating the expected number of mutations present in the sequencing data due to sequencing errors, where determining includes using the first subset of the sequence reads to identify a plurality of trinucleotide context (TNC) error rates for each of a plurality of TNC error types, grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, using at least some of the plurality of TNC error rates to identify a TNC group error rate for each of the plurality of TNC error rate groups, and using the TNC group error rate to determine a first value indicating the expected number of mutations present in the sequencing data; (C) use at least a second subset of the sequence reads to identify a second value indicating the actual number of mutations present at positions monitored for mutations; and (D) use the first value indicating the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicating the actual number of mutations present in the sequencing data at positions monitored for mutations to determine whether the sequencing data provides an indicator that the subject has minimal residual disease.
[0057] The following is a more detailed disclosure of various concepts and embodiments related to indicating MRD described herein and methods and configurations for indicating MRD described herein.
[0058] FIG. 1 is a diagram showing an exemplary technique 100 for identifying and outputting an MRD 107 indicator of a subject 101. The technique 100 includes collecting a biological sample 102 (e.g., plasma) from the subject 101, sequencing a polynucleotide (e.g., ctDNA) from the biological sample using a sequencing device 103 to generate sequencing data 104, and inputting the sequencing data 104 and positions being monitored for mutations 106 into a computing device 105 (e.g., a laptop computer, a desktop computer, one or more servers, a cloud computing device, a smartphone, a tablet, and / or any other suitable computing device) configured to execute software 108 that executes the techniques described herein for detecting an MRD indicator on the computing device 105 when executed.
[0059] Subject In some embodiments, the subject can be a mammal (e.g., a human, non-human primate, dog, cat, horse, goat, sheep, mouse, or rat), a bird, an insect, an amphibian, a fish, or a research model organism (e.g., a mouse and a rat). In some embodiments, the subject can be a human. In some embodiments, the subject can be an adult human (e.g., 19 years old or older). In some embodiments, the subject can be a human child. In some embodiments, the subject can be a human infant.
[0060] In some embodiments, the subject may be in remission from a disease. In some embodiments, the subject may have been treated for a disease (e.g., cancer). In some embodiments, the subject may have been previously treated using one or more of surgery, chemotherapy, radiation therapy, immunotherapy, and / or hormone therapy. In some embodiments, the subject may be in remission from cancer. In some embodiments, the subject may be in remission from lung cancer, brain tumor, liver cancer, kidney cancer, cancer of the immune system, breast cancer, skin cancer, osteosarcoma, uterine cancer, prostate cancer, testicular cancer, or colon cancer. In some embodiments, the subject may be in remission from non-small cell lung cancer (NSCLC). In some embodiments, the subject may be in remission from small cell lung cancer (SCLC). In some embodiments, the subject may be in remission from lung adenocarcinoma. In some embodiments, the subject may be in remission from squamous cell carcinoma. In some embodiments, the subject may be in remission from melanoma. In some embodiments, the biological sample comprises ctDNA released into the subject's body fluid by cancer cells. In some embodiments, the subject may be in remission from a cancer selected from NSCLC, colorectal cancer (CRC), bladder cancer, pancreatic cancer, head and neck squamous cell carcinoma (HNSCC), breast cancer, and blood cancers (e.g., leukemia, lymphoma, and multiple myeloma). These cancers may in particular have a high likelihood of releasing DNA into the body fluid.
[0061] In some embodiments, the subject has or is suspected of having cancer or a tumor.
[0062] Biological sample As shown in FIG. 1, the biological sample 102 is collected from a subject. In some embodiments, the biological sample includes any cells, tissues, biological fluids, or bones from the subject or any other part of the subject. In some embodiments, the biological sample includes ctDNA of the subject. In some embodiments, the biological sample can be a tissue biopsy (e.g., a tumor biopsy). In some embodiments, the tissue can be brain tissue, lung tissue, liver tissue, kidney tissue, skin tissue, pancreatic tissue, connective tissue, muscle tissue, or nerve tissue. In some embodiments, the biological sample can be a biological fluid from the subject. In some embodiments, the biological fluid can be saliva, semen, vaginal secretions, urine, excrement, nasal mucus, sweat, earwax, cerebrospinal fluid, blood, serum, or plasma from the subject. In some embodiments, the biological sample can be plasma from the subject. In some embodiments, the biological sample from the subject can be blood or a blood product (e.g., serum or plasma), and the blood or blood product contains tumor cells and / or ctDNA. In some embodiments, the biological sample (e.g., a blood product, blood, plasma, or serum) contains cell-free DNA (cfDNA) or ctDNA. Cell-free DNA from a subject having cancer or a tumor often contains ctDNA.
[0063] In some embodiments, the biological sample includes, but is not limited to, any part of the subject's body including hair, skin (including parts of the epidermis, dermis, and / or subcutaneous tissue), pharynx, larynx, esophagus, stomach, bronchi, salivary glands, tongue, oral cavity, nasal cavity, vaginal cavity, anal cavity, bone, bone marrow, brain, thymus, spleen, small intestine, appendix, colon, rectum, anus, liver, bile duct, pancreas, kidney, ureter, bladder, urethra, uterus, vagina, vulva, ovaries, cervix, scrotum, penis, prostate, testicles, seminal vesicles, any fluid [blood (whole blood, serum, or plasma), saliva, tears, synovial fluid, cerebrospinal fluid, pleural effusion, pericardial fluid, ascites, and / or urine, etc.], and / or any type of tissue (e.g., can be collected from muscle tissue, epithelial tissue, connective tissue, or nerve tissue).
[0064] Any of the biological samples described in this specification can be collected using any suitable technique, including those as described in Biospecimens and biorepositories: from afterthought to science by Vaught et al. (Cancer Epidemiol Biomarkers Prev. 2012 Feb;21(2):253-5) and Biological sample collection, processing, storage and information management by Vaught and Henderson (IARC Sci Publ. 2011;(163):23-42).
[0065] In some embodiments, a plurality of biological samples may be collected from a subject, sequenced, and sequencing data obtained. In some embodiments, the plurality of biological samples may be sequentially collected from the subject over a specified time period, sequenced, and sequencing data obtained. In some embodiments, the specified time period may begin after the end of cancer treatment and may continue over the remaining lifespan of the subject. In some embodiments, the frequency at which biological samples are collected from the subject may be any frequency suitable for MRD monitoring. In some embodiments, biological samples may be collected from the subject weekly. In some embodiments, biological samples may be collected from the subject approximately twice a month. In some embodiments, biological samples may be collected from the subject approximately once a month. In some embodiments, biological samples may be collected from the subject approximately once every three months. In some embodiments, biological samples may be collected from the subject approximately once every six months. In some embodiments, biological samples may be collected from the subject at least twice a month. In some embodiments, biological samples may be collected from the subject at least once a month. In some embodiments, biological samples may be collected from the subject at least once every three months. In some embodiments, biological samples may be collected from the subject at least once every six months. In some embodiments, the frequency at which biological samples may be collected from the subject may be based on the type of disease being monitored in the subject (e.g., type of cancer), the expected likelihood of recurrence, and the rate of progression of the disease after recurrence.
[0066] Biological samples can be preserved using suitable methods. In some embodiments, biological samples can be preserved using cryopreservation. Non-limiting examples of cryopreservation include, but are not limited to, step-down freezing, blast freezing, direct plunge freezing, snap freezing, slow freezing using a programmable freezer, and vitrification. In some embodiments, biological samples can be preserved using lyophilization. In some embodiments, after a biological sample is collected from a subject, the biological sample can be placed in a container that already contains a preservative (e.g., RNALater for preserving RNA) and then frozen (e.g., by snap freezing). In some embodiments, such preservation in a frozen state can be performed immediately after collection of the biological body sample. In some embodiments, biological samples can be stored for a short time (e.g., up to 1 hour, up to 8 hours, or up to 1 day, or for several days) at either room temperature or 4°C in a preservative or in a buffer without a preservative before being frozen.
[0067] Non-limiting examples of preservatives include formalin solution, formaldehyde solution, RNALater or other equivalent solutions, TriZol or other equivalent solutions, DNA / RNA Shield or equivalent solutions, EDTA (e.g., Buffer AE (10 mM Tris·Cl; 0.5 mM EDTA, pH 9.0)), and other coagulants, and dextrose citrate (e.g., for blood samples).
[0068] In some embodiments, special containers may be used to collect and / or preserve biological samples. For example, a vacutainer may be used to preserve blood. In some embodiments, the vacutainer may contain a preservative (e.g., a coagulant or an anticoagulant). In some embodiments, the container in which a biological sample is preserved may be housed within a secondary container for better preservation purposes or for contamination avoidance purposes.
[0069] Any of the biological samples from the subjects described in this specification may be stored under any conditions that maintain the stability of the biological sample. In some embodiments, the biological sample may be stored at a temperature that maintains the stability of the biological sample. In some embodiments, the sample may be stored at 18°C to 28°C (e.g., 25°C). In some embodiments, the sample may be stored under refrigeration (e.g., 4°C). In some embodiments, the sample is stored under freezing conditions (e.g., -20°C). In some embodiments, the sample may be stored under ultra-low temperature conditions (e.g., -50°C to -800°C). In some embodiments, the sample may be stored under liquid nitrogen (e.g., -1700°C). In some embodiments, the biological sample may be stored at -60°C to -80°C (e.g., -70°C) for up to 5 years. In some embodiments, the biological sample may be stored at -60°C to -80°C (e.g., -70°C) for up to 1 month, up to 2 months, up to 3 months, up to 4 months, up to 5 months, up to 6 months, up to 7 months, up to 8 months, up to 9 months, up to 10 months, up to 11 months, up to 1 year, up to 2 years, up to 3 years, up to 4 years, or up to 5 years. In some embodiments, the biological sample may be stored for up to 5 years, up to 10 years, up to 15 years, or up to 20 years as described by any of the methods described in this specification.
[0070] Sequencing device As shown in FIG. 1, biological sample 102 can be sequenced using sequencing device 103. In some embodiments, sequencing device 103 can be sequenced using any suitable next-generation sequencing device or any high-throughput or ultra-parallel sequencing device. In some embodiments, sequencing device 103 can include any suitable sequencing device and / or any sequencing system including one or more devices. In some embodiments, the sequencing device used for sequencing a biological sample can be selected from any suitable platform known in the art including, but not limited to, Illumina®, SOLid®, Ion Torrent®, PacBio®, nanopore-based, Sanger sequencing, or 454TM. In some embodiments, the sequencing device used for sequencing a biological sample is an Illumina sequencing device (e.g., NovaSeq®, NextSeq®, HiSeq®, MiSeq®, or MiniSeq®).
[0071] Sequencing data As shown in FIG. 1, sequencing platform 103 processes sample 102 to generate sequencing data 104.
[0072] In some embodiments, sequencing data 104 can include sequence reads (e.g., plus and minus strands) of polynucleotide sequences from a biological sample of a subject. In some embodiments, the sequence reads can include nucleotide representations of the nucleotides of the polynucleotide sequence. In some embodiments, the nucleotide representation can be in any suitable form including, but not limited to, alphabetic representation, numeric representation, alphanumeric representation, or symbolic representation. In some embodiments, the sequence reads include A (representing adenosine), C (representing cytosine), G (representing guanosine), and T (representing thymidine).
[0073] In some embodiments, the array data includes array reads annotated with additional information (e.g., mapped location, length, sample source, acquisition date, etc.). In some embodiments, the array reads are mapped to a reference array (e.g., a reference genome), thereby providing the mapped array reads. Thus, in some embodiments, the array data includes multiple annotated array reads. The annotated array reads and / or the array data can be provided in a suitable digital format (e.g., a BAM file).
[0074] In some embodiments, the sequencing data can include sequence reads of any suitable polynucleotide of a biological sample. In some embodiments, the sequencing data includes sequence reads of RNA of a biological sample. In some embodiments, the sequencing data includes sequence reads of DNA of a biological sample. In some embodiments, the sequencing data includes sequence reads of tumor DNA of a biological sample. In some embodiments, it includes sequence reads of cell-free DNA (e.g., from healthy cells and cancer cells (e.g., ctDNA)). In some embodiments, the sequencing data includes sequence reads of circulating tumor DNA (ctDNA) of a biological sample. In some embodiments, the sequencing data includes sequence reads of whole exome sequencing of a biological sample. In some embodiments, the sequencing data includes sequence reads of whole genome sequencing of a biological sample. In some embodiments, the sequencing data includes sequence reads covering positions being monitored for mutations (e.g., positions associated with MRD). In some embodiments, the sequencing data includes sequence reads obtained using a targeted gene sequencing panel.
[0075] It should be understood that array reads are data generated by sequencing a polynucleotide (e.g., by using a sequencing device). Thus, array reads do not contain physical molecules but contain data representing physical molecules. Thus, references to nucleotides within an array read are references to information about the nucleotides (e.g., information representing the type of nucleotide, e.g., "A" or "G" or "C" or "T").
[0076] In some embodiments, array reads include a plus-strand primer binding sequence and a minus-strand primer binding sequence. In some embodiments, the plus-strand primer binding sequence and the minus-strand primer binding sequence can be complementary to primers used for amplification of the polynucleotide. In some embodiments, the plus-strand primer binding sequence and the minus-strand primer binding sequence are determined for use in designing plus-strand sequence primers and minus-strand sequence primers for amplifying a particular polynucleotide (e.g., a polynucleotide including positions being monitored for mutations) (e.g., when generating sequencing data). For example, for each polynucleotide including positions being monitored for mutations, a plus-strand primer binding sequence and a minus-strand primer binding sequence can be determined. A plus-strand primer binding sequence and a minus-strand primer binding sequence adjacent to the positions being monitored for mutations (e.g., 3' and 5' to the positions) can be determined.
[0077] In some embodiments, the plus-strand primer binding sequence is within 50 nucleotides, 40 nucleotides, 30 nucleotides, 20 nucleotides, 10 nucleotides, or 5 nucleotides of the 3' end of the sequence read. In some embodiments, the plus-strand primer binding sequence is within 50 nucleotides, 40 nucleotides, 30 nucleotides, 20 nucleotides, 10 nucleotides, or 5 nucleotides of the 5' end of the sequence read. In some embodiments, the minus-strand primer binding sequence is within 50 nucleotides, 40 nucleotides, 30 nucleotides, 20 nucleotides, 10 nucleotides, or 5 nucleotides of the 3' end of the sequence read. In some embodiments, the minus-strand primer binding sequence is within 50 nucleotides, 40 nucleotides, 30 nucleotides, 20 nucleotides, 10 nucleotides, or 5 nucleotides of the 5' end of the sequence read. In some embodiments, the plus-strand primer binding sequence and the minus-strand primer binding sequence are 15 to 30 nucleotides in length.
[0078] In some embodiments, a target gene sequencing panel can be used to specifically sequence only a particular polynucleotide (e.g., a polynucleotide that includes the positions being monitored for mutations) from a biological sample. In some embodiments, the target gene sequencing panel can include a representation of the polynucleotide sequence that includes the positions being monitored for mutations. In some embodiments, the representation of the polynucleotide sequence can be used to determine the polynucleotide to amplify (e.g., using polymerase chain reaction (PCR)). In some embodiments, amplification can be achieved using PCR primers (e.g., a plus-strand sequence primer and a minus-strand sequence primer) that are complementary to the polynucleotide that includes the positions being monitored for mutations. In some embodiments, the plus-strand sequence primer is complementary to a plus-strand primer binding sequence corresponding to the polynucleotide that includes the positions being monitored for mutations. In some embodiments, the minus-strand sequence primer is complementary to a minus-strand primer binding sequence corresponding to the polynucleotide that includes the positions being monitored for mutations. Methods for designing primers for amplifying a particular polynucleotide are well known in the art, such as, for example, as described in Untergasser, Andreas et al., Nucleic Acids Research 40.15 (2012): e115-e115. In some embodiments, the primers are designed using the ArcherDx panel design algorithm. The primers may be designed using other suitable panel design algorithms and can be used without departing from the scope of the invention.
[0079] In some embodiments, the amplification method used in conjunction with the target gene sequencing panel is anchored multiplex PCR (AMP). AMP is a multiplex PCR enrichment chemistry that incorporates strand-specific priming and the incorporation of unique molecular identifiers (UMIs) into sequence reads, and is well known in the art, for example, as described in Zheng Z et al. Nature medicine 20.12 (2014): 1479-1484.
[0080] Positions being monitored for mutations As shown in FIG. 1, the indicators of the positions being monitored for mutation 104 are input into computer 105 and software 108. In some embodiments, the indicators of the positions being monitored for mutations may include the representation of polynucleotides including disease-related mutations (such as cancer-related mutations). In some embodiments, the disease-related mutations may be the positions being monitored for mutations. In embodiments, the positions being monitored for mutations are a predetermined set of positions in a reference genome (such as hg19) or a portion thereof. In some embodiments, the positions being monitored for mutations are positions in a reference sequence, and the reference sequence may be completely arbitrary, may be a reference genome assembly, or may be a custom reference. In some embodiments, the positions being monitored for mutations correspond to positions that mutate in the genome of the subject's tumor. In some embodiments, the positions being monitored for mutations may be positions related to MRD. In some aspects, the positions being monitored for mutations may be positions that overlap with MRD. In some embodiments, the positions being monitored for mutations may be monitored to identify an indicator of MRD.
[0081] In some embodiments, the positions being monitored for mutations can be determined by methods well known in the art. In some embodiments, the positions being monitored for mutations can be determined prior to identifying an MRD indicator. In some embodiments, determining the positions being monitored for mutations includes sequencing the polynucleotides of diseased cells from a subject. In some embodiments, determining the positions being monitored for mutations includes sequencing the polynucleotides of cancer cells from a subject. In some embodiments, determining the positions being monitored for mutations includes sequencing a tumor from a subject. In some embodiments, determining the positions being monitored for mutations includes sequencing a primary tumor from a subject. In some embodiments, determining the positions being monitored for mutations includes sequencing a metastatic tumor from a subject. In some embodiments, determining the positions being monitored for mutations includes sequencing ctDNA from a subject prior to treatment (e.g., pre - surgery). Thus, in some embodiments, the positions being monitored for mutations can be specific to a given subject. In some embodiments, the positions being monitored for mutations can be determined using the ArcherDx panel design algorithm. Other suitable panel design algorithms may be used without departing from the scope of the invention.
[0082] In some embodiments, the number of positions being monitored for mutations can include at least 10 positions, at least 25 positions, at least 50 positions, at least 75 positions, at least 100 positions, at least 125 positions, at least 150 positions, at least 175 positions, at least 200 positions, at least 250 positions, or at least 300 positions. In some embodiments, the number of positions being monitored for mutations can be from 10 to 200 positions. In some embodiments, the number of positions being monitored for mutations can be from 25 to 200 positions. In some embodiments, the number of positions being monitored for mutations can be from 50 to 200 positions. In some embodiments, the number of positions being monitored for mutations can be from 75 to 200 positions. In some embodiments, the number of positions being monitored for mutations can be from 100 to 200 positions. In some embodiments, the number of positions being monitored for mutations can be from 10 to 150. In some embodiments, the number of positions being monitored for mutations can be from 25 to 150 positions. In some embodiments, the number of positions being monitored for mutations can be from 50 to 150 positions. In some embodiments, the number of positions being monitored for mutations can be from 75 to 150 positions. In some embodiments, the number of positions being monitored for mutations can be from 100 to 150 positions.
[0083] As shown in FIG. 1, the computing device 105 and the software 108 obtain sequencing data, obtain indicators of positions being monitored for mutations, and identify indicators of MRD. In some embodiments, the computing device 105 can be one or more computing devices of any suitable type. In some embodiments, the computing device 105 can be, but is not limited to, a laptop, desktop, cloud computer, server, phone, or tablet. In some embodiments, the computing device 105 may be located in a single physical location or may be distributed across multiple physical locations. In some embodiments, the computing device 105 may be located in a facility operated by an entity (e.g., a hospital or research institution). In some embodiments, the software 105 can identify indicators of MRD based on the trinucleotide context (TNC) in the array reads of the sequencing data.
[0084] In some embodiments, the computing device 105 can be operated by a user such as a researcher, patient, physician, clinician, or other individual. For example, the user can provide the sequencing data 104 to the computing device 105 as input (e.g., by uploading a file).
[0085] Trinucleotide context (TNC) In some embodiments, a TNC can be a series of three consecutive nucleotides in an array read. Generally, the nucleotides in an array read can include, but are not limited to, A (representing adenine), C (representing cytosine), G (representing guanine), and T (representing thymine). In some embodiments, the trinucleotide context can be any permutation of any three of A, C, T, G. For example, a TNC can be, but is not limited to, any one of GAT, TTC, TAT, CAT, TTT, and TAA. In some embodiments, an array read includes multiple TNCs. In some embodiments, a given nucleotide representation associated with a given position included in a first TNC may not be included in a second TNC. In some embodiments, a TNC corresponds to an amino acid codon and / or an anticodon. In some embodiments, a TNC is used to determine a first value indicative of the expected number of mutations present in sequencing data due to sequencing errors, and the first value can be used in identifying an indicator of MRD. One advantage of using a TNC over using a single position to calculate an error rate is that, when errors in three nucleotides in a TNC are available compared to more data, such as errors in a single nucleotide, the error rate can be estimated better (Deng, Shibing et al. BMC Bioinformatics 19.1 (2018): 1-7).
[0086] Minimal residual disease (MRD) Minimal residual disease refers to any residual disease (e.g., diseased cells or ctDNA) that may be present in a subject after or following completion of treatment of the disease. For example, minimal residual disease associated with cancer may exist when cancer cells, cancer RNA, and / or circulating tumor DNA (e.g., ctDNA) are present in the subject after treatment. In some embodiments, MRD can be detected based on ctDNA detection using standard surveillance imaging (e.g., computed tomography (CT), magnetic resonance imaging (MRI), or positron emission tomography (PET)) before cancer recurrence is detected. Some cancer types shed DNA, which can ultimately end up in the subject's bloodstream (e.g., ctDNA). Thus, in some embodiments, minimal residual disease can be monitored based on sequencing of ctDNA from a blood-based biological sample (e.g., plasma). In some embodiments, the indicator of minimal residual disease can increase in likelihood or probability over time. For example, cancer cells that have evaded treatment can continue to replicate and / or metastasize, leading to additional ctDNA release.
[0087] Indicator of Minimal Residual Disease (MRD) Any of the techniques described in this specification can be used to identify an MRD indicator. In some embodiments, an MRD indicator can include any information that provides an estimate of the likelihood and / or probability that MRD may be present. For example, an MRD indicator can be an alphabetic, numeric, symbolic, or alphanumeric representation of the likelihood and / or probability that MRD is present. In a further example, an MRD indicator can be based on a scale from 0 to 1, where 0 indicates the lowest likelihood or probability that a subject may have MRD, and 1 indicates the highest likelihood or probability that a subject may have MRD. In some embodiments, an indicator that a subject has MRD can be binary (e.g., "yes" or "no", "true" or "false", etc.). In some embodiments, an MRD indicator can be an instruction to perform additional testing to confirm the MRD indicator. In some embodiments, an MRD indicator can be based on sequencing data from two or more biological samples from the same subject. For example, an MRD indicator can be identified if the analysis of sequencing data from at least two biological samples from a subject both reveals the same result. In another example, an MRD indicator can be identified if the analysis of sequencing data from at least one of two or more biological samples from a subject indicates MRD.
[0088] In some embodiments, an MRD indicator can be an estimate of the likelihood and / or probability that MRD is present in a subject's biological sample, based on a statistical test that compares a sequencing error (e.g., a first value) in the analysis of the biological sample to the actual number of mutations (e.g., a second value) at positions being monitored for mutations in the biological sample. In some embodiments, an indicator of minimal residual disease can be based on the number of positions being monitored for mutations.
[0089] In some embodiments, statistical tests can be used to identify an indicator of MRD. In some embodiments, the statistical test can be selected from the group consisting of a Poisson test, a binomial test, or any other suitable statistical test. In some embodiments, the statistical test may be a nonparametric test (e.g., Wilcoxon rank-sum test). In some embodiments, the statistical test may be a one-sided test. In some embodiments, the statistical test may be a one-sided Poisson test. In some embodiments, the statistical test may compare a null distribution based on a first value and an alternative hypothesis based on a second value. In some embodiments, the indicator of MRD may be based on the p-value of the statistical test being less than a predetermined value alpha. In some embodiments, alpha may be at most 0.2, at most 0.15, at most 0.1, at most 0.05, at most 0.01, at most 0.005, or at most 0.001. In some embodiments, alpha is 0.2, 0.15, 0.1, 0.05, 0.01, 0.005, or 0.001. In some embodiments, alpha may be 0.01. Aspects of identifying an indicator of MRD are described herein, including the description with reference to FIG. 6.
[0090] FIG. 2A shows a flowchart of an exemplary process 200 for identifying minimal residual disease (MRD) based on sequencing data from a biological sample of a subject. Process 200 has the following acts: Act 201: Obtain sequencing data including sequence reads that cover positions being monitored for mutations; Act 202: Determine a first value indicative of the expected number of mutations present in the sequencing data due to background error; Act 203: Identify a second value indicative of the actual number of mutations present at positions being monitored for mutations; and Act 204: Use the first value and the second value to identify an indicator of minimal residual disease. In some embodiments, process 200 is embodied as part of program code (e.g., processor-executable instructions) of software 108 and is executed using computing device 105.
[0091] As shown in FIG. 2A, process 200 begins at act 201, and sequencing data is obtained. In some embodiments, the sequencing data can be generated from a biological sample collected from a subject (e.g., a subject who has received cancer treatment), as described above. In some embodiments, the sequencing data includes sequence reads that cover positions being monitored for mutations. In some embodiments, the sequencing data can be obtained from a data store, cloud storage, compact disk, external storage, sequencing device, public repository, desktop computer, laptop computer, phone, tablet, flash drive, or any other suitable source. In some embodiments, the sequencing data can be obtained by generating the sequencing data using a sequencing system (e.g., sequencing a biological sample of a subject). Aspects of sequencing data are described herein, including the description with reference to FIG. 1.
[0092] In some embodiments, low-quality sequence reads can be identified in the sequencing data before identifying the first and second values. In some embodiments, low-quality sequence reads can be identified and removed from the sequencing data before performing the first and second values. Methods for removing low-quality sequence reads from sequence data are well known in the art, such as, for example, as described in Chen, Shifu et al. BMC Bioinformatics 18.3 (2017): 91-100.
[0093] In some embodiments, consensus array reads can be generated after act 201 and before acts 202 and 203. In some embodiments, consensus array reads can be generated based on unique molecular identifiers (UMIs) associated with the consensus array reads. A UMI can be a plurality of unique nucleotide sequences that can be ligated to a polynucleotide fragment (e.g., a DNA fragment from a biological sample) to detect sample preparation bias. Prior to sequencing, polynucleotides extracted from a biological sample (e.g., ctDNA) are often amplified using polymerase chain reaction to produce sufficient material for sequencing. However, polymerase chain reaction can introduce amplification bias because some polynucleotides can be amplified more efficiently than others. To control for this, UMIs can be ligated to the polynucleotides prior to amplification. During amplification, each polynucleotide UMI molecule can be copied multiple times, and the UMI can indicate which copy is from the same original polynucleotide, regardless of the number of copies made. The polynucleotide UMIs can be sequenced to produce sequence reads. In some embodiments, consensus array reads can be generated by aligning all sequence reads with the same UMI and generating a single consensus array read. Methods for generating consensus array reads are well known in the art, for example, as described in Chen, Shifu et al. BMC Bioinformatics 20.23 (2019): 1-8. In some embodiments, each of the consensus array reads can be generated from at least a threshold number of sequence reads associated with each common UMI. In some embodiments, the threshold number of sequence reads associated with each common UMI can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any other suitable threshold.
[0094] In some embodiments, prior to acts 202 and 203, a subset of the consensus array reads can be selected. In some embodiments, the subset of the consensus array reads can be selected to facilitate the use of high-quality data when calculating the first value and the second value. In some embodiments, selecting a subset of the array reads can be based on a measure of similarity between the plus-strand consensus array reads and the minus-strand consensus array reads. In some embodiments, selecting the subset can be performed based on the relative numbers of the plus-strand consensus array reads and the corresponding minus-strand consensus array reads. For example, if less than 95% of the reads associated with a given array are from plus-strand consensus array reads or minus-strand consensus array reads, a subset of the array reads can be selected. In some embodiments, selecting the subset can be performed based on the relative numbers of the read 1 consensus array reads and the corresponding read 2 consensus array reads. Those skilled in the art will understand that some types of sequencing techniques (e.g., Illumina®) can sequence a given polynucleotide (e.g., plus-strand or minus-strand) from the 3' end and the 5' end of the polynucleotide, and thus produce two array reads, namely read 1 and read 2.
[0095] Next, process 200 proceeds to acts 202 and 203. As shown in FIG. 2A, these acts can be performed in parallel. However, in some embodiments these acts can be performed sequentially (e.g., 202 is performed before 203 or 203 is performed before 202), so this need not be the case and the aspects of the techniques described herein are not limited in this regard. Nevertheless, for clarity, act 202 is described first.
[0096] In operation 202, the computing device executing process 200 determines a first value indicative of the expected number of mutations present in the sequencing data due to a sequencing error. An example of the first value indicative of the expected number of mutations is a first value representing the expected number of mutations. In another example, the first value indicative of the expected number of mutations is a first value corresponding to the expected number of mutations present in the sequencing data. In some embodiments, determining the first value may include, but is not limited to, determining an indicator of the error rate of observing mutations at positions being monitored for mutations. In some embodiments, the first value is determined using a subset of the array reads of the sequencing data.
[0097] The first subset of array reads may include one, some, or all of the array reads. In some embodiments, the first subset of array reads includes array reads that cover positions being monitored for mutations. It should be understood that array reads that cover positions being monitored for mutations may also cover background regions. In some embodiments, the first subset of array reads consists of reads selected based on the data quality filters described herein. In some embodiments, the first subset of array reads includes reads associated with positions that indicate an MRD if mutated. In some embodiments, the first subset of array reads may be the same as the second subset of array reads.
[0098] Including the description with reference to FIG. 2B, the manner of determining the first value will be further described herein. In act 203, a computing device executing process 200 identifies a second value indicative of the actual number of mutations present at the positions being monitored for mutations. In some embodiments, the second value can be identified based on the number of times a mutation is observed at the positions being monitored for mutations. In some embodiments, the second value can be identified based on the number of times a mutation is observed at a plurality of positions being monitored for mutations. For example, the second value can be identified based on the number of times a mutation is observed at at least 10 positions being monitored for mutations. In other examples, the second value can be identified based on the number of times a mutation is observed at at least 16 positions, at least 25 positions, at least 50 positions, at least 75 positions, at least 100 positions, at least 125 positions, at least 150 positions, at least 175 positions, at least 200 positions, at least 250 positions, or at least 300 positions being monitored for mutations. In some embodiments, the second value can be identified based on the number of times a mutation is observed at positions from 10 to 200, 16 to 200, 25 to 200, 50 to 200, 75 to 200, 100 to 200, 125 to 200, 150 to 200, 75 to 200, 25 to 150, 50 to 150, 75 to 150, 100 to 150, 25 to 100, 50 to 100, 75 to 100 being monitored for mutations. In some embodiments, the second value can be identified based on the number of times a mutation is observed at positions from 16 to 100 being monitored for mutations. In some embodiments, the second value can be identified based on the number of times a mutation is observed at positions from 50 to 100 being monitored for mutations. In some embodiments, the second value can be identified based on the number of times a mutation is observed at positions from 50 to 200 being monitored for mutations. In some embodiments, identifying the second value includes counting the number of mutations present at each position being monitored for mutations. In some embodiments, identifying the second value can be based on the number of mutations present at each position being monitored for mutations.In some embodiments, identifying the second value can be based on the frequency of the presence of a mutation at each position being monitored for the mutation. In some embodiments, identifying the second value can be performed using consensus array reads. Aspects of the consensus read array are further described herein.
[0099] In some embodiments, the second value can be identified based on a second subset of the array reads. The second subset of the array reads can include one, some, or all of the array reads. In some embodiments, the second value of the array reads includes only the array reads that cover the positions being monitored for the mutation. In some embodiments, the reads that cover the positions being monitored for the mutation may cover a background region. In some embodiments, the second subset of the array reads consists of reads selected based on the data quality filters described herein. In some embodiments, the second subset of the array reads includes the reads associated with the positions indicating MRD when mutated. In some embodiments, the second subset of the array reads can be the same as the first subset of the array reads. In some embodiments, the second value is based on the number of times the positions being monitored for the mutation mutate. In some embodiments, the second value is the sum of the number of times at least some of the positions being monitored for the mutation mutate. In some embodiments, the second value is the sum of the number of times each position being monitored for the mutation mutates.
[0100] After acts 202 and 203 are completed, process 200 proceeds to act 204, and the information identified during those acts (e.g., the first and second values) can be used to identify an indicator that the subject has MRD. Examples of identifying an indicator that a subject has MRD are provided herein, including the above description and the description with reference to FIG. 6 in the section called "Indicator of Minimal Residual Disease (MRD)".
[0101] FIG. 6 shows a manner of using a first value (calculated in act 202 and indicating the expected number of mutations present due to sequencing errors) and a second value (calculated in act 204 and indicating the actual number of mutations present) to identify the likelihood that MRD may be present using a statistical test (e.g., a one-sided Poisson test). In some embodiments, using a statistical test may include identifying the likelihood of observing a second value (indicating the actual number of mutations present) assuming that the distribution of mutations due to sequencing errors follows the distribution associated with the null hypothesis of the statistical test. The distribution may be parameterized by one or more parameters. For example, the distribution may be parameterized by one or more parameters indicating the expected number of mutations present in the sequencing data due to sequencing errors. For example, in some embodiments, the distribution may be a Poisson distribution, and its mean value (λ) may be set to the first value indicating the expected number of mutations present in the sequencing data.
[0102] For example, as shown in FIG. 6, a one-sided Poisson test is used to identify the likelihood of an observation 602 (e.g., the actual number of observations present) under the null hypothesis that the expected number of mutations present follows a Poisson distribution (601) with a mean (λ) 603 set to the first value (calculated in act 202 of process 200) indicating the expected number of mutations due to sequencing errors. In some embodiments, the identification of the presence of MRD may be performed when the p-value of the statistical test is less than a threshold value (e.g., 0.1, 0.01, 0.001, 0.0001, etc.).
[0103] FIG. 2B illustrates a flowchart of an exemplary process 210 for identifying the expected number of mutations due to sequencing errors. Process 201 includes the following acts: Act 211: identifying a plurality of TNC error rates for each of a plurality of trinucleotide context (TNC) error types; Act 212: grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups; Act 213: using at least some of the plurality of TNC error rates to identify a TNC group error rate for the plurality of TNC error rate groups; and Act 214: determining a first value indicative of the expected number of mutations present in the sequencing data using the TNC group error rate.
[0104] In act 211 of process 210, the computing device identifies a plurality of TNC error rates for each of a plurality of trinucleotide context (TNC) error types. Examples of trinucleotide contexts are provided herein, including the above description in the section entitled "trinucleotide context (TNC)". In some embodiments, the TNC error rate may be based on the number of times the TNC error type is observed in the sequencing data. In some embodiments, the TNC error rate may be based on the frequency at which the TNC error type is observed in the sequencing data. In some embodiments, identifying a plurality of TNC error rates for each of a plurality of TNC error types using a first subset of the sequence reads includes identifying the occurrence frequency of each of the TNC error types in the first subset of the sequence reads.
[0105] In some embodiments, the TNC error type can be a mutated variant of wild-type TNC. In some embodiments, wild-type TNC is a TNC that can occur naturally at a given position in the genome or a portion thereof (e.g., a reference genome such as GRCh38 or hg19). In some embodiments, the mutated variant TNC of wild-type TNC can be a TNC having a mutation from wild-type TNC at at least one position of the TNC. For example, if the reference genome contains the following sequence 367-ATGGTACTGCGTACG-381, the wild-type TNC at positions 367 to 369 is ATG. In this example, the TNC error types of ATG can be, but are not limited to, TTG, AAG, and ATC. In some embodiments, the TNC error type can depend on the wild-type TNC. For example, the TNC error type from ATG to AAG can be distinct from the TNC error type from ACG to AAG, despite the wild-type TNC having the same sequence in each mutated variant. In some embodiments, the TNC error type corresponds to a mutation at any position of a given TNC. In some embodiments, the TNC error type includes a mutation at the central position of the TNC and does not include a mutation at the first position or the third position of the TNC (central position TNC error type). Aspects of the central position TNC error type are described herein, including the description with reference to FIG. 3. In some embodiments, the TNC error rate of each of the plurality of TNC error types can be determined.
[0106] Figure 3 shows an example of identifying the TNC error rate. In Figure 3, a consensus array read 304 (described below) that covers all the same positions being monitored for mutation 313 can be aligned with a reference array 303 (e.g., a known array including the positions being monitored for the mutation). In some embodiments, the reference array is the human reference genome assembly hg38. In some embodiments, the TNCs of the aligned consensus array reads can be identified within predetermined background regions 308 and 309. In some embodiments, background region 308 is specific to the minus-strand (-) consensus array reads, and background region 309 is specific to the plus-strand (+) consensus array reads. Background region 308 is specified by threshold 1 (305) and threshold 3 (307). Background region 309 is specified by threshold 1 (305) and threshold 2 (306).
[0107] In some embodiments, identifying the TNC error rate includes counting the number of occurrences of wild-type TNCs 300 and TNC error types 309 in the consensus array reads. In some embodiments, the counting can be based on the wild-type sequence 301. For each TNC of each consensus array read, the identification can be performed as to whether the TNC is a wild-type TNC or a TNC error type by comparing the TNC of the consensus array read with the TNC at the same position in the wild-type sequence 301. In some embodiments, the count of each occurrence 310 of the wild-type TNC and each occurrence 310 of each TNC error type is identified (314). In some embodiments, the total number 311 of the occurrences of the wild-type TNC and the occurrences of the TNC error type is identified.
[0108] In some embodiments, the TNC error types that are counted are TNC error types 309 that have a mutation at the central position of the TNC (from the wild-type TNC) but do not have a mutation at the first (5’) or last (3’) position of the TNC (e.g., central position TNC error types). For example, central position TNC error types of ATG can include, but are not limited to, AAG, ACG, and AGG. In some embodiments, there are 192 different central position TNC error types. The values of the 192 different central TNC error types can be specified as follows: Since each position of the TNC (first, central, and last) can have any one of A, T, G, or C at that position (e.g., 4 common nucleotide ^3 positions), there are 64 possible TNCs based on the 4 common nucleotides A, T, G, and C. The central position of each TNC can mutate in 3 different ways (e.g., from ATG to AAG, ACG, or AGG). Thus, in some embodiments, there are 64 TNCs × 3 possible central position mutations (e.g., 192 possible different central position TNC error types).
[0109] In some embodiments, the plurality of TNC error types includes any combination of any number of central position TNC error types 309. In some embodiments, the plurality of TNC error types includes 10, 25, 50, 75, 96, 100, 125, 150, 176, 192, or any other suitable number of TNC error types. In some embodiments, the plurality of TNC error types includes at least 10, at least 25 TNC error types, at least 50 TNC error types, at least 75 TNC error types, at least 96 TNC error types, at least 100 TNC error types, at least 125 TNC error types, at least 150 TNC error types, at least 176 TNC error types, or at least 192 TNC error types. In some embodiments, the plurality of TNC error types includes 192 central position TNC error types. In some embodiments, the plurality of TNC error types refers to 96 central position TNC error types. In some embodiments, the plurality of TNC error types refers to TNC error types of 10-25, 10-75, 10-125, 10-176, 10-192, 25-50, 25-75, 25-125, 25-176, 25-192, 50-75, 50-125, 50-176, 50-192. In some embodiments, the plurality of TNC error types may be based on the TNC error types present in the array reads. In some embodiments, the TNC error type is specified for each TNC error type of the plurality of TNC error types.
[0110] The TNC error rate 312 can be a function of the occurrences 310 and the total number 311 in some embodiments. In some embodiments, the TNC error rate 312 for a given TNC error type of 309 can be the number obtained by dividing the occurrence 310 of that TNC error type by the total number 311. In some embodiments, the TNC error rate 312 for a given TNC error type of 309 can be determined based on the occurrence 310 of the TNC error type across all of the same wild-type TNCs and the total number 311 across all of the same wild-type TNCs. For example, the wild-type TNC “ATG” can occur at two different positions of the wild-type sequence 300, namely position 1 and position 2. The occurrence of the wild-type TNC in the consensus sequence read at position 1 can be 15, and the occurrence of the wild-type TNC at position 1 can be 19. Further, the occurrence of the TNC error type “AGG” in the consensus sequence read can be 2 at position 1 and 3 at position 2. In this example, the TNC error rate can be calculated as (2 + 3) / (15 + 19). In some aspects, the TNC error rate 312 for a given TNC error type of 309 can be determined using the following formula:
Number
[0111] In some embodiments, the TNC error rate includes the error rate of each TNC error type. In some embodiments, the TNC error rate includes the error rate of each central position TNC error type. In some embodiments, the TNC error rate includes 10, 25, 50, 75, 100, 125, 150, 176, 192, or any other suitable number of TNC error rates. In some embodiments, the TNC error rate includes 10 - 25, 10 - 75, 10 - 125, 10 - 176, 10 - 192, 25 - 50, 25 - 75, 25 - 125, 25 - 176, 25 - 192, 50 - 75, 50 - 125, 50 - 176, 50 - 192, or any other suitable number of TNC error rates. In some embodiments, the plurality of trinucleotide context (TNC) error rates includes a TNC error rate of 192. In some embodiments, the plurality of trinucleotide context (TNC) error rates includes a TNC error rate of 96. In some embodiments, identifying the plurality of TNC error rates of each of the plurality of TNC error types using a first subset of the sequence reads includes identifying the occurrence frequency of each TNC error type in the background region of the first subset of the sequence reads.
[0112] In some embodiments, the sequence reads described herein include a background region. The background region, in some aspects, can identify a region of the sequence that can be used to determine a first value indicative of the expected number of mutations present in the sequencing data due to background errors. In some embodiments, the background region does not include any position that is being monitored for mutations.
[0113] In some embodiments, the background regions of the plurality of background regions can include at least 25 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 125 nucleotides, at least 150 nucleotides, at least 175 nucleotides, at least 200 nucleotides, at least 250 nucleotides, at least 300 nucleotides, at least 400 nucleotides, at least 500 nucleotides, at least 750 nucleotides, at least 1000 nucleotides, at least 2000 nucleotides, at least 5000 nucleotides, or at least 10000 nucleotides. In some embodiments, the background region includes at least 400 nucleotides. In some embodiments, the background region can include at least 10 trinucleotides, at least 25 trinucleotides, at least 50 trinucleotides, at least 100 trinucleotides, at least 125 trinucleotides, at least 150 trinucleotides, at least 175 trinucleotides, at least 200 trinucleotides, at least 250 trinucleotides, at least 300 trinucleotides, at least 400 trinucleotides, at least 500 trinucleotides, at least 750 trinucleotides, at least 1000 trinucleotides, at least 2000 trinucleotides, at least 3000 trinucleotides, at least 5000 trinucleotides, or at least 10,000 trinucleotides. In some embodiments, the background region includes at least 130 trinucleotides.
[0114] In some embodiments, the background region can be identified based on a first threshold, a second threshold, and / or a third threshold.
[0115] In some embodiments, the background region can be identified based on a first threshold. In some embodiments, the first threshold sets a nucleotide distance between the position being monitored for the mutation and the start of the background region. In some embodiments, the first threshold is 0 nucleotides, indicating that the position being monitored for the mutation is included in the background region. In some embodiments, the first threshold is 1 nucleotide, indicating that the position being monitored for the mutation is not included in the background region. In some embodiments, the first threshold is 2 nucleotides, indicating that the position being monitored for the mutation and the nucleotide on one side of that position are not included in the background region. In some embodiments, the first threshold is 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 50, 100, 150, 200, or more nucleotides. In some embodiments, the first threshold can be 2 to 5, 2 to 10, 2 to 20, 2 to 50, 2 to 100, or 2 to 200 nucleotides. In some embodiments, the first threshold can be 5 to 10, 5 to 20, 5 to 50, 5 to 100, or 5 to 200 nucleotides. In some embodiments, the first threshold can be 20 to 50, 20 to 100, or 20 to 200 nucleotides.
[0116] In some embodiments, the background region can be identified based on a second threshold. In some embodiments, the second threshold can be specific to the plus-strand array read. In some embodiments, the second threshold can be the distance from the beginning of the plus-strand array read (e.g., plus-strand primer binding array). For example, if the second threshold is 100 nucleotides, any nucleotide that falls within the first 100 nucleotides of the plus-strand array read can be included in the background region. In some embodiments, the second threshold can be 50, 100, 150, 200, 250, 300, or 400 nucleotides. In some embodiments, the second threshold can be at least 50 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 250 nucleotides, at least 300 nucleotides, or at least 400 nucleotides. In some embodiments, the second threshold can be 50-100 nucleotides, 100-150 nucleotides, 150-200 nucleotides, 200-250 nucleotides, 250-300 nucleotides, 300-400 nucleotides, or 100-400 nucleotides. In some embodiments, the second threshold can be based on the sequencing quality associated with the nucleotides of the array read. It has been found that sequencing near the end of the array read can produce lower quality data than the sequencing at the beginning of the array read. Thus, one of ordinary skill in the art can set the second threshold based on the data quality of the array read to remove the identification of low-quality nucleotides near the end of the array read.
[0117] In some embodiments, the background region can be identified based on a third threshold. In some embodiments, the third threshold can be specific to the minus strand array read. In some embodiments, the third threshold can be the distance from the start of the minus strand array read (e.g., in the minus strand primer binding array). For example, if the third threshold is 100 nucleotides, any nucleotide that falls within the first 100 nucleotides of the third array read can be included in the background region. In some embodiments, the third threshold can be 50, 100, 150, 200, 250, 300, or 400 nucleotides. In some embodiments, the third threshold can be at least 50 nucleotides, at least 100 nucleotides, at least 150 nucleotides, at least 200 nucleotides, at least 250 nucleotides, at least 300 nucleotides, or at least 400 nucleotides. In some embodiments, the third threshold can be 50-100 nucleotides, 100-150 nucleotides, 150-200 nucleotides, 200-250 nucleotides, 250-300 nucleotides, 300-400 nucleotides, or 100-400 nucleotides. In some embodiments, the third threshold can be based on the sequencing quality associated with the nucleotides of the array read. It has been found that sequencing near the end of the array read can produce lower quality data than the sequencing at the start of the array read. Thus, one of ordinary skill in the art can set a second threshold based on the data quality of the array read to remove the identification of low quality nucleotides near the end of the array read. In some embodiments, the first threshold and the second threshold are the same number of nucleotides from their corresponding primer binding arrays.
[0118] In some embodiments, the grouped TNC error rate can be selected after act 211 and before act 212. In some embodiments, selecting the grouped TNC error rate can be based on an error rate threshold (e.g., 401 in FIG. 4). For example, in some embodiments, if the TNC error rate does not exceed the error rate threshold, the grouped TNC error rate can be selected. In some embodiments, the error rate threshold is from 0.0001% to 1%, from 0.001% to 0.1%, or from 0.005% to 0.05%. In some embodiments, the error rate threshold is 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, or 0.09%, 0.1%. In some embodiments, the error rate threshold is 0.01%. In some embodiments, if the upper limit of the confidence interval of the TNC error rate is less than the error rate threshold, the TNC error rate can be selected.
[0119] In some aspects, FIG. 4 shows selecting the grouped TNC error rates (e.g., 403 and 405) based on the upper limit of each 99% binomial confidence interval (CI) (e.g., 402&404). In some embodiments, the error rate threshold 401 is a static threshold (e.g., 0.001% to 0.10%). In some embodiments, for each of the individual TNC error rates (e.g., 402 and 404), a binomial confidence interval 406 (e.g., 99% confidence interval) is also calculated. In some embodiments, for each individual TNC error rate binomial CI, if the upper limit of the TNC error rate binomial CI (e.g., 404) is less than the threshold 401, the grouped TNC error rate is selected. Thus, by selecting the TNC error rate, error rates that are abnormally high (e.g., exceeding the confidence interval 406) and that may reduce accurate error estimation (e.g., determining a first value) can be removed.
[0120] Next, process 210 proceeds to act 212, "grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups". In some embodiments, the grouping can be performed using any known grouping method. In some embodiments, the grouping can be performed using a clustering algorithm. In some embodiments, the clustering algorithm can be selected from the group consisting of affinity propagation, BIRCH (Balanced Iterative Reducing and Clustering using Hierarchies), DBSCAN (Density-Based Spatial Clustering of Applications with Noise), K-means, mini-batch K-means, mean shift, OPTICS (Ordering Points To Identify the Clustering Structure), spectral clustering, Gaussian mixture, and partitioning around medoids (PAM) clustering. In some embodiments, the grouping is performed by a process that includes nearest neighbor clustering or agglomerative hierarchical clustering.
[0121] In some embodiments, the grouping can be performed using partitioning around medoids (PAM) clustering.
[0122] In some embodiments, the TNC error rates can be grouped into 2, 3, 4, 5, 6, 7, 8, 9, 10, or any other suitable number of TNC error rate groups. In some embodiments, the suitable number of TNC error rate groups can be determined based on obtaining sufficient read depth within each group to increase statistical power. In some embodiments, different numbers of TNC error rates are associated with different TNC error rate groups. For example, 40 TNC error rates can be associated with group 1 and 25 TNC error rates can be associated with group 2.
[0123] Next, process 210 proceeds to act 213, "identifying the TNC group error rate of the plurality of TNC error rate groups using at least some of the plurality of TNC error rates." In some embodiments, the TNC error rate includes the selected TNC error rate, as described above. In some embodiments, the TNC error rate includes the TNC error rate associated with 10, 25, 50, 75, 96, 100, 125, 150, 176, 192, or any other possible number of TNC error types. In some embodiments, the TNC error rate includes the TNC error rate associated with at least 10 TNC error types, at least 25 TNC error types, at least 50 TNC error types, at least 75 TNC error types, at least 96 TNC error types, at least 100 TNC error types, at least 125 TNC error types, at least 150 TNC error types, at least 176 TNC error types, or at least 192 TNC error types. In some embodiments, the TNC error rate includes the TNC error rate associated with 192 TNC error types. In some embodiments, the TNC error rate includes the TNC error rate associated with 96 TNC error types. In some embodiments, the TNC error rate includes the TNC error rate associated with the TNC error types of 10 - 25, 10 - 75, 10 - 125, 10 - 176, 10 - 192, 25 - 50, 25 - 75, 25 - 125, 25 - 176, 25 - 192, 50 - 75, 50 - 125, 50 - 176, 50 - 192. In some embodiments, the TNC error rate includes the TNC error rate based on the TNC error types present in the array reads.
[0124] In some embodiments, identifying the TNC group error rate is based on the TNC error rates associated with the groups. In some embodiments, the TNC group error rate can be identified as the average of the TNC error rates associated with the TNC error rate groups. In some embodiments, the TNC group error rate can be identified using a weighted average of the TNC error rates associated with the TNC error rate groups. In some embodiments, the TNC group error rate can be identified using a population weighted average of the TNC error rates associated with the TNC error rate groups. In some embodiments, the TNC group error rate can be identified using the mean or median of the TNC error rates associated with the TNC error rate groups.
[0125] In some aspects, FIG. 5 shows the TNC group error rates 501 - 504 for each of the plurality of TNC error rate groups 509. FIG. 5 shows TNC error rate confidence intervals, e.g., 505 and 506 (e.g., 99% binomial confidence intervals), grouped into any one of groups 1 - 4. The TNC group error rates 501 - 504 can be identified, in some embodiments, by calculating the population weighted average (501(μ1), 502(μ2), 503(μ3), 504(μ4)) for each group. In some embodiments, the TNC group error rates 501 - 504 can be identified by calculating the median for each group. In some embodiments, the TNC group error rate can be identified by calculating the mean of each of the groups 501 - 504.
[0126] Next, process 210 proceeds to act 214, which "determines a first value indicative of the expected number of mutations present in the sequencing data using the TNC group error rate." In some embodiments, the first value may be based on at least one of the TNC group error rates. In some embodiments, the first value may be based on each of the TNC group error rates. In some embodiments, the first value may be based on the number of times each position being monitored for mutations is covered by sequence reads. In some embodiments, the first value may be based on the number of times each position being monitored for mutations is covered by sequence reads within each TNC error group. In some embodiments, the first value may be based on a function of the TNC group error rate and the number of times the positions being monitored for mutations are covered by sequence reads within each TNC error rate group. In some embodiments, the first value may be based on a linear combination of the TNC group error rate and the number of times each position being monitored for mutations is covered by sequence reads within each TNC error rate group. In some embodiments, the first value may be determined as follows: First value = (μ1 × r1) + (μ2 × r2) ··· + ··· (μn × rn) Where n can be the total number of TNC error rate groups, μ1 through μn can be the TNC group error rates associated with TNC error rate groups 1 through n (e.g., the population weighted average of the TNC group error rates associated with the TNC error rate groups), and r1 through rn can be the sum of the number of times each position being monitored for mutations is covered by sequence reads (both mutated and non-mutated positions) within a given TNC error rate group 1 through n. Or, put another way, r1 through rn can be the sum of the number of sequences (e.g., consensus sequences) that cover the positions being monitored for mutations within a given TNC error rate group 1 through n.
[0127] Computer Implementations An exemplary embodiment of a computer system 1300 that can be used in combination with any of the embodiments of the technology described herein (e.g., the methods of FIGS. 2A and 2B, etc.) is shown in FIG. 12. The computer system 1300 includes one or more processors 1302 and one or more products including non-transitory computer-readable storage media (e.g., memories 1305 and one or more non-volatile storage media 1303). The processor 1302 may control the reading and writing of data to and from the memory 1305 and the non-volatile storage device 1303 in any suitable manner, and the aspects of the technology described herein are not limited to any particular technique for writing or reading data. To perform any of the functions described herein, the processor 1302 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., memory 1305) that can function as a non-transitory computer-readable storage media for storing processor-executable instructions for execution by the processor 1302.
[0128] The computing device 1300 may include a network input / output (I / O) interface 1301 that enables the communication device to communicate with other computing devices (e.g., via a network), and may include one or more user I / O interfaces 1306 that enable the computing device to provide outputs to and receive inputs from a user. The user I / O interface may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or a touch screen), a speaker, a camera, and / or various other types of I / O devices.
[0129] The above-described embodiments can be implemented in any number of ways. For example, an embodiment can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed by any suitable processor (e.g., a microprocessor) or collection of processors, whether provided on a single computing device or distributed among a plurality of computing devices. It should be understood that any component or collection of components that performs the above-described functions can generally be regarded as one or more controllers that control the above functions. The one or more controllers can be implemented in many ways, such as using dedicated hardware or general-purpose hardware (e.g., one or more processors) programmed to execute the above functions using microcode or software.
[0130] In this regard, one embodiment of the embodiments described herein is that when executed by one or more processors, a computer program (i.e., a plurality of executable instructions) that executes the above functions of one or more embodiments is encoded in at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage device, or other tangible non-transitory computer-readable storage medium). The computer-readable medium is transportable to load the stored program onto any computing device to implement aspects of the techniques described herein. Additionally, it should be understood that references to a computer program that, when executed, performs any of the above functions are not limited to an application program running on a host computer. Rather, the terms computer program and software are used herein in a generic sense to refer to any kind of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be employed to program one or more processors to implement aspects of the techniques described herein.
[0131] The foregoing description of the embodiments provides illustration and description, but is not intended to be exhaustive, i.e., not intended to limit the embodiments to the precise forms disclosed. Modifications and variations are possible in light of the above teachings, or may be obtained from practice of the embodiments. In other embodiments, the methods shown in these figures may include fewer operations, different operations, operations in a different order, and / or additional operations. Further, the non-dependent blocks may be executed in parallel.
[0132] As described above, it will be apparent that the exemplary aspects may be implemented in many different forms of software, firmware, and hardware in the illustrated embodiments. Further, certain portions of the embodiments may be implemented as "modules" that perform one or more functions. Such modules may include hardware such as a processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or a combination of hardware and software.
[0133] Some aspects and embodiments of the technology described in this disclosure have thus been described. It is to be understood that various alternatives, modifications, and improvements will be readily apparent to those of ordinary skill in the art. Such alternatives, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. For example, those of ordinary skill in the art will readily devise various other means and / or structures for performing the functions described herein and / or for obtaining one or more of the results and / or advantages described herein, and each such variation and / or modification is to be regarded as being within the scope of the embodiments described herein. It is possible for those of ordinary skill in the art to recognize or ascertain many equivalents to the specific embodiments described herein using only routine experimentation. Accordingly, it is to be understood that the above embodiments are presented by way of example only, and that embodiments of the invention may be practiced otherwise than as particularly described within the scope of the appended claims and their equivalents. In addition, any combination of two or more features, systems, articles, materials, kits, and / or methods described herein is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0134] The above-described embodiments can be implemented in any of many modes. One or more aspects and embodiments of the present disclosure, including the execution of a process or method, may utilize program instructions executable by a device (e.g., a computer, a processor, or other device) to execute or control the execution of the process or method. In this regard, various concepts of the present invention can be implemented as a computer-readable storage medium (or plural computer-readable storage media) (e.g., computer memory, one or more floppy disks, compact disks, optical disks, magnetic tapes, flash memory, circuit configurations within a field programmable gate array or other semiconductor device, or other tangible computer storage media) encoded with one or more programs that, when executed on one or more computers or other processors, execute a method for implementing one or more of the various embodiments described above. One or more computer-readable media can be transportable such that the stored one or more programs can be loaded onto one or more different computers or other processors to implement the various aspects of the above-described aspects. In some embodiments, the computer-readable media can be non-transitory media.
[0135] The terms "computer" or "software" are used herein in a generic sense to refer to any kind of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement the various aspects as described above. Further, according to one aspect, it should be understood that one or more computer programs that, when executed, implement the methods of the present disclosure need not be present on a single computer or processor, but may be distributed in a modular fashion across several different computers or processors to implement the various aspects of the present disclosure.
[0136] Computer-executable instructions can be in many forms such as program modules to be executed by one or more computers or other devices. Generally, a program module includes routines, programs, objects, components, data structures, etc. that execute a particular task or implement a particular abstract data type. Typically, the functions of program modules can be combined or distributed as desired in various embodiments.
[0137] Also, a data structure can be stored in a computer-readable medium in any suitable form. For the sake of brevity, a data structure can be shown as having fields related through locations in the data structure. Such relationships can likewise be achieved by allocating storage for the fields at locations in the computer-readable medium that convey the relationships between the fields. However, any suitable mechanism can be used to establish relationships between the information within the fields of a data structure, including the use of pointers, tags, or other mechanisms to establish relationships between data elements.
[0138] When implemented in software, the software code can be executed by any suitable processor or collection of processors, whether provided on a single computer or distributed among multiple computers. Further, it should be understood that the computer can be implemented in any of several forms, such as, by way of non-limiting example, a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, the computer can be implemented by a device that is generally not considered a computer, including a personal digital assistant (PDA), a smartphone, a tablet, or any other suitable portable or fixed electronic device, provided it has suitable processing capabilities.
[0139] In addition, a computer may have one or more input devices and output devices. These devices can be used, in particular, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display for visual presentation of output and a speaker or other sound generating device for audible presentation of output. Examples of input devices that can be used for a user interface include pointing devices such as a keyboard and a mouse, a touchpad, and a digital tablet. As another example, a computer may receive input information through speech recognition or in other audible formats.
[0140] Such a computer may be interconnected by one or more networks in any suitable form, including a corporate network and a local area network or wide area network such as an intelligent network (IN) or the Internet. Such networks may be based on any suitable technology, may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0141] Also, as described above, some aspects may be implemented as one or more methods. The acts performed as part of a method may be arranged in any suitable manner. Thus, although exemplary embodiments show acts as sequential acts, embodiments may be constructed in which acts are performed in an order different from the illustrated order, including performing some acts simultaneously.
[0142] It is to be understood that all definitions defined and used in this specification take precedence over dictionary definitions, definitions in incorporated references, and / or the ordinary meaning of the defined terms.
[0143] The indefinite articles "a" and "an" are to be understood to mean "at least one" as used in this specification and the claims, unless the contrary is expressly stated.
[0144] When the phrase "and / or" is used in this specification and the claims, it should be understood to mean "either or both" of the elements so conjoined, i.e., elements that in some cases coexist conjunctively and in other cases exist disjunctively. Multiple elements listed by "and / or" should likewise be construed as "one or more" of the elements so conjoined. Elements other than those specifically recited by the "and / or" clause may optionally be present, whether or not they are related to those specifically recited elements. Thus, by way of non-limiting example, reference to "A and / or B", when used in conjunction with open-ended terms such as "comprising", may refer in one embodiment to only A (optionally including elements other than B); in another embodiment to only B (optionally including elements other than A); in yet another embodiment to both A and B (optionally including other elements), and so on.
[0145] As used in this specification and the claims, the phrase "at least one" referring to a list of one or more elements means at least one selected from any one or more of the elements in the list of those elements, but is not necessarily meant to include at least one of each of the elements specifically recited in the list of those elements, and is not meant to exclude any combination of the elements in the list of those elements. It should be understood that this definition also allows for the possibility that elements other than those specifically identified in the list of elements referred to by the phrase "at least one" may optionally be present, whether or not they are related to those specifically identified elements. Thus, by way of non-limiting example, "at least one of A and B" (or equivalently "at least one of A or B" or equivalently "at least one of A and / or B") can, in one embodiment, refer to at least one A, optionally including two or more, where no B is present (and optionally including elements other than B); in another embodiment, refer to at least one B, optionally including two or more, where no A is present (and optionally including elements other than A); in yet another embodiment, refer to at least one A, optionally including two or more, and at least one B, optionally including two or more (and optionally including other elements), and so on.
[0146] In the claims and the above specification, all transitional phrases such as "comprising", "including", "carrying", "having", "containing", "involving", "holding", "composed of", etc. are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" are to be regarded as closed-ended or semi-closed-ended transitional phrases, respectively.
[0147] The exemplary methods disclosed herein may be practiced as appropriate without any element not specifically disclosed herein. Thus, for example, in each instance, the term "comprising" may be replaced in this specification with "consisting essentially of" or "consisting of".
[0148] The terms "about", "substantially", and "approximately" may be used in some embodiments to mean within ±20% of a target value, in some embodiments within ±10% of a target value, in some embodiments within ±5% of a target value, and in some embodiments within ±2% of a target value. The terms "about", "substantially", and "approximately" may include the target value.
[0149] Example Example A1. A method for determining whether sequencing data of a biological sample of a subject provides an indicator that the subject has minimal residual disease, using at least one computer hardware processor, (A) obtaining sequencing data, the sequencing data having been previously generated by sequencing a biological sample of the subject, the sequencing data including sequence reads that cover positions to be monitored for mutations, and (B) using at least a first subset of the sequence reads to determine a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors, the determining comprising using the first subset of the sequence reads to identify a plurality of NC error rates for each of a plurality of nucleotide context (NC) error types, grouping at least some of the plurality of NC error rates into a plurality of NC error rate groups, using at least some of the plurality of NC error rates to identify an NC group error rate for the plurality of NC error rate groups, and Determining a first value indicative of an expected number of mutations present in sequencing data using the NC group error rate including the determining, (C) identifying a second value indicative of an actual number of mutations present at positions being monitored for mutations using at least a second subset of the array reads, (D) using the first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicative of the actual number of mutations present in the sequencing data at positions being monitored for mutations to determine whether the sequencing data provides an indicator that the subject has minimal residual disease, A method comprising performing the above.
[0150] Example A2.NC is the method according to Example A1, selected from a single nucleotide context (SNC), dinucleotide context (DNC), trinucleotide context (TNC), 4-nucleotide context (4NC), 5-nucleotide context (5NC), 6-nucleotide context (6NC), 7-nucleotide context (7NC), and 8-nucleotide context (8NC).
[0151] Example A3. The array reads cover at least 10 positions being monitored for mutations, and the method is as described in Example A1 or A2.
[0152] Example A4. The array reads cover 10 to 200 positions being monitored for mutations, and the method is as described in any one of Examples A1 to A3.
[0153] Example A5. The array reads cover 50 to 200 positions being monitored for mutations, and the method is as described in any one of Examples A1 to A4.
[0154] Example A6. The method according to any one of Examples A1 to A5, further comprising obtaining sequencing data by sequencing a biological sample.
[0155] Example A7. The method according to any one of Examples A1 - A6, wherein the biological sample is a body fluid or a sample obtained from a body fluid.
[0156] Example A8. The method according to any one of Examples A1 - A7, wherein the sequencing data includes sequence reads from circulating tumor DNA (ctDNA).
[0157] Example A9. The method according to any one of Examples A1 - A8, wherein each of the sequence reads covers at least one of the positions being monitored for mutations.
[0158] Example A10. The method according to any one of Examples A1 - A9, wherein the sequence reads are obtained using whole exome sequencing.
[0159] Example A11. The method according to any one of Examples A1 - A11, wherein the sequence reads are obtained using a targeted gene sequencing panel.
[0160] Example A12. The method according to Example A11, wherein the targeted gene sequencing panel targets sequences covering the positions being monitored for mutations.
[0161] Example A13. The method according to Example A11, wherein primers are used to amplify sequences including the positions being monitored for mutations.
[0162] Example A14. The method according to Example A11, wherein the sequences targeted by the targeted gene sequencing panel are determined using sequence data from the primary tumor of the subject.
[0163] Example A15. The method according to any one of Examples A1 - A14, wherein the first subset of sequence reads and the second subset of sequence reads are the same.
[0164] Example 16A.(B) is the method according to any one of Examples A1 - A15, which is executed using at least a first subset of the array reads and one or more array reads in the sequencing data that do not cover the positions being monitored for mutations.
[0165] Executing Example 17A.(B) further includes generating a consensus array read using at least a first subset of the array reads, and each of the consensus array reads is generated from the array reads within at least the first subset of the array reads, each associated with its respective common unique molecular identifier (UMI). Identifying the plurality of NC error rates for each of the plurality of NC error types is the method according to any one of Examples A1 - A16, which is executed using the generated consensus array reads.
[0166] Example A18. The method according to Example A17, wherein each of the consensus array reads is generated from at least a threshold number of array reads associated with their respective common UMI.
[0167] Example A19. The method according to Example A18, wherein the threshold number of array reads is between 2 and 20.
[0168] Example A20. Further includes selecting a subset of the consensus array reads. Identifying the plurality of NC error rates for each of the plurality of NC error types is the method according to any one of Examples A17 - A19, which is executed using only the selected subset of the consensus array reads.
[0169] Example A21. The consensus array reads include a plus - strand consensus array read and a minus - strand consensus array read, and selecting the subset is executed using a criterion for applying a similarity measure between the corresponding plus - strand consensus array read and the minus - strand consensus array read. The method according to Example A20.
[0170] Example A22. The consensus array reads include a plus-strand consensus array read and a minus-strand consensus array read, and selecting a subset of the consensus reads is the method described in Example A20 that uses one or more criteria applied to the plus-strand consensus array read and the minus-strand consensus array read.
[0171] Example A23. The consensus array reads include a plus-strand consensus array read and a minus-strand consensus array read, and selecting a subset is performed using criteria applied to the relative numbers of the plus-strand consensus array read and the corresponding minus-strand consensus array read, the method described in any one of Examples A20 - A22.
[0172] Example A24. Identifying the plurality of NC error rates for each of the plurality of NC error types using consensus array reads includes identifying the plurality of NC error rates using the background region of the consensus array reads, the positions being monitored for mutations include a first position, the consensus array reads include a first consensus array read covering the first position, and the background region includes the first background region of the first consensus array read, the first background region includes nucleotides within the first consensus array read that are at least a first threshold distance from the first position, the method described in any one of Examples A17 - A23.
[0173] Example A25. The background region does not include the positions being monitored for mutations, the method described in Example A24.
[0174] Example A26. The consensus array reads include a first group of plus-strand consensus array reads associated with a plus-strand primer binding sequence at the 3' end of each of the plus-strand consensus array reads for a first position of a position being monitored for mutations, and a second group of minus-strand consensus array reads associated with a minus-strand primer binding sequence at the 3' end of the minus-strand consensus array reads in a second group, Identifying a plurality of NC error rates for each of a plurality of NC error types using consensus array reads is identifying a plurality of NC error rates, nucleotides in any array read within a first group of plus-strand consensus array reads that are located within a second threshold distance of the plus-strand primer binding sequence, and nucleotides in any array read within a second group of minus-strand consensus array reads that are located within a third threshold distance of the minus-strand primer binding sequence The method according to any one of Examples A17 to A23, including identifying a plurality of NC error rates using
[0175] Example A27. Identifying a plurality of NC error rates using consensus array reads is the method according to any one of Examples A17 to A26, including identifying the occurrence frequency of each of the TNC error types within the consensus array reads.
[0176] Example A28. Identifying a plurality of NC error rates for each of a plurality of NC error types using consensus array reads is including identifying a plurality of NC error rates from a background region of the consensus array reads, the consensus array reads include a first consensus array read, and the background region includes a first background region of the first consensus array read, The NC error rate is identified based on the frequency at which each of the NC error types occurs in the first background region of the first consensus array read, according to the method described in Example A21.
[0177] Example A29. The NC error type is a method according to any one of Examples A1 to A28 corresponding to a mutation at any position of a given NC.
[0178] Example A30. Each of the NC error types is a method according to any one of Examples A1 to A29 corresponding to a specific mutation of a central nucleotide within a given NC.
[0179] Example A31. Further comprising identifying a confidence interval of the NC error rate after identifying a plurality of NC error rates and before grouping at least some of the plurality of NC error rates into a plurality of NC error rate groups, and selecting at least some of the plurality of NC error rates for grouping using a criterion applied to the confidence interval of the NC error rate, a method according to any one of Examples A1 to A30.
[0180] Example A32. Grouping at least some of the plurality of NC error rates into a plurality of NC error rate groups includes clustering the plurality of NC error rates, a method according to any one of Examples A1 to A31.
[0181] Example A33. Grouping at least some of the plurality of NC error rates into a plurality of NC error rate groups includes grouping using partitioning around medoids (PAM) clustering, a method according to any one of Examples A1 to A32.
[0182] Example A34. Grouping at least some of the plurality of NC error rates includes grouping into four NC error rate groups, a method according to any one of Examples A1 to A33.
[0183] Example A35. Determining a first value indicating the expected number of mutations present in the sequencing data is performed using at least some of the NC group error rates and the number of times at least some of the positions monitored for mutations are covered by sequence reads within a first subset of sequence reads, a method according to any one of Examples A1 to A34.
[0184] Determining a first value indicative of an expected number of mutations present in sequencing data comprises determining the first value as a weighted linear combination of the NC error rates of each of one specific NC error rate of a first subset of array reads weighted by the number of times the monitored position is covered by array reads corresponding to the NC error type belonging to the NC error group, wherein the first value is determined as a weighted linear combination of the NC error rates of each of one specific NC error rate of a first subset of array reads weighted by the number of times the monitored position is covered by array reads corresponding to the NC error type belonging to the NC error group, the method according to any one of Examples A1 to A35.
[0185] Performing (C) further comprises generating a second consensus array read using at least a second subset of array reads, each of the second consensus array reads being generated from array reads within at least a second subset of array reads associated with respective common unique molecular identifiers (UMIs), identifying a second value indicative of an actual number of mutations present at positions monitored for mutations is performed using the second consensus array read, the method according to any one of Examples A1 to A36.
[0186] Performing (D) is performed using a statistical hypothesis test with a null hypothesis by comparing the second value to a distribution associated with the null hypothesis, the distribution having one or more parameters that depend on the first value, the method according to any one of Examples A1 to A37.
[0187] The distribution is a Poisson distribution having a mean value (λ) set to the first value, the method according to Example A38.
[0188] Using a statistical hypothesis test comprises determining a measure of the likelihood of observing the actual number of mutations indicated by the second value under the null hypothesis, the method according to Example A38 or A39.
[0189] Performing (D) is performed using a one-sided Poisson hypothesis test, the method according to any one of Examples A1 to A40.
[0190] Example A42. Using a one-sided Poisson hypothesis test involves setting the mean value (λ) of the Poisson distribution to a first value and determining a measure of the likelihood of observing the actual number of mutations indicated by a second value under the Poisson distribution, the method described in Example A41.
[0191] Example A43. Using the measure of likelihood to determine whether the sequencing data provides an indication that the subject has minimal residual disease, the method described in Example A42.
[0192] Example A44. If the second value indicates that the null hypothesis can be rejected, the subject is likely to have minimal residual disease, the method described in any one of Examples A39 - A43.
[0193] Example A45. (D) further includes providing an indication that the subject has minimal residual disease, the method described in any one of Examples A1 - A44.
[0194] Example A46. Using at least one computer hardware processor, obtaining one or more of additional sequencing data previously generated by sequencing one or more additional biological samples of a subject, wherein each of the one or more of additional sequencing data includes additional sequence reads that cover positions being monitored for mutations, obtaining; and for each of the additional sequence reads of the one or more of additional sequencing data, using at least a first subset of the additional sequence reads to determine an additional first value indicating the expected number of mutations present in each of the sequencing data due to sequencing errors, and determining includes identifying a plurality of additional NC error rates for each of a plurality of NC error types, grouping at least some of the plurality of additional NC error rates into a plurality of additional NC error rate groups, Identifying a further NC group error rate of a further plurality of NC error rate groups using at least some of the further plurality of NC error rates, and Determining a further first value indicative of the expected number of mutations present in each of the sequencing data using the further NC group error rate including, determining; Identifying a further second value indicative of the actual number of mutations present at a position using at least a second subset of the further sequence reads; and Using the further first value indicative of the expected number of mutations present in each of the sequencing data due to sequencing errors and the further second value indicative of the actual number of mutations present at the positions monitored for mutations in each of the sequencing data, to determine whether each of the sequencing data provides an indicator that the subject has minimal residual disease; and The method according to any one of Examples A1 to A45, further comprising performing.
[0195] Example 1. A method for determining whether sequencing data of a biological sample of a subject provides an indicator that the subject has minimal residual disease, comprising: Using at least one computer hardware processor, (A) Obtaining sequencing data, the sequencing data having been previously generated by sequencing a biological sample of the subject, the sequencing data including sequence reads covering positions monitored for mutations; (B) Determining a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors using at least a first subset of the sequence reads, the determining comprising: Identifying a plurality of TNC error rates of each of a plurality of trinucleotide context (TNC) error types using the first subset of the sequence reads; Grouping at least some of a plurality of TNC error rates into a plurality of TNC error rate groups, using at least some of the plurality of TNC error rates to identify a TNC group error rate for the plurality of TNC error rate groups, and using the TNC group error rate to determine a first value indicative of an expected number of mutations present in the sequencing data including determining, (C) using at least a second subset of the array reads to identify a second value indicative of an actual number of mutations present at positions being monitored for mutations, and (D) using the first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicative of the actual number of mutations present in the sequencing data at positions being monitored for mutations to determine whether the sequencing data provides an indicator that the subject has minimal residual disease A method including performing.
[0196] Example 2. The method of Example 1, wherein the array reads cover at least 10 positions being monitored for mutations.
[0197] Example 3. The method of Example 1 or 2, wherein the array reads cover 10 to 200 positions being monitored for mutations.
[0198] Example 4. The method according to any one of Examples 1 to 3, wherein the array reads cover 50 to 200 positions being monitored for mutations.
[0199] Example 5. The method according to any one of Examples 1 to 4, further comprising obtaining sequencing data by sequencing a biological sample.
[0200] Example 6. The method according to any one of Examples 1 to 5, wherein the biological sample is a body fluid or a sample obtained from a body fluid.
[0201] Example 7. The sequencing data is the method according to any one of Examples 1 to 6, comprising sequence reads from circulating tumor DNA (ctDNA).
[0202] Example 8. The method according to any one of Examples 1 to 7, wherein each of the sequence reads covers at least one of the positions being monitored for mutations.
[0203] Example 9. The method according to any one of Examples 1 to 8, wherein the sequence reads are obtained using whole exome sequencing.
[0204] Example 10. The method according to any one of Examples 1 to 9, wherein the sequence reads are obtained using a targeted gene sequencing panel.
[0205] Example 11. The method according to Example 10, wherein the targeted gene sequencing panel targets sequences covering the positions being monitored for mutations.
[0206] Example 12. The method according to Example 10, wherein primers are used to amplify sequences containing the positions being monitored for mutations.
[0207] Example 13. The method according to any one of Example 10, wherein the sequences targeted by the targeted gene sequencing panel are determined using sequence data from the primary tumor of the subject. Claims 10 to 27 are related to identifying the expected number of mutations due to (B) sequencing errors.
[0208] Example 14. The method according to any one of Examples 1 to 13, wherein the first subset of sequence reads and the second subset of sequence reads are the same.
[0209] Example 15. (B) is performed using at least the first subset of sequence reads and one or more sequence reads in the sequencing data that do not cover the positions being monitored for mutations, according to the method of any one of Examples 1 to 14.
[0210] Performing Example 16.(B) further includes generating consensus array reads using at least a first subset of the array reads, each of the consensus array reads being generated from array reads within at least the first subset of the array reads associated with respective common unique molecular identifiers (UMIs), identifying the respective plurality of TNC error rates for each of the plurality of trinucleotide context (TNC) error types is the method according to any one of Examples 1 to 15, performed using the generated consensus array reads.
[0211] Example 17. The method according to Example 16, wherein each of the consensus array reads is generated from at least a threshold number of array reads associated with respective common UMIs.
[0212] Example 18. The method according to Example 17, wherein the threshold number of array reads is from 2 to 20.
[0213] Example 19. Further including selecting a subset of the consensus array reads, identifying the respective plurality of TNC error rates for each of the plurality of trinucleotide context error types is performed using only the selected subset of the consensus array reads, the method according to Example 16.
[0214] Example 20. The consensus array reads include plus-strand consensus array reads and minus-strand consensus array reads, and selecting the subset is performed using criteria for applying a similarity measure between the corresponding plus-strand consensus array reads and minus-strand consensus array reads, the method according to Example 19.
[0215] Example 21. The consensus array reads include plus-strand consensus array reads and minus-strand consensus array reads, and selecting a subset of the consensus reads is performed using one or more criteria applied to the plus-strand consensus array reads and minus-strand consensus array reads, the method according to Example 19.
[0216] Example 22. A consensus array read includes a plus-strand consensus array read and a minus-strand consensus array read, and selecting a subset is performed using criteria applied to the relative numbers of the plus-strand consensus array read and the corresponding minus-strand consensus array read, the method according to any one of Examples 19 to 21.
[0217] Example 23. Identifying the respective plurality of TNC error rates for a plurality of trinucleotide context (TNC) error types using consensus array reads includes identifying the plurality of TNC error rates using the background region of the consensus array reads, the positions being monitored for mutations include a first position, the consensus array reads include a first consensus array read covering the first position, the background region includes the first background region of the first consensus array read, the first background region includes nucleotides within the first consensus array read that are at least a first threshold distance from the first position, the method according to any one of Examples 16 to 22.
[0218] Example 24. The background region does not include the positions being monitored for mutations, the method according to any one of Examples 16 to 22.
[0219] Example 25. The consensus array reads include a first group of plus-strand consensus array reads associated with plus-strand primer binding sequences at the 3' ends of each of the plus-strand consensus array reads in the first group and a second group of minus-strand consensus array reads associated with minus-strand primer binding sequences at the 3' ends of the minus-strand consensus array reads in the second group for a first position among the positions being monitored for mutations, identifying the respective plurality of TNC error rates for a plurality of trinucleotide context (TNC) error types using consensus array reads identifying a plurality of TNC error rates, nucleotides in any array read within a first group of plus-strand consensus array reads located within a second threshold distance of the plus-strand primer binding array, and nucleotides in any array read within a second group of minus-strand consensus array reads located within a third threshold distance of the minus-strand primer binding array using to identify a plurality of TNC error rates, the method according to any one of Examples 16 to 22.
[0220] Example 26. Identifying a plurality of trinucleotide context (TNC) error rates using consensus array reads includes identifying the frequency of occurrence of each of the TNC error types within the consensus array reads, the method according to any one of Examples 16 to 25.
[0221] Example 27. Identifying the respective plurality of TNC error rates for each of a plurality of trinucleotide context (TNC) error types using consensus array reads includes identifying a plurality of TNC error rates from a background region of the consensus array reads, the consensus array reads include a first consensus array read, the background region includes a first background region of the first consensus array read, the TNC error rates are identified based on the frequency with which each of the TNC error types occurs in the first background region of the first consensus array read, the method according to Example 20.
[0222] Example 28. A TNC error type corresponds to a mutation at any position of a given TNC, the method according to any one of Examples 1 to 27.
[0223] Example 29. Each of the TNC error types corresponds to a specific mutation of the central nucleotide within a given TNC, the method according to any one of Examples 1 to 28.
[0224] Example 30. After identifying a plurality of trinucleotide context (TNC) error rates and before clustering at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, further comprising identifying a confidence interval for the TNC error rate and selecting at least some of the plurality of TNC error rates for clustering using a criterion applied to the confidence interval for the TNC error rate, the method according to any one of Examples 1 to 29.
[0225] Example 31. Clustering at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, the method according to any one of Examples 1 to 30, including clustering the plurality of TNC error rates.
[0226] Example 32. Clustering at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, the method according to any one of Examples 1 to 31, including clustering using partitioning around medoids (PAM) clustering.
[0227] Example 33. Clustering at least some of the plurality of TNC error rates, the method according to any one of Examples 1 to 32, including clustering into four TNC error rate groups.
[0228] Example 34. Determining a first value indicative of the expected number of mutations present in the sequencing data is performed using at least some of the TNC group error rates and the number of times at least some of the positions being monitored for mutations are covered by sequence reads within a first subset of the sequence reads, the method according to any one of Examples 1 to 33.
[0229] Example 35. Determining a first value indicative of the expected number of mutations present in the sequencing data is Determining a first value as a weighted linear combination of the TNC error group rates with each one of a particular one of the number of times covered by array reads corresponding to the TNC error type belonging to the TNC error group, where the monitored position is covered by array reads within a first subset of array reads, as described in any one of Examples 1 to 34.
[0230] Performing Example 36.(C) further includes generating a second consensus array read using at least a second subset of array reads, where each of the second consensus array reads is generated from array reads within at least a second subset of array reads associated with respective common unique molecular identifiers (UMIs), Identifying a second value indicative of the actual number of mutations present at the position being monitored for mutations, as described in any one of Examples 1 to 35, which is performed using the second consensus array read.
[0231] Example 37.(D) is performed using a statistical hypothesis test with a null hypothesis by comparing the second value to a distribution associated with the null hypothesis, where the distribution has one or more parameters that depend on the first value, as described in any one of Examples 1 to 36.
[0232] Example 38. The distribution is a Poisson distribution having a mean value (λ) set to the first value, as described in Example 37.
[0233] Using a statistical hypothesis test, as described in Example 37 or 38, includes determining a measure of the likelihood of observing the actual number of mutations indicated by the second value under the null hypothesis.
[0234] Example 40.(D) is performed using a one-sided Poisson hypothesis test, as described in any one of Examples 1 to 39.
[0235] Example 41. Using a one-sided Poisson hypothesis test, which includes setting the mean value (λ) of the Poisson distribution to a first value and determining a measure of the likelihood of observing the actual number of mutations indicated by a second value under the Poisson distribution, the method described in Example 40.
[0236] Example 42. Using the measure of likelihood to determine whether the sequencing data provides an indication that the subject has minimal residual disease, the method described in Example 41.
[0237] Example 43. If the second value indicates that the null hypothesis can be rejected, the subject is likely to have minimal residual disease, the method described in any one of Examples 38 - 42.
[0238] Example 44. (D) further includes providing an indication that the subject has minimal residual disease, the method described in any one of Examples 1 - 43.
[0239] Example 45. Using at least one computer hardware processor, obtaining one or more of additional sequencing data previously generated by sequencing one or more additional biological samples of a subject, wherein each of the one or more additional sequencing data includes additional sequence reads covering positions being monitored for mutations, and for each of the additional sequence reads of the one or more additional sequencing data, determining a further first value indicating the expected number of mutations present in each of the sequencing data due to sequencing errors using at least a first subset of the additional sequence reads, and the determining includes identifying a further plurality of TNC error rates for each of a plurality of trinucleotide context (TNC) error types, grouping at least some of the further plurality of TNC error rates into further plurality of TNC error rate groups, Using at least some of the additional plurality of TNC error rates to identify an additional TNC group error rate for the additional plurality of TNC error rate groups, and using the additional TNC group error rate to determine an additional first value indicative of the expected number of mutations present in each of the sequencing data including determining, identifying an additional second value indicative of the actual number of mutations present at a position using at least a second subset of the additional array reads, and using the additional first value indicative of the expected number of mutations present in each of the sequencing data due to sequencing errors and the additional second value indicative of the actual number of mutations present at the positions monitored for mutations in each of the sequencing data to determine whether each of the sequencing data provides an indicator that the subject has minimal residual disease, and The method according to any one of Examples 1 to 44, further comprising performing.
[0240] Example 46. A system for determining whether sequencing data of a biological sample of a subject provides an indicator that the subject has minimal residual disease, comprising at least one computer hardware processor, and at least one non-transitory computer-readable storage medium storing processor-executable instructions, and wherein the processor-executable instructions, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to (A) obtaining sequencing data, the sequencing data having been previously generated by sequencing a biological sample of the subject, the sequencing data including array reads covering positions monitored for mutations, obtaining, and (B) Using at least a first subset of the array reads to determine a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors, the determining comprising: Using a first subset of the array reads to identify a plurality of TNC error rates for each of a plurality of trinucleotide context (TNC) error types; Grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups; Using at least some of the plurality of TNC error rates to identify a TNC group error rate for the plurality of TNC error rate groups, and Using the TNC group error rate to determine a first value indicative of the expected number of mutations present in the sequencing data. Including the determining; (C) Using at least a second subset of the array reads to identify a second value indicative of the actual number of mutations present at positions being monitored for mutations; and (D) Using the first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicative of the actual number of mutations present in the sequencing data at positions being monitored for mutations to determine whether the sequencing data provides an indicator that the subject has minimal residual disease. A system for causing the above to be performed.
[0241] Example 47. A system according to Example 46, wherein at least one computer hardware processor stores processor-executable instructions that cause at least one computer hardware processor to perform the method according to any one of Examples 2 - 45.
[0242] Example 48. At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause at least one computer hardware processor to: (A) Obtaining sequencing data, wherein the sequencing data has been previously generated by sequencing a biological sample of a subject, the sequencing data includes array reads covering positions to be monitored for mutations, and obtaining; (B) Using at least a first subset of the array reads to determine a first value indicating an expected number of mutations present in the sequencing data due to sequencing errors, and determining includes: Using the first subset of the array reads to identify a plurality of TNC error rates for each of a plurality of trinucleotide context (TNC) error types; Grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups; Using at least some of the plurality of TNC error rates to identify a TNC group error rate for the plurality of TNC error rate groups, and Using the TNC group error rate to determine a first value indicating the expected number of mutations present in the sequencing data. Determining includes the above; (C) Using at least a second subset of the array reads to identify a second value indicating the actual number of mutations present at positions being monitored for mutations; (D) Using the first value indicating the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicating the actual number of mutations present in the sequencing data at positions being monitored for mutations, to determine whether the sequencing data provides an indicator that the subject has minimal residual disease; At least one non - transient computer - readable storage medium storing processor - executable instructions for causing the above to be executed.
[0243] Example 49. At least one computer hardware processor stores processor-executable instructions that cause the at least one computer hardware processor to execute any of the methods of Examples 2 to 45, the at least one non-transitory computer-readable storage medium storing the processor-executable instructions described in Example 48.
[0244] Example 50. Identify indicators of minimal residual disease (MRD) in non-small cell lung cancer (NSCLC).
[0245] Overview Data are presented from a patient-specific tumor information-based approach of ctDNA detection (e.g., an indicator of MRD) that examines the median of 200 somatic variants per patient from surgically resected tumors. 712 unnatural samples were analyzed to assay performance in parallel with 1093 clinically annotated plasma samples from 197 prospectively recruited patients with early operable NSCLC, of which 76 had lesion recurrence. Through the analysis of 383 extracranial surveillance scans, it was found that MRD status can assist in the interpretation of indeterminate findings on surveillance imaging with the potential to guide early definitive intervention at these sites.
[0246] First Circulating tumor DNA (ctDNA) is a multifaceted biomarker with the potential to accelerate innovation, especially in the early intervention trial space. Postoperative ctDNA detection indicating minimal residual disease (MRD) is a specific indicator of impending NSCLC recurrence. Here, an anchor multiplex PCR (AMP) approach based on personalized tumor information that identifies indicators of MRD in the Evolutionary NSCLC study (NCT01888601) is reported.
[0247] MRD Detection Compared with previous studies, here we adopted a method that increases the number of mutations tracked by a patient-specific enrichment panel (e.g., a PSP or a targeted gene sequencing panel) and enables a higher-resolution characterization of NSCLC recurrence (Figure 7). The PSP targeted the median of 200 mutations (range 72 - 201). The algorithm for selecting variants from the tumor exome data for the PSP design incorporated parameters including prediction of low-background sequencing errors and elevated tumor copy number status to maximize sensitivity (Figures 8 and 9). To minimize sequencing errors, only deep reads were considered for cfDNA analysis (see Methods). Deep reads are reads supported by a minimum number (e.g., at least 5) of sequence reads associated with a single unique identifier [UMI].
[0248] The MRD detection algorithm evaluated background (non-variant) sequencing positions to estimate the in-library trinucleotide context error rate and enabled ctDNA detection at the single-library unit level (see Methods, Figure 7). Details regarding the determination of the MRD call threshold in 10 patient pilot data, the analytical validation of variant DNA detection sensitivity in 712 suspicious samples analyzed using a 50-variant PSP, and the orthogonal validation of NSCLC pre-operative ctDNA positive calls using digital droplet PRC are presented below in the “Analytical validation experiments” including the description referring to Figure 11E).
[0249] Pre-operative ctDNA detection is specific for recurrent NSCLC Pre-operative cfDNA samples from 45 recurrence-free patients were analyzed (Figure 10A). All samples had p-values greater than 0.1 (Figure 11A). Pre-operative cfDNA samples from 20 patients who developed new secondary primary cancers during follow-up were also analyzed (the PSP is specific for the resected primary NSCLC and detection of secondary primary cancers was not expected, Figure 10B). Of the 472 post-operative samples from these patients, 462 (97.9%) were negative for ctDNA detection, and of the 65 patients, 62 (95%) had no detectable post-operative ctDNA.
[0250] Regarding FIGS. 10A to 10C, the circle on the left side on day 0 is the pre-operative time point since the patient's tumor was still in-situ. The circle on the right side on day 0 is after surgical resection of primary NSCLC. When the circle is black, it reflects positive ctDNA detection. The light gray rectangles (rectangles 1 and 3) represent whether the patient received chemotherapy, the dark gray rectangles (rectangles 2 and 4) represent whether the patient received radiotherapy, and the medium gray shaded rectangle (rectangle 5) represents the case when the patient received surgery after recurrence. The triangles represent standard post-treatment CT, PET, or MRI contrast classified as no lesion (medium gray, triangle 8), uncertain image (fairly light gray, triangle 9), or clear contrast evidence of extracranial recurrence (dark gray, triangle 10). The medium gray triangle (triangle 6) represents no evidence of intracranial recurrence, and the very dark gray triangle (triangle 7) indicates intracranial recurrence. The black vertical line represents the patient's event day (in the case of death, secondary primary, or NSCLC recurrence), and in other cases, the vertical line represents the NSCLC study (NCT01888601) follow-up sensor ship day of that patient.
[0251] ctDNA Detection in Recurred NSCLC Patients Four hundred postoperative plasma samples from 76 patients with recurrent NSCLC were analyzed (Figure 10C). Among postoperative inpatients with recurrent NSCLC, one of 13 calls was made with an MRD p-value between 0.1 and 0.01 (Figure 11B). The rest were made with p-values less than 0.01. Among patients with detectable ctDNA preoperatively, postoperative ctDNA was detected in 54 of 61 patients (89%). In contrast, among patients without detectable ctDNA preoperatively, 10 of 14 cases (71%, Figure 10C) showed postoperative ctDNA detection. The lead time (number of days from the first postoperative ctDNA detection to confirmation of recurrence on imaging) was assessed. A lead time of 0 days was assigned to patients in whom postoperative ctDNA was not detected or who had an initial ctDNA detection after clinical recurrence, and 67 of 75 patients were lead time evaluable (patients without preclinical recurrence plasma samples [CRUK0048, 0557, 0516, 0674, 0640] or patients with incomplete resection of lesions on postoperative imaging [CRUK0230, 0234, 0291, and 0387] were excluded). The median lead time across all patients with detectable ctDNA preoperatively was 119 days and the mean lead time was 236 days (range 0 to 1137 days, n = 54, Figure 10C). The median lead time across patients without detectable ctDNA preoperatively was 0 days and the mean lead time was 114 days (range 0 to 589 days, n = 14, Figure 10C). Twenty-six of 52 patients were ctDNA positive at initial postoperative ctDNA sample collection, and 26 patients became ctDNA positive during surveillance after a median of 2 (range 1 to 9) ctDNA-negative time points (Figure 10C).
[0252] For the reference of preoperative ctDNA calls from the pilot cohort, 7 patients had positive ctDNA in plasma preoperatively, and all calls were made with p-values less than 0.01 (Figure 11C). To assess the specificity of the ctDNA caller, an in-silico simulation analysis was performed, generating 3157 false MRD panels within the pilot patient library and assessing the p-values of the ctDNA caller (Figure 11D). At a threshold of less than 0.1, 121 out of 3157 simulated false panels were ctDNA positive (in-silico specificity 96.2%), and at a p-value of less than 0.01, 22 out of 3157 simulated false panels were ctDNA positive (in-silico specificity 99.3%).
[0253] Imaging and MRD Detection Three hundred and eighty-six extracranial surveillance imaging reports from 131 patients who had plasma samples collected postoperatively (covering sites including the neck, chest, abdomen, pelvis, colon, kidney, bladder, and spine, 7 magnetic resonance imaging [MRI] spine, liver, or femur scans, and 343 CT scans including 36 whole-body positron emission tomography scans) were reviewed. MRD detection prior to scans showing no new abnormalities occurred in 15 cases in 11 patients, 9 of whom subsequently relapsed with NSCLC (Figure 10D). ctDNA detection prior to scans showing new suspicious abnormalities was common. These data suggest that MRD status can guide early definitive therapeutic intervention (e.g., surgery, radiation, or ablation) at suspicious anatomical sites.
[0254] Discussion ctDNA detection in the post-curative treatment setting indicated impending disease recurrence. A total of 1,096 pre- and post-operative plasma samples from 197 patients were analyzed. Superior performance of MRD surveillance was observed in patients with pre-operative ctDNA detection, suggesting that MRD surveillance could be prioritized over standard radiation surveillance in these patients (median lead time 119 days vs 0 days in pre-operative ctDNA positive vs pre-operative ctDNA negative patients). Of the 12 patients with MRD detected prior to adjuvant treatment, only 1 did not recur after adjuvant radiotherapy (CRUK0086 received postoperative radiotherapy to the mediastinal bed and died of bowel perforation at 880 days postoperatively without recurrence). The remaining patients all had NSCLC recurrence by 2.8 years after ctDNA detection. This suggests that in NSCLC, MRI detection at the assay limit based on the presented tumor information of detection may fail to identify 8% of patients who receive a 5-year survival benefit from adjuvant chemotherapy; perhaps because the amount of metastatic tumor required to yield a positive pre-adjuvant MRD result exceeds that which is curable with systemic therapy.
[0255] Method Library Preparation Using Anchor Multiplex PCR|Anchor multiplex PCR (AMP) is a nested multiplex PCR enrichment chemistry that incorporates strand-specific priming and the incorporation of unique molecular identifiers (UMIs) into the reads to be sequenced. Cell-free DNA, fragmented peripheral blood mononuclear cell (PBMC) DNA, or fragmented normal tissue DNA was end-repaired, phosphorylated, and A-tailed. An adapter containing a universal priming site, an index for multiplexing, and a UMI was then ligated to the DNA. One round of target-specific PCR was performed using a gene-specific primer 1 (GSP1) that amplifies against the P5 primer in the adapter, and additional rounds of PCR were then performed using a primer incorporating a second nested gene-specific primer (GSP2) and a P7 index. Strand-specific priming was performed in both rounds of amplification to facilitate the identification of positive and negative strand input DNA molecules during information-based analysis.
[0256] The method aimed to sequence each library to 100 million reads. The on-target deduplication ratio, which describes the ratio of raw on-target reads to on-target reads with unique molecular identifier [UMI] support (UMI-supported reads contain five or more raw reads with matching molecular barcodes), was then evaluated. For samples with an on-target deduplication ratio of less than 10:1 at the initial sequencing depth, additional sequencing was performed to maximize the recovery of unique molecular barcode (UMI) families. The PBMC and normal tissue libraries were sequenced on either a NovaSeq® 6000 System (Illumina®) or a NextSeq® System (Illumina®).
[0257] MRD Calling Algorithm| An MRD caller was generated to examine background sequencing noise with in-library bases (Figure 7). The MRD caller utilized the Archer® informatics pipeline to clean the input reads and generate duplicate UIM support reads. The cleaned, duplicated, and error-corrected UMI support reads were aligned to hg19 and used to evaluate alternative observations at predefined positions (positions based on tumor information) where tumor-specific variants were present in the patient's tumor. Only "deep" consensus reads supported by five or more PCR duplicates were used to infer the expected sequencing noise and calculate the signal of the MRD calling algorithm.
[0258] Alternative bases at positions based on tumor information were subject to a stringent set of quality filters consisting of an off-target filter, a read strand bias filter, a sequencing strand bias filter, a background error rate filter, and a variant allele frequency outlier filter to remove artifact signals. The variant allele frequency outlier filter functioned by performing PAM (Partitioning Around Medoids) clustering of the variant allele frequencies at positions based on tumor information that passed the above-described filters. K was set to 2 in the clustering algorithm, thus resulting in a high VAF group and a low VAF group. If one of the two clusters had a much higher VAF and contained two or fewer tumor-specific variants, as indicated by the non-overlapping confidence intervals between the highest VAF of the low VAF cluster and the lowest VAF of the high VAF cluster, those variants were removed from consideration downstream in the algorithm.
[0259] Next, the in-library background error rate (ER) was calculated. Using the ER, the noise level present in each library that must be reliably exceeded to enable MRD calls was established. To calculate the background library ER, the number of UMI-supported alternate observations (DAO (deep alternate observation)) across the region of interest (ROI) of the assay at each possible alternate position was counted based on each trinucleotide context (TNC) and the plus strand of the reference sequence. The ER corresponding to each TNC alternate was calculated as DAO / DDP (DDP (deep UMI-corrected depth across a TNC alternate)). To measure only PCR and sequencing errors, if the VAF at that position of a particular alternate exceeded 1% (as a reference, this could represent a clonal hematopoiesis-related variant or a single nucleotide polymorphism), the position in the ROI was not included in the calculation of the TNC ER.
[0260] A mapping of tumor observed variants and their associated TNC ERs was generated. Any tumor observed variant with a corresponding TNC ER upper confidence interval exceeding 0.01% was filtered from the MRD call algorithm. Using PAM clustering, four 'D groups' of TNC error rates were generated from the eligible TNCs. The population-weighted mean TNC error rate for each of the four D groups was calculated based on the product of the TNC error rate included in each D group cluster and the total DDP of each TNC. The generation of the four D groups ensured that there was sufficient in-library DDP coverage for each D group to make a precise estimate regarding the ER of the variants within each group.
[0261] To determine whether ctDNA is present in the sample, the total observed DAO summed over the tumor-specific positions remaining after filtering was compared to the number of DAO expected due to background ER determined by group D. A one-tailed exact Poisson test was applied where the total remaining observed DAO functioned as the value being tested and the expected number of DAO due to error functioned as the lambda of the Poisson distribution. If the generated p-value of the test was below a pre-specified alpha threshold set at 0.01, the sample was classified as MRD positive. Figure 11 provides data regarding how the pre-specified alpha threshold of 0.01 used in these analyses was generated.
[0262] To examine whether a single mutation targeted by the panel was present, a one-tailed Poisson test was utilized to assess whether the specific nucleotide error rate corresponding to the mutation of interest and the number of DAO over the mutation of interest exceeded the background ER from which the DAO were expected. If the number of DAO was higher than the background error expected using an alpha threshold of 0.01, the variant was considered to be reliably detected.
[0263] Design of the AMP-MRD Enrichment Panel|An individualized AMP-MRD enrichment panel based on tumor information was designed for 197 NSCLC study (NCT01888601) patients. Using the ArcherDx design algorithm, a median of 50 variants per panel (range 0 to 50) was selected, and using variants selected from NSCLC study (NCT01888601) multi-region exome sequencing data, a median of 150 variants (range 34 to 153) was selected. For Archer variant selection, WES sequencing data from the highest purity tumor regions and paired germline DNA were used. The algorithm then identified highly reliable variants that are tumor-specific and not artifacts. The algorithm then determined which variants could be targeted using the ArcherDX AMP panel, and based on these criteria: the quality of the primers targeting the variant (to ensure high sequencing coverage of the target variant), the error rate predicted for the variant in the error correction bin, and the mapping likelihood, the 50 most informative mutations from this set of variants were targeted. The predicted error rate for each variant is based on analysis of the AMP cfDNA library sequenced on the NovaSeq instrument. This error rate analysis was performed by making target variant calls for every possible SNV within the set of Archer LiquidPlex cfDNA libraries. The NSCLC primary tumor WES pipeline 7 was used for ranking to select NSCLC study variants.
[0264] Each personalized enrichment panel also included 90 primers targeting 45 common single nucleotide polymorphisms (SNPs). During analysis, the zygosity of these SNPs in the cfDNA library was compared to their zygosity in the patient's whole exome sequencing data to confirm that no sample swaps occurred. Additionally, the coverage provided by these primers helps to ascertain the background PCR and sequencing error rates of the library. To maximize the utility in detecting sample swaps, these 45 SNPs were selected based on their presence in each Gnomad sub-population at a frequency of 25% - 75%.
[0265] Analytical Validation Experiments | For Experiment LOD1, 634 samples of fragmented DNA with a known SNP profile (GIAB (Genome in a Bottle) DNA, NA24385) were added to the background of four other fragmented GIAB inputs (NA24149, NA24631, NA24694, and NA24695). Six AMP enrichment panels were generated targeting 50 SNPs heterozygous in NA24385 but not in the other four cell lines. To generate unnatural samples, NA24385 DNA was spiked into the background of the other four samples at ratios from 0.006% to 0.2% by mass, targeting variant allele frequencies in the range of 0.003% to 0.1% (since there are 50% heterozygous variants in neat NA24538). As part of the same dilution series, mixes with target allele frequencies of 1%, 5%, and 10% were made. These mixtures were used as inputs for AMP library preparation to confirm that mass-based mixing achieved the desired target allele frequencies. The proportion of spike-in variants in these high AF libraries was measured by adding the number of deep substitute reads across the target SNPs and dividing by the combined coverage of all deep reads across the target SNPs. This analysis confirmed that the spike-in achieved the target AF. Fragmented DNA inputs from 2 ng to 80 ng were used in the experiments to reflect the range of DNA inputs faced in a clinical setting. Overall, 564 out of 634 samples were considered evaluable for LOD1 analysis (62 samples failed because incorrect DNA inputs, determined by off-target reads / primer / ng input less than 30 or greater than 400, were used, and 8 samples failed because the reads were less than 10 million). Clinical samples were used for the validation of AMP-MRD (LOD2), and the clinical samples were prepared using the GIAB mixture method. Whole exome sequencing data from four patients were used to design patient-specific panels using the ArcherDx panel design algorithm containing 50 SNVs.Using the panel, libraries were prepared using cfDNA from each patient, the total number of deep unique reads containing the target tumor-specific variants was added, and the overall tumor variant AF for each sample was calculated by dividing by the sum of the deep unique coverage across all target tumor variants. The cfDNA libraries of all four patients had a total AF greater than 1%. A single mixture was made using cfDNA from healthy donors and used to dilute the patients' cfDNA. These dilutions were performed as a serial dilution method. First, dilutions were made targeting 1% total AF, and libraries were prepared using this mixture. The total AF was measured in this sample, and a dilution correction factor was calculated taking into account the conversion efficiency difference between the background dfDNA. For example, if 1% AF was targeted and 1.3% AF was observed, this indicates that the patient cfDNA was converted to the library more efficiently than the background and more background DNA needed to be used. Next, mixtures were made to achieve AFs of 0.1%, 0.05%, 0.01%, 0.008%, and 0.005%. A total of 100 libraries were prepared with 5 AFs and 3 input masses. Forty-eight blank samples (DNA provided from 22 healthy donors) were analyzed to assay the assay specificity. The panel observed allele frequency was calculated by taking the number of significant deep alternative reads across the AMP panel, removing the estimated background errors, and dividing by the deep depth across the panel. For assay sensitivity in specific spike-in categories, Clopper-Pearson binomial two-sided 95% confidence intervals were calculated using the R package DescTools and the BinomCI function in Supplementary 2e-f.
[0266] Simulation analysis for assessing specificity | The nucleotide context of tumor-specific SNVs within the AMP-MRD pilot cohort panel of each NSCLC study (NCT01888601) was assayed. Based on these data, pseudo-tumor signatures (genomic positions covered by enrichment primers with similar expected error rates at the positions of the target SNVs) were generated. Pseudo-variants were added to the pseudo-signature if the following criteria were met: covered bidirectionally by primers aimed at MRD detection, included the same TNC group error rate as the true MRD variant, substitution was occurring, not a known population SNP variant determined by Ensemble’s Variant Effect Predictor version 94.5, had an error-corrected coverage delta of 2,000 or less compared to the true MRD variant, and not used within any other pseudo-tumor signature including itself. Thus, the generated pseudo-signatures targeted bases that were not mutated in the primary tumor, and any positive MRD calls from these pseudo-signatures were, by default, false positives. For MRD positive calls, 3157 pseudo-signatures across 91 pilot cfDNA libraries were interrogated. The simulated ctDNA fragments for each sample were estimated by taking the number of significant deep substitution reads across the pseudo-signatures, removing the estimated background error, and dividing by the deep depth across the pseudo-signatures.
[0267] Digital droplet PCR orthogonal validation|Preoperative plasma analysis was also performed on 30 preoperative plasma samples and 8 negative controls (preoperative plasma from patients diagnosed with non-malignant diseases after surgery) from NSCLC study (NCT01888601) patients by a method based on personalized tumor information. Digital droplet polymerase chain reaction (ddPCR) orthogonal analysis was performed. NSCLC study (NCT01888601) patients were selected as those having clonal driver mutations that could be targeted by a single ddPCR assay. The ddPCR assay used was the SAGAsafe® assay (SAGA diagnostics), designed and developed on the BioRad QX200 droplet digital PCR system. The ddPCR analysis was performed by SAGA, which received plasma (median 4.8 ml, range 2.5 to 5.2 ml). cfDNA was extracted using the QiaAMP MinElute ccfDNA Midi kit (Qiagen). The cfDNA was diluted with 40 ul of buffer EB. All of the cfDNA material was input per case, and the ddPCR analysis was performed in 4 replicate reaction wells per sample.
Claims
1. A method for determining whether sequencing data of a biological sample of a subject provides an indicator that the subject has minimal residual disease, comprising using at least one computer hardware processor to: (A) obtain the sequencing data, which has been previously generated by sequencing the biological sample of the subject and which comprises sequence reads covering positions monitored for mutations; (B) use at least a first subset of the sequence reads to determine a first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors, said determining comprising: using the first subset of the sequence reads to identify a plurality of trinucleotide context (TNC) error rates for each of a plurality of TNC error types; grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups; using at least some of the plurality of TNC error rates to identify a TNC group error rate for the plurality of TNC error rate groups; and using the TNC group error rate to determine the first value indicative of the expected number of mutations present in the sequencing data; (C) use at least a second subset of the sequence reads to identify a second value indicative of the actual number of mutations present at the positions monitored for mutations; and (D) use the first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicative of the actual number of mutations present in the sequencing data at the positions monitored for mutations to determine whether the sequencing data provides the indicator that the subject has minimal residual disease. A method comprising performing the above.
2. The method according to claim 1, wherein the sequence reads cover at least 10 positions monitored for mutations.
3. The method according to claim 1 or 2, wherein the sequence reads cover 10 to 200 positions monitored for mutations.
4. The method according to claim 1 or 2, wherein the sequence reads cover 10 to 200 positions monitored for mutations.
4. The method according to any one of claims 1 to 3, wherein the array reads cover positions 50 to 200 that are being monitored for mutations.
5. The method according to any one of claims 1 to 4, further comprising obtaining the sequencing data by sequencing the biological sample.
6. The method according to any one of claims 1 to 5, wherein the biological sample is a body fluid or a sample obtained from a body fluid.
7. The method according to any one of claims 1 to 6, wherein the sequencing data includes sequence reads from circulating tumor DNA (ctDNA).
8. The method according to any one of claims 1 to 7, wherein each of the sequence reads covers at least one of the positions being monitored for mutations.
9. The method according to any one of claims 1 to 8, wherein the sequence reads are obtained using whole exome sequencing.
10. The method according to any one of claims 1 to 9, wherein the sequence reads are obtained using a targeted gene sequencing panel.
11. The method according to claim 10, wherein the targeted gene sequencing panel targets sequences that cover the positions being monitored for mutations.
12. The method according to claim 10, wherein primers are used to amplify the sequences including the positions being monitored for mutations.
13. The method according to claim 10, wherein the sequences targeted by the targeted gene sequencing panel are determined using sequence data from the primary tumor of the subject.
14. The method according to any one of claims 1 to 13, wherein the first subset of the sequence reads and the second subset of the sequence reads are the same.
15. (B) is performed using at least the first subset of the sequence reads and one or more sequence reads in the sequencing data that do not cover the positions being monitored for mutations, according to any one of claims 1 to 14.
16. (B) is performed Further comprising generating a consensus array read using at least the first subset of the array reads, each of the consensus array reads being generated from array reads within at least the first subset of the array reads associated with respective common unique molecular identifiers (UMIs). The method according to any one of claims 1 to 15, wherein identifying each of the plurality of TNC error rates of the plurality of trinucleotide context (TNC) error types is performed using the generated consensus array reads.
17. The method according to claim 16, wherein each of the consensus array reads is generated from at least a threshold number of array reads associated with respective common UMIs.
18. The method according to claim 17, wherein the threshold number of the array reads is 2 to 20.
19. Further comprising selecting a subset of the consensus array reads, The method according to claim 16, wherein identifying each of the plurality of TNC error rates of the plurality of trinucleotide context (TNC) error types is performed using only the selected subset of the consensus array reads.
20. The method according to claim 19, wherein the consensus array reads include a plus-strand consensus array read and a minus-strand consensus array read, and selecting the subset is performed using a criterion for applying a similarity measure between the corresponding plus-strand consensus array read and the minus-strand consensus array read.
21. The method according to claim 19, wherein the consensus array reads include a plus-strand consensus array read and a minus-strand consensus array read, and selecting the subset of the consensus reads uses one or more criteria applied to the plus-strand consensus array read and the minus-strand consensus array read.
22. The method according to any one of claims 19 to 21, wherein the consensus array reads include a plus-strand consensus array read and a minus-strand consensus array read, and selecting the subset is performed using a criterion applied to the relative numbers of the plus-strand consensus array read and the corresponding minus-strand consensus array read.
23. Using the consensus array reads to identify each of the plurality of trinucleotide context (TNC) error rates of the plurality of TNC error types, includes identifying the plurality of TNC error rates using a background region of the consensus array reads, where the positions being monitored for the mutation include a first position, the consensus array reads include a first consensus array read covering the first position, and the background region includes a first background region of the first consensus array read, The method according to any one of claims 16 to 22, wherein the first background region includes nucleotides within the first consensus array read that are at least a first threshold distance from the first position. **Claim 24** The method according to claim 23, wherein the background region does not include the position being monitored for the mutation. **Claim 25** The consensus array reads include a first group of plus-strand consensus array reads associated with a plus-strand primer binding sequence at the 3' end of each of the plus-strand consensus array reads in the first group for a first position of the positions being monitored for the mutation, and a second group of minus-strand consensus array reads associated with a minus-strand primer binding sequence at the 3' end of the minus-strand consensus array reads in the second group, Using the consensus array reads to identify each of the plurality of trinucleotide context (TNC) error rates of the plurality of TNC error types, which is to identify the plurality of TNC error rates, nucleotides in any array read within the first group of plus-strand consensus array reads that are located within a second threshold distance of the plus-strand primer binding sequence, and nucleotides in any array read within the second group of minus-strand consensus array reads that are located within a third threshold distance of the minus-strand primer binding sequence The method according to any one of claims 16 to 22, including using to identify the plurality of TNC error rates. **Claim 26** Identifying the plurality of trinucleotide context (TNC) error rates using the consensus array reads includes identifying the occurrence frequency of each of the TNC error types within the consensus array reads, the method according to any one of claims 16 to 25. **Claim 27** Identifying the plurality of TNC error rates for each of the plurality of trinucleotide context (TNC) error types using the consensus array reads includes identifying the plurality of TNC error rates from a background region of the consensus array reads, the consensus array reads include a first consensus array read, and the background region includes a first background region of the first consensus array read, the TNC error rate is identified based on the frequency with which each of the TNC error types occurs in the first background region of the first consensus array read, the method according to claim 20. **Claim 28** A TNC error type corresponds to a mutation at any position of a given TNC, the method according to any one of claims 1 to 27. **Claim 29** Each of the TNC error types corresponds to a specific mutation of a central nucleotide within a given TNC, the method according to any one of claims 1 to 28. **Claim 30** After identifying the plurality of trinucleotide context (TNC) error rates and before clustering at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, identifying a confidence interval for the TNC error rates and using a criterion applied to the confidence interval of the TNC error rates to select the at least some of the plurality of TNC error rates for clustering, the method according to any one of claims 1 to 29. **Claim 31** Clustering the at least some of the plurality of TNC error rates into a plurality of TNC error rate groups includes clustering the plurality of TNC error rates, the method according to any one of claims 1 to 30. **Claim 32** Clustering the at least some of the plurality of TNC error rates into a plurality of TNC error rate groups includes clustering using partitioning around medoids (PAM) clustering, the method according to any one of claims 1 to 31. **Claim 33** Grouping at least some of the plurality of TNC error rates includes grouping into four TNC error rate groups, the method according to any one of claims 1 to 32.
34. Determining the first value indicative of the expected number of mutations present in the sequencing data is performed using at least some of the TNC group error rates and the number of times at least some of the positions being monitored for the mutations are covered by sequence reads within a first subset of the sequence reads, the method according to any one of claims 1 to 33.
35. Determining the first value indicative of the expected number of mutations present in the sequencing data is determining the first value as a weighted linear combination of the TNC error group rates with each one of specific ones of the TNC error group rates weighted by the number of times the positions being monitored are covered by sequence reads within the first subset of the sequence reads corresponding to the TNC error type belonging to the TNC error group, the method according to any one of claims 1 to 34.
36. Performing (C) further includes generating a second consensus sequence read using at least the second subset of the sequence reads, each of the second consensus sequence reads being generated from sequence reads within at least the second subset of the sequence reads associated with respective common unique molecular identifiers (UMIs), identifying the second value indicative of the actual number of mutations present at the positions being monitored for mutations is performed using the second consensus sequence read, the method according to any one of claims 1 to 35.
37. (D) is performed using a statistical hypothesis test having the null hypothesis by comparing the second value with a distribution associated with the null hypothesis, the distribution having one or more parameters that depend on the first value, the method according to any one of claims 1 to 36.
38. The distribution is a Poisson distribution having a mean value (λ) set to the first value, the method according to claim 37.
39. Using the statistical hypothesis test includes determining a measure of the likelihood of observing the actual number of mutations indicated by the second value under the null hypothesis, the method according to claim 37 or 38.
40. The method according to any one of claims 1 to 39, wherein (D) is performed using a one-sided Poisson hypothesis test. **Claim 41** Using the one-sided Poisson hypothesis test includes setting the mean value (λ) of the Poisson distribution to the first value and determining a measure of the likelihood of observing the actual number of mutations indicated by the second value under the Poisson distribution. The method according to claim 40. **Claim 42** Using the measure of likelihood to determine whether the sequencing data provides the indicator that the subject has minimal residual disease. The method according to claim 41. **Claim 43** If the second value indicates that the null hypothesis can be rejected, the subject is likely to have minimal residual disease. The method according to any one of claims 38 to 42. **Claim 44** The method according to any one of claims 1 to 43, wherein (D) further includes providing the indicator that the subject has minimal residual disease. **Claim 45** Using the at least one computer hardware processor, obtaining one or more of additional sequencing data previously generated by sequencing one or more additional biological samples of the subject, wherein each of the one or more additional sequencing data includes additional sequence reads covering the positions being monitored for mutations; and for each of the one or more additional sequence reads of the additional sequencing data, using at least a first subset of the additional sequence reads to determine an additional first value indicating the expected number of mutations present in each sequencing data due to sequencing errors, and the determining includes identifying a plurality of additional trinucleotide context (TNC) error rates for each of a plurality of TNC error types, grouping at least some of the plurality of additional TNC error rates into a plurality of additional TNC error rate groups, using the at least some of the plurality of additional TNC error rates to identify an additional TNC group error rate for the plurality of additional TNC error rate groups, and Determining the further first value indicating the expected number of mutations present in each of the sequencing data, using the further TNC group error rate including, determining; identifying a further second value indicating the actual number of mutations present at the position, using at least a second subset of the further array reads; using the further first value indicating the expected number of mutations present in each of the sequencing data due to sequencing errors and the further second value indicating the actual number of mutations present at the positions being monitored for mutations in each of the sequencing data, to determine whether each of the sequencing data provides an indication that the subject has minimal residual disease; The method according to any one of claims 1 to 44, further comprising performing.
46. A system for determining whether sequencing data of a biological sample of a subject provides an indication that the subject has minimal residual disease, comprising: at least one computer hardware processor; at least one non-transitory computer-readable storage medium storing processor-executable instructions; wherein the processor-executable instructions, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to: (A) obtain the sequencing data, the sequencing data having been previously generated by sequencing the biological sample of the subject, the sequencing data including array reads covering positions monitored for mutations; (B) determine a first value indicating the expected number of mutations present in the sequencing data due to sequencing errors, using at least a first subset of the array reads, the determining comprising: identifying a plurality of TNC error rates for each of a plurality of trinucleotide context (TNC) error types, using the first subset of the array reads; grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups; Identifying a TNC group error rate of the plurality of TNC error rate groups using at least some of the plurality of TNC error rates, and Determining a first value indicating an expected number of mutations present in the sequencing data using the TNC group error rate Including, determining; (C) Identifying a second value indicating an actual number of mutations present at positions being monitored for mutations using at least a second subset of the sequence reads; and (D) Using the first value indicating the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicating the actual number of mutations present in the sequencing data at positions being monitored for mutations, determining whether the sequencing data provides an indication that the subject has minimal residual disease; A system for causing execution.
47. The at least one computer hardware processor stores processor-executable instructions that cause the at least one computer hardware processor to execute the method according to any one of claims 2 to 45, the system according to claim 46.
48. At least one non-transitory computer-readable storage medium storing processor-executable instructions, the processor-executable instructions, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to, (A) Obtaining the sequencing data, the sequencing data having been previously generated by sequencing a biological sample of the subject, the sequencing data including sequence reads that cover positions being monitored for mutations; obtaining; (B) Determining a first value indicating an expected number of mutations present in the sequencing data due to sequencing errors using at least a first subset of the sequence reads, the determining comprising: Identifying a plurality of TNC error rates of respective plurality of trinucleotide context (TNC) error types using a first subset of the sequence reads; Grouping at least some of the plurality of TNC error rates into a plurality of TNC error rate groups, identifying a TNC group error rate for the plurality of TNC error rate groups using the at least some of the TNC error rates of the plurality of TNC error rates, and determining the first value indicative of the expected number of mutations present in the sequencing data using the TNC group error rate including determining; (C) identifying a second value indicative of the actual number of mutations present at the positions being monitored for mutations using at least a second subset of the sequence reads; and (D) using the first value indicative of the expected number of mutations present in the sequencing data due to sequencing errors and the second value indicative of the actual number of mutations present in the sequencing data at the positions being monitored for mutations, determining whether the sequencing data provides an indication that the subject has minimal residual disease; at least one non-transitory computer-readable storage medium storing processor-executable instructions for causing a processor to perform the above. **Claim 49** The at least one computer hardware processor stores processor-executable instructions for causing the at least one computer hardware processor to perform the method according to any one of claims 2 to 45, the at least one non-transitory computer-readable storage medium storing the processor-executable instructions according to claim 48.