A highly sensitive method for detecting cancer DNA in a sample.
By employing sample sequencing methods and statistical models, the problems of insufficient sensitivity and high false positive rate in detecting trace amounts of cancer DNA in existing technologies have been solved, achieving high-sensitivity, low-false-positive cancer DNA detection and improving the early warning capability for cancer recurrence.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INIVATA LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-10
AI Technical Summary
Existing detection methods cannot detect trace amounts of residual cancer cells with high sensitivity, resulting in insufficient early warning capabilities for cancer recurrence. Furthermore, conventional sequencing methods are susceptible to interference from PCR errors and sequence variations, leading to false positive results.
The aliquot-based sequencing method is used to identify and exclude statistically impossible variants by repeatedly sequencing the same target region in multiple samples. Combined with statistical models and signal-to-noise ratio enhancement techniques, the accuracy of the test results is ensured.
It improves the detection sensitivity of trace amounts of cancer DNA, reduces the false positive rate, and can reliably detect cancer DNA at extremely low concentrations, reducing misjudgments and improving the early warning capability for cancer recurrence.
Smart Images

Figure 2026063177000001 
Figure 2026063177000002 
Figure 2026063177000003
Abstract
Description
[Technical Field]
[0001] cross reference This application is U.S. Provisional Patent Application No. 63 / 061,568, filed on August 5, 2020. This application claims the interest of the patent, and that application is incorporated herein by reference. [Background technology]
[0002] In most cases, cancer treatment involves at least two steps: removing tumor cells. If the intended first treatment, and then the initial treatment, are not completely successful, the remaining substances in the patient's body A second treatment may be needed, aimed at eradicating all remaining cancer cells. The treatments used to eradicate cancer cells are often different from the first treatment.
[0003] The small number of cancer cells remaining in a patient after initial treatment, when the patient may be clearly in remission, are often In this case, it is called "minimal residual disease" (MRD) or residual disease. These residual cells are the most Ultimately, it will likely cause recurrence in many cancers. Additional treatment may be necessary. The patients with the highest risk can receive additional treatment, while those who do not require additional treatment can avoid it. This reduces harm to patients and lowers treatment costs, and after initial treatment, The likelihood of patients experiencing disease recurrence and relapsing. It is important to determine this. Therefore, an effective method for detecting minimal residual disease is needed. This is highly desirable. Also, current (for example, usually performed by imaging or clinical analysis) It is important to have a highly sensitive method that can detect the risk of cancer recurrence earlier than any other method. .
[0004] MRD can analyze relatively large amounts of DNA and the frequency of common tumor-specific fusions. Because it can be measured in a simple way, it has been successfully used in several hematological malignancies. It has been detected. Currently, regarding circulating tumor DNA (ctDNA), cell-free DNA (cfD) By evaluating NA, MRD can be detected in many solid tumors. There is strong evidence to that effect. However, when detecting minimal residual disease in cfDNA The problem is that many of the tests used to detect sequence variations in a sample are... The problem is that it is not sufficiently sensitive. Many of today's molecular tests are based on known genes. This is done by sequencing the cfDNA of the panel. The problem with detecting minimal residual disease by doing so is that tumor DNA is present in cell-free DNA. The amount is often far below the detection limit of such methods. Specifically, Individual tumor sequence variations expected to occur in the cfDNA of patients with minimal residual disease The frequency of these artifacts is typically due to sequencing artifacts such as PCR errors and base mismatches. This is considerably less frequent than when generated by DNA damage. In some cases, the level of mutant DNA may be very low, and on average, it is analyzed. The fact that there is less than one copy of each mutation being evaluated in the cfDNA sample This worsens the condition. In addition, relatively small amounts of mutant DNA derived from white blood cells dissolved in the bloodstream This can lead to incorrect results. Therefore, sequencing-based approaches to minimal residual disease Detecting anomalies remains difficult.
[0005] This disclosure provides a highly sensitive method for detecting tumor DNA. The method includes, among other things, It can be used to diagnose minimal residual disease.
Summary of the Invention
[0006] A method for detecting cancer DNA in a test sample of DNA from a patient is described below. . In some embodiments, the method comprises: (a) sequencing a plurality of aliquots of the test sample to generate sequence reads corresponding to two or more target regions each having a sequence variation present in the patient's cancer; and (b) for each aliquot and for each target region: i. determining the number of sequence reads having the sequence variation, ii . determining the total number of sequence reads, and iii. comparing i. and ii. to one or more error probability distribution models of the sequence variation, wherein one or more models are obtained from DNA not containing the sequence variation for comparison; and (c) integrating the collective results of step (b) to determine whether cancer DNA is present in the test sample. The method may include: (c) integrating the collective results of step (b) to determine whether cancer DNA is present in the test sample. The method may include: (c) integrating the collective results of step (b) to determine whether cancer DNA is present in the test sample. The method may include: (c) integrating the collective results of step (b) to determine whether cancer DNA is present in the test sample. In any embodiment, step (b) may include iv. excluding variants that exceed a threshold in a statistically unlikely number of aliquots. These variants (i.e., variants in a statistically unlikely number of aliquots) can be identified by measuring the amount of test sample DNA added to each aliquot, calculating the fraction of cancer DNA in the test sample, and estimating the probability of observing the number of aliquots containing variants that exceed the threshold based on i and ii. That is, the variants (i.e., the variants in a statistically unlikely number of aliquots) can be identified by measuring the amount of test sample DNA added to each aliquot, calculating the fraction of cancer DNA in the test sample, and estimating the probability of observing the number of aliquots containing variants that exceed the threshold based on i and ii. ii. The method may include: (c) integrating the collective results of step (b) to determine whether cancer DNA is present in the test sample.
[0007] The method has two features: (i) aliquot-based sequencing (i.e., sequencing the same target regions in a plurality of aliquots of the same sample, i.e., split or distributed samples) i.e., sequencing the same target regions in a plurality of aliquots of the same sample, i.e., split or distributed samples) and (ii) evaluating signals in any of the aliquots (identifying variant DNA in one aliquot, and then determining that the same variant may be found in another aliquot, in contrast to determining that the sample definitely contains cancer DNA), depends on the analysis of multiple variants of all of the data after statistically unlikely data points have been removed. One problem solved by this method is that for some samples (i.e., samples containing a small fraction of cancer DNA, e.g., samples containing less than 0.01% tDNA), the number of sequence reads containing a particular sequence variation is substantially indistinguishable from variations caused by noise (i.e., combinations such as base miscalls, PCR errors, damaged DNA, etc.). Thus, in many cases, it is simply impossible to reliably determine that a sample contains cancer DNA by conventional sequencing approaches. As described above, the present invention is aliquot-based. For example, in some embodiments, the method may include sequencing at least 10 target regions in at least 3 aliquots of a test sample, and in practice, the method may include sequencing at least 24 target regions in at least 4 aliquots of a test sample. Although the aliquot-based sequencing may initially seem like a waste of effort because the same number of wild-type and variant molecules are still being sequenced (although divided over multiple aliquots), the signal-to-noise ratio is actually increased in the aliquot-based method. 、統計的にありそうもないデータ点が除去された後、データのうちの全てを分析する複数 のバリアントの分析に依存する。
[0008] この方法によって解決される1つの問題は、いくつかの試料(すなわち、がんDNAの 小画分、例えば、0.01%未満のtDNAを含む試料)について、特定の配列バリエー ションを含む配列リードの数が、ノイズ(すなわち、塩基ミスコール、PCRエラー、損 傷したDNAなどの組み合わせ)によって引き起こされるバリエーションと実質的には区 別できないことである。そのため、多くの場合では、試料ががんDNAを含むことを従来 の配列決定アプローチによって確実に決定することは単純に不可能である。
[0009] 上記のように、本発明は、アリコートベースである。例えば、いくつかの実施形態では 、方法は、試験試料の少なくとも3つのアリコートにおいて少なくとも10個の標的領域 を配列決定することを含み得、実際には、方法は、試験試料の少なくとも4つのアリコー トにおいて少なくとも24個の標的領域を配列決定することを含み得る。アリコートベー スの配列は、同じ数の野生型分子及びバリアント分子が依然として配列決定されている( ただし、複数のアリコートにわたって分割されている)ため、当初は労力の無駄のように 見える場合があるが、シグナル対ノイズ比は、アリコートベースの方法において実際に増 Specifically, a very small number of variant molecules (e.g., one or two) in the sample are added. In a situation where a variant molecule is present, the ratio of variant molecule to wild-type molecule is such that the variant molecule It would be much higher in aliquots. This, in turn, excludes miscalls. This makes the data more reliable. In addition to increasing the signal-to-noise ratio, the method is conventional This approach generates more data than previous approaches, which in turn leads to more refined data. This allows for analysis by statistical and / or threshold-based methods. For example: (i ) So-called "noisy bases" (i.e., high intrinsic backgrounds that are frequently miscalled) (Positions with a sound) where the signal is present in almost all or all aliquots (background (ii) It is consistently high (relatively) for the round, so it can be identified and excluded. Variants associated with an improbably high signal (for example, one aliquot) Three times the expected number of sequence reads for a single variant molecule, and other Variants with a background number of sequence reads in recoating, or other variants When the ant is present in only 1 or 0 of the aliquots, 3 of the 4 aliquots Variants that appear to be in two can be identified and excluded. Various other advantages The following is a description of the details.
[0010] Depending on how the method is implemented, the method may have certain advantages over conventional methods. For example, the method can be used even when the cancer DNA fraction in the sample is less than 0.01%. This can be used to consistently and reliably determine whether sample A contains cancer DNA. The sensitivity level is far below that of conventional methods, and errors in sequencing artifacts The frequency at which cts can be generated is far below the expected rate. Evaluate several sequence variations. By doing so, the method also copies less than one of each individual sequence variation on average. Cancer DNA can be detected in DNA samples that contain the presence of [unclear].
[0011] The method does not sacrifice specificity (i.e., does not produce many false positive results), It can be carried out to reach a certain level. The presence of ctDNA is used in DNA sequencing. Not at the level of the variant lead after determination, but at the level of the variant molecule attached to each aliquot. This can be estimated. This is true in several situations (for example, DNA at high sequencing depth). With a low initial molecular input, false positives can be reduced, and the overall fraction of cancer DNA can be reduced. It provides a more accurate estimate.
[0012] In addition, in some embodiments, this method optionally determines whether the sample contains cancer DNA. Rather than calculating the number of positives (the number of aliquots with clear evidence of ctDNA), In a probabilistic continuum (i.e., a probability distribution over the number of observed molecules), all allicor All variations in the test are scored, and positive or negative results are determined by applying a simple rule. This is determined by determining a negative result. This is not important when obtained individually. However, it can be combined with strong evidence of ctDNA across multiple variants, It increases the degree of accuracy and enables investigation of boundary signals. It also offers flexible solutions based on confidence levels. Reports and other data, such as prior probabilities of disease recurrence based on cancer type or stage. It enables the potential for combinations.
[0013] In addition, rare errors such as DNA damage before amplification or initial cycle PCR errors can occur. This can be directly modeled by the approach described in the previous paragraph. Based on the estimation process, these would appear to be actual signals. These effects are, Most models of DNA sequencing errors fail to capture and therefore fail to explain them. If present, this can lead to false positives. Alternatively, these can be detected in aliquots. This can be addressed by requiring a signal (in a single sample, 2 This reduces sensitivity because such an event is highly unlikely. The method is to estimate By considering factors such as the type of cancer DNA fraction or DNA base alteration, each ant The molecules detected in the coat are likely to originate from ctDNA, or are rare errors. This effect can be modeled by considering whether it is likely to originate from [a specific source]. ru.
[0014] The method identifies abnormally high levels in multiple aliquots based on the estimated cancer DNA fraction. By excluding variants that show the signal, further error reduction strategies are employed. It is possible to detect only a small number of variant molecules in the sample as a whole. If they are released, all of these (unless amplified or copy number changed) exist in a single location. This is unlikely. This is a clonal hematopoietic (CHIP) mutation with undetermined potential. This can be caused by contamination or similar errors. It can also be explained in background models. It is due to a single DNA base that generates far more sequencing errors than is revealed. This method allows for the first sequencing of a panel of normal samples without the need for sequencing. To make it suitable for "single-shot" use.
[0015] These and other advantages may become apparent in light of the following discussion.
[0016] Those skilled in the art will understand that the drawings shown below are for illustrative purposes only. This instruction is not intended to limit the scope in any way. [Brief explanation of the drawing]
[0017] [Figure 1] This flowchart shows a method for performing aliquot-based sequencing. As will be apparent, different aliquots of a test sample can be barcoded with different aliquot identifier sequences and then combined before sequencing. [Figure 2] This flowchart follows from the flowchart in Figure 1. Figure 2 shows a method for processing sequence reads to determine the number of sequence reads with sequence variations and the total number of sequence reads for each aliquot and each target region. [Figure 3] This flowchart shows an example of how the workflow shown in Figure 2 can be implemented. The steps shown in Figure 3 can be performed in any convenient order. [Figure 4] This flowchart follows the one in Figure 2. Figure 4 shows a method for determining whether cancer DNA is present in a sample by analyzing the variants and total read counts for each sequence variation and aliquot, along with the probability distribution for each sequence variation, and then integrating them. [Figure 5] This flowchart shows a method for generating probability distribution models for each sequence variation. The probability distributions include binomial, overdispersed binomial, beta, normal, exponential, or gamma probability distribution models. Such models may not be necessary in embodiments using molecular indices. [Figure 6] This flowchart shows a threshold-based approach for analyzing data on each sequence variation in each aliquot. [Figure 7] Figure 6 is a flowchart showing a method for integrating the results of the threshold-based method. [Figure 8] This flowchart shows a statistical approach for analyzing data for each sequence variation in each aliquot. [Figure 9] Figure 8 is a flowchart showing a method for integrating the statistical results shown. [Figure 10] Figure 1 is a flowchart illustrating the final step, showing two approaches that allow the results of one test sample to be compared with one or more additional samples. [Figure 11] Some of the principles of the embodiments of this method are outlined below. [Figure 12] This section demonstrates the principle of probability distributions for estimating the number of variant molecules. [Figure 13A-B] Examples of error probability distributions are shown. In the model shown in Figure 13A, data corresponding to low-frequency, high-signal events are hatched. The model shown in Figure 13B is a mixed model. "VAF" refers to variant allele frequencies. Such models are obtained from DNA that does not contain sequence variations, and they represent the probability (or number of variant reads relative to total wt reads) of different variant allele fractions in this normal DNA. Such distributions may differ between variant classes and between sequence depths. In some cases, two or more distributions are needed to explain different types of errors. In some cases, a threshold can be established at which one can reasonably be confident that a sequence variation identified in a sequence read is not an error. [Figure 14] This paper demonstrates a method for identifying and excluding data from "noisy" bases using an aliquot approach. [Figure 15]This illustrates some of the difficulties in detecting cancer DNA by using a method that scores whether individual aliquots contain a particular variant. [Figure 16] This document describes a method for calculating the fraction of cancer DNA. [Figure 17] The results of an experiment in which more than 40 sequence variations were evaluated in four aliquots of each of three different samples containing fluctuating levels of circulating tumor (ctDNA) are shown. [Figure 17-1] Continuing from Figure 17.
[0018] definition Unless otherwise defined, all technical and scientific terms used herein are in accordance with the present invention. It has the same meaning as is generally understood by those skilled in the art to which it belongs. To make references clear and easy, specific elements are defined.
[0019] The terms and symbols used herein in nucleic acid chemistry, biochemistry, genetics, and molecular biology are defined as follows: Standard papers and documents in the relevant technical field, e.g., Kornberg and Ba ker,DNA Replication,Second Edition(WHF reeman, New York, 1992), Lehninger, Biochemi stry,Second Edition(Worth Publishers,New York, 1975), Strachan and Read, Human Mole cular Genetics,Second Edition(Wiley-Liss , New York, 1999), Eckstein, editor, Oligonuc. leotides and analogs: A Practical Approac. h(Oxford University Press, New York, 1991) , Gait, editor, Oligonucleotide Synthesis:A Practical Approach(IRL Press,Oxford,198 Follow 4) etc.
[0020] The term "nucleotide" refers not only to known purine and pyrimidine bases, but also to modified bases. It is intended to include a moiety that also contains other heterocyclic bases. Such modifications are Methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated ribose or It includes other heterocycles. In addition, the term "nucleotide" can refer to hapten or fluorescently labeled nucleotides. It contains the components and includes not only conventional ribose and deoxyribose sugars, but also other sugars. Modified nucleosides or nucleotides can also have, for example, a hydroxyl group. One or more of the atoms are replaced by halogen atoms or aliphatic groups, or ethers, ammonium compounds. This includes modifications to the sugar moiety, such as functionalization as 'n'.
[0021] The terms "nucleic acid" and "polynucleotide" are used interchangeably herein. A creotide, for example, any consisting of a deoxyribonucleotide or ribonucleotide. Length, for example, more than approximately 2 bases, more than approximately 10 bases, more than approximately 100 bases, more than approximately 50 Over 0 bases, over 1,000 bases, over 10,000 bases, 100,000 salts Over 1,000,000, with a maximum of approximately 10 10 or polymers of a higher base are described. This can be produced by enzymes or by synthesis (for example, in the United States). PNA, No. 5,948,902 and the references cited therein Hybridization of naturally occurring nucleic acids in a sequence-specific manner similar to that of naturally occurring nucleic acids. It can be soybeans, for example, involved in Watson-Crick base pairing interactions. This can be achieved. Naturally occurring nucleotides include guanine, cytosine, adenine, thymine, It contains uracil (G, C, A, T, and U, respectively). DNA and RNA are also included. Each has a deoxyribose and ribose sugar backbone, and the backbone of PNA is It consists of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. In PNA, various purine and pyrimidine bases are linked by methylene carbonyl bonds. It is bound to the backbone. Often referred to as inaccessible RNA, or locked nucleus. Acids (LNA) are modified RNA nucleotides. Ribose of LNA nucleotides The portion is modified with additional crosslinks linking the 2' oxygen and 4' carbon atoms. The crosslinks are often In the 3'-end (North) stereochemistry observed in type A double helix, ribose is "ro "Click". LNA nucleotides can be added to oligonucleotides whenever desired. It can be mixed with NA or RNA residues. The term "unstructured nucleic acid" or "UNA" is used. The term refers to nucleic acids containing non-natural nucleotides that bind to each other with reduced stability. In other words, unstructured nucleic acids may contain G' and C' residues, and these residues have reduced stability. They form base pairs with each other, but each forms base pairs with naturally occurring C and G residues. These correspond to non-natural forms, i.e., analogues, of G and C that retain their capabilities. Structured nucleic acids are described in US2005 / 0233340, which is a UNA disclosure. It is incorporated herein by reference.
[0022] As used herein, the term "nucleic acid sample" refers to a sample containing nucleic acids. The nucleic acid samples used in the book contain multiple different molecules that contain sequences. It can be a complex at a point. Genomic DNA from mammals (e.g., mice or humans). The sample is a composite sample type. A composite sample is approximately 10 4 , 10 5 , 10 6 , or 10 7 , 10 8 , 10 9 , or 10 10 It may have more than one different nucleic acid molecule. Nucleic acids, for example, Any sample containing genomic DNA from tissue cultured cells or a tissue sample is collected as described herein. It can be used.
[0023] As used herein, the term "oligonucleotide" refers to a group of approximately 2 to 200 nucleotides. This shows a single-stranded multimer of nucleotides with a maximum length of 500 nucleotides. Otide may be synthesized or enzymatically produced, and in some embodiments, 30 They are approximately 150 nucleotides long. Oligonucleotides are ribonucleotide monomers ( That is, it may be an oligoribonucleotide, or a deoxyribonucleotide monomer. or containing both ribonucleotide monomers and deoxyribonucleotide monomers Obtain. Oligonucleotides are, for example, 10-20, 21-30, 31-40, 41-5 0, 51-60, 61-70, 71-80, 80-100, 100-150, or 150 It can be approximately 200 nucleotides in length.
[0024] A "primer" is the starting point for nucleic acid synthesis when it forms a double helix with a polynucleotide template. It can act as a template along its 3' end so that an extended double helix is formed. This refers to natural or synthetic oligonucleotides that can be extended from their ends. The sequence of nucleotides added during the process is determined by the sequence of the template polynucleotide. The primers are extended by DNA polymerase. The length of the primer extension product is suitable for use in the synthesis of primer extension products, and is usually 8-2. 00 nucleotide lengths, for example, 10 to 100 nucleotide lengths or 15 to 80 nucleotides. It is within the length range. The primer may include a 5' tail portion that does not hybridize to the mold.
[0025] Primers are usually single-stranded for maximum amplification efficiency, but alternatively, two-stranded primers are used. It may be a chain or partially double-stranded. Incorporated herein by reference, Zhang As described in et al (Nature Chemistry 2012 4:208-214). As shown, toehold replacement primers are also included in this definition.
[0026] Therefore, the "primer" is complementary to the template and forms hydrogen bonds or hybridization with the template. The primer forms a complex through synthesis, initiating synthesis by polymerase. / This results in a template complex, which in the process of DNA synthesis has a 3' end that is complementary to the template. It is extended by the addition of covalently bonded bases.
[0027] The term "hybridization" or "to hybridize" refers to the relationship between nucleic acid chains. The region anneals to the second complementary nucleic acid chain under normal hybridization conditions, and this It forms a stable double helix, either homo-double helix or hetero-double helix, and the same normal hive This refers to the process under redilation conditions where unrelated nucleic acid molecules do not form stable double strands. The formation of a double helix occurs when two complementary nucleic acid chain regions are combined in a hybridization reaction. This is achieved by Neil. Hybridization is a reaction in which two nucleic acid strands are It does not contain a specific number of nucleotides in a particular sequence that is substantially or completely complementary. As long as two nucleic acid strands are stable double-stranded, for example, under normal stringency conditions, Hybridization reactions occur in a way that prevents the formation of a double chain that retains the region in the main chain state. This can be made highly specific by adjusting the hybridization conditions. "Normal hybridization or normal stringency conditions" are any given high Bridization reactions are easily determined. For example, Ausubel et al. l.,Current Protocols in Molecular Biolog y, John Wiley & Sons, Inc., New York, or Sambro ok et al.,Molecular Cloning: A Laboratory Manual,Cold Spring Harbor Laboratory Pr See ess. Where used herein, “to hybridize” or “ The term "hybridization" refers to the process by which a nucleic acid chain combines with a complementary strand through base pairing. It refers to any process.
[0028] Nucleic acids are hybridized under moderate to high stringency conditions. When they hybridize specifically with each other, the reference nucleic acid sequence must be "selectively hybridizable". It is considered to be "possible". Moderate and high stringency hybridization conditions The matter is known (for example, Ausubel, et al., Short Protocol ls in Molecular Biology,3rd ed.,Wiley&So ns 1995 and Sambrook et al., Molecular Cloni ng:A Laboratory Manual,Third Edition,200 (See 1 Cold Spring Harbor, NY).
[0029] As used herein, the terms “double-stranded” or “double-stranded” refer to base pairing. In other words, it represents two complementary polynucleotide regions that hybridize together.
[0030] "Genetic locus," "gene locus," "object" related to the genome or target polynucleotide A target locus, region, or segment is a genome or target polynucleotide. It means a continuous sub-region or segment. As used herein, it refers to genetic inheritance. A locus, gene locus, or target gene locus is a nucleotide, gene, or This can refer to the location of a part of a gene, or it may be within a gene, for example, a coding sequence. Whether or not, any continuous part of the genome sequence It can refer to a single nucleoty. A genetic locus, gene locus, or locus of interest is a single nucleoty. The segments can range from a few hundred or thousands of nucleotides in length. Generally, the purpose is The gene locus will have an associated reference sequence (see the "reference sequence" below). (See reference). The terms "plurality," "population," and "collection" are used interchangeably to refer to something that contains at least two members. In certain cases, a plurality, population, or collection can have at least 5, at least 10, at least 100, at least 1,000, at least 10,000, at least 100,000, at least 10 6, at least 10 7, at least 10 8, or at least 10 6 9 or more members. 7 8 9
[0031] The terms "sample identifier array," "sample index," "multiplex identifier," or "MID" are sequences of nucleotides that are added to a target polynucleotide, and the sequences identify the source of the target polynucleotide (i.e., the sample from which the target polynucleotide is derived). In use, each sample is tagged with a different sample identifier array (e.g., one array is added to each sample, and different samples are added to different arrays), and the tagged samples are pooled. After sequencing the pooled samples, the sample identifier array can be used to identify the source of the sequences. The sample identifier array can be added to the 5' end or the 3' end of the polynucleotide. In certain cases, some of the sample identifier arrays can be at the 5' end of the polynucleotide, and the remainder of the sample identifier arrays can be at the 3' end of the polynucleotide. When the elements of the sample identifier have sequences at each end, together, the 3' and 5' sample identifier arrays identify the sample. In many instances, the sample identifier array is only a subset of the bases added to the target oligonucleotide. The identifier array is by ligation Alternatively, it can be added to polynucleotides by primer extension. Morphologically, the identifier sequence may be in the 5' tail or in the primer used for primer extension. In such embodiments, the target polynucleotide is a copy of the original target polynucleotide. It's a beep.
[0032] The term "aliquot identifier sequence" refers to sequence reads from different aliquots interacting with each other. It refers to an added sequence that makes it possible to distinguish between different elements. Aliquot identifier sequences are used to distinguish between different elements. Except for the fact that they are used on aliquots of the sample, rather than as ingredients, the above-mentioned sample identifiers It functions similarly to a sequence. A single sequence functions as both a sample identifier and an aliquot identifier. obtain.
[0033] The term "variable" is used in the context of two or more nucleic acid sequences that are variable relative to one another. This refers to two or more nucleic acids that have different nucleotide sequences. In other words, a polynucleus of a group. If rheotides have variable sequences, the nucleotide sequences of the polynucleotide molecules in the population are: It can vary from molecule to molecule. The term "variable" means that all molecules in a population are different from those in the population. It should not be interpreted as requiring that it have a different sequence from other molecules.
[0034] The term "effectively" is used in relation to Hamming distance, Levenstein distance, and Jackard distance. This includes, but is not limited to, distance, cosine distance, etc., and is measured by similarity functions. This refers to sequences that are duplicates (generally, Kemena et al., Bioinformat (See ics 2009 25:2455-65). The exact threshold is determined by performing the analysis. The error rate of the sample preparation and sequencing used depends on the error rate of the sample preparation and sequencing used, and a higher error rate is A lower threshold for similarity is required. In certain cases, substantially identical sequences are less than It also has 98% or at least 99% sequence identity.
[0035] As used herein, the term “arrangement variation” refers to somatic burrs such as cheek swabs. Reference sequences, such as reference genomes or sequences from patient samples that are not expected to contain ants. It is a different variant. In many cases, "sequence variation" is a variant found in other samples. It is a variant that exists at a frequency of less than 50% compared to the molecule. Many sequence variations Sequences, such as indel and nucleotide substitutions, do not involve sequence variations in molecules. It is virtually identical. In some cases, a particular sequence variation is less than 20%. , less than 10%, less than 5%, less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, and It may be present in the sample at a frequency of less than 0.01%.
[0036] The term "nucleic acid template" is intended to refer to the initial nucleic acid molecule that is copied during amplification. Yes. In this context, copying involves the formation of complements for a particular single-stranded nucleic acid. This can be done. The "initial" nucleic acid is already processed nucleic acid, for example, amplified and extended with an adapter. This may include labeled nucleic acids, etc.
[0037] The term "tail" is used in the context of tail primers or primers having a 5' tail. Therefore, it does not hybridize to the same target as the 3' end of the primer, or it partially hybridizes. Soybeans, the region at their 5' end (for example, at least 12-50 nucleotides) This refers to a primer that has a region.
[0038] The term "initial template" refers to a sample containing the target sequence to be amplified. The term "amplification" as used refers to using a target nucleic acid as a template to amplify one aspect of the target nucleic acid. This refers to generating more than one copy.
[0039] As used herein, the term "amplicon" refers to a specific amplicon in a PCR reaction. This refers to the product (or "band") amplified by the imer pair.
[0040] As used herein, "replicated amplicon" refers to a method using different parts or aliquots of a sample. It refers to the same amplicon amplified using a mold. A duplicate amplicon is typically made using a mold. Sequence variations, PCR errors, and the selection of primers used for each aliquot. Differences in the sequence (for example, differences at the 5' end of primers such as aliquot identifier sequences) They have almost identical sequences, except for (etc.).
[0041] Polymerase chain reaction, or PCR, is a process in which one or more sequence-specific template DNAs are used. This is an enzymatic reaction that is amplified using a target primer pair.
[0042] "PCR conditions" are the conditions under which PCR is performed, as is well known in the field of technology. , reagents (e.g., nucleotides, buffers, polymerases, etc.), and temperature cycles (e.g.) This includes the presence of a temperature cycle suitable for denaturation, resaturation, and extension.
[0043] Multiple polymerase chain reaction (PCR) or multiple PCR is a method that uses two or more different target templates. This is an enzymatic reaction that employs the primer pair shown above. If the target template is present during the reaction, multiple plastination occurs. Rimerase chain reactions co-amplify in a single reaction using a corresponding number of sequence-specific primer pairs. This results in two or more amplified DNA products that are widened.
[0044] The term "next-generation sequencing" refers to the process of performing nucleic acid sequencing using a highly parallelized method. This refers to a method developed by Illumina, Life Technologies, and Pacif. Currently adopted by ic Biosciences and Roche, etc. This includes sequencing by sequencing or sequencing by a ligation platform. The column determination method is a nanopore arrangement such as that provided by Oxford Nanopore. Column determination method, or Ion T commercialized by Life Technologies This may include, but is not limited to, electronic detection-based methods such as orrent technology.
[0045] The term "sequence read" refers to the output of a sequencer. Sequence reads are typically , containing a sequence of G, A, T, and C with a length of 50 to 1000 bases or more, in most cases Each base in a sequence read may be associated with a score indicating the quality of the base call.
[0046] "Assessing the existence of ~" and "Assessing the existence of ~" The term "evaluate whether an element exists" means determining whether an element exists and the element Includes any form of measurement, including estimating a quantity. "Determine," "Measure," "Judge." "To conclude," "to evaluate," and "to analyze" are used interchangeably, and both quantitative and qualitative decisions are made. Includes. The evaluation may be relative or absolute. "To evaluate the existence of ~" is an evaluation of existence. To determine the quantity of something, and / or whether it exists or not. This includes doing so.
[0047] If two nucleic acids are "complementary," they will be compatible with each other under high stringency conditions. To hybridize. The term "completely complementary" means that each of the nucleic acid bases is It is used to describe double strands that pair with complementary nucleotides in other nucleic acids. In some cases, the two complementary sequences are at least 10, for example, at least 12 or 1 It has 5 nucleotide complements.
[0048] The "oligonucleotide binding site" is where the oligonucleotide is located within the target polynucleotide. This refers to the site of hybridization. The oligonucleotide "provides" the primer's binding site. If this occurs, the primer may hybridize to its oligonucleotide or its complement. ru.
[0049] As used herein, the term "chain" refers to a covalent bond, such as a phosphodiester bond. It refers to nucleic acids, which are composed of nucleotides that are covalently bonded together by a molecule. In cells, DNA is They usually exist in a double-stranded form and are therefore referred to as the "upper" and "lower" strands in this specification. It has two complementary strands of nucleic acid. In certain cases, the complementary strands of a chromosomal region are "plus" and the "minus" chain, the "first" and "second" chains, and the "encoded" and "non-encoded" chains. It may be referred to as the "Watson" and "Click" chain, or the "Sense" and "AntiSense" chain. The assignment of the chains as upper or lower is arbitrary and depends on the specific orientation. Does not show function or structure. Several exemplary mammalian chromosome regions (e.g., BAC, A The nucleotide sequences of the first strand (e.g., Swertia japonica, chromosomes) are known, for example, NCB. It can be found in the Genbank database.
[0050] As used herein, the term “extension” refers to the process of nucleation using polymerase. This refers to the extension of the primer by the addition of ostides. The primer that anneals to the nucleic acid extends. When extended, nucleic acids act as templates for the extension reaction.
[0051] As used herein, the term “sequencing” refers to a small number of polynucleotides. Identity of 10 consecutive nucleotides (for example, at least 20, at least 5 0, at least 100, or at least 200 consecutive identical nucleotides This refers to a method of obtaining sex.
[0052] As used herein, the term “pooling” refers to those samples or aliquots. Combine two or more samples or aliquots of samples so that the molecules within are dispersed among each other in the solution. It refers to combining or mixing things together.
[0053] As used herein, the term “pooled sample” refers to the product of pooling. vinegar.
[0054] As used herein, the term “part” refers to different parts of the same sample. This refers to an aliquot or portion of a sample. For example, 1 microliter of a 100 µl sample is If added to each of 10 different PCR reactions, each of those reactions will be the same sample. It contains different parts.
[0055] As used herein, the term "cell-free DNA" ("cfDNA") refers to a cell. It refers to DNA that is free in bodily fluids, not DNA that is already present. cfDNA can be found, for example, in plasma, serum, and brain. It can be isolated from cerebrospinal fluid, urine, saliva, or feces. "Cell-free DNA from bloodstream" and "Circulating cell-free DNA" refers to DNA circulating in the peripheral blood of a patient. DNA molecules have a median size of less than 1kb (e.g., 50bp-500bp, 80bp-). It may have 400 bp, or within the range of 100 to 1,000 bp, but the central signal outside this range may be Fragments containing is may exist. Cell-free DNA is tumor DNA (tDNA), for example, It may contain tumor DNA that circulates freely in the patient's blood. cfDNA is obtained by centrifuging the sample. Then, all cells are removed, and DNA is isolated from the remaining liquid (e.g., plasma or serum). This can be obtained by doing so. Such methods are well known (for example, Lo e See t al, Am J Hum Genet 1998;62:768-75. i) Circulating cell-free DNA can be double-stranded or single-stranded. This term refers to DNA circulating in the bloodstream. Free DNA molecules, and extracellular vesicles (such as exosomes) circulating in the bloodstream It is intended to include existing DNA molecules.
[0056] As used herein, the term "tumor DNA" (or "tDNA") refers to tumor DNA. It is the DNA of origin. tDNA can be identified because it contains mutations. tDNA is From tissue biopsy, from circulating tumor cells (CTCs), from urine or stool samples, etc. Although they are no longer part of the tissue, they can be directly isolated from other non-circulating cells. It may be a part (or "fraction") of the patient's cfDNA. tDNA is clonal. This includes both heterochromia and subclonal mutations. Tumor evolution involves clonal and subclonal mutations. There is a transition between different types. Subclonal mutations are present only in a subset of cells within the tumor. These arise after the most recent common ancestor of all cancer cells in the tumor sample. In contrast, Clonal mutations occurred before the closest common ancestor of all cancer cells. Clonal mutations are Therefore, mutations, for example, in that case the entire gene locus is lost in a subset of cells Unless there is some mechanism that eliminates the structural variations, all of the tumor It is present in cells. ctDNA originates from tumors and is found in tumors or circulating tumor cells (CTCs). These originate directly from the primary tumor and can leak out and enter the bloodstream or lymphatic system. These are viable and intact tumor cells. The precise mechanism by which ctDNA is released. The reason is unclear, but it could be apoptosis and necrosis from dying cells, or surviving tumor cells. It is assumed that this is accompanied by the release of activity. Circulating tDNA (ctDNA) can be highly fragmented. In some cases, the length is approximately 100-250 bp, for example, a flat length of 150-200 bp. It may have a uniform fragment size. ctDN in a sample of acircumcirculating cell DNA isolated from cancer patients The amount of A varies greatly: a typical sample contains less than 10% ctDNA, but M Many samples from patients being evaluated for RD contain less than 0.01% ctDNA. Some samples may contain more than 10% ctDNA. The ctDNA molecule is multiply In some cases, these can be identified because they contain tumorigenic mutations.
[0057] As used herein, the term “sequence variation” refers to the location and nature of the sequence modification. It refers to a combination of types. For example, an array variation refers to the position of the variation. Therefore, which type of substitution (for example, G to A, G to T, G to C, A to G, etc.) It can be named by whether an insertion / deletion of G, A, T, or C is present in the position. Sequence variations can be substitutions, deletions, insertions, or rearrangements of one or more nucleotides. In the context of this method, sequence variations include, for example, PCR errors and sequencing errors. It can be generated by errors or genetic variations.
[0058] When used in this specification, 「 The term "genetic variation" refers to the presence of [something] in nucleic acid samples. Variations that exist or are considered likely to exist (e.g., Nucleochi) This refers to genetic variations (substitutions, indels, or rearrangements). Genetic variations are derived from any source. It can be that genetic variation occurs through mutation (e.g., somatic variation). It can be generated, or it may be part of the germline in organ transplantation or pregnancy. When sequence variations are called as genetic variations, the call is in the sample. This indicates that it is likely to contain variations, and in some cases, "call" This can be inaccurate. In many cases, the term "genetic variation" is used to mean "variation." The term can be replaced with "different". For example, the method is caused by cancer or mutation. When used to detect sequence variations associated with other diseases that may occur The term "genetic variation" can be replaced with the term "mutation."
[0059] Where used herein, depending on the context, the term “calling” refers to a specific genetic trait. Whether the sequence contains genetic variation, or whether the sample contains genetic variation, or the sample This could indicate whether it contains cancer DNA.
[0060] As used herein, the term "threshold" means the threshold required to make a call. It refers to the level of evidence (e.g., ratio).
[0061] As used herein, the term “value” means a numerical value that can indicate the strength of evidence. , letters, words (e.g., "high", "medium", or "low"), or descriptors (e.g., "+ This refers to "++" or "++"). The value is one depending on how the value is analyzed. It may contain one component (for example, a single component) or two or more components.
[0062] As used herein, the term "aliquot" refers to a portion of a sample. If three volumes are removed independently from the same sample, each volume is called an aliquot. It is possible. Aliquots do not need to be of the same volume.
[0063] As used herein, the term “cancer-associated cells” refers to a portion of the patient’s cancer cells. It means cells that are or are genetically related to cancer. Cancer-related cells are one of the cells of a solid tumor. It may be a hematological cancer or a solid tumor. The presence of cancer-associated cells in the patient is This could be evidence that not all cancer cells were removed or killed during treatment. Cancer-related cells They have substantially the same somatic mutations as the patient's cancer cells, and in some cases, one of the cancers These cells may be descendants of the cells above. Cancer-related cells may result from minimal residual disease, or from tumors. Incomplete removal of the tumor, incomplete treatment, or recurrence of cancer at the primary or distal site. rence) or relapse, and / or tumor metastasis (including micrometastasis) It can be generated by.
[0064] When used herein, "a sequence variation associated with (or present within) the patient's cancer" The term "treatment" refers to a condition found in the genome of a patient's cancer cells or prior to any cancer treatment. This is intended to mean a somatic mutation that was present in the genome of the patient's cancer cells. This could also refer to epigenetic changes present within the cancer sample.
[0065] As used herein, “minimum residual disease” (MRD) refers to the condition after treatment for the purpose of cure. MRD refers to the presence of malignant cells. MRD is also referred to as "molecular residual lesion" or "residual" in some publications. It can be referred to as a "lesion."
[0066] As used herein, the term “detecting recurrence” refers to the specific characteristics of mutant DNA. This refers to detecting tumor recurrence through monitoring. In this context, the term "early detection" is used. The term refers to the detection of tumor recurrence through conventional standard treatment / monitoring methods such as radiography. This refers to the detection of mutant DNA before it can actually be detected. This is described, for example, below. As such, the presence of ctDNA in cfDNA was collected sequentially at multiple time points. This can be achieved by monitoring blood samples.
[0067] The term "cancer" as used herein refers to any cancer characterized by uncontrolled cell division. It is used to refer to a disease. Cancer refers to cancer of the blood (i.e., hematological cancer), for example. , leukemia, lymphoma, or multiple myeloma, or cancer is neoplastic, for example In other words, abnormal histoma is a condition in which cells proliferate and divide excessively, or fail to die when they should. It may be related to tumors. Neoplastic cancers, such as lung cancer, breast cancer, or liver cancer, are related to solid tumors. They are connected.
[0068] The term "cancer DNA" refers to DNA derived from cancer cells. If a person has blood cancer, a collection of cells isolated from the patient's lymph, bone marrow, or circulating blood is used. It can be present in DNA isolated from tumors. Cancer DNA derived from solid tumors has cfDNA odor This can be done by referring to tDNA or ctDNA.
[0069] The terms "error probability distribution" and "error probability distribution model" refer to observations (typically, A distribution to estimate or model the probability that the rian allele fraction is due to error. These terms refer to "high signal background events" (DNA damage or very early events). (Possible cause: cycle PCR errors during the period) and "Estimated background error rate" It captures both sequencers and PCR polymerase "errors". Examples of fabrics are shown in Figures 13A and 13B.
[0070] In the context of analyzing "collective outcomes," the term "collective" refers to positive outcomes. This refers not only to the results for all variants and aliquots (any statistically). Outliers, or for example, they are not present in tumor DNA, or Buffy Coat D (Excluded because it exists in NA; other variants are excluded).
[0071] Other definitions of terms may appear throughout the specification. It should be noted that in some cases, the draft may have been prepared to exclude certain elements. Therefore Therefore, this description is related to the enumeration of elements of the claims, such as "simply" and "only". It is intended to serve as an antecedent for the use of derogatory terms or for the use of "negative" definitive terms. It is being done. [Modes for carrying out the invention]
[0072] Before describing the present invention in more detail, it should be noted that the present invention is not limited to the specific embodiments described, Therefore, it should be understood that this is naturally subject to change. The scope of this invention is as follows: Since the terms used herein are limited solely to the claims of the patent, the terms used herein are specific to the patent. This is solely for the purpose of explaining the implementation method and is not intended to be limiting. This should be understood.
[0073] If a range of values is provided, the range between the upper and lower limits of that range, unless otherwise specified in the context. Unless otherwise specified, each intermediary value up to one-tenth of the lower limit unit, and the stated It should be understood that any other stated or intervening values within that range are included in the present invention.
[0074] Unless otherwise defined, all technical and scientific terms used herein are defined in accordance with the present invention. It has the same meaning as generally understood by those skilled in the art to which it belongs. Any method and substance similar to or equivalent to those used in the present invention may also be used in the implementation or testing of this invention. Methods and substances that can be used, but are preferred, are described here.
[0075] All publications and patents referenced herein are subject to the terms of each individual publication or patent. This specification incorporates, by reference, as if specifically and individually indicated to be incorporated. To disclose and describe the relevant methods and / or materials by which the publication is cited, This is incorporated herein by reference. Any citation of any publication is based on its disclosure prior to the filing date. Therefore, it is acknowledged that the present invention does not have the authority to precede the publication of the prior art. It should not be interpreted as such. Furthermore, the published date shown may differ from the actual published date. This may be the case, and it may be necessary to check separately.
[0076] As used herein and in the appended claims, the singular forms "a", "an", and "The" can refer to multiple things unless the context explicitly indicates otherwise. This must be noted. The claims exclude any optional elements. It should be noted that it may have been drafted in a different way. Therefore, this description is a patent The use of exclusive terms such as "simply" or "only" in relation to the enumeration of elements of a claim, or It is intended to serve as an antecedent for the use of "negative" modifiers.
[0077] As will be apparent to those skilled in the art when reading this disclosure, the individual Each embodiment may be adapted to several other embodiments without departing from the scope or spirit of the present invention. It can be easily separated from any of the characteristics of the state, or combined with such characteristics. It has individual elements and characteristics that enable it. Any enumerated method is in the order of the enumerated events. They may be carried out in order, or in any other logically possible order.
[0078] As may be obvious, each of the multiple aliquots evaluated for two or more target regions The assay reliably detects cancer DNA, sometimes referred to as the limit of detection (LOD). It may have different lower limits. It may also be called the limit of quantification or LOQ. However, there may be different limitations in accurately quantifying the amount of cancer DNA. For the assay to be most useful, in some cases, either LOD or LOQ is required. Alternatively, it may be important to obtain accurate estimates of both. Such estimates may be related to clonality. Mapping possibility, estimated error rate, estimated proportion of high-signal background events, region The presence of copy number increase within the sequence, or each sequence variation associated with cancer in the targeted patient. It can be obtained by combining factors that may include amplification of . The number of aliquots, the total number of sequencing reads for the targeted region, and the inputs to each aliquot. It may include factors specific to library preparation and sequencing runs, which may include the number of molecules involved.
[0079] As described above, cancer DNA in DNA test samples from patients (for example, cancer patients) A method for obtaining is provided. In some embodiments, the method involves multiple a of the test sample. Recoating (for example, at least two, at least three, at least four, or fewer of the test samples) Arrange at least five, or at least six, aliquots, and for each aliquot... , each of two or more target regions having sequence variations present within the patient's cancer (for example If so, at least 3, at least 5, at least 10, at least 20, and at least 50, at least 100, at least 1000, or at least 5000 labels The method may include generating sequence reads corresponding to the target region of the test DNA. The sequence of 3 to 10 aliquots of the sample is determined, and for each aliquot, 8 to 100 This may include generating sequence reads corresponding to the target region. Very generally speaking, sensitivity This is achieved by increasing the number of aliquots, thereby increasing the number of variants. It can be increased by increasing the number of aliquots and variants. It is possible. For example, in some embodiments, the method involves at least two test samples (e.g., Determine the arrangement of 3 or 4 aliquots, and for each aliquot, each has an arrangement variation. This may include generating sequence reads corresponding to 10 or more target regions having a specific configuration. In another embodiment, the method involves arranging at least 10 aliquots of the test sample. For each aliquot, there are two (for example, three or four) arrangement variations. This may include generating sequence reads corresponding to more than 1) target regions. In fact, the method is: If a sufficient number of sequence variations are to be analyzed, do so using a single aliquot. It is possible.
[0080] The method is as follows: (a) Arrange multiple aliquots of the test sample, and for each aliquot, Each corresponds to two or more target regions with sequence variations present within the patient's cancer. (b) For each aliquot, for each target region: ii. Determine the number of sequence reads that have sequence variations. And, iii.i. and ii. are one or more error probability distribution models of sequence variations. The comparison is made by comparing one or more models from DNA that does not contain sequence variations. The results obtained and compared, and the collective results of step (b) are combined, to determine the test sample This may include determining whether cancer DNA is present within it.
[0081] In these embodiments, different aliquots are different aliquots of the same sample (i.e., (including parts). As you can see, different barcode sequences are attached to different samples. This allows for different samples to be pooled before sequencing.
[0082] flowchart Some of the workflows of this method are shown in the attached flowcharts (Figures 1-10). The flowchart is considered to be mostly self-evident.
[0083] Before describing the method in more detail, let me explain that this method can be used for both solid tumors and hematological cancers. It should be noted that cancer DNA can be detected from this. Therefore, this claim When the term "cancer" is used, it refers to blood cancers and solid tumors. In this embodiment, the method involves cancer DNA (or) in cfDNA (e.g., circulating cfDNA) More precisely, tumor DNA can be identified. In embodiments of hematological cancers, the method involves bone marrow, DNA extracted from cells taken from lymph nodes or circulating leukocytes, or cfD Cancer DNA can be identified in NA. For example, in embodiments of hematological cancers, AML patients (treatment (Previously) Bone marrow aspirate can be collected, and variants in AML can be identified, next Then, after treatment, further bone marrow aspirates, cell-free DNA, or urine tests are performed to check if the patient still... It can determine whether or not someone has cancer DNA.
[0084] In addition, the nucleic acids analyzed by the method may be DNA or RNA. This disclosure applies to DNA ( Specifically, it describes embodiments that utilize ctDNA. However, The law also works when using RNA (or cDNA) made from the same source. It is.
[0085] In addition, this method is described in detail using an example that utilizes "amplicon" sequencing. However, this method involves attaching a molecular barcode or index, for example, to the nucleic acid before amplification. This can be easily applied to methods that utilize random sequences. Molecular barcode sequences are of size and The composition can vary widely, and the following references are barcode sequences suitable for specific embodiments. Provides guidance for selecting a set: Casbon (Nuc.Acids Res.2011,22 e81), Brenner, U.S. Patent No. 5,635,400 ,Brenner et al,Proc.Natl.Acad.Sci.,97:16 65-1670(2000), Shoemaker et al, Nature Gen. etics,14:450-456(1996), Morris et al, European Patent Public notice 0799897A1, Wallace, U.S. Patent No. 5,981,179, etc. Specific In this embodiment, the barcode sequence is 2 to 36 nucleotides, or 6 to 30 nucleotides. or may have a length in the range of 8 to 20 nucleotides. For example, an aliquot-based sequence. The determination can be made against indexed DNA, and the number of molecules / the number of molecules present The rate can be estimated using the index sequence in each aliquot.
[0086] In the pre-calibration method shown in Figure 5, the type of variant from which the error probability distribution is generated and Note that classes can differ. For example, a particular variant may differ from its surroundings. It can be analyzed within the context of the sequence. This is DNA that is not expected to contain variants (e.g., The target region is sequenced using DNA from a healthy donor, or the wild-type sequence is used. Synthetic DNA / RNA is applied to the target region (outside the variant region) including the barcode. By incorporating a spike, it becomes possible to separate the barcode and spike from the test reaction. This can be achieved. In another example, a particular variant is in the context of the variant's class. It can be analyzed internally. The class of variants is the same type of variant (for example, A > T). Indels such as SNVs, TTTT insertions, dinucleotide substitutions such as CT>AA, and transpositions. Or conversion, single-nucleotide variant and 1 to 5 bases of either 3', 5', or both (for example) , A has 5'TTCA, A>T (TTCAA>TTCAT), or A has 5'T and It has 3'G and includes A>T(TAG>TTG). Alternatively, the variant is the above They can be grouped into classes, but some or all of the variant's bases 3' and / or 5' are It can be one of several bases described by the IUPAC degenerate nucleotide code. (For example, if A has 5'K and 3'S, then A>T(KAS>KTS)(K=G / T) (and S=C / G). In an alternative embodiment, the local array context is N is 1 to 100 This involves selecting a window of N3' and / or 5' bases around the desired variant. And, the base change at each position, the type of base change at each position (e.g., rearrangement or transposition), By extracting different sequence descriptors such as the distance from the end of the imager and the distance from the repeating sequence, These are then explored, and these are subsequently analyzed using heuristic combinatorial scoring or machine learning methods. By using (unsupervised or supervised) categorical error rate classes (e.g.) For example, they are combined to predict high, medium, or low numerical error rate values. This method involves a penalty score in the form of a multiplicative factor, specifically a mononucleotide repeat. The probing of variants that are close to predefined sequence features, such as repeating regions or similar elements. It is assigned to a constant error rate. This analysis is not expected to include variant classes. This can be done by sequencing DNA (for example, DNA from a healthy donor). Yes, it is possible. In this embodiment, each variant class is at least once (and ideally, for example, For example, a sufficient area is targeted, as expressed by 10 times, 50 times, or more than 100 times. The sequence must be determined.
[0087] In addition, the number and types of error probability distributions can differ for each variant (or class). In some versions, there is a single distribution for all errors. In other embodiments, There are multiple distributions that separate different types of errors. In some embodiments, each vari There are two error distributions for Ant, one of which is "Estimated background error These are "Lar rate" values. These are typically from the later stages of library preparation (e.g., PC These are sequencing errors and PCR errors that occur after the first few cycles of R. Although the occurrence rate is quite low, when it does occur, it is at a fairly high level, typically in the sample. There are events that are at the same level as the actual variant (in terms of variant allele frequency). These "high signal background events" occur during the first few months of library preparation. This includes DNA damage and polymerase errors in the cut or pre-amplification process. These are related to the second... This can be captured by the distribution of (for example, one or two in the estimated background error rate) (One term distribution and one high-signal background event). In some embodiments, different The distribution is used for the estimated background error rate and high-signal background events. (For example, beta distribution and high signal background in the estimated background error rate) (Binomial distribution for round events).
[0088] In some embodiments of each variant, the same variant class (e.g., 2bp3') And 2bp5') are used for both distributions. However, in some embodiments, The conclusion of two different distributions for different error processes (e.g., DNA damage and PCR errors) Since this can be the case, for each variant, there are two different variants for the two distributions. The class is used.
[0089] The control substances and methods for generating one or more distributions may also differ. For example, The rate distribution was generated in the same library preparation and tested using control DNA beforehand. When performing the test as a material, or when evaluating the test sample in advance, it is expected that it will contain variants. It can be prepared using all bases except for the base specified.
[0090] In all cases, the same sequencing process (library preparation) is used to generate the model. (including sequencers) and preferably the same sample type and extraction method (e.g., cfDN) (cfDNA extracted from blood collected in a blood collection tube) should be used.
[0091] In some cases, different models are generated for a set of different DNA inputs, and most A test sample is analyzed using a model with a matched DNA input. For example, for each ali Define the maximum, minimum, and central DNA inputs for each coat, and then obtain one or more distributions for all three of all classes of variants to be tested. The test sample When evaluated, its DNA input is compared to the distribution with the closest match. Preferably, there will be dozens, hundreds, or thousands of samples tested to build the model.
[0092] There will be. The distributions can be stored in a database and / or downloaded from a public database.
[0093] In some embodiments, the amount of cancer DNA (as shown, for example, in FIG. 8) can be quantified using the method. In these embodiments, the average or central variant allele fraction (across variants and aliquots), the corrected average or central variant allele fraction (generated by subtracting a previously determined offset or base line error rate), the maximum likelihood (testing a range of levels and determining the most likely one), estimating the tumor fraction: the maximum likelihood, a grid search or expectation maximization search method that selects the tumor fraction giving the Bayesian posterior, or the estimated variant allele number for each variant (and optionally each aliquot) can be used singly or in combination to determine the amount of cancer DNA in the test sample, the likely range of amounts in the test sample or estimated tumor fraction. In another embodiment, the amount of cancer DNA is counted as the number of variant positive target regions (target regions exceeding a threshold) in each aliquot, and this is compared to the total number of target regions multiplied by the aliquot. It can be done.
[0094] In some embodiments, the amount of cancer DNA (as shown, for example, in FIG. 8) can be quantified using the method. In these embodiments, (across mutants and aliquots) Average or central variant allele fraction, corrected average or central variant allele fraction (generated by subtracting a previously determined offset or base line error rate), maximum likelihood (testing a range of levels and determining the most likely one), Estimating the tumor fraction: maximum likelihood, a grid search or expectation maximization search method that selects the tumor fraction giving the Bayesian posterior, or the estimated variant allele number for each variant (and optionally each aliquot) Of the maximum likelihood, selecting the tumor fraction that gives the Bayesian posterior, Grid search or expectation maximization search method, or the estimated variant allele number for each variant (and optionally each aliquot) One or a combination of summing the number of molecules can be used to determine the amount of cancer DNA in the test sample, the likely range of amounts in the test sample or estimated tumor fraction. In another embodiment, the amount of cancer DNA is the number of variant positive target regions (target regions exceeding a threshold) in each aliquot, and this is compared to the total number of target regions multiplied by the aliquot. One or a combination of summing the number of variant molecules for each variant (and optionally each aliquot) can be used to determine the amount of cancer DNA in the test sample, the likely range of amounts in the test sample or estimated tumor fraction. In another embodiment, the amount of cancer DNA is the number of variant positive target regions (target regions exceeding a threshold) in each aliquot, and this is compared to the total number of target regions multiplied by the aliquot. The amount of cancer DNA in the test sample, the likely range of amounts in the test sample or estimated tumor fraction can be determined. In another embodiment, the amount of cancer DNA is the number of variant positive target regions (target regions exceeding a threshold) in each aliquot, and this is compared to the total number of target regions multiplied by the aliquot. In another embodiment, the amount of cancer DNA is the number of variant positive target regions (target regions exceeding a threshold) in each aliquot, and this is compared to the total number of target regions multiplied by the aliquot. Count the number of target regions (target regions exceeding the threshold), and divide this by the total number of target regions multiplied by the aliquot. By comparing and applying Poisson correction to the positive result fraction, the standard per aliquot is obtained. This can be determined by quantifying the average number of variants containing the target sequence per target region. In some embodiments, to perform more accurate quantification, the entire set of variants is used. The estimated proportion of high-signal background events can also be used in Poisson correction. ru.
[0095] General methodology In some embodiments, (a) multiple aliquots of the test sample are arranged, and each aliquot Regarding the marker, each has two or more labels with sequence variations present within the patient's cancer. (b) For each aliquot, generate sequence reads corresponding to the target region For the region, we will derive an estimate of the number of molecules with sequence variations, or sequence variations Calculate the probability that at least one molecule with a sequence exists, or the total number of sequence reads The frequency of sequence reads in (a) that have sequence variations compared to the number exceeds the threshold. (c) Determine the value of (b) the estimate, or the probability or frequency used in the test. A method comprising determining whether cancer DNA is present in a sample. Several embodiments Step (b) can be performed by the thresholding approach described below, alternatively. In this embodiment, step (a) aliquots as long as there are a sufficient number of target regions. It can be done without any effort.
[0096] In some embodiments, for each aliquot and each target region, the sequence in the test sample The number of molecules having relation, or at least one molecule having sequence variation. The probability that exists is (b)(i) the number of sequence reads of (a) that have sequence variations, (ii) Total number of sequence reads in (a), and (iii) Estimated background of sequence variations It is estimated using the round error rate. (iii) The background error rate is It can be represented by a Ra probability distribution. In addition, at least one molecule has sequence variations. The probability of existence is estimated using the number of molecules entered into each aliquot in (a). (iii) The estimated background error rate is, excluding the target variant, step ( a) Data of the control base obtained in the preceding sequencing reaction, for example, Or, for example, from publicly available information from prior sequencing reactions and / or current The estimated background is derived from the sequencing reaction using any convenient method. For example, the estimated background The Derrler rate is estimated by analyzing the control sequencing reads generated in step (a). obtain.
[0097] In any embodiment, the background error rate is estimated using a probability distribution. This is possible. In some embodiments, two distributions of the same family (for example, two binomials) (Distribution) is possible, or if two different families are used, background One distribution for the error rate, and another for the estimated proportion of high-signal background events. There may be such things. As described above, in any embodiment, the estimate is the number of existing variants This is a probability distribution over the number of offspring.
[0098] In any embodiment, (c) if cancer DNA is present in the sample, (ii) ) If cancer DNA is not present, calculate the likelihood ratio between the likelihoods that observe the estimate in (b). can be performed by doing so. Similarly, in any embodiment, (c) is for each target region and aliquots: (i) when cancer DNA is present, (ii) when cancer DNA is not present calculate the likelihood ratio (LR i ) between the likelihoods of observing the estimated value in (b). In these embodiments, the individual likelihood ratios LR i are combined into a cumulative LR score (equal to the sum of the log-likelihoods, the product of the LR ) across all regions and aliquots of the sample. i In these embodiments, when cancer DNA is present in the test sample, the likelihood of observing the estimated value in (b) can be calculated based on (i) the estimated value or probability in step (b), and optionally (ii) the estimated value of the cancer DNA fraction in the test sample. Similarly, when cancer DNA is not present in the test sample, the likelihood of observing the estimated value in (b) can be calculated based on (i) the estimated value or probability in step (b), and (ii) the estimated proportion of high-signal background events. In any embodiment, step (c) can be calculated by using a mixture model that incorporates (i) the estimated value or probability in step (b), (ii) the estimated proportion of high-signal background events, and optionally (iii) the estimated value of the cancer DNA fraction in the test sample. For example, in some cases, step (c) can further include comparing the output or likelihood ratio of the mixture model to a threshold, and an output above the threshold indicates that the test sample contains tumor DNA. The threshold is at least 10 or at least 100 or at least 1000 or at least 10,000 samples that do not contain cancer DNA (or at least where it is not known to have cancer DNA) are assayed, and the threshold exceeds the identified signal in the control sample. In any embodiment, step (c) can be calculated by using a mixture model that incorporates (i) the estimated value or probability in step (b), (ii) the estimated proportion of high-signal background events, and optionally (iii) the estimated value of the cancer DNA fraction in the test sample. For example, in some cases, step (c) can further include comparing the output or likelihood ratio of the mixture model to a threshold, and an output above the threshold indicates that the test sample contains tumor DNA. The threshold is at least 10 or at least 100 or at least 1000 or at least 10,000 samples that do not contain cancer DNA (or at least where it is not known to have cancer DNA) are assayed, and the threshold exceeds the identified signal in the control sample. For example, in some cases, step (c) can further include comparing the output or likelihood ratio of the mixture model to a threshold, and an output above the threshold indicates that the test sample contains tumor DNA. The threshold is at least 10 or at least 100 or at least 1000 or at least 10,000 samples that do not contain cancer DNA (or at least where it is not known to have cancer DNA) are assayed, and the threshold exceeds the identified signal in the control sample. 10,000 samples that do not contain cancer DNA (or at least where it is not known to have cancer DNA) are assayed, and the threshold exceeds the identified signal in the control sample. is at least 10 or at least 100 or at least 1000 or at least 10,000 samples that do not contain cancer DNA (or at least where it is not known to have cancer DNA) are assayed, and the threshold exceeds the identified signal in the control sample. Alternatively, the false positive rate determined using a control sample is 1% or less, 0.1% or less, or 0.01%. This can be determined by selecting a threshold that is estimated to be below the following. The method involves identifying a patient as having cancer cells if the result is above a threshold. , and may further include, for example, administering therapy to the patient. In these embodiments, the patient In some cases, the patient may have previously received the first treatment. In these cases, the method may differ from the first treatment. This includes administering a second therapy to the patient.
[0099] In any embodiment, the method determines the cancer D in the test sample based on the estimate from step (b). This may further include determining the range of the amount of NA or the amount of cancer DNA likely to be present. For example, (i) calculate the mean or central variant allele fraction, ( ii) Maximum likelihood analysis, (iii) Bayesian posterior analysis, (iv) For each variant and each aliquot (v) Each aliquot The number of variant-positive target regions is counted, and this is used to determine the aliquot target region. By comparing it to the total number of regions and applying Poisson correction to the fraction of positive results, the aliquot By quantifying the average number of variants containing the target sequence per target region per area, This can be done. This type of analysis involves calculating the number of starting molecules in digital PCR. It is done for that purpose, and can be adapted from there.
[0100] In any embodiment, the method obtains from the patient between at least a first and a second time point This procedure can be performed on a sample, with the first time point being before treatment and the second time point being after treatment. Yes, the method involves measuring the amount of cancer DNA or the presence of cancer DNA between the first and second time points. This includes determining whether there is a change in the range of the quantity. This change is determined by point estimates, confidence intervals, or Both can be used to determine a significant decrease, and a significant change indicates that the therapy is effective. The absence or increase indicates that the therapy is ineffective. In these cases, at least 20 %, at least 30%, at least 50%, at least 70%, or at least 90% The change may be considered significant. In some embodiments, the change is greater than a threshold, such as 50%. Furthermore, when quantifying cancer DNA at the first and second time points, the confidence intervals do not overlap. In addition, the change is considered significant. In these embodiments, a significant decrease indicates that the therapy is effective. This indicates that there is no significant change or an increase, which indicates that the therapy is not effective.
[0101] In any embodiment, the estimated cancer DNA fraction, the number of DNA molecules attached to each aliquot. , and optionally, the number of times each variant appears in individual cancer calls (copy number analysis). In a statistically unlikely number of aliquots (which can be determined through) The sequence variations that are to be performed are excluded from the result of step (b) before step (c). In any embodiment, step (a) is to obtain at least three aliquots, for example, 3 , 4, 5, 6, 7, 8, 9, 10, 11, or 12 or more aliquots are sequenced. It may include the following.
[0102] In some cases, if the variant is amplified in cancer cells, it is all It can be expected that recoating is present. Therefore, this part of the method is each barrier in cancer cells. Enter the number of copies of the ant, and use this to determine which of the variants should exceed the threshold. Further improvements can be made by estimating the likely number of recoats.
[0103] In some embodiments, step (a) also includes aspirates, biopsies, and other samples derived from the same patient. Or cancer DNA from surgical specimens, buffy coat DNA, cheek swab DNA, whole blood DNA, Adjacent normal DNA, i.e., sets adjacent to the tumor that look like normal or reference DNA. This may include sequencing positive and / or negative controls that may include at least one of the fabrics. The sequencing of these samples may be performed simultaneously with the sequencing of the test samples, or the sequencing of the test samples may be performed separately. This can be done before or after the decision.
[0104] In any embodiment, variants not detected in cancer DNA are excluded. Detected together or individually in buffy coat, cheek swab, adjacent normal blood, or whole blood. The variant may be excluded.
[0105] In any embodiment, two or more target regions are at least 2, at least 4, and at least 10, at least 20, at least 50, at least 100, at least 500, less There are at least 1,000, or at least 5,000, target regions. In many embodiments, 2 to 200, for example, 10 to 100, target regions may be investigated. Step (a) Sequence variations can be independently single-nucleotide variants, indels, and dinucleotide substitutions (DBS). ), transposition, rearrangement, variable number tandem repeats, short tandem repeats, or integration into the patient genome. It could be an embedded viral genome (such as HPV).
[0106] In some embodiments, the variant is 5-methylcytosine (5mC) or 5-Hyd Epigenetic variants, not sequence variants like roxymethylcytosine. This may be the case. In certain embodiments, sequence variants and epigenetic variants Two or more elements exist that are less than 10bp apart, less than 50bp apart, or less than 100bp apart. Selected when available.
[0107] As described above, the sequence variations analyzed in the method are pre-identified sequence variations. This is a variation. For example, sequence variations include (i) tissue biopsy containing cancer cells or (ii) DNA or RNA isolated from, or cancer tissue obtained from surgery containing cancer cells. (iii) Sequencing of isolated DNA or RNA samples, or (iii) Cell-free D (iv) NA or RNA, or DNA or RNA isolated from circulating cancer cells The sample can be identified by sequencing, for example, from the same patient before any treatment. Regarding blood cancers, sequence variations include, for example, bone marrow and circulating blood cells. Alternatively, it can be identified by sequencing a sample of DNA or RNA from a lymph node. In some embodiments, both DNA and RNA are sequenced and combined. These variants are identified. These sequence variations can be determined by sequencing the entire genome. By or whole exome, genes that are frequently mutated in cancer (e.g., C OSMIC (found in the Cancer Gene Census), mitochondrial gene M, regions of common structural rearrangement (e.g., common fusions or common amplifications such as MYC) Common amplification regions, common rearrangement regions (e.g., chromothripsis), common Localized hypermutation regions (e.g., cataegis), or 80% or 9% of the target patient population. 0% or more than 95% have enough identified mutations to reach the required sensitivity. The region of the genome that has been identified as typically containing a sufficient number of mutations for a given target cancer type. It can be identified by sequencing one or more of the regions (the required sensitivity is this sensitivity) The number of variants required to satisfy the condition is predetermined, and this is the target To determine the number of Mb in the genome, the rate of mutations per megabase (Mb) and the target... (This is compared to the variability among patients in cancer type.)
[0108] In some embodiments, the viral sequence is one that has been incorporated into the human genome, and They are targeted to identify the location where they are incorporated. In some embodiments, the whole geno Either the genotype or a specific region of the genome is subject to epigenetic changes, for example, Whole-genome bisulfite sequencing, TET-assisted pyridineborane sequencing, enzymatic sequencing Reduced display of Chill sequencing, bisulfite sequencing, and methylated DNA immunoprecipitation sequencing. or evaluated by target bisulfite sequencing. Epigenetic and genetic Both types of transmission can also be identified by the array. In some embodiments, Assays that utilize either methylation changes and / or sequence variants use ctDNA This assay is performed to detect cancer early through the identification of these changes. In such embodiments, the patient may have ctDNA and therefore cancer. When identified as such, epigenetic and / or present in the patient's ctDNA sample Sequence variants are identified and selected for targeting.
[0109] Hotspots can also be sequenced. Alternatively, sequence variations can be , which can be identified by RNA-seq, and optionally, to target a specific type of RNA. RNA selection / depletion, such as poly(A) selection or ribosomal RNA depletion, is used.
[0110] In some embodiments, multiple candidate sequence variations are first identified, and then... Sequence variations may be selected. In some embodiments, the variations are It may be ranked, then the "best" variation may be selected, and the variant may be filtered. Any variants that may be processed and are not ideal for tracking are removed or filtered first. They can be analyzed and then ranked. In some embodiments, the sequence variations are Filtering, scoring, or ranking based on one or more of the following: i) Variants present throughout the tumor are preferred, clonal, ii) The variant read is a region or a pre-annotated blacklister - Tried allies of predictive PCR amplicons designed to amplify presence within a region Mapping based on the comments is difficult, and there are overlapping repeats and homopolymer regions. Notation should be avoided, mapping possibility, iii) Variants with a high error rate are disadvantaged or filtered out. It should be, estimated background error rate, iv) Estimated proportion of high-signal background events, where bases with a low proportion are preferred. , v) Distance from another selected variant. In some embodiments, the variant is a G They should be placed at equal intervals throughout the entire Nomu, not densely packed, for example, any dye 10% of all variants are present on the chromosome, or any chromosomal arm, or any 1Mb region. The following exists. This is a region of the genome where many variants for tracking no longer exist. This is to prevent the loss of regions (for example, due to the loss of chromosome arms during evolution). In this application, two variants are targeted by a single sequencing read and reside on the same chromosome. If it is close enough to be present, such a variant is preferred. vi) Predictive ability to determine sequences, vii) Variants present in multiple copies in a single cancer cell are preferred, copy number Existence within the region of increase or amplification, viii) Any germline variant that can be used to enrich mutant alleles Proximity, ix) The likelihood is somatic, x) Somatic but clonal hematopoiesis with undetermined potential, etc., not originating from the target cancer. Likelihood, xi) It is preferable to avoid such areas, which are frequently found in the cancer type being tested. An existence in a realm that is lost, xii) Likelihood that the variant is a common SNP / polymorphism xiii) Variants that arise from specific protocols / sequencing methods / capture kits Likelihood is a fact This is through the proliferation of variants in current and / or previous reaction / sequencing batches. This includes variant profiles that match those of known FFPE / other errors. nothing.
[0111] In some embodiments, all or a combination of these factors are scored, Ants are ranked by score and then selected. In some embodiments, Regions of the genome are ranked rather than being specific variants. In such embodiments, The genome can be divided into overlapping or non-overlapping windows. These windows can be, for example, 10bp, 50bp, or 100bp in length. The 5bp, 25bp, and 50bp values may overlap, or they may not overlap at all. As will be clear, the window is smaller than the typical length of DNA from the test sample. The read length should be shorter than the sequencing read length of the intended sequencing platform. Therefore, in high molecular weight DNA and long-read sequencers, the window is, for example, It can be 100, 1000, or 10,000 bp. Illumina sequencer - And in cfDNA, the window is always 160 bp (typical length of cfDNA). It should be full. In a preferred embodiment, the window is 20 to 100 bp. It involves overlaps that are half the length of the entire window. After scoring each variant, the size of each region The core combines the scores of all variants within the domain and can optionally map them. This may include the ability to predict sequences, and the presence within a region of copy number increase or amplification. It is generated by combining one or more scores of region-specific features. In one embodiment, regions can be ranked and the best region can be selected, and the assay can be performed. It is designed to target these areas. The advantage of such a method is that the information is tested in the DN. Multiple variants from a single molecule of A (when the variants are in cis form on the same chromosome) To assign weights to the genomic regions that can be obtained, and to determine if variants are in the same genomic region. However, when trans, that is, when located on another chromosome, it targets a single region. It simply involves obtaining more information from it.
[0112] In some embodiments, different sets of PCR primer pairs (forward and reverse) The combination is designed to target multiple identified candidate sequence variations or regions. These are based on the following characteristics for each variation or domain: To identify a single best primer pair, selection, scoring, filtering, or r To be ranked: i) Presence of repeating regions within the primer sequence (e.g., homopolymer of >= 6 nucleotides) (Avoiding the area), ii) The presence of known single nucleotide polymorphisms in the primer sequence (this is avoided or the SNP is Tumor sequencing is used to confirm its presence, iii) In silico PC between primers and / or between primers and amplicon regions Based on R and / or local alignment and / or 3'-based alignment Since it is generated using one forward and one reverse primer, sequencing is possible. Predicted formation of unintended PCR products is highly likely to occur (such primers (There are high penalties for pairings.) iv) Similar to iii), but unintended PCR production which is highly likely to be impossible to sequence. Predicted formation of the substance (either two forward primers or two reverse primers) Either is produced, and such a product does not contain both necessary sequencer adapters. Therefore, compared to (iii), such a primer would not enable sequence determination. (There is a low penalty for this combination.) v) Total amplicon size in nucleotides, vi) The number of times the predicted PCR product aligns with a region of the genome beyond the predicted target. (Ranking scores can be based on multiple mappings.) vii) The number of times the primer sequence aligns to a genomic region other than the intended target. viii) Close (i.e., based on a predefined threshold, 50, 100, or (less than 150 nucleotides) By using forward and reverse primers other than the intended ones The alignment of the primer pair configured in this way is performed a certain number of times. ix) The combined score of all variants present within the target amplicon.
[0113] In some embodiments, if the score exceeds a threshold, the primers are used to define these features. Filtering is performed based on some or all of the features. In some embodiments, the features Composite scoring based on several or all combinations of linear or polynomial formulas is optimized for multiples. Used for selection. In some embodiments, cancer DNA containing a sample or cell line is used. A large number of variants were selected, and multiple multiplicative PCR panels were applied to these variants. The process is designed. A dilution series of cancer DNA into normal DNA is generated, followed by multiple multiplex PCR. The assay is used to generate a sequencing library from DNA. The process is optimally repeated with at least 10 or at least 100 samples. Some or all of these signs, along with sequencing signals, detect cancer DNA in the test sample. To determine the optimal combination of primers for this purpose, a machine learning system or neural network is used. It is input into the network.
[0114] In some embodiments, a reagent that targets the variant (e.g., capture bait or multiplex) is used. PCR primers can be designed for all variants, and then for each variant or region Rather than selecting a specific area, the best combination of primer or bait is chosen. Each primer or bait consists of a primer or bait pair, and other primers. Alternatively, amplify and / or enrich and / or sequence targeted variants or regions within the bait. The combination of scores for all variants or regions targeted by the predicted capability. They can be ranked and selected based on the criteria. As will be clear, variants or territories Rather than by region, it is advantageous to select and rank primers or baits using this method. This is possible. This is because the assay output is an integrated analysis of the collective results of multiple variants. Therefore, in some embodiments, the score is high, but it is difficult to duplicate with others. Evaluating a larger number of variants at the expense of some potentially difficult ones. This is because there are cases where this is preferable.
[0115] In one embodiment, the best multiple assay is designed after the top variants have been selected. ru.
[0116] In any embodiment, the patient has cancer, has had cancer, or does not yet have cancer However, it has clonal proliferation that has the potential to transform. In some embodiments The patient has received or is receiving treatment for cancer.
[0117] In any embodiment, the DNA is cell-free DNA, for example, the cell-free DNA is plasma. It is isolated from serum, cerebrospinal fluid, urine, saliva, or feces. In other embodiments, DNA is isolated from serum, cerebrospinal fluid, urine, saliva, or feces. Cells, for example, bone marrow cells, lymph node-derived cells, and in the case of blood cancer, circulating leukocytes, Cells derived from tumor nodes, cells derived from tumor margins, or cells derived from solid tumors by other means. Other samples such as CSF and whole blood are currently being screened for the presence of cancer cells. It can be isolated from the type.
[0118] The fraction of cancer DNA in the DNA test sample was 0.01% or less, 0.005% or less, and 0. It may be 0.002% or less, or 0.001% or less, and in some embodiments, the test sample is DNA with a genome equivalent of less than 25,000, e.g., less than 20,000, less than 10,000 or containing DNA with a genome equivalent of less than 5,000.
[0119] In some embodiments, the number of aliquots and the maximum number of molecules per aliquot are single If the number of input molecules in one aliquot is sufficiently low and a single variant molecule exists The input will generate a signal that is significantly different from the background. It is adjusted based on the total number of offspring and the estimated background error rate.
[0120] In any embodiment, for each aliquot of each sequence variation, step (a) The read depth is at least 10,000, at least 25,000, at least 50, It may be 000, or at least 100,000, or at least 500,000. In one embodiment, the method involves measuring the amount of DNA in the test sample before step (a). It may include the following.
[0121] In any embodiment, the sequence of the target region is obtained by PCR prior to step (a), and This is achieved by hybridization to nucleic acid probes, or by hybridization to one side of the target DNA molecule. There is a universal sequence, and at least one, and optionally further nested primers, Using a one-sided PCR approach, which is used to target the other side of the child, the test sample It can be concentrated from the material. Linked Target Capture, molecular inversion probe Other methods known to those skilled in the art, such as ATOM Seq, may also be used.
[0122] As described above, this method can be performed using a threshold-based approach. In the application method, any target region in any aliquot is arranged in i) step b ii) In step b If the calculated probability exceeds the specificity threshold (e.g., 95%, 99%, 99.9%) iii) If the frequency exceeds the threshold, or iv) (i) If cancer DNA is present , and (ii) if cancer DNA is not present, observe the estimated value in (b) in the sample. By calculating the likelihood ratio for each variant in each aliquot between likelihoods, It can be determined that it contains at least one variant molecule. The target region contains two variants. In some embodiments, the region is defined as the region where the signals of both variants are located within the same sequence. In that case, it can be determined that it contains at least one mutant molecule.
[0123] In some embodiments, cancer DNA is reduced in step (c) of the method: i) less In any aliquot determined to contain at least one mutant molecule, the target region exceeds a threshold number. If a region exists, and / or ii) at least one mutant molecule having At least two or at least three aliquots determined to include one target region If present, this can be determined. In these embodiments, the number of thresholds for the target region is i) small Two or more mutant molecules in any aliquot that are determined to contain at least one mutant molecule (for example, ii) may be a target region of 3, 4, 5, or 10 or more, or ii) high signal background The number of win events was 5%, 0.5%, 0.1%, or 0.01%, or 0.001%. If you expect it to occur with a probability of less than 1%, determine the threshold by considering all target areas and ants. By combining the estimated proportion of high-signal background events for the court, Can this be determined (for example, if there are 4 aliquots and 48 target regions, Furthermore, for specific combinations of target regions and variants within these regions, 0.01% It is estimated that there is a probability of obtaining four or more high-signal events across all aliquots. (The threshold will be set, and then a threshold of 4 will be set), or iii) Target region or barrier It can be a score rather than a fixed number of points, and the threshold score is either 2 or 3. Positive target regions or variants are determined according to their proportion of high-signal background events. This contributes to different scores. In one embodiment, high signal background events are never used. Any variant or class of a variant that is not present is given a score of 1, and the remaining variants The class of a variant is based on the proportion of high-signal background events. They are then divided into one or more groups, and the lower the score, the lower the group. For example, there could be two groups. 50% of the variant or variant class with the lowest proportion of high-signal events Receiving a score of 0.75, the highest percentage of those who are positive is always 50%, which is 0.5 Receive the score.
[0124] In any embodiment, the threshold frequency of step (b) is a variable for array variations. Binomial, overdispersed binomial, beta, normal, exponential, or gamma probability distributions of the ground error rate. The frequency can be determined using a model, and if no variant molecule is present, the variant specificity Depending on the desired predefined per unit, the percentages are 5%, 2%, 1%, 0.1%, and 0.01%. Alternatively, the selection is made such that a signal is observed with a probability of less than 0.001%.
[0125] Further details, alternative steps, and embodiments of the present invention are described below.
[0126] Sequence variations related to cancer in patients This method involves analyzing multiple sequence variations in a sample that are associated with the patient's cancer. Therefore, such sequence variations are thought to exist in the cancer cells of patients. Each individual sequence variation can be a driver mutation or a passenger mutation, and Column variations may be clonal or non-clonal. Sequence used in this method The variation is thought to exist only in cancer cells, not in normal cells, in patients. In that sense, it is related to cancer. The set of mutations that define a patient's cancer is different for each patient. While it is patient-specific in the sense that it fluctuates, some mutations (e.g., KRAS) It can occur in several patients and / or several different types of cancer. The location of the passenger mutation is difficult to predict in advance (some hot (There may be spots), but the position of the sequence variations differs from patient to patient, therefore this method The sequence variations to be analyzed can be identified on a patient-to-patient basis. In some embodiments, Sequence variation is higher in samples with a higher cancer fraction, such as bone marrow aspirates and tissue biopsy samples. It can be identified from samples or isolated circulating cancer cells. For example, sequence variations The tumor can be obtained from bone marrow aspirate, tumor tissue biopsy, or surgical excision, or from circulating tumor cells (CTCs). Other cells that are not part of the tumor tissue but are not circulating, such as those found in urine or stool samples. Alternatively, when the sample from which DNA is extracted is likely to have a high ctDNA level, cancer The DNA isolated from cell-free DNA collected from the patient before treatment is sequenced. This may be determined by the following. In some embodiments, clonality is determined by To achieve this, multiple sample types or multiple regions from the same sample can be sequenced. The sequence determination step is, as described above, whole genome sequencing, exome sequencing, or targeting Sequencing (for example, by sequencing a panel of oncogenes or mutation hotspots) This can be done by (for example, by determining the arrangement of the panels in the array, which are the spots). To make it clearer, the patient may be a cancer patient, and the patient may have received cancer treatment. In some cases, they may be experiencing or seeking treatment. In other words, array variance Aspects are present in samples where they are present at relatively high levels, for example, when any cancer treatment is initiated. This can be identified in samples collected before the procedure.
[0127] Depending on how the method is performed, sequence variations are analyzed in the test samples. This can be identified beforehand or simultaneously with the analysis of the test sample. Therefore, some of the results of this method One embodiment of this method uses a “pre-specified” sequence variation, Sequence variations, for example, are related to the patient's cancer before or during treatment. This is a previously identified sequence variation. In other embodiments, the sequence variation This is not predetermined, and instead, sequence variations are derived from the sequence of the test sample. The dots were obtained from control samples (for example, positive and negative control samples as described below). This can be identified by comparing it with sequence reads.
[0128] The sequence variations analyzed by this method are independently single-nucleotide variations. These can be indels, transpositions, or rearrangements. Generally, sequence variations can affect cancer cells. DNA isolated from tissue samples (e.g., biopsy, surgical excision, or fine / coarse needle aspiration) Sequencing of whole genotypes, or sequencing of cell-free DNA from patients (for example, whole genotypes) Identifying by (membrane sequencing, exome sequencing, or targeted sequencing approaches) This allows for the arrangement of multiple regions. For example, in some embodiments, an array barrier The list of entries is based on sequencing at least 50kb of cancer DNA, which is used to study genomes. Cancer DNA can be obtained through targeted sequencing of large regions or whole-genome sequencing. Tumor tissue (e.g., biopsy) or a sample that is expected to contain high levels of cancer DNA. It is obtained from one of the following (such as a plasma DNA sample before treatment). In some embodiments, Only the DNA is sequenced. In an alternative embodiment, whole blood, buffy coat, and adjacent to the tumor are used. In contact with clearly normal tissue, or a cheek swab, etc., cancer DNA and tissue expected to be normal. Both DNAs can be sequenced. The variant evaluates cancer and normal DNA. By means of evaluating only cancer DNA, or by arbitrarily evaluating other features known in the art. In addition to using the variant allele fraction, either by using it or by using the variant allele fraction, It can be classified as a germline somatic characteristic.
[0129] In some cases, analysis of early cancer DNA samples is used to list candidate sequence variations. This can result in some candidate sequence variations being pre-identified sequence variations. It is excluded in order to generate a list of. In some embodiments, this method is used as an experiment. The sample is somatic from the patient being evaluated (for example, by sequencing a biopsy). This includes obtaining a list of possible candidate variants and then prioritizing the variants. It is possible. In these embodiments, prioritization is, for example, not sequence determination artifacts. In contrast, the probability of it being an actual variant, the probability of it being a somatic genetic abnormality, and the probability of it being a clonal mutation. Probability, error rate estimate, fit estimate multiplexed with other variants, and / or Mapping possibilities for riant and surrounding regions, increase or amplification of episomes or In each type of cancer, such as in the presence of double microchromosomes or ring chromosome break fusion regions, This can be based on an estimated number of copies of the variant. Candidate variations can be prioritized. In addition, one or more of the candidate sequence variations may be excluded, and the candidate sequence variations Only a subset of the candidates may be selected for subsequent analysis. For example, candidate sequence variations After the variation is identified, the target region containing those sequence variations is used in normal cells (variations). DNA from the cheek coat, leukocytes, cheek swab, or adjacent tissue can be sequenced. This sequencing uses the same approach as that used to sequence tumor DNA. The procedure may be performed using or sequencing to detect variants identified in tumor DNA. This can be done using assays designed to identify these normal cells. Any variant identified is likely to be a germline polymorphism or clonal hematopoietic, Some candidates may be excluded, and the remaining sequence variations may be preferred. For example, several actual In the administration method, at least some of the target regions in the DNA of patient-derived leukocytes are distributed. The method may further include determining the number of candidate genetic variations. This includes comparing the results with the genetic variations called using leukocyte DNA. It is possible. If variations are identified in both samples, it is a pre-determined distribution. This embodiment can be excluded as it is a column variation. This embodiment is excluded from subsequent analysis. This is possible in clonal hematopoiesis (CHIP) with uncertain potential. Possible variations (generally, Funari et al, Blood 2016) 128:3176 and Heuser et al,Dtsch.Arztebl.Int. (See .2016 113:317-322) and identify germline variants. The present invention provides a method for applying candidate genetic variations to a tumor. In an alternative embodiment, the method applies candidate genetic variations to a tumor. By comparing the called genetic variation with adjacent, clearly normal tissue, This may include the following. If variations are identified in both samples, they are identified in advance. This embodiment can be excluded as it is a sequence variation. Variations that may potentially be caused by cancer field effects can be excluded. The present invention provides a method for identifying germline variants.
[0130] Therefore, in any embodiment, the method uses one or more positive and / or negative control samples. This may include sequencing (which can be performed before or simultaneously with the test sample). As expected, this assay requires that the early cancer DNA sample, control sample, and test sample be the same individual It is "individualized" in the sense that it is obtained from the body. Positive and negative control samples are limited to these. It is not determined, but the tumor is found in biopsies or surgical specimens from either the primary tumor or metastases. DNA, buffy coat DNA, cheek swab DNA, whole blood DNA, normal tissue (e.g., adjacent tissue) The DNA is isolated from the tissue, or a reference DNA. In these embodiments, tumor D Sequence variations not detected in NA may be excluded, buffy coat, cheek swab, adjacent Sequence variations detected in contact with normal blood or whole blood are excluded. Optional implementation In terms of morphology, sequence variations include clonality, mapping potential, estimated error rate, and other factors. When designing a multiplier PCR or hybrid capture panel based on the distance from the selected variant, Compatibility with other variants, predictive ability to determine sequences, and within the region of copy number increase or amplification. The presence of either cis or trans alleles, which may be used to enrich the mutant allele. Prioritized based on one or more factors that may include proximity of any germline variant in To obtain. This would allow for the enrichment of sequence variations that are close to germline variants. The method is such that at least one of the primers is specific to the germline altered chain, Perform allele-specific PCR on alleles that are on the same strand (cis), or on the wild-type strand. To remove a variant when it is on the opposite chain (or trans), for example, restricting yeast This includes targeting germline changes using elements, Cas9, or similar methods. Other implementations In this case, sequence variations are detected by allele-specific PCR, cold-PCR, or other methods. Based on its suitability for variant enrichment methods such as other methods known to the public, It is possible to be ahead of the curve.
[0131] As may be apparent, the sequence variations analyzed in the method are in the method The sequence variations analyzed are "customized" for each patient, so that differences between patients are not apparent. It is possible. Therefore, in many embodiments, the method is obtained from a DNA sample from a first patient. The first set of sequence variations, sequence variations from DNA samples from the second patient. The second set of DNA samples from the third patient, and the third set of sequence variations from the DNA samples from the third patient. This may include identifying such things.
[0132] Aliquot-based sequencing Aliquote-based sequencing methods can be carried out in various different ways. Morphologically, target regions with sequence variations are those with pre-identified sequence variations. A target fragment containing this material is amplified directly from the sample by PCR, in an "amplicon-based" manner. The sequence can be determined using the approach. In some embodiments, the test sample is first, for example For example, adapter ligation and PCR targeting the ligated adapter. Pre-amplification can be performed by carrying out the following. In these embodiments, the sequencing adapter is It can be added during amplification or ligated after amplification. In other embodiments, it can be added beforehand. Target regions with identified sequence variations are ligated to the sample by the adapter. Before amplification, using a primer in which the fragment containing the target region hybridizes to the adapter. "Targeted enrichment-based" apps, enriched by hybridization to nucleic acid probes. Roach can be used for sequencing. A reaction may be performed, or an adapter having multiple barcodes may be used on DNA. The molecules are gated, allowing for effective separation into separate barcode groups or "aliquots." It is possible to make it possible. Therefore, the sequence of the target region is obtained by PCR or nuclear The sample can be concentrated by hybridization with an acid probe. A shrinkage method may be used. In other embodiments, physical duplication or the use of molecular barcodes may be used. Any other method involving either, for example, Molecule Inversion Prob es(MIP) or Anchored Multiplex PCR (AMP) is used. It is possible. Some of the principles of the amplicon-based method are described below. Similar concepts are It can be applied to targeted enrichment approaches. In some embodiments, variants The columns are COLD-PCR, allele-specific PCR targeting variants, and adjacent cells. Allele-specific PCR targeting germline changes, and utilization of adjacent germline changes A targeted step is obtained by digesting a bio-type sequence or by any other method known to those skilled in the art. It can be concentrated during the process.
[0133] In embodiments that use pre-specified sequence variations, the pre-specified sequence variations After the ation is identified, multiple primer pairs are obtained, and each primer pair is specially selected beforehand. Amplify a target region having one or more of the defined sequence variations. In the embodiment, the length of each amplicon is independently in the range of 50 bp to 500 bp, for example It can be 70-150bp, but in some implementations, longer or shorter amplifiers are used. Recon may be used. In some embodiments, some of the variants are rearranged. In these embodiments, the primers are one primer on the 3' side and 5' side of the rearrangement. Designed using one primer on each side, the rearranged sequence designs a primer pair. Used for this purpose, the primers are specifically designed to amplify the rearranged sequence. After the primer pairs are obtained, the method is to use each of the same portion of the sample (i.e., the same sample) A multiplex PCR reaction (e.g., 2, 3, 4) containing different aliquots. Up to 10 multiplex PCR reactions, such as 5, 6, 7, 8, 9, or 10 multiplex PCR reactions. This step may include setting the reaction. In this step, all reactions use the same primer and Multiple PCR reactions can be identical to each other in that they involve different parts of the same sample. In this method, the number of aliquots and the maximum number of molecules per aliquot are determined in a single aliquot. If the number of input molecules is sufficiently low and a single variant molecule is present, the background The total number of input molecules and the estimated signal will generate a signal that is significantly different from the original. It can be adjusted based on a constant background error rate. As will be clear, each multiple P CR should include a compatible primer, and the compatible primer is such that the reaction is directed to the primer. When subjected to appropriate thermal circulation conditions using a suitable mold, the primer dimer and intended Responding to PCR primer pairs while minimizing the generation of non-specific or non-absent PCR products. It is designed to specifically amplify the region where the amplicon is to be generated. Typically, While not always the case, each primer pair is used in multiplex PCR reactions to target a single region. Amplify the region. Implement a program for multiplex PCR and designing suitable primers. The conditions for this are well known (e.g., Sint et al, Methods Ecol Evol.2012 3:898-90 and Shen et al BMC Bioi See Nformatics 2010 11:143. Suitable primer pair These are several primer pairs specifically designed for designing primer pairs for multiplex PCR. It can be designed using one of several different programs. For example, a primer pair is Y amada et al.(Nucleic Acids Res.2006 34:W 665-9), Lee et al. (Appl.Bioinformatics 20 06 5:99-109), Vallone et al. s.2004 37:226-31), Rachlin et al.BMC Geno mics.2005 6:102, or Gorelenkov et al. (Biot It can be designed using the method described in echniques.2001 31:1326-30. In some embodiments, the method uses at least five pairs of compatible primers, for example, a minimum of At least 10 pairs, at least 50 pairs, at least 100 pairs, at least 1000 pairs, or a small number At the very least, 5000 pairs of compatible primers can be used. The amplified amplicon is optional. It may be of a preferred length, and the length may vary. In some embodiments, the sequence variation The choice may be prioritized based on the potential suitability of the primer design in multiplex PCR. ru.
[0134] Next, the amplicon generated by thermal circulation of the reaction, or its amplified product (e.g.) For example, the amplicon hybridizes to the 5' tail in the primer, a universal primer (When re-amplified by a meter) the sequence is determined and sequence reads are generated. Various A The recoated PCR reaction should produce a replicated amplicon, and the "replicated" amplicon This is an amplicon amplified by the same primer in the aliquot. Plicons generally have the same sequence (PCR error, genetic variation in the sample) This excludes variations corresponding to the , and any variations in PCR primers. Ku).
[0135] When sequencing amplicons, the amplicons derived from each different multiplex PCR reaction are They can be sequenced separately from each other, or the amplicons are barcoded with aliquot identifiers. It can be denatured and then pooled before sequencing. In some embodiments, multiplex PCR The primer in the reaction may have a 5' tail containing an aliquot identifier, and therefore, After the PCR reaction is complete, the 5' tail sequence of the primer is present in the amplicon. In this embodiment, a primer having a 5' tail containing an aliquot identifier is not used. Multiple PCR reactions can be performed. In these embodiments, the PCR product is alicol For the second round of amplification, use PCR primers with a 5' tail containing a T identifier. In this case, it can be barcoded with an aliquot identifier. The adapter array can also be used to create the product. It can be gated. In any case, the amplicon is a specific sequence determination platform. A primer with a 5' tail that provides compatibility with the sequence is used for amplification before sequencing. It may be possible. In certain embodiments, in addition to the aliquot identifier, used in this step One or more of the primers may additionally contain a sample identifier. In the application method, one or both of the primers may contain a barcode, which may be applied independently or in combination. Either method can be used to identify both the sample and the aliquot. If the lymer has a sample identifier, products from different samples are pooled before sequencing. This is possible. In some embodiments, the target-specific primer extends from 5' to 3'. A universal "tagging" array, an arbitrary aliquot barcode array, then designed to target the desired object. It contains the sequence. The primers used to further amplify the initial product are specific A 5' tail providing compatibility with the sequencing platform, a sample barcode, and optionally Aliquot barcodes or barcodes that identify both the sample and the aliquot, and target It binds to some or all of the reverse complement of the tagging sequence present on the specific primer. It may contain sequences that can perform. Typically, the forward and reverse primers are , will have different tagging sequences. As will be obvious, for the amplification step The primers used are those used in any next-generation sequencing platform where primer extension is used. Forms, for example, Illumina's reversible terminator method, Roche's pyrosy Quensing method (454), Life Technologies by ligation Sequence determination (SOLiD platform), Life Technologies' Io n Torrent platform, or Pacific Biosciences Fluorescent base cleavage, and any other platform, e.g., Oxford Nanop It may be suitable for use in ore. An example of such a method is found in the following reference: Marguli es et al(Nature 2005 437:376-80), Ronaghi et al(Analytical Biochemistry 1996 242: 84-9), Shendure(Science 2005 309:1728), Im elfort et al(Brief Bioinform.2009 10:609 -18), Fox et al(Methods Mol Biol.2009;553 :79-108), Appleby et al(Methods Mol Biol. 2009;513:19-39), English(PLoS One.2012 7: e47768), and Morozova (Genomics.2008 92:255- As described in 64), these include all starting products, reagents, and in each step A general description of the method and specific steps of the method, including the final product, are provided by reference. To be absorbed.
[0136] In an alternative embodiment, aliquot-based sequencing is performed on a panel of mutation hotspots. A panel of oncogenes can be targeted. Alternatively, the sequencing step can be performed on an exo By sequencing or whole genome sequencing, or by at least 1, at least 5, or a small number of genomes. At the very least, this can be done by sequencing a 10MB genome to a suitable depth. In these embodiments, the sequence variations do not need to be "pre-specified". Rather Sequence variations occur in the same assay in which the test sample is sequenced, i.e., Comparison of data with a control performed in the same assay (e.g., the same sequencing run). This can be identified by the following: Once sequence variations are identified using a control sample... These sequence variations can be analyzed in the test sample.
[0137] The sequencing step can be performed using any convenient next-generation sequencing method per reaction. at least 100,000, at least 500,000, at least 1M, at least Array reads of 10M, at least 100M, at least 1B, or at least 10B It is possible. In some cases, the reeds can be paired-end reeds.
[0138] Processing sequences, predicting variant molecules, and determining the presence of cancer DNA. to Next, the sequence reads are computationally processed. The initial processing step is barcode (sample identification). Identifying child or aliquot identifier sequences, and removing low-quality or adapter sequences. This may include trimming the reads. In addition, the dataset must be of acceptable quality. To ensure this, quality evaluation metrics can be implemented.
[0139] After the sequence reads have undergone initial processing, they determine which reads correspond to the target region. These sequences can be analyzed to determine if they are identical or nearly identical to the sequence of the target region. Therefore, it can be identified. It is identical or nearly identical to the target area so that it will be recognized. The sequence reads are analyzed to determine if there are potential variations in the target sequence. This is possible. The sequence is aligned with a reference sequence, for example, a genome sequence, in this method. It can be found, or it can be matched to a database of expected sequences.
[0140] After the sequence reads have been processed, the method is as follows for each aliquot and each sequence variation. , counting the number of sequence reads that have sequence variations, and the total number of sequence reads This may include counting numbers. A method for counting leads is, for example, For shew et al(Sci.Transl.Med.2012 4:136ra68 ), Gale et al (PLoS One 2018 13:e0194630), and Weaver et al (Nat. Genet. 2014 46:837-843) It can be adapted from what is described by ). Similar results can be obtained by adopting a molecular index. This can be obtained using the following approaches. These methods yield the total number of sequenced molecules. The number of variant molecules and their corresponding numbers can be estimated using an index. The molecular identifier sequence is a sequence of other features of the fragments that distinguish them from each other (e.g., a break point that defines a break). It can be used in conjunction with the terminal sequence of the fragment. The molecular identifier sequence is (Casbon Nucl. This is described in Acids Res. 2011, 22 e81).
[0141] As shown in Figure 11, the number of sequence reads with variations is counted, and the sequence After counting the total number of reads, the fractions in the original sample before amplification that had sequence variations were counted. The estimated number of offspring can be determined for each aliquot in each target region. Alternatively, For each aliquot of each target region, at least one segment having sequence variations The probability of a child existing can be calculated. The latter, for example, is calculated by considering all non-zero numbers in the numerator ( That is, it can be derived by summing the individual probabilities of all positive integers. In these embodiments, the estimates may be stochastic estimates, and the estimates may not be point estimates. This means that it is a probability distribution. This step is the variant in aliquot. This can be done by assigning a probability to each possible number of children, which is done via a probability density function. This can be done, and an example is shown in Figure 12. In these embodiments, each aliquot and mark Regarding the region, the number of molecules having sequence variations, or the number of molecules having sequence variations The probability that at least one molecule exists is (i) a sequence with sequence variation (ii) the number of alphanumeric characters, (ii) the total number of sequence reads, (iii) the number of molecules entered into each aliquot, and (iv) Can be calculated using the estimated background error rate of sequence variations. In these embodiments, the sequence of the target region is defined by the number of sequence reads (e.g., at least 10). There are 0,000 reads, but this number varies depending on the number of aliquots to be sequenced. These are represented by (obtained), and some of those reads may contain sequence variations. These reads can be counted to provide input values (i) and (ii). The input value (iii) is the amount of DNA in the DNA sample to be measured before starting the method. This can be calculated by, for example, the total amount of DNA, the total amount of double-stranded DNA, The total amount of double-stranded and single-stranded DNA, the total amount of DNA within a specific size range, or amplicons. DN can be amplified using primers with specific parameters such as size. This can be done by measuring the total amount of A. This step is performed using digital PCR. By qPCR, fluorescence analysis, electrophoresis, or various kits or other strategies This can be done using either method. Estimated background amount for each sequence variation The Lar rate, i.e., the input value (iv), is the result of a prior sequence determination reaction, e.g., sequence variation. Samples that are known not to have cancer, or that are not known to have cancer, therefore This procedure is performed on samples from individuals that are not expected to have a large number of somatic variants. This can be determined from the column determination response. Specifically, the background of each variation The error rate is calculated based on the same execution, past executions, or using past executions. Adjust the use of the control base (or base not known to contain variants) selected by [the relevant method]. DNs that are not expected to contain somatic variation in similar variants evaluated by any of the following methods It can be estimated through sequencing of similar variants in A, and the variant is salt The types of group changes, base changes (rearrangement / conversion), and trinucleotide contexts, penta Nucleotide context, position in the amplicon referencing the primer, insertion status The size, type and number of inserted bases, size of deletions, type and number of deleted bases, They are considered similar based on a rearrangement class, for example, a feature that may include tandem overlap. The hypothetical error model is shown as a frequency distribution in Figure 13A, or as a mixture model in Figure 13B. These examples show multiple trials that are not known to contain somatic variants. A sample (e.g., several hundred samples) is sequenced and has a specific type of sequence variation. The fraction of sequence reads can be calculated for each sample. Variant sequence reads are the main ones. This includes pre-PCR events such as errors, base miscalls, and DNA damage that occur during PCR. This is caused by (for example, when A base pairs with another base, and in the sequence read, the base pairs from G to T) The oxidation of guanine to 8-oxoguanine appears as a reaction. These fractions This can be plotted as a frequency distribution, which in turn allows us to observe the results in sequence reads. To calculate the probability that the sequence variations are actually genetic variations It can be used for this purpose.
[0142] Next, the presence or absence of cancer DNA in the sample is determined from each target from each aliquot of the original sample. This can be determined using estimates (or probabilities) of variant molecules in the region. In some cases, the data is also used to estimate the overall cancer DNA fraction in the sample. This can be used. This estimate is the most likely amount of cancer DNA in the test sample or It could be within the range of possible amounts of DNA, and the variant read fraction or burr in the original sample. Based on the estimated values of the ant molecule, for example, the mean or central variant allele fraction, It can be estimated by maximum likelihood or Bayesian posterior method.
[0143] In one embodiment, the presence or absence of cancer DNA in the sample yields the same result regardless of the cancer DNA. Considering the likelihood that cancer DNA may be present in a sample that does not contain A, The likelihood of observing the result can be determined through the likelihood ratio by comparing the likelihoods of observing the result. If the likelihood is higher that the same data can be generated using a sample that does not contain this cancer DNA The sample may not contain any cancer DNA. First likelihood (cancer DNA is present) The likelihood of (i) is calculated above for each aliquot of each target region, and the sequence variation The estimated number of molecules having a probability or likelihood, and optionally, (ii) the estimated cancer DNA in the sample. It can be calculated using fractions. The second likelihood (likelihood of the null hypothesis) is (i) calculated above. (ii) a probabilistic estimate or probability, and (ii) a "high signal background event" This is an event that cannot be explained by a simple model of the background error rate for each hit. This can be calculated using the estimated rate of high-signal background events in the sample. After calculating the likelihood of existence and the likelihood of the null hypothesis, they are compared to obtain the likelihood ratio. In order, the likelihood ratio can be compared to a threshold. In some embodiments, The likelihood ratio is determined for each aliquot of each target region. The individual likelihood ratios are then used for the sample. Combined with the cumulative likelihood ratio score across all regions and aliquots. The degree ratio indicates that the DNA sample contains cancer DNA. Alternatively, the likelihood ratio can be used directly or The sample contains cancer DNA, either by comparing it to a reference distribution calculated using a control sample. This can be interpreted as the probability of something happening.
[0144] Specifically, as described above, the models in Figures 13A and 13B have at least three types of Errors include: errors that occur during PCR, base miscalls during sequencing, and DNA damage. Pre-PCR events such as injuries. Pre-PCR errors are considered "high signal" in the sense that they are rare. It does occur (though not in all samples), and when it does, it is present in the original sample. The fraction of variant reads that matches the variant molecule is much higher than other errors. This results in, i.e., mimicking the appearance of a true positive ctDNA variant. In some cases, Errors that occur in the first one, two, or three cycles of PCR also cause high signal events. It can result in such errors. The rate of such errors can be determined using various different methods. Yes, it is possible. In some cases, an error distribution or error probability distribution can be used. In this embodiment, the error distorts the distribution as shown in Figures 13A and 13B. Analysis of error distribution allows high-signal events to be identified as distinct events. For example, in some cases, the event is defined by a threshold (e.g., average), as shown in Figure 13A. It can be identified using an event that is one, two, or three standard deviations from the value or median. Yes, it is possible. Such thresholds can vary between variations, but generally, These are identified as having a frequency exceeding the defined threshold as shown in Figure 13A. These high-signal events can be modeled separately, and each sequence variation It can be used to determine the proportion of high-signal background events.
[0145] In another embodiment, the determination of whether a test sample contains cancer DNA is made by (i) each target region Estimates or probabilities of variant molecules in each aliquot, high signal background A mixed model that incorporates the estimated proportion of elephants and, optionally, the prior estimate of the cancer DNA fraction in the test sample. It is calculated using (Figure 13B). The output of the mixed model is compared to the threshold. It is possible to do so, and an output above the threshold indicates that the test sample contains cancer DNA. Such a threshold for that method is multiple samples that are not known to contain cancer DNA. We analyze the results, determine the distribution of the results, and then determine the probability of false positives being less than 0.01% and less than 0.1%. The threshold is such that the rate is expected to be less than 0.5%, less than 1%, or less than 5%. This can be determined by setting [something].
[0146] In some embodiments, before calculating the likelihood of the presence of cancer DNA in the sample, or Before evaluating a sample with a mixed model of cancer DNA, or to indicate the presence of cancer DNA Before determining whether a sufficient target area, variant, and / or aliquot exceeds the threshold Based on the estimated cancer DNA fraction, it was identified in a statistically unlikely number of aliquots. The probabilistic estimates or probabilities of the sequence variations are excluded. For example, relatively high even Except for the original aliquot, for most aliquots of most variations The estimate or probability is relatively low, indicating that it is unlikely to contain variant DNA. In this case, one sequence variation is present in all or almost all aliquots with a relatively high probability. It is statistically unlikely to exist. As a further example, an implementation with four aliquots Morphologically, evidence for most variants is that they contain zero or one variant DNA. If you support any of the aliquots, the evidence for all four aliquots is barrier Any variant that supports the presence of nucleotide DNA is likely to be an outlier. Outliers ("noise-based," or, for example, non-cancer-specific changes originating from CHIP) (which may be caused by) can be identified and excluded from the calculation. In another example, each Alico The number of test DNA molecules attached to the test and all variants (or subsets) are used. Using the estimated tumor fraction calculated, each alias containing at least one cancer molecule The probability of each individual variant in the set can be calculated. Then, the probability of a variant exceeding a threshold can be calculated. The number of licotes, compared to the total number of aliquots, gives a result that suggests there are unlikely to be variants. It is possible to determine whether the number of copies of each variant is This is corrected during the calculation. This concept is shown in Figure 14.
[0147] In this method, high signal is predicted (taking into account cfDNA concentration and estimated ctDNA fraction). Identify and exclude variant-containing regions that would yield more aliquots than would otherwise be produced. This can be done by considering the known cfDNA concentration and putative ctDNA fraction. Using the probability of sampling at least one ctDNA molecule per sample It can be calculated. A variant that is statistically unlikely (e.g., p<0.05) is: It can be excluded. For example, each of the four partitions has a probability of 0.2 that it contains a variant. If present (based on the estimated ctDNA fraction and the number of input molecules), 2 have a high score. The likelihood of representing two partitions can be calculated.
[0148] For clarity, some embodiments of this method are variations in different aliquots. It does not include specifying (or calling) the method. Specifically, the method In some embodiments, the frequency of potential sequence variations in each aliquot is such that the threshold It does not include determining whether it is above or below. Rather, these embodiments are the whole and It depends on the analysis of the data.
[0149] The method can be performed with any type of sample that contains cancer DNA, The method is limited to cases where the cancer DNA fraction is less than 0.01% (i.e., less than 100 ppm). It is most commonly used for analyzing samples containing cancer DNA, which other assays cannot detect. This is because it becomes impossible to distinguish a sample containing cancer DNA from a sample that does not contain cancer DNA. For example, several In the embodiment, the method is used in concentrations of 0.0001% (1 ppm) to 0.001% (10 ppm). It can be used to detect cancer DNA in samples containing cancer DNA, and the sample is 25 DNA with a genome equivalent of less than 1,000 (e.g., 100-10,000, 500-5,000) It contains (or 2,000 to 20,000 genome equivalents of DNA), but these numbers can vary. Furthermore, in order to obtain statistically significant results, each aliquot of each target region may be used as desired. , at least 5,000, at least 10,000, at least 20,000, or less Sequence determination is possible up to a read depth of 100,000.
[0150] Estimating the amount of cancer DNA In some embodiments, the amount of cancer DNA is measured as the total number of variant-containing molecules. In another embodiment, the amount of cancer DNA is the putative variant allele fraction (VAF). ) can be measured as the mean or median VAF (i.e., analytical The mean or median of all variants may be generated, and in other embodiments, corrected The mean or median VAF can be determined (i.e., for each variant, in advance) The average across variants after subtracting the determined offset or baseline error rate. (Average or median level). In some embodiments, the VAF and c added to the sequencing reaction. The total number of fDNA molecules estimates the total number of variant tumor molecules added to the sequencing reaction. They can be multiplied together as a method for achieving this.
[0151] In other embodiments, information obtained through sequencing of tumor tissue is used to identify single cancer cells. This information can be used to estimate the copy number of each variant within a cell, and this information can be used to estimate the tumor it represents. To determine the number of tumor cells, i.e., "represented cancer cells," the cells detected in the sample were... It can be used in combination with riants and their frequencies.
[0152] In some embodiments, the measured number of variant-containing molecules, or the estimated number of cancer cells, is used in the trial. To estimate the number of molecules per 1 ml of liquid, such as plasma from which DNA has been extracted, It can be combined with the number of liters. In an example of such analysis, Varian per 1 ml of plasma Mean value of 1 molecule, median of variant molecules per 1 ml of plasma, tumor per 1 ml of plasma A series of outputs are measured, such as the median number of cells or the median number of variant molecules per 1 ml of CSF. It can be calculated.
[0153] In some embodiments, this calculation is performed to determine the amount of DNA lost between blood collection and sequencing analysis. This may include a step to correct for the cfDNA extraction efficiency. Alternatively, this may include correcting for library preparation efficiency. For example, plasma 1 When calculating the median of variant molecules per ml, first, it is possible to detect them in the sample. The number of mutant molecules and the volume of plasma from which the cfDNA sample used was extracted. It will be determined. Then, this number will typically be recovered by the extraction chemistry used. A known number of molecules and / or changes in such molecules during sequencing library preparation and analysis The conversion and subsequent sequencing rates are corrected. In some embodiments, known sequences At least one synthetic spike DNA sequence having this is added to the sample before extraction, The sequences are analyzed during sequencing to determine the efficiency of extraction and library preparation, and then... This is then applied to correct the previously described mutant molecule estimates. In certain embodiments, The spike sequence contains molecular barcodes and counts the number of molecules that are successfully read. It can make it possible to do so.
[0154] Estimating the detection limit As will be apparent to those skilled in the art, many factors affect the sensitivity of such methods. Depending on the approach, these factors are added to the library preparation reaction and sequenced. The amount of DNA from the test sample, the number of aliquots, the number of target regions and variants, and the back Ground error rate, and high signal background events for each variant. It may include a proportion.
[0155] In some embodiments, the detection limit is determined each time the sample is analyzed. In the application method, sequencing is used to determine the number of DNA molecules to be evaluated for variants. The amount of DNA from the sample added to the reaction is multiplied by the number of target regions. During the verification study, a wide range of samples with different numbers of molecules evaluated for variants were found to be... The detection limit is tested to be determined empirically. In some settings, in addition, a barrier The test results are divided into classes, and the influence of each class is determined. When a sample is tested, the results are determined. The limiting factors are then the number of variants, the amount of DNA added to each aliquot, and the variants The number of molecules evaluated, or at least one of the classes of variants evaluated. It is estimated based on this.
[0156] By utilizing cancer signatures A series of mutation processes drive somatic mutation in the cancer genome, and each of these processes is particularly It is known in the field to generate distinctive mutation signatures (Alexand rov, Nature 2020 578:94-101). Some of these processes And therefore those signatures are common to many cancers, but others are different. It is specific to certain cancers. It targets a sufficiently large region of the genome, such as the exome or the whole genome. By sequencing, these signatures can be detected in tumor DNA. It is possible. In one embodiment of this method, once tumor DNA from the patient is sequenced, It can be analyzed to determine the signature of the tumor. When the primary origin of the tumor is unknown, the origin of the cancer Signatures may be used to infer the source. For example, SBS7 present within a tumor. The signature (Alexandrov, above) is consistent with the primary tumor, which is melanoma. It is likely.
[0157] In another embodiment, variants identified in a tumor are artifacts, germline To determine the likelihood that it is a cancer-specific somatic change, rather than just one of the CHIPs, Signatures may be used. In such embodiments, multiple potential tumor-specific somatic Variants are identified by sequencing tumor DNA. Tumor type (for example) For example, melanoma has common signatures present in its tumor type (for example, TCN is mainly It is identified as SBS7a, where C>T. Common signature of cancer type and Matching variants are included and prioritized, or used for targeted sequencing. When selecting, ranking, or scoring riants, it is possible that they represent actual physical changes. A score indicating a high degree of symmetry is given, but variants that do not match the main signature are... Whether it is excluded by filtering, or given a lower priority or score. It is one of the two.
[0158] Methods for evaluating cfDNA quality The test sample is cell-free DNA, and before sequencing the cell-free DNA from plasma, A method used to evaluate the amount or ratio of DNA that is of high molecular weight. Cell-free DNA is Typically, it is short (about 160 bp). This can occur when blood samples are handled improperly or during transport. White blood cells can lyse, and upon lysis, they can release high molecular weight DNA, which masks cfDNA. Therefore, a high proportion of long DNA molecules can lead to poor quality tests with a risk of false negatives. It can indicate the amount. The ratio between the number of short DNA molecules and the number of long DNA molecules is determined, and the short ones The available options are 50bp, 60bp, 70bp, 80bp, 90bp, 100bp, 110bp, It can be 120bp, 130bp, 140bp, 150bp, or less than 160bp, and is long These are 320bp, 480bp, 1000bp, or over 2000bp. 1:10 If the DNA is longer than 1:5, 1:4, 1:3, or 1:2, the sample will be treated after blood collection. Potentially containing high levels of long DNA molecules, which could be evidence of leukocyte DNA release. A flag is set regarding this.
[0159] The ratio is determined by electrophoresis such as agarose gel analysis, or by fragment analyzer or tape analysis. A method of measurement using commercially available systems such as stations. The ratio is PCR-based. A method measured using roach. Examples include targeting both long and short regions of the genome. This includes using digital PCR or qPCR with primers and probes. Either a long region or a single short region can be targeted, or the assay can be varied Multiple markers of different sizes or multiple markers of one size and multiple markers of another size can be used. It can be modified. The advantage of such a method is that some regions of the genome can be modified by copy number changes. This includes the ability to compensate when affected. Alternatively, the assay targets repeating sequences. This allows for targeting of short regions of the repetitive sequence and long regions of the repetitive sequence. An advantage of such embodiments is that less test DNA is needed to measure the ratio. In another embodiment, two or more primer pairs targeting short regions of the genome are used. Although the two regions are located on the same chromosome, they are larger than 320 bp, larger than 480 bp, larger than 1000 bp, and They are separated by just over 2000 bp. In a replication PCR reaction, both regions are amplified in the same reaction. The number of times, the number of times in which only one region is amplified in the reaction, and the number of times in which neither region is amplified. To determine this, typically, dilute so that there is less than one copy of genome per reaction. This is done on the test DNA. The frequencies of these three events are used to analyze long molecules and The number of short molecules can be estimated. In another embodiment, next-generation sequencing is used. Obtain. In one embodiment, the standard library is ligated on the sequencer adapter. It is produced from cfDNA by arbitrarily amplifying DNA. In an alternative example, One or more plastics target one or more repeat regions to amplify cfDNA before column determination. Imagers are used. Then, sequencing reads are aligned to the genome, and the molecular size is determined. This is determined by identifying the start and end of each sequencing read. Then, the sequence The sequence determination reads are divided into groups based on the length of the determined reads, and then the ratio is determined. This allows us to obtain the ratio between short and long molecules. In such a setting, PCR and Both next-generation sequencing methods typically have a bias towards shorter DNA molecules. Therefore, it may be important to use a correction factor. At least one side of the cfDNA molecule An alternative method for ligating an adapter, as well as one or more targeting primers and adapters PCR, which uses primers that target the cfDNA, followed by NGS, measures the length of cfDNA. It can be used to obtain a constant value. In some embodiments, the test sample is cell-free. It is DNA, and before generating a sequencing library, we enrich the shorter cfDNA molecules. To increase the fraction of ctDNA, size selection is used, and this enrichment is done by beads or This can be done using size selection on a gel, with shorter molecules being 160bp or 150bp. These are those with a length of less than 140 bp.
[0160] usefulness If a DNA sample from a patient contains cancer DNA, the patient may, for example, undergo minimal residual disease testing. It may contain cancer-related cells resulting from early recurrence or metastasis. ctDNA has a half-life of approximately 1 hour. Because it has a period, it is a particularly strong biomarker in this setting, and the tumor is completely removed. In that case, any remaining ctDNA should be rapidly removed.
[0161] In some cases, cell-free DNA collected from patients after treatment is used to study minimal residual disease. When testing for abnormalities, the tumor may have a high level of ctDNA sufficient for accurate detection of minimal residual disease. It may be important to first confirm whether it releases. In one embodiment, a cell-free DNA sample It is collected and tested before treatment for curative purposes, and does not have detectable ctDNA before treatment. If the probability of any patient or sample containing tumor DNA before treatment falls below a certain threshold It releases too little ctDNA for accurate detection of minimal residual disease, therefore further analysis is required. It may be excluded. In an alternative embodiment, the patient's pre-treatment ctDNA is 0.01% VAF. If it is estimated to be below a threshold such as 0.005%VAF or 0.001%VAF and may be excluded from further analysis. In another embodiment, the pre-treatment ctDNA level is set Estimated amount of ctDNA released by a defined volume of tumor, therefore tumor ctDNA To provide a standardized measure of release, pre-treatment tumor volume and the amount of tumor volume assessed by imaging are used. They correlate. In patients where this standardized measurement falls below a set threshold, for example, 1 cm 3 of The tumor is expected to release ctDNA at levels below the assay's predetermined detection limit. The combination may be excluded. Alternatively, changes in ctDNA levels after treatment may predict changes in tumor volume. Measure and determine if it corresponds to the complete removal of the tumor, or if it remains constant even if residual lesions remain. This estimate can be combined with other factors to determine whether it exists.
[0162] Patients providing test samples may have cancer and have a history of it (for example, at least two weeks ago, or less). (At least 3 months ago, at least 6 months ago, at least 1 year ago) You have received treatment for cancer In some cases, the condition may be in complete remission and / or have the potential to transform. It has lone growths (e.g., neoplastic growths such as nodules, polyps, and cysts or lumps) obtain.
[0163] Similarly, the source of cancer DNA in a sample can vary. For example, cancer DNA can be malignant. MR occurs as a result of clonal proliferation, tumor metastasis, incomplete tumor removal, or ineffective treatment. This could be the result of D.
[0164] In some embodiments, the method generates a report indicating whether cancer DNA is present in the sample. This may include providing. In some embodiments, the report may include likelihood ratios, mixed models, The score, or the threshold number of variants, and the above aliquot outputs, or another number representing them. Furthermore, by comparing the likelihood ratio or mixed model results, it is determined whether the sample contains cancer DNA. It may include a threshold that allows for the treatment of residual lesions. In some embodiments, the report is for the treatment of residual lesions. Approved therapies (e.g., FDA approved), such as chemotherapy or immunotherapy, are also available. This information can be enumerated. This information is used for the diagnosis of the disease (e.g., whether the patient has MRD) and / or for medical This can be helpful in the therapist's decision-making regarding treatment.
[0165] In some embodiments, the report may be in electronic form, and the method may remotely transmit the report. For example, refer the case to a doctor or other medical professional to determine the appropriate course of action. This includes assisting in diagnosing the subject or identifying the most appropriate therapy for the subject. For example, to determine whether a subject is susceptible to the effects of therapy, metrics from other patients are used. It can be used together with
[0166] In any embodiment, the report can be forwarded to a “remote location,” and the “remote location” is This refers to locations other than where the sequence is being analyzed. For example, a remote location is another location within the same city. For example, an office, a laboratory, etc., different locations in different cities, different locations in different states, different This could be in a different part of the country, for example. Therefore, if one item is "remote" from another item... When indicated, it means that the two items are in the same room but separate, or at least Located in different rooms or different buildings, at least 1 mile, 10 miles, or less It also means that they are 100 miles apart. To "communicate" information means to represent that information. A suitable communication channel (e.g., private or public network) to use the data as an electrical signal. This refers to sending via (c). "Transferring" an item means physically transporting that item. Whether by or by another means (if possible), place the item in one location. This refers to any means of moving from one location to the next, and at least in the case of data, a medium having data. This includes physically transporting a body or communicating data. Examples of communication media include wireless or Infrared transmission channel, and network to another computer or network device Internet connections, as well as information recorded on websites and email transmissions, etc. Including the internet. In certain embodiments, the report is made by an MD or other qualified medical professional. The sample may be analyzed, and a report based on the results of the sequence analysis will be forwarded to the patient from whom the sample was obtained. obtain.
[0167] In some embodiments, the sample is placed in a first location, for example, a hospital or a doctor's office. In any clinical setting, the sample may be collected from the patient, and the sample may be processed at a second location, for example, where the sample is processed. The above method is carried out to generate the report, which may then be transferred to the laboratory. The “report” described in the book refers to a test that may indicate the presence and / or amount of cancer DNA in the sample. An electronic or tangible document containing report elements that provide results. Once generated, the report As part of clinical judgment, it is the responsibility of a medical professional (e.g., a clinician, a clinical laboratory technician, or It can be interpreted by a physician (such as an oncologist, surgeon, pathologist, or virologist). It can be transferred to another location (which may be the same location as the first location).
[0168] Patients analyzed using this method may have any type of cancer, or have previously had any type of cancer. Patients may be receiving treatment for various types of cancer. For example, patients may be receiving treatment for melanoma, carcinoma, or rheumatoid arthritis. You may have, or may have had, a carcinoma, sarcoma, or glioma. For example, cancer Among many others, including other solid tumors and blood cancers, melanoma, lung cancer (e.g., non-small lung cancer) Cellular lung cancer, breast cancer, head and neck cancer, bladder cancer, Merkel cell carcinoma, cervical cancer, hepatocyte cancer Cancer, stomach cancer, cutaneous squamous cell carcinoma, classical Hodgkin lymphoma, B-cell lymphoma, colorectal cancer It could be cancer, pancreatic cancer, stomach cancer, or breast cancer.
[0169] In some embodiments, the method may be used to guide treatment decisions. In terms of treatment methods, the patient should be treated again, for example, with the same therapy or a second therapy. It can be used to determine if a patient has been previously treated with a first cancer therapy. If a patient is identified as having MRD using this method, the patient will undergo the first cancer therapy. The patient may be treated with the same or a different second cancer therapy. For example, if the patient has previously undergone surgery or immunotherapy. If a patient is being treated with a checkpoint inhibitor and is identified as having MRD, then the patient This may involve further surgery, the same or a different immune checkpoint inhibitor, or another type of treatment. It can be treated by law, and immune checkpoint therapy targets CTLA-4, PD1, PD-L1, and T Administration of IM-3, VISTA, LAG-3, IDO, or KIR checkpoint inhibitors. Other types of therapies include, for example, (a) anthracycline therapy (for example, Dauno (b) By administering mycin, doxorubicin, or mitoxantrone Killing agent therapy (e.g., mechloretan, cyclophosphamide, ifosfamide, melfa) Lan, cisplatin, carboplatin, nitrosourea, dacarbazine and procarbazine (c) Topoisomerase II inhibitor therapy (by administering , or busulfan) (For example, by administering etoposide or teniposide), (d) bleomycin therapy (e) Antimetabolite therapy (e.g., methotrexate, 5-fluorosyl, cytarabine) (f) Vinca (by administering 6-mercaptopurine or 6-thioguanine) Alkiloid treatment (e.g., by administering vincrisen or vinblastine), (g) Steroid therapy (for example, by administering prednisone or dexamethasone) (h) including radiotherapy, etc. Alternative therapies include targeted therapy and untargeted chemotherapy. Targeted therapy, including erroci, can be administered to patients with activating mutations in EGFR. Nib (Tarceva), afatinib (Gilotrif), gefitinib (Ires) sa) or osimertinib (Tagrisso), administered to patients with ALK conjugates. Possible medications include crizotinib (Xalkori), ceritinib (Zykadia), and alectinib. Nib (Alecensa), or Brigatin nib (Alunbrig), ROS1 fusion Crizotinib (Xalkori) and entrectinib (R) may be administered to patients with a condition. XDX-101), Lorlatinib (PF-06463922), Crizotinib (Xalk ori), Entrectinib (RXDX-101), Lorlatinib (PF-064639 22) Lopotrectinib (TPX-0005), DS-6051b, Ceritinib, En Saltinib, cabozantinib, or patients with activating mutations in BRAF The possible medications to be administered are dabrafenib (Tafinlar) or trametinib (Mekini This includes treatment at st). Many other feasible mutations are known. The patient undergoes non-targeted chemotherapy. If a switch to therapy is possible, that therapy could be, for example, platinum-based doublet chemotherapy. (Platinum-based doublet chemotherapy includes cisplatin (CDDP), carboplatin ( Platinum-based drugs selected from CBDCA and nedaplatin (CDGP), as well as One third-generation drug (docetaxel (DTX), paclitaxel (PTX), vinorelvin) (VNR), gemcitabine (GEM), irinotecan (CPT-11), pemetrexate It may include (selected from PEM and tegafurgimeracyloteracil (S1)). (ru).
[0170] In some embodiments, the method may be used to monitor treatment. For example, The method involves analyzing the sample obtained at the first time point using the method, and then analyzing the second time point using the method. Analyzing the sample obtained at that point and comparing the results, that is, if cancer is present in the sample To determine whether DNA is present, or to determine the presence of cancer DNA between the first and second time points. This may include determining whether there has been a change in the quantity or the range of quantities likely to contain cancer DNA. In some embodiments, such changes are determined using point estimates or confidence intervals. A significant decrease may indicate that the therapy is effective, while the absence of a significant decrease or an increase may indicate that the therapy is effective. , indicating that the therapy is ineffective. The first and second time points are before and after treatment, or treatment It could be the following two points in time. For example, comparing the results obtained from one point in time with those from another. The method is such that variations previously identified during the course of treatment are no longer present in the subject. It can be used to determine whether or not it exists. The period between the first time point and the second time point is less It could be at least one month, at least six months, or at least one year, and in some cases Patients may receive treatment regularly, for example every three months, every six months, or annually, for several years, for example, more than five years. It can be tested over a period of time.
[0171] This method also helps determine whether the subject is disease-free or if the disease has recurred. It can be used. As described above, the method can be used for the analysis of minimal residual disease and for detecting recurrence. In these embodiments, the primer pairs used in the method are used to target carcinogens at an earlier time. Either sequence the cfDNA of or sequence another suitable sample. Through this process, the sequence is amplified, including variations previously identified in the patient's cancer. It can be designed to be so.
[0172] In some embodiments, when testing for minimal residual disease or recurrence detection, the patient The DNA test sample will be cell-free DNA. This cell-free DNA will be available at any time after treatment. It can be collected from the patient at a single point. In some embodiments, this cell-free DNA is successfully used in cancer treatment. If treated, the remaining ctDNA from the cancer would have been removed at the time of collection. Obtained. This point in time may depend on factors such as the initial amount of ctDNA and the treatment modality. Regarding methods for removing multiple tumors at once, for example, the timing of surgery is one of the treatments aimed at a cure. It may be after a week, two weeks, three weeks, or four weeks. In addition, these periods can be longer, such as one or two months. Other DNA extracted from alternative sources is also used to assess the presence or amount of cancer DNA. This is possible. Examples include the cellular fraction of cerebrospinal fluid, the cellular fraction and acellular fraction of cerebrospinal fluid. Examples include fecal samples, cells present in urine, biopsies, or fine-needle aspiration samples. Not limited.
[0173] In some embodiments, the method also involves biopsy or fine-needle aspiration of material from lymph nodes or the like. It can be used to assess the presence of residual cancer cells. As will be clear, Such methods allow the number of tumor cells in a biopsy sample to be analyzed by a pathologist through histopathological analysis. The probability of being able to examine enough cells during a biopsy to identify the remaining cancer is practically zero. It would be especially powerful when it's possible to reach a certain level.
[0174] In some embodiments, the method is also used to track multiple variants in parallel. It can be used, for example, to code predicted neoantigens after immunotherapy or personalized vaccines. Track the mutations.
[0175] In some embodiments, the method may be used in clinical trials. For example, the method may be used in clinical trials. To identify a specific group of patients for record-keeping purposes, or to use a new drug (e.g., nonspecific or otherwise) This refers to neoadjuvant therapy or adjuvant therapy that can be targeted to the patient's cancer, or It can potentially be used to evaluate the effectiveness of combination therapies. In some embodiments Therefore, the amount of ctDNA in the patient's bloodstream can be estimated at multiple time points, thereby For example, it would allow for modifications to the dosage of drugs administered to patients during a trial. In this embodiment, the amount of ctDNA in the patient's bloodstream is estimated at multiple time points during the clinical trial. , a specific therapy, level of treatment, duration of treatment, or combination of treatment type and patient may work It can be used to determine whether it is. Many steps of the method are easy to understand. For example, a report showing the presence of cancer DNA in a DNA test sample after a sequencing step. The generation of the data can be performed on a computer. Therefore, in some embodiments, The method involves the patient having DNA samples taken from the patient based on the analysis of sequence reads. This involves executing an algorithm to calculate the likelihood of having cancer DNA and outputting the likelihood. This may include forcing the computer. In some embodiments, this method may involve the computer to process the array. We implement an algorithm that can input data and calculate the likelihood using the input measurements. This may include the act of doing something.
[0176] As should be clear, the calculation steps described could be performed on a computer. Therefore, the instructions for performing the steps are recorded on a suitable physical computer-readable storage medium. The resulting programming can be described as such. Sequence determination reads can be analyzed computationally.
[0177] Embodiment Embodiment 1. A method for detecting tumor DNA in a DNA test sample from a patient. (a) Arrange multiple aliquots of the test sample, and for each aliquot, Sequences corresponding to two or more target regions with sequence variations related to the patient's tumor (b) generating a sequence variant for each aliquot and for each target region To derive an estimate of the number of molecules with sequence variations, or to estimate the number of molecules with sequence variations. (c) Calculate the probability that at least one molecule exists, and the estimate of step (b) or A method including determining whether tumor DNA is present in a test sample using probability. .
[0178] Embodiment 2. For each aliquot, for each target region, the sequence variation in the test sample The number of molecules having a sequence, or the presence of at least one molecule having a sequence variation. The probability of (b)(i) the number of sequence reads of (a) that have sequence variations, (ii (a) Total number of sequence reads, (iii) Number of molecules entered into each aliquot of (a), (iv) Estimated using the estimated background error rate of sequence variations The method of Embodiment 1.
[0179] The estimated background error rate of Embodiment 3(iv) is estimated from the prior sequencing reaction. The method of Embodiment 2 is performed.
[0180] The estimated background error rate of Embodiment 4.(iv) is obtained in step (a). Embodiment 3, estimated from a prior sequencing reaction, adjusted using data from a control base. The method.
[0181] The estimated background error rate of Embodiment 5(iv) is generated in step (a) The method of Embodiment 2, estimated by analysis of control sequencing reads. Embodiment 6. Estimation The value is not a point estimate, but a probability distribution over the number of existing variant molecules, in practice. Method 1.
[0182] Embodiment 7(c) is when (i) ctDNA is present in the sample, (ii) ctD If no NA exists, calculate the likelihood ratio between the likelihoods that observe the estimate in (b). A method of any prior embodiment, performed by...
[0183] Embodiment 8. Likelihood of observing the estimated value of (b) when tumor DNA is present in the test sample. However, (i) the estimated value or probability of step (b), and optionally (ii) the tumor fraction in the test sample The method of Embodiment 7, calculated based on the estimated value.
[0184] Embodiment 9. Observe the estimated value of (b) when tumor DNA is not present in the test sample. The degree is (i) an estimate or probability of step (b), and (ii) high signal background A method according to embodiment 7 or 8, calculated based on the estimated proportion of events.
[0185] Embodiment 10.(c) provides (i) an estimate or probability of step (b), and (ii) a high signal. (iii) Estimated rate of background events, and optionally (iii) Estimated value of tumor fraction in the test sample The calculation is performed by using a hybrid model that incorporates one of the prior embodiments. Law.
[0186] Embodiment 11. Step (c) is to compare the output or likelihood ratio of the mixed model with a threshold. The embodiment further includes an output above a threshold indicating that the test sample contains tumor DNA. 7 or 10 methods.
[0187] Embodiment 12. If the result is above a threshold, the patient is identified as having tumor-associated cells. The method of Embodiment 11 further includes the following:
[0188] Embodiment 13. The method of Embodiment 12, further comprising administering therapy to a patient.
[0189] Embodiment 14. The patient has previously received the first therapy, and the method is different from the first therapy. The method of Embodiment 13, comprising administering a second therapy to the patient.
[0190] Embodiment 15. The method is used to determine the tumor DNA in the test sample based on the estimate in step (b). The amount or range of the likely amount of tumor DNA, for example, the mean or median variant antagonist. Any prior determination which may further include determining by gene fractionation, maximum likelihood, or Bayesian posterior. Method of an embodiment.
[0191] Embodiment 16. A method obtained from a patient between at least a first and second time point. The procedure was performed on a sample, with the first time point being before treatment and the second time point being after treatment. The law determines the amount of tumor DNA or the amount of tumor DNA that is likely to be present between the first and second time points. A method of Embodiment 15, which includes determining whether there is a change in range.
[0192] Embodiment 17. The change is determined using point estimates or confidence intervals, and a significant decrease is observed. The law indicates that it is effective, and the absence of a significant change or increase indicates that the therapy is not effective. The method of Embodiment 16, which demonstrates this.
[0193] Embodiment 18. An embodiment further comprising generating a report indicating whether or not the therapy is effective. Method of application 17.
[0194] Embodiment 19. A statistically unlikely number of aliquot odors based on the estimated tumor fraction. The estimated value of the sequence variation identified is obtained by the result of step (b) before step (c). A method of any prior embodiment that is excluded from the results.
[0195] Embodiment 20. Step (a) determines the arrangement of at least three aliquots. Including the methods of any prior embodiment.
[0196] Embodiment 21. Step (a) also involves tumor DNA from a biopsy or surgical specimen, Buffy Of the following: coat DNA, cheek swab DNA, whole blood DNA, adjacent normal DNA, and reference DNA The process involves sequencing positive and / or negative controls, which may include at least one of the following: A method of any prior embodiment.
[0197] Embodiment 22. Variants not detected in tumor DNA are excluded and buffy coat is applied. Variants detected in the cheek swab, adjacent normal blood, or whole blood are excluded. Method of application 21.
[0198] Embodiment 23. Two or more target regions are any of at least 10 target regions. A method of a prior embodiment.
[0199] Embodiment 24. The sequence variation of step (a) independently represents a single nucleotide. A method of any prior embodiment, which is a variant, indel, dislocation, or rearrangement.
[0200] Embodiment 25. The sequence variation is a sequence variation that has been specified in advance. A method of any of the prior embodiments.
[0201] Embodiment 26. Sequence variations are (i) isolated from a tissue biopsy containing tumor cells. (ii) DNA, and DNA isolated from tumor tissue obtained from surgery containing tumor cells, and sequencing To determine, or (iii) cell-free DNA or (iv) isolated from circulating tumor cells A method of any prior embodiment, identified by sequencing DNA.
[0202] Embodiment 27. Sequence variations are common to the whole genome, whole exome, or cancer mutations. The implementation is identified by sequencing the genome regions selected to include it. Method 26.
[0203] Embodiment 28. Multiple candidate sequence variations are first identified, and the sequence variations However, clonality, mapping potential, estimated error rate, distance from another selected variant Separation, predictive ability to determine sequencing, presence within a region of copy number increase or amplification, and variant alleles. One or more of the proximity of any germline variants that may be used to enrich the genes Methods of embodiments 26-27, selected based on the above.
[0204] Embodiment 29. A patient has cancer, or has had cancer, or does not yet have cancer but has a form of cancer. A method of any prior embodiment having clonal propagation with the potential for qualitative transformation.
[0205] Embodiment 30. The patient has received or is receiving cancer treatment, either prior implementation. Method of form.
[0206] Embodiment 31. A method of any of the prior embodiments, wherein the DNA is cell-free DNA.
[0207] Embodiment 32. Cell-free DNA is isolated from plasma, serum, cerebrospinal fluid, urine, saliva, or feces. The method of embodiment 31.
[0208] Embodiment 33. The tumor DNA fraction in the DNA test sample is 0.01% or less. Method of any of the prior embodiments. Embodiment 34. The test sample is less than 25,000 genome equivalents. A method of any prior embodiment, comprising DNA.
[0209] Embodiment 35. The number of aliquots and the maximum number of molecules per aliquot are as follows: If the number of input molecules in the backing is sufficiently low and a single variant molecule is present, the backing The total number of input molecules and the signal will be significantly different from the round. A method of any prior embodiment that is adjusted based on the estimated background error rate. .
[0210] Embodiment 36. For each aliquot of each sequence variation, step (a) A method of any prior embodiment, wherein the depth of field is at least 10,000.
[0211] Embodiment 37 further includes measuring the amount of DNA in the test sample before step (a). The method of any of the prior embodiments.
[0212] Embodiment 38. The sequence of the target region is obtained by PCR or nucleic acid before step (a). Hybridization to the probe concentrates one of the preceding elements from the test sample. Method of an embodiment. [Examples]
[0213] The following examples provide a complete disclosure and explanation of how the present invention is made and used for those skilled in the art. It has been proposed to provide clarity and to define the scope of what the inventors consider to be their invention. It is not intended to restrict.
[0214] Figure 15 shows that for samples with particularly low tumor fractions, the sample is considered to contain tumor DNA. This shows why it may be difficult to do. As shown in the panel above, a high tumor fraction (T Samples with F) yield several positive signals in multiple aliquots. This allows for easy recall of tumor DNA, which eliminates most false positives. As shown in the panel, samples with a low tumor fraction have data background It can be explained by the rate of Ra, making it more difficult to call. For example, each positive Varian If the sequence has an 80% probability of corresponding to the actual sequence variation, then the low in Figure 15 The evidence presented regarding tumor fraction samples is insufficient to conclude that the samples contain tumor DNA. However, when evidence is aggregated across multiple variants and aliquots... In total, there may be sufficient evidence to consider the sample to contain tumor DNA.
[0215] Figure 11 shows how evidence can be combined across multiple variants. For dilute samples (<<0.1% tumor fraction), due to the dropout effect, the individual components in each sample The variant read fractions of each variant are not expected to approximate the overall tumor fraction. For example, many variants and aliquots will contain 0 molecules instead. The effect of obtaining n / input reads per aliquot as a discrete distribution is modeled. In this example, the tumor fraction is not measured directly. Rather, it is measured against all possible inputs. Excluding it, it provides an accurate estimate of the tumor fraction of the sample. Specifically, variant fraction Instead of guessing the number of offspring, the probability of all possible values is (i) having a variation in the array. (ii) the number of sequencing reads, (ii) the total number of sequencing reads, (iii) the input for each aliquot. (iv) Based on the number of molecules being analyzed and the estimated background error rate of sequence variations. The values with the highest probability are then calculated and identified. This avoids guesswork. In Figure 15, variants are shown as present or absent for each aliquot. However, these actually include many things such as tumor fraction and noise estimates per base. This is the probability considering the factors. A ground truth line (Figure 16) can be constructed. Figure 14 shows particularly noisy variations, i.e., statistically unlikely numbers. The variation identified in the aliquot can be excluded from the analysis. show.
[0216] Figure 17 shows the results of three different samples containing varying levels of circulating tumor DNA (ctDNA). In each of the four aliquots, more than 40 sequence variations were found using this method. The results of the analysis and experiment are shown. The 52 ppm and 544 ppm samples contain ctDNA. It was identified that this was beneficial for combining evidence across multiple aliquots and variants. The points are shown. In this figure, color intensity correlates with VAF (Variant Allele Fraction), and the brightest point is... The lighter color represents >=1%. Some variant names indicate their absence in the original tumor sample. It is displayed in gray to indicate its presence.
[0217] Example 1 To construct the optimal assay for detecting residual lesions, the target cancer type, this In this case, breast cancer was initially selected. The cancer mutation rate was examined, and in approximately 90% of patients... It was identified as a mutation of more than 0.5 per 1 Mb, and the average patient has 1 per 1 Mb. It had mutations exceeding (Martincorena and Campbell, Sci. (nce 2015 349:1483-9). A pilot study of 22 early-stage breast cancer patients. Therefore, ctDNA is detected from a median of 0.06% VAF to 0.0007% VAF. It was determined that...
[0218] A study diluting three cancer cell lines with normal DNA tracked 48 variants. Performed using a separate assay, the analysis combined 48 variants yielded the following results. While cancer DNA can be consistently detected at 0.001% VAF, variants... We demonstrated that the level of sensitivity is halved each time the number is halved.
[0219] The mutation rate of breast cancer, ctDNA, is detected with a probability of approximately 50% and less than 0.06%, and pilot Based on observations in the study that VAF can be detected down to 0.0007%, We set targets for at least 90% of breast cancer samples that have a detection limit of 0.001% VAF. With a mutation rate of 0.5 mutations per 1Mb, sequencing in breast cancer requires 96 of the genome. A MB area was required.
[0220] The main advantage of this approach is that at least 90% of patients have ≥48 variants. To identify the target cancer type, the required level of sensitivity must be reproducibly achieved. This includes. Another advantage is that when samples with lower mutation rates are targeted, the sequencing cost is reduced. It can be reduced.
[0221] Example 2 To design the optimal MRD assay, the system incorporates as many high-quality barriers as possible. The procedure is designed to investigate the tumor. To do this, a tumor biopsy is first obtained, and 50% of the tumor Macroscopic dissection targeting the ulcer content, exome capture, and then Illumina The sample is sequenced using a sequencer. All potential variants are standard Illu Identify using the mina pipeline, then 1) likelihood of being real, and 2) somatically 3) likelihood, 4) background error rate of the variant, 5) high signal background Dweller rate, 5) probability of clonality, 6) level of variant amplification or copy number increase It gives a combined score based on the following. The genome is divided into 50bp windows, These windows overlap by 25bp. Each window contains 1) elements within the window 1) Score of all variants, 2) Score of ability to uniquely align the area (uniquely align Penalties are given for areas where alignment is not possible, and the number of misalignments is large. (The larger the value, the higher the penalty), 3) Score of the ability to amplify and arrange the region (repetition) Features that are known to challenge sequencing decisions that include are penalized. A combined score is given. Then, regions are selected by the score and PCR primer is applied. Select the top 100 to design. Two overlapping areas are in the list of the top 100. If this is the case, the region with the highest score is maintained, and the region with the weaker score is destroyed. It is discarded. Next, the 101st region is added to the list, and so on. Multiplex PCR is Designed for the top 48 variants. Insilico PCR is used for all plastics. This is done using primer pairs. The primer combinations that produce a nonspecific region of ≥2 are identified. This breaks down the primer in the lowest-scoring region that is causing this nonspecific product. Discard and design alternative primers. If the nonspecific PCR problem is not overcome, discard the region. The following areas are added to the primer design.
[0222] One challenge with this tumor information-based method for detecting cancer DNA in test samples is its robustness. This is the number of areas that can be targeted cost-effectively. This strategy ranks the areas. This can maximize the number of variants successfully investigated in the test DNA sample. When riants are in cis (adjacent on the same chromosome), they are read together. This allows for increased ability to isolate signals from noise. Although it is in the same state, it is still read with the same primer pair (or other targeting reagents such as bait). When available, the amount of information from a single targeted region should double. The approach should also limit the number of reads wasted on non-specific products.
[0223] Example 3 To detect cancer DNA in test samples with high sensitivity, multiple variants must be targeted. This is advantageous. For some cancer types, targeting only one type of variant is beneficial. It is sufficient to target it. However, it is better to target multiple types of variants. In some cases, multiple structural variants exist for a particular breast cancer patient. Other patients are identified to have more SNVs and indels. Breast cancer tumor D Large panel to sequence NAs and evaluate for SNVs, indels, and rearrangements Design the variant. Identify the optimal variant that includes the region. Target these regions. Design the primers for the region. If the region contains one or more SNVs / indels, the primers Design to be adjacent to all SNVs and indels. Identify that the "region" contains rearrangements. In such cases, it is likely that two different parts of the same chromosome or two different chromosomes are joined together. The rearranged sequence is used in primer design, and one primer is 3' of the rearranged sequence. The other is 5'. SNV, indel, or other variant (e.g., DBS) In cases of cis with rearrangement, the primer uses the rearranged sequence obtained from the tumor. It is designed to be used to be adjacent to both rearranged and other variants. The advantage of this method is that it allows for the consistent acquisition of a large number of variants for evaluating cancer DNA in test samples. It is power.
[0224] Example 4 Determine both the background error rate and the percentage of high-signal background events. To achieve this, design 50 different panels, each with 48 amplicons. Each of these targets the exome of patients with lung cancer, CRC, or breast cancer. Design. Each amplicon in the panel has an average length of approximately 100bp, and within this, There is a sequence (i.e., a non-primer sequence) that can be read from approximately 60 bp of test DNA. Blood samples were obtained from 200 healthy donors. Each donor's blood was collected using Streck cell-free DNA collection. It is drawn into the blood vessels. The blood is rotated to plasma, cell-free DNA is extracted, and then it is processed using a digital PC. DNA is quantified using R. Each panel is tested with cfDNA from four donors. Using the panel and cfDNA, set up multiplex PCR with multiple aliquots (3). Process. This PCR is barcoded. The barcoded product from the patient is processed together. These are executed using an Illumina NovaSeq sequencer. Evaluation is then performed. The variant types are agreed to be SNV and indel. Divide into the following classes: SNV type (e.g., C>A, T>A, or G>A), in Dell type and size (e.g., 1bp, 2bp, 3bp Dell). From the donor. Based on the results of digital PCR quantification of cfDNA, the results were divided into three groups (low DNA input, medium D Divide into NA input and high DNA input. Primer sequence, 3 bp buffer, latent Except for all locations where specific germline variants have been reported in gnomAD, the remaining germline variants at each location For each base, the total number of reads, the number of each non-reference base, and the different types of indels Obtain the number of p / size. For each change (e.g., C > A), fit a beta distribution to the data. This will yield both the mean and the coefficient of variation (CV). The cumulative distribution function (CDF) of a specific base change will be used. Using a threshold of 0.9999, the contrasting residue that must be considered positive in a sample is determined by the threshold value used. Determine the gene fraction cutoff. This is the background error rate. High signal bar To determine the proportion of ground events, for each change (e.g., C>A), the test was conducted. Evaluate all instances of change in the panel and set the threshold for the allele fraction determined by the CDF. Calculate the percentage of signals that exceed the limit.
[0225] Example 5 A biopsy sample is obtained, 96 Mb of the tumor genome is sequenced, and then 48 regions are amplified. By selecting primers, we design a panel for tumors in breast cancer patients, and in total , 48 regions are somatic and 50 variants are thought to be tumor-specific. Includes SNVs and indels). Patient-specific primers are multiplexed and tumor DNA is used. Set up multiplex PCR. Barcode the PCR products, then Illumina The sequence is determined using a sequencer. Variants not detected in tumor DNA are bioinjected. Filter using morphology. Apply the same panel to buffy coat DNA from patients. Apply. Generate a library and determine the sequence. All VAFs identified in over 40% Flag each variant as germline and filter them. Variant type and exceed the allele fraction cutoff as determined by the background error rate. All variants identified as having less than 40% potential are given a Claw with undetermined potential. Flagged as potentially hematopoiesis and filtered. After filtering 1 If more than two variants remain, apply the panel to cfDNA extracted from the patient. If there is little remaining, attempt to redesign the panel. Divide the CfDNA into three aliquots, Multiplex PCR is performed using patient-specific primers in all three aliquots. PCR product The samples are barcoded, beads are cleaned up, then the samples are pooled and sequenced. Upon completion of sequencing, the reads are demultiplexed, trimmed, filtered based on quality, and Align to the reference genome. For each target region, all variants within that target region are examined. Then, count the number of wild-type leads and the total number of leads.
[0226] Example 6 After completing the sequencing of three aliquots of cfDNA from breast cancer patients, the total number of variants and The total amount of all aliquots of all variants except for filtered variants Obtain the variant allele fraction (mutant / total reads), and then determine this variant. The target allele fraction is compared to a threshold generated using the background error rate. Evaluate all aliquots of all variants and determine whether they are positive or negative (the threshold is set). Determine if it exceeds the limit. The tumor fraction is first analyzed using the background error rate. Correct the VAF and then average it over all aliquots of all variants. This is estimated by comparing the number of DNA molecules attached to each library preparation with the average VAF. In comparison, it is possible to predict at least one variant molecule in each aliquot of each variant. The possibility is determined. Then, each variant is evaluated to see if there are more positives than expected by chance. Determine whether aliquots are present and identify an unlikely number of positive aliquots (P<0.05). Filter those that are determined to have the following characteristics. Then, assign a score of 1 to the high signal bar. This applies to any variant that does not have a ground event (e.g., typically an indel). For the remaining variants, assign them a high percentage of "high signal background events". Those with a combination of the above (top 50%) and those with a low percentage of "high-signal background events" (These are in the bottom 50% excluding those that do not have "high-signal background events") Separate into all of them. All variants with a low percentage will be grouped towards a score of 0.75. Those with a high proportion contribute to a score of 0.5. If there are two or more test DNA samples... If it is determined that there is a total score, and at least two aliquots are 0.5 or higher If a score is obtained, the test sample is considered to contain cancer DNA. There are many advantages to this approach. In some approaches, if enough variants exceed the threshold It is possible to simply determine whether there are dolphins (for example, two variants that exceed a threshold). While some variants generally result in high-signal background events, others... Because it never brings about such a result, it is limited. Therefore, this approach is limited by these barriers. When the signal never produces a high-signal background event, the two variants When detected, it enables reliable calling with high specificity. Identified Varian If the system is more prone to high-signal background events, the scoring approach is Therefore, to be more careful and to ensure that the assay maintains high specificity This requires 3-4 variants. It requires scoring in multiple aliquots. By doing so, the assay prevents false positives due to contamination of a single aliquot, and Baffyco It is present in more aliquots than is likely based on the estimated tumor fraction. Filter out variants that are either present or not, while checking the CHIP and error Common sources of false positives containing bases prone to generating false positives are eliminated.
[0227] Example 6 After completing the sequencing of three aliquots of cfDNA from breast cancer patients, the total number of variants and The total amount of all aliquots of all variants except for filtered variants Obtain the variant allele fraction (mutant / total reads), and then determine this variant. The target allele fraction is compared to a threshold generated using the background error rate. Evaluate all aliquots of all variants and determine whether they are positive or negative (the threshold is set). Determine if it exceeds the limit. The tumor fraction is first analyzed using the background error rate. Correct the VAF and then average it over all aliquots of all variants. This is estimated by comparing the number of DNA molecules attached to each library preparation with the average VAF. In comparison, it is possible to predict at least one variant molecule in each aliquot of each variant. The possibility is determined. Then, each variant is evaluated to see if there are more positives than expected by chance. Determine whether aliquots are present and identify an unlikely number of positive aliquots (P<0.05). Filter those that are determined to have the following characteristics. Then, call the number of variants. The threshold is the high signal background of all remaining unfiltered variants. Obtain the estimated proportion of events, then the high signaling across all remaining aliquots and variants. This is determined by calculating the distribution of the likely number of background events. A threshold number of positive variants is obtained, and positive variants are detected purely through high-signal background events. There is a change of less than 0.01% in the number of sexual events. Positive variants (variants exceeding the VAF threshold) If the total number of riants exceeds the threshold number of positive variants, and at least two ariants If the recoat has a positive variant, the sample is then considered positive. There are many advantages to the approach. In some approaches, a sufficient number of variants exceed the threshold. It is possible to simply determine whether (for example, two variants that exceed a threshold). This is because some variants generally result in high-signal background events, Because it is never generated by others, it is limited. Therefore, this approach is high signal bus To estimate how common and distributed ground events are. This enables reliable calling. Next, it determines how noisy the variant is. Set individualized thresholds depending on how many variants exist. This enables extremely high sensitivity, but also allows for a balance with specificity (for example, common high When multiple variants with signal background events are tested, the threshold is high. When a small number of variants that rarely have signal background events are being tested. (The rate is also high). By requiring positive results in multiple aliquots, the assay is simple. To prevent false positives due to contamination of one aliquot, and to detect the presence of or presumed tumor in the buffy coat. The variance is likely to be present in more aliquots than the ulcer fraction. While filtering out ants, false positives containing CHIP and error-prone bases are filtered out. The common source of sex will be eradicated.
[0228] Example 7 Obtain FFPE tumor material. Section the tissue and extract total RNA from 10 slides. Prepare the library by performing ribosomal RNA depletion, reverse transcription, and sequencing. The library is barcoded, and then duplicated with other libraries from the patient. Sequencing is performed on the Ilumina NovaSeq platform. Reads are demultiplexed. It is transformed, aligned, and then the variant is called. The variant is SNV, indel , and gene fusion products. These variants are then used for primer design. The RNA transcripts are mapped to the correct genomic DNA coordinates. The present invention provides, for example, the following items: (Item 1) A method for detecting cancer DNA in a DNA test sample from a patient, (a) Sequence multiple aliquots of the test sample, and for each aliquot, generate sequence reads corresponding to two or more target regions having sequence variations present in the patient's cancer, (b) For each aliquot, for each target region: i. Determine the number of sequence reads having the sequence variation, ii. Determine the total number of sequence reads, iii.i. and ii. are compared with one or more error probability distribution models of the sequence variation, and if the one or more models are obtained from DNA that does not contain the sequence variation, iv. Excluding variants that exceed a threshold in a statistically unlikely number of aliquots, (c) A method comprising integrating the collective results of step (b) to determine whether cancer DNA is present in the test sample. (Item 2) A statistically unlikely number of aliquots, The amount of test sample DNA added to each aliquot was measured. Using sequencing data of all or a subset of the aforementioned variants, the fraction of cancer DNA in the test sample is calculated. The method according to item 1, which is identified by estimating the probability of observing the number of aliquots containing the sequence variation exceeding a threshold, based on i. and ii. (Item 3) The method according to any of the preceding items, wherein the fraction of cancer DNA in the DNA test sample is 0.01% or less. (Item 4) The method according to any of the preceding items, wherein step (a) involves sequencing at least 10 target regions in at least 3 aliquots of the test sample. (Item 5) The method of any of the preceding items, wherein the method includes identifying a set of sequence variations present in the patient's cancer prior to step (a). (Item 6) The method according to any of the preceding items, wherein the cancer is a blood cancer, and the test sample comprises cellular DNA isolated from cells from peripheral blood, lymph nodes, or bone marrow. (Item 7) The method according to any one of items 1 to 5, wherein the cancer is a solid tumor and the test sample contains cfDNA. (Item 8) Step (b) is, v. (i) Derive an estimate of the number of molecules having the sequence variation, (ii) Calculate the probability that there is at least one molecule having the sequence variation, (iii) Determine whether the frequency of sequence reads having the sequence variation, compared to the total number of sequence reads, exceeds a threshold. (iv)Calculate the likelihood ratio of (i), and / or (v) A method for any of the preceding items, comprising determining whether any of (i), (ii), or (iv) exceeds a threshold. (Item 9) The method of any of the preceding items, further comprising calculating the fraction or total amount of cancer DNA in the test sample based on the result of step (b). (Item 10) (b)(iv) is the sample: (i) If cancer DNA is present (ii) If cancer DNA is not present In (b)(i), calculate the likelihood ratio between the likelihoods that observe the result obtained in (b)(i). The method described in item 8, which is performed by combining individual likelihood ratios with cumulative likelihood ratio scores across all sequence variations and aliquots of the test sample. (Item 11) The method of any of the preceding items, further comprising identifying the patient as having cancer if the result of step (c) is greater than or equal to the threshold. (Item 12) A method of any of the preceding items, further comprising administering therapy to the aforementioned patient. (Item 13) The method of any of the preceding items, wherein the patient has previously received a first therapy, and the method comprises administering a second therapy to the patient, which is different from the first therapy, based on the outcome of step (c). (Item 14) The method according to any of the preceding items, wherein the patient has or has had cancer, or has clonal proliferation that is not yet cancerous but has the potential to transform. (Item 15) The method described in either of the preceding items, wherein the patient has received or is receiving treatment for the cancer.
Claims
[Claim 1] The invention described in the specification.