Normalization of Tumor Gene Mutation Quantity
By determining and normalizing the number of mutations and minor allele ratios in cell-free nucleic acids, this method addresses the challenges of cancer detection and immunotherapy selection, providing a more accurate indicator of tumor gene mutations for improved patient treatment outcomes.
Patent Information
- Application Number
- JP2023181005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-11-03
- Filing Date
- 2023-10-20
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2038-11-02
AI Technical Summary
Current methods for detecting cancer through cell-free nucleic acids in body fluids are hindered by the small amount of nucleic acid released, variability in its presence, and challenges in accurately determining tumor gene mutations for predicting response to immunotherapy.
A method involving the determination of the number of mutations and minor allele ratio in cell-free nucleic acids, followed by normalization using control samples to establish an indicator of tumor gene mutations, which can help identify subjects likely to respond positively to immunotherapy.
This approach provides a more accurate and reliable indicator of tumor gene mutations, enabling better selection of patients for immunotherapy by normalizing mutation data across different samples and cancer types.
Smart Images

Figure 0007696975000001 
Figure 0007696975000002
Abstract
Description
Technical Field
[0001] Cross-reference This application claims priority based on U.S. Provisional Application No. 62 / 581,563, filed on November 3, 2017, which is hereby incorporated by reference in its entirety for all purposes.
Background Art
[0002] Background A tumor is an abnormal growth of cells. When cells, such as tumor cells, die, fragmented DNA is often released into the body fluid. Thus, a part of the cell-free DNA in the body fluid is tumor DNA. Tumors can be either benign or malignant. Malignant tumors are often referred to as cancer.
[0003] Cancer is a major cause of disease worldwide. Each year, tens of millions of people are diagnosed with cancer around the world, and more than half ultimately die from cancer. In many countries, cancer is ranked as the second leading cause of death after cardiovascular disease. For many cancers, early detection is associated with improved outcomes.
[0004] Cancer is usually caused by the accumulation of mutations in the normal cells of an individual, at least some of which result in unregulated cell division. Such mutations generally include single nucleotide variants (SNVs), gene fusions, insertions and deletions (indels), transversions, translocations, and inversions. The number of mutations within a cancer is an indicator of cancer susceptibility to immunotherapy.
[0005] Cancer is often detected by biopsy of a tumor followed by analysis of cytopathology, biomarkers or DNA extracted from cells. However, more recently, it has been proposed that cancer can also be detected from body fluids, such as cell-free nucleic acids (e.g., circulating nucleic acids in blood, circulating tumor nucleic acids in blood, exosomes, nucleic acids derived from apoptotic cells and / or necrotic cells) in blood or urine (see, e.g., Siravegna et al., Nature Reviews 2017). Such tests have the advantage that they are non-invasive and can be performed without identifying cancer-suspected cells by biopsy and sample nucleic acids from each part of the cancer. However, such tests are complicated by the fact that the amount of nucleic acid released into body fluids is small and varies as well as recovering nucleic acids from such liquids in an analyzable form. These variable factors can make the predicted values for comparing the amount of tumor gene mutations (TMB) between samples ambiguous.
[0006] TMB is a measure of the mutations carried by tumor cells in the tumor genome. TMB is a type of biomarker that can be used to evaluate whether a subject diagnosed with or suspected of having cancer symptoms will benefit from cancer therapy, such as cancer immunotherapy (I-O). SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0007] One aspect of the present disclosure is a method for providing an indicator of the amount of tumor gene mutations in a test sample of cell-free nucleic acids from a subject having a cancer type or a sign of a cancer type, the method comprising: (a) determining the number of mutations present in the cell-free nucleic acids of the test sample and the minor allele ratio based on one or more mutations most frequently shown in the cell-free nucleic acids of the test sample; and (b) normalizing the number of mutations present in the sample to the number of mutations present in a control sample from another subject having the same cancer type and to the minor allele ratio within a bin of minor allele ratios including the minor allele ratio of the test sample to determine an indicator of the amount of cancer gene mutations in the test sample.
[0008] In some embodiments, the number of mutations present in the control sample is an average.
[0009] In some embodiments, the bin has a width of 20% or less, 10% or less, or 5% or less.
[0010] In some embodiments, the method further comprises determining whether the number of mutations present in the sample exceeds a threshold, the threshold being set to indicate a subject who tends to respond positively to immunotherapy.
[0011] In some embodiments, the normalizing step comprises dividing the number of mutations in the test sample by the average number of mutations in the control sample.
[0012] In some embodiments, the normalizing step comprises subtracting the average number of mutations in the control sample within the bin from the determined number of mutations in the test sample of cell-free nucleic acids.
[0013] In some embodiments, the method further comprises dividing the number of mutations in the test sample of cell-free nucleic acids minus the average number of mutations in the control sample by the standard deviation of the number of mutations in the control sample to calculate a Z-score. The average can be a mean.
[0014] In some embodiments, the normalizing step includes determining the mean and spread of the number of mutations in at least 10, 50, 100, or 500 control samples, determining a standard score of the deviation from the mean in the test sample, and determining whether the standard score exceeds a threshold number. The mean can be a mean value, median, or mode. The spread can be represented as variance, standard deviation, or interquartile range. The standard score of the deviation can be a Z-score.
[0015] In some embodiments, the normalizing step further includes dividing the determined number of mutations in the test sample of cell-free nucleic acid by the average number of mutations present in the control samples within the same bin.
[0016] In some embodiments, the normalizing step is performed on a computer programmed to store values of the number of mutations present in a plurality of bins of minor allele frequencies. The values stored can be the mean value and standard deviation of the number of mutations present in each of the plurality of bins.
[0017] In some embodiments, it includes determining a standard score of the amount of tumor gene mutations in a subject and determining whether the standard score exceeds a threshold for a control subject, in accordance with responsiveness to immunotherapy.
[0018] In some embodiments, (a) includes determining the sequence of cell-free nucleic acid molecules in a test sample and comparing the obtained sequence to a corresponding reference sequence to identify the number of mutations and minor allele frequencies present in the sample. The reference sequence is hG19 or hG38.
[0019] In some embodiments, the control sample includes at least 25, 50, 100, 200, or 500 control samples.
[0020] In some embodiments, at least 50,000, 100,000, or 150,000 nucleotides are sequenced in a segment of the nucleic acid.
[0021] In some embodiments, (a) involves determining the presence or absence of a panel of predetermined mutations known to occur in a type of cancer present or suspected to be present in the sample, and optionally, the mutations are somatic mutations that affect the sequence of the encoded protein.
[0022] In some embodiments, step (a) includes ligating an adapter to the cell-free nucleic acid, amplifying the cell-free nucleic acid from primers that bind to the adapter, and sequencing the amplified nucleic acid.
[0023] In some embodiments, sequencing is bridge amplification sequencing, pyrosequencing, ion semiconductor sequencing, paired-end sequencing, sequencing by ligation, or single molecule real-time sequencing.
[0024] In one aspect, the present disclosure relates to a method of treating a subject, comprising: (a) determining the number of mutations present in the cell-free nucleic acid of a test sample and the minor allele ratio based on one or more mutations most frequently shown in the cell-free nucleic acid of the test sample; (b) normalizing the number of mutations present in the sample to the minor allele ratio within a bin of minor allele ratios including the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio of the test sample to determine an indicator of the amount of cancer gene mutations in the test sample; and (c) administering immunotherapy to the subject if the indicator of the amount of tumor gene mutations exceeds a threshold.
[0025] In some embodiments, the method is performed in a plurality of subjects to determine an indicator of the amount of tumor gene mutations in each subject, and a higher proportion of subjects having an indicator of the amount of cancer gene mutations exceeding the threshold compared to subjects having an indicator of tumor gene mutations below the threshold receive immunotherapy for cancer.
[0026] In some embodiments, all subjects with an indicator exceeding a first threshold receive immunotherapy, and all subjects with an indicator below a second threshold do not receive immunotherapy.
[0027] In some embodiments, the indicator is a Z-score.
[0028] In some embodiments, immunotherapy includes the administration of checkpoint inhibitor antibodies.
[0029] In some embodiments, immunotherapy includes the administration of antibodies against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40.
[0030] In some embodiments, immunotherapy includes the administration of pro-inflammatory cytokines.
[0031] In some embodiments, immunotherapy includes the administration of T cells against the cancer type.
[0032] In some embodiments, the cancer type is a solid tumor.
[0033] In some embodiments, the cancer type is renal cancer, mesothelioma, soft tissue cancer, primary CNS cancer, thyroid cancer, liver cancer, prostate cancer, pancreatic cancer, CUP, neuroendocrine cancer, NSCLC, gastroesophageal cancer, head and neck cancer, SCLC, breast cancer, melanoma, cholangiocarcinoma, gynecological cancer, colorectal cancer or urothelial cancer.
[0034] In some embodiments, the cancer type is a hematological malignancy.
[0035] In some embodiments, the cancer type is leukemia or lymphoma.
[0036] In one aspect, the present disclosure is a method of treating a subject having cancer, the method comprising administering an immunotherapeutic agent to the subject, wherein the subject is identified for immunotherapy from an index of the amount of cancer gene mutations of the subject determined by: (a) determining the number of mutations present in the cell-free nucleic acids of a sample from the subject and the minor allele ratio for the mutations most frequently represented in the cell-free nucleic acids of the test sample; and (b) normalizing the number of mutations present in the sample to the minor allele ratio within a bin of minor allele ratios that includes the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio of the test sample to determine an index of the amount of tumor gene mutations in the sample of the subject; and wherein the subject is determined to have a tumor gene mutation amount that exceeds a threshold.
[0037] The present disclosure provides (1) a communication interface that receives, via a communication network, sequencing reads obtained by sequencing cell-free nucleic acids in a test sample, and (2) a computer that communicates with the communication interface, the computer comprising one or more computer processors and, when executed by the one or more computer processors: (a) receiving, via the communication network, the sequencing reads obtained by a nucleic acid sequencer; (b) determining the number of mutations present in the sequencing reads from the test sample and the minor allele ratio based on one or more mutations most frequently represented in the sequencing reads from the test sample; and (c) normalizing the number of mutations present in the test sample to the minor allele ratio within a bin of minor allele ratios that includes the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio of the test sample to determine an index of the amount of cancer gene mutations in the test sample and comprising a computer-readable medium containing machine-executable code for implementing a method comprising the foregoing. The disclosure further provides a system comprising the foregoing.
[0038] In some embodiments, a sequencing library obtained from cell-free DNA molecules derived from a subject is sequenced by a nucleic acid sequencer, and the sequencing library includes cell-free DNA molecules and adapters containing barcodes. In some embodiments, sequencing by synthesis is performed on the sequencing library by a nucleic acid sequencer to obtain sequencing reads. In some embodiments, pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, or sequencing by hybridization is performed on the sequencing library by a nucleic acid sequencer to obtain sequencing reads. In some embodiments, a clonal single molecule array derived from the sequencing library is used by a nucleic acid sequencer to obtain sequencing reads. In some embodiments, the nucleic acid sequencer includes a chip having an array of microwells for sequencing the sequencing library to obtain sequencing reads. In some embodiments, the computer-readable medium includes a memory, a hard drive, or a computer server. In some embodiments, the communication network includes a long-distance communication network, the Internet, an extranet, or an intranet. In some embodiments, the communication network includes one or more computer servers capable of distributed computing. In some embodiments, the distributed computing is cloud computing. In some methods, the computer is installed on a computer server remotely located from the nucleic acid sequencer. In some embodiments, the sequencing library further includes a sample barcode that identifies the sample from one or more samples. In some embodiments, the system further includes an electronic display that communicates with the computer through a network, the electronic display including a user interface for displaying the results when (a)-(c) are performed.In some embodiments, the user interface is a graphical user interface (GUI) or a web-based user interface. In some embodiments, the electronic display is present in a personal computer. In some embodiments, the electronic display is present in an Internet-enabled computer. In some embodiments, the Internet-enabled computer is installed at a location remote from the computer.
[0039] In some embodiments, the results of the systems and methods disclosed herein are used as input and a report is created in paper form. For example, this report can provide an indication of the called variants and / or variants considered to be deamination errors.
[0040] The various steps of the methods disclosed herein, or the steps performed by the systems disclosed herein, can be performed at the same or different times, at the same or different geographical locations, e.g., in a country, and / or by the same or different people.
Brief Description of the Drawings
[0041]
Figure 1
[0042]
Figure 2
Modes for Carrying Out the Invention
[0043] Definitions The subject refers to an animal, such as a mammalian species (preferably a human) or an avian (e.g., a bird) species, or another organism, such as a plant. More specifically, the subject is a vertebrate, such as a mammal like a mouse, a primate, an ape or a human. Animals include livestock, sport animals, and pets. The subject can be a healthy individual, an individual having or suspected of having a disease or a predisposition to a disease, or an individual in need of or suspected of being in need of treatment.
[0044] For example, the subject is an individual who has been diagnosed with having cancer, an individual who is scheduled to receive cancer treatment, and / or an individual who has received at least one cancer treatment. The subject may be in remission from cancer. As another example, the subject is an individual diagnosed with having an autoimmune disease. As another example, the subject can be a pregnant individual or an individual planning to become pregnant, and this subject may have been diagnosed with or suspected of having a disease, such as cancer, an autoimmune disease.
[0045] A cancer marker is a genetic variant associated with the presence of cancer or the risk of developing cancer. A cancer marker can be an indication that a subject has a higher risk of developing cancer than an age- and sex-matched subject of the same species who does not have cancer or does not have the cancer marker. A cancer marker may or may not be the cause of cancer.
[0046] Barcodes can be attached to one or both ends of a nucleic acid. The barcode can be decoded to reveal information such as the original sample, nucleic acid form, or processing. Using barcodes enables pooling and parallel processing of multiple samples containing nucleic acids with different barcodes (which are then deconvolved by reading the barcode) related to the nucleic acids. Barcodes can also be referred to as molecular identifiers, sample identifiers, tags, or index tags. Barcodes can be used to identify samples (sample identifiers). Additionally or alternatively, barcodes can be used to identify different molecules within the same sample. This includes both uniquely barcoding each different molecule in the sample and using non-unique barcode addition to each molecule. In the case of non-unique barcode addition, a limited number of barcodes may be used to barcode each molecule, such that different molecules can be identified based on their start / stop positions where they map onto the reference genome in combination with at least one tag. Then, typically, using a sufficient number of different barcodes results in a low probability (e.g., <10%, <5%, <1%, or <0.1%) that the barcodes of any two molecules with the same start / stop are the same. Some barcodes contain multiple molecular identifiers to label multiple samples, multiple forms of molecules within one sample, and multiple molecules within one form with the same starting and stopping points. Such barcodes can be present in form A1i, where the letters indicate the sample type, the Arabic numerals indicate the form of the molecule within the sample, and the Roman numerals indicate the molecule within the form.
[0047] An adapter is usually a short nucleic acid that is at least partially double-stranded (e.g., less than 500, 100, or 50 nucleotides in length) for ligation to one or both ends of a sample nucleic acid molecule. The adapter may include a primer binding site that enables amplification of the nucleic acid molecule adjacent to the adapter at both ends, and / or a sequencing primer binding site that includes a primer binding site for next-generation sequencing (NGS). The adapter may also include a binding site for a capture probe, such as an oligonucleotide attached to a flow cell support. The adapter may also include the barcode described above. Preferably, the barcode is positioned relative to the primer and the sequencing primer binding site such that it is included in the amplicon and sequencing reads of the nucleic acid molecule. The same or different adapters can be ligated to each end of the nucleic acid molecule. The same adapter may be ligated to each end, except that the barcode may be different. A preferred adapter is a Y-shaped adapter that is blunt-ended or has a tail as described herein for joining to a nucleic acid molecule that is also blunt-ended or has a tail with respect to one or more complementary nucleotides. Another preferred adapter is a bell-shaped adapter that similarly has a blunt or tailed end for joining to the nucleic acid to be analyzed.
[0048] As used herein, the term "sequencing" refers to any of several techniques used to determine the sequence of a biomolecule, such as a nucleic acid molecule, e.g., DNA or RNA. Exemplary sequencing methods include, but are not limited to, target sequencing, single molecule real-time sequencing, exon sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy terminator sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single nucleotide extension sequencing, solid-phase sequencing, high-throughput sequencing, ultraparallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature (COLD-PCR), multiplex PCR, sequencing by reversible terminator dyes, paired-end sequencing, short read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a gene analyzer, such as a gene analyzer commercially available from Illumina or Applied Biosystems.
[0049] The term "next-generation sequencing" or NGS refers to sequencing technologies with increased throughput compared to traditional Sanger and capillary electrophoresis-based methods, and has, for example, the ability to generate hundreds of thousands of relatively small sequence reads at once. Some examples of next-generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.
[0050] The term "sequencing run" refers to any step or portion of a sequencing experiment that is performed to determine some information about at least one biomolecule (e.g., a nucleic acid molecule such as DNA or RNA).
[0051] DNA (deoxyribonucleic acid) is a nucleotide chain containing four types of nucleotides: adenine (A), thymine (T), cytosine (C), and guanine (G). RNA (ribonucleic acid) is a nucleotide chain containing four types of nucleotides: A, uracil (U), G, and C. Specific nucleotide pairs bind specifically to each other in a complementary manner (referred to as complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides that are complementary to the nucleotides in the first strand, the two strands bind to form a double strand. As used herein, "nucleic acid sequencing data", "nucleic acid sequencing information", "nucleic acid sequence", "nucleotide sequence", "genomic sequence", "gene sequence", or "fragment sequence", or "nucleic acid sequencing read" refers to any information or data that indicates the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a nucleic acid molecule such as DNA or RNA (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment). It should be understood that, according to the teachings of the present invention, sequence information obtained using all available variations of techniques, platforms or technologies including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion or pH-based detection systems, and electronic signature-based systems is contemplated.
[0052] "Polynucleotide", "nucleic acid", "nucleic acid molecule", or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by linkages between the nucleosides. Typically, a polynucleotide contains at least 3 nucleosides. The size of an oligonucleotide often ranges from a few monomer units, e.g., 3 - 4, up to several hundred monomer units. When a polynucleotide is represented by a character sequence such as "ATGCCTG", the nucleotides are always in the 5' to 3' order from left to right, unless otherwise noted, where "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents thymidine. As is standard in the art, the letters A, C, G, and T can be used to refer to the bases themselves, the nucleosides containing the bases, or the nucleotides.
[0053] A reference sequence is a known sequence used for comparison with experimentally determined sequences. For example, the known sequence can be an entire genome, a chromosome, or any segment thereof. A reference typically contains at least 20, 50, 100, 200, 250, 300, 350, 400, 450, 500, 1000, or more nucleotides. A reference sequence can be aligned with a single continuous sequence of a genome or chromosome or can contain non - continuous segments aligned with different regions of a genome or chromosome. Examples of reference human genomes include hG19 and hG38.
[0054] When the first nucleic acid sequence or its complement and the second nucleic acid sequence or its complement are aligned in overlapping fashion, excluding non - homologous segments of a continuous reference sequence such as a human chromosome sequence, the first single - stranded nucleic acid sequence overlaps the second single - stranded nucleic acid sequence. A nucleic acid that is wholly or partially double - stranded overlaps with another wholly or partially double - stranded nucleic acid when either strand of the former overlaps with a strand of the latter.
[0055] "Average", without limitation, refers to any statistical measure of central tendency including the mean value, median, and mode.
[0056] "Spread", without limitation, refers to any statistical measure of variability including variance, standard deviation, and interquartile range.
[0057] "Standard score", without limitation, refers to any statistical measure of distance from the mean including a normalized score or Z-score (number of standard deviations from the mean).
[0058] The normalized amount of tumor gene mutations refers to the standard score of tumor gene mutations compared to a control subject. This includes an indicator of the amount of tumor gene mutations in a test nucleic acid molecule sample adjusted to account for random variations between samples in factors affecting the detection of such mutations, such as the release of nucleic acids from cancer cells into body fluids and the recovery of nucleic acids from body fluids in an analyzable form.
[0059] A mutation refers to a variation from a known reference sequence and includes mutations such as SNVs, copy number variations / abnormalities, indels, and gene fusions. Mutations can be germline or somatic mutations. A preferred reference sequence for comparison is the wild-type genomic sequence of the species of the subject giving rise to the test sample, typically the human genome.
[0060] A variant may also be referred to as an allele. Variants typically occur at frequencies of 50% (0.5) or 100% (1) depending on whether the allele is heterozygous or homozygous. For example, germline variants are hereditary and typically have a frequency of 0.5 or 1. However, somatic variants are acquired variants and typically have a frequency of less than 0.5.
[0061] The major and minor alleles at a locus refer to nucleic acids having the locus, occupied by nucleotides of the reference sequence and variant nucleotides that differ from the reference sequence, respectively. The measurement at a locus can take the form of an allele frequency (AF) that measures the frequency at which an allele is observed in a sample.
[0062] The term "minor allele frequency" can refer to the frequency at which a minor allele (e.g., not the most frequent allele) occurs in a given nucleic acid population, e.g., a sample. Genetic variants with low minor allele frequencies may be relatively infrequent in a sample.
[0063] "Minor allele frequency (MAF)" refers to the proportion of DNA molecules having an allele change (e.g., a mutation) at a given genomic position in a given sample. The MAF of a somatic variant can be less than 0.5, 0.1, 0.05, or 0.01 of all somatic variants or alleles present at a given locus. For example, the MAF of a somatic variant is less than 0.05. Minor allele frequency can also be used interchangeably with "variant allele frequency".
[0064] The terms "neoplasm" and "tumor" are used interchangeably. These refer to abnormal proliferation of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. Malignant tumors are referred to as cancers or cancerous tumors.
[0065] The terms "tumor mutation burden (TMB)", "tumor mutational burden (TMB)", or "cancer gene mutation amount" are used interchangeably. These refer to the total number of mutations, such as somatic mutations, present in the sequenced portion of the tumor genome. TMB may refer to the number of coding, base substitution, and indel mutations per megabase of the tumor genome being investigated. These can be indicated to detect, evaluate, calculate, or predict sensitivity and / or resistance to cancer therapeutics or drugs, such as immune checkpoint inhibitors, antibodies. Tumors with higher levels of TMB can express more neoantigens, i.e., certain types of cancer-specific antigens, and may enable a stronger immune response and thus a more sustained response to immunotherapy. The immune system relies on a sufficient number of neoantigens to respond appropriately, and the number of somatic mutations may act as a proxy for determining the number of neoantigens in a tumor. TMB can be used to infer the strength of the immune response to drug treatment and the effectiveness of drug treatment in a subject. Germline and somatic variants can be bioinformatically identified to identify antigenic somatic variants as described in PCT / US2018 / 52087, which is incorporated herein by reference.
[0066] A threshold is a predetermined value used to characterize the values of the same parameter experimentally determined for various samples according to their relationships to that threshold.
[0067] The terms "processing", "calculating", and "comparing" are used interchangeably. This term can refer to determining a difference, such as a difference in numbers or sequences. For example, gene expression, copy number variation (CNV), indels, and / or single nucleotide variant (SNV) values or sequences can be processed.
[0068] "Cancer type" refers to a type or subtype defined, for example, by histopathology. Cancer types can be defined by any conventional criteria, such as cancers of the same tissue (e.g., blood cancers, CNS, brain cancer, lung cancer (small cell and non-small cell), skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, small intestine cancer, soft tissue cancer, thyroid cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer, solid cancer, heterogeneous cancer, homogeneous cancer), cancers of unknown primary origin, etc., and / or cancers of the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma) and / or cancer markers, such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors, and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether they are primary or secondary. Detailed Description I. General
[0069] This disclosure is predicated in part on the result that values of tumor gene mutation amounts from various samples are particularly easy to compare with each other or to a control reference substance by a normalization regime that takes into account the minor allele ratio of the mutations that are highly proportioned in the sample. Such an analysis can provide an indication of whether, when the amount of tumor gene mutations in a test sample lies on the distribution of the amount of tumor gene mutations in a control population, and thus, whether the individual providing the test sample may be suitable for immunotherapy for treating cancer. II. Determination and Normalization of Tumor Gene Mutation Amounts
[0070] The nucleic acids present in the sample can be processed and sequenced as further described below. By sequencing, the total number of mutations present in and detected in the sample, i.e., preferably, the minor allele, is statistically low such that it is unlikely to represent a sequencing artifact (e.g., p ≤ 0.05), and the total number of loci detected at a sufficient frequency in various nucleic acid molecules in the sample becomes apparent. The total number of determined mutations can indicate mutations present anywhere in the genome of the sample, or any ratio thereof, e.g., a particular chromosome, or a set of non - contiguous genomic segments, e.g., a set of segments known to have loci where mutations associated with cancer occur. The determined mutations can exclusively be mutations that result in a change in the sequence of the encoded protein, e.g., SNV, indel, fusion, or any kind of mutation that does not result in a change in the sequence of the encoded protein, e.g., copy number variation, copy number aberration. Regardless of the type of mutation determined, mutations that change the amino acid sequence of the encoded protein can be selected prior to the next processing. Mutations can include germline mutations, somatic mutations, or both.
[0071] Mutations encoding amino acid changes in the encoded protein tend to correlate more with eligibility for immunotherapy than other mutations, but sampling all mutations can be more correlated with the number of mutations that change the sequence in the encoded protein sequence than counting such mutations, which is directly due to the loss of some such mutations below the detection level. Thus, both approaches have advantages and both can be used.
[0072] Sequencing also provides the minor allele frequency of any or all of the detected mutations (or that subset thereof selected for subsequent processing). The minor allele frequency means the ratio of all sequenced nucleic acids in a sample that contains a locus of a mutation having a minor allele (distinct from the wild-type allele). Thus, the minor allele frequency can represent a number between 0 and 1. If more than one minor allele can occur at a locus, the minor allele frequency can be defined as the ratio of any of the minor alleles or the aggregate ratio of all or any subset of the minor alleles.
[0073] The minor allele frequency of the most frequently represented variant or the average minor allele frequency of the set of most frequently represented variants is used in the following normalization. If a set of the most frequently represented variants is used, the set can represent, for example, the top 2, 3, 5, or 10 most frequently represented variants.
[0074] The analysis described for the test sample can also be performed on a population of control samples to provide a data set for comparison. The control population can include, for example, samples from at least 10, 20, 25, 50, 100, 200, 250, 500, 1,000, 5,000, 10,000, 50,000, or more individuals. The control sample can be a sample from a subject having the same cancer type as the test sample. Each control sample is similarly analyzed for the total tumor gene mutation amount and the minor allele frequency of the most frequently represented mutation or set of mutations. Preferably, the minor allele frequency is determined in the same manner between the test sample and the control sample (i.e., based on the same most frequently represented mutation or the same set of most frequently represented mutations). The same is true for the case of mutation counting. For example, if mutations occurring anywhere in the genome are counted in the test sample, it is preferably the same in the control sample. Similarly, if only mutations affecting the coding protein sequence are counted for the test sample, it is also the same for the control sample.
[0075] Subsequently, the control samples can be sorted into bins according to the determined minor allele ratio. The bins may be of equal size (e.g., 0.05 - 0.1, 0.1 - 0.15, 0.15 - 0.2, 0.2 - 0.25), or the bins can be sized to equalize, for example, the number of control samples fitting into each bin. The bins may also be defined as a percentage of the overall variation in the minor allele ratio. For example, if the minor allele ratio of the most frequent variant(s) in the control population varies between 0.1 and 0.5, the bins can be defined at a percentage of that range (e.g., 5%, 10%, or 20% per bin). The average tumor gene mutation amount is then determined for the control samples in each bin. For example, if there are three control samples with gene mutation amounts of 3, 4, and 5 in the 0.1 - 0.15 bin, the average cancer gene mutation amount for that bin is 4. The standard deviation may be calculated for the cancer mutation values within the bin. Such a collection of bins can populate data for comparison with test samples of the same cancer type. The bins may have a width of 20% or less, 10% or less, or 5% or less.
[0076] The control population need only be analyzed once, and the data obtained can serve for comparison with any number of test samples. However, the control population may be supplemented with data from additional individuals having the same cancer type.
[0077] The same type of analysis may be performed in additional control populations having other types of cancer for comparison with test samples having these other forms of cancer.
[0078] Next, the number of tumor gene mutations measured in the test sample can be compared to the average number of mutations in the bins of the control sample defined by the range of minor allele frequencies that includes that of the test sample. For example, if the test sample has a minor allele frequency of 0.125 for the most frequently shown minor allele, a bin that includes minor allele frequencies from 0.1 to 0.15 can be selected for comparison. A simple numerical comparison (e.g., subtracting the average tumor gene mutation amount of the control sample from the tumor gene mutation amount of the test sample, or dividing the tumor gene mutation amount of the test sample by the average value of the control sample) indicates whether the tumor gene mutation amount of the sample is average, above average, or below average. For example, if the tumor gene mutation amount of the test sample is 5 and the average of the bin indicating minor allele conformity is 3, the tumor gene mutation amount of the test sample can be represented as 2 mutations more than the average or 5 / 3, i.e., 167% of the average.
[0079] However, a more quantitative comparison can be performed by calculating the Z-score. The Z-score is calculated by subtracting the average tumor gene mutation amount of the matched bin from the tumor gene mutation amount of the test sample and dividing the result by the standard deviation of the variation of the tumor gene mutations within the bin. The Z-score can be positive (higher than the average gene mutation amount), negative (lower than the average cancer gene mutation amount), or zero (average gene mutation amount). The magnitude of the Z-score (positive or negative) is an indication of the degree to which the test sample is above or below the average tumor gene mutation amount.
[0080] The normalized amount of tumor gene mutations (e.g., indicated by a Z-score) of a test sample from a subject provides an indication of the subject's suitability for immunotherapy. Generally, the greater the normalized amount of tumor gene mutations (which can be indicated by a higher positive Z-score), the more suitable the subject is for immunotherapy. Without being bound by any theory, more mutations indicate the presence of more neoepitopes that form non-self targets for immunotherapy. Conversely, the lower the normalized mutations (e.g., indicated by a negative Z-score), the lower the subject's suitability for immunotherapy.
[0081] One or more thresholds of the normalized amount of tumor gene mutations can be set or at least provide an indication (which can be used in combination with other factors) for determining whether a subject should receive or continue to receive immunotherapy or stop receiving it. For example, the threshold can be set such that subjects at or above the threshold receive or continue to receive immunotherapy, and subjects below the threshold do not receive or stop receiving immunotherapy. Alternatively, two thresholds can be set for subjects at or above the higher threshold who are receiving or continuing to receive immunotherapy, and for subjects at or below the lower threshold who are not receiving or have stopped receiving immunotherapy. Subjects between the thresholds can be evaluated by additional factors regarding whether the subject should receive or continue to receive immunotherapy.
[0082] The threshold can be determined empirically by observing the response to immunotherapy in subjects characterized by the normalized amount of tumor gene mutations in order to determine the threshold most correlated with a beneficial response to immunotherapy or the lack thereof. Alternatively, the threshold can be set at a predetermined point on the scale, for example, for subjects with a positive Z-score who are receiving or continuing to receive therapy, and / or for subjects with a negative Z-score who are not receiving or have discontinued immunotherapy. As another example, subjects with a Z-score greater than 1, 2, or 3 can receive or continue to receive therapy. As another example, subjects with a Z-score less than 1 may discontinue receiving therapy. As another example, subjects with a positive Z-score and subjects with the same cancer type having a Z-score, for example, at least 75%, 50%, 25%, 15%, 10%, or 5% of the maximum Z-score can receive or continue to receive immunotherapy, and other subjects cannot receive immunotherapy.
[0083] As described above, the normalized amount of tumor gene mutations can be used with or without other factors when determining whether immunotherapy is to be administered or continued. Such other factors can include, among other factors, the condition of the subject, the response of the subject to other therapies previously attempted, and the availability of other therapies not yet attempted with respect to the subject. Thus, not all subjects above the threshold will receive or continue to receive immunotherapy, or not all subjects below the threshold will receive or continue to receive immunotherapy, but generally a higher proportion of subjects with a normalized amount of tumor gene mutations above the threshold will receive or continue to receive immunotherapy than cases for subjects with a normalized amount of tumor gene mutations below the threshold. III. Immunotherapy
[0084] Immunotherapy refers to treatment with one or more agents that act to stimulate the immune system to kill cancer cells or at least inhibit the growth of cancer cells, preferably to reduce the size of cancer, reduce further growth of cancer, and / or eliminate cancer. Some such agents bind to targets present on cancer cells; some bind to targets present on cells that are immune cells and not cancer cells; some bind to targets present on both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of immune system pathways that maintain self-tolerance and regulate the duration and amplitude of the physiological immune response in peripheral tissues to minimize collateral tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Exemplary agents include antibodies to any of PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. Other exemplary agents include pro-inflammatory cytokines such as IL-1β, IL-6, and TNF-α. Other exemplary agents are, for example, T cells activated against a tumor by expression of a chimeric antigen that targets a tumor antigen derived from the T cell. In some embodiments, immunotherapy stimulates the immune system to attack tumor antigens that are distinguished from wild-type counterparts by the presence of a mutation. IV. Other Applications
[0085] Using the normalized amount of tumor gene mutations determined by the method of the present invention, the condition of a subject, particularly the presence of cancer, is diagnosed, the condition is characterized (e.g., classifying the stage of cancer or determining the heterogeneity of cancer), the response to treatment of the condition is monitored, and the prognostic risk of developing the condition or the next course of the condition is brought about. The normalized amount of tumor gene mutations can also be used to characterize specific forms of cancer. Cancer is often heterogeneous both in composition and stage classification. Gene profile data can enable the characterization of specific subtypes that may be important in the diagnosis or treatment of a particular cancer subtype. This information provides clues regarding the prognosis of a particular cancer type to the subject or the treating physician, enabling either the subject or the treating physician to adopt treatment options that follow the progression of the disease. Some cancers progress and become more aggressive and genetically unstable. Other cancers may remain benign, inactive or in a quiescent state. The normalized amount of tumor gene mutations can be useful in determining disease progression.
[0086] The normalized amount of tumor gene mutations can also be used in selecting treatments beyond immunotherapy and in determining the effectiveness of specific treatment options. If a treatment is successful, more cancer cells die and nucleic acids are removed, so a successful treatment option may initially increase the normalized amount of tumor gene mutations and subsequently decrease the normalized amount of tumor gene mutations as the cancer regresses or dies. A successful treatment can also make it possible to decrease the amount of tumor gene mutations and / or the minor allele ratio without causing an initial increase. Further, if it is observed that the cancer is in a remission state after treatment, the normalized amount of tumor gene mutations can be used to monitor residual disease or disease recurrence as indicated by the count of normalized mutations in body fluids. V. Computer-Implemented Execution
[0087] This method can be implemented on a computer such that any or all of the steps described in this specification or the appended claims, other than the wet chemistry steps, can be implemented on a suitable programmed computer. The computer can be a mainframe, personal computer, tablet, smartphone, cloud, online data storage, remote data storage, etc. The computer can be operated at one or more locations.
[0088] A computer program for analyzing a nucleic acid population includes code for performing any of the steps described in this specification or the appended claims, other than the wet chemistry steps; for example, code for receiving raw sequencing data, code for determining a nucleic acid sequence from such data, code for determining the number of mutations present in a given sequence, code for classifying mutations that affect the encoded protein or other sequence, code for determining the minor allele ratio of any of the mutations, and code for determining an indicator of the amount of cancer gene mutations in a test sample by comparing the number of mutations present in the test sample with the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio within bins of minor allele ratios including the minor allele ratio of the test sample, and optionally code for outputting the amount of gene mutations normalized using a related immunotherapy treatment.
[0089] This method can be implemented in a system (e.g., a data processing system) for analyzing a nucleic acid population. The system includes a processor, a system bus, and, for example, the following steps: receiving raw sequencing data, determining a nucleic acid sequence from such data, identifying mutations within the determined sequence, classifying mutations that affect the encoded protein sequence or otherwise, determining the minor allele ratio for any determined mutation, and determining the amount of cancer gene mutations in a test sample by comparing the number of mutations present in the test sample to the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio within a bin of minor allele ratios that includes the minor allele ratio of the test sample, and optionally outputting a normalized amount of gene mutations by treatment with immunotherapy, and main memory connected to each other to perform one or more steps described herein or in the appended claims, and may include auxiliary memory as needed. The memory of this system can also store control data from various populations having different cancer types. In any such population, the data can include bins of minor allele frequencies characterized by the number of mutations present in the subject, the minor allele frequency of some or all of such mutations, and the standard deviation of the average mutation frequency for the number of mutations in the sample that fall within the bin. This system can also include a display or printer for outputting results such as the amount of cancer gene mutations in the sample, expressed as a Z-score, and / or recommended future treatments such as administering or continuing immunotherapy. The system can also include a keyboard and / or a pointer for providing user input, among other accessories, such as defining the cancer type, the set of mutations for which the analysis is to be performed, or a set threshold. The system can also include a sequencing device connected to a memory for providing raw sequencing data.
[0090] The various steps of the method of the present invention utilize information and / or programs to create results that can be stored on a computer-readable medium (e.g., hard drive, auxiliary memory, external memory, server; database, portable memory device (e.g., CD-R, DVD, ZIP disk, flash memory card), etc.). For example, the information used in this method and the results created thereby that can be stored on a computer-readable medium include control data, reference sequences, raw sequencing data, sequenced nucleic acids, mutations, minor allele ratios, indicators of the amount of normalized gene mutations, such as Z-scores, thresholds, and immunotherapy treatment regimens related to the amount of gene mutations normalized against the thresholds in various cancer types, from various populations having the different cancer types described above.
[0091] The present disclosure also includes a kit containing instructions for use for providing an indicator of the amount of tumor gene mutations in a sample. The kit may include a machine-readable medium containing one or more programs that, when executed, perform the steps of the method. The kit may not include a physical machine-readable medium, but rather may access a cloud or online data storage that provides a platform through which a user can perform an analysis of the amount of tumor gene mutations in a sample.
[0092] The present disclosure can be implemented in hardware and / or software. For example, different aspects of the present disclosure can be implemented either on the client side logic or the server side logic. The present disclosure or its components can be embodied in a fixed media program component that contains logic instructions and / or data for causing the device to implement according to the present disclosure when loaded into a properly configured computing device. The fixed media containing the logic instructions can be delivered to a viewer of the fixed media for physically loading onto the viewer's computer, or the fixed media containing the logic instructions may be present on a remote server that the viewer accesses via a communication medium to download the program component.
[0093] The present disclosure provides a computer control system programmed to implement the method of the present disclosure. FIG. 2 shows a computer system 901 programmed to implement the method of the present disclosure or otherwise configured to implement the method of the present disclosure. The computer system 901 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 905, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 901 also includes a memory or memory location 910 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 915 (e.g., hard disk), a communication interface 920 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 925 such as cache memory, other memory, data storage, and / or an electronic display adapter. The memory 910, storage unit 915, interface 920, and peripheral devices 925 are
[0094] It communicates with the CPU 905 through a communication bus (solid line) such as a motherboard. The storage unit 915 can be a data storage unit (or data repository) for storing data. The computer system 901 can be operably connected to a computer network ("network") 930 with the assistance of the communication interface 920. The network 930 can be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. In some cases, the network 930 is a telecommunications and / or data network. The network 930 can be a local area network. The network 930 can include one or more computer servers that enable distributed computing such as cloud computing. In some cases, the network 930 can implement a peer-to-peer network that enables devices connected to the computer system 901 to function as clients or servers with the assistance of the computer system 901.
[0095] The CPU 905 can execute a sequence of machine-readable instructions that can be embodied in a program or software. The instructions can be stored at a memory location such as the memory 910. The instructions can be targeted at the CPU 905, which can then be programmed or otherwise configured to implement the method of the present disclosure. Examples of operations performed by the CPU 905 can include fetching, decoding, executing, and writing back.
[0096] The CPU 905 can be part of a circuit, for example, an integrated circuit. One or more other components of the system 901 can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0097] The storage unit 915 can store files such as drivers, libraries, and saved programs. The storage unit 915 can store user data, such as user preferences and user programs. In some cases, the computer system 901 may include one or more additional data storage units external to the computer system 901, such as being located on a remote server that communicates with the computer system 901 through an intranet or the Internet.
[0098] The computer system 901 can communicate with one or more remote computer systems through the network 930. For example, the computer system 901 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slates or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), phones, smartphones (e.g., Apple® iPhone®, Android®-capable devices, Blackberry®), or personal digital assistants. A user can access the computer system 901 via the network 930.
[0099] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 901, such as on the memory 910 or the electronic storage unit 915. The machine executable or machine readable code can be provided in the form of software. In use, the code can be executed by the processor 905. In some cases, the code can be retrieved from the storage unit 915 and stored on the memory 910 for easy access by the processor 905. In some situations, the electronic storage unit 915 may be excluded and the machine executable instructions may be stored on the memory 910.
[0100] The code can be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or can be compiled during runtime. The code can be provided in a programming language that can be selected to enable the code to be executed in a pre-compiled or just-in-compiled manner.
[0101] Aspects of the systems and methods provided herein, such as computer system 901, can be embodied in programming. Various aspects of the technology can typically be thought of as a "product" or "manufactured article" in the form of machine (or processor) executable code and / or associated data carried or embodied in a type of machine-readable medium. The machine executable code can be stored on an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory), or on a hard disk.
[0102] A "storage" medium can include any or all of the tangible memories such as computers, processors, or their associated modules (such as various semiconductor memories, tape drives, disk drives, etc.), which can always provide non-transitory storage of software programming. All or part of the software can sometimes be communicated through the Internet or various other telecommunications networks. For example, such communication can enable the loading of software from one computer or processor to another, such as from an administrative server or host computer to the computer platform of an application server. Thus, another type of medium that can have software elements includes optical waves, radio waves, and electromagnetic waves used over wired and optical terrestrial communication networks, as well as over various wireless links, such as the physical interface between local devices. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that hold software. As used herein, non-transitory tangible
[0103] Unless limited to "storage" media, terms such as computer or machine "readable media" refer to any medium involved in providing instructions for execution to a processor.
[0104] Thus, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media includes, for example, any optical or magnetic disk such as any storage device of a computer that can be used to implement, for example, a database shown in the drawings. Volatile storage media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cable; wires including a bus within a computer system, including copper wire and optical fiber. Carrier wave transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves such as those generated during high frequency (RF) infrared (IR) data communication. Thus, common forms of computer-readable media include, for example, floppy (registered trademark) disk, flexible disk, hard disk, magnetic tape, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punch cards, paper tape, any other physical storage media having a pattern of holes, RAM, ROM, PROM and EPROM, FLASH (registered trademark)-EPROM, any other memory chip or cartridge, carrier wave transporting data or instructions, a cable or link transporting such a carrier wave, or any other media that a computer can read programming code and / or data from. Many of these forms of computer-readable media can be involved in holding one or more sequences of one or more instructions for a processor for execution.
[0105] Computer system 901 can include, or be communicable with, an electronic display 935 that includes, for example, a user interface (UI) 940 for providing reports. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0106] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit 905. VI. General Features of the Method 1. Sample
[0107] The sample can be any biological sample isolated from a subject. Examples of samples include body tissues such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (leucocytes), endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid, fluid in the intercellular space including gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions such as plasma and serum, and urine. Such samples contain nucleic acids extracted from tumors. The nucleic acids may include DNA and RNA and may be in double-stranded and / or single-stranded form. The sample may be in the form isolated from the subject or may be subjected to further processing to remove or add components such as cells, concentrate one component relative to another, or convert one form of nucleic acid to another, such as converting RNA to DNA or single-stranded nucleic acid to double-stranded. Thus, for example, the body fluid for analysis is plasma or serum containing cell-free nucleic acids, such as cell-free DNA (cfDNA).
[0108] The volume of the body fluid varies depending on the desired read depth for the sequenced region. Exemplary volumes are 0.4 - 40 ml, 5 - 20 ml, 10 - 20 ml. For example, the volume can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, or 40 ml. The volume of the sampled plasma can be 5 to 20 ml.
[0109] The sample can contain various amounts of nucleic acids containing genomic equivalents. For example, a sample of about 30 ng of DNA is about 10,000 (10 4) can contain a haploid human genome equivalent, and in the case of cfDNA, can contain about 200 billion (2×10 4 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents and, in the case of cfDNA, can contain about 600 billion individual molecules.
[0110] Samples can include nucleic acids from different sources, such as cells and cell-free sources. Samples can include nucleic acids having mutations. For example, a sample can include DNA having germline mutations and / or somatic mutations. A sample can include DNA having cancer-related mutations (e.g., cancer-related somatic mutations).
[0111] Exemplary amounts of cell-free nucleic acids in a sample prior to amplification range from about 1 fg to about 1 μg, such as from 1 pg to 200 ng, from 1 ng to 100 ng, from 10 ng to 1000 ng. For example, this amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. This amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. This amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. This method can include obtaining from 1 femtogram (fg) to 200 ng.
[0112] A cell-free nucleic acid sample refers to a sample containing cell-free nucleic acids. Cell-free nucleic acids are nucleic acids that are not contained within cells or are otherwise not bound to cells, i.e., nucleic acids remaining in a sample from which intact cells have been removed. Cell-free nucleic acids can refer to all unencapsulated nucleic acids sourced from a subject's body fluids (e.g., blood, urine, CSF, etc.). Examples of cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and their hybrids (including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or any fragments thereof). Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids. Cell-free nucleic acids can be released into body fluids via processes of secretion or cell death, such as necrosis and apoptosis of cells. Some cell-free nucleic acids are released from cancer cells into body fluids (e.g., circulating tumor DNA (ctDNA)). Others are released from healthy cells. CtDNA can be fragmented DNA derived from unencapsulated tumors. Cell-free fetal DNA (cffDNA) is fetal DNA that circulates freely in the mother's bloodstream.
[0113] Cell-free nucleic acids or proteins related thereto can have one or more epigenetic modifications. For example, cell-free nucleic acids can be acetylated, 5-methylated, ubiquitinated, phosphorylated, SUMOylated, ribosylated, and / or citrullinated.
[0114] Cell-free nucleic acids have an exemplary size distribution of about 100 to 500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of the molecules, having a mode of about 168 nucleotides and a second minor peak in the range of 240 to 440 nucleotides in humans. Cell-free nucleic acids can be about 160 to about 180 nucleotides, or about 320 to about 360 nucleotides, or about 440 to about 480 nucleotides.
[0115] Cell-free nucleic acids can be isolated from body fluids via a separation step in which the cell-free nucleic acids found in solution are separated from intact cells and other insoluble constituents of the body fluid. The separation step can include techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid can be lysed and the cell-free nucleic acids and intracellular nucleic acids are processed together. Generally, after the addition of buffer and washing steps, the cell-free nucleic acids can be precipitated with alcohol. Further purification steps such as silica-based columns to remove contaminants or salts may be used. For example, non-specific bulk carrier nucleic acids can be added throughout the reaction to optimize certain aspects of the procedure such as yield.
[0116] After such treatment, the sample can contain nucleic acids in various forms including double-stranded DNA, single-stranded DNA and / or single-stranded RNA. If desired, the single-stranded DNA and / or single-stranded RNA may be converted to the double-stranded form so as to be included in the subsequent processing and analysis steps. 2. Amplification
[0117] Sample nucleic acids adjacent to the adapter can be amplified by PCR and typically other amplification methods that prime up to the primer binding site of the adapter adjacent to the DNA molecule amplified from the bound primer. The amplification method can include cycles of annealing resulting from extension, denaturation and thermocycling and can be isothermal as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication.
[0118] Using conventional nucleic acid amplification methods, one or more amplifications can be applied to introduce barcodes into nucleic acid molecules. The amplification may be performed in one or more reaction mixtures. The barcodes may be introduced simultaneously or in any sequential order. The barcodes may be introduced before and / or after sequence capture. In some cases, only the barcodes that label individual nucleic acid molecules are introduced before probe capture, while the barcodes that label the sample are introduced after sequence capture. In some cases, both the barcodes that label individual nucleic acid molecules and the barcodes that label the sample are introduced before probe capture. In some cases, the barcodes that label the sample are introduced after sequence capture. Typically, sequence capture involves introducing single-stranded nucleic acid molecules complementary to the target sequences, e.g., genomic regions and the coding sequences of mutations in such regions are relevant to cancer types. Typically, amplification results in a plurality of nucleic acid amplicons that are non-uniquely or uniquely barcoded with individual nucleic acids and / or barcodes that label the sample in the size range of 200 nt to 700 nt, 250 nt to 350 nt, or 320 nt to 550 nt. In some embodiments, the amplicon has a size of about 300 nt. In some embodiments, the amplicon has a size of about 500 nt. 3. Barcodes
[0119] Barcodes can be introduced into or otherwise ligated to adapters by, among other methods, chemical synthesis, ligation, and overlap extension PCR. Generally, the assignment of unique or non-unique barcodes in a reaction follows the methods and systems described in U.S. Patent Application Nos. 20010053519, 20110160078, and U.S. Patents 6,582,908, 7,537,898, and US9,598,731.
[0120] Barcodes can be ligated to sample nucleic acids either randomly or non-randomly. In some cases, these are introduced at an expected ratio of barcodes (e.g., combinations of unique or non-unique barcodes) to microwells. For example, barcodes can be loaded such that about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or more than 1,000,000,000 barcodes are loaded per genomic sample. In some cases, barcodes can be loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 barcodes are loaded per genomic sample. In some cases, the average number of barcodes loaded per sample genome is less than or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 barcodes per genomic sample. Barcodes can be either unique or non-unique.
[0121] A preferred format uses 20 - 50 different barcodes, e.g., 400 - 2500 barcodes, ligated to both ends of the target molecule creating 20 - 50×20 - 50 tags. The number of such barcodes is sufficient in that different molecules having the same start and stop points are likely (e.g., at least 94%, 99.5%, 99.99%, 99.999%) to receive different combinations of tags.
[0122] In some cases, the barcode may be an oligonucleotide of a predetermined or random or semi-random sequence. In other cases, multiple barcodes may be used such that the barcodes are not necessarily unique to each other among the multiple barcodes. In this example, the barcode may be attached to an individual molecule (e.g., by ligation or PCR amplification), resulting in the creation of a unique sequence that may be individually tracked by the combination of the barcode and the sequence to which it may be attached. As described herein, the detection of non-unique barcodes in combination with the sequence data at the beginning (start) and end (stop) portions of the sequence read may enable the assignment of unique identity to a particular molecule. The unique identity of such a molecule may be assigned using the length, number of base pairs, of an individual sequence read. As described herein, a fragment derived from a single strand of a nucleic acid to which a unique identity has been assigned may thereby enable the subsequent identification of fragments derived from the parental strand and / or complementary strand. 4. Sequencing
[0123] Regardless of the presence or absence of pre-amplification, the sample nucleic acid adjacent to the adapter can be the target of sequencing as needed. Examples of sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), ultra-parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent or nanopore platforms. The sequencing reaction can be carried out in various sample processing units, and the sample processing unit may be a multi-lane, multi-channel, multi-well, or other means of processing multiple sample sets substantially simultaneously. The sample processing unit may include multiple sample chambers that enable the processing of multiple runs simultaneously.
[0124] The sequencing reaction can be carried out with one or more fragments of a type known to contain markers for cancer or other diseases. The sequencing reaction can also be carried out with any nucleic acid fragment present in the sample. The sequencing reaction can provide a sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome. In other cases, the sequence coverage of the genome can be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100%.
[0125] Sequencing reactions that occur simultaneously may be carried out using multiplex sequencing. In some cases, cell-free polynucleotides can be sequenced with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, cell-free polynucleotides can be sequenced with less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. The sequencing reactions can be carried out sequentially or simultaneously. The next data analysis can be carried out on all or part of the sequencing reactions. In some cases, the data analysis can be carried out with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, the data analysis can be carried out with less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. An exemplary read depth is 1000 - 50000 reads per locus (base).
[0126] In one approach, the nucleic acid population is prepared for sequencing by enzymatic blunt-ending of double-stranded nucleic acids having single-stranded overhangs at one or both ends. This population can be treated with a protein having 5'-3' polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Exemplary proteins are Klenow large fragment and T4 polymerase. For 5' overhangs, the protein extends the recessed 3' end of the corresponding strand until it is flush with the 5' end, generating a blunt end. For 3' overhangs, the protein digests from the 3' end to the 5' end of the corresponding strand and sometimes beyond. If digestion proceeds beyond the 5' end of the corresponding strand, the gap can be filled by polymerase activity in the case of a 5' overhang. Blunt-ending of double-stranded nucleic acids facilitates adapter attachment and subsequent amplification.
[0127] The nucleic acid population can be subject to additional processing such as conversion of single-stranded nucleic acids to double-stranded and / or conversion of RNA to DNA. These forms of nucleic acids can also be ligated to adapters and amplified.
[0128] With or without pre-amplification, the nucleic acid subject to blunt-ending as described above, and optionally other nucleic acids in the sample, can be sequenced to generate sequenced nucleic acids. The sequenced nucleic acids can refer to the sequence of the nucleic acids or the sequence of the nucleic acids whose sequence has been determined. Sequencing can be performed to provide sequence data for individual nucleic acid molecules in the sample that directly or indirectly derive from the consensus sequence of the amplification products of the individual nucleic acid molecules in the sample.
[0129] In one method, a double-stranded nucleic acid having single-stranded overhangs in the sample after blunt-ending is ligated to both ends of an adapter containing a barcode, and the nucleic acid sequence and the inline barcode introduced by the adapter are determined by sequencing. The blunt-ended DNA molecule can be blunt-end ligated to the blunt end of at least a partial double-stranded adapter (e.g., a Y-shaped or bell-shaped adapter). Alternatively, the blunt ends of the sample nucleic acid and the adapter can be tailed with complementary nucleotides to facilitate ligation.
[0130] A sufficient number of adapters can be contacted with the sample such that the probability that any two instances of the same nucleic acid receive the same combination of adapter barcodes from the adapters ligated to both ends is low (e.g., less than 1 or 0.1%). By using adapters in this manner, it becomes possible to identify a family of nucleic acid sequences that have the same starting and stopping points in the reference nucleic acid and are ligated to the same combination of barcodes. Such a family represents the sequences of the amplification products of the nucleic acids in the sample prior to amplification. The sequences of the family members can be compiled to derive the consensus nucleotide or the complete consensus sequence for the nucleic acid molecules in the original sample when modified by blunt-ending and adapter attachment. In other words, it is determined that the nucleotide occupying a particular position in the nucleic acid in the sample is the consensus of the nucleotides occupying the corresponding positions in the sequences of the family members. A family can include the sequences of one or both strands of the double-stranded nucleic acid. When the members of the family include the sequences of both strands derived from the double-stranded nucleic acid, the single-stranded sequences are converted to their complements to compile all the sequences from which the consensus nucleotide or sequence is derived. Some families contain only the sequence of a single member. In this case, this sequence can be obtained as the sequence of the nucleic acid in the sample prior to amplification. Alternatively, families containing only the sequence of a single member can be excluded from subsequent analysis.
[0131] Nucleotide variations in a sequenced nucleic acid can be determined by comparing the sequenced nucleic acid to a reference sequence. The reference sequence is often a known sequence, such as a known full genomic sequence or partial genomic sequence from an object, or the full genomic sequence of a human subject. The reference sequence may be hG19. The sequenced nucleic acid can represent, as described above, the sequence directly determined for the nucleic acid in the sample or the consensus of the sequences of amplification products of such nucleic acids. The comparison can be performed at one or more specified positions of the reference sequence. A subset of the sequenced nucleic acid containing the positions corresponding to the specified positions of the reference sequence can be identified when the sequences are maximally aligned. Within such a subset, if present, which of the sequenced nucleic acids contain a nucleotide variation at the specified position and, if necessary, if present, which contain the reference nucleotide (i.e., the same as in the reference sequence) can be determined. If the number of sequenced nucleic acids in the subset containing the nucleotide variation exceeds a threshold, the variant nucleotide can be called at the specified position. The threshold can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequenced nucleic acids within the subset containing the nucleotide variant, or a ratio, such as at least 0.5, and 1, 2, 3, 4, 5, 10, 15, or 20 of the sequenced nucleic acids within the subset, among other possibilities, contain the nucleotide variant. The comparison can be repeated for any specified position of interest in the reference sequence. The comparison can also be performed for specified positions occupying at least 20, 100, 200, or 300 consecutive positions in the reference sequence, such as 20 to 500, or 50 to 300 consecutive positions.
[0132] All patent applications, websites, other publications, accession numbers, etc., cited above or below are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference. Where different sequence versions are associated with accession numbers at different times, the version associated with the accession number of the effective filing date of this application is meant. The effective filing date means, where applicable, the earlier of the actual filing date or the filing date of the priority application that refers to the accession number. Similarly, where different versions of publications, websites, etc. are published at different times, unless otherwise indicated, the version published closest to the effective filing date of this application is meant. Any configuration, step, element, embodiment, or aspect of this disclosure can be used in combination with any other, unless specifically indicated otherwise. This disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, but it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims.
Example
[0133] This example determines the distribution of Z - scores in individuals with various cancer types. Blood collection, transportation, and plasma isolation
[0134] All cfDNA extraction, processing, and sequencing were performed in a CLIA - certified, CAP - accredited laboratory. Briefly, for clinical samples, plasma was isolated from 10 ml of whole blood collected in a cell - free blood collection tube by double centrifugation, cfDNA was extracted therefrom, labeled with non - random barcodes, 5 - 30 ng was used to prepare a sequencing library, which was then enriched by hybrid capture, pooled, and sequenced by paired - end synthesis (NextSeq 500 and / or HiSeq 2500, Illumina, Inc.). Artificial analytical samples were prepared using cfDNA similarly prepared from healthy donors and cfDNA isolated as described above from the culture supernatants of model cell lines and continuously size-selected using Agencourt Ampure XP beads (Beckman Coulter, Inc.) until no detectable gDNA remained. Bioinformatics analysis and variant detection
[0135] All variant detection analyses were performed using the locked clinical Guardant360 bioinformatics pipeline, and no changes were reported by post hoc analysis. All decision thresholds were determined using an independent training cohort that was predictively locked and applied to all validation and clinical samples. As previously described [PMID 26474073], basecall files created by Illumina's RTA software (v2.12) were demultiplexed using bcl2fastq (v2.19) and processed with a custom pipeline for molecular barcode detection, sequencing adapter trimming, and base quality trimming (discarding bases with Q20 or less at the ends of reads). The processed reads were then aligned against BWA-MEM [Li et al. 2013 arXiv:1303.3997v2], and this was used to construct a duplex consensus representation of the original unique cfDNA molecules using both hG19 and the putative barcode and read start / stop positions. SNVs were detected by comparing the characteristics of reads and consensus molecules against sequencing platform and position-specific reference error noise profiles independently determined for each position on the panel by sequencing a training set of 62 healthy donors on both the NextSeq 500 and HiSeq 2500. The observed SNV error profiles at positions were used to define the calling cutoffs for SNV detection in terms of the number and characteristics of variant molecules, which varied by position but were the most common unique molecules and corresponded to a detection limit of allele ratio -0.04% in an average sample (unique molecule coverage of 5,000). To detect indels, a generative background noise model was constructed to account for PCR artifacts that frequently occur in homopolymer or repeat contexts, which are strand-specific and anticipate late PCR errors.Next, detection was determined by the likelihood ratio score for the observed feature variant molecule support relative to the background noise distribution. The reporting threshold, when determined by performance in training samples, is event-specific but is most commonly two or more unique molecules for clinically actionable indels and corresponds to a detection limit of allele ratio -0.02% in the average sample. Fusion events were detected by merging overlapping paired-end reads, forming a representation of the sequenced cfDNA molecules, which were then aligned and mapped to the initial unique cfDNA molecules based on barcode addition and alignment information including soft clipping. Soft clip reads were analyzed using directionality and breakpoint proximity to identify clusters of molecules indicative of candidate fusion events, which were then used to construct a fusion reference to which the reads soft clipped by the first pass aligner were realigned. The specific reporting threshold was determined by retrospective and training set analysis and was generally one or more unique molecules after realignment meeting quality requirements, which corresponds to a detection limit of allele ratio -0.04% in the average sample. To detect CNA, probe-level unique molecule coverage was normalized against overall unique molecule throughput, probe efficiency, GC content, and signal saturation and aggregated firmly at the gene level. CNA determination was based on a determination threshold established in the training set for both the deviation of the absolute copy number from the diploid baseline per sample and the deviation from the baseline variation of the signal normalized at the probe level in the context of the background variation within the diploid baseline of each sample itself. The normalized tumor amount per sample was determined by normalization against the expected amount of genetic variation for the tumor type and ctDNA ratio and reported as a Z-score.
[0136] FIG. 1 plots the Z-score against the amount of cancer gene mutations in samples from different individuals having one of the cancer types shown on the X-axis. The distribution of the Z-scores varies for different cancer types, but generally, it is typically asymmetric with a mode value less than zero, although a small number of individuals show highly positive Z-scores. The present invention provides, for example, the following items. (Item 1) A method for providing an indicator of the amount of tumor gene mutations in a test sample of cell-free nucleic acid from a subject having a cancer type or a symptom of a cancer type, comprising: (a) determining the number of mutations present in the test sample of cell-free nucleic acid and the minor allele ratio based on one or more mutations most frequently shown in the test sample of cell-free nucleic acid; and (b) normalizing the number of mutations present in the sample to the number of mutations present in a control sample from another subject having the same cancer type and to the minor allele ratio within a bin of minor allele ratios including the minor allele ratio of the test sample to determine an indicator of the amount of cancer gene mutations in the test sample. A method comprising the above. (Item 2) The method according to Item 1, wherein the number of mutations present in the control sample is the average. (Item 3) The method according to any one of the preceding items, wherein the bin has a width of 20% or less, 10% or less, or 5% or less. (Item 4) The method according to any one of the preceding items, further comprising determining whether the number of mutations present in the sample exceeds a threshold value, wherein the threshold value is set to indicate a subject having a tendency to respond positively to immunotherapy. (Item 5) The method according to any one of the preceding items, wherein the normalizing step includes dividing the number of mutations in the test sample by the average number of mutations in the control sample. (Item 6) The method according to any one of items 1 to 4, wherein the normalizing step comprises subtracting the average number of mutations in the control samples in the bin from the determined number of mutations in the test sample of cell-free nucleic acid. (Item 7) The method according to item 6, further comprising the step of calculating a Z-score by dividing the number of mutations in the test sample of cell-free nucleic acid, from which the average number of mutations present in the control sample has been subtracted, by the standard deviation of the number of mutations present in the control sample. (Item 8) The method according to item 7, wherein the average is the mean value. (Item 9) The method according to any one of items 1 to 4, wherein the normalizing step comprises determining the average and spread of the number of mutations in at least 10, 50, 100 or 500 control samples, determining the standard score of the deviation from the average in the test sample, and determining whether the standard score exceeds a threshold number. (Item 10) The method according to item 9, wherein the average is the mean value, median or mode. (Item 11) The method according to item 9, wherein the spread is represented as variance, standard deviation, or interquartile range. (Item 12) The method according to item 9, wherein the standard score of the deviation is the Z-score. (Item 13) The method according to any one of items 1 to 4, wherein the normalizing step further comprises dividing the determined number of mutations in the test sample of cell-free nucleic acid by the average number of mutations present in the control samples in the same bin. (Item 14) The method according to any one of the preceding items, wherein the normalizing step is performed on a computer programmed to store values of the number of mutations present in a plurality of bins of minor allele frequencies. (Item 15) The method according to item 13, wherein the stored values are the mean value and standard deviation of the number of mutations present in each of the plurality of bins. (Item 16) The method according to any one of the preceding items, comprising the step of determining a standard score of the amount of tumor gene mutation in the subject and whether the standard score exceeds a threshold for a control subject, which is consistent with responsiveness to immunotherapy. (Item 17) The method according to item 1, wherein (a) comprises determining the sequence of cell-free nucleic acid molecules in the test sample and comparing the obtained sequence with a corresponding reference sequence to identify the number of mutations and the minor allele ratio present in the sample. (Item 18) The method according to item 17, wherein the reference sequence is hG19 or hG38. (Item 19) The method according to item 7, wherein the control sample comprises at least 25, 50, 100, 200 or 500 control samples. (Item 20) The method according to item 15, wherein at least 50,000, 100,000 or 150,000 nucleotides are sequenced in the segment of the cell-free nucleic acid. (Item 21) The method according to item 1, wherein step (a) comprises determining the presence or absence of a panel of predetermined mutations known to occur in the type of cancer present or suspected to be present in the sample, and optionally, the mutations are somatic mutations that affect the sequence of the encoded protein. (Item 22) The method according to item 1, wherein step (a) comprises ligating an adapter to the cell-free nucleic acid, amplifying the cell-free nucleic acid from a primer that binds to the adapter, and sequencing the amplified nucleic acid. (Item 23) The method according to item 1, wherein the step of sequencing is bridge amplification sequencing, pyrosequencing, ion semiconductor sequencing, paired-end sequencing, sequencing by ligation or single molecule real-time sequencing. (Item 24) A method of treating a subject, comprising: (a) determining the number of mutations present in a test sample of cell-free nucleic acid and the minor allele ratio based on one or more mutations most frequently represented in said test sample of cell-free nucleic acid; (b) normalizing the number of mutations present in said sample to the number of mutations present in a control sample from another subject having the same cancer type and to the minor allele ratio within a bin of minor allele ratios that includes the minor allele ratio of said test sample, to determine an indicator of the amount of cancer gene mutations in said test sample; (c) administering immunotherapy to said subject if said indicator of the amount of tumor gene mutations exceeds a threshold. A method comprising the steps above. (Item 25) The method according to item 24, which is performed in a plurality of subjects to determine an indicator of the amount of tumor gene mutations in each subject, and in which a higher proportion of subjects having an indicator of the amount of cancer gene mutations exceeding the threshold than subjects having an indicator of the amount of tumor gene mutations below the threshold receive immunotherapy for said cancer. (Item 26) The method according to item 25, in which all subjects in which said indicator exceeds a first threshold receive immunotherapy and all subjects in which said indicator is below a second threshold do not receive immunotherapy. (Item 27) The method according to item 24 or 25, in which said indicator is a Z-score. (Item 28) The method according to item 24 or 25, in which said immunotherapy comprises administration of a checkpoint inhibitor antibody. (Item 29) The method according to item 24 or 25, in which said immunotherapy comprises administration of an antibody against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. (Item 30) The method according to item 24 or 25, in which said immunotherapy comprises administration of a pro-inflammatory cytokine. (Item 31) The method according to item 24 or 25, wherein the immunotherapy comprises administration of T cells against the cancer type. (Item 32) The method according to any one of the preceding items, wherein the cancer type is a solid tumor. (Item 33) The method according to any one of items 1 to 32, wherein the cancer type is renal cancer, mesothelioma, soft tissue cancer, primary CNS cancer, thyroid cancer, liver cancer, prostate cancer, pancreatic cancer, CUP, neuroendocrine cancer, NSCLC, gastroesophageal cancer, head and neck cancer, SCLC, breast cancer, melanoma, cholangiocarcinoma, gynecological cancer, colorectal cancer or urothelial cancer. (Item 34) The method according to any one of items 1 to 32, wherein the cancer type is a hematopoietic malignancy. (Item 35) The method according to item 33, wherein the cancer type is leukemia or lymphoma. (Item 36) A method of treating a subject having cancer, comprising the step of administering an immunotherapeutic agent to the subject, wherein the subject (a) determining the number of mutations present in a test sample of cell-free nucleic acid from the subject and the minor allele ratio for the mutation most frequently shown in the test sample of cell-free nucleic acid; (b) normalizing the number of mutations present in the sample to the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio within a bin of minor allele ratios including the minor allele ratio of the test sample, to determine an indicator of the amount of tumor gene mutations in the sample of the subject, wherein the subject is determined to have an amount of tumor gene mutations exceeding a threshold; A method specified for immunotherapy from the indicator of the amount of cancer gene mutations of the subject determined by (Item 37) A method of treating a subject having cancer, comprising for each subject, receiving an indicator of the amount of tumor gene mutations in a sample from the subject, wherein the indicator of the amount of tumor gene mutations is (a) Determining the number of mutations present in a test sample of cell-free nucleic acid from the subject, and the minor allele ratio for the mutation most frequently shown in the test sample of cell-free nucleic acid in the test sample; (b) Normalizing the number of mutations present in the sample to the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio within a bin of minor allele ratios including the minor allele ratio of the test sample, to determine the indicator of the amount of tumor gene mutations in the sample of the subject; (c) Administering immunotherapy to at least one subject determined to have an amount of tumor gene mutations exceeding a threshold; A method determined thereby. (Item 38) A communication interface that receives sequencing reads obtained by sequencing cell-free nucleic acid in a test sample through a communication network, and A computer that communicates with the communication interface, having one or more computer processors, and when executed by the one or more computer processors, (a) Receiving the sequencing reads obtained by a nucleic acid sequencer through the communication network; (b) Determining the number of mutations present in the sequencing reads from the test sample, and the minor allele ratio based on one or more mutations most frequently shown in the sequencing reads from the test sample; (c) Normalizing the number of mutations present in the test sample to the number of mutations present in a control sample from another subject having the same cancer type and the minor allele ratio within a bin of minor allele ratios including the minor allele ratio of the test sample, to determine an indicator of the amount of cancer gene mutations in the test sample; A computer-readable medium including machine-executable code for implementing a method including the above steps, and A system including the above. (Item 39) The system according to item 38, wherein the nucleic acid sequencer sequences a sequencing library obtained from cell-free DNA molecules derived from a subject, and the sequencing library includes the cell-free DNA molecules and adapters containing barcodes. (Item 40) The system according to item 38, wherein the nucleic acid sequencer performs sequencing by synthesis on the sequencing library to obtain the sequencing reads. (Item 41) The system according to item 38, wherein the nucleic acid sequencer performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation or sequencing by hybridization on the sequencing library to obtain the sequencing reads. (Item 42) The system according to item 38, wherein the nucleic acid sequencer uses a clonal single molecule array derived from the sequencing library to obtain the sequencing reads. (Item 43) The system according to item 38, wherein the nucleic acid sequencer includes a chip having an array of microwells for sequencing the sequencing library to obtain the sequencing reads. (Item 44) The system according to item 38, wherein the computer-readable medium includes a memory, a hard drive or a computer server. (Item 45) The system according to item 38, wherein the communication network includes a long-distance communication network, the Internet, an extranet, or an intranet. (Item 46) The system according to item 38, wherein the communication network includes one or more computer servers capable of distributed computing. (Item 47) The system according to item 46, wherein the distributed computing is cloud computing. (Item 48) The system according to item 38, wherein the computer is installed on a computer server remotely located from the nucleic acid sequencer. (Item 49) The system according to item 39, wherein the sequencing library further includes a sample barcode for identifying a sample from one or more samples. (Item 50) The system according to item 38, further including an electronic display that communicates with the computer through a network and includes a user interface for displaying the results when (a) to (c) are performed. (Item 51) The system according to item 50, wherein the user interface is a graphical user interface (GUI) or a web-based user interface. (Item 52) The system according to item 51, wherein the electronic display is present in a personal computer. (Item 53) The system according to item 51, wherein the electronic display is present in an Internet-enabled computer. (Item 54) The system according to item 53, wherein the Internet-enabled computer is installed at a location remote from the computer.
Claims
1. A method for obtaining a Z-score in a test sample of cell-free nucleic acids from a subject having a cancer type or a sign of a cancer type as an indicator of whether the subject has a tendency to respond positively to treatment, comprising: (a) determining the number of mutations present in the test sample of cell-free nucleic acids and the minor allele ratio for one or more mutations most frequently shown in the test sample of cell-free nucleic acids; (b) normalizing the number of mutations present in the sample to the number of mutations present in control samples from other subjects having the same cancer type within a bin of control samples, wherein the control samples have a minor allele ratio within a certain range including the minor allele ratio of the test sample, and the normalizing step includes subtracting the mean, median or mode of the number of mutations in the control samples within the bin from the determined number of mutations in the test sample of cell-free nucleic acids, and further dividing the number of mutations in the test sample of cell-free nucleic acids minus the mean, median or mode of the number of mutations in the control samples by the standard deviation of the number of mutations in the control samples to calculate a Z-score; (c) determining whether the Z-score is at, above or below one or more thresholds, wherein the one or more thresholds are set to indicate whether the subject has a tendency to respond positively to the treatment or not, the bin of control samples having a minor allele ratio within a certain range including the minor allele ratio of the test sample is selected from a set of bins of control samples, each of the set of bins of control samples has a minor allele ratio within a certain range of control samples, each of the bins of control samples is defined as a proportion of the entire range of minor allele ratios of the control samples, and the minor allele ratio refers to the ratio of DNA molecules having mutations at a given genomic position in a given sample. A method comprising the above steps.
2. The method according to claim 1, wherein the vial has a width of 20% or less, 10% or less, or 5% or less of the entire range of the minor allele ratio of the control sample.
3. One threshold value is set, and a subject whose Z-score is at or above the threshold value tends to respond positively to the treatment, and a subject whose Z-score is below the threshold value does not tend to respond positively to the treatment. The method according to claim 1 or 2.
4. A subject with a positive Z-score tends to respond positively to the treatment, and a subject with a negative Z-score does not tend to respond positively to the treatment. The method according to claim 1 or 2.
5. A subject whose Z-score is at or above the first threshold value tends to respond positively to the treatment, and a subject whose Z-score is at or below the second threshold value does not tend to respond positively to the treatment. The method according to claim 1 or 2.
6. A subject whose Z-score exceeds 1, 2, or 3 tends to respond positively to the treatment. The method according to claim 1 or 2.
7. A subject whose Z-score is less than 1 does not tend to respond positively to the treatment. The method according to claim 1 or 2.
8. The subject who tends to respond positively to the treatment is a candidate for receiving the treatment. The method according to any one of claims 1 to 6.
9. The treatment is immunotherapy. The method according to any one of claims 1 to 8.
10. The immunotherapy includes administration of a checkpoint inhibitor antibody. The method according to claim 9.
11. The method according to claim 9, wherein the immunotherapy comprises an antibody against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40, a pro-inflammatory cytokine, and / or administration of T cells against the cancer type.
12. The method according to any one of claims 1 to 11, wherein the control sample used in the normalizing step described in (b) comprises at least 25, 50, 100, 200, or 500 control samples.
13. The method according to any one of claims 1 to 12, wherein the normalizing step is performed on a computer programmed to store the values of the number of mutations present in a plurality of bins of minor allele ratios.
14. The method according to claim 13, wherein the stored values are the mean value and standard deviation of the number of mutations present in each of the plurality of bins.
15. The method according to claim 13 or 14, wherein at least 50,000, 100,000 or 150,000 nucleotides are sequenced in the segment of the cell-free nucleic acid.
16. (a) is (i) determining the sequence of cell-free nucleic acid molecules in the test sample and comparing the obtained sequence with a corresponding reference sequence to identify the number of mutations and the minor allele ratio present in the sample; (ii) determining the presence or absence of a panel of predetermined mutations known to occur in a type of cancer present or suspected to be present in the sample; or (iii) ligating an adapter to the cell-free nucleic acid, amplifying the cell-free nucleic acid from a primer that binds to the adapter, and sequencing the amplified nucleic acid The method according to any one of claims 1 to 15, comprising.
17. The method according to claim 16, wherein the reference array described in (a)(i) is derived from hG19 or hG38.
18. The method according to claim 16 or 17, wherein the predetermined mutation described in (a)(ii) is a somatic mutation that affects the sequence of the encoded protein.
19. The method according to any one of claims 16 to 18, wherein the sequencing is bridge amplification sequencing, pyrosequencing, ion semiconductor sequencing, paired-end sequencing, sequencing by ligation, or single molecule real-time sequencing.
20. The cancer type is (a) solid tumor; (b) renal cancer, mesothelioma, soft tissue cancer, primary CNS cancer, thyroid cancer, liver cancer, prostate cancer, pancreatic cancer, CUP, neuroendocrine cancer, NSCLC, gastroesophageal cancer, head and neck cancer, SCLC, breast cancer, melanoma, cholangiocarcinoma, gynecological cancer, colorectal cancer, or urothelial cancer; (c) leukemia or lymphoma; or (d) hematopoietic malignancy The method according to any one of claims 1 to 19.
21. A system comprising A communication interface that receives, through a communication network, sequencing reads generated by sequencing cell-free nucleic acids in a test sample from a subject having a cancer type or a sign of a cancer type, and A computer that communicates with the communication interface, the computer comprising one or more computer processors, and when executed by the one or more computer processors: (a) receiving, through the communication network, sequencing reads generated by a nucleic acid sequencer; and (b) determining the number of mutations present in the sequencing reads from the test sample and the minor allele ratio for one or more mutations most frequently shown in the sequencing reads from the test sample; (c) normalizing the number of mutations present in the sample relative to the number of mutations present in a control sample from another subject having the same cancer type within a bin of control samples, wherein the control sample has a minor allele ratio within a range that includes the minor allele ratio of the test sample, and the normalizing step includes subtracting from the determined number of mutations in the test sample of cell-free nucleic acids an average, median, or mode number of mutations in the control samples within the bin, and further dividing the number of mutations in the test sample of cell-free nucleic acids minus the average, median, or mode number of mutations present in the control samples by the standard deviation of the number of mutations present in the control samples to calculate a Z-score; A computer including a computer-readable medium including machine-executable code for performing a method including including the bin of control samples having a minor allele ratio within a range that includes the minor allele ratio of the test sample is selected from a set of bins of control samples, each of the set of bins of control samples has a minor allele ratio within a range of control samples, each of the bins of control samples is defined as a proportion of the entire range of minor allele ratios of the control samples, and the minor allele ratio refers to the ratio of DNA molecules having a mutation at a given genomic location in a given sample. System.
22. The system of claim 21, wherein the nucleic acid sequencer sequences a sequencing library purified from cell-free DNA molecules from a subject, and the sequencing library includes the cell-free DNA molecules and adapters including barcodes.
23. The system according to claim 22, wherein the sequencing library further comprises a sample barcode for identifying a sample from one or more samples. **Claim 24** (a) the computer-readable medium includes a memory, a hard drive, or a computer server; (b) the communication network includes a long-distance communication network, the Internet, an extranet, or an intranet; (c) the communication network includes one or more computer servers capable of distributed computing; and / or (d) the computer is installed on a computer server remotely located from the nucleic acid sequencer. The system according to any one of claims 21 to 23. **Claim 25** The system according to claim 24, wherein the distributed computing described in (c) is cloud computing. **Claim 26** The nucleic acid sequencer (a) performs sequencing bisynthesis on the sequencing library to generate the sequencing reads; (b) performs pyrosequencing, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, or sequencing by hybridization on the sequencing library to generate the sequencing reads; (c) uses a clonal single molecule array derived from the sequencing library to generate the sequencing reads; and / or (d) includes a chip having an array of microwells for sequencing the sequencing library to generate the sequencing reads. The system according to any one of claims 22 to 23 and claims 24 to 25 when directly or indirectly citing claim 22.
27. An electronic display that communicates with the computer through a network, further comprising an electronic display including a user interface for displaying the results when (a) to (c) are implemented, according to any one of claims 21 to 26.
28. The system according to claim 27, wherein the user interface is a graphical user interface (GUI) or a web-based user interface.
29. The system according to claim 27 or claim 28, wherein the electronic display exists in a personal computer and / or an Internet-enabled computer.
30. The system according to claim 29, wherein the Internet-enabled computer is installed at a location remote from the computer.
Citation Information
Patent Citations
Detection and treatment of disease exhibiting disease cell heterogeneity and systems and methods for communicating test results
WO2016109452A1