Detection of homologous recombination repair defects
A method using HRD scores derived from cfDNA analysis improves cancer detection and therapy guidance by enhancing sensitivity and specificity in identifying HRD, facilitating targeted treatments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-14
- Publication Date
- 2026-04-08
AI Technical Summary
Current methods lack sensitivity and specificity in detecting homologous recombination repair deficiency (HRD) in patients, which is crucial for diagnosing diseases like cancer and guiding targeted therapies such as PARP inhibitor treatments.
A method is developed to generate HRD scores using sequence information from cell-free nucleic acid (cfDNA) to determine HRD status, involving the use of computer algorithms to analyze HRD nucleic acid variants and generate scores for improved cancer detection and therapy guidance.
Enhances the sensitivity and specificity of cancer detection and identifies patients who can benefit from PARP therapy, providing accurate treatment strategies.
Smart Images

Figure 0007842700000007 
Figure 0007842700000008 
Figure 0007842700000009
Abstract
Description
Technical Field
[0001] Cross - reference to Related Patent Applications This application claims priority based on U.S. Provisional Patent Application No. 63 / 041,721, filed on June 19, 2020, and U.S. Provisional Patent Application No. 63 / 025,126, filed on May 14, 2020, the entire disclosures of which are incorporated herein by reference.
Background Art
[0002] Background There is a complex network of molecular pathways that function to repair DNA damage in order to maintain genomic stability. For example, homologous recombination repair (HRR) acts as a pathway to correct double-strand breaks in DNA during the S and G2 phases of the cell cycle [Lupo et al., "Inhibition of poly(ADP-ribosyl)ation in cancer: old and new paradigms revisited," Biochim Biophys Acta, 1846:201-15 (2014); Moschetta et al., "BRCA somatic mutations and epigenetic BRCA modifications in serous ovarian cancer," Ann Oncol., 27:1449-55 (2016)]. Disruptions to the HRR pathway, known as HRR deficiency (HRD), are thought to result in the loss or duplication of chromosomal regions, known as loss of heterozygosity (LOH) of the genome, and increase the number of tumor mutations and neoantigen rates [Solinas et al., "BRCA gene mutations do not shape the extent and organization of tumor infiltrating lymphocytes in triple negative breast cancer," Cancer Lett, 450:88-97 (2019)]. When cells have HRD, other repair pathways, such as non-homologous end joining (NHEJ), may be used to repair damaged DNA [Wang et al., "PARP-1 and Ku compete for repair of DNA double strand breaks by distinct NHEJ pathways," Nucleic Acids Res, 34:6170-82 (2006)].NHEJ is more error-prone than HRR and often leads to further mutation accumulation and chromosomal instability, thereby increasing the likelihood of tumorigenesis [Hoeijmakers, "Genome maintenance mechanisms for preventing cancer," Nature, 411:366-74 (2001)]. Patients with germline or somatic HRD may be candidates for targeted therapies, including DNA damage response (DDR) inhibitors, such as poly(ADP-ribose) polymerase (PARPi) inhibitors [Fong et al., "Poly(ADP)-ribose polymerase inhibition: frequent durable responses in BRCA carrier ovarian cancer correlating with platinum-free interval," J Clin Oncol, 28:2512-9 (2010); Audeh et al., "Oral poly(ADP-ribose) polymerase inhibitor olaparib in patients with BRCA1 or BRCA2 mutations and recurrent ovarian cancer: a proof-of-concept trial," Lancet, 376:245-51 (2010)]. Therefore, in order to diagnose diseases such as cancer and / or to obtain guidance for the treatment of these diseases, it is necessary to detect or classify HRD in patients, particularly from cell-free nucleic acid (cfDNA) samples. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Lupo et al., "Inhibition of poly(ADP-ribosyl)ation in cancer: old and new paradigms revisited," Biochim Biophys Acta, 1846:201-15 (2014) [Non-Patent Document 2] Moschetta et al., "BRCA somatic mutations and epigenetic BRCA modifications in serous ovarian cancer," Ann Oncol.,27:1449-55 (2016) [Non-Patent Document 3] Solinas et al., "BRCA gene mutations do not shape the extent and organization of tumor infiltrating lymphocytes in triple negative breast cancer," Cancer Lett, 450:88-97 (2019) [Non-Patent Document 4] Wang et al., "PARP-1 and Ku compete for repair of DNA double strand breaks by distinct NHEJ pathways," Nucleic Acids Res, 34:6170-82 (2006) [Non-Patent Document 5] Hoeijmakers, "Genome maintenance mechanisms for preventing cancer," Nature, 411:366-74 (2001) [Non-Patent Document 6] Fong et al., "Poly(ADP)-ribose polymerase inhibition: frequent durable responses in BRCA carrier ovarian cancer correlating with platinum-free interval," J Clin Oncol, 28:2512-9 (2010) [Non-Patent Document 7] Audeh et al., "Oral poly(ADP-ribose) polymerase inhibitor olaparib in patients with BRCA1 or BRCA2 mutations and recurrent ovarian cancer: a proof-of-concept trial," Lancet, 376:245-51 (2010) [Overview of the project] [Means for solving the problem]
[0004] Abstract This disclosure provides a method for generating homologous recombination repair deficiency (HRD) scores and a method for determining the HRD status of a test subject having a condition (e.g., cancer). The disclosed method improves the sensitivity and specificity of cancer detection assays and improves the sensitivity and specificity of identifying patients who may benefit from poly(ADP-ribose) polymerase (PARP) therapy. The disclosed method can be used to obtain guidance for treatment strategies. Additional methods, as well as related systems and computer-readable media, are also provided.
[0005] In one embodiment, a method is provided for generating homologous recombination repair deficiency (HRD) scores using at least partially a computer, comprising the steps of: generating a set of reference HRD scores by computer for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from one or more reference subjects having one or more oncologies, wherein a given reference HRD score includes the prevalence of a given HRD nucleic acid variant; and generating reference HRD scores from the set of reference HRD scores.
[0006] In one embodiment, a method is provided for determining the homologous recombination repair deficiency (HRD) status of a test subject having one or more cancer types, the method comprising the steps of: generating test HRD scores for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from the test subject to produce a set of test HRD scores, wherein a given test HRD score includes the prevalence of a given HRD nucleic acid variant; generating test HRD scores from the set of test HRD scores; and comparing the test HRD scores to a reference HRD score, wherein test HRD scores that are higher than the reference HRD score indicate that those test HRD scores are from a test subject having HRD, and test HRD scores that are at or below the reference HRD score indicate that those test HRD scores are from a test subject lacking HRD, thereby determining the HRD status of the test subject having one or more cancer types.
[0007] In one embodiment, a method is provided for detecting the presence or absence of homologous recombination repair deficiency (HRD) in a subject, which at least partially uses a computer, and includes the step of detecting HRD in a subject by determining by computer the presence or absence of at least one HRD nucleic acid variant in sequence information associated with one or more genes in a set of homologous recombination repair (HRR) genes derived from cell-free nucleic acid (cfDNA) obtained from the subject, using (i) a first probability that the sequence information includes a first state and a second probability that the sequence information includes a second state, wherein the first or second state includes at least the first HRD nucleic acid variant, and / or (ii) one or more aligned contigs generated from the sequence information, wherein the aligned contigs include at least the second HRD nucleic acid variant.
[0008] In one embodiment, a method is provided for treating a disease, comprising the step of administering one or more therapies to a subject having a disease and a homologous recombination repair deficiency (HRD) associated with the disease, wherein the presence or absence of at least one HRD nucleic acid variant in sequence information associated with one or more genes in a set of homologous recombination repair (HRR) genes derived from cell-free nucleic acid (cfDNA) obtained from the subject is determined by (i) a first probability that the sequence information includes a first state and a second probability that the sequence information includes a second state, wherein the first or second state includes at least the first HRD nucleic acid variant, and / or (ii) one or more aligned consecutive sequences (contigs) generated from the sequence information, wherein the HRD is detected and the disease is treated.
[0009] In one embodiment, a method is provided that includes the step of determining sequence data for a biological sample. The biological sample may include cell-free DNA (cfDNA). The method may include the step of determining coverage data based on the sequence data. The method may include the step of determining one or more breakpoints associated with one or more fusion events based on the coverage data. The method may include the step of determining one or more deletions associated with one or more genes based on the coverage data. The method may include the step of determining a homologous recombination deficiency (HRD) score based on one or more breakpoints and one or more deletions. The method may include the step of classifying the biological sample as HRD-positive based on the HRD score.
[0010] In one embodiment, a method is provided that includes the step of determining sequence data for a biological sample. The biological sample may include cell-free DNA (cfDNA). The method may include the step of determining coverage data based on the sequence data. The method may include the step of determining one or more breakpoints associated with one or more fusion events based on the coverage data. The method may include the step of determining one or more deletions associated with one or more genes based on the coverage data. The method may include the step of determining a homologous recombination deficiency (HRD) score based on one or more breakpoints and one or more deletions. The method may include the step of classifying the biological sample as HRD-negative based on the HRD score.
[0011] In some embodiments, targeted therapy may be administered to subjects having HRD as determined by one of the disclosed methods. Targeted therapy may include PARP inhibitors. Examples of PARP inhibitors that may be administered include one or more of veliparib, olaparib, talazoparib, lucaparib, niraparib, pamiparib, CEP 9722 (Cephalon), E7016 (Eisai), E7449 (Eisai, a PARP 1 / 2 and tankirase 1 / 2 inhibitor), or 3-aminobenzamides. In some embodiments, targeted therapy may include at least one base excision repair (BER) inhibitor. For example, olaparib may inhibit BER. In certain embodiments, targeted therapy may include a combination of PARP inhibitor and radiotherapy. In some embodiments, the combination of PARP inhibitor and radiotherapy makes it possible for the PARP inhibitor to cause the formation of double-strand breaks from single-strand breaks generated by radiotherapy in tumor tissue (e.g., tissue with BRCA1 / BRCA2 mutations). This combination can provide a more potent therapy per unit of radiation dose.
[0012] In some embodiments, the results of the systems and methods disclosed herein are used as input data for generating a report. The report may be in paper or electronic format. For example, if the determination of whether a subject has an HRD according to an HRD score is determined by the method or system disclosed herein, this can be directly shown in such a report.
[0013] Various steps of the methods disclosed herein, or steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographical locations, for example, in different countries, and / or by the same or different persons.
[0014] The accompanying drawings, which are incorporated in and constitute a part of this application, illustrate certain specific embodiments and, together with the description herein, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings, which are included by way of example and not limitation. Unless the context indicates otherwise, like reference numerals are understood to identify like components throughout these drawings. Also, it is understood that some or all of the drawings may be schematic representations for purposes of illustration and do not necessarily show the actual relative sizes or positions of the elements shown. **Brief Description of the Drawings**
[0015] [Figure 1] FIG. 1 (Panels A and B) schematically shows that cells having a defect in the homologous recombination repair (HRR) pathway are vulnerable to increased DNA damage and have increased sensitivity to DNA damage repair inhibitors (e.g., PARP inhibitors, etc.) and / or other therapies (modified from Peng et al. Exploiting the homologous recombination DNA repair network for targeted cancer therapy. World J Clin Oncol 2011; 2(2): 73 - 79 [PMID: 21603316]).
[0016] [Figure 2] FIG. 2 shows an example of a system including an HRD scoring module according to an embodiment of the present disclosure.
[0017] [Figure 3] FIG. 3 shows a schematic diagram of an HRD scoring module according to an embodiment of the present disclosure.
[0018] [Figure 4] FIG. 4 shows a schematic diagram of a fusion roller according to an embodiment of the present disclosure.
[0019] [Figure 5] Figure 5 shows a schematic diagram of a missing caller according to one embodiment of the present disclosure.
[0020] [Figure 6] Figure 6 shows a schematic diagram of an annotation module according to one embodiment of the present disclosure.
[0021] [Figure 7] Figure 7 shows a schematic diagram of a method for determining a return according to one embodiment of the present disclosure.
[0022] [Figure 8] Figure 8 shows a schematic diagram of a method for HRD scoring according to one embodiment of the present disclosure.
[0023] [Figure 9] Figure 9 shows a histogram of HRD scores for various cancer types.
[0024] [Figure 10] Figure 10 is a schematic flowchart illustrating exemplary method steps for generating a homologous recombination DNA repair deficiency (HRD) score and detecting HRD in a test subject, according to some embodiments.
[0025] [Figure 11] Figure 11 is a schematic flowchart illustrating exemplary method steps for determining homologous recombination DNA repair deficiency (HRD) status in a test subject having a given cancer type, according to some embodiments.
[0026] [Figure 12] Figure 12 is a schematic flowchart illustrating exemplary method steps for detecting homologous recombination DNA repair defects (HRDs) in a subject, according to some embodiments.
[0027] [Figure 13]Figure 13 is a flowchart illustrating exemplary method steps for treating a disease in a subject, according to some embodiments.
[0028] [Figure 14] Figure 14 is a schematic flowchart illustrating exemplary method steps for detecting DNA damage repair defects (DDRD) in a subject, according to some embodiments.
[0029] [Figure 15] Figure 15 (Panels A-C) is a plot of data showing the limit of detection (LoD) of GuardantOMNI® RUO for HRR deletions and fusions. Panel A shows the LoD for homozygous BRCA2 deletions, Panel B shows the LoD2 for LOH, and Panel C shows the LoD for long BRCA1 deletions. In the figures, the y-axis plots the probability of detection, and the x-axis plots the tumor percentage (TF).
[0030] [Figure 16] Figure 16 shows the oncoprint of HRR mutations in the prostate cancer cohort.
[0031] [Figure 17] Figure 17 (Panels A-C) plots the prevalence of HRR mutations by variant class detected in the prostate cohort. [Modes for carrying out the invention]
[0032] definition To facilitate understanding of this disclosure, certain terms are first defined below. Further definitions of the following terms and other terms may be provided throughout this specification. If any definition of a term provided below conflicts with a definition in a patent application or granted patent incorporated by reference, the definition provided here should be used to understand the meaning of the term.
[0033] Where used herein and in the appended claims, the singular forms “a,” “an,” and “the” include multiple subjects unless explicitly indicated otherwise by the context. Thus, for example, a reference to “a method” includes one or more methods and / or steps of the type described herein and / or of the type that would be apparent to those skilled in the art by reading this disclosure, etc. It will also be understood that there is an implicit “about” before temperatures, concentrations, times, number of bases or base pairs, coverage, etc., discussed herein, and therefore only a few, non-substantial equivalents are also within the scope of this disclosure. In this application, the use of the singular form includes the plural unless specifically stated otherwise. Also, “comprise,” “comprises,” “comprising,” “contain,” “contains,” “containing,” “include,” “includes,” and “including” are not intended to be limiting.
[0034] It should also be understood that the terminology used herein is intended solely to describe specific embodiments and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure pertains. The following terms, and their grammatical variants, will be used in descriptions of methods, computer-readable media, and systems, and in the claims, according to the definitions set forth below.
[0035] About: As used herein, “about” or “approximately” when applied to one or more values or elements of interest means a value or element that is similar to the reference value or element being referred to. In certain embodiments, unless otherwise stated or the context makes it clear that the term “about” or “approximately” means a range of values or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less than 1% of the reference value or element being referred to (except where such a number exceeds 100% of the possible values or elements).
[0036] Adapter: As used herein, “adapter” refers to a short nucleic acid (e.g., less than approximately 500 nucleotides, less than approximately 100 nucleotides, or less than approximately 50 nucleotides in length) that is typically at least partially double-stranded and used to ligate to either or both ends of a given sample nucleic acid molecule. An adapter may include a nucleic acid primer binding site for enabling amplification of the nucleic acid molecule whose ends are adjacent to the adapter, and / or a sequencing primer binding site, including a primer binding site for sequencing applications such as various next-generation sequencing (NGS) applications. An adapter may also include a binding site for a capture probe, such as an oligonucleotide bound to a flow cell support or similar. An adapter may also include nucleic acid tags as described herein. Nucleic acid tags are typically positioned relative to amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequencing read of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of a nucleic acid molecule. In certain embodiments, the same adapter is attached to each end of a nucleic acid molecule, except that the nucleic acid tag differs in its sequence. In some embodiments, the adapter is a Y-shaped adapter having one end blunt-ended or having a tail as described herein, for binding to a nucleic acid molecule that is also blunt-ended, or to a nucleic acid molecule having a tail with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter including an end with a blunt end or tail for binding to the nucleic acid to be analyzed. Other exemplary adapters include an adapter having a T tail and an adapter having a C tail.
[0037] To administer: As used herein, “to administer” or “to administer” a therapeutic agent (e.g., an immunological therapeutic agent, a DNA damage response (DDR) inhibitor (e.g., a poly(ADP-ribose) polymerase (PARP) inhibitor (PARPi))) to a subject means to give, apply, or bring into contact with the subject the composition. Administration may be carried out by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal routes.
[0038] Align: As used herein, “align,” “alignment,” and “to align” in the context of nucleic acids refer to aligning DNA or RNA sequences to identify regions of similarity. Similarity may relate to functional, structural, and / or evolutionary relationships between sequences. DNA sequence alignment includes alignment of the genomic DNA of one sequence with the genomic DNA of at least one other sequence. Such alignment may exclude non-genomic DNA, such as molecular barcodes, padding bases, and similar elements. For example, the genomic DNA of a sequence read may be aligned with the genomic DNA of a reference DNA sequence, excluding any molecular tags that may be bound to the sequence read.
[0039] Allele: As used herein, “allele” or “allele variant” refers to a specific genetic variant at a defined gene location or genomic locus. Allele variants are typically expressed with a frequency of 50% (0.5) or 100%, depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and typically have a frequency of 0.5 or 1. Somatic variants, however, are acquired variants and typically have a frequency of <0.5. Major and minor alleles of a gene locus refer to nucleic acids that have loci occupied by a reference sequence nucleotide and a variant nucleotide different from the reference sequence, respectively. Measurements at a gene locus may take the form of allele proportions (AF), which are measures of the frequency of the allele found in a sample.
[0040] Amplification: As used herein, "amplification" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or a portion of a polynucleotide, usually starting with a small amount of polynucleotide (e.g., a single polynucleotide molecule), and the amplified product or amplicon is generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes.
[0041] Barcode: As used herein, “barcode” in the context of nucleic acids refers to a nucleic acid molecule having a sequence that can function as a molecular identifier. For example, individual “barcode” sequences are typically added to DNA fragments during next-generation sequencing (NGS) library preparation so that each read can be identified and sorted before final data analysis.
[0042] Base excision repair inhibitor: As used herein, “base excision repair inhibitor” or “BER inhibitor” refers to a therapeutic agent that inhibits the base excision repair (BER) pathway, mechanism, or process.
[0043] Breakpoint: As used herein, “breakpoint” in the context of a nucleic acid fusion molecule or a corresponding sequencing read refers to a terminal nucleotide position at the junction between fused subsequences of a nucleic acid fusion, or a terminal nucleotide position represented by the corresponding sequencing read. For example, a given split sequence read contains a first subsequence that is contiguous with a second subsequence in the split sequence read and is located 5' to the second subsequence, and this first subsequence is located at a first gene locus in a reference sequence, which is discontinuous with a second gene locus in the reference sequence where the second subsequence is located. In this example, the first subsequence of the split sequence read contains a breakpoint at its 3' terminal nucleotide, while the second subsequence of the split sequence read contains a breakpoint at its 5' terminal nucleotide. In certain applications, breakpoints, for example, these breakpoints, are referred to as a “breakpoint pair.”
[0044] Oncology type: As used herein, “cancer,” “oncology type,” or “tumor type” refers to a type or subtype of cancer as defined, for example, by histopathological diagnosis. Oncology type is defined by any conventional criteria, for example, by the presence in a given tissue (e.g., hematological cancer, central nervous system (CNS) cancer, brain cancer, lung cancer (small cell and non-small cell), skin cancer, nasal cancer, pharyngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer). Cancer can be defined based on cancers that present cancer markers, such as solid cancers, heterologous cancers, allologous cancers, cancers of unknown cause, and similar cancers, as well as cancers of the same cell lineage (e.g., carcinomas, sarcomas, lymphomas, cholangiocarcinomas, leukemias, mesotheliomas, melanomas, or glioblastomas), and cancers that present cancer markers, such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, KRAS, BRAF, NRAS, hormone receptors, and NMP-22. Cancer can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether it is primary or secondary.
[0045] Cell-free nucleic acids: As used herein, “cell-free nucleic acids” refers to nucleic acids that are not contained within cells and are not otherwise bound to cells. In some embodiments, “cell-free nucleic acids” refers to nucleic acids that, at the time of isolation from the subject, are not contained within cells and are not otherwise bound to cells. Cell-free nucleic acids may include, for example, all unencapsulated nucleic acids supplied from bodily fluids from the subject (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, nucleolar small RNA (snoRNA), Piwi-binding RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids may be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids by secretion or cell death processes, such as cell necrosis, apoptosis, or similar processes. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. ctDNA can be unencapsulated tumor-derived fragmented DNA. Another example of cell-free nucleic acids is fetal DNA that circulates freely in the maternal bloodstream, also known as cell-free fetal DNA (cffDNA). Cell-free nucleic acids may have one or more epigenetic modifications; for example, cell-free nucleic acids may be acetylated, 5-methylated, ubiquitinated, phosphorylated, SUMOylated, ribosylated, and / or citrullinated.
[0046] Classifier: As used herein, “classifier” generally refers to an algorithmic computer code that receives test data as input data and produces output data of the classification of the input data as belonging to one or another class (e.g., tumor DNA or non-tumor DNA, having or not having DNA damage repair deficiency (DDRD)).
[0047] Contiguous Sequence: As used herein, “contiguous sequence” or “contig” refers to a set of overlapping nucleic acid segments that together represent a consensus region of nucleic acids.
[0048] Copy number variant: As used herein, “copy number variant,” “CNV,” or “copy number diversity” refers to the phenomenon in which genomic compartments are repeated and the number of repeats within the genome differs among individuals in the population under consideration.
[0049] Coverage: As used herein, “coverage” refers to the number of nucleic acid molecules that represent a particular base position.
[0050] De novo fusion caller: As used herein, “de novo fusion caller,” “fusion caller,” or “de novo method” refers to a fusion caller, either a DNA fusion caller or an RNA fusion caller, that identifies a fusion event de novo, i.e., without prior knowledge such as that which can be obtained from a database of known gene fusion events.
[0051] Deoxyribonucleic acid or ribonucleic acid: As used herein, “deoxyribonucleic acid” or “DNA” refers to a natural or modified nucleotide having a hydrogen group at the 2' position of the sugar moiety. Typically, DNA refers to a nucleotide chain containing a deoxyribonucleoside with one of four types of nucleic acid bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, “ribonucleic acid” or “RNA” refers to a natural or modified nucleotide having a hydroxyl group at the 2' position of the sugar moiety. Typically, RNA refers to a nucleotide chain containing a ribonucleoside with one of four types of nucleic acid bases: A, uracil (U), G, and C. As used herein, the term “nucleotide” refers to a natural or modified nucleotide. Certain nucleotide pairs bind specifically to each other in a complementary manner (this is called complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand is joined to a second nucleic acid strand composed of nucleotides complementary to those in the first strand, these two strands join to form a double helix. As used herein, “nucleic acid sequencing data,” “nucleic acid sequencing information,” “sequence information,” “nucleic acid sequence,” “nucleotide sequence,” “genome sequence,” “gene sequence,” or “fragment sequence,” or “nucleic acid sequencing read” means any information or data indicating the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a nucleic acid molecule such as DNA or RNA (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment).It should be understood that this instruction intends to describe sequence information obtained using any available type of technique, platform, or technology, including but not limited to capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion or pH-based detection stems, and digital signature-based systems.
[0052] To detect: As used herein, “to detect,” “to detect,” or “detect” means the act of determining the existence or presence of one or more target nucleic acids (e.g., nucleic acids having a target mutation or other marker) in a sample.
[0053] DNA Damage Repair: As used herein, “DNA damage repair” or “DDR” refers to a biochemical pathway, mechanism, or process that repairs DNA damage during the cell cycle. Directional reversal DNA damage repair mechanisms do not include a template because the underlying damage does not involve a cleavage of the phosphodiester backbone in the affected DNA. Other DNA damage repair pathways, mechanisms, or processes include a template. These include single-strand damage repair mechanisms that act to repair DNA when only one of the two strands of a given double helix is damaged, e.g., base excision repair (BER), nucleotide excision repair (NER), and mismatch repair (MMR). Template-dependent DNA repair processes also include double-strand damage repair mechanisms that act to repair DNA when both strands of a given double helix are damaged (e.g., cleaved), e.g., homologous recombination (HR), microhomology-mediated end joining (MMEJ), and non-homologous end joining (NHEJ).
[0054] DNA damage repair deficiency: As used herein, “DNA damage repair deficiency” or “DDRD” refers to a mutation or set of mutations that partially or completely disrupts a DNA damage repair pathway, mechanism, or process.
[0055] DNA damage repair genes: As used herein, “DNA damage repair genes” or “DDR genes” refer to genes that encode polypeptides involved in DNA damage repair pathways, mechanisms, or processes.
[0056] Fusion event: As used herein, “fusion event” refers to the fusion of at least two separate genes at a specific location. Examples of causes of fusion events include translocations, intermediate deletions, or chromosomal inversions.
[0057] Gene: As used herein, “gene” refers to any segment of DNA associated with a biological function. Thus, a gene includes coding sequences and, if necessary, regulatory sequences required for their expression. A gene also includes, if necessary, non-expressed DNA segments that form recognition sequences for, for example, other proteins.
[0058] Homologous Recombination Repair Deficiency Score: As used herein, “Homologous Recombination Repair Deficiency Score” or “HRD Score” refers to a value representing the number or set of mutations or other measures associated with DNA damage repair deficiencies (DDRD), such as homologous recombination repair deficiencies (HRD), that are found in or otherwise known to be present in one or more genomic regions of a given subject or in one or more genomic regions of a given subject population.
[0059] Germline mutation: As used herein, “germline mutation” means a mutation in germ cells, and therefore a mutation that can be passed on to offspring.
[0060] Homologous recombination repair: As used herein, “homologous recombination repair” or “HRR” refers to a template-dependent DNA repair process that occurs during DNA replication. Typically, homologous regions on sister chromatids act as templates as part of the process to repair damaged DNA strands.
[0061] Homologous recombination repair defect: As used herein, “homologous recombination repair defect” or “HRD” refers to a mutation or set of mutations that partially or completely disrupts a homologous recombination repair pathway, mechanism, or process.
[0062] Homologous recombination repair gene: As used herein, “homologous recombination repair gene” or “HRR gene” refers to a gene that encodes a polypeptide involved in a homologous recombination repair pathway, mechanism, or process.
[0063] Homozygous deletion: As used herein, “homozygous deletion” or “biallelic inactivation” refers to a mutation or nucleic acid variant that results in the loss of both alleles of a given gene.
[0064] Hemizygous deletion: As used herein, “hemizygous deletion” or “monoallelic inactivation” refers to a mutation or nucleic acid variant that results in the loss of one of two alleles of a given gene. “Heterozygous deletion” is a hemizygous deletion in which the original or initial two alleles of a given gene are distinct from each other.
[0065] Indel: As used herein, “indel” refers to a mutation that involves the insertion or deletion of a nucleotide in the genome in question.
[0066] Loss of Function: As used herein, “loss of function” or “LoF” in the context of a biochemical pathway, mechanism, or process means a mutation or set of mutations (e.g., in a given sample) that renders a biochemical pathway, mechanism, or process nonfunctional. For example, a loss of function (LoF) DNA damage repair deficiency (DDRD) is a mutation or set of mutations that renders a given DNA damage repair (DDR) pathway, mechanism, or process (e.g., base excision repair (BER) pathway, mechanism, or process, nucleotide excision repair (NER) pathway, mechanism, or process, mismatch repair (MMR) pathway, mechanism, or process, homologous recombination repair (HRR) pathway, mechanism, or process, non-homologous end joining (NHEJ) pathway, mechanism, or process, and / or similar).
[0067] Loss of Heterozygosity: As used herein, “loss of heterozygosity” or “LOH” refers to a mutational event that results in the loss of one parent’s contribution to a given cell or a given group of cell clones (e.g., an entire gene and surrounding chromosomal regions). LOH may be caused, for example, by gene conversion, direct deletion, mitotic recombination, deletion due to unbalanced organization, or loss of chromosome (monosomy).
[0068] Minor allele frequency: As used herein, “minor allele frequency” refers to the frequency at which a minor allele (e.g., not the most frequent allele) is present in a given nucleic acid population, such as a sample obtained from a subject. Genetic variants with low minor allele frequencies are typically relatively infrequent in the sample.
[0069] Mutant Allele Proportion: As used herein, “mutant allele proportion” or “MAF” refers to the proportion of nucleic acid molecules in a given sample that contain an allele modification or mutation relative to a given genomic location. MAF is generally expressed as a proportion or percentage. For example, MAF is typically about 0.5, 0.1, 0.05, or less than 0.01 (i.e., about 50%, 10%, 5%, or less than 1%) of all somatic variants or alleles present at a given locus.
[0070] Maximum Mutagenic Allele Proportion: As used herein, “Maximum Mutagenic Allele Proportion,” “Maximum MAF,” or “MAX MAF” refers to the maximum or largest MAF of all somatic variants present or observed in a given sample.
[0071] Mutation: As used herein, “mutation,” “nucleic acid variant,” “variant,” or “genetic abnormality” refers to a variation from a known reference sequence, including, for example, single nucleotide variants (SNVs), copy number variants or variations (CNVs) / abnormalities, insertions or deletions (indels), shortenings, gene fusions, transversions, translocations, frameshifts, duplications, repeat extensions, and epigenetic variants. Mutations can be germline or somatic mutations. In some embodiments, the reference sequence for comparison is the wild-type genome sequence of the species under consideration, typically the human genome, which provides the test sample. In certain cases, the mutation or variant is a “tumor-associated genetic variant” that causes tumorigenesis or at least contributes to tumorigenesis.
[0072] Next-generation sequencing: As used herein, “next-generation sequencing” or “NGS” refers to sequencing techniques that offer increased throughput compared to conventional Sanger and capillary electrophoresis-based methods, such as the ability to simultaneously generate hundreds of thousands of relatively short sequence reads. Some examples of next-generation sequencing techniques include, but are not limited to, single-nucleotide synthesis, ligation sequencing, and hybridization sequencing.
[0073] Nucleic acid tags: As used herein, “nucleic acid tags” means short nucleic acids (e.g., less than approximately 500, 100, 50, or 10 nucleotides in length) used to label nucleic acid molecules to distinguish nucleic acids from different samples of different types or those that have undergone different treatments (e.g., representing a sample index), or nucleic acids used to label nucleic acid molecules to distinguish different nucleic acid molecules in the same sample of different types or those that have undergone different treatments (e.g., representing a molecular tag). Nucleic acid tags may be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags may have the same length or a variety of lengths, as necessary. Nucleic acid tags may contain a double-stranded molecule with one or more blunt ends, a 5' or 3' single-stranded region (e.g., an overhang), and / or one or more other single-stranded regions at other positions within a given molecule. Nucleic acid tags may be conjugated to one or both ends of another nucleic acid (e.g., a sample nucleic acid to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information about a given nucleic acid, such as its origin, morphology, or processing. Nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different nucleic acid tags and / or sample indices, which are then deconvoluted by reading the nucleic acid tags. Nucleic acid tags are sometimes called molecular identifiers or tags, sample identifiers, index tags, and / or barcodes. In addition or alternatively, nucleic acid tags can be used to distinguish different molecules within the same sample. This includes, for example, uniquely tagging different nucleic acid molecules within a given sample, or non-uniquely tagging such molecules. In the case of non-unique tagging applications, nucleic acid molecules can be tagged using tags with a limited number of different sequences, and thus different molecules can be distinguished, for example, based on their start and / or stop locations in a selected reference genome in combination with at least one nucleic acid tag.Typically, a sufficient number of different nucleic acid tags are used, and therefore, the probability that any two molecules will have the same start / stop position and also have the same nucleic acid tag is low (e.g., less than approximately 10%, less than approximately 5%, less than approximately 1%, or less than approximately 0.1%). Some nucleic acid tags include multiple molecular identifiers to label a sample, the morphology of the nucleic acid molecule in the sample, and the nucleic acid molecule in a morphology that has the same start and stop position. Such nucleic acid tags can be referred to using the exemplary morphology "A1i", where uppercase letters indicate the sample type, Arabic numerals indicate the morphology of the molecule in the sample, and lowercase Roman numerals indicate the molecule in a particular morphology.
[0074] Poly-ADP-ribose polymerase inhibitors: As used herein, “poly-ADP-ribose polymerase inhibitors,” “PARP inhibitors,” or “PARPi” refer to therapeutic agents that inhibit the action of the enzyme poly-ADP-ribose polymerase (PARP).
[0075] Polynucleotides: As used herein, “polynucleotides,” “nucleic acids,” “nucleic acid molecules,” or “oligonucleotides” refer to linear polymers of nucleosides linked by nucleoside linkages (deoxyribonucleosides, ribonucleosides, or analogs thereof). Typically, polynucleotides contain at least three nucleosides. Oligonucleotides often range in size from a small number of monomer units, e.g., 3-4, to several hundred. Whenever polynucleotides are represented by a sequence of letters such as “ATGCCTG,” it will be understood that, unless otherwise noted, the nucleotides are in 5'→3' order from left to right, and in the case of DNA, “A” represents deoxyadenosine, “C” represents deoxycytidine, “G” represents deoxyguanosine, and “T” represents deoxythymidine. The letters A, C, G, and T may be used to refer to the base itself, to refer to a nucleoside, or to refer to a nucleotide containing a base, as is common in the art.
[0076] Prevalence: As used herein, “prevalence” in the context of nucleic acid variants means the extent, prevalence, or frequency to which a given nucleic acid variant is found or observed in a given sample (e.g., a given body fluid sample, a given non-body fluid sample, etc.) or other population (e.g., a given population of body fluid samples, a given population of non-body fluid samples, etc.).
[0077] Reference Sample: As used herein, “reference sample” or “reference cfDNA sample” refers to a sample of known composition and / or known to have, or to have or lack, certain characteristics (e.g., known nucleic acid variants, known cellular origin, known tumor percentage, known coverage, and / or similar) that is analyzed together with or compared to a test sample to evaluate the accuracy of an analytical procedure. A reference sample dataset typically contains at least about 25 to at least about 30,000 or more reference samples. In some embodiments, the reference sample dataset includes approximately 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000, or more reference samples.
[0078] Reference Sequence: As used herein, “reference sequence” or “reference genome” refers to a known sequence used for comparison with an experimentally determined sequence. For example, a known sequence may be an entire genome, a chromosome, or any segment thereof. A reference sequence typically contains at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1,000, at least about 10,000, at least about 100,000, at least about 1,000,000, at least about 10,000,000, at least about 1,000,000, or more nucleotides. A reference sequence may be aligned with a single contiguous sequence of genome or chromosome, or it may contain discontinuous segments that align with different regions of genome or chromosome. Examples of reference sequences include, for instance, the human genome, specifically hG19 and hG38.
[0079] Sample: As used herein, “Sample” means any biological sample that can be analyzed by the methods and / or systems disclosed herein. In certain aspects of this disclosure, Sample is a fluid sample, in particular whole blood or a fraction thereof, lymph, urine, and / or cerebrospinal fluid, among the types of fluids supplied with cell-free (circulating, not contained within cells or otherwise bound to cells) nucleic acids. In certain implementations, a fluid sample is a plasma sample, which is the fluid portion of whole blood excluding cells such as red blood cells and white blood cells. In some implementations, a fluid sample is a serum sample, i.e., fibrinogen-free plasma. In certain aspects of this disclosure, Sample is a “non-fluid sample” or “non-plasma sample,” i.e., a biological sample other than a “fluid sample,” such as a cell and / or tissue sample supplied with nucleic acids other than cell-free nucleic acids.
[0080] Sensitivity: As used herein, “sensitivity” in the context of a given assay or method refers to the ability of the assay or method to detect and distinguish between target (e.g., nucleic acid variant) analytes and non-target analytes.
[0081] Sequencing: As used herein, “sequencing” refers to any of several techniques used to determine the sequence (e.g., identity and order of monomeric units) of a biomolecule, such as nucleic acids, such as DNA or RNA. Exemplary sequencing methods include targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, sungardideoxytermination sequencing, whole-genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, and single-nucleotide extension sequencing. Examples of sequencing methods include, but are not limited to, solid-phase sequencing, high-throughput sequencing, large-scale parallel signature sequencing, emulsion PCR, co-amplification-PCR at lower denaturation temperatures (COLD-PCR), multiplex PCR, sequencing by reversible dye terminators, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, single-nucleotide synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD® sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed using a gene analyzer, for example, a gene analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among others.
[0082] Sequence information: As used herein, “sequence information” in the context of nucleic acid polymers means the order and / or identity of monomer units (e.g., nucleotides, etc.) within the polymer.
[0083] Single nucleotide variant: As used herein, “single nucleotide variant” or “SNV” means a variation or diversity of a single nucleotide at a specific location within the genome.
[0084] Somatic mutation: As used herein, “somatic mutation” means a given mutation in the genome that occurs after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0085] Specificity: As used herein, “specificity” in the context of a diagnostic analysis or assay refers to the degree to which the analysis or assay detects the intended target analyte while excluding other components of a given sample.
[0086] Status: As used herein, “status” in the context of an object means one or more statuses of a given object, for example, whether the object has a DNA damage repair deficiency (DDRD) (e.g., homologous recombination repair deficiency (HRD) and / or similar).
[0087] Subject: As used herein, “subject” or “test subject” means an animal, e.g., a mammalian species (e.g., human) or a bird species (e.g., bird), or another organism, e.g., a plant. More specifically, a subject may be a vertebrate, e.g., a mammal, e.g., a mouse, a primate, a monkey, or a human. Animals include livestock (e.g., production cattle, dairy cows, poultry, horses, pigs, and similar), competition animals, and companion animals (e.g., pets or support animals). A subject may be a healthy individual, an individual with or suspected to have a disease or disease predisposition, or an individual in need of therapy or suspected to need therapy. The terms “individual” or “patient” are intended to be synonymous with “subject.” In some embodiments, a subject is a human being who has or is suspected to have cancer. For example, a subject may be an individual diagnosed with cancer, an individual who will undergo cancer therapy, and / or an individual who has received at least one cancer therapy. A subject may also be in remission from cancer. In another example, a subject may be an individual diagnosed with an autoimmune disease. In another example, the subject may be a pregnant woman who has been diagnosed with, or is suspected of having, a disease such as cancer or an autoimmune disease, or a woman who is planning to become pregnant. A “reference subject” refers to a subject that is known to have or lack certain characteristics (e.g., known DDRD or HRD status, known nucleic acid variant, known cellular origin, known tumor rate, known coverage, and / or similar).
[0088] Threshold: As used herein, “threshold” refers to a separately determined value used to characterize or classify experimentally determined values. In certain embodiments, for example, “threshold value” refers to a selected value from which a quantitative value is compared to determine whether a given target nucleic acid variant is absent at a given gene locus.
[0089] Tumor Percentage: As used herein, “tumor percentage” refers to an estimate of the proportion of nucleic acid molecules in a given sample that originate from tumors. For example, the tumor percentage of a sample may be the maximum variant allele frequency (MAX MAF) of the sample, or the coverage of the sample, or a measure derived from the length, epigenetic state, or other properties of cfDNA fragments in the sample, or any other selected feature of the sample. The term “MAX MAF” refers to the maximum or greatest MAF of all somatic variants present in a given sample. In some embodiments, the tumor percentage of a sample is equal to the MAX MAF of the sample.
[0090] Value: As used herein, “Value” or “Score” generally refers to any entry in a dataset that may characterize the feature pointed to by the Value. This includes, but is not limited to, numbers, words or phrases, symbols (e.g., + or -) or degrees. Detailed explanation
[0091] Introduction
[0092] DNA damage repair (DDR) is a cellular process that plays a role in maintaining the integrity or stability of the genome. A defect or deficiency in a given DDR mechanism can lead to tumorigenesis or other diseases, and these defects or deficiencies can be used to identify test subjects or patients who may benefit from a given targeted therapy. For example, homologous recombination repair deficiency (HRD) is a cellular phenotype that may make a patient a candidate for administration of therapeutic agents such as PARP inhibitors. To illustrate with an example, Figure 1 (Panels A and B) schematically shows that cells with a deficiency in the homologous recombination repair (HRR) pathway are more susceptible to increased DNA damage and are more sensitive to DNA damage repair inhibitors (e.g., PARP inhibitors) and / or other therapies. As shown, normal cells with DNA damage (Panel A) will often survive even if a patient is administered a PARP inhibitor, because the PARP inhibitor only inhibits PARP-mediated repair of single-strand breaks (SSBs). During DNA replication, as the DNA helix unwinds, these SSBs result in double-strand breaks (DSBs). In these normal cells, this DNA damage can be repaired by homologous recombination (HR)-mediated repair pathway, which repairs DSBs, and therefore the normal cells survive. In contrast, in the case of HR-deficient cancer cells (Panel B), for example, the HR-mediated repair pathway does not function, and therefore the administered PARP inhibitor inhibits the remaining PARP-mediated repair pathway, leading to the death of the cancer cells.
[0093] There are various classes of alterations or mutations that inactivate HRD. Some of these include, among many other alterations, SNVs and / or indels of the HRR gene, homozygous deletions, gene-specific LOH, copy-number-less LOH, genome-wide LOH, shortened rearrangements, and multi-exon (long) deletions. In certain embodiments disclosed herein, SNVs and indels are identified using pathogenicity annotation techniques, homozygous deletions, gene-specific LOH, copy-number-less LOH, and genome-wide LOH are identified using homozygous deletion / LOH CNV callers, and shortened rearrangements and multi-exon deletions are identified using rearrangement or de novo fusion callers.
[0094] In some embodiments, the methods and related embodiments disclosed herein are used to identify HRR pathway deficiencies to guide PARP inhibitor treatment in patients with ovarian cancer, prostate cancer, breast cancer, or other cancers. In some of these embodiments, the HRD workflow provides information on copy number reduction, rearrangement, and pathogenic SNVs and indels of the HRR gene for identifying samples with HRD, and thus for identifying candidates for targeted therapies including PARP inhibitors. In some of these embodiments, this is achieved by system modules such as SNV / indel, fusion, and CNV callers. In some embodiments, the reports generated as output data of these processes reveal values for variant types indicating loss of function (LoF) of the relevant HRR gene.
[0095] Essentially any DDR (e.g., HRR) gene or biomarker can be evaluated for associated mutations that cause the corresponding DDR (e.g., HRR) pathway to malfunction or become non-functional in a given sample. This information can be used as a selection criterion for administering targeted therapies (e.g., PARP inhibitors, BER inhibitors, etc.) to a patient. In certain embodiments, targeted therapy may include PARP inhibitors. Examples of PARP inhibitors that may be administered include one or more of veliparib, olaparib, talazoparib, lucaparib, niraparib, pamiparib, CEP 9722 (Cephalon), E7016 (Eisai), E7449 (Eisai, a PARP 1 / 2 and tankirase 1 / 2 inhibitor), or 3-aminobenzamides. In some embodiments, targeted therapy may include at least one base excision repair (BER) inhibitor. For example, olaparib can inhibit BIR. In certain embodiments, targeted therapy may include a combination of PARP inhibitors and radiotherapy. In some embodiments, a combination of a PARP inhibitor and radiotherapy allows the PARP inhibitor to induce the formation of double-strand breaks from single-strand breaks generated by radiotherapy in tumor tissue (e.g., tissue with BRCA1 / BRCA2 mutations). This combination can provide a more potent therapy per radiation dose. Using the methods and related embodiments of this disclosure, essentially any number of genes can be evaluated as needed. In some embodiments, for example, the set of DDR genes (e.g., HRR genes) subject to analysis as described herein includes at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 100, 1,000, 10,000, or more genes. A non-exclusive list of HRR genes is provided in Table 1, and one or more of these HRR genes may be selected for evaluation using the methods and related embodiments disclosed herein, as necessary. [Table 1]
[0096] Table 2 contains an exemplary set of HRR genes that can be evaluated as described herein to identify patients who are candidates for specific targeted therapies. [Table 2]
[0097] Exemplary Systems and Methods
[0098] Figure 2 illustrates an example of a system 100 for determining the DNA damage repair deficiency (DDRD) status (e.g., HRD status or similar) of a test subject 111 according to embodiments of the present disclosure. System 100 can process one or more samples 101 from the subject 111 to generate sequence reads for variant detection and DDRD status determination. System 100 may include a laboratory system 102, a computer system 110, and / or other components. Note that the laboratory system 102 and the computer system 110 may be located far apart from each other and may be connected to each other by a computer network (not shown). The laboratory system 102 may include a sample collection and preparation pipeline 103, a sequencing pipeline 105, a sequence read data store 109, and / or other components. The sequencing pipeline 105 may include one or more sequencing devices 107 (illustrated in Figure 2 as sequencing devices 107a...n).
[0099] The methods of this disclosure can be used in a wide variety of applications in the manipulation, preparation, identification, quantification, and / or analysis of cell-free nucleic acids. As shown in Figure 2, the sample collection and preparation pipeline 103 includes obtaining a cfDNA reference sample 101 from one or more reference subjects and a cfDNA test sample 111 from a test subject. As described herein, polynucleotides may comprise any type of nucleic acid, such as DNA and / or RNA. For example, if the polynucleotide is DNA, it may be genomic DNA, complementary DNA (cDNA), or any other deoxyribonucleic acid. Polynucleotides may also comprise cell-free nucleic acids, such as cell-free DNA (cfDNA). For example, polynucleotides may comprise circulating cfDNA. Circulating cfDNA may comprise DNA that is expelled from cells of the body by apoptosis or necrosis. cfDNA expelled by apoptosis or necrosis may originate from normal (e.g., healthy) cells of the body. Tumor DNA may be expelled if there is abnormal tissue growth, such as abnormal tissue growth for cancer. Circulating cfDNA may comprise circulating tumor DNA (ctDNA). i. Sample
[0100] Cell-free polynucleotides can be isolated and extracted by collecting samples using various techniques. The sample may be any biological sample isolated from the subject. Samples may include body tissues, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsy material (e.g., biopsy material from known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial or extracellular fluid (e.g., fluid from the interstitial space), gingival exudate, gingival crevicular exudate, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Preferably, the sample is a body fluid, particularly blood and its fractions, as well as urine. Such samples may contain nucleic acids excreted from tumors. Nucleic acids may include DNA and RNA and may be in double-stranded or single-stranded form. A sample may be the form in which it was initially isolated from the subject, or it may have undergone further processing to remove or add components such as cells, to concentrate one component compared to another, or to convert one form of nucleic acid to another, for example, RNA to DNA, or single-stranded nucleic acid to double-stranded nucleic acid. Therefore, for example, a bodily fluid sample for analysis may be plasma or serum containing cell-free nucleic acid, such as cell-free DNA (cfDNA).
[0101] In some embodiments, the sample volume of bodily fluids taken from the subject depends on the desired reading depth for the region being sequenced. Exemplary volumes are approximately 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, volumes are approximately 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, 40 ml, or more milliliters. The volume of plasma sampled is typically between approximately 5 ml and 20 ml.
[0102] A sample may contain varying amounts of nucleic acids. Typically, the amount of nucleic acid in a given sample is considered equivalent to several genome equivalents. For example, a sample with approximately 30 ng of DNA is equivalent to approximately 10,000 (10) 4 ) In the case of haploid human genome equivalents and cfDNA, approximately 200,000,000,000 (2 × 10⁻¹⁶). 11) may contain individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, about 600,000,000,000 individual molecules.
[0103] In some embodiments, the sample includes nucleic acids from different sources, e.g., from cells and from cell-free sources (e.g., blood samples). Typically, the sample includes nucleic acids that carry mutations. For example, the sample may include DNA that carries germline mutations and / or somatic mutations, as is appropriate. Typically, the sample includes DNA that carries cancer-related mutations (e.g., cancer-related somatic mutations). In some embodiments of this disclosure, cell-free nucleic acids in the subject may be derived from a tumor. For example, cell-free DNA isolated from the sample may include ctDNA.
[0104] Exemplary amounts of cell-free nucleic acids in a sample before amplification are typically in the range of about 1 femtogram (fg) to about 1 microgram (μg), for example, about 1 picogram (pg) to about 200 nanograms (ng), about 1 ng to about 100 ng, or about 10 ng to about 1000 ng. In some embodiments, the sample contains cell-free nucleic acid molecules in amounts of about 600 ng or less, about 500 ng or less, about 400 ng or less, about 300 ng or less, about 200 ng or less, about 100 ng or less, about 50 ng or less, or about 20 ng or less. If necessary, the amounts are at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the amounts are approximately 1 fg, 10 fg, 100 fg, 1 pg, 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or less than 200 ng of cell-free nucleic acid molecules. In some embodiments, the method includes the step of obtaining approximately 1 fg to approximately 200 ng of cell-free nucleic acid molecules from a sample.
[0105] Cell-free nucleic acids typically have a size distribution between approximately 100 and 500 nucleotides in length, with molecules between approximately 110 and 230 nucleotides accounting for about 90% of the molecules in the sample. The mode is approximately 168 nucleotides, and a second small peak is in the range of approximately 240 to 440 nucleotides. In certain embodiments, cell-free nucleic acids are approximately 160 to 180 nucleotides, or approximately 320 to 360 nucleotides, or approximately 440 to 480 nucleotides.
[0106] In some embodiments, cell-free nucleic acids are isolated from body fluids by a fractionation step, in which cell-free nucleic acids, if present in solution, are separated from intact cells and other insoluble components in the body fluids. In some of these embodiments, fractionation includes techniques such as centrifugation or filtration. Alternatively, cells in the body fluids are lysed, and cell-free nucleic acids and cellular nucleic acids are processed together. Generally, after the addition of buffers and washing steps, cell-free nucleic acids are precipitated, for example, with alcohol. In certain embodiments, additional purification steps are used, such as silica-based columns to remove impurities or salts. For example, non-specific bulk carrier nucleic acids are added as needed through the reaction to optimize certain aspects of the exemplary procedure, such as yield. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. If necessary, single-stranded DNA and / or single-stranded RNA are converted to double-stranded form for inclusion in subsequent processing and analysis steps. Further details regarding cfDNA fractionation and the relevant analysis of epigenetic modifications as necessary for use in carrying out the methods disclosed herein are described, for example, in WO2018 / 119452, filed on 22 December 2017, which is incorporated herein by reference. ii. Nucleic acid tags
[0107] In certain embodiments, a tag providing a molecular identifier or barcode is incorporated into or otherwise attached to an adapter as part of a sample collection and preparation pipeline 103, by a number of methods, particularly chemical synthesis, ligation, or overlap extension PCR. In some embodiments, the assignment of a unique or non-unique identifier or molecular barcode in a reaction utilizes the methods and systems described in, for example, U.S. Patent Application Publications 20010053519, 20030152490, 20110160078, and U.S. Patents 6,582,908, 7,537,898, and 9,598,731, each of which is incorporated herein by reference.
[0108] The tags are attached to the sample nucleic acids randomly or intentionally. In some embodiments, the tags are introduced into the microwells in a ratio of expected identifiers (e.g., combinations of unique and / or non-unique barcodes). For example, identifiers can be loaded so that approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or more identifiers are loaded per genomic sample. In some embodiments, identifiers are loaded such that each genome sample is loaded with approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or fewer than 1,000,000,000 identifiers. In certain embodiments, the average number of identifiers loaded per sample genome is less than or greater than approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers per genome sample. Identifiers are generally unique and / or non-unique.
[0109] One exemplary form uses approximately 2 to 1,000,000 different tags, or approximately 5 to 150 different tags, or approximately 20 to 50 different tags, ligated to both ends of a target nucleic acid molecule. In the case of 20 to 50 × 20 to 50 tags, a total of 400 to 2,500 tags are produced. Typically, such a number of tags is sufficient to ensure that different molecules with the same start and stop points have a high probability of receiving different combinations of tags (e.g., at least 94%, 99.5%, 99.99%, 99.999%).
[0110] In some embodiments, the identifier is an oligonucleotide of a predetermined, random, or semi-random sequence. In other embodiments, multiple barcodes may be used such that the barcodes in the multiple barcodes are not necessarily unique to one another. In these embodiments, the barcodes are generally bound to individual molecules (e.g., by ligation or PCR amplification), resulting in a unique sequence that can be individually tracked by the combination of the barcode and the sequence to which it can be bound. As described herein, by detecting the non-uniquely tagged barcode in combination with sequence data of the start (departure) and end (stop) portions of the sequence read, it is usually possible to assign a unique identity to a particular molecule. The length or number of base pairs of individual sequence reads may also be used to assign a unique identity to a given molecule, if necessary. As described herein, a fragment from a single strand of nucleic acid to which a unique identity has been assigned may then allow for the subsequent identification of fragments from the parent strand and / or complementary strand. iii. Nucleic acid amplification
[0111] The sample nucleic acid adjacent to the adapter is typically amplified by PCR and other amplification methods, using the binding of nucleic acid primers to primer binding sites in the adapter adjacent to the DNA molecule to be amplified, as part of the sample collection and preparation pipeline 103. In some embodiments, the amplification method may involve a cycle of extension, denaturation, and annealing resulting from thermal circulation, or it may be isothermal, for example, in the case of transcription-mediated amplification. Other exemplary amplification methods that may be used as needed include, among many others, ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and auto-persistent sequence-based replication.
[0112] Conventional nucleic acid amplification methods typically involve one or more rounds of amplification cycles to introduce molecular tags and / or sample index / tags into nucleic acid molecules. Amplification is usually performed using one or more reaction mixtures. Molecular tags and sample index / tags are introduced simultaneously or in any order, as needed. In some embodiments, molecular tags and sample index / tags are introduced before and / or after the sequence capture step. In some embodiments, only the molecular tag is introduced before probe capture, and the sample index / tag is introduced after the sequence capture step. In certain embodiments, both the molecular tag and sample index / tag are introduced before the probe-based capture step. In some embodiments, the sample index / tag is introduced after the sequence capture step. Typically, the sequence capture protocol involves introducing a single-stranded nucleic acid molecule complementary to a target nucleic acid sequence, e.g., the coding sequence of a genomic region, and the coding sequence of a mutation in such a region associated with an oncogenic type. Typically, the amplification reaction produces multiple nucleic acid amplicons non-uniquely or uniquely tagged with molecular tags and sample index / tags, ranging in size from approximately 200 nucleotides (nt) to approximately 700 nt, 250 nt to approximately 350 nt, or approximately 320 nt to approximately 550 nt. In some embodiments, the amplicons have a size of approximately 300 nt. In some embodiments, the amplicons have a size of approximately 500 nt. iv. Nucleic acid enrichment
[0113] In some embodiments, sequences are enriched as part of a sample collection and preparation pipeline 103 before sequencing the nucleic acids. The enrichment may be performed on specific target regions or non-specifically (on “target sequences”), as required. In some embodiments, the target regions of interest can be enriched using nucleic acid capture probes (“baits”) selected for one or more bait set panels using a differential tiling and capture scheme. A differential tiling and capture scheme typically involves using bait sets of different relative concentrations to tile differentially (e.g., at different “resolutions”) across genomic regions associated with the baits, imposing a set of constraints (e.g., sequencer constraints, e.g., sequencing load, utilization of each bait, etc.) to capture the target nucleic acids at the desired level for downstream sequencing. These target genomic regions of interest may optionally include native or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, the target sequences may be captured using biotin-labeled beads having probes on one or more compartments of interest, and these compartments may subsequently be amplified to enrich the regions of interest, as required.
[0114] Sequence capture typically involves the use of oligonucleotide probes that hybridize to a target nucleic acid sequence. In certain embodiments, a probe set strategy involves tiling probes across a desired compartment. Such probes may be, for example, about 60 to about 120 nucleotides in length. This set may have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x, or more. Generally, the effectiveness of sequence capture depends, in part, on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the probe sequence. b. Nucleic acid sequencing
[0115] As shown in Figure 2, after extraction and isolation of cfDNA from the sample by the sample collection and preparation pipeline 103, the cfDNA can be sequenced by a sequencing pipeline 105 including one or more sequencing devices 107. Sample nucleic acids, pre-amplified or unamplified, and optionally adjacent to adapters, are generally subjected to sequencing. Sequencing methods or commercially available formats used as needed include, for example, Sanger sequencing, high-throughput sequencing, bisulfite sequencing, pyrosequencing, single-nucleotide synthesis, single-molecule sequencing, nanopore-based sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), large-scale parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking; and sequencing using PacBio, SOLiD, Ion Torrent, or nanopore platforms. Sequencing reactions can be carried out in various sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. The sample processing unit may also include multiple sample chambers to enable simultaneous processing of multiple runs.
[0116] Sequencing can be performed on one or more nucleic acid fragment types or compartments known to contain markers for cancer or other diseases. Sequencing can also be performed on any nucleic acid fragment present in the sample. Sequencing can provide genome sequence coverage of at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome. In other cases, genome sequence coverage may be less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome.
[0117] Simultaneous sequencing reactions can be performed using multiplex sequencing techniques. In some embodiments, cell-free polynucleotides are sequenced in at least about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. In other embodiments, cell-free polynucleotides are sequenced in less than about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. Sequencing reactions are typically performed sequentially or simultaneously. Subsequent data analysis is generally performed with respect to all or some of the sequencing reactions. In some embodiments, data analysis is performed for at least about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. In other embodiments, data analysis may be performed for fewer than about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, or 100,000 sequencing reactions. An example read depth is about 1,000 to about 50,000 reads per locus (base position).
[0118] In some embodiments, nucleic acid populations are prepared for sequencing by enzymatically forming blunt ends on double-stranded nucleic acids having single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Examples of enzymes or catalytic fragments of which may be used as needed include Krenow large fragments and T4 polymerase. At the 5' overhang, the enzyme typically extends the recessed 3' end on the opposite strand until its 3' end aligns perfectly with the 5' end to produce a blunt end. At the 3' overhang, the enzyme generally digests from the 3' end to the 5' end of the opposite strand, and may also digest beyond the 5' end. If this digestion proceeds beyond the 5' end of the opposite strand, an enzyme with the same polymerase activity used for the 5' overhang can fill the gap. The formation of blunt ends on double-stranded nucleic acids facilitates, for example, adapter binding and subsequent amplification.
[0119] In some embodiments, the nucleic acid population is subjected to additional processing, such as the conversion of single-stranded nucleic acids to double-stranded nucleic acids and / or RNA to DNA. These forms of nucleic acids are also ligated to adapters and amplified as needed.
[0120] Sequenced nucleic acids can be produced by sequencing nucleic acids subjected to the blunt-end formation process described above, whether previously amplified or not, and, if necessary, other nucleic acids in the sample. Sequentially, sequenced nucleic acids may refer to the sequence of a nucleic acid (i.e., sequence information), or to a nucleic acid whose sequence has been determined. Sequence data for individual nucleic acid molecules in a sample can be obtained, directly or indirectly, from the consensus sequences of the amplified products of individual nucleic acid molecules in the sample.
[0121] In some embodiments, double-stranded nucleic acids with single-stranded overhangs in a sample are blunt-ended, and both ends are ligated to an adapter containing a barcode. Sequencing then determines not only the nucleic acid sequence but also the inline barcode induced by the adapter. The blunt-ended DNA molecule is ligated, as necessary, to the blunt ends of an adapter that is at least partially double-stranded (e.g., Y-shaped or bell-shaped). Alternatively, complementary nucleotide tails can be attached to the blunt ends of the sample nucleic acid and the adapter to facilitate ligation (e.g., adherent-end ligation).
[0122] Typically, a nucleic acid sample is brought into contact with a sufficient number of adapters, and therefore, the probability that any two copies of the same nucleic acid will receive the same combination of adapter barcodes from adapters ligated to both ends is low (e.g., <1 or 0.1%). Using adapters in this way allows for the identification of families of nucleic acid sequences that have the same start and stop points on the reference nucleic acid and are ligated to the same combination of barcodes. Such families represent the sequences of the amplified products of the nucleic acids in the sample before amplification. When modified by blunt end formation and adapter ligation, the sequences of family members can be compiled to derive the consensus nucleotide or complete consensus sequence of the nucleic acid molecule in the original sample. In other words, a nucleotide occupying a specific position in the nucleic acid in the sample is determined to be the consensus of the nucleotide occupying the corresponding position in the family member sequence. A family may include sequences from one or both strands of a double-stranded nucleic acid. If a family member includes sequences from both strands of a double-stranded nucleic acid, the sequence from one strand is converted to its complementary strand for the purpose of compiling all sequences to derive a consensus nucleotide or sequence. Some families contain only a single member sequence. In this case, this sequence can be considered the nucleic acid sequence in the sample before amplification. Alternatively, families containing only a single member sequence can be excluded from subsequent analyses.
[0123] For further details regarding nucleic acid sequencing, including the forms and applications described herein, see, for example, Levy et al., Annual Review of Genomics and Human Genetics, 17: 95-115 (2016), Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012), Voelkerding et al., Clinical Chem., 55: 641-658 (2009), MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009), and Astier et al., J Am Chem Soc., 128(5):1705-10. (2006), U.S. Patent Nos. 6,210,891, 6,258,568, 6,833,246, 7,115,400, 6,969,488, 5,912,148, 6,130,073, 7,169,560, 7,282,337, 7,482,120, and 7,501,245 These references are also provided in U.S. Patent Nos. 6,818,395, 6,911,345, 7,501,245, 7,329,492, 7,170,050, 7,302,146, 7,313,308, and 7,476,503, and these references are incorporated herein by reference in their entirety. i. Sequencing Panel
[0124] To increase the likelihood of detecting a target genomic region and, if necessary, the likelihood of detecting a tumor exhibiting the mutation, the DNA parcels to be sequenced may include a panel of genes or genomic parcels containing known genomic regions. By selecting a limited number of parcels (e.g., a limited panel) for sequencing, the total sequencing required (e.g., the total amount of nucleotides to be sequenced) can be reduced. The sequencing panel may target multiple different genes or regions, for example, to detect a single cancer, a set of cancers, or all cancers. Alternatively, DNA can be sequenced without using a sequencing panel by whole-genome sequencing (WGS) or other unbiased sequencing methods. Examples of suitable panels and targets for use in panels can be found in the epigenetic targets described in international application WO2020160414, filed on 31 January 2020, the aforementioned reference patent document, is incorporated herein by reference in its entirety.
[0125] In some embodiments, a panel targeting multiple different genes or genomic regions (e.g., transcription factor binding regions, distal regulatory elements (DREs), repeat elements, intron-exon junctions, transcription start sites (TSSs), and / or similar) is selected such that a predetermined proportion of subjects with cancer exhibit a genetic variant or tumor marker in one or more genes within the panel. The panel may be selected to limit the region to be sequenced to a fixed number of base pairs. The panel may be selected to sequence a desired amount of DNA. The panel may further be selected to achieve a desired sequence reading depth. The panel may be selected to achieve a desired sequence reading depth or sequence reading coverage with respect to the amount of base pairs being sequenced. The panel may be selected to achieve theoretical sensitivity, theoretical specificity, and / or theoretical precision for detecting one or more genetic variants in a sample.
[0126] The HRR genes included in this panel may include one or more of the following: ATM, ATR, BAP1, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, HDAC2, MRE11, NBN, PALB2, RAD50, RAD51, RAD51B, RAD51C, RAD51D, RAD54L, XRCC2, and XRCC3.
[0127] Probes for detecting a panel of regions can include not only those for detecting target genomic regions (hotspot regions) but also nucleosome recognition probes (e.g., KRAS codons 12 and 13), and such probes can be designed to optimize capture based on an analysis of cfDNA coverage and fragment size variations influenced by nucleosome binding patterns and GC sequence composition. In this case, the regions used may also include non-hotspot regions optimized based on nucleosome location and GC model. The panel may include multiple subpanels, including a subpanel for identifying primary tissue (e.g., using published literature to define 50-100 baits representing genes (not necessarily promoters) with the most diverse transcriptional profiles across tissue), a subpanel for identifying whole-genome scaffolds (e.g., for identifying hyperconservative genomic content and sparsely tiling across chromosomes with a small number of probes for copy-number-based lining), and a subpanel for identifying transcription start sites (TSS) / CpG islands (e.g., for capturing differential methylation regions (e.g., variable methylation regions (DMRs)) in the promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer)). In some embodiments, the marker for primary tissue is a tissue-specific epigenetic marker.
[0128] In some embodiments, one or more regions within the panel include one or more loci from one or more genes to detect residual cancer after surgery. This detection may be earlier than that possible with existing cancer detection methods. In some embodiments, one or more gene locations within the panel include one or more loci from one or more genes to detect cancer in high-risk patient populations. For example, smokers have a much higher rate of lung cancer than the general population. Furthermore, smokers may develop other lung conditions that make cancer detection more difficult, such as the development of irregular nodules in the lungs. In some embodiments, the methods described herein can detect cancer in high-risk patients earlier than that possible with existing cancer detection methods.
[0129] Gene locations to be included in the sequencing panel can be selected based on the number of subjects with cancer in which the tumor marker is located in that gene or region. Gene locations to be included in the sequencing panel can also be selected based on the prevalence of cancer and subjects with cancer in which the tumor marker is located in that gene. The presence of a tumor marker within a region may indicate a subject with cancer.
[0130] In some cases, information from one or more databases can be used to select a panel. Information about cancer may be derived from cancer tumor biopsies or cfDNA assays. A database may contain information describing a population of sequenced tumor samples. A database may contain information about mRNA expression in tumor samples. A database may contain information about regulatory elements or genomic regions in tumor samples. Information about sequenced tumor samples may include the frequencies of various genetic variants and may describe the genes or regions in which these genetic variants exist. Genetic variants may be tumor markers. A non-limiting example of such a database is COSMIC, a catalog of somatic mutations found in various cancers. For a particular cancer, COSMIC ranks genes based on their mutation frequency. Genes with high mutation frequencies among given genes can be selected for inclusion in a panel. For example, COSMIC shows that 33% of a population of sequenced breast cancer samples have the TP53 mutation, and 22% of a population of sampled breast cancers have the KRAS mutation. Other ranked genes, including APC, have mutations found in only about 4% of the population of sequenced breast cancer samples. TP53 and KRAS can be included in the sequencing panel based on their relatively high frequency in the sampled breast cancers (compared to APC, for example, which is present at a frequency of about 4%). While COSMIC is provided as a non-limiting example, any database or set of information that associates cancer with tumor markers located in genes or genetic regions can be used. In another example, as provided by COSMIC, 380 out of 1156 biliary tract cancer samples (33%) carry the TP53 mutation. Several other genes, such as APC, have mutations in 4–8% of all samples. Therefore, TP53 can be selected for inclusion in the panel based on its relatively high frequency in the population of biliary tract cancer samples.
[0131] A gene or genomic compartment may be selected for a panel if the frequency of a tumor marker is significantly higher in the sampled tumor tissue or circulating tumor DNA than in a given background population. A combination of gene locations may be selected to include in the panel such that at least the majority of subjects with cancer have a tumor marker or genomic region located in at least one gene location or gene within the panel. For a particular cancer or set of cancers, a combination of gene locations may be selected based on data showing that the majority of subjects have one or more tumor markers in one or more selected regions. For example, to detect cancer 1, a panel including regions A, B, C, and / or D may be selected based on data showing that 90% of subjects with cancer 1 have tumor markers in regions A, B, C, and / or D of the panel. Alternatively, a tumor marker may be shown to be independently present in two or more regions of subjects with cancer, and therefore, combined, tumor markers in two or more regions are present in the majority of the population of subjects with cancer. For example, to detect cancer 2, a panel including regions X, Y, and Z can be selected based on data showing that 90% of subjects have tumor markers in one or more regions, and that in 30% of such subjects the tumor marker is detected only in region X, while in the remaining subjects where the tumor marker is detected, it is detected only in regions Y and / or Z. Tumor markers located at one or more gene locations previously proven to be associated with one or more cancers can indicate or predict that a subject has cancer if the tumor marker is detected in one or more regions in 50% or more of those regions at that time. Computational methods, such as models that utilize the conditional probability of detecting cancer based on the frequency of cancer for a set of tumor markers in one or more regions, can be used to predict which regions, individually or in combination, may predict cancer.Other methods for panel selection include the use of databases describing information from studies using comprehensive genomic profiling and / or whole-genome sequencing (WGS, RNA-seq, Chip-seq, bisulfite sequencing, ATAC-seq, etc.) of tumors in large panels. Information gathered from the literature may also describe pathways that are generally affected and mutate in certain cancers. The use of ontologs describing genetic information can provide further information for panel selection.
[0132] The genes included in the sequencing panel may include the complete transcription region, promoter region, enhancer region, regulatory elements, and / or downstream sequences. To further increase the likelihood of detecting tumors exhibiting mutations, only exons may be included in the panel. The panel may include all exons of a selected gene, or one or more exons of a selected gene. The panel may include exons from each of several different genes. The panel may also include at least one exon from each of several different genes.
[0133] In some embodiments, a panel of exons from each of several different genes is selected such that a predetermined proportion of subjects having cancer exhibits a genetic variant in at least one exon within the panel of exons.
[0134] At least one whole exon from each of the different genes within a gene panel can be sequenced. The panel being sequenced may contain exons from multiple genes. The panel may contain exons from 2 to 100 different genes, 2 to 70 genes, 2 to 50 genes, 2 to 30 genes, 2 to 15 genes, or 2 to 10 genes.
[0135] The selected panel may contain a variety of exons. The panel may contain 2 to 3000 exons. The panel may contain 2 to 1000 exons. The panel may contain 2 to 500 exons. The panel may contain 2 to 100 exons. The panel may contain 2 to 50 exons. The panel may contain 300 or fewer exons. The panel may contain 200 or fewer exons. The panel may contain 100 or fewer exons. The panel may contain 50 or fewer exons. The panel may contain 40 or fewer exons. The panel may contain 30 or fewer exons. The panel may contain 25 or fewer exons. The panel may contain 20 or fewer exons. The panel may contain 15 or fewer exons. The panel may contain 10 or fewer exons. The panel may contain 9 or fewer exons. The panel may contain 8 or fewer exons. The panel may contain 7 or fewer exons.
[0136] The panel may contain one or more exons from multiple different genes. The panel may contain one or more exons from each of a certain proportion of multiple different genes. The panel may contain at least two exons from each of at least 25%, 50%, 75%, or 90% of different genes. The panel may contain at least three exons from each of at least 25%, 50%, 75%, or 90% of different genes. The panel may contain at least four exons from each of at least 25%, 50%, 75%, or 90% of different genes.
[0137] The size of a sequencing panel can vary. For example, the size of a sequencing panel can be larger or smaller (in terms of nucleotide size) depending on several factors, including the total amount of nucleotides being sequenced, or the number of unique molecules being sequenced in a particular region of the panel. Sequencing panels can be 5kb to 50kb in size. Sequencing panels can be 10kb to 30kb in size. Sequencing panels can be 12kb to 20kb in size. Sequencing panels can be 12kb to 60kb in size. The sequencing panel may be at least 10kb, 12kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 110kb, 120kb, 130kb, 140kb, or 150kb in size. The sequencing panel may also be less than 100kb, 90kb, 80kb, 70kb, 60kb, or 50kb in size.
[0138] The panel selected for sequencing may contain at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 gene locations (e.g., each containing a genomic region of interest). In some cases, gene locations within the panel that are relatively small in size are selected. In some cases, the regions within the panel have a size of approximately 10kb or less, approximately 8kb or less, approximately 6kb or less, approximately 5kb or less, approximately 4kb or less, approximately 3kb or less, approximately 2.5kb or less, approximately 2kb or less, approximately 1.5kb or less, or approximately 1kb or less or less. In some cases, gene locations within a panel have sizes ranging from approximately 0.5kb to 10kb, 0.5kb to 6kb, 1kb to 11kb, 1kb to 15kb, 1kb to 20kb, 0.1kb to 10kb, or 0.2kb to 1kb. For example, regions within a panel may have sizes ranging from approximately 0.1kb to 5kb.
[0139] The panels selected herein may enable deep sequencing sufficient to detect low-frequency genetic variants (e.g., in cell-free nucleic acid molecules obtained from a sample). The amount of a genetic variant in a sample may be referred to in terms of the minor allele frequency for a given genetic variant. Minor allele frequency may refer to the frequency at which a minor allele (e.g., not the most frequent allele) is present in a given nucleic acid population, such as a sample. Genetic variants with low minor allele frequencies may be relatively infrequent in a sample. In some cases, the panel may enable the detection of genetic variants with minor allele frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, or 0.5%. The panel may enable the detection of genetic variants with minor allele frequencies of 0.001% or greater. The panel may enable the detection of genetic variants with minor allele frequencies of 0.01% or greater. The panel enables the detection of genetic variants present in the sample at frequencies as low as 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel enables the detection of tumor markers present in the sample at frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. The panel may enable the detection of tumor markers present in the sample at frequencies as low as 1.0%. The panel may enable the detection of tumor markers present in the sample at frequencies as low as 0.75%. The panel can enable the detection of tumor markers at frequencies as low as 0.5% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.25% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.1% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.075% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.05% in the sample. The panel can enable the detection of tumor markers at frequencies as low as 0.025% in the sample.The panel may enable the detection of tumor markers at frequencies as low as 0.01% in a sample. The panel may enable the detection of tumor markers at frequencies as low as 0.005% in a sample. The panel may enable the detection of tumor markers at frequencies as low as 0.001% in a sample. The panel may enable the detection of tumor markers at frequencies as low as 0.0001% in a sample. The panel may enable the detection of tumor markers in sequenced cfDNA at frequencies as low as 1.0% to 0.0001% in a sample. The panel may enable the detection of tumor markers in sequenced cfDNA at frequencies as low as 0.01% to 0.0001% in a sample.
[0140] Genetic variants can be expressed as a percentage of the target population having a disease (e.g., cancer). In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the population with cancer exhibit one or more genetic variants in at least one region within the panel. For example, at least 80% of the population with cancer may exhibit one or more genetic variants in at least one genomic location within the panel.
[0141] The panel may include one or more locations containing the target genomic region from each of one or more genes. In some cases, the panel may include one or more locations containing the target genomic region from each of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the panel may include one or more locations containing the target genomic region from each of at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, the panel may include one or more locations containing the target genomic region from each of about 1 to about 80, 1 to about 50, about 3 to about 40, 5 to about 30, or 10 to about 20 different genes.
[0142] You can select locations within the panel that include genomic regions so that one or more epigenetically modified regions are detected. These epigenetically modified regions may be acetylated, methylated, ubiquitinated, phosphorylated, SUMOylated, ribosylated, and / or citrullinated. For example, you can select regions within the panel so that one or more methylated regions are detected.
[0143] Regions within a panel can be selected to contain sequences that are differentially transcribed across one or more tissues. In some cases, locations containing genomic regions may contain sequences that are transcribed at a higher level in certain tissues compared to others. For example, locations containing genomic regions may contain transcribed sequences in certain tissues but not in others.
[0144] Gene locations within a panel may include coding and / or non-coding sequences. For example, gene locations within a panel may include one or more sequences in exons, introns, 3' untranslated regions, 5' untranslated regions, regulatory elements, transcription start sites, and / or splice sites. In some cases, regions within a panel may include other non-coding sequences, including pseudogenes, repetitive sequences, transposons, viral elements, and telomeres. In some cases, gene locations within a panel may include sequences in non-coding RNA, such as ribosomal RNA, transfer RNA, Piwi-binding RNA, and microRNA.
[0145] Gene locations within the panel can be selected to detect (diagnose) cancer at a desired sensitivity level (e.g., by detecting one or more genetic variants). For example, regions within the panel can be selected to detect cancer with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% (e.g., by detecting one or more genetic variants). Gene locations within the panel can also be selected to detect cancer with 100% sensitivity.
[0146] Gene locations within the panel can be selected to detect (diagnose) cancer at a desired specificity level (e.g., by detecting one or more genetic variants). For example, gene locations within the panel can be selected to detect cancer with a specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% (e.g., by detecting one or more genetic variants). Gene locations within the panel can also be selected to detect one or more genetic variants with 100% specificity.
[0147] Genetic locations within a panel can be selected to detect (diagnose) cancer with a desired positive predictive value. Positive predictive values can be increased by improving sensitivity (the chance of detecting an actual positive) and / or specificity (the chance of not mistaking an actual negative for a positive). As a non-limiting example, genetic locations within a panel can be selected to detect one or more genetic variants with positive predictive values of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within a panel can be selected to detect one or more genetic variants with a 100% positive predictive value.
[0148] Gene locations within the panel can be selected to detect (diagnose) cancer with the desired accuracy. As used herein, the term “accuracy” refers to the ability of a test to discriminate between a diseased state (e.g., cancer) and a healthy state. Accuracy can be quantified using measures such as sensitivity and specificity, predictive value, likelihood ratio, area under the ROC curve, Joden index, and / or diagnostic odds ratio.
[0149] Accuracy can be presented as a percentage, which refers to the ratio of the number of trials that yielded correct results to the total number of trials performed. Regions within the panel can be selected to detect cancer with an accuracy of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Gene locations within the panel can be selected to detect cancer with 100% accuracy.
[0150] Panels can be selected to achieve high sensitivity and to detect low-frequency genetic variants. For example, a panel can be selected to detect genetic variants or tumor markers present at frequencies as low as 0.01%, 0.05%, or 0.001% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Gene locations within the panel can be selected to detect tumor markers present at frequencies of 1% or less in a sample with a sensitivity of 70% or higher. The panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.1%, with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.01%, with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.001% with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0151] Panels can be selected to achieve high specificity and to detect low-frequency genetic variants. For example, a panel can be selected to detect genetic variants or tumor markers present at frequencies as low as 0.01%, 0.05%, or 0.001% in a sample with a specificity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Gene locations within the panel can be selected to detect tumor markers present at frequencies of 1% or less in a sample with a specificity of 70% or higher. A panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.1%, with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.01%, with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. A panel can be selected to detect tumor markers present in the sample at a frequency as low as 0.001%, with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0152] Panels can be selected to achieve high accuracy and to detect low-frequency genetic variants. Panels can be selected to detect genetic variants or tumor markers present in the sample at frequencies as low as 0.01%, 0.05%, or 0.001% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Gene locations within the panel can be selected to detect tumor markers present in the sample at frequencies of 1% or less with 70% or higher accuracy. Panels can be selected to detect tumor markers present in the sample at frequencies as low as 0.1% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The panel can be selected to detect tumor markers present in the sample at frequencies as low as 0.01% with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers present in the sample at frequencies as low as 0.001% with an accuracy of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0153] Panels can be selected to provide high predictive capabilities and to detect low-frequency genetic variants. Panels can be selected so that genetic variants or tumor markers present at frequencies as low as 0.01%, 0.05%, or 0.001% in the sample can have positive predictive values of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0154] The concentration of the probe or bait used in the panel can be increased (from 2 to 6 ng / μL) to capture more nucleic acid molecules in the sample. The concentration of the probe or bait used in the panel may be at least 2 ng / μL, 3 ng / μL, 4 ng / μL, 5 ng / μL, 6 ng / μL, or higher. The probe concentration may be approximately 2 ng / μL to 3 ng / μL, approximately 2 ng / μL to 4 ng / μL, approximately 2 ng / μL to 5 ng / μL, or approximately 2 ng / μL to 6 ng / μL. The concentration of the probe or bait used in the panel may be 2 ng / μL or higher to 6 ng / μL or lower. In some cases, this allows for the analysis of more molecules in the biologic, thereby enabling the detection of lower-frequency alleles.
[0155] In one embodiment, after sequencing, the sequence reads can be stored in a sequence read data store 109. The sequence reads can be stored in any format. The sequence read store 109 can be local and / or remote from the location where sequencing is performed.
[0156] As shown in Figure 2, stored sequence reads can be fed into a sequence analysis pipeline 112. The sequence analysis pipeline 112 may include a sequence quality control (QC) component 113 that can filter sequence reads from the laboratory system 102. The sequence quality control (QC) component 113 can assign a quality score to one or more sequence reads. The quality score may be an indication of the sequence reads, indicating whether those sequence reads may be useful for subsequent analysis based on a threshold. In some cases, some sequence reads are neither of sufficient quality nor length to perform subsequent mapping steps. Sequence reads with a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the sequence read dataset. In other cases, sequence reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the dataset.
[0157] Sequence reads that meet a specific quality score threshold can be mapped to a reference genome by copy number module 115. After mapping alignment, a mapping score can be assigned to the sequence reads. The mapping score may be an indication of the sequence read remapped to the reference sequence, showing whether each position is uniquely mappable. Sequence reads with a mapping score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the dataset. In other cases, sequence reads assigned a mapping score of less than 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered and removed from the dataset of sequence reads.
[0158] After filtering, multiple sequence reads generate a chromosomal coverage region. The copy number module 115 can divide this chromosomal region into variable-length windows or bins. The windows or bins may be at least 5kb, 10kb, 25kb, 30kb, 35kb, 40kb, 50kb, 60kb, 75kb, 100kb, 150kb, 200kb, 500kb, or 1000kb. The windows or bins may also have bases less than or equal to 5kb, 10kb, 25kb, 30kb, 35kb, 40kb, 50kb, 60kb, 75kb, 100kb, 150kb, 200kb, 500kb, or 1000kb. The window or bin can also be approximately 5kb, 10kb, 25kb, 30kb, 35kb, 40kb, 50kb, 60kb, 75kb, 100kb, 150kb, 200kb, 500kb, or 1000kb.
[0159] Copy number module 115 can normalize coverage by having windows or bins contain approximately the same number of mappable bases. In some cases, each window or bin within a chromosomal region may contain exactly the same number of mappable bases. In other cases, each window or bin may contain a different number of mappable bases. In addition, each window or bin may not overlap with any adjacent windows or bins. In other cases, a window or bin may overlap with another adjacent window or bin. In some cases, windows or bins may overlap by at least 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 10 bp, 20 bp, 25 bp, 50 bp, 100 bp, 200 bp, 250 bp, 500 bp, or 1000 bp. In other cases, windows or bins may overlap by 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 10 bp, 20 bp, 25 bp, 50 bp, 100 bp, 200 bp, 250 bp, 500 bp, or 1000 bp or less. In some cases, windows or bins may overlap by approximately 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 10 bp, 20 bp, 25 bp, 50 bp, 100 bp, 200 bp, 250 bp, 500 bp, or 1000 bp.
[0160] In some cases, each window region may have a specified size, and therefore they contain approximately the same number of uniquely mappable bases. The mappability of each base constituting a window region is determined and used to generate a mappability file, which contains a display of reads from references that are remapped to the references in each file. The mappability file contains one line per position, indicating whether each position is uniquely mappable or not.
[0161] In addition, predefined windows can be filtered from the dataset that are known to be difficult to sequence across the entire genome or to contain a significant amount of GC bias. For example, regions known to be near the centromere of a chromosome (i.e., centromere DNA) are known to contain highly repetitive sequences that can produce false positive results. These regions can be filtered out. Other regions of the genome that contain abnormally high concentrations of other highly repetitive sequences, such as microsatellite DNA, can also be filtered from the dataset.
[0162] The number of windows analyzed can vary. In some cases, at least 10, 20, 30, 40, 50, 100, 200, 500, 1000, 2000, 5,000, 10,000, 20,000, 50,000, or 100,000 windows are analyzed. In other cases, the number of windows analyzed is 10, 20, 30, 40, 50, 100, 200, 500, 1000, 2000, 5,000, 10,000, 20,000, 50,000, or 100,000 or less, and these windows are analyzed.
[0163] The copy number module 115 can determine read coverage for each window / bin region. This determination can be made using either barcoded or non-barcoded reads. If barcoded reads are not used, the previous mapping step will provide coverage for different base positions. Sequence reads that have sufficient mapping and quality scores and fall within the unfiltered chromosomal window can be counted. A score is assigned to each mappable position in the number of coverage reads.
[0164] In one embodiment, a quantitative measure of sequencing read coverage is a measure indicating the number of reads (e.g., specific locations, bases, regions, genes, or chromosomes from a reference genome) derived from DNA molecules corresponding to a gene locus. To associate reads with gene loci, reads can be mapped or aligned to a reference. Software for mapping or aligning (e.g., Bowtie, BWA, mrsFAST, BLAST, BLAT) can associate sequencing reads with gene loci. During the mapping process, certain parameters can be optimized. Non-limiting examples of mapping process optimization include masking repeating regions; utilizing mapping quality (e.g., MAPQ) score cutoffs; using different seed lengths to generate alignments; and limiting edit distances between genomic locations.
[0165] Quantitative measures related to sequencing read coverage may include read counts associated with gene loci. In some cases, these counts are converted to novel metrics to mitigate the effects of different sequencing depths, library complexity, or gene locus sizes. Exemplary metrics include Reads Per Kilobase per Million (RPKM), Fragments Per Kilobase per Million (FPKM), Trimmed Mean of M values (TMM), variance-stabilized raw counts, and log-transformed raw counts. Other conversions are also known to those skilled in the art and can be used for specific applications.
[0166] A quantitative measure can be determined using the number of read families or folded reads, where each read family or folded read corresponds to an initial template DNA molecule. Methods for folding and quantifying read families can be found in PCT / US2013 / 058061 and PCT / US2014 / 000048, each of which is incorporated herein by reference in its entirety. In particular, a method for quantifying and / or folding read families can be used, which sorts reads into families using barcode and sequence information from sequencing reads such that each family shares a barcode sequence and a sequencing read sequence and / or at least a portion of the same genomic coordinates when mapped to a reference sequence. Thus, for the majority of families, each family originates from a single initial template DNA molecule. The count derived from the mapping of sequences from families can be called a “Unique Molecular Count” (UMC). In some cases, determining a quantitative measure related to sequencing read coverage involves normalizing UMC by a library size-related metric to obtain a normalized UMC ("normalized UMC"). Exemplary methods include dividing the UMC of a locus by the sum of all UMCs; or dividing the UMC of a locus by the sum of all autosomal UMCs. When comparing multiple sequencing read datasets, UMC can be normalized, for example, by the median UMC of the loci in two sequencing read datasets. In some cases, a quantitative measure related to sequencing read coverage may be a normalized UMC, which is further normalized as follows: (i) determine the normalized UMC for the corresponding loci from sequencing reads derived from the training sample; (ii) for each locus, normalize the sample's normalized UMC by the median normalized UMC of the training sample at the corresponding locus to obtain the relative abundance (RA) of the locus.
[0167] Consensus sequences can be identified by folding sequencing reads based on those sequences, for example, based on identical sequences within the first 5, 10, 15, 20, or 25 bases. In some cases, folding allows for one, two, three, four, or five differences in reads that are otherwise identical. In some cases, folding uses the mapping position of the reads, for example, the mapping position of the initial bases of the sequencing read. In some cases, folding uses barcodes, and sequencing reads that share a barcode sequence are folded into the consensus sequence. In some cases, folding uses both barcodes and the sequence of the initial template molecule. For example, all reads that share a barcode and map to the same position in the reference genome can be folded. In another example, all reads that share a barcode and the sequence of the initial template molecule (or a percentage of identity with respect to the sequence of the initial template molecule) can be folded.
[0168] In some cases, the quantitative measure of sequencing read coverage is determined for specific subregions of the genome. These regions may be bins, target genes, exons, regions corresponding to sequence probes, regions corresponding to primer amplification products, or regions corresponding to primer binding sites. In some cases, the subregion of the genome is the region corresponding to a sequence capture probe. A read may be located in a region corresponding to a sequence capture probe if at least a portion of the read maps to at least a portion of that region. A read may be located in a region corresponding to a sequence capture probe if at least a portion of the read lies in a large portion of that region. A read may be located in a region corresponding to a sequence capture probe if at least a portion of the read lies across the center point of that region.
[0169] In another embodiment involving barcodes, all sequences having the same barcode, physical properties, or a combination of both, if they originate from the sample parent molecule, can be folded into a single read to reduce any bias that may have been introduced during amplification. For example, if one molecule is amplified 10 times while another is amplified 1000 times, the effect of the uneven amplification is canceled out by presenting each molecule only once after folding. Only reads with unique barcodes may be counted for each mappable position and may influence the assigned score.
[0170] The consensus sequence can be generated from a family of sequence reads by any method known in the art. Such methods include, for example, linear or nonlinear methods for constructing the consensus sequence derived from digital communications theory, information theory, or bioinformatics (e.g., voting, averaging, statistical, maximum aposterior or maximum likelihood detection, dynamic programming, Bayesian, hidden Markov, or support vector machine methods).
[0171] After determining the sequence read coverage, probabilistic modeling algorithms can be applied to convert the normalized nucleic acid sequence read coverage for each window / bin region into separate copy number states. In some cases, this algorithm may include one or more of the following: hidden Markov models, dynamic programming, support vector machines, Bayesian networks, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering methodology, and neural networks. By utilizing the separate copy number states for each window region, copy number diversity in chromosomal regions can be identified. In some cases, all adjacent window / bin regions with the same copy number can be merged into a segment to report the presence or absence of copy number diversity states. In some cases, different windows / bins can be filtered before merging them with other segments.
[0172] The data analyzed and / or output by the sequence analysis pipeline 112 can be stored in the analysis data store 117.
[0173] The variant detection pipeline 130 can take in / receive data from the analysis data store 117. For example, the variant detection pipeline 130 can take in / receive data representing multiple sequence reads. Multiple sequence reads can be analyzed by the copy number module 115 and / or the HRD module 300 to determine one or more variants. Variants may include, for example, single nucleotide variants (SNVs), indels, fusions, and copy number diversity. Any known technique for variant calling can be used. In one embodiment, nucleotide diversity in a sequenced nucleic acid can be determined by comparing the sequenced nucleic acid with a reference sequence. The reference sequence is often a known sequence, e.g., a known whole or partial genome sequence from a subject (e.g., a whole genome sequence from a human subject). The reference sequence may be, for example, hG19 or hG38. The sequenced nucleic acid may represent a sequence directly determined for the nucleic acid in the sample, as described above, or it may be a consensus of sequences of amplified products of such nucleic acids. Comparisons can be made at one or more specified positions on the reference sequence. A subset of sequenced nucleic acid sequences can be identified that contain positions corresponding to the specified positions on the reference sequence when each sequence is maximally aligned. Within such a subset, it can be determined which sequenced nucleic acids, if any, contain nucleotide variants at the specified positions; the length of a given cfDNA fragment based on the offset of the midpoint of the given cfDNA fragment from the midpoint of the genomic region within the cfDNA fragment, if its endpoints (i.e., its 5' and 3' terminal nucleotides) are located in the reference sequence; and, if necessary, which contain the reference nucleotide (i.e., the same as that in the reference sequence). If the number of sequenced nucleic acids in the subset containing nucleotide variants exceeds a selected threshold, the variant nucleotides can be called at the specified positions.The threshold may be a simple number such as at least 1, 2, 3, 4, 5, 6, 7, 9, or 10 of the sequenced nucleic acids in the subset containing the nucleotide variant, or the threshold may be a ratio such as at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20 of the sequenced nucleic acids in the subset containing the nucleotide variant, among many other possibilities. The comparison can be repeated for any designated position of interest in the reference sequence. Sometimes, the comparison can be made for designated positions occupying at least approximately 20, 100, 200, or 300 consecutive positions on the reference sequence, for example, approximately 20–500 or approximately 50–300 consecutive positions.
[0174] Generally speaking, the processor 120 may implement (and be programmed by) various components of the variant detection pipeline 130, such as the copy number module 115, the HRD module 300, and / or other components. Alternatively, note that these components of the variant detection pipeline 130 may include hardware modules. Although illustrated separately for convenience, one or more of the various components or instructions, such as the copy number module 115 and the HRD module 300, can be integrated with each other. In any case, the variant detection pipeline 130 can cause the computer system 110 to identify variants, diseases (accurate diagnoses) from the variants, HRDs, and / or treatment regimens. The accurate diagnoses and treatment regimens can be stored in a repository such as the clinical results store 160 or the diagnostic results store 150.
[0175] The HRD module 300 can be configured to analyze output data from the sequence analysis pipeline 112. The HRD module 300 can be configured to generate one or more of the following: de novo fusion rearrangement calls, deletion calls, SNV / indel pathogenicity annotations, and / or HRD scores. The HRD module 300 may include an HRD aggregator configured to generate a summary of sample-level HRD status.
[0176] As shown in Figure 3, the HRD module 300 can be configured to perform one or more of the following: fusion caller 301, deletion caller 302, annotation module 303, HRD scoring module 304, aggregator 305, and / or output module 306.
[0177] The fusion caller 301 can be configured to generate one or more candidate fusion calls by analyzing data received from the sequence analysis pipeline 112. The fusion caller 301 can be configured to assemble candidate fusion reads into a de Bruijn graph, call candidate fusion events, filter candidate fusion calls, and remove technical false positives. The fusion caller 301 can be configured to select fusion candidate reads, cluster the fusion candidate reads into packets, assemble the clustered fusion candidate reads to generate a de Bruijn graph assembly. The fusion caller 301 can be configured to unpack the de Bruijn graph assembly into a fusion candidate contig, align the fusion candidate contig with a reference having a decoy, and generate a candidate fusion call.
[0178] In one embodiment, the fusion caller 301 can be configured to select candidate fusion reads, create an undirected graph G of reads by joining reads with consistent break points, and save the output data as packets to a file. The fusion caller 301 can be configured to hybridize reads and assemble them to generate a de Bruijn graph assembly. The fusion caller 301 can be configured to unfold the assembly into a linear contig. The fusion caller 301 can be configured to align a contig on a reference with decoys and call estimated fusion.
[0179] The fusion caller 301 can be configured to filter candidate fusion calls based on one or more criteria. One or more criteria may include filtering a candidate fusion if any of its breakpoints are 350 bases or more away from one of the probes in the probe set. One or more criteria may include filtering a candidate fusion if none of its breakpoints belong to one of the genes in the gene list. One or more criteria may include filtering a candidate fusion if it consists of two deletions (96 bases or less) that are 48 bases or less away from each other. One or more criteria may include filtering a candidate fusion if it is a deletion of exactly less than 60 bases. One or more criteria may include filtering a candidate fusion if it does not have at least one double-stranded molecular support and its average family size is <1 / 0.9 (about 1.10). The average family size may be defined as the number of supporting reads divided by the number of supporting molecules. One or more criteria may include filtering a candidate fusion if the alignment of a 120-base reference segment centered at a first breakpoint with a 120-base reference segment centered at a second breakpoint has an alignment score of 50 or greater. One or more criteria may include filtering a candidate fusion if the alignment of a segment 36 units away from the first breakpoint with a segment 36 units away from the second breakpoint has an alignment score of 20 or greater. One or more criteria may include filtering a candidate fusion if it does not have a robust support. A molecule may be considered robust if it has a family size of 2 or greater (where family size refers to the number of reads supporting the molecule).
[0180] The fusion caller 301 can be configured to upgrade a candidate fusion to pass if the candidate fusion is reciprocal with one or more other candidate fusions, and the other candidate fusions have not been removed by any other filtering criteria. The fusion caller 301 can also be configured to tag a candidate fusion if one or more of its cut points are in an amplified region.
[0181] The fusion caller 301 can be configured to output fusion data containing filtered fusion calls. The fusion data may include auxiliary data for in-depth analysis of fusion events. The data output by the fusion caller 301 can be read by the annotation module 303 and / or the aggregator 305.
[0182] Figure 4 shows a schematic diagram of a de novo fusion caller (e.g., the fusion (reorganization) caller 301 referenced in Figure 3) according to embodiments of the present disclosure. In the embodiments shown, the fusion caller determines the location of reads and whether a given sample contains an HRD nucleic acid variant by examining a shared break point (step 1), assembling a locibly located bug (step 2), linearizing the locibly located bug into a contig (step 3), and aligning the contig to a reference sequence (step 4). In some embodiments, the de novo fusion caller includes instructions configured to align a plurality of sequence reads to a reference sequence, instructions configured to determine a break point in the alignment of at least one of the plurality of sequence reads to a reference sequence, instructions configured to identify any sequence read associated with a break point in the alignment as a candidate fusion sequence read, and instructions configured to determine a candidate fusion sequence read associated with a common break point among one or more break points. These embodiments may also typically include grouping candidate fusion sequence reads based on one or more common break points. In one embodiment, reads can be grouped (e.g., clustered) based on having a cleavage point within a nucleotide threshold window. Furthermore, instructions can be configured to assemble candidate fusion sequence reads within a group to generate one or more contigs, to align contigs from a group to a reference sequence, to determine one or more candidate fusion events based on the alignment of contigs from a group, to apply one or more criteria to one or more candidate fusion events, and to determine one or more fusion events containing a second LoF HRD nucleic acid variant based on the application of one or more criteria to one or more candidate fusion events.In certain embodiments, the criteria include filtering criteria, e.g., the absence of breakpoints near the probe, the absence of breakpoints in the gene to be reported, the exclusion of small indel and intron events, the performance of a "pc_molecules" test ((pc_molecules=n_molecules / n_reads)), and, in some embodiments, the discarding of fusions having an average family size less than 1.7, artifacts known to be related to address switching, the exclusion of events where they may be "template crossovers," the application of a minimally robust molecule test, and / or similar. Further details regarding de novo fusion callers adapted as necessary for use in performing the methods and related embodiments of this disclosure are described, for example, in U.S. Patent Application No. 16 / 803,680, filed on February 27, 2020, which is incorporated herein by reference in its entirety.
[0183] Returning to Figure 3, the deletion caller 302 can be configured to determine homozygous deletions and heterozygous loss (LOH) at the gene and genome-wide levels by analyzing data received from the sequence analysis pipeline 112 (e.g., copy number data). The deletion caller 302 can also be configured to detect deletions by comparing the coverage of the region of interest with a reference profile generated from a cancer sample that does not contain a deletion.
[0184] In one embodiment, the deletion caller 302 can segment copy number data and identify genomic regions with abnormal copy numbers using a segmentation algorithm (e.g., a circular binary segmentation (CBS) algorithm). The segmentation algorithm can segment copy number data into estimated copy number regions by recursively dividing the chromosome into either two or three subsegments based on the maximum t-statistic. The reference distribution used to determine whether or not to divide can be estimated by permutation. Thus, the segmentation algorithm can find change points in the copy number data. A change point may refer to a point after the (log) sample-to-reference ratio has changed position. Therefore, the change point corresponds to the position where the underlying DNA copy number has changed. Thus, change points can be used to identify copy number gain and loss regions. The output data of the segmentation algorithm may include a table in which rows show the sample, chromosome, start and end map locations, number of markers, and mean values for each segment. Further details regarding the segmentation algorithm are, for example, described in Olshen, AB, Venkatraman, ES, Lucito, R., Wigler, M. (2004). Circular binary segmentation for the analysis of array-based DNA copy number data. Biostatistics 5: 557-572, and Venkatraman, ES, Olshen, AB (2007) A faster circular binary segmentation algorithm for the analysis of array CGH data. Bioinformatics 23: 657-63, the contents of which are incorporated herein by reference in their entirety.
[0185] A deletion caller 302 can be configured to label the deletion once it is detected. For example, deletion caller 302 can be configured to label deletions as "cov_del", "loh", "loh_cn_neutral", or "homdel". Deletion caller 302 can use the mutant allele proportion (MAF) of a predefined list of germline SNPs to determine whether cancer cells have a single copy of the region of interest (labeled "LOH") or whether both copies are deleted (labeled "homdel"). The discriminator for these two cases lies in the observation that a single copy in tumor cells results in an allele imbalance of heterozygous SNP MAFs. The decision rule can be based on a likelihood ratio test in which the likelihood ratios of two models representing the two types of deletions are compared to a threshold estimated from a training set of "undetected target" (TND) samples. The likelihood models can be based on the trained MAF distribution of heterozygous SNPs for all three possible SNP genotypes. The output label used for deletions where no heterozygous SNP is observed may be "cov_del," which represents a third label indicating the predicted deletion.
[0186] In one embodiment shown in Figure 5, the detection caller 302 may generate a deletion call by determining whether a deletion is observed based on comparing the coverage of the region of interest with a reference profile generated from a cancer sample without a deletion in 501. If a deletion is observed in 501, the baseline MAF-adjusted gene coverage and z-score can be analyzed in 502 to determine whether the somatic cell has a gene deletion. If no deletion is reported in 502, the resulting call may be "no deletion" or "no amplification," or "no cnv" in 503. If a deletion is reported in 502, a heterozygous SNP overlapping the gene can be identified in 504. If no heterozygous SNP overlapping the gene is observed in 504, the gene may be called "no heterozygous SNP" to differentiate it, even if there was a deletion and it was "cov_del" in 505. If a heterozygous SNP that overlaps the gene is observed at 504, a heterozygous SNP indicating allele disequilibrium can be determined at 506. If no heterozygous SNP indicating allele disequilibrium is determined at 506, the gene may be called homozygous deletion, "homdel," at 507. If a heterozygous SNP indicating allele disequilibrium is determined for one copy at 506, the gene may be called heterozygosity loss, "LOH," at 508. If a heterozygous SNP indicating allele disequilibrium is determined for two copies at 506, the gene may be called heterozygosity loss without copy number change, "LOH CN NEUTRAL," at 509.
[0187] The deletion caller 302 can be configured to determine which genes to report and / or effective genes. Effective genes can be included in the final deletion call output table. Effective genes may be those that meet the following criteria: the number of good probes is >30, and the 95% detection limit (LoD) of the LOH is <0.3 (with the exception that this threshold may be met for biologically important genes). LOH or homozygous deletion genes to report may be those with a heterozygous SNP detection probability >50%. If the heterozygous SNP detection probability is less than or equal to 50%, the gene may be reported as a coverage-based detection ("cov_del").
[0188] In certain embodiments, a homozygous deletion / LOH fusion (CNV) caller (e.g., deletion caller 302) detects LOH at the gene level, detects homozygous deletions, and / or estimates genome-wide LOH levels in a given sample. In some of these implementations, the homozygous deletion / LOH fusion (CNV) caller typically achieves a detection limit of about 5% and a specificity higher than 99% for detecting HRR gene deletions. In some embodiments, the CNV caller is molecular coverage-based and uses SNP information to distinguish between LOH (allele-imbalanced) and homozygous deletions (50% MAF for SNPs). In certain embodiments, the CNV caller is fragment-size information-based. These implementations involve the use of fragment-size distributions, which can improve the sensitivity / specificity of detecting genes or genomic regions and associated deletions.
[0189] To illustrate with further examples, in some embodiments, the CNV caller includes the use of a first probability that the sequence information under consideration includes a first state, and a second probability that the sequence information includes a second state, wherein the first or second state includes a LoF HRD nucleic acid variant. Typically, these embodiments include generating a first probability that the sequence information includes a first state, generating a second probability that the sequence information includes a second state, comparing the first and second probabilities, and, based on the comparison step, generating a prediction of whether the sequence information includes a first state or a second state. In some of these embodiments, the CNV caller includes generating a first model of allele counts based on one or more germline single nucleotide polymorphism (SNP) locations associated with at least one gene locus in the sequence information, by a first probability distribution. The first model typically represents at least one somatic homozygous deletion. These embodiments also include generating a second model of allele counts based on one or more germline SNP locations associated with a gene locus in the sequence information, using a second probability distribution. The second model generally represents at least one somatic heterozygous deletion. These implementations also include comparing a first output data of the first model with a second output data of the second model, and, based on the comparison, generating a prediction that a somatic homozygous deletion for a gene locus is present in the sequence information. In certain embodiments, the CNV caller generates a first probability that the sequence information under consideration contains a somatic homozygous deletion, and a second probability that the sequence information contains a somatic heterozygous deletion, compares the first and second probabilities, and, based on the comparison, generates a prediction that the sequence information contains either a somatic homozygous deletion or a somatic heterozygous deletion. Further details regarding CNV callers as may be adapted for use in carrying out the methods and related embodiments of this disclosure are, for example, described in Ordinary U.S. Patent Application No. 16 / 803,680, filed on 27 February 2020, which is incorporated herein by reference in its entirety.
[0190] The deletion caller 302 can be configured to output deletion call data. Deletion call data may, for example, show labels (e.g., no_cnv, cov_del, loh, loh_cn_neutral, homdel, no_call, and similar). Deletion call data may, for example, show loh / homdel for reported genes and when the prediction is based on a single heterozygous SNP. Deletion call data may, for example, show when multiple genes on different chromosomes are labeled "homdel" (e.g., a latent baseline MAF error not limited to the gene to be reported). Deletion call data may, for example, show genes with somatic SNVs / indels that are labeled "homdel" (e.g., a latent label error). Deletion call data may include all CNV calls (amplification, locality, deletion, aneuploidy, etc.) for all genes on the panel. Deletion call data may include all genes on the panel that have one of the following states: homozygous deletion, LOH, coverage-based deletion, LOH without copy number change, or no call. Deletion call data may include all genes to be reported that have one of the following states: homozygous deletion, LOH, coverage-based deletion, LOH without copy number change, or no call. Deletion caller 302 can be configured to determine and output one or more plots to summarize gene-level LOH or deletions, including coverage and SNPs. Plots may include genome-wide (all linked regions) and chromosome-level CNV plots, highlighted by probes by the HRR gene and whether the HRR gene is reported as cov_del, homdel, or loh deletion. Deletion call data may include percentages of segments with LOH or deletions. Deletion call data may include segments used for deletion calling. The data output by the deletion caller 302 can be read by the HRD scoring module 304 and / or the aggregator 305.
[0191] Returning to Figure 3, annotation module 303 can be configured to provide data demonstrating clinical significance, such as relationships between variants and phenotypes (e.g., ClinVar data). Annotation module 303 can be configured to provide annotations of functional effects for some or all SNVs / indels. Annotation module 303 can be configured to indicate somatic calls (e.g., somatic or germline calls) and / or functional effects of fusions being called (e.g., those considered harmful due to GH or revertant mutations).
[0192] The annotation module 303 can analyze data received from the sequence analysis pipeline 112 and / or fused data received from the fusion caller 301. The annotation module 303 can perform the annotation method 600 shown in Figure 6.
[0193] Annotation module 303 can determine clinically significant data in 601. Annotation module 303 can determine clinically significant data by retrieving data related to variants in a sample from a data source. For example, clinically significant data may be ClinVar data retrieved from a remote or local data source. For example, data from the "CLNSIG" field for that variant in the ClinVar VCF file can be determined by matching for chromosome ("chrom"), location ("pos"), variant nucleotide ("mut_nt"), and gene name ("gene"). Clinically significant data can be associated with the variant. Clinically significant data may include, for example, benign, likely benign, uncertain positive, likely pathogenic, pathogenicity, drug response, association, risk factor, protective, emotional, conflicting data from submitters, and / or may not be provided. Data demonstrating clinical significance may also relate to the review status indicating the accuracy of the data, such as no definitive conclusion, no definitive criteria provided, no definitive conclusion for individual variants, criteria provided (single submitter), criteria provided (contradictory interpretations), criteria provided (multiple submitters, no contradictions), reviewed by an expert panel, and / or clinical practice guidelines.
[0194] Annotation module 303 can be configured in 602 to determine the molecular consequences for some or all SNVs / indels. Molecular consequences can be assigned to some or all variants in a sample based on the application of a set of rules. If no rule applies, the molecular consequence may be "NULL". These rules can be applied in order of priority, i.e., starting from the highest priority, and if one rule applies, the remaining rules can be ignored. Molecular consequences may be, for example, nonsense, frameshift, stop loss, start loss, in-frame insertion, in-frame deletion, in-frame duplication, missense, splice acceptor, splice donor, synonym, noncode, utr, promoter, splice region, splice event, and similar.
[0195] Annotation module 303 can be configured to determine the annotation of functional effects for some or all SNVs / indels in 603. Variants (SNVs / indels) in a sample may be annotated as “harmful” or “NULL”. For example, SNVs / indels in any gene associated with clinically significant data showing that the mutation may be causative or strongly associated with disease (e.g., “pathogenic” or “likely pathogenic”) and / or strongly associated with a response to therapy (e.g., “drug response”). SNVs / indels in known HRR genes, or in known tumor suppressor genes that are nonsense, frameshift, splice acceptor, or splice donor, are uncommon (e.g., population allele frequency <0.001), and are not present in BRCA2 with more than 3326 codons, may be annotated as harmful. Any remaining variants may be annotated as “NULL”.
[0196] Annotation module 303 can be configured to determine revertant variants (e.g., small variants) in 604. Revertant variants can repair the function of genes disrupted by pathogenic alleles. Generally, a revertant variant is determined if any of the following criteria are met: • SNVs that are located within the same codon as a harmful SNV, and that restore the codon to a non-harmful state. • In-frame indels that extend to the entire location of harmful SNVs / indels • Indels that are within the distance threshold along with frameshift indels and that return proteins to the frame. • Indels that are far from the frameshift indel and whose proteins are identified as cis or harmful trans, then return to the frameshift indel. • A long somatic deletion that extends to another harmful SNV or indel within the same gene.
[0197] Figure 7 shows an example method 700 for determining reversion variants. According to method 700, variants for SNVs and small indels that are classified as germline but meet the requirements for reversion may be flagged for manual review.
[0198] An indel that is upstream or downstream of another frameshift indel and brings a protein back into frame (the sum of inserted or deleted nucleotides is a multiple of 3) is called "reversion_cis" if it can be confirmed that the indel is cis (shares the same molecule) with the second frameshift indel. If the variant is confirmed to be trans (on a different molecule), the annotation is called "deleterious_trans". If the supporting molecule does not span both indels to determine the cis / trans annotation and the protein returns to frame, the annotation is simply "reversion".
[0199] An SNV that is in the same codon as a harmful SNV (e.g., nonsense, pathogenic missense) and reverts the codon to a non-harmful state (e.g., nonsense to missense, pathogenic missense to synonymous) may be labeled as reverted. A determination may be made as to whether an SNV is in the same genomic location as a harmful SNV, and if the SNV is in the same genomic location as a harmful SNV and the result is non-harmful (e.g., "synonymous"), then the SNV may be labeled as reverted.
[0200] An indel that is in-frame and spans another harmful SNV or indel may be reverted and labeled.
[0201] Annotation module 303 may also be configured to determine reversion variants (e.g., long deletion variants) in 604. Fusions and long deletions from fusion data received from fusion caller 301 may be used to annotate somatic cell calls and functional effects. A long deletion may be defined as a large genomic rearrangement resulting in the loss of a DNA sequence within the same gene. Somatic classification of long deletions may be performed by annotation module 303. Fusions and long deletions may be annotated as somatic if the variant percentage is below a constitutable threshold (e.g., <15%), and as germline if the variant percentage is above a constitutable threshold. Long deletions that meet the requirements for reversion may have a somatic cell call of “somatic” regardless of the variant percentage. A long deletion may have a “reversion” functional effect if the long deletion is somatic and spans another harmful SNV or indel within the same gene.
[0202] A long deletion having a variant percentage above a configurable threshold (e.g., ≥15%) that satisfies the requirements for reversion may have a “reversion” functional effect and a “somatic” somatic call. This somatic status overrides the original germline call determined by the configurable threshold. In some embodiments, all long deletions may have a “harmful” functional effect annotation unless the long deletion is considered reversion. In some embodiments, all fusions occurring between at least two different genes may have a “harmful” functional effect annotation.
[0203] Annotation module 303 may be configured to output annotation data. Annotation data may include, for example, fusion tables showing somatic calls and functional effects. Annotation data may include SNV call data including some or all SNV results and / or some or all indel results from copy number module 115, with, for example, clinical significance data, adverseness annotations, and / or reversion annotations. Annotation data may include data associated with SNV / indel calls, with, for example, clinical significance annotations, molecular consequences, and / or functional effects. Data output from annotation module 303 may be read by HRD scoring module 304 and / or aggregator 305.
[0204] Returning to Figure 3, the HRD scoring module 304 may be configured to analyze and / or summarize data from the fusion caller 301, deletion caller 302, and / or annotation module 303 to generate an HRD score (e.g., a metric for measuring HRD). The HRD scoring module 304 may also be configured to analyze data from some or all of the somatic variant outputs from the sequence analysis pipeline 112 and / or variant detection pipeline 130 to calculate the maximum somatic allele proportion (msaf).
[0205] The HRD scoring module 304 may be configured to generate an HRD score by utilizing, at least in part, the number and / or nature of the rearrangements and / or the array context around the rearrangement breakpoints, as determined by the fusion caller 301.
[0206] The HRD scoring module 304 may be configured to generate an HRD score by summarizing, at least in part, the number of segments with breakpoints and / or deletions per sample, as determined by the deletion caller 302. Segments may represent different copy number states in the genome, more copy number states, more genomic instability, and potentially, underlying HRD deficiencies. Generally, the HRD scoring module 304 may be configured for: • Smooth segments less than 3MB in length between segments that are within 1.5 standard deviations of the segment mean. • Remove the segment that accounts for 90% of the chromosome's length. • Filter out segments that are less than a configurable length threshold (10 Mb) to eliminate segments that may be products of non-HRD mechanisms. • Count the number of break points between adjacent segments. • Filter out segments from the deletion caller for repeating segments (which may not be informative) found in normal or tumor-free samples. • Count the number of segments with deletions (LOH, homdel, or coverage-based deletions). • Add up the number of cutting points and missing segments.
[0207] In one embodiment, the determination of the HRD score may be based on one or more metrics that correlate with the HRD status. For example, one or more metrics may include one or more of the following: loss of heterozygosity (LOH), telomere allele imbalance (TAI), large-scale transitions (LST), or a combination thereof.
[0208] LSTs can refer to breakpoints between adjacent regions of at least 10 Mb after filtering out a 3 MB region. The number of LSTs correlates with the number of adjacent breakpoints and the gene mutation status. The 3 MB cutoff can often be used to remove smaller variations unrelated to HRD from large variations that usually indicate interchromosomal translocations.
[0209] In one embodiment, the HRD scoring module 304 may be configured to determine the HRD score according to method 800 shown in Figure 8. The HRD scoring module 304 may have access to deletion call data, including segments used for deletion calling. Segments spanning the length of a chromosome (>90%) may be removed in 801. These segments are more likely to result from chromosomal nondisjunction rather than HRD. Segments less than 3 MB in length may be smoothed in 802. Segment smoothing may involve combining segments that are not more than 3 MB apart, with the second segment having a segment mean (e.g., normalized coverage in the segment) within 1.5 standard deviations of the previous segment. Segments may be filtered based on size in 803. For example, segments less than 10 Mb may be removed. Segments with small diversity typically correspond to intrachromosomal rearrangements unrelated to HRD, in contrast to LSTs, which mostly indicate interchromosomal translocations. Segments in repeated (e.g., >10 times) bins in a large cohort of samples (including patient samples, samples without detected tumors, and healthy normal samples) that are likely to exhibit technical artifacts may be removed in 804. The bin size may be, for example, 500 bp from either the start or end point of the segment. The number of breakpoints between adjacent segments within a certain bp distance may be determined in 805. The number of segments with LOH, Homdel, and / or deletion labels may be determined in 806. The number of breakpoints and the number of segments may be summed in 807 to determine the HRD score.
[0210] The HRD scoring module 304 may be configured to determine the tumor percentage from a sample, for example, using the maximum somatic allele percentage (MSAF). MSAF may include the maximum percentage of variants in the sample, expressed as a percentage, including any somatic variants not annotated as having a clonal hematopoietic origin, and including any somatic variants that are fusions or non-synonymous SNVs or indels. MSAF may be 0 if no somatic variants are present in SNVs, indels, fusions, or de novo fusion outputs. For variants present on amplified genes on the X chromosome in male samples, the percentages may be adjusted to account for haploid chromosomes, as follows: Adjusted percentage = Variant percentage / log2(gene CN) * 2) Here, sex can be predicted by the sequence analysis pipeline 112. For all other variants present on the amplified gene, the percentages can be adjusted as follows: Adjusted percentage = Variant percentage / log2(gene CN). The tumor percentage may be used in the HRD scoring module 304 to obtain an adjusted estimate of the HRD score for samples with low tumor shedding.
[0211] Figure 9 shows a histogram of HRD scores representing examples across cancer types. Clinical patient samples from the OMNI 2.12 panel were evaluated for HRD scores (n=200 for all cancer types except ovarian and cutaneous, with n=139 and n=113 for ovarian and cutaneous cancers, respectively). Longer tails (>100) in HRD scores were observed in breast and genitourinary cancer types.
[0212] The HRD score module 305 may be configured to output HRD score data. The HRD score data may include, for example, the HRD score and / or MSAF. In one embodiment, the MSAF may provide information about the score (for example, a low tumor percentage and a low score may exceed a threshold under certain conditions, while a high tumor percentage and a low score may not).
[0213] The HRD score can be compared to a threshold. If a sample's HRD score exceeds the threshold, the sample can be determined to be HRD-positive. The threshold can be determined empirically through analysis of tumor type populations considered HRD+ve and HRD-ve by the presence of specific loss-of-function genomic biomarkers (e.g., BRCA1 / 2 biallelelic inactivation), or based on tumor type populations that clinically responded to or did not respond to PARP inhibitors (responders are likely to have higher HRD scores).
[0214] Returning to Figure 3, the aggregator 305 may be configured to provide a sample-level summary from other HRD modules (fusion caller 301, deletion caller 302, annotation 303, and / or HRD scoring module 304) and / or to determine HRR genes in the sample that have biallelic inactivation. The aggregator 305 may be configured to analyze data from fusion caller 301, data from deletion caller 302 (e.g., deletion call data), data from annotation module 303 (e.g., annotation data), and / or data from HRD scoring module 304 (e.g., HRD score data).
[0215] In one embodiment, the aggregator 305 can receive / retrieve annotation data including de novo fusion calls and somatic calls with functional effects; SNV calls including some or all SNV results from copy number module 115 with clinical significance annotations, harmful annotations, and / or reversion annotations; and / or data including SNV / indel calls with clinical significance annotations, molecular consequences, and functional effects. In one embodiment, the aggregator 305 can receive / retrieve deletion call data including some or all reportable genes having one of the following states: homozygous deletion, LOH, coverage-based deletion, LOH without copy number change, or no call; and / or genome-wide LOH call data including genome-wide LOH calls based on segments. In one embodiment, the aggregator 305 can receive / retrieve HRD score data including HRD scores and / or maximum somatic allele percentages. In one embodiment, the aggregator 305 can receive / retrieve fusion data, including some or all of the fusions detected by the fusion caller 301, as well as auxiliary data for in-depth analysis of fusion events.
[0216] In one embodiment, the aggregator 305 may determine, for example, the total number of rearrangements in the sample, often attributed to loss of BRCA1 / 2 function, or a subset of these rearrangements having HRD features / signature features, such as tandem duplication or deletion, clustered and unclustered deletions (>100kb), inversions, and interchromosomal translocations. In another embodiment, the aggregator 305 may determine the total number of indels in the sample, also previously attributed to loss of BRCA1 / 2 function, or a subset of these indels having adjacent sequence contexts with microhomology indicating the underlying HRD.
[0217] Biallelic inactivation occurs when both copies of a gene exhibit loss of function; this can occur through the presence of pathogenic variants, deletions, or rearrangements in both alleles of the gene. Patients with biallelic inactivation have been shown to have a stronger HRD phenotype compared to patients with only one allele inactivated, and may show improved clinical benefits compared to uniallelic inactivation when treated with PARP. Aggregator 305 may be configured to determine whether a gene in a sample is associated with biallelic inactivation, and if the gene is an HRR gene, at least one of the following is true: • The gene has at least two different harmful SNVs or indels. • The gene is in at least one fusion / reorganization state and has at least one harmful SNV or indel. • The gene has at least one harmful SNV or indel and has a LOH or coverage base deletion. • The gene is in at least one fusion / reorganization state and has a LOH or coverage base deletion. • Genes exist in two different fusion / reorganization states (this does not include reciprocal fusions or identical fusion gene pairs with different cleavage points). • The gene has a homozygous deletion.
[0218] Aggregator 305 may be configured to determine / search / receive a list of HRR genes. Aggregator 305 may be configured to flag samples that have homozygous deletions and fusions detected in the same HRR gene.
[0219] The aggregator 305 may be configured to generate data summarizing sample-level HRD information, including (but not limited to) the number of biallelelic mutations, HRD score, and maximum somatic allele percentage (MSAF).
[0220] Output module 306 may be configured to output a user-friendly summary of some or all variants called by copy number module 115 and HRD module 300 for manual review. Output module 306 may be configured to generate sample-level data including metrics from the outputs of both copy number module 115 and HRD module 300. Output module 306 may be configured to generate a report summarizing the manual review flags from HRD module 300. The report may indicate samples and variants that require manual review. Output module 306 may be configured to generate a report including sample-level and variant-level QC metrics. Output module 306 may be configured to obtain sample-level and variant-level QC metrics to generate a report. Output module 306 may be configured to generate a report summarizing variant calls and manual review comments from both copy number module 115 and HRD module 300. The output module 306 may be configured to generate a table of missing calls and corresponding cutoff thresholds.
[0221] Figure 10 is a flowchart illustrating exemplary method steps for generating one or more HRD scores (e.g., using the HRD module 300 of System 100) and detecting HRD in a test subject, according to some embodiments. As shown, Method 1000 includes the step (Step 1001) of generating a set of reference HRD scores for genes in a set of HRR genes (e.g., homologous recombination repair (HRR) genes) from sequence information derived from cell-free nucleic acid (cfDNA) obtained from a reference subject having a given oncology type. In some embodiments, the set of HRR genes is selected from those listed in Table 1. A given reference HRD score typically includes the prevalence of a given HRD nucleic acid variant. The reference HRD scores can be generated based on the set of reference HRD scores (Step 1002). The reference HRD scores can then be used to detect HRD in a test subject. As shown in Method 1000, this generally includes the step (Step 1003) of generating test HRD scores for genes in a set of HRR genes from sequence information derived from cfDNA obtained from a test subject having a given cancer type, thereby producing a set of test HRD scores. A given test HRD score typically includes the prevalence of a given HRD nucleic acid variant. In some embodiments, a given HRD nucleic acid variant causes monoallelic or biallelic inactivation of the corresponding HRR gene. To detect HRD in a test subject, Method 1000 also includes the step (Step 1004) of generating test HRD scores from the set of test HRD scores and the step (Step 905) of detecting HRD in the test subject if the test HRD scores exceed a reference HRD score.
[0222] Figure 11 is a flowchart illustrating exemplary method steps for determining the HRD status of a test subject having a given cancer type (e.g., using the HRD module 300 of system 100) according to some embodiments. As shown, method 1100 includes the step (step 1101) of generating test HRD scores for genes in a set of HRR genes (e.g., homologous recombination repair (HRR) genes) from sequence information derived from cell-free nucleic acid (cfDNA) obtained from the test subject to produce a set of test HRD scores. A given test HRD score generally includes the prevalence of a given HRD nucleic acid variant. In some embodiments, the set of HRR genes is selected from those listed in Table 1. Method 1100 also includes the step (step 1102) of generating test HRD scores from the set of test HRD scores. Furthermore, Method 1100 also includes a step (step 1103) of comparing study HRD scores with reference HRD scores, wherein study HRD scores that are higher than the reference HRD score indicate that those study HRD scores are from study subjects who have HRD, and study HRD scores that are at or below the reference HRD score indicate that those study HRD scores are from study subjects who do not have HRD, thereby determining the HRD status of a study subject having a given cancer type.
[0223] Method 1100 is a method that generates a set of reference HRD scores, comprising the steps of generating reference HRD scores for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from one or more reference subjects having one or more oncologies, wherein a given reference HRD score includes the prevalence of a given HRD nucleic acid variant, and generating reference HRD scores from the set of reference HRD scores.
[0224] To further illustrate, Figure 12 is a flowchart illustrating exemplary method steps for detecting HRD in a subject (e.g., using the HRD module 300 of System 100) according to some embodiments. As shown, Method 1200 includes the step (Step 1201) of detecting HRD in a subject by determining the presence or absence of at least one HRD nucleic acid variant in sequence information derived from cell-free nucleic acid (cfDNA) obtained from the subject using (i) a first probability that the sequence information includes a first state, and a second probability that the sequence information includes a second state, wherein the first or second state includes at least the first HRD nucleic acid variant (e.g., using a CNV caller as described herein), and / or (ii) one or more aligned contigs generated from the sequence information, wherein the aligned contigs include at least the second HRD nucleic acid variant (e.g., using a de novo fusion caller as described herein). Some embodiments of Method 1200 involve using only one of steps (i) to (ii), while other embodiments involve using each of steps (i) to (ii).
[0225] Figure 13 is a flowchart illustrating exemplary method steps for treating a disease in a subject, according to some embodiments. As shown, Method 1300 is detected by step 1301, which includes the step of administering one or more therapies (e.g., PARP inhibitors, BER inhibitors, etc.) to a subject having a disease (e.g., a given cancer type) and a disease-associated DNA damage repair deficiency (DDRD) (e.g., HRD), wherein the DDRD is detected by determining the presence of at least one HRD nucleic acid variant in sequence information derived from cell-free nucleic acid (cfDNA) obtained from the subject using (i) a first probability that the sequence information includes a first state and a second probability that the sequence information includes a second state, wherein the first or second state includes at least the first HRD nucleic acid variant (e.g., using a CNV caller as described herein), and / or (ii) one or more aligned contigs generated from the sequence information, wherein the aligned contigs include at least the second HRD nucleic acid variant. Some embodiments of Method 1300 involve using only one of steps (i) to (ii), while other embodiments involve using each of steps (i) to (ii).
[0226] In some embodiments, the first HRD nucleic acid variant includes homozygous deletions, heterozygous loss-of-heterozygous (LOH) variants (e.g., gene-specific LOH variants, LOH variants without copy number changes, and / or genome-wide LOH variants), copy number variations (CNVs), etc. In certain embodiments, the second HRD nucleic acid variant includes structural rearrangements (e.g., shortened rearrangements, multi-exon deletions, etc.). In some embodiments, the first and / or second HRD nucleic acid variants include single nucleotide variations (SNVs), indels, etc.
[0227] The techniques in steps (i) and / or (ii) of these methods may include at least aligning a segment of sequence information to at least one reference sequence. These methods may include using only one of steps (i) to (ii). These methods may include using each of steps (i) to (ii).
[0228] At least one homologous recombination repair (HRR) gene in these methods may contain an HRD nucleic acid variant. The HRR gene in these methods may be selected from the group consisting of ATM, ATR, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, NBN, PALB2, RAD51, RAD51B, RAD51C, RAD51D, RAD54L, HDAC2, MRE11, PPP2R2A, XRCC5, WRN, MLH1, FANCC, BAP1, XRCC2, XRCC3, and RAD50. A set of HRR genes may contain at least approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more genes.
[0229] One or more HRD nucleic acid variants in these methods may cause biallelelic inactivation of a given HRR gene. One or more HRD nucleic acid variants in these methods may cause monoallelelic inactivation of a given HRR gene. HRD nucleic acid variants in these methods may correlate with subjects having the disease. Subjects may be unknown as to whether they have the disease. Subjects may be known to have the disease. The disease may be cancer.
[0230] These methods may include the step of administering one or more therapies to a subject to treat a disease. Therapies may include at least one poly-ADP-ribose polymerase (PARP) inhibitor. Therapies may include at least one base excision repair (BER) inhibitor.
[0231] Step (i) of these methods may include generating a first probability that the sequence information contains a first state, generating a second probability that the sequence information contains a second state; comparing the first and second probabilities; and, based on the comparison, generating a prediction about whether the sequence information contains a first state or a second state.
[0232] Step (i) of these methods may include generating a first model of allele counts based on one or more germline single nucleotide polymorphism (SNP) locations associated with at least one gene locus in sequence information, the first model representing at least one somatic homozygous deletion, by a first probability distribution; generating a second model of allele counts based on one or more germline SNP locations associated with gene locus in sequence information, the second model representing at least one somatic heterozygous deletion, by a second probability distribution; comparing a first output of the first model with a second output of the second model; and, based on the comparison, generating a prediction that a somatic heterozygous deletion for a gene locus is present in the sequence information.
[0233] Step (i) of these methods may include generating a first probability that the sequence information contains a somatic homozygous deletion, generating a second probability that the sequence information contains a somatic heterozygous deletion, comparing the first and second probabilities, and, based on the comparison, generating a prediction of whether the sequence information contains a somatic homozygous deletion or a somatic heterozygous deletion.
[0234] Step (ii) of these methods may include aligning multiple sequence reads to a reference sequence, determining one or more breakpoints in the alignment of at least one of the multiple sequence reads to the reference sequence, identifying any sequence reads associated with one or more breakpoints in the alignment as candidate fusion sequence reads, determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints, grouping the candidate fusion sequence reads based on one or more common breakpoints, assembling the candidate fusion sequence reads within each group to generate one or more contigs, aligning the contigs from each group to a reference sequence, determining one or more candidate fusion events based on the alignment of the contigs from each group, applying one or more criteria to one or more candidate fusion events, and determining one or more fusion events containing a second HRD nucleic acid variant based on the application of one or more criteria to one or more candidate fusion events.
[0235] These methods may include a step of detecting HRD, or HRD in a subject, using a CNV and / or de novo fusion caller. The gene may contain an HRD nucleic acid variant.
[0236] In one embodiment, a method 1400 for determining HRD status is disclosed, as shown in Figure 14. In one embodiment, the sequence QC component 113, copy number module 115, and / or HRD module 300 may be configured, individually and / or in combination, to access the sequence read data store 150 and / or analysis data store 117 to perform method 1400 in whole and / or in part. Method 1400 may be performed in whole or in part by a single computer device, multiple computer devices, etc. Method 1400 may include, in step 1401, the step of determining sequence data of a biological sample. The biological sample may include cell-free DNA (cfDNA). Method 1400 may include, in step 1402, the step of determining coverage data based on the sequence data. Method 1400 may include, in step 1403, the step of determining one or more breakpoints associated with one or more fusion events based on the coverage data. Method 1400 may include, in step 1404, a step of determining one or more deletions associated with one or more genes based on coverage data. Method 1400 may include, in step 1405, a step of determining a homologous recombination deletion (HRD) score based on one or more breakpoints and one or more deletions. Method 1400 may include, in step 1406, a step of classifying biological samples based on the HRD score. Method 1400 may include, in step 1406, a step of classifying biological samples as HRD positive based on the HRD score. Method 1400 may include, in step 1406, a step of classifying biological samples as HRD negative based on the HRD score.
[0237] The step of determining the sequence data of a biological sample may include sequencing a panel of one or more HRR genes. One or more HRR genes may be selected from the group consisting of ATM, ATR, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, NBN, PALB2, RAD51, RAD51B, RAD51C, RAD51D, RAD54L, HDAC2, MRE11, PPP2R2A, XRCC5, WRN, MLH1, FANCC, BAP1, XRCC2, XRCC3, and RAD50.
[0238] Biological samples may be associated with subjects having a disease. The disease may be cancer. Coverage data may be associated with multiple bins. Multiple bins may represent regions of a chromosome.
[0239] The step of determining one or more breakpoints associated with one or more fusion events based on coverage data may include aligning multiple sequence reads from sequence data to a reference sequence; determining one or more breakpoints in the alignment of multiple sequence reads to the reference sequence for multiple sequence reads; identifying any sequence reads associated with one or more breakpoints in the alignment as candidate fusion sequence reads; determining candidate fusion sequence reads associated with a common breakpoint among the one or more breakpoints; grouping the candidate fusion sequence reads based on one or more common breakpoints; assembling the candidate fusion sequence reads in the group to generate one or more contigs; aligning the contigs from one of the groups to a reference sequence; determining one or more candidate fusion events based on the alignment of the contigs from the groups; applying one or more criteria to one or more candidate fusion events; and determining one or more fusion events based on the application of one or more criteria to one or more candidate fusion events.
[0240] The step of determining one or more deletions associated with one or more genes based on coverage data may include determining multiple segments based on coverage data, the multiple segments being separated by change points. Determining multiple segments may include applying a segmentation algorithm, which may include a circular binary segmentation algorithm. Change points may correspond to locations where coverage data indicates a change in the underlying DNA copy number. One or more deletions may include one or more homozygous deletions or heterozygous loss of heterozygosity (LOH) deletions.
[0241] Method 1400 may further include the steps of: comparing multiple segments with a reference sequence to identify a subset of multiple segments containing at least one deletion; removing any segments over the length of a chromosome from the subset of multiple segments; combining any segments in the subset of multiple segments that are not separated by more than a threshold distance; removing any segments from the subset of multiple segments that are less than a threshold length; and removing any segments from the subset of multiple segments that are related to technical artifacts. Method 1400 may further include the step of determining the number of breakpoints between adjacent segments within a threshold based on one or more remaining segments in the subset of multiple segments and based on one or more breakpoints related to one or more fusion events.
[0242] Method 1400 may further include the step of determining the number of segments related to a single copy of the region of interest, or related to both copies of the region of interest that are to be deleted, based on one or more remaining segments from a subset of multiple segments.
[0243] The step of determining the HRD score based on one or more breakpoints and one or more deletions may include summing the number of breakpoints and the number of segments.
[0244] Method 1400 may further include a step of determining the presence of one or more genomic rearrangements based on sequencing data. The step of determining the HRD score may further be based on one or more genomic rearrangements. The step of determining the HRD score may include summing the number of breakpoints, the number of segments, and the number of genomic rearrangements.
[0245] Method 1400 may further include a step of determining the maximum somatic allele percentage (MSAF). The step of determining the MSAF may further include determining the maximum percentage of variants in a biological sample, based on sequence data, including any somatic variants not annotated as having a clonal hematopoietic origin and including any somatic variants that are fused or non-synonymous SNVs or indels.
[0246] Method 1400 may further include the step of annotating one or more variants contained in the sequence data. The step of annotating one or more variants contained in the sequence data may include determining annotations of clinical significance related to the effect of one or more variants on human health.
[0247] Method 1400 may further include the step of aggregating sequence data, coverage data, one or more cuts, one or more deletions, and HRD scores. Method 1400 may further include the step of outputting the aggregated sequence data, coverage data, one or more cuts, one or more deletions, and HRD scores.
[0248] The step of classifying a biological sample as HRD-positive based on the HRD score may include determining that the HRD score exceeds a threshold. The step of classifying a biological sample as HRD-negative based on the HRD score may include determining that the HRD score does not exceed a threshold. Method 1400 may further include determining a threshold based on one or more reference HRD scores. The threshold may include reference HRD scores. Sequence information derived from cell-free nucleic acid (cfDNA) obtained from one or more reference subjects may be used to generate a set of reference HRD scores. The reference subjects may have the same condition as the subject from which the biological sample was taken. For example, the reference subject and the subject from which the biological sample was taken may have the same disease (e.g., cancer and / or cancer type). The reference HRD score may be generated from a set of reference HRD scores, for example, by taking the mean (or other statistical analysis) of the set of reference HRD scores.
[0249] Method 1400 may further include a step of administering therapy based on the step of classifying a biological sample as HRD-positive. The therapy may be a poly-ADP-ribose polymerase (PARP) inhibitor or a base excision repair (BER) inhibitor. The PARP inhibitor may be at least one of veliparib, olaparib, talazoparib, lucaparib, niraparib, pamiparib, CEP9722, E7016, E7449, or 3-aminobenzamide. The therapy may be a combination of a PARP inhibitor and radiotherapy.
[0250] The various processing operations and / or methods shown in the drawings may be achieved using some or all of the system components described in detail herein, and in some implementations, the various operations may be performed in different sequences, and some operations may be omitted. Further operations may be performed in conjunction with some or all of the operations shown in the flow diagrams provided. One or more operations may be performed simultaneously. Thus, the operations exemplified (described in further detail herein) are provided as examples and should not be considered limiting.
[0251] Computer implementation
[0252] The methods of the present invention may be carried out on a computer, such that any or all operations described herein and in the appended claims, other than the wet chemical steps, can be performed on a appropriately programmed computer. The computer may be a mainframe, personal computer, tablet, smartphone, cloud, online data storage, remote data storage, etc. The computer may be operated in one or more locations.
[0253] Various operations of the present invention can utilize information and / or programs stored on computer-readable media (e.g., hard drives, auxiliary memory, external memory, servers; databases, portable memory devices (e.g., CD-R, DVD, ZIP disk, flash memory card), etc.) and / or generate results.
[0254] The disclosure also includes a product for analyzing nucleic acid populations, comprising a machine-readable medium containing one or more programs that, if executed, perform steps of the method of the present invention.
[0255] This disclosure may be implemented in hardware and / or software. For example, different embodiments of this disclosure may be implemented in either client-side logic circuits or server-side logic circuits. This disclosure or its components may be embodied in fixed-media program components that, when loaded into a properly configured computer device, cause that device to perform in accordance with this disclosure. The fixed media containing the logic instructions may be delivered to the viewer on a fixed medium for physical loading into the viewer's computer, or the fixed media containing the logic instructions may reside on a remote server accessed by the viewer via a communication medium for downloading the program components.
[0256] This disclosure provides a computer control system programmed to carry out the methods of this disclosure. Returning to Figure 1, the processor 120 may include a single-core or multi-core processor, or multiple processors for parallel processing. The storage device 122 may include random-access memory, read-only memory, flash memory, hard disk, and / or other types of storage. The computer system 110 may include a communication interface (e.g., a network adapter) for communicating with one or more other systems and peripheral devices, such as caches, other memories, data storage, and / or electronic display adapters. The components of the computer system 110 may communicate with each other via an internal communication bus, such as a motherboard. The storage device 122 may be a data storage unit (or data repository) for storing data. The computer system 110 may be operably connected to a computer network ("Network") with the help of the communication interface. The Network may be the Internet, the Internet and / or extranet, or an intranet and / or extranet communicating over the Internet. In some cases, the Network is a telecommunications and / or data network. The Network may include a local area network. The network may include one or more computer servers that can enable distributed computing, such as cloud computing. In some cases, the network may implement a peer-to-peer network that, with the help of computer system 110, allows devices connected to computer system 120 to act as clients or servers.
[0257] The processor 120 can execute a sequence of machine-readable instructions that may be embodied in a program or software. Instructions may be stored in a memory location, for example, in a storage device 122. Instructions may be directed to the processor 120, thereby allowing it to be subsequently programmed or otherwise configured to perform the methods of this disclosure. Examples of operations performed by the processor 120 may include fetching, decoding, executing, and writing back.
[0258] The processor 120 may be part of a circuit, such as an integrated circuit. One or more other components of the system 100 may be included in the circuit. In some cases, the circuit may include an application-specific integrated circuit (ASIC).
[0259] The storage device 122 may store files, such as drivers, libraries, and saved programs. The storage device 122 may also store user data, such as user preferences and user programs. In some cases, the computer system 110 may include one or more additional data storage units located on a remote server that communicates with the computer system 110 externally, for example, via an intranet or the internet.
[0260] Computer system 110 can communicate with one or more remote computer systems via a network. For example, computer system 110 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad®, Samsung® Galaxy Tab), phones, smartphones (e.g., Apple® iPhone®, Android®-enabled devices, Blackberry®), or personal digital assistants. A user can access computer system 110 via a network.
[0261] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of computer system 110, such as storage device 122. The machine executable or machine readable code can be provided in the form of software (e.g., a computer readable medium). During use, the code can be executed by processor 120. In some cases, the code can be retrieved from storage device 122 for immediate access by processor 120 and stored on storage device 122.
[0262] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language selected to enable execution of the code in a pre-compiled or as-compiled manner.
[0263] Aspects of the systems and methods provided herein, for example, computer system 110, may be embodied in programming. Various aspects of the technology may typically be considered a "product" or "manufacture" in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine executable code may be stored on an electronic storage unit, for example, memory (e.g., read only memory, random access memory, flash memory) or a hard disk.
[0264] "Storage" type media can include any or all tangible memories of a computer, processors, etc., or associated modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage for software programming at any time. All or portions of the software may be communicated intermittently via the Internet or various other remote communication networks. Such communication can enable, for example, the loading of software from one computer or processor to another, such as from a management server or host computer to an application server computer platform. Thus, another type of media that can carry software elements includes light waves, radio waves, and electromagnetic waves such as those used across physical interfaces between local devices, via wired and optical fixed telephone line networks, and through various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that carry software. As used herein, "media" can include other types of (intangible) media, unless limited to non-transitory tangible storage media.
[0265] Terms such as "storage" media, computer or machine "readable media" refer to any tangible (e.g., physical) non-transitory media involved in providing instructions to a processor for execution.
[0266] Therefore, machine-readable media, such as computer executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, any storage device in any computer, such as optical or magnetic disks, which may be used to implement databases, for example, shown in drawings. Volatile storage media include dynamic memory, for example, the main memory of such a computer platform. Tangible transmission media include copper wires and optical fibers, including coaxial cables and wires including buses in computer systems. Carrier media can take the form of electrical or electromagnetic signals, or sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example,: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card paper tapes, any other physical storage media having a pattern of holes, RAM, ROMs, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carriers for transporting data or instructions, cables or links for transporting such carriers, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in transporting one or more sequences of one or more instructions to a processor for execution.
[0267] The computer system 110 may include, for example, an electronic display 935 that includes a user interface (UI) for providing reports, or may communicate with such a display. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0268] The methods and systems of this disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by the processor 120.
[0269] The methods of the present invention may be used to diagnose the presence or absence of a condition in a subject, particularly cancer; to characterize a condition (e.g., to stage cancer or determine the heterogeneity of cancer); to select a treatment for a condition; to monitor the response to a treatment for a condition; or to generate a prognostic risk of developing a condition or a subsequent process of a condition.
[0270] Various types of cancer can be detected using the methods of the present invention. Cancer cells, like most cells, can be characterized by the rate of turnover, in which old cells die and are replaced by new cells. Generally, dead cells in contact with vascular structures in a given object can release DNA or DNA fragments into the bloodstream. This is also true for cancer cells between different stages of disease. Depending on the stage of the disease, cancer cells can also be characterized by various genetic abnormalities, such as copy number diversity and rare mutations. This phenomenon can be used to detect the presence or absence of cancer in an individual using the methods and systems described herein.
[0271] In certain embodiments, the methods and embodiments disclosed herein are used to diagnose a given disease, disorder, or condition in a patient. Typically, the disease considered is a type of cancer. Non-limited examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), and chronic myelomonocytic leukemia (CMM). L) This includes liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, progenitor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal cancer (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.
[0272] Other gene-based diseases, disorders, or conditions that may be evaluated using the methods and systems disclosed herein include, but are not limited to, DNA damage repair deficiencies, achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), feline meowing, Crohn's disease, cystic fibrosis, Darkham's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombotic diathesis, familial hypercholesterolemia, familial Mediterranean fever, and fragile X syndrome. This includes syndromes such as Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland syndrome, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs syndrome, thalassemia, trimethylaminuria, Turner syndrome, palatocardiafacial syndrome, WAGR syndrome, and Wilson's disease.
[0273] Cancer can be detected from genetic diversity, which includes mutations, rare mutations, indels, copy number diversity, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, chromosome fusions, gene shortening, gene amplification, gene duplication, chromosomal lesions, DNA lesions, abnormal changes in the chemical modification of nucleic acids, and abnormal changes in epigenetic patterns.
[0274] Genetic data can also be used to characterize specific types of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profiling data can enable the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that particular subtype. This information can also provide subjects or practitioners with clues about the prognosis of specific types of cancer, allowing either subjects or practitioners to tailor treatment options as the disease progresses. Some cancers become more aggressive and genetically unstable as they progress. Other cancers may remain benign, inactive, or quiescent. The systems and methods of this disclosure may be useful in determining disease progression.
[0275] The analysis of the present invention is also useful in determining the effectiveness of a particular treatment option. A successful treatment option may increase the amount of copy number diversity or rare mutations detected in the blood of the subject, as more cancer cells may die and shed DNA when the treatment is successful. In other cases, this may not occur. In another case, a particular treatment option may correlate over time with the genetic profile of the cancer. This correlation may be useful in selecting a therapy. Furthermore, if the cancer is observed to be in remission after treatment, the method of the present invention may be used to monitor residual disease or disease recurrence.
[0276] The method of the present invention may also be used to detect genetic diversity in conditions other than cancer. Immune cells, such as B cells, may undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number diversity detection, and certain immune states can be monitored. In this example, copy number diversity analysis may be performed over time to produce a profile of how a particular disease may be progressing.
[0277] Furthermore, the methods of this disclosure may be used to characterize heterogeneity of abnormal conditions in a subject, the methods comprising the step of generating a gene profile of extracellular polynucleotides in the subject, the gene profile comprising multiple data obtained from the analysis of copy number diversity and rare mutations. Diseases, including but not limited to cancer, may be heterogeneous. Disease cells may not be identical. In the case of cancer, it is known that some tumors contain different types of tumor cells, and it is known that some cells are in different stages of cancer. In other cases, heterogeneity may include multiple lesions of the disease. Again, in the case of cancer, multiple tumor lesions may be present, and perhaps one or more lesions are the result of metastasis spreading from the primary site.
[0278] The methods of the present invention can be used to generate or profile fingerprints or sets of data, which are summaries of genetic information derived from different cells in heterogeneous diseases. This set of data may include, alone or in combination with, analyses of copy number diversity and rare mutations.
[0279] Example of precision treatment
[0280] The detailed diagnostic methods provided by the improved computer system 110 may result in detailed treatment plans, which can be identified by the computer system 110 (and / or curated by medical professionals). For example, one type of detailed diagnosis and treatment may relate to a gene in the homologous recombination repair (HRR) pathway.
[0281] Homologous recombination is a type of genetic recombination in which nucleotide sequences are exchanged between two similar or identical molecules of DNA. It is most widely used by cells to precisely repair harmful breaks occurring on both strands of DNA, known as double-strand breaks (DSBs). HRR provides a mechanism for error-free removal of damage present in replicated DNA (S phase and G2 phase) to eliminate chromosomal breaks before cell division occurs. The initial model for how homologous recombination repairs double-strand breaks in DNA is the homologous recombination repair pathway, which mediates the double-strand break repair (DSBR) pathway and the synthesis-dependent strand annealing (SDSA) pathway. Germline and somatic cell defects in homologous recombination genes have been strongly associated with breast, ovarian, and prostate cancers.
[0282] The number and type of variant nucleotides in a sample can provide an indicator of the subject providing the sample's compliance with treatment, i.e., therapeutic intervention. For example, various poly-ADP-ribose polymerase (PARP) inhibitors have been shown to halt tumor growth from breast, ovarian, and prostate cancers caused by hereditary mutations in the BRCA1 or BRCA2 genes. Some of these therapeutic agents may inhibit base excision repair (BER), which may counteract HRR deficiencies.
[0283] On the other hand, certain BRCA and HRR wild-type patients may not clinically benefit from treatment with PARP inhibitors. Furthermore, not all ovarian cancer patients with BRCA mutations respond to PARP inhibitors. Moreover, different types of mutations may lead to different therapies. For example, somatic heterozygous deletions in the HRR gene may lead to different therapies than somatic homozygous deletions. Therefore, the state of the genetic material can influence therapies. For example, PARP inhibitors may be administered to individuals with somatic homozygous deletions in the HRR gene, but not to individuals with wild-type alleles or somatic heterozygous deletions in the HRR gene.
[0284] Nucleotide diversity in a sequenced nucleic acid can be determined by comparing the sequenced nucleic acid to a reference sequence. The reference sequence is often a known sequence, such as a known whole or partial genomic sequence from a subject, often the whole genomic sequence of a human subject. The reference sequence can be hG19. The sequenced nucleic acid can represent the sequence directly determined for the nucleic acid in a sample, or the consensus of the sequences of amplification products of such nucleic acids, as described above. The comparison can be made at one or more designated positions on the reference sequence. When the respective sequences are maximally aligned, a subset of the sequenced nucleic acid containing the positions corresponding to the designated positions on the reference sequence can be identified. Within such a subset, if present, it can be determined that the sequenced nucleic acid contains nucleotide diversity at the designated position and, optionally, if present, contains the reference nucleotide (i.e., the same nucleotide as in the reference sequence). If the number of sequenced nucleic acids in the subset containing the nucleotide variant exceeds a threshold, the variant nucleotide can be called at the designated position. The threshold can be, among other possibilities, a simple number within the subset containing the nucleotide variant, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequenced nucleic acids, or a ratio of the sequenced nucleic acids within the subset containing the nucleotide variant, such as at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20. The comparison can be repeated for any designated position of interest in the reference sequence. Sometimes, the comparison can be made for designated positions occupying at least 20, 100, 200, or 300 consecutive positions on the reference sequence, such as 20 - 500, or 50 - 300 consecutive positions.
Example
[0285] (Example 1) Status of homologous recombination repair (HRR) mutations in prostate cancer profiled by ctDNA next-generation sequencing
[0286] Background
[0287] PARP inhibition can cause synthetic lethality and increased therapeutic sensitivity in patients with HRR deficiency (HRD), which can be detected via molecular profiling of the HRR gene. For example, the FDA recently approved the use of the PARP inhibitors olaparib (de Bono et al., "Olaparib for Metastatic Castration-Resistant Prostate Cancer," N Engl J Med., 382(22):2091-2102 (2020)) and rucaparib (Abida et al., "Non-BRCA DNA Damage Repair Gene Alterations and Response to the PARP Inhibitor Rucaparib in Metastatic Castration-Resistant Prostate Cancer: Analysis From the Phase II TRITON2 Study," Clin Cancer Res., 10.1158 / 1078-0432.CCR-20-0394 (2020)) in patients with metastatic castration-resistant prostate cancer (mCRPC) who have mutations in the HRR gene.Prostate cancer is the most common malignant disease in men (Siegel et al., "Cancer statistics, 2020," CA Cancer J Clin, 70(1):7-30 (2020)), and men with advanced prostate cancer have a high prevalence of HRD (20-30%), Athie et al., "Targeting DNA Repair Defects for Precision Medicine in Prostate Cancer" Curr Oncol Rep, 21(5):42 (2019); Mateo et al., "Olaparib in patients with metastatic castration-resistant prostate cancer with DNA repair gene aberrations (TOPARP-B): a multicentre, open-label, randomised, phase 2 trial," Lancet Oncol., 21(1):162-174 (2020); Robinson et al., "Integrative clinical genomics of advanced prostate cancer" [published correction appears in Cell. 2015 Jul [16;162(2):454]. Cell, 161(5):1215-1228 (2015))High failure rates of tissue biopsy in patients with metastatic prostate cancer (25-75%, or even higher (e.g., ≥90%)) (Ross et al., "Predictors of prostate cancer tissue acquisition by an undirected core bone marrow biopsy in metastatic castration-resistant prostate cancer--a Cancer and Leukemia Group B study," Clin Cancer Res, 11(22):8109-13 (2005); Spritzer et al., "Bone marrow biopsy: RNA isolation with expression profiling in men with metastatic castration-resistant prostate cancer--factors affecting diagnostic success," Radiology, 269(3):816-23 (2013); Sailer et al., "Bone biopsy protocol for advanced prostate cancer in the era of precision medicine," Cancer, 124(5):1008-1015) (2018) highlights the challenges in HRD profiling and emphasizes the need for non-invasive ctDNA alternatives. With ctDNA, detecting copy number loss, a frequent cause of HRD, is even more difficult due to signal dilution by cell-free leukocyte DNA (Barbacioru et al., "Abstract 435: Cell-free circulating tumor DNA (ctDNA) detects somatic copy number loss in homologous recombination repair genes," Proceedings: AACR Annual Meeting 2019; March 29-April 3, 2019; Atlanta, GA).Therefore, we developed a pipeline to identify HRDs on GuardantOMNI, a 500-gene fluid biopsy panel, by detecting loss-of-function SNVs / indels, structural rearrangements, and gene deletions. This example demonstrates its performance across more than 650 prostate cancer GuardantOMNI samples.
[0288] method
[0289] Samples from 687 prostate cancer patients were processed on GuardantOMNI RUO (Table 3 shows some of the product characteristics), and a median of approximately 4600 unique molecules were sequenced to a read depth of 20,000×. Somatic and germline SNVs and small indels were called using the Guardant bioinformatics pipeline (Helman et al., "Cell-Free DNA Next-Generation Sequencing Prediction of Response and Resistance to Third-Generation EGFR Inhibitor," Clin Lung Cancer, 19(6):518-530 (2018)). We developed a novel HRD module to annotate pathogenic SNVs / indels and identified structural rearrangements, gene-level homozygous deletions, heterozygous loss (LOH), and genome-wide LOH, consisting of novel CNVs (Barbacioru et al., "Abstract 435: Cell-free circulating tumor DNA (ctDNA) detects somatic copy number loss in homologous recombination repair genes," Proceedings: AACR Annual Meeting 2019; March 29-April 3, 2019; Atlanta, GA) and de novo fusion callers (Yablonovitch et al., "Identification of FGFR2 / 3 fusions from clinical cfDNA NGS using a de novo fusion caller," 2020 May. ASCO Poster).LOH deletion was determined based on the allele frequency predicted considering the loss of wild-type alleles (Barbacioru et al., "Abstract 435: Cell-free circulating tumor DNA (ctDNA) detects somatic copy number loss in homologous recombination repair genes," Proceedings: AACR Annual Meeting 2019; March 29-April 3, 2019; Atlanta, GA). Loss-of-function variants were analyzed in 24 HRR genes: ATM, ATR, BAP1, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, HDAC2, MRE11, NBN, PALB2, RAD51, RAD50, RAD51B, RAD51C, RAD51D, RAD54L, XRCC2, and XRCC3. [Table 3]
[0290] result
[0291] Pathogenic alterations in the HRR gene were called in 300 / 687 (43.6%) of prostate cancer samples in which ctDNA was detected: 23% of all samples had pathogenic somatic or germline SNVs / indels, 7.8% had homozygous deletions, and 3.0% had HRR gene-associated rearrangements. The majority of SNVs / indels were present in BRCA2 (32% of all 159 harmful SNVs / indels) and ATM (35%) as well as in tissue (Dhawan et al., "DNA Repair Deficiency Is Common in Advanced Prostate Cancer: New Therapeutic Opportunities," Oncologist, 21(8):940-5 (2016)), but mutations were also present across 21 additional genes, including CDK12 (13%), CHEK2 (8%), and NBN (6%). Among prostate patients with germline BRCA1 / 2 SNVs / indels and sufficient tumor shedding (maximum MAF > 20%) to detect LOH, 6 / 12 (50%) also had LOH, compared to 86% in tissue (Jonsson et al., "Tumor lineage shapes BRCA-mediated phenotypes," Nature, 571(7766):576-579 (2019)). Homozygous deletions were enriched in BRCA2 (12% of all samples), ATM (6%), and FANCA (5%). Rearrangements, including fusions and multi-exon deletions, accounted for 6.5% of the detected inactivating HRD mutations. In total, 6.8% of prostate samples had biallelic inactivation involving SNVs, indels, or deletions.
[0292] Table 4 further summarizes the GuardantOMNI RUO performance metrics. In particular, the range listed for 95% LoD ( * The ranges listed for 95%LoD include clinically actionable variants and clinically actionable variants, respectively. **The terms (indicated by ) are for homozygous and heterozygous deletions, respectively. All metrics were based on a 30 ng input using clinical samples of cfDNA, with the exception of HRR deletions based on in-silico simulations, which will be discussed further below. Specificity is based on the detection of false-negative variants across a large cohort of normal samples. [Table 4]
[0293] Figure 15 (Panels A-C) shows plots of data indicating the limit of detection (LoD) of GuardantOMNI RUO for HRR deletions and fusions. More specifically, in-silico simulations demonstrated that a 95% sensitivity for detecting BRCA2 deletions in a sample resulted in a 12.5% tumor percentage (TF) for homozygous deletions (Panel A) and a 25% tumor percentage (TF) for LOH (Panel B). The LoD for deletions is shown where zygosity is deterministic (DET) and indeterminate (INDET). Experiments using clinical cfDNA containing known fusions and long deletions were evaluated using a probit model to determine the 95% LoD for MAF 0.15% (Panel C).
[0294] Table 5 shows a comparison of HRR gene mutation prevalence in tissues using GuardantOMNI RUO plasma. As shown, a larger proportion of samples in GuardantOMNI RUO compared to FFPE tissue yielded reportable results (PROFOUND - FoundationOne (de Bono et al., "Olaparib for Metastatic Castration-Resistant Prostate Cancer," N Engl J Med., 382(22):2091-2102 (2020)), TOPARP - Institute of Cancer Research (Mateo et al., "Olaparib in patients with metastatic castration-resistant prostate cancer with DNA repair gene aberrations (TOPARP-B): a multicenter, open-label, randomized, phase 2 trial," Lancet Oncol., 21(1):162-174 (2020))). * Note that the input material for the GuardantOMNI RUO assay was plasma with a variable input volume (average = 2.78 mL). A higher success rate was expected for the Laboratory Diagnostic Test (LDT), considering the requirement of 10 mL of whole blood. FACF and RANCM ( ** Except for those highlighted, the highlighted genes indicate genes that are not currently present on the GuardantOMNI RUO HRR gene list (without deletion and fusion output) but are covered on the OMNI panel. A comparative study of MSKCC is described in Jonsson et al., "Tumour lineage shapes BRCA-mediated phenotypes," Nature, 571(7766):576-579 (2019). [Table 5-1] [Table 5-2]
[0295] Figure 16 shows the oncoprint of HRR mutations in a prostate cancer cohort. Only homozygous copy number deletions are shown.
[0296] Figure 17 (Panels A-C) plots the prevalence of HRR mutations by variant class detected in the prostate cohort. Panel A shows HRR mutations by variant type, where "deletion" indicates a deletion with insufficient allele information to determine zygosity. The left side of Panel B shows harmful SNVs / indels by gene and somatic status. Consistent with the low prevalence in the clinical prostate cohort (3 / 1000 samples), no reversions were detected in this cohort (data not shown). The right side shows fusions and long deletions by gene. HRR fusions and long deletions were found in 3.2% (21 / 654) of samples, none of which contained harmful HRR SNVs / indels. Panel C shows an example of a BRCA2 deletion in exons 24-26. The black line in the center indicates the discontinuous axis, and the bottom shows the distance of the detected multi-exons.
[0297] conclusion
[0298] This example demonstrates that in a prostate cancer cohort, GuardantOMNI ctDNA profiling calls all classes of mutations contributing to HRD, and that the relative prevalence of alterations is consistent with that in tissue. cfDNA offers an alternative method for identifying patients who may benefit from PARP or cisplatin / platinum therapy, raising prevalence from 28% with small variants to 42% with the complete HRD biomarker set.
[0299] All patent applications, websites, other publications, accession numbers, etc., cited above or below are incorporated by reference in their entirety for all purposes to the same extent that each individual item is specifically and individually indicated as being incorporated by reference in that manner. Where different versions of a sequence are associated with an accession number at different points in time, the version associated with that accession number as of the effective filing date of this application is meant. The effective filing date means, where applicable, the earlier of the actual filing date or the filing date of the priority application referring to that accession number. Similarly, where different versions of a publication, website, etc., are published at different points in time, the most recently published version as of the effective filing date of this application is meant, unless otherwise indicated. Any feature, step, element, embodiment, or aspect of this disclosure may be used in combination with any other feature, step, element, embodiment, or aspect unless specifically indicated otherwise. While this disclosure has been described in some detail for illustrative and illustrative purposes for clarity and understanding, it will be clear that certain changes and modifications may be implemented within the scope of the appended claims. In certain embodiments, for example, the following items are provided: (Item 1) A step of determining sequence data of a biological sample, wherein the biological sample contains cell-free DNA (cfDNA); A step of determining coverage data based on the aforementioned sequence data; A step of determining one or more breakpoints associated with one or more fusion events based on the coverage data; A step of determining one or more deletions associated with one or more genes based on the coverage data; A step of determining a homologous recombination deletion (HRD) score based on the one or more breakpoints and the one or more deletions; and A step of classifying the biological sample as HRD positive based on the HRD score. A method that includes this. (Item 2) The method according to item 1, wherein the step of determining the sequence data of the biological sample includes sequencing a panel of one or more HRR genes. (Item 3) The method according to item 2, wherein one or more HRR genes are selected from the group consisting of ATM, ATR, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, NBN, PALB2, RAD51, RAD51B, RAD51C, RAD51D, RAD54L, HDAC2, MRE11, PPP2R2A, XRCC5, WRN, MLH1, FANCC, BAP1, XRCC2, XRCC3, and RAD50. (Item 4) The method according to any one of the above items, wherein the biological sample is associated with a subject having a disease. (Item 5) The method described in item 4, wherein the disease is cancer. (Item 6) The method according to any one of the items, wherein the coverage data is associated with a plurality of bins, and the plurality of bins represent regions of a chromosome. (Item 7) The step of determining one or more breakpoints associated with one or more fusion events based on the coverage data is: Aligning multiple array reads from the aforementioned array data to a reference array; Determining one or more break points in the alignment of multiple sequence reads to the reference sequence for multiple sequence reads; Identifying any sequence read associated with one or more breakpoints in the alignment as a candidate fusion sequence read; Determining candidate fusion sequence reads associated with a common cleavage point among one or more cleavage points; Grouping the candidate fusion sequence reads based on one or more common breakpoints; Assembling the candidate fusion sequence reads within the group to generate one or more contigs; Aligning the contig from one of several groups to the reference sequence; Determine one or more candidate fusion events based on the alignment of the contigs from the group; Applying one or more criteria to the one or more candidate fusion events; and Determining one or more fusion events based on the application of one or more criteria to the one or more candidate fusion events. The method described in any one of the preceding items, including (Item 8) The method according to any one of the above items, wherein the step of determining one or more deletions associated with one or more genes based on the coverage data includes determining a plurality of segments based on the coverage data, wherein the plurality of segments are separated by change points. (Item 9) The method according to item 8, wherein determining the plurality of segments includes applying a segmentation algorithm. (Item 10) The method according to item 9, wherein the segmentation algorithm includes a cyclic binary segmentation algorithm. (Item 11) The method according to item 8, wherein the change point corresponds to a position in the coverage data indicating a change in the basal DNA copy number. (Item 12) The method according to any one of the above items, wherein the one or more deletions include one or more homozygous deletions or heterozygous loss (LOH) deletions. (Item 13) A step of comparing the plurality of segments with a reference sequence to identify a subset of the plurality of segments containing at least one deletion; A step of removing any segment over the length of a chromosome from the subset of the plurality of segments; A step of combining any segments from the subset of the plurality of segments that are not separated by a threshold distance; The steps of removing any segment having a length less than a threshold length from the subset of the plurality of segments; and Steps to remove any segments related to technical artifacts from the subset of the plurality of segments. The method described in item 8, further including the method described in item 8. (Item 14) The method of item 13, further comprising the step of determining the number of breakpoints between adjacent segments within a threshold based on one or more remaining segments in the subset of the plurality of segments and based on one or more breakpoints associated with the one or more fusion events. (Item 15) The method according to item 13, further comprising the step of determining the number of segments relating to a single copy of the desired region to be deleted, or relating to copies of both of the desired regions, based on one or more remaining segments from the subset of the plurality of segments. (Item 16) The method according to item 8, wherein the step of determining the HRD score based on the one or more breakpoints and the one or more deletions includes summing the number of breakpoints and the number of segments. (Item 17) The method according to any one of the above items, further comprising the step of determining the presence of one or more genomic rearrangements based on the sequencing data. (Item 18) The method according to item 17, wherein the step of determining the HRD score is further based on one or more genome rearrangements. (Item 19) The method according to item 17, wherein the step of determining the HRD score includes summing the number of breakpoints, the number of segments, and the number of genomic rearrangements. (Item 20) The method according to any one of the preceding items, further comprising the step of determining the maximum somatic allele proportion (MSAF). (Item 21) The method according to item 20, wherein the step of determining the MSAF further comprises determining the maximum percentage of variants in the biological sample, based on the sequence data, any somatic cell variant not annotated as having clonal hematopoietic origin, and including any somatic cell variant that is fused or non-synonymous SNV or indel. (Item 22) The method according to any one of the above items, further comprising the step of annotating one or more variants contained in the sequence data. (Item 23) The method according to item 22, wherein the step of annotating one or more variants contained in the sequence data includes determining an annotation of clinical significance related to the effect of the one or more variants on human health. (Item 24) The method according to any one of the items, further comprising the step of aggregating the sequence data, the coverage data, one or more breakpoints, one or more deletions, and the HRD score. (Item 25) The method according to item 24, further comprising the step of outputting aggregated sequence data, coverage data, one or more breakpoints, one or more deletions, and an HRD score. (Item 26) The method according to any one of the items, wherein the step of classifying the biological sample as HRD-positive based on the HRD score includes determining that the HRD score exceeds a threshold. (Item 27) The method according to item 26, further comprising determining the threshold based on one or more reference HRD scores. (Item 28) The method according to any one of the above items, further comprising the step of administering a therapy based on the step of classifying the biological sample as HRD-positive. (Item 29) The method according to item 28, wherein the therapy is a poly-ADP-ribose polymerase (PARP) inhibitor or a base excision repair (BER) inhibitor. (Item 30) The method according to item 29, wherein the PARP inhibitor is at least one of veliparib, olaparib, talazoparib, lucaparib, niraparib, pamiparib, CEP 9722, E7016, E7449, or 3-aminobenzamide. (Item 31) The method according to item 28, wherein the therapy is a combination of a PARP inhibitor and radiotherapy. (Item 32) A method for generating homologous recombination repair loss (HRD) scores using a computer at least partially, A step of generating a set of reference HRD scores by the computer for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from one or more reference subjects having one or more oncologies, wherein a given reference HRD score includes the prevalence of a given HRD nucleic acid variant; and Steps to generate a reference HRD score from the aforementioned set of reference HRD scores. A method that includes this. (Item 33) A step of generating a set of test HRD scores for one or more genes in the set of HRR genes from sequence information derived from cfDNA obtained from a test subject having one or more cancer types, wherein a given test HRD score includes the prevalence of the given HRD nucleic acid variant; A step of generating a test HRD score from the aforementioned set of test HRD scores; and The step of detecting the HRD in the test subject when the test HRD score exceeds the reference HRD score. The method described in item 32, further including the method described in item 32. (Item 34) A method for determining the homologous recombination repair deficiency (HRD) status of a test subject having one or more cancer types, using a computer at least partially; A step of generating a set of test HRD scores for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from the test subject, wherein a given test HRD score includes the prevalence of a given HRD nucleic acid variant; A step of generating a test HRD score from the aforementioned set of test HRD scores; and A step of comparing the aforementioned test HRD scores with reference HRD scores, wherein test HRD scores that are higher than the reference HRD scores indicate that those test HRD scores are from test subjects having HRD, and test HRD scores that are at or below the reference HRD scores indicate that those test HRD scores are from test subjects lacking HRD, thereby determining the HRD status of the test subject having one or more cancer types. A method that includes this. (Item 35) A step of generating a set of reference HRD scores by the computer for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from one or more reference subjects having one or more oncologies, wherein a given reference HRD score includes the prevalence of a given HRD nucleic acid variant; and Steps to generate a reference HRD score from the aforementioned set of reference HRD scores. The method described in item 34, further including the method described in item 34. (Item 36) A method for detecting the presence or absence of homologous recombination repair deficiency (HRD) in a subject, using a computer at least partially, comprising the step of detecting the HRD in the subject by determining, by the computer, the presence or absence of at least one HRD nucleic acid variant in sequence information relating to one or more genes in a set of homologous recombination repair (HRR) genes derived from cell-free nucleic acid (cfDNA) obtained from the subject, with (i) a first probability that the sequence information includes a first state and a second probability that the sequence information includes a second state, wherein the first or second state includes at least a first HRD nucleic acid variant, and / or (ii) one or more aligned contigs generated from the sequence information, wherein the aligned contigs include at least a second HRD nucleic acid variant. (Item 37) A method for treating a disease, comprising the step of administering one or more therapies to a subject having the disease and a homologous recombination repair deficiency (HRD) associated with the disease, wherein the presence or absence of at least one HRD nucleic acid variant in sequence information relating to one or more genes in a set of homologous recombination repair (HRR) genes derived from cell-free nucleic acid (cfDNA) obtained from the subject is determined by (i) a first probability that the sequence information includes a first state and a second probability that the sequence information includes a second state, wherein the first or second state includes at least a first HRD nucleic acid variant, and / or (ii) one or more aligned contigs generated from the sequence information, wherein the aligned contigs include at least a second HRD nucleic acid variant, thereby detecting HRD and treating the disease. (Item 38) The method according to any one of the above items, wherein the first HRD nucleic acid variant comprises at least one homozygous deletion, at least one heterozygous loss (LOH) variant and / or at least one copy number variation (CNV). (Item 39) The method according to item 38, wherein the LOH variant comprises at least one gene-specific LOH variant, at least one LOH variant without copy number change, and / or at least one LOH variant. (Item 40) The method according to any one of the above items, wherein the second HRD nucleic acid variant comprises at least one structural rearrangement. (Item 41) The method according to item 40, wherein the structural reorganization includes at least one shortening reorganization and / or at least one multiexon deletion. (Item 42) The method according to any one of the above items, wherein the first and / or second HRD nucleic acid variant comprises at least one single nucleotide polymorphism (SNV) and / or at least one indel. (Item 43) The method according to any one of the items, wherein the technique of step (i) and / or (ii) includes at least aligning the segment of the sequence information to at least one reference sequence. (Item 44) The method described in any one of the preceding items, which includes using only one of steps (i) to (ii). (Item 45) The method described in any one of the preceding items, comprising using each of steps (i) to (ii). (Item 46) The method according to any one of the above items, wherein at least one homologous recombination repair (HRR) gene comprises the HRD nucleic acid variant. (Item 47) The method according to item 46, wherein the HRR gene is selected from the group consisting of ATM, ATR, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, NBN, PALB2, RAD51, RAD51B, RAD51C, RAD51D, RAD54L, HDAC2, MRE11, PPP2R2A, XRCC5, WRN, MLH1, FANCC, BAP1, XRCC2, XRCC3, and RAD50. (Item 48) The method according to any one of the preceding items, wherein the set of HRR genes comprises at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more genes. (Item 49) The method according to any one of the above items, wherein one or more of the HRD nucleic acid variants cause biallelic inactivation of a given HRR gene. (Item 50) The method according to any one of the above items, wherein one or more of the HRD nucleic acid variants cause uniallelic inactivation of a given HRR gene. (Item 51) The method according to any one of the items, wherein the HRD nucleic acid variant correlates with the subject having the disease. (Item 52) The method according to any one of the above items, wherein the subject is not known to have a disease. (Item 53) The method according to any one of the above items, wherein the subject is known to have a disease. (Item 54) The method according to any one of the above items, wherein the disease is cancer. (Item 55) The method according to any one of the items, comprising the step of administering one or more therapies to the subject in order to treat the disease. (Item 56) The method according to any one of the above items, wherein the therapy comprises at least one poly-ADP-ribose polymerase (PARP) inhibitor. (Item 57) The method according to any one of the above items, wherein the therapy comprises at least one base excision repair (BER) inhibitor. (Item 58) Step (i) is, The sequence information generates the first probability that includes the first state; The sequence information generates the second probability that includes the second state; Comparing the first probability and the second probability; and Based on the above comparison, a prediction is generated as to whether the sequence information includes the first state or the second state. The method described in any one of the preceding items, including (Item 59) Step (i) is, A first model of allele counts based on one or more germline single nucleotide polymorphism (SNP) locations associated with at least one gene locus in the sequence information, wherein the first model representing at least one somatic homozygous deletion is generated by a first probability distribution; A second model of allele counts based on one or more germline SNP locations associated with the gene locus in the sequence information, wherein the second model representing at least one somatic heterozygous deletion is generated by a second probability distribution; Comparing the first output data of the first model with the second output data of the second model; and Based on the above comparison, a prediction is generated that the somatic homozygous deletion for the gene locus is present in the sequence information. The method described in any one of the preceding items, including (Item 60) Step (i) is, The sequence information generates the first probability, which includes somatic homozygous deletions; The sequence information generates the second probability, which includes somatic heterozygous deletions; Comparing the first probability and the second probability; and Based on the above comparison, a prediction is generated as to whether the sequence information includes the somatic homozygous deletion or the somatic heterozygous deletion. The method described in any one of the preceding items, including (Item 61) Step (ii) is, Aligning multiple array reads to a reference array; Determining one or more break points in the alignment of at least one of the plurality of sequence reads to the reference sequence; Identifying any sequence read associated with one or more breakpoints in the alignment as a candidate fusion sequence read; Determining candidate fusion sequence reads associated with a common cleavage point among one or more cleavage points; Grouping the candidate fusion sequence reads based on one or more common breakpoints; Assembling the candidate fusion sequence reads within each group to generate one or more contigs; Align the contigs from each group to the reference sequence; Determine one or more candidate fusion events based on the alignment of the contigs from each group; Applying one or more criteria to the one or more candidate fusion events; and Determining one or more fusion events containing the second HRD nucleic acid variant based on the application of the one or more criteria to the one or more candidate fusion events. The method described in any one of the preceding items, including (Item 62) The method according to any one of the items, comprising the step of detecting the HRD, or the HRD in the subject, using a CNV and / or a de novo fusion caller. (Item 63) The method according to any one of the above items, wherein the gene contains the HRD nucleic acid variant. (Item 64) One or more non-temporary computer-readable media that themselves store processor-executable instructions causing the processor to perform any of the methods described in items 1 to 31 when executed by the processor. (Item 65) A computer device configured to perform any of the methods described in items 1 through 31, An output device configured to output the aforementioned HRD score and A system that includes this. (Item 66) One or more processors, A memory that stores processor-executable instructions that cause the device to perform any of the methods described in items 1 to 31 when executed by one or more of the aforementioned processors, A device that includes this. (Item 67) One or more non-temporary computer-readable media that themselves store processor-executable instructions causing the processor to perform the actions described in any of items 32 to 33 or 38 to 63 when executed by the processor. (Item 68) A computer device configured to perform any of the methods described in items 32 to 33 or 38 to 63, An output device configured to output the aforementioned reference HRD score and A system that includes this. (Item 69) One or more processors, A memory that stores processor-executable instructions that cause the device to perform the method described in any of items 32 to 33 or 38 to 63 when executed by one or more processors, A device that includes this. (Item 70) One or more non-temporary computer-readable media that themselves store processor-executable instructions causing the processor to perform the actions described in any of items 34 to 35 or 38 to 63 when executed by the processor. (Item 71) A computer device configured to perform any of the methods described in items 34 to 35 or 38 to 63, An output device configured to output the HRD status of the subject under test, A system that includes this. (Item 72) One or more processors, A memory that stores processor-executable instructions that cause the device to perform the method described in any of items 34 to 35 or 38 to 63 when executed by one or more processors, A device that includes this. (Item 73) One or more non-temporary computer-readable media that themselves store processor-executable instructions causing the processor to perform the actions described in any of items 36 or 38 to 63 when executed by the processor. (Item 74) A computer device configured to perform any of the methods described in item 36 or 38 to 63, An output device configured to output an indication of the presence or absence of HRD in the aforementioned target, A system that includes this. (Item 75) One or more processors, A memory that stores processor-executable instructions that cause the device to perform the method described in any of items 36 or 38 to 63 when executed by one or more processors, A device that includes this.
Claims
1. A step of determining sequence data of a biological sample, wherein the biological sample contains cell-free DNA (cfDNA); A step of determining coverage data based on the aforementioned sequence data; A step of determining one or more breakpoints associated with one or more fusion events based on the coverage data; A step of determining one or more deletions associated with one or more genes based on the coverage data; A step of generating a homologous recombination deletion (HRD) score based on the one or more breakpoints and the one or more deletions; and A step of classifying the biological sample as HRD positive based on the HRD score. A method for determining HRD status, including [specific details omitted].
2. The method according to claim 1, wherein the step of determining the sequence data of the biological sample includes sequencing a panel of one or more HRR genes.
3. The method according to claim 2, wherein one or more HRR genes are selected from the group consisting of ATM, ATR, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCA, FANCL, NBN, PALB2, RAD51, RAD51B, RAD51C, RAD51D, RAD54L, HDAC2, MRE11, PPP2R2A, XRCC5, WRN, MLH1, FANCC, BAP1, XRCC2, XRCC3, and RAD50.
4. The method according to any one of claims 1 to 3, wherein the biological sample is associated with a subject having a disease.
5. The method according to claim 4, wherein the disease is cancer.
6. The method according to any one of claims 1 to 5, wherein the coverage data is associated with a region of a chromosome.
7. The step of determining one or more breakpoints associated with one or more fusion events based on the coverage data is: Aligning multiple array reads from the aforementioned array data to a reference array; Determining one or more break points in the alignment of multiple sequence reads to the reference sequence for multiple sequence reads; Identifying any sequence read associated with one or more breakpoints in the alignment as a candidate fusion sequence read; Determining candidate fusion sequence reads associated with a common cleavage point among one or more cleavage points; Grouping the candidate fusion sequence reads based on one or more common breakpoints; Assembling the candidate fusion sequence reads within the group to generate one or more contigs; Aligning the contig from one of several groups to the reference sequence; Determining one or more candidate fusion events based on the alignment of the contigs from the group; Applying one or more criteria to the one or more candidate fusion events; and Determining one or more fusion events based on the application of one or more criteria to one or more candidate fusion events. The method according to any one of claims 1 to 6, including the method described in any one of claims 1 to 6.
8. The method according to any one of claims 1 to 7, wherein the step of determining one or more deletions associated with one or more genes based on the coverage data includes determining a plurality of segments based on the coverage data, wherein the plurality of segments are separated by change points.
9. The method according to claim 8, wherein determining the plurality of segments includes applying a segmentation algorithm.
10. The method according to claim 9, wherein the segmentation algorithm includes a cyclic binary segmentation algorithm.
11. The method according to claim 8, wherein the change point corresponds to a position in the coverage data indicating a change in the basal DNA copy number.
12. The method according to any one of claims 1 to 11, wherein the one or more deletions include one or more homozygous deletions or heterozygous loss (LOH) deletions.
13. A step of comparing the plurality of segments with a reference sequence to identify a subset of the plurality of segments containing at least one deletion; A step of removing any segment from the subset of the plurality of segments that spans a portion of the length of the chromosome; A step of combining any segments from the subset of the plurality of segments that are not separated by a threshold distance; The steps of removing any segment having a length less than a threshold length from the subset of the plurality of segments; and Steps to remove any segments related to technical artifacts from the subset of the plurality of segments. The method according to claim 8, further comprising:
14. The method of claim 13, further comprising the step of determining the number of breakpoints between adjacent segments within a threshold based on one or more remaining segments in the subset of the plurality of segments and based on one or more breakpoints associated with the one or more fusion events.
15. The method according to claim 13, further comprising the step of determining the number of segments relating to a single copy of the target area to be deleted, or relating to copies of both of the target areas, based on one or more remaining segments from the subset of the plurality of segments.
16. The method according to claim 8, wherein the step of generating the HRD score based on the one or more breakpoints and the one or more deletions includes summing the number of breakpoints and the number of segments.
17. The method according to any one of claims 1 to 16, further comprising the step of determining the presence of one or more genomic rearrangements based on the sequencing data.
18. The method according to claim 17, wherein the step of generating the HRD score is further based on one or more genome rearrangements.
19. The method according to claim 17, wherein the step of generating the HRD score includes summing the number of cuts, the number of segments, and the number of genomic rearrangements.
20. The method according to any one of claims 1 to 19, further comprising the step of determining the maximum somatic allele proportion (MSAF).
21. The method according to claim 20, wherein the step of determining the MSAF further comprises determining the maximum percentage of variants in the biological sample, based on the sequence data, any somatic cell variant not annotated as having clonal hematopoietic origin, and including any somatic cell variant that is fused or non-synonymous SNV or indel.
22. The method according to any one of claims 1 to 21, further comprising the step of annotating one or more variants contained in the sequence data.
23. The method according to claim 22, wherein the step of annotating one or more variants contained in the sequence data includes determining an annotation of clinical significance related to the effect of the one or more variants on human health.
24. The method according to any one of claims 1 to 23, further comprising the step of aggregating the sequence data, the coverage data, one or more breakpoints, one or more deletions, and the HRD score.
25. The method according to claim 24, further comprising the step of outputting aggregated sequence data, coverage data, one or more breakpoints, one or more deletions, and an HRD score.
26. The method according to any one of claims 1 to 25, wherein the step of classifying the biological sample as HRD positive based on the HRD score includes determining that the HRD score exceeds a threshold.
27. The further includes determining the threshold based on one or more reference HRD scores, The method according to claim 26.
28. The method according to any one of claims 1 to 27, further comprising the step of administering a therapy based on the step of classifying the biological sample as HRD-positive.
29. The method according to claim 28, wherein the therapy is a poly-ADP-ribose polymerase (PARP) inhibitor or a base excision repair (BER) inhibitor.
30. The method according to claim 29, wherein the PARP inhibitor is at least one of veliparib, olaparib, talazoparib, lucaparib, niraparib, pamiparib, CEP 9722, E7016, E7449, or 3-aminobenzamide.
31. The method according to claim 28, wherein the therapy is a combination of a PARP inhibitor and radiotherapy.
32. A method for generating a reference homologous recombination repair loss (HRD) score using a computer at least partially, A step of generating a set of HRD scores by the computer for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from one or more reference subjects having one or more oncologies, wherein a given HRD score includes the prevalence of a given HRD nucleic acid variant; and Steps to generate a reference HRD score from the aforementioned set of HRD scores. A method that includes this.
33. A step of generating a set of test HRD scores for one or more genes in the set of HRR genes from sequence information derived from cfDNA obtained from a test subject having one or more cancer types, wherein a given test HRD score includes the prevalence of the given HRD nucleic acid variant; A step of generating a test HRD score from the set of test HRD scores; and The step of detecting the HRD in the test subject when the test HRD score exceeds the reference HRD score. The method according to claim 32, further comprising:
34. A method for determining homologous recombination repair deficit (HRD) status of a test subject having one or more cancer types, using a computer at least partially; A step of generating a set of test HRD scores for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from the test subject, wherein a given test HRD score includes the prevalence of a given HRD nucleic acid variant; A step of generating a test HRD score from the set of test HRD scores; and A step of comparing the test HRD scores with reference HRD scores, wherein test HRD scores that are higher than the reference HRD scores indicate that those test HRD scores are from test subjects having HRD, and test HRD scores that are at or below the reference HRD scores indicate that those test HRD scores are from test subjects lacking HRD, thereby determining the HRD status of the test subject having one or more cancer types. A method that includes this.
35. A step of generating a set of reference HRD scores by the computer for one or more genes in a set of homologous recombination repair (HRR) genes from sequence information derived from cell-free nucleic acid (cfDNA) obtained from one or more reference subjects having one or more oncologies, wherein a given reference HRD score includes the prevalence of a given HRD nucleic acid variant; and Steps to generate a reference HRD score from the set of reference HRD scores. The method according to claim 34, further comprising:
Citation Information
Patent Citations
Methods and Materials for Assessing Homologous Recombination Defects
JP2017533693A
Methods and applications of gene fusion detection in cell-free DNA analysis
JP2018535481A
Methods and systems for detecting insertions and deletions - Patents.com
JP2020521216A