An integrated machine learning framework for inferring homologous recombination defects
By employing machine learning algorithms trained on DNA sequencing data, the method effectively predicts homologous recombination deficiency in cancer patients, addressing the lack of computational resources and improving treatment stratification.
Patent Information
- Application Number
- JP2023176962
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-10
- Filing Date
- 2023-10-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-02-12
AI Technical Summary
Current methods lack effective computational resources for predicting homologous recombination deficiency (HRD) in cancer patients, which is crucial for stratifying treatment with PARP inhibitors.
The development of systems and methods using machine learning algorithms trained on DNA sequencing data from cancerous and non-cancerous tissues to predict the homologous recombination status of cancers.
This approach improves the accuracy of predicting HRD status, enabling better stratification of patients for treatment with PARP inhibitors and other therapies.
Smart Images

Figure 0007689557000006 
Figure 0007689557000007 
Figure 0007689557000008
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 804,730, filed February 12, 2019, and U.S. Provisional Patent Application No. 62 / 946,347, filed December 10, 2019, which are hereby incorporated by reference in their entireties for all purposes.
[0002] The present disclosure relates generally to the use of machine learning classifiers trained on DNA sequencing of cancerous tissue to predict homologous recombination deficiencies. [Background technology]
[0003] Precision oncology is the practice of tailoring cancer therapy to the unique genomic, epigenetic, and / or transcriptomic profile of an individual tumor. This is in contrast to the traditional approach of treating cancer patients solely based on the type of cancer they suffer from, for example, treating all breast cancer patients with one line of therapy and all lung cancer patients with a second line of therapy. Precision oncology arose from the many observations that different patients diagnosed with the same type of cancer, for example breast cancer, responded very differently to common treatment regimens. Over time, researchers have identified genomic, epigenetic, and transcriptomic markers that facilitate a level of prediction about how individual cancers will respond to specific treatment modalities.
[0004] Therapies targeting specific genomic alterations are already standard of care in some tumor types (e.g., as suggested by the NCCN (National Comprehensive Cancer Network) guidelines for melanoma, colorectal cancer, and non-small cell lung cancer). These few well-known mutations in the NCCN guidelines may be addressed with individual assays or small next-generation sequencing (NGS) panels. However, for the greatest number of patients to benefit from personalized oncology, molecular alterations that could be targeted for off-label drug indications, combination therapy, or tissue-independent immunotherapy must be evaluated. See Schwaederle et al. 2016 JAMA Oncol. 2, 1452-1459, Schwaederle et al. 2015 J Clin Oncol. 32, 3817-3825, and Wheler et al. 2016 Cancer Res. 76, 3690-3701. Large-panel NGS assays also cast a wider net for clinical trial enrollment. See Coyne et al. 2017 Curr. Probl. Cancer 41,182-193, and Markman 2017 Oncology 31,158,168.
[0005] Genomic analysis of tumors is rapidly becoming routine clinical practice to provide patient-tailored treatment and improve outcomes. See Fernandes et al. 2017 Clinics 72, 588-594. In fact, recent studies have shown that clinical care is guided by NGS assay results in 30-40% of patients undergoing such testing. See Hirshfield et al. 2016 Oncologist 21, 1315-1325, Groisberg et al. 2017 Oncotarget 8, 39254-39267, Ross et al. JAMA Oncol. 1, 40-49, and Ross et al. 2015 Arch. Pathol. Lab Med. 139, 642-649. There is growing evidence that patients who receive genetically guided treatment advice have better outcomes. See, for example, Wheler et al. (2016 Cancer Res. 76, 3690-3701), who used a matching score (e.g., a score based on the number of therapy associations and genomic aberrations per patient) to show that patients with higher matching scores had a higher frequency of stable disease, longer time to treatment failure, and greater overall survival. Such methods may be particularly useful for patients who have already failed multiple lines of therapy.
[0006] Targeted therapy has shown significant improvements in patient outcomes, especially with regard to progression-free survival. See Radovich et al. 2016 Oncotarget 7,56491-56500. Recent evidence is reported from the IMPACT trial with genetic testing of advanced stage tumors from 3,743 patients, where approximately 19% of patients received targeted therapy matched based on tumor biology, and showed that patients who received matched therapy had a 16.2% response rate compared to 5.2% for patients who received mismatched therapy. See Bankhead. "IMPACT Trial: Support for Targeted Cancer Tx Approaches." MedPageToday. June 5, 2018. The IMPACT study further found that patients who received molecularly matched therapy had a 3-year overall survival that was more than three times longer than those who received mismatched therapy (15% vs. 7%). See ibid. and ASCO Post. "2018 ASCO:IMPACT Trial Matches Treatment to Genetic Changes in the Tumor to Improve Survival Across Multiple Cancer conditions." The ASCO POST. June 6, 2018. Estimates of the proportion of patients whose care trajectory will change as a result of genetic testing vary widely, from about 10% to over 50%. See Fernandes et al. 2017 Clinics 72,588-594.
[0007] One example of a genomic trait linked to the efficacy of a particular treatment is a mutation in the BRCA1, BRCA2, or PALB2 homologous recombination genes. A class of pharmacological inhibitors of poly ADP-ribose polymerase 1 (PARP1), known as PARP inhibitors (PARPi), have therapeutic efficacy for treating several cancers that contain mutations in the BRCA1, BRCA2, or PALB2 homologous recombination genes. PARP1 is an essential enzyme in the error-prone microhomology-mediated end joining (MMEJ) DNA repair pathway. Sharma S.et al.,Cell Death Dis.6(3):e1697(2015). In the absence of PARP1 activity, DNA replication forks stall when they encounter a single-strand break. Fork stalling ultimately leads to a double-stranded chromosome break that can be repaired by homologous recombination (HR) repair, which is less error prone than the MMEJ pathway.
[0008] Unlike other DNA repair proteins that are commonly deficient in cancer cells, PARP1 has been shown to be overexpressed in certain cancer types. It is theorized that increased MMEJ DNA repair compared to homologous repair may lead to the accumulation of genomic mutations and cancer development. However, the efficacy of PARP inhibitors is not fully understood. For example, not all cancers with BRCA1, BRCA2, or PALB2 mutations are sensitive to PARP inhibitors. In addition, some cancers that do not have mutations in homologous recombination proteins are sensitive to PARP inhibitors.
[0009] Homologous recombination (HR) is a normal, highly conserved DNA repair process that allows the exchange of genetic information between identical or closely related DNA molecules. It is most widely used by cells to precisely repair harmful breaks (i.e., damage) that occur in both strands of DNA. DNA damage can occur from exogenous (outside) sources, such as UV light, radiation, or chemical damage, or from endogenous (inside) sources, such as errors in DNA replication or other cellular processes that cause DNA damage. Double-strand breaks are one type of DNA damage.
[0010] The use of poly(ADP-ribose) polymerase (PARP) inhibitors in patients with HRD impairs two pathways of DNA repair, leading to cell death (apoptosis). The efficacy of PARP inhibitors improves not only ovarian cancers that display germline or somatic BRCA mutations, but also cancers in which HRD is caused by other underlying etiologies.
[0011] Poly(ADP-ribose) polymerases (PARPs) are a family of proteins involved in many cellular processes, including DNA repair, genomic stability, and programmed cell death. Homologous recombination deficiency ("HR deficiency" or "HRD") is a deficiency that has been shown to enhance the efficacy of PARP inhibitors (PARPi) and platinum-based therapies in patients. The most common lesions in cellular DNA are single-strand breaks (SSBs), which occur in the tens of thousands per cell per day. PARPs are DNA repair enzymes that help repair single-strand breaks. When these PARPs are not functioning or are blocked (e.g., by PARP inhibitor therapy), this often leads to so-called double-strand breaks (DSBs). Homologous recombination repair (HRR) is the main way the body repairs these DSBs. When cancer cells have HRD (or in other words, a deficiency in HRR), the cell's chances of recovering from DSBs are reduced, instead of the cell continuing to grow, leading the cell to undergo apoptosis (programmed cell death). Causing cancer cells to die is one way to stop the growth of a person's cancer.
[0012] Some consider HRD to be a disease state that arises in tumors through loss of the homologous recombination DNA repair pathway, commonly caused by biallelic inactivation of BRCA1 / 2. Although the deficiency is often manifested by mutations in the BRCA genes, there are other ways in which tumors can have HR deficiencies, as is common in cancer.
[0013] Overall, HRD occurs with a frequency of approximately 6% in cancer. The incidence may be as high as 30% in ovarian cancer and intermediate (12-13%) in breast, pancreatic, and prostate cancer. HRD can be caused by biallelic inactivation of BRCA1, BRCA2, RAD51C, and PALB2. Loss of heterozygosity (LOH) and deletions, especially of BRCA2, are also thought to be major causes. Summary of the Invention
[0014] Given the above background, what is needed in the art are improved methods for predicting which cancers are homologous repair deficient (HRD), for example, identifying which cancer patients are likely to respond favorably to PARP inhibitors. The present disclosure addresses these and other needs by providing systems and methods for evaluating DNA sequencing results from cancerous tissues using machine learning algorithms trained to predict the homologous recombination status of cancers.
[0015] Loss of homologous recombination is a widely recognized determinant of cancer progression. However, few computational resources exist for estimating homologous recombination deficiency (HRD) from patient genomes. Genomics-based HRD testing can aid in cancer diagnosis and can be used, for example, to stratify patients for treatment with PARPi. Systems and methods for estimating the HRD status of human cancers are disclosed.
[0016] In one aspect, the present disclosure provides a method for determining the homologous recombination pathway status of cancer in a test subject. The method includes obtaining a first plurality of sequence reads of a first DNA sample from the test subject in electronic format, the first DNA sample comprising DNA molecules from a cancerous tissue of the subject. The method includes obtaining a second plurality of sequence reads of a second DNA sample from the test subject in electronic format, the second DNA sample comprising DNA molecules from a non-cancerous tissue of the subject. The method then includes generating a genome data construct of the subject based on the first plurality of sequence reads and the second plurality of sequence reads, the genome data construct comprising one or more features of the genome of the cancerous tissue and the non-cancerous tissue of the subject. In some embodiments, the plurality of features comprises (i) a heterozygosity state of a first plurality of DNA damage repair genes in the cancerous tissue of the subject, (ii) a measure of loss of heterozygosity across the genome of the cancerous tissue of the subject, (iii) a measure of mutant alleles detected in a second plurality of DNA damage repair genes in the genome of the cancerous tissue of the subject, and (iv) a measure of mutant alleles detected in a second plurality of DNA damage repair genes in the genome of the non-cancerous tissue of the subject.The method then includes inputting the genomic data construct into a classifier trained to distinguish between cancers with homologous recombination pathway defects and cancers without homologous recombination pathway defects, thereby determining the homologous recombination pathway status of the test subject.
[0017] In another aspect, the present disclosure provides a method for training an algorithm for determining a homologous recombination pathway status of cancer. The method includes, for each training subject in a plurality of training subjects having cancer, obtaining a corresponding genomic data construct for each training subject. The corresponding genomic training construct includes (a) the homologous recombination pathway status of the cancer of each training subject, and (b) one or more features of the genome of the cancerous tissue and non-cancerous tissue of each training subject. In some embodiments, the one or more features include (i) the heterozygosity status of a first plurality of DNA damage repair genes in the cancerous tissue of each training subject, (ii) a measure of loss of heterozygosity throughout the genome of the cancerous tissue of each training subject, (iii) a measure of mutant alleles detected in a second plurality of DNA damage repair genes of the genome of the cancerous tissue of each training subject, and (iv) a measure of mutant alleles detected in a second plurality of DNA damage repair genes of the genome of the non-cancerous tissue of each training subject. Next, the method includes training, for each training subject, a classification algorithm on at least (a) the homologous recombination pathway status of the cancer of each training subject, and (b) a plurality of features determined from a corresponding DNA sample from the cancerous tissue of each training subject.
[0018] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only exemplary embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other different embodiments, and its several details can be modified in various obvious respects, all without departing from the present disclosure. Thus, the drawings and description should be regarded as illustrative in nature, and not as restrictive. [Brief description of the drawings]
[0019] [Figure 1A] 1A-1C collectively show block diagrams of example computing devices for predicting the homologous recombination status of cancer using information derived from DNA sequencing of cancerous tissue, in accordance with some embodiments of the present disclosure. [Figure 1B] 1A-1C collectively show block diagrams of example computing devices for predicting the homologous recombination status of cancer using information derived from DNA sequencing of cancerous tissue, in accordance with some embodiments of the present disclosure. [Diagram 2] 1 provides a flowchart of an exemplary method for predicting the homologous recombination status of a cancer using information derived from DNA sequencing of cancerous tissue, according to some embodiments of the present disclosure. [Diagram 3] SUMMARY OF THE DISCLOSURE An example method for generating a clinical report based on information generated from the analysis of one or more patient samples is provided. [Figure 4] 1 illustrates example inputs for an HRD classification model, according to some embodiments of the present disclosure. [Diagram 5] FIG. 1 shows an exemplary bioinformatics pipeline for tumor-normal match variant calling and tumor-only calling according to some embodiments of the present disclosure. [Figure 6] FIG. 13 shows that paired-end reads from tumor and normal isolates are compressed and stored separately under the same sequence identifier, according to some embodiments of the present disclosure. [Figure 7] 1 illustrates quality modification of a FASTQ file according to some embodiments of the present disclosure. [Figure 8] 1 shows steps for obtaining tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 9] FIG. 1 shows steps for calling variants from tumor and normal BAM alignment files according to some embodiments of the present disclosure. [Figure 10] 1 illustrates an exemplary system for generating HRD calls and required outputs, according to some embodiments of the present disclosure. [Figure 11] 1 illustrates an exemplary display of text and images showing HRD information, according to some embodiments of the present disclosure. [Figure 12]1 shows an exemplary display of a list of genetic variants associated with genes in the homologous recombination DNA repair pathway and / or genes that interact with this pathway, according to some embodiments of the present disclosure.
[0020] Like reference numerals refer to corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] The present disclosure provides systems and methods for predicting the homologous recombination status of cancer using information derived from DNA sequencing of cancerous tissues, and improving treatment prediction and outcome. In some embodiments, sequencing data from matched cancerous tissues and germline tissues are used together to improve the accuracy of prediction.
[0022] definition The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to be limiting of the invention. When used in the description of the invention and in the claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly dictates otherwise. It will also be understood that, as used herein, the term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items. It will also be understood that, as used herein, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, to the extent that the terms "including", "include", "having", "has", "with" or variations thereof are used in either the detailed description and / or the claims, such terms are inclusive in a manner similar to the term "comprising".
[0023] As used herein, the term "if" may be interpreted to mean "if" or "when," or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if determined" or "if (a stated condition or event) is detected" may be interpreted to mean "when determining" or "in response to determining," or "when (a stated condition or event) is detected," or "in response to detecting (a stated condition or event)," depending on the context.
[0024] It will also be understood that, although terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject can be referred to as a second subject, and similarly, a second subject can be referred to as a first subject, without departing from the scope of the present disclosure. A first subject and a second subject are both subjects, but are not the same subject. Furthermore, the terms "subject", "user", and "patient" are used interchangeably herein.
[0025] As used herein, the term "subject" refers to a living or non-living human being. In some embodiments, the subject is a human being of any stage, male or female (e.g., a man, a woman, or a child).
[0026] As used herein, the terms "control", "control sample", "reference", "reference sample", "normal" and "normal sample" describe a sample from an otherwise healthy subject that does not have a particular condition. In one example, the methods disclosed herein can be performed on a subject with a tumor, and the reference sample is a sample taken from the subject's healthy tissue. The reference sample can be obtained from the subject or a database. The reference can be, for example, a reference genome used to map sequence reads obtained from sequencing a sample from the subject. The reference genome can refer to a haploid or diploid genome where sequences are read from a biological sample and the constituent samples can be aligned and compared. An example of a constituent sample can be the DNA of white blood cells obtained from a subject. In the case of a haploid genome, only one nucleotide can be present at each locus. In the case of a diploid genome, heterozygous loci can be identified, and each heterozygous locus can have two alleles, and either allele can allow for alignment matches to the locus.
[0027] As used herein, the term "locus" refers to a position (e.g., site) in a genome, for example, on a particular chromosome. In some embodiments, a locus refers to a single nucleotide position in a genome, i.e., on a particular chromosome. In some embodiments, a locus refers to a small group of nucleotide positions in a genome, for example, as defined by mutations (e.g., substitutions, insertions, or deletions) of consecutive nucleotides in a cancer genome. Because normal mammalian cells have a diploid genome, a normal mammalian genome (e.g., a human genome) generally has two copies of every locus in the genome, or at least two copies of every locus located on an autosome, for example, one copy on a maternal autosome and one copy on a paternal autosome.
[0028] As used herein, the term "allele" refers to a particular sequence of one or more nucleotides at a chromosomal locus.
[0029] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either the predominant allele represented at that chromosomal locus within a population of a species (e.g., a "wild type" sequence) or an allele that is predefined within a reference genome for the species.
[0030] As used herein, the term "variant allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either not the predominant allele represented at that chromosomal locus within a population of a species (e.g., not a "wild type" sequence) or is not a predefined allele within the reference genome of the species.
[0031] As used herein, the term "single base variant" or "SNV" refers to the substitution of a nucleotide at a position (e.g., site) of a nucleotide sequence, e.g., a sequence read from an individual, with a different nucleotide. The substitution of a first nucleobase X with a second nucleobase Y can be indicated as "X>Y". For example, a SNV of cytosine to thymine can be indicated as "C>T".
[0032] As used herein, the term "mutation" or "variant" refers to a detectable change in the genetic material of one or more cells. In certain instances, one or more mutations can be found in cancer cells and identify the cancer cells (e.g., driver and passenger mutations). Mutations can be transmitted from the apparent cell to daughter cells. One of skill in the art will understand that a genetic mutation in a parent cell (e.g., a driver mutation) can induce additional, different mutations in daughter cells (e.g., passenger mutations). Mutations generally occur in nucleic acids. In certain instances, a mutation can be a detectable change in one or more deoxyribonucleic acids or fragments thereof. Mutations generally refer to nucleotides that have been added, deleted, substituted, inverted, or converted to a new position in a nucleic acid. Mutations can be natural or experimentally induced. A mutation in a sequence in a particular tissue is an example of a "tissue-specific allele." For example, a tumor can have a mutation that results in an allele at a locus that does not occur in normal cells. Another example of a "tissue-specific allele" is a fetal-specific allele that occurs in fetal tissue but not in maternal tissue.
[0033] As used herein, the term "loss of heterozygosity" refers to the loss of one copy of a segment (e.g., including part or all of one or more genes) of the genome of a diploid subject (e.g., human) or the loss of one copy of a sequence encoding a functional gene product in the genome of a diploid subject, a tissue of a subject, e.g., a cancerous tissue. As used herein, when referring to a metric that represents the loss of heterozygosity across the genome of a subject, the loss of heterozygosity is caused by the loss of one copy of various segments in the genome of the subject. The loss of heterozygosity across the genome can be estimated without sequencing the entire genome of the subject, and such methods for such estimation based on gene panel targeting-based sequencing methodology are described in the art. Thus, in some embodiments, the metric that represents the loss of heterozygosity across the genome of a tissue of a subject is expressed as a single value, e.g., a percentage or fraction of the genome. In some cases, a tumor is composed of various subclonal populations, each of which may have different degrees of loss of heterozygosity across the entire genome of the respective genome. Thus, in some embodiments, loss of heterozygosity across the genome of a cancerous tissue refers to the average loss of heterozygosity across a heterogeneous tumor population. As used herein, when referring to a metric of loss of heterozygosity in a particular gene, for example, a DNA repair protein, such as a protein involved in the homologous DNA recombination pathway (e.g., BRCA1 or BRCA2), loss of heterozygosity refers to the complete or partial loss of one copy of the gene that codes for the protein in the genome of the tissue, and / or a mutation in one copy of the gene that prevents translation of the full-length gene product, for example, a frameshift or truncation (creating a premature stop codon) mutation in the gene of interest. In some cases, a tumor is composed of various subclonal populations, each of which may have a different mutation status in the gene of interest. Thus, in some embodiments, loss of heterozygosity of a particular gene of interest is represented by the average value of loss of heterozygosity of the gene across all sequenced subclonal populations of the cancerous tissue.In other embodiments, loss of heterozygosity of a particular gene of interest is represented by counting the unique occurrences of loss of heterozygosity in the gene of interest across all sequenced subclonal populations of cancerous tissue (e.g., the number of unique frameshift and / or truncating mutations in the gene identified in the sequencing data).
[0034] As used herein, the term "cancer," "cancerous tissue," or "tumor" refers to an abnormal mass of tissue in which the growth of the mass exceeds and is uncoordinated with that of normal tissue. A cancer or tumor can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, rate of growth, local invasion, and metastasis. "Benign" tumors can be well differentiated and are characterized by slower growth than malignant tumors, remaining localized at the primary site. Additionally, in some cases, benign tumors lack the ability to invade, invade, or metastasize to distant sites. "Malignant" tumors can be poorly differentiated (anaplastic) and have characteristically rapid growth with progressive infiltration, invasion, and destruction of surrounding tissue. Additionally, malignant tumors can have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is uncoordinated with that of normal tissue. Thus, a "tumor sample" refers to a biological sample obtained or derived from a tumor of a subject, as described herein.
[0035] As used herein, the terms "sequencing," "sequence determination," and the like generally refer to any and all biochemical processes that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as an mRNA transcript or a genomic locus.
[0036] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence generated by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), or in some cases, from both ends of the nucleic acid (e.g., paired-end read, double-end read). The length of the sequence read is often related to the particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are of an average, median, or arithmetic mean length between about 15 bp and 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp. In some embodiments, the sequence reads are of an average, median, or arithmetic mean length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads that do not vary much, e.g., most sequence reads can be less than 200 bp. A sequence read (or sequencing read) can refer to sequence information that corresponds to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides throughout the entire nucleic acid fragment.Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, for example, hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR), linear amplification using a single primer, or isothermal amplification.
[0037] As used herein, the term "read segment" or "read" refers to any nucleotide sequence including sequence reads obtained from an individual and / or nucleotide sequence derived from an initial sequence read from a sample obtained from an individual. For example, a read segment can refer to an aligned sequence read, a folded sequence read, or a stitched read. Additionally, a read segment can refer to an individual nucleotide base, such as a single base variant.
[0038] As used herein, the term "reference exome" refers to any particular known, sequenced, or characterized exome, whether partial or complete, of any tissue from any organism or pathogen that can be used to reference sequences identified from a subject. Exemplary reference exomes used for human subjects, as well as many other organisms, are provided in the online genome browser hosted by NCBI (National Center for Biotechnology Information).
[0039] As used herein, the term "reference genome" refers to any particular known, sequenced, or characterized genome, whether partial or complete, of any organism or pathogen that can be used to reference sequences identified from a subject. Exemplary reference genomes used for human subjects and many other organisms are provided in online genome browsers hosted by NCBI (National Center for Biotechnology Information) or UCSC (University of California, Santa Cruz). "Genome" refers to the complete genetic information of an organism or pathogen, expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genomic sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genomic sequence from one or more human individuals. A reference genome can be considered a representative example of a gene set for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).
[0040] As used herein, the term "assay" refers to a technique for determining a characteristic of a substance, e.g., a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining the variation in copy number of a nucleic acid in a sample, the methylation state of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation state of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those skilled in the art can be used to detect any characteristic of a nucleic acid described herein. Characteristics of a nucleic acid can include sequence, genomic identity, copy number, methylation state at one or more nucleotide positions, size of the nucleic acid, the presence or absence of a mutation of a nucleic acid at one or more nucleotide positions, and the pattern of fragmentation of a nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). An assay or method can have a particular sensitivity and / or specificity, and their relative utility as a diagnostic tool can be measured using the ROC-AUC statistic.
[0041] The term "classification" can refer to any number or other letter associated with a particular property of a sample. For example, in some embodiments, the term "classification" can refer to the type of cancer in a subject or sample, the stage of cancer in a subject or sample, the prognosis of cancer in a subject or sample, the tumor burden of a subject, the presence of tumor metastases in a subject, etc. The classification can be binary (e.g., positive or negative) or more levels of classification (e.g., a scale of 1-10 or 0-1). The terms "cutoff" and "threshold" can refer to a predetermined number used in an operation. For example, a cutoff size can refer to a size above which fragments are excluded. A threshold can be a value above or below which a particular classification is applied. Any of these terms can be used in any of these contexts.
[0042] Some aspects are described below with reference to exemplary applications for illustration. It should be understood that numerous specific details, relationships, and methods are described to provide a complete understanding of the features described herein. However, those skilled in the art will readily recognize that the features described herein can be implemented without one or more specific details, or in other ways. The features described herein are not limited by the described order of acts or events, as some acts may occur in different orders and / or simultaneously with other acts or events. Furthermore, not all described acts or events are required to implement a methodology in accordance with the features described herein.
[0043] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0044] Exemplary System Embodiments A detailed description of a system 100 for determining a homologous recombination pathway status of a cancer in a test subject and / or training an algorithm for determining a homologous recombination pathway status of a cancer is described in conjunction with Figures 1A-1B, which thus collectively show the topology of the system according to an embodiment of the present disclosure.
[0045] Referring to FIG. 1A, in a typical embodiment, the system 100 includes one or more computers. For illustrative purposes, in FIG. 1A, the system 100 is represented as a single computer that includes all the functionality for identifying interactions in a complex biological system using data from a cell-based assay. However, in some embodiments, the functionality for determining the homologous recombination pathway status of cancer in a test subject is distributed across any number of networked computers, and / or is present on each of a plurality of networked computers, and / or is hosted on one or more virtual machines at a remote location accessible via a communication network 105. Those skilled in the art will understand that any of a wide variety of different computer topologies may be used in the present application, and all such topologies are within the scope of the present disclosure.
[0046] Details of an exemplary system are now described in conjunction with Figure 1. Figure 1 is a block diagram illustrating a system 100 according to some implementations. The device 100 in some implementations includes at least one or more processing units CPU 102 (also called a processor), one or more network interfaces 104, a user interface 106 including, for example, a display 108 and / or a keyboard 110, a memory 111, and one or more communication buses 114 for interconnecting these components, which optionally include circuitry (sometimes called a chipset) that interconnects and controls communication between the system components. The memory 111 may be non-persistent memory, persistent memory 112, or any combination thereof. Non-persistent memory typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, while persistent memory typically includes CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid-state storage device. Regardless of its particular implementation, memory 111 includes at least one non-transitory computer-readable storage medium and stores thereon computer-executable executable instructions, which may be in the form of programs, modules, and data structures.
[0047] In some embodiments, as shown in FIG. 1A, the memory 111 stores: · An operating system 116 that handles various basic system services and contains instructions for performing hardware-dependent tasks. · An optional network communications module (or instructions) 118 for connecting the system 100 to other devices and / or communications networks 105. A first test dataset 120-1 comprising, in electronic form, a first plurality of sequence reads 122 (e.g., 122-1-1, ..., 122-1-N) of a first DNA sample from a test subject, the first DNA sample comprising DNA molecules from cancerous tissue of the subject. A second test dataset 120-2 including a second plurality of sequence reads 122 (e.g., 122-2-1, ..., 122-2-M) in electronic form of a second DNA sample from the test subject, the second DNA sample consisting of DNA molecules from non-cancerous tissue of the subject. a test genomic data construct 128 including one or more features of genomes of cancerous and non-cancerous tissues of a subject, generated based on the first plurality of sequence reads and the second plurality of sequence reads, and input to a classifier trained to distinguish between cancers with a homologous recombination pathway deficiency and cancers without a homologous recombination pathway deficiency, the test genomic data construct 128 including: o As shown in FIG. 1B, the heterozygous state in the genome of the cancerous tissue of the subject (e.g., a first dataset) 132 for a first plurality of DNA damage repair genes 130-1. o A genome-wide measure of loss of heterozygosity of the subject's cancerous tissue (e.g., a first dataset) 134, wherein the genome-wide measure of loss of heterozygosity of the subject's cancerous tissue is optionally determined by determining genomic loss of heterozygosity in a first plurality of sequence reads 136 and normalizing the determined loss of heterozygosity by an estimate of tumor purity for the first plurality of sequence reads 138, the measure of loss of heterozygosity 134. - A measure of detected mutant alleles in the genome of the cancerous tissue of the subject (e.g., a first dataset) 140-1 for a second plurality of DNA damage repair genes 130-2. o A measure of detected mutant alleles in the genome of the non-cancerous tissue of interest (e.g., a second dataset) 140-2 for a second plurality of DNA damage repair genes 130-2. A classifier training module 170 for training a disease classifier 173 to distinguish disease states, for example using training data stored in a training genome data structure 176. A disease classifier 173, e.g., one or more homologous recombination pathway classifiers 174 for distinguishing between cancers with homologous recombination pathway deficiencies and cancers without homologous recombination pathway deficiencies. · A classifier evaluation module 171 for evaluating a disease classifier. A disease classification module 172 for determining the homologous recombination pathway status of the test subject, for example, by evaluating the test genome data construct 128 with a trained disease classifier 173. a training genomic data structure 176 that stores training genomic data that can be used to train an algorithm, e.g., a disease classifier 173, for determining a homologous recombination pathway state of a cancer for each training subject, the training genomic data structure 176 including a homologous recombination pathway state 190 for one or more features of the genome of the cancer of each training subject and the genome of a non-cancerous tissue of each training subject, the training genomic data structure 176 including: o As shown in FIG. 1B, a heterozygous state 180 in the genome of the subject's cancerous tissue for a first plurality of DNA damage repair genes 178-1. 〇 A measure of loss of heterozygosity 182 across the genome of a cancerous tissue of a subject, wherein the measure of loss of heterozygosity across the genome of a cancerous tissue of a subject is optionally determined by determining loss of genomic heterozygosity in a first plurality of sequence reads 184 and normalizing the determined loss of heterozygosity by an estimate of tumor purity 186 for the first plurality of sequence reads. A measure of detected mutant alleles in the genome of the cancerous tissue of interest for a second multiple DNA damage repair gene 178-2 188-1. A measure of detected mutant alleles in the genome of the subject's non-cancerous tissue for a second multiple DNA damage repair gene 178-2 188-2.
[0048] In some implementations, modules 118, 170, 171 and / or 172 and / or data stores 120, 128 and / or 176 are accessible within any browser (e.g., installed on a phone, tablet, or laptop / desktop system). In some embodiments, modules 118, 120, 170, 171 and / or 172 run on a native device framework and are downloadable to system 100 running operating system 116, such as Windows, macOS, Linux operating system, Android OS, or iOS.
[0049] In some implementations, one or more of the above data elements or modules of system 100 are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above data, modules, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise reconfigured in various implementations. In some implementations, memory 111 optionally stores a subset of the above modules and data structures. Furthermore, in some embodiments, memory 111 stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than elements of system 100, which are addressable by system 100 so that system 100 can retrieve all or a portion of such data when needed.
[0050] While FIG. 1 illustrates a "system 100," this diagram is intended as a functional description of various features that may be present in a computer system, rather than as a structural schematic of the implementations described herein. In practice, and as will be recognized by those skilled in the art, items shown separately may be combined and some items may be separate. Additionally, while FIG. 1 illustrates certain data and modules in memory 111 (which may be non-persistent 111 or persistent memory 112), it should be understood that these data and modules, or portions thereof, may be stored in two or more memories.
[0051] Exemplary Methods Having disclosed details of system 100 for determining a homologous recombination pathway status of a cancer in a test subject and / or training an algorithm for determining a homologous recombination pathway status of a cancer, details regarding the processes and features of the system are disclosed below in accordance with various embodiments of the present disclosure. In particular, an exemplary process is described below with reference to FIG. 2. In some embodiments, such processes and features of the system are performed by modules 118, 120, 170, 171 and / or 172 as shown in FIG. 1. With reference to these methods, the systems described herein (e.g., system 100) include instructions for determining a homologous recombination pathway status of a cancer in a test subject and / or training an algorithm for determining a homologous recombination pathway status of a cancer.
[0052] 2 shows an exemplary workflow 200 for determining the homologous recombination pathway status of a cancer in a test subject, according to various embodiments of the present disclosure. Further details regarding various implementations of the steps shown in workflow 200 are described in more detail below. Those skilled in the art will be aware of suitable alternatives for carrying out each step shown in workflow 200.
[0053] In one aspect, the disclosure provides a method 200 for determining a homologous recombination pathway status of a cancer in a test subject. The method includes obtaining (202) a first plurality of sequence reads of a first DNA sample from the test subject in electronic format, the first DNA sample comprising DNA molecules from a cancerous tissue of the subject. The method includes obtaining (204) a second plurality of sequence reads of a second DNA sample from the test subject in electronic format, the second DNA sample consisting of DNA molecules from a non-cancerous tissue of the subject.
[0054] In some embodiments, the first DNA sample is from a solid tumor biopsy of the cancerous tissue of the subject.In other embodiments, the second DNA sample is from a liquid sample, for example, a liquid biopsy.Generally, the cancerous biological sample of the subject is a biopsy.Methods for obtaining samples of cancerous tissue are known in the art and depend on the type of cancer to be sampled. For example, bone marrow biopsies and isolates of circulating tumor cells can be used to obtain samples of blood cancers, endoscopic biopsies can be used to obtain samples of gastrointestinal, bladder, and lung cancers, needle biopsies (e.g., fine needle aspiration, core needle aspiration, vacuum assisted biopsy, and image guided biopsy) can be used to obtain samples of subcutaneous tumors, skin biopsies, e.g., shave biopsy, punch biopsy, incision biopsy, and excision biopsy can be used to obtain samples of skin cancers, and surgical biopsies can be used to obtain samples of cancers affecting the patient's internal organs. In some embodiments, the biological sample is a solid biopsy. In some embodiments, the solid biopsy is a macro-dissected formalin-fixed paraffin-embedded (FFPE) tissue section. In some embodiments, the biological sample comprises blood or saliva.
[0055] In some embodiments, the first plurality of sequence reads is generated by targeted sequencing using a plurality of nucleic acid probes to enrich nucleic acid from the cancerous tissue of the subject for a panel of genomic regions.In some embodiments, the first plurality of sequence reads is generated by whole genome sequencing of nucleic acid from the cancerous tissue of the subject.In some embodiments, the first plurality of sequence reads is generated by whole or partial exome sequencing of nucleic acid from the cancerous tissue of the subject.
[0056] In some embodiments, the second DNA sample is from the buffy coat preparation of the blood sample from the subject.In other embodiments, the second DNA sample is from the saliva of the subject.In general, any sample that contains genomic or exomic material that is substantially all derived from non-cancerous tissue can be used to generate the second multiple sequence reads.
[0057] In some embodiments, the second plurality of sequence reads is generated by targeted sequencing using a plurality of nucleic acid probes to enrich nucleic acid from the subject's non-cancerous tissue for a panel of genomic regions.In some embodiments, the second plurality of sequence reads is generated by whole genome sequencing of nucleic acid from the subject's non-cancerous tissue.In some embodiments, the second plurality of sequence reads is generated by whole or partial exome sequencing of nucleic acid from the subject's non-cancerous tissue.
[0058] The method then includes generating 206 a genomic data construct for the subject based on the first plurality of sequence reads and the second plurality of sequence reads, the genomic data construct including one or more features of the genomes of the cancerous and non-cancerous tissues of the subject. In some embodiments, the plurality of features includes (i) a heterozygosity state of the first plurality of DNA damage repair genes in the cancerous tissue of the subject, (ii) a measure of loss of heterozygosity across the genome of the cancerous tissue of the subject, (iii) a measure of mutant alleles detected in the second plurality of DNA damage repair genes of the genome of the cancerous tissue of the subject, and (iv) a measure of mutant alleles detected in the second plurality of DNA damage repair genes of the genome of the non-cancerous tissue of the subject.
[0059] In some embodiments, a measure of genome-wide loss of heterozygosity of a cancerous tissue of a subject is obtained by determining the loss of genome heterozygosity in a first plurality of sequence reads and normalizing the determined loss of heterozygosity by an estimate of tumor purity for the first plurality of sequence reads. That is, many "tumor biopsies" contain a residual percentage of non-cancerous cells. When estimating loss of heterozygosity from nucleic acid isolated from a tumor biopsy, the presence of nucleic acid from non-cancerous cells will skew the overall loss of heterozygosity downward. Estimating the tumor purity of a sample, e.g., the percentage of nucleic acid derived from cancerous cells rather than non-cancerous cells, can account for the presence of non-cancerous contributions to the sequencing data, providing a more accurate analysis of the loss of heterozygosity of the whole cancer genome of a subject.
[0060] In some embodiments, the heterozygous state of the first plurality of DNA damage repair genes comprises the count of the number of unique frameshift mutations detected in the first plurality of DNA damage repair genes.In some embodiments, the heterozygous state of the first plurality of DNA damage repair genes comprises the count of the number of unique truncation mutations detected in the first plurality of DNA damage repair genes.In some embodiments, the first plurality of DNA damage repair genes are genes involved in homologous recombination pathway.In some embodiments, the first plurality of DNA damage repair genes comprises BRCA1 and BRCA2.
[0061] In some embodiments, the measure of mutant alleles detected in the second plurality of DNA damage repair genes in the genome of the cancerous tissue of the subject comprises a count of the number of unique mutations associated with loss of homologous recombination detected in the first plurality of sequence reads.In some embodiments, the measure of mutant alleles detected in the second plurality of DNA damage repair genes in the genome of the non-cancerous tissue of the subject comprises a count of the number of unique mutations associated with loss of homologous recombination detected in the second plurality of sequence reads.
[0062] In some embodiments, the second plurality of DNA damage repair genes are genes involved in homologous recombination pathway.In some embodiments, the second plurality of DNA damage repair genes comprises BRCA1 and BRCA2.In some embodiments, the unique mutation associated with loss of homologous recombination in BRCA1 and BRCA2 comprises at least 25, 50, 75, 100, 125 or all of the mutations listed in Table 1.
[0063] The method then includes inputting the mutant allele genome data construct into a classifier trained to distinguish between cancers with and without homologous recombination pathway deficiencies, thereby determining the homologous recombination pathway status of the test subject (208). In some embodiments, the classifier is a neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm, as described in more detail below.
[0064] In some embodiments, the method 200 also includes treating the subject based on the HRD prediction made by the classifier. For example, in some embodiments, when the cancer of the test subject is determined to be homologous recombination deficient, the cancer is treated by administering a poly ADP ribose polymerase (PARP) inhibitor to the test subject, and when the cancer of the test subject is determined not to be homologous recombination deficient, the cancer is treated with a therapy that does not include administering a PARP inhibitor to the test subject. In some embodiments, the PARP inhibitor is selected from olaparib, veliparib, rucaparib, niraparib, and talazoparib. A summary of current FDA approvals for various PARP inhibitors is provided in Table 2 below.
[0065] In another aspect, the present disclosure provides a method for training an algorithm for determining a homologous recombination pathway status of cancer. The method includes, for each training subject in a plurality of training subjects having cancer, obtaining a corresponding genomic data construct for each training subject. The corresponding genomic training construct includes (a) the homologous recombination pathway status of the cancer of each training subject, and (b) one or more features of the genome of the cancerous tissue and non-cancerous tissue of each training subject. In some embodiments, the one or more features include (i) the heterozygosity status of a first plurality of DNA damage repair genes in the cancerous tissue of each training subject, (ii) a measure of loss of heterozygosity throughout the genome of the cancerous tissue of each training subject, (iii) a measure of mutant alleles detected in a second plurality of DNA damage repair genes of the genome of the cancerous tissue of each training subject, and (iv) a measure of mutant alleles detected in a second plurality of DNA damage repair genes of the genome of the non-cancerous tissue of each training subject. Next, the method includes training, for each training subject, a classification algorithm on at least (a) the homologous recombination pathway status of the cancer of each training subject, and (b) a plurality of features determined from a corresponding DNA sample from the cancerous tissue of each training subject.
[0066] FIG. 3 displays a flow chart of an exemplary method for generating a clinical report not based on information generated from the analysis of one or more patient specimens and the ingestion of patient health information. A clinical laboratory may receive an order, such as an order for comprehensive genomic profiling or an order for a test that provides an estimate of HRD status. The physical specimen may be provided to a laboratory for processing and analysis. The processing and analysis may include analysis that may include nucleotide and clinical information that may include an estimate of HRD status. The one or more specimens may be processed through the laboratory, which may include steps of accession, pathology review, extraction, library preparation, capture and hybridization, pooling, and sequencing. Sequencing may be performed using next generation sequencing technologies, such as short read technologies. Alternately, other sequencing methods, such as long read sequencing or other sequencing methods known in the art, may be used. The results of the sequencing may be provided to a bioinformatics pipeline. The results of the bioinformatics pipeline may be provided to variant science analysis, including interpretation of variants (including somatic and germline variants, if applicable) for pathogenicity and biological significance. The variant science analysis may also estimate microsatellite instability (MSI) or mutational burden of the tumor. Targeted treatments may be identified based on the gene, variant, and cancer type for further consideration and review by the ordering physician. In some embodiments, clinical trials for which the patient may be eligible may be identified based on the mutation, cancer type, and / or medical history. A validation step may occur, after which a report may be completed for sign-out and delivery. In some embodiments, the report includes an estimate of HRD status. In other embodiments, a second report with an estimate of HRD status may be delivered based on the information generated in the portion of the method presented in FIG. 3.
[0067] Biological samples In some embodiments, the estimated HRD status may be generated based on information about the nucleotides of the cancer and / or normal specimens. The cancer specimens may be from different subtypes of cancer, including hematological and solid tumors. In some embodiments, the sample type utilized for comprehensive genomic profiling may be fixed formalin, paraffin-embedded (FFPE) slides, peripheral blood, or bone marrow aspirates. The samples may be collected in a repository, such as potassium ethylenediaminetetraacetate (EDTA) tubes. The specimens may be tissue blocks or multiple FFPE slides, for example, up to 3 slides, up to 5 slides, up to 10 slides, or up to 20 slides. In some embodiments, the matched normal specimen is peripheral blood or saliva.
[0068] Features In some embodiments, the information used to generate the inferred HRD status may be generated by sequencing performed by a multi-gene comprehensive genomic profiling panel. The panel may analyze more than 10, more than 100, or more than 1,000 genes. The panel may be a whole-exome panel that analyzes the specimen's exome. The panel may be a whole-genome panel that analyzes the specimen's genome. In some embodiments, the information used to generate the inferred HRD status may be generated as part of a comprehensive genomic profiling test, such as a DNA-based test. The panel may identify single nucleotide polymorphisms (SNVs), insertions / deletions, copy number variations (CNVs), and gene rearrangements.
[0069] The systems and methods may take into account the mutational status of specific genes. For example, the systems and methods may take into account the mutational status of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 genes. The systems and methods may take into account the mutational status of 15-30 genes, 30-45 genes, 45-60 genes, 60-75 genes, 75-105 genes, and 1-700 genes. The systems and methods may take into account commonly mutated genes in pathways such as the HR pathway (homologous recombination repair mutated HRRm).
[0070] The systems and methods may use panels with at least 99% sensitivity for base substitutions in at least 5% of the mutant allele fraction, at least 98% sensitivity for indels in at least 5% of the mutant allele fraction, at least 95% sensitivity for CNVs from 8 or more gene copies in 30% or more tumor nuclei, and / or at least 99% sensitivity for gene sequences.
[0071] The panel may have an average sequencing depth of 500x for the tumor. The panel may have an average sequencing depth of 150x for the matched normal.
[0072] In some embodiments, a report may be sent back to the clinician with information regarding the mutation status of the patient's cancer, and the comprehensive genomic profiling information, such as a presumption of HRD status. In some embodiments, genes reported in the comprehensive genomic profiling information may be highlighted as underlying or otherwise relevant to a presumption of HRD status. The number of such genes may be 1-5, 1-10, 1-20, 1-30, 1-40, 1-50, etc. In some embodiments, genes reported as mutations in the comprehensive genomic profiling information may be highlighted as being germline or somatic alterations when detected.
[0073] In some embodiments, the systems and methods are scalable and allow for integration with other data types, such as other genes in DNA damage repair pathways, or RNA expression, and can be utilized to provide clinical decision support regarding treatment options, such as PARP inhibitor treatment options.
[0074] The bioinformatics pipeline may generate a variety of features that can be fed into the HRD prediction engine. In some embodiments, some or all of the copy number segments, truncating and stop-gain effect pathogenic variants in the BRCA genes of interest, genome-wide LOH fraction, tumor purity, and LOH in the BRCA genes are used to infer HRD status.
[0075] Tumor-normal matched sequencing analysis of patient samples in gene sequencing panels followed by a bioinformatics pipeline may be used to call SNPs and copy number variants for each patient and store them in a DNA variant dataset.
[0076] Each DNA variant dataset may be generated by processing cancer and non-cancer samples from the same patient with DNA whole-exome next-generation sequencing (NGS) to generate DNA sequencing data, which may be processed by a bioinformatics pipeline to generate DNA variant call files (among other outputs) for each sample. The cancer samples may be tissue samples or blood samples containing cancer cells. In some cases, tumor organoid samples may be processed in place of the patient's cancer samples.
[0077] More specifically, germline ("normal", non-cancerous) DNA can be extracted from either blood (e.g., if the patient has a cancer that is not a blood cancer) or saliva (e.g., if the patient has a blood cancer). A normal blood sample can be collected from the patient (e.g., with PAXgene Blood DNA Tubes) and a saliva sample can be collected from the patient (e.g., with an Oragene DNA Saliva kit).
[0078] Blood cancer samples can be collected from patients (e.g., in EDTA collection tubes). Macro-dissected FFPE tissue sections from solid tumor samples (which may be mounted on histopathology slides) can be analyzed by a pathologist to determine the overall tumor burden in the sample and the ratio of tumor cellularity as a ratio of tumor to normal nuclei. For each section, background tissue can be excluded or removed so that the section meets a threshold of tumor purity (in one example, at least 20% of the nuclei in the section are tumor nuclei).
[0079] DNA can then be isolated from blood samples, saliva samples, and tissue sections using commercially available reagents that contain proteinase K to produce a liquid solution of DNA.
[0080] Each solution of separated DNA may be subjected to a quality control protocol to determine the concentration and / or quantity of DNA molecules in the solution, which may include the use of fluorescent dyes and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0081] For each cancer sample and each normal sample, the separated DNA molecules can be mechanically sheared to an average length using an ultrasonicator (e.g., a Covaris ultrasonicator). The DNA molecules can also be analyzed to determine fragment size, which can be done via gel electrophoresis techniques and may include the use of a device such as a LabChip GX Touch.
[0082] A DNA library can be prepared from the isolated DNA, for example, using KAPA Hyper Prep kit, New England Biolabs (NEB) kit, or similar kit. Preparation of a DNA library can include ligating adapters to DNA molecules. For example, UDI adapters, including Roche SeqCap dual-end adapters, or UMI adapters (e.g., full-length or chunky Y adapters) can be ligated to DNA molecules.
[0083] In this example, the adapters are nucleic acid molecules that can serve as barcodes to identify DNA molecules according to the sample they originate from and / or to facilitate downstream bioinformatics processing and / or next generation sequencing reactions. The sequence of nucleotides in the adapters may be unique to the samples to differentiate them. The adapters can act as seeds for the sequencing process by facilitating the binding of DNA molecules to immobilize oligonucleotide molecules on a sequencer flow cell and providing a starting point for the sequencing reaction.
[0084] The DNA library can be amplified and purified using reagents such as Axygen MAG PCR clean-up beads. The concentration and / or amount of DNA molecules can then be quantified using fluorescent dyes and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0085] DNA libraries can be pooled (two or more DNA libraries can be mixed to create a pool) and treated with reagents to reduce off-target capture (e.g., Human COT-1 and / or IDT xGen Universal Blockers). Pools can be dried under vacuum and resuspended. DNA libraries or pools can be hybridized to probe sets (e.g., probe sets specific to a panel containing approximately 100, 600, 1,000, 10,000, etc. of the 19,000 known human genes) and amplified with commercially available reagents (e.g., KAPA HiFi HotStart ReadyMix).
[0086] The pool can be incubated in an incubator, PCR machine, water bath, or other temperature regulating device to allow the probes to hybridize. The pool can then be mixed with Streptavidin coated beads or another means for capturing the hybridized DNA probe molecules, such as DNA molecules representing exons of the human genome and / or genes selected for a gene panel.
[0087] Pools can be amplified and purified two or more times using commercially available reagents, e.g., KAPA HiFi Library Amplification kit and Axygen MAG PCR clean-up beads, respectively. Pools or DNA libraries can be analyzed to determine the concentration or quantity of DNA molecules, e.g., by using fluorescent dyes (e.g., PicoGreen pool quantification) and a fluorescent microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0088] In one example, the DNA library preparation and / or whole exome capture steps can be performed in an automated system using a liquid handling robot (e.g., SciClone NGSx).
[0089] Library amplification can be performed on a device, e.g., an Illumina C-Bot2, and the resulting flow cells containing the amplified target capture DNA libraries are sequenced on a next-generation sequencer, e.g., an Illumina HiSeq 4000 or NovaSeq 6000, to a unique on-target depth selected by the user (e.g., 300x, 400x, 500x, 10,000x, etc.). Samples can be further evaluated for uniformity with each sample requiring 95% of all target bp to be sequenced to a user-selected minimum depth (e.g., 300x). The next-generation sequencer may generate FASTQ, BCL, or other files for each flow cell or each patient sample.
[0090] Bioinformatics Pipeline In certain embodiments, the bioinformatics pipeline comprises the systems and methods disclosed in this document.
[0091] FASTQ and alignment A tumor-normal match sequencing run is performed when matched normal tissue becomes available for a patient. DNA is extracted from the normal tissue, usually blood or saliva. This is then sequenced in addition to the DNA extracted from the tumor tissue. These two sequencing runs (one for the tumor tissue and one for the normal tissue) generate two FASTQ output files. The FASTQ format is a text-based format for storing both biological sequences, such as nucleotide sequences, and their corresponding quality scores. These FASTQ files are analyzed to determine the genetic variants or copy number changes present in the sample. A "matched" panel-specific workflow is run to jointly analyze the tumor-normal match FASTQ files. If a matched normal is not available, the FASTQ files from the tumor tissue are analyzed in "tumor only" mode. See Figure 5 for an example.
[0092] When two or more patient samples are processed simultaneously on the same sequencer flow cell, differences in the sequences of the adapters used for each patient sample can serve the purpose of barcoding to facilitate associating each read with the correct patient sample and placing it in the correct FASTQ file.
[0093] For efficiency, the paired-end sequencing results for each isolate are included in a split pair of FASTQ files. The forward (read 1) and reverse (read 2) sequences for each tumor and normal isolate are stored separately, but in the same order and under the same identifier. See Figure 6 for an example.
[0094] In various embodiments, the bioinformatics pipeline can filter the FASTQ data from each isolate. Such filtering includes correcting or masking sequencer errors, removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, over-represented sequences, biases caused by library preparation, amplification, or capture, and other errors (Figure 7). Whole reads, individual nucleotides, or multiple nucleotides that may be erroneous may be discarded based on a quality assessment associated with the read in the FASTQ file, the known error rate of the sequencer, and / or a comparison of each nucleotide in the read to one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be done in part or in whole by various software tools, such as software tools such as Skewer (see https: / / doi.org / 10.1186 / 1471-2105-15-182). FASTQ files may be analyzed by sequencing data QC software such as AfterQC, Kraken, RNA-SeQC, FastQC (Illumina, BaseSpace Labs, or https: / / www.illumina.com / products / by-type / informatics-products / basespace-sequence-hub / apps / fastqc.html), or another similar software program, for quality control and rapid assessment of the reads. In case of paired-end reads, the reads can be merged.
[0095] In matched panel-specific tumor-normal analysis, two FASTQ files are analyzed from each, one for the tumor and one for the normal (if available). In tumor-only analysis, only the tumor FASTQs are available for analysis.
[0096] Each read from FASTQ can be aligned to the location in the human genome that has a sequence that best matches the sequence of the nucleotides in the read. There are many software programs designed to align reads, such as Novoalign (Novocraft, Inc.), Bowtie, Burrows Wheeler Aligner (BWA), and programs that use the Smith-Waterman algorithm. Alignment is directed to the use of a reference genome (e.g., hg19, GRCh38, hg38, GRCh37, and other reference genomes developed by the Genome Reference Consortium, etc.) by determining the portion of the reference genome sequence that most likely corresponds to the sequence of the read by comparing the nucleotide sequence in each read with the portion of the nucleotide sequence in the reference genome. Alignment may generate a SAM file that stores the start and end positions of each read according to the coordinates of the reference genome and the coverage (number of reads) of each nucleotide in the reference genome. SAM files can be converted to BAM files, BAM files can be sorted, and duplicate reads can be marked for deletion to create a BAM file without duplicates. This process generates a tumor BAM file and a normal BAM file (when available) (e.g., as shown in FIG. 8). In various embodiments, the BAM files may be analyzed to detect genetic variants and other genetic features, including single nucleotide variants (SNVs), copy number variants (CNVs), gene rearrangements, etc. In various embodiments, the detected genetic variants and genetic features are analyzed as a form of quality control. For example, a pattern of detected genetic variants or features may indicate problems associated with the sample, the sequencing procedure, and / or the bioinformatics pipeline, e.g., contamination of the sample, incorrect labeling of the sample, alteration of reagents, problems with the sequencing procedure and / or the bioinformatics pipeline, etc.
[0097] SNV and indel calling Following alignment, tools like SamBAMBA can be used to mark and filter duplicates in the sorted BAMs. Software packages like freebayes and pindel are used to call variants using the sorted BAM files as input and the genome and panelbed files containing the gene targets to be analyzed as references. A raw VCF file (variant call format) file is output, indicating where the nucleotide base in the sample is not the same as the nucleotide base at that position in the reference genome. Software packages like vcfbreakmulti and vt are used to normalize the multi-nucleotide polymorphism variants in the raw VCF file, and a variant normalized VCF file is output. SNVs in the VCF are annotated using SNPEff for transcription information, mutation effect, and prevalence in the 1000 genomes database. EGFR variants are called separately through realigning tumor and normal fastq files on chr 7 using speedseq. Duplicates are marked using tools like Sambamba, and variant calling is done similar to the steps described for other chromosomes. See for example, Figure 9.
[0098] Copy number variant determination In various embodiments, the systems and methods include copy number analysis methods for calculating genomic features used to estimate HRD status. For example, in some embodiments, to assess copy number, the VCF generated from the deduplicated BAM file and variant calling pipeline can be used to calculate the read depth and variation of heterozygous germline SNVs between tumor and normal samples. If a matched normal sample is not available, a comparison of the tumor sample to a pool of process-matched normal controls can be utilized. A circular binary segmentation can be applied, and segments can be selected with highly different log2 ratios between the tumor and its comparator (matched normal or normal pool). Approximate integer copy numbers can be assessed from a combination of differential coverage in the segmented regions and estimates of stromal mixtures (e.g., tumor purity, or the portion of the sample that is tumor vs. non-tumor) generated by the analysis of heterozygous germline SNVs.
[0099] Determining loss of heterozygosity In some embodiments, LOH can be determined through the use of copy number calling algorithm. First, the tumor purity and copy state of tumor genome can be estimated using expectation maximization algorithm (EM). Estimation of copy state and tumor purity may involve the following steps: 1) read alignment and normalization; 2) calculation of B allele frequency and deviation; 3) preliminary estimation of tumor purity; 4) genome segmentation; and 5) refinement of initial tumor purity estimation by EM algorithm to estimate copy state and LOH.
[0100] Read alignment and normalization To calculate probe target coverage, sequenced reads from tumors can be aligned to the human reference genome and normalized by length and depth, as well as GC content. Reads from normal tissues can be treated similarly when available. If a matched normal is not available, a normal pool consisting of read coverage from normal healthy individuals not known to have cancer can be used. To select a gender-matched normal pool, a gender estimation step can be performed by mapping the variants to the X chromosome along with the X chromosome coverage. From the normal pool, the nearest neighbors can be selected, for example through applying a PCA selection step. Their coverage values can be used to normalize the tumor coverage. This PCA selection increases the sensitivity of somatic CNV detection. Finally, read coverage can be expressed as the ratio of tumor coverage to normal coverage and log2 transformed.
[0101] Calculation of B allele frequency and deviation Heterozygous variants contain useful information about copy number and LOH. These variants can be mined from somatic and germline variant calls made using freebayes and pindel. B allele frequency (BAF) deviation from expected normal is calculated for each heterozygous SNP and is also expressed as BAF log odds ratio. If the variant is normal germline, the BAF deviation from normal should be close to 0. For variants showing LOH, the BAF will deviate significantly from 0.
[0102] Preliminary estimate of tumor purity An initial estimate of tumor purity can be obtained from somatic variant and BAF data to be used as input for the EM algorithm. The maximum VAF of a somatic variant should theoretically be equal to the tumor purity. This is the somatic estimate of tumor purity. From the BAF data, for variants showing a log odds ratio greater than 2, it is clearly LOH, and such significant deviations are expected only when a copy is lost or the copy is neutral. Twice the maximum possible VAF of such a variant should theoretically be equal to the tumor purity, and corresponds to the estimate of the BAF. These two estimates are averaged to form the initial estimate of tumor purity.
[0103] Genome segmentation Bivariate segmentation of the genome is performed using tumor-to-normal coverage ratio and BAF log-odds data. A series of rolling T tests are performed across the genome using an algorithm similar to circular binary segmentation to identify sections of the genome where significant copy number switching is observed. This results in the aggregation of the whole genome into segments, each with a distinct copy number profile. Segmentation branching and pruning threshold parameters control how well segmentation and focal segment detection are possible and are optimized for Tempus data.
[0104] Refining the initial tumor purity estimate and estimating copy states and LOH with the EM algorithmFrom the initial estimate of tumor purity, tumor purity values ranging from half the tumor purity to the maximum possible value are iterated to estimate the optimal copy state for each genomic segment. For each tumor purity estimate and genomic segment, expected log ratios and BAFs are calculated for each copy state ranging from 0 to 20, allowing only meaningful copy state combinations. The likelihood of the observed coverage and BAF is then calculated given these expectations from the bivariate probability density function to create a likelihood matrix. The most likely copy state is returned from this matrix. This process is repeated for all segments, building the segments into an optimal copy state map. Repeating this step for all tumor purities produces a tumor purity likelihood matrix, and the most likely tumor purity with the least model error is returned as the final estimate. Once copy state assignments are available for all genomic segments, segments with a minor copy number of 0 are assigned LOH. These segments can be either one copy loss, copy neutral, or high-order LOH, depending on the tumor purity.
[0105] Tumor purity To calculate tumor purity, an initial tumor purity estimate is obtained from somatic variants and germline B allele frequencies, which is refined using a greedy algorithm that evaluates the likelihood of tumor purity given the log ratio of tumor-normal coverage to tumor-normal coverage and the B allele frequency deviation from normal expectation. The algorithm iterates through a set of tumor purity ranges surrounding the initial estimate and returns the tumor purity in a maximum likelihood fashion.
[0106] Loss of heterozygosity For genome-wide loss of heterozygosity (LOH) estimation, each SNP was assessed for LOH based on the germline variant allele fraction and the deviation of the B allele frequency from normal expectation. A binary 0 / 1 system was used to assign no LOH / yes LOH to obtain the average fraction of genomic bases under LOH. The number of bases undergoing LOH can be divided by the total number of bases analyzed using copy number methods such as those disclosed in this patent to determine an estimate of the genome-wide LOH fraction. In one example, the estimate of the genome-wide LOH fraction may represent LOH in somatic (cancer) samples that may not be present in germline (normal) samples.
[0107] The average LOH of the BRCA1 and BRCA2 genes can be determined in a similar manner, but considering only the coordinates of the two genes. In one example, the LOH of the BRCA1 / 2 genes may represent LOH in somatic (cancer) samples that may not be present in germline (normal) samples.
[0108] Counting pathogenic variants To count the number of pathogenic variants for a given gene, we used all SNPs called for each patient and matched them against a curated reference mutation list that contains a list of known pathogenic and truncating BRCA variants (e.g., BRCA1 and BRCA2). We then obtained the number of pathogenic variants based on the overlap of SNP positions. Separate counts of somatic and germline variants are also output for BRCA. A sum of the two counts can also be generated.
[0109] In some embodiments, the pathogenic variants used in the systems and methods described herein include one or more of the variants listed in Table 1. In some embodiments, the pathogenic variants used in the systems and methods described herein include at least 5, 10, 15, 20, 25, 30, 40, 50, 75, 100, 125, or all of the variants listed in Table 1. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]
[0110] Positive HRD call based on HRD markers In various embodiments, if certain markers of HRD are detected, the system and method disclosed herein will return a positive HRD call.In one example, if pathogenic stop-gain or frameshift variants exist in BRCA1 or BRCA2, a positive HRD call will be returned.In another example, if the ratio of genome-wide loss of heterozygosity combined with the loss of heterozygosity of BRCA1 or BRCA2 exceeds the threshold value indicating BRCA mutation, a positive HRD call will be returned.
[0111] classifier In general, it is understood that many different classification algorithms may be used in the systems and methods described herein. For example, in some embodiments, the model is a neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, a multinomial logistic regression algorithm, a linear model, or a linear regression algorithm.
[0112] In some embodiments, the classification algorithm used in the systems and methods described herein is a random forest algorithm. In some embodiments, the trained classification method includes a trained classifier stream. In some embodiments, as a non-limiting example, the trained classifier stream is a decision tree. Decision tree algorithms suitable for use as classification models described herein are described, for example, in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, 395-396, which is incorporated herein by reference. Tree-based methods divide the feature space into a set of rectangles and fit a model (such as a constant) to each one. In some embodiments, the decision tree is a random forest regression. One specific algorithm that can be used as a classification model is a classification and regression tree (CART). Other examples of specific decision tree algorithms that can be used as classifiers include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York. 396-408, 411-412, which is incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests--Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.
[0113] In some embodiments, tumor organoids with various BRCA LOH status, pathogenic mutations, and genome-wide LOH measurements can be grown and treated with PARP inhibitors to obtain in vitro PARP drug responses. Samples can span a broad cancer cohort. Tumor cell lines expected to be sensitive to PARP can be tested alongside negative controls that do not have HRD mutations. PARP outcome data can be used to refine input features of a random forest classifier. Additional information can be gleaned from mutation signatures of the HRD pathway and other genes. See, for example, Gulhan DC, Lee JJ, Melloni GEM, Cortes-Ciriano I, Park PJ, "Detecting the mutational signature of homologous recombination deficiency in clinical samples," Nat Genet., 51(5):912-19 (2019), incorporated herein by reference.
[0114] In alternative embodiments, instead of or in addition to training a random forest classifier to generate HRD calls, the systems and methods use business logic. For example, in some embodiments, a business rule set such as that shown in FIG. 10 is used in the systems and methods described herein.
[0115] In some embodiments, the classification algorithm using the systems and methods described herein is a regression algorithm. The regression algorithm can be any type of regression. For example, in some embodiments, the regression algorithm is logistic regression. Logistic regression algorithms are disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated herein by reference. In some embodiments, the regression algorithm is logistic regression with Lasso, L2, or elastic net regularization.
[0116] In some embodiments, the classification algorithm using the systems and methods described herein is a neural network. Examples of neural network algorithms, including convolutional neural network algorithms, are disclosed, for example, in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.
[0117] In some embodiments, the classification algorithm used in the systems and methods described herein is a support vector machine (SVM). Examples of SVM algorithms are, for example, Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, Second Edition, 2001, John Wiley&Sons, Inc., pp. 259, 262-265; Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al. al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVMs separate a particular set of binary labeled data training sets using a hyperplane that is maximally far from the labeled data. When linear separation is not possible, SVMs work in conjunction with a "kernel" approach that automatically achieves a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.
[0118] In some embodiments, the machine learning model includes a logistic regression classifier. In other embodiments, the machine learning or deep learning model can be one of a decision tree, an ensemble (e.g., bagging, boosting, random forest), a gradient boosting machine, a linear regression, Naive Bayes, or a neural network. The HRD model includes learned weights of features that are adjusted during training. Here, the term "weight" is used generically to represent the amount of learning associated with any given feature of the model, regardless of the particular machine learning technique being used. In some embodiments, the cancer index score is determined by inputting feature values derived from one or more DNA sequences (or DNA sequence reads thereof) into a machine learning or deep learning model.
[0119] In some embodiments, for example, when the HRD assessment model is a neural network (e.g., a conventional or convolutional neural network), the output of the disease classifier is, for example, a classification, either cancer positive or cancer negative. However, in some embodiments, to provide a continuous or semi-continuous value for the output of the model rather than a classification, a hidden layer of the neural network, for example, a hidden layer immediately before the output layer, is used as the output of the classification model.
[0120] Thus, in some embodiments, the model includes: (i) an input layer for receiving values of a plurality of genotypic features, the plurality of genotypic features having a first dimensionality; (ii) an embedding layer including a set of weights, the embedding layer directly or indirectly receiving an output of the input layer, the output of the embedding layer being a model score set having a second dimensionality less than the first dimensionality; and (iii) an output layer that directly or indirectly receives the model score set from the embedding layer. In some embodiments, the output of the classifier is the output of a set of neurons associated with a hidden layer in the neural network, called an embedding layer. In such embodiments, each such neuron in the embedding layer is associated with a weight and an activation function, and the output consists of the output of each such activation function. In some embodiments, the activation function of the neurons in the embedding layer is a rectified linear unit (ReLU), tanh, or sigmoid activation function. In some such embodiments, the neurons of the embedding layer are fully connected to each of the inputs of the input layer. In some such embodiments, each neuron of the output layer is fully connected to each neuron of the embedding layer. In some embodiments, each neuron of the output layer is associated with a softmax activation function. In some embodiments, one or more of the embedding layers and the output layer are not fully connected.
[0121] Patient Reports In some embodiments, a patient report is generated based on the output of the classifier. The report can be presented to the patient, physician, medical professional, or researcher as a digital copy (e.g., a JSON object, a pdf file, or an image on a website or portal), a hard copy (e.g., printed on paper or another tangible medium), or another format.
[0122] In some embodiments, the report includes information related to the specimen's HRD status, the genetic variants detected, other characteristics of the patient's sample, and / or clinical records. The report may include clinical trials for which the patient is eligible, therapies to which the patient may be matched, and / or expected side effects if the patient receives a given therapy based on the HRD status, the genetic variants detected, other characteristics of the sample, and / or clinical records. In one example, if a patient specimen is predicted to have HRD, the patient may be matched to a PARP inhibitor, platinum-based chemotherapy, and / or additional DNA damaging therapy.
[0123] The results contained in the report and / or additional results (e.g., from a bioinformatics pipeline) can be used to analyze a database of clinical data, in particular to determine whether there is a trend indicating that the treatment has slowed the progression of cancer in other patients with the same or similar results as the specimen. The results can also be used to design tumor organoid experiments. For example, organoids may be genetically engineered to have the same characteristics as the specimen and observed after exposure to the treatment to determine that the treatment can reduce the growth rate of the organoids and therefore is likely to reduce the growth rate of the patient associated with the specimen.
[0124] In this example, the HRD information can be stored in a report object, such as a JSON object, for further processing and / or display. For example, information from the report object can be used to prepare a lab test report for return to an ordering physician. The information may be provided as a combination of text, images, and / or audio. An exemplary display of text and images showing the HRD information is presented as FIG. 11.
[0125] In some embodiments, the report also includes a list of genetic variants associated with genes in the homologous recombination DNA repair pathway and / or genes that interact with this pathway. An exemplary display of this list is presented as FIG.
[0126] treatment In some aspects, the systems and methods disclosed herein may be used as companion diagnostics. For example, in some embodiments, the predicted HRD status may be used by clinicians to make decisions to treat cancer with PARP inhibitors.
[0127] Table 2 lists several PARP inhibitors and their FDA approval or clinical trial status for various cancer types in 2019. This table illustrates the broad potential utility of PARP inhibitors for patients who test positive for HRD. [Table 2]
[0128] In some embodiments, the estimated HRD status may be used by clinicians to make a decision to treat a cancer by adding platinum to standard neoadjuvant chemotherapy. Because adding a platinum agent to a standard combination chemotherapy increases the toxicity of the treatment, patients would benefit from an estimated HRD that indicates whether the cancer is likely to be treated through the combination of a platinum agent and a standard combination chemotherapy.
[0129] In some embodiments, PARP inhibitors are specifically approved for the treatment of cancers that harbor germline alterations.For example, olaparib is approved for germline BRCA (gBRCA) positive ovarian cancer that has been treated with at least three chemotherapy regimens, and talozaparib is approved for gBRCA positive, HER2 negative localized or metastatic breast cancer.Detecting germline variants in BRCA or other genes related to DNA repair pathways may help doctors decide to prescribe PARPi.
[0130] Implementation using digital and laboratory healthcare platforms The methods and systems described herein can be utilized in conjunction with or as part of a digital and laboratory healthcare platform generally directed to medical care and research. It should be understood that many of the above-mentioned methods and systems can be used in conjunction with such platforms. One example of such a platform is described in U.S. Patent Application No. 16 / 657,804, filed October 18, 2019, entitled "Data Based Cancer Research and Treatment Systems and Methods," which is incorporated herein by reference in its entirety for all purposes.
[0131] For example, an implementation of one or more embodiments of the above-described methods and systems may include microservices that constitute a digital and laboratory healthcare platform supporting HRD detection. An embodiment may include a single microservice for performing and delivering a ____, or may include multiple microservices, each with a specific role that together implements one or more of the above-described embodiments. In one example, a first microservice may perform computation of genomic features to deliver features to a second microservice for training an HRD model. Similarly, the second microservice may perform training of an HRD model and deliver the trained HRD model to a third microservice, according to one embodiment above. The third microservice may use the trained HRD model to analyze data associated with a specimen and determine the likelihood that the specimen has HRD.
[0132] When the above embodiments are implemented in one or more microservices in conjunction with or as part of a digital and laboratory healthcare platform, one or more of such microservices may be part of an order management system that orchestrates the sequence of events as necessary at the appropriate time and in the appropriate order necessary to instantiate the above embodiments. A microservices-based order management system is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 873,693, filed July 12, 2019, entitled "Adaptive Order Fulfillment and Tracking Methods and Systems," which is incorporated by reference herein in its entirety for all purposes.
[0133] For example, continuing with the first and second microservices described above, the order management system may notify the first microservice that an order for _______ has been received and is ready for processing. When delivery of the ________ is ready for the second microservice, the first microservice executes and notifies the order management system. Further, the order management system may determine that the execution parameters (preconditions) of the second microservice have been met, including that the first microservice has completed, and notify the second microservice that it can continue to process the order for ________, according to one embodiment described above.
[0134] When the digital and laboratory healthcare platform further includes a genetic analysis system, the genetic analysis system can include a targeting panel and / or a sequencing probe. Examples of targeting panels are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 902,950, filed September 19, 2019, entitled "System and Method for Expanding Clinical Options for Cancer Patients using Integrated Genomic Profiling," which is incorporated herein by reference in its entirety for all purposes. In one example, the targeting panel can enable delivery of __ next-generation sequencing results according to one embodiment described above. Examples of next-generation sequencing probe designs are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 924,073, filed October 21, 2019, entitled "Systems and Methods for Next Generation Sequencing Uniform Probe Design," which is incorporated herein by reference in its entirety for all purposes.
[0135] If the digital and laboratory healthcare platform further includes a bioinformatics pipeline, the above-described methods and systems can be utilized after completion or substantial completion of the systems and methods utilized in the bioinformatics pipeline. As an example, the bioinformatics pipeline may receive next-generation genetic sequencing results and return a set of binary files, such as one or more BAM files, reflecting DNA and / or RNA read counts aligned to a reference genome. The above-described methods and systems can be utilized, for example, to ingest the DNA and / or RNA read counts and generate a result.
[0136] If the digital and laboratory healthcare platform further includes an RNA data normalizer, any RNA read counts can be normalized before processing the embodiment as described above. Examples of RNA data normalizers are disclosed, for example, in U.S. Patent Application No. 16 / 581,706, entitled "Methods of Normalizing and Correcting RNA Expression Data," filed September 24, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0137] Where the digital and laboratory healthcare platform further includes a genetic data deconvolutor, any of the systems and methods for deconvolution may be utilized to analyze genetic data associated with a specimen having two or more biological components to determine the contribution of each component to the genetic data and / or to determine what genetic data is associated with any component of the specimen when the specimen is purified. Examples of genetic data deconvolutors are disclosed, for example, in U.S. Patent Application Nos. 16 / 732,229 and PCT19 / 69191, both filed December 31, 2019, entitled "Transcriptome Deconvolution of Metastatic Tissue Samples," U.S. Provisional Patent Application No. 62 / 924,054, filed October 21, 2019, entitled "Calculating Cell-type RNA Profiles for Diagnosis and Treatment," and U.S. Provisional Patent Application No. 62 / 944,995, filed December 6, 2019, entitled "Rapid Deconvolution of Bulk RNA Transcriptomes for Large Data Sets (Including Transcriptomes of Specimens Having Two or More Tissue Types)," which are incorporated by reference herein in their entireties for all purposes.
[0138] If the digital and laboratory healthcare platform further includes an automated RNA expression caller, the RNA expression levels can be adjusted to be expressed as values relative to a reference expression level, which is often done to prepare multiple RNA expression datasets for analysis, to avoid artifacts that occur when datasets have differences because they were not generated using the same methods, equipment, and / or reagents. Examples of automated RNA expression callers are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 943,712, filed December 4, 2019, entitled "Systems and Methods for Automating RNA Expression Calls in a Cancer Prediction Pipeline," which is incorporated by reference herein in its entirety for all purposes.
[0139] The digital and laboratory healthcare platform may further include one or more insight engines for delivering information, characteristics, or determinations related to a disease state that may be based on genetic and / or clinical data associated with the patient and / or specimen. Exemplary insight engines may include a tumor of unknown origin engine, a human leukocyte antigen (HLA) loss of homozygosity (LOH) engine, a tumor mutation burden engine, a PD-L1 status engine, a homologous recombination deficiency engine, a cell pathway activation report engine, an immune infiltration engine, a microsatellite instability engine, a pathogen infection status engine, and the like. Examples of tumor of unknown origin engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 855,750, filed May 31, 2019, entitled "Systems and Methods for Multi-Label Cancer Classification," which is incorporated by reference herein in its entirety for all purposes. Examples of HLA LOH engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 889,510, filed August 20, 2019, entitled "Detection of Human Leukocyte Antigen Loss of Heterozygosity," which is incorporated by reference herein in its entirety for all purposes. Examples of tumor mutation burden (TMB) engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 804,458, filed February 12, 2019, entitled "Assessment of Tumor Burden Methodologies for Targeted Panel Sequencing," which is incorporated by reference herein in its entirety for all purposes.Examples of PD-L1 status engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 854,400, filed May 30, 2019, entitled "A Pan-Cancer Model to Predict The PD-L1 Status of a Cancer Cell Sample Using RNA Expression Data and Other Patient Data," which is incorporated by reference herein in its entirety for all purposes. Additional examples of PD-L1 status engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 824,039, filed March 26, 2019, entitled "PD-L1 Prediction Using H&E Slide Images," which is incorporated by reference herein in its entirety for all purposes. The systems and methods disclosed herein are an example of a homologous recombination deficiency engine. Alternative homologous recombination deficiency engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 804,730, filed February 12, 2019, entitled "An Integrative Machine-Learning Framework to Predict Homologous Recombination Deficiency," which is incorporated herein by reference in its entirety for all purposes. Examples of cellular pathway activation report engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 888,163, filed August 16, 2019, entitled "Cellular Pathway Report," which is incorporated herein by reference in its entirety for all purposes. Examples of immune infiltration engines are disclosed, for example, in U.S. Patent Application No. 16 / 533,676, filed August 6, 2019, entitled "A Multi-Modal Approach to Predicting Immune Infiltration Based on Integrated RNA Expression and Imaging Features," which is incorporated herein by reference in its entirety for all purposes.Additional examples of immune infiltration engines are disclosed, for example, in U.S. Patent Application No. 62 / 804,509, filed February 12, 2019, entitled "Comprehensive Evaluation of RNA Immune System for the Identification of Patients with an Immunologically Active Tumor Microenvironment," which is incorporated by reference in its entirety for all purposes. Examples of MSI engines are disclosed, for example, in U.S. Patent Application No. 16 / 653,868, filed October 15, 2019, entitled "Microsatellite Instability Determination System and Related Methods," which is incorporated by reference in its entirety for all purposes. Additional examples of MSI engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 931,600, filed November 6, 2019, entitled "Systems and Methods for Detecting Microsatellite Instability of a Cancer Using a Liquid Biopsy," which is incorporated by reference in its entirety and for all purposes.
[0140] When the digital and laboratory healthcare platform further includes a report generation engine, the above-described methods and systems can be utilized to generate a summary report of the patient's genetic profile and the results of one or more insight engines for presentation to a physician. For example, the report can provide the physician with information about how much of the sequenced specimen contained tumor or normal tissue from a first organ, a second organ, a third organ, etc. For example, the report may provide a genetic profile of each of the tissue types, tumors, or organs in the specimen. The genetic profile may represent the gene sequences present in the tissue type, tumor, or organ and can include information about variants, expression levels, gene products, or other information that may be derived from genetic analysis of the tissue, tumor, or organ. The report can include treatments and / or clinical trials that were matched based on some or all of the genetic profile or insight engine results and summary. For example, therapies may be matched according to the systems and methods disclosed in U.S. Provisional Patent Application No. 62 / 804,724, filed February 12, 2019, entitled "Therapeutic Suggestion Improvements Gained Through Genomic Biomarker Matching Plus Clinical History," which is incorporated by reference herein in its entirety for all purposes. For example, clinical trials may be matched according to the systems and methods disclosed in U.S. Provisional Patent Application No. 62 / 855,913, filed May 31, 2019, entitled "Systems and Methods of Clinical Trial Evaluation," which is incorporated by reference herein in its entirety for all purposes.
[0141] The report may include a comparison of the results to a database of results from many specimens. Examples of methods and systems for comparing the results to a database of results are disclosed in U.S. Provisional Patent Application No. 62 / 786,739, filed December 31, 2018, entitled "A Method and Process for Predicting and Analyzing Patient Cohort Response, Progression and Survival," which is incorporated by reference herein in its entirety for all purposes. This information may be used in combination with similar information and / or clinical response information from additional specimens, in some cases, to discover biomarkers or design clinical trials.
[0142] When the digital and laboratory healthcare platform further includes application of one or more embodiments herein to organoids developed in association with the platform, the method and system can be used to further evaluate the genetic sequencing data derived from the organoids to provide information regarding the extent to which the sequenced organoids include a first cell type, a second cell type, a third cell type, etc. For example, the report can provide a genetic profile of each of the cell types in the specimen. The genetic profile can represent the gene sequences present in a given cell type and can include information regarding variants, expression levels, gene products, or other information that can be derived from genetic analysis of the cells. The report can include treatments that are matched based on some or all of the deconvoluted information. These treatments can be tested on the organoid, derivatives of the organoid, and / or similar organoids to determine the sensitivity of the organoids to those treatments. For example, organoids can be cultured and tested according to the systems and methods disclosed in U.S. Patent Application No. 16 / 693,117, filed November 22, 2019, entitled "Tumor Organoid Culture Compositions, Systems, and Methods," U.S. Provisional Patent Application No. 62 / 924,621, filed October 22, 2019, entitled "Systems and Methods for Predicting Therapeutic Sensitivity," and U.S. Provisional Patent Application No. 62 / 944,292, filed December 5, 2019, entitled "Large Scale Phenotypic Organoid Analysis," which are incorporated by reference herein in their entireties for all purposes.
[0143] When the digital and laboratory healthcare platform further includes one or more of the above applications in combination with or as part of a medical device or a laboratory-developed test targeted to medical care and research in general, the results of such laboratory-developed test or medical device can be improved and personalized through the use of artificial intelligence. Examples of laboratory-developed tests, particularly those that may be improved by artificial intelligence, are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 924,515, filed October 22, 2019, entitled "Artificial Intelligence Assisted Precision Medicine Enhancements to Standardized Laboratory Diagnostic Testing," which is incorporated by reference herein in its entirety for all purposes.
[0144] It should be understood that the examples given above are illustrative and not limiting of the use of the systems and methods described herein in combination with digital and laboratory healthcare platforms. EXAMPLES
[0145] Example 1 - Analysis of initial HRD prediction models The accuracy of the initial HRD prediction algorithm was evaluated using a small 40-sample training set curated with samples with known pathogenic mutations in BRCA, as described herein. All genomic features required for HRD prediction were computed on the training samples using CONA. The sklearn “train_test_split” method was used to create training and test sets for initial validation. The sklearn “standardscaler” and “fit_transform” methods were used to normalize the mean and variance of the training samples, and the scale of the future test data was also kept the same. The “RandomForestClassifier” method was used to create a random forest classifier with the number of genomic features set as “n_estimators”. A simple 5-fold cross-validation score metric was calculated using “compute_simple_cross_val_score” to obtain a classification accuracy of 99%. The top k features were obtained using the standard Gini criterion. The classification model was dumped to a file using pickle, and the model was loaded to make predictions for each test sample. For each patient, we first computed HRD features using CONA and standardized the features using the same scaling function used for the training samples. Then, the probability of HRD was obtained given these standardized features using the "model.predict_proba" function implemented in sklearn. The confidence of the HRD prediction is the model predicted probability, with a positive call defined for samples with probability >0.5. Any new features can be easily incorporated into the model, and the training set can be easily extended for retraining and prediction.
[0146] Example 2 - Analysis of initial HRD prediction models The HRD status of 1000 patient samples across 35 different cancer types was analyzed using the HRD classifier as described herein. The analysis identified a total of 6.4% HRD positive calls. Pathogenic variants in BRCA genes were significantly greater in HRD positive calls than negative calls (P<4.1e-219, Mann-Whitney test), but LOH in BRCA was not enriched (P<0.06, Mann-Whitney test). Ovarian cancer (12% HRD positive, n=57), breast cancer (14.6%, n=89), and colorectal cancer (10%, n=285) were some of the most represented cancer types. In contrast to previously published results, most of the pancreatic (2.3%, n=295) and prostate (2.7%, n=37) patients did not have predicted HRD.
[0147] References and Alternative Embodiments All references cited in this specification are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety for all purposes.
[0148] The present invention can be implemented as a computer program product that includes a computer program mechanism embedded in a non-transitory computer-readable storage medium. For example, the computer program product can include the program modules shown in any combination in FIG. 1 and / or described elsewhere in this application. These program modules can be stored on a CD-ROM, DVD, magnetic disk storage product, USB key, or any other non-transitory computer-readable data or program storage product.
[0149] As will be apparent to those skilled in the art, many modifications and variations of the present disclosure can be made without departing from its spirit and scope. The specific embodiments described herein are provided only as examples. The embodiments have been selected and described in order to best explain the principles of the invention and its practical use, so as to enable those skilled in the art to best utilize the invention and various embodiments with various modifications suited to the particular applications contemplated. The present disclosure should be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A method for determining whether a cancer in a test subject has a homologous recombination pathway deficiency (HRD), comprising:
1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, (A) obtaining in electronic form a first plurality of sequence reads of a first DNA sample from the test subject, the first DNA sample comprising DNA molecules from a cancerous tissue of the subject; (B) obtaining in electronic form a second plurality of sequence reads of a second DNA sample from the test subject, the second DNA sample consisting of DNA molecules from a non-cancerous tissue of the test subject; and (C) generating a genomic data construct for the test subject based on the first plurality of sequence reads and the second plurality of sequence reads, wherein the genomic data construct includes: (i) a heterozygosity state of a first plurality of DNA damage repair genes in the genome of the cancerous tissue of the test subject; (ii) a measure of loss of heterozygosity across the genome of the cancerous tissue of the test subject; (iii) a measure of mutant alleles detected in a second plurality of DNA damage repair genes in the genome of the cancerous tissue of the test subject; and (iv) a measure of mutant alleles detected in the second plurality of DNA damage repair genes in the genome of the non-cancerous tissue of the subject; (D) inputting the genomic data construct into a classifier trained to distinguish between cancers with homologous recombination pathway deficiencies and cancers without homologous recombination pathway deficiencies, thereby determining whether the test subject's cancer has HRD.
2. 2. The method of claim 1, wherein the first DNA sample is from a solid tumor biopsy of the cancerous tissue of the test subject.
3. 2. The method of claim 1, wherein the second DNA sample is from a buffy coat preparation of a blood sample from the test subject.
4. the first plurality of sequence reads are generated by (i) targeted sequencing using a plurality of nucleic acid probes to enrich nucleic acids from the cancerous tissue of the subject for a panel of genomic regions, or (ii) by whole genome sequencing of nucleic acids from the cancerous tissue of the test subject; 2. The method of claim 1, wherein the second plurality of sequence reads were generated by (i) targeted sequencing using a plurality of nucleic acid probes to enrich nucleic acid from the non-cancerous tissue of the test subject for a panel of genomic regions, or (ii) by whole genome sequencing of nucleic acid from the non-cancerous tissue of the test subject.
5. The measure of loss of heterozygosity across the genome of the cancerous tissue of the test subject is determining loss of genomic heterozygosity in the first plurality of sequence reads; and 2. The method of claim 1, wherein the determined loss of heterozygosity is determined by normalizing the determined loss of heterozygosity by an estimate of tumor purity for the first plurality of sequence reads.
6. 2. The method of claim 1, wherein the heterozygosity status of the first plurality of DNA damage repair genes comprises a count of the number of unique frameshift mutations detected in the first plurality of DNA damage repair genes.
7. 2. The method of claim 1, wherein the heterozygosity status of the first plurality of DNA damage repair genes comprises a count of the number of unique truncating mutations detected in the first plurality of DNA damage repair genes.
8. 2. The method of claim 1, wherein the first plurality of DNA damage repair genes comprises BRCA1 and BRCA2.
9. 2. The method of claim 1, wherein the measure of mutant alleles detected in the second plurality of DNA damage repair genes in the genome of the cancerous tissue of the test subject comprises a count of the number of unique mutations associated with loss of homologous recombination detected in the first plurality of sequence reads.
10. 2. The method of claim 1, wherein the measure of mutant alleles detected in the second plurality of DNA damage repair genes in the genome of the non-cancerous tissue of the test subject comprises a count of the number of unique mutations associated with loss of homologous recombination detected in the second plurality of sequence reads.
11. 2. The method of claim 1, wherein the second plurality of DNA damage repair genes comprises BRCA1 and BRCA2.
12. 12. The method of claim 11, wherein the unique mutations associated with loss of homologous recombination in BRCA1 and BRCA2 include at least 25 of the mutations listed in Table 1.
13. 12. The method of claim 11, wherein the unique mutations associated with loss of homologous recombination in BRCA1 and BRCA2 include the mutations listed in Table 1.
14. The method further comprising: recommending treating the cancer by administering a poly ADP ribose polymerase (PARP) inhibitor to the test subject when the cancer of the test subject is determined to be homologous recombination deficient; The method of claim 1, further comprising, when the cancer of the test subject is determined to not be homologous recombination deficient, recommending treating the cancer with a therapy that does not include administering a PARP inhibitor to the test subject.
15. 15. The method of claim 14, wherein the PARP inhibitor is selected from the group consisting of olaparib, veliparib, rucaparib, niraparib, and talazoparib.
16. 10. The method of claim 1, wherein the cancer is breast cancer, ovarian cancer, or colorectal cancer.
17. 2. The method of claim 1, wherein the classifier is a neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.
18. The method of claim 1 , wherein the classifier is a random forest algorithm.
19. the first plurality of sequence reads are generated by exome sequencing of cDNA molecules generated from the cancerous tissue of the test subject; 2. The method of claim 1, wherein the second plurality of sequence reads are generated by exome sequencing of cDNA molecules generated from the non-cancerous tissue of the test subject.
20. 1. A computer system comprising: one or more processors; and a non-transitory computer readable medium comprising computer executable instructions that, when executed by the one or more processors, cause the processors to perform a method for determining whether a cancer in a test subject has a homologous recombination pathway deficiency, the method comprising: (A) obtaining in electronic form a first plurality of sequence reads of a first DNA sample from the test subject, the first DNA sample comprising DNA molecules from a cancerous tissue of the subject; (B) obtaining in electronic form a second plurality of sequence reads of a second DNA sample from the test subject, the second DNA sample consisting of DNA molecules from a non-cancerous tissue of the test subject; and (C) generating a genomic data construct for the test subject based on the first plurality of sequence reads and the second plurality of sequence reads, wherein the genomic data construct includes: (i) a heterozygosity state of a first plurality of DNA damage repair genes in the genome of the cancerous tissue of the test subject; (ii) a measure of loss of heterozygosity across the genome of the cancerous tissue of the test subject; (iii) a measure of mutant alleles detected in a second plurality of DNA damage repair genes in the genome of the cancerous tissue of the test subject; and (iv) a measure of mutant alleles detected in the second plurality of DNA damage repair genes in the genome of the non-cancerous tissue of the subject; (D) inputting the genomic data construct into a classifier trained to distinguish between cancers with homologous recombination pathway deficiencies and cancers without homologous recombination pathway deficiencies, thereby determining whether the test cancer has HRD.
21. 1. A non-transitory computer readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method for determining whether a cancer in a test subject has a homologous recombination pathway deficiency, the method comprising: (A) obtaining in electronic form a first plurality of sequence reads of a first DNA sample from the test subject, the first DNA sample comprising DNA molecules from a cancerous tissue of the subject; (B) obtaining in electronic form a second plurality of sequence reads of a second DNA sample from the test subject, the second DNA sample consisting of DNA molecules from a non-cancerous tissue of the test subject; and (C) generating a genomic data construct for the test subject based on the first plurality of sequence reads and the second plurality of sequence reads, wherein the genomic data construct includes: (i) a heterozygosity state of a first plurality of DNA damage repair genes in the genome of the cancerous tissue of the test subject; (ii) a measure of loss of heterozygosity across the genome of the cancerous tissue of the test subject; (iii) a measure of mutant alleles detected in a second plurality of DNA damage repair genes in the genome of the cancerous tissue of the test subject; and (iv) a measure of mutant alleles detected in the second plurality of DNA damage repair genes in the genome of the non-cancerous tissue of the subject; (D) inputting the genomic data construct into a classifier trained to distinguish between cancers with homologous recombination pathway deficiencies and cancers without homologous recombination pathway deficiencies, thereby determining whether the test subject's cancer has HRD.
Citation Information
Patent Citations
A method for evaluating HRD score based on low-depth WGS
CN113257346B
Analysis system of t cell receptor repertoire and b cell receptor repertoire, and utilization of the same for treatment and diagnosis
JP2017212988A
Methods for detecting mutational signatures in samples
JP2019519872A
Methods and materials for assessing homologous recombination deficiency
US20170283879A1