Determining relative timing of mutation and amplification
By determining the timing of gene mutations relative to amplification, the techniques enhance predictive accuracy of cancer treatment effectiveness, preventing ineffective treatments and reducing patient harm.
Patent Information
- Application Number
- PCT/US2025/035547
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-28
- Filing Date
- 2025-06-26
- Publication Date
- 2026-01-02
AI Technical Summary
Existing cancer treatment methods often fail to account for the timing of gene mutations relative to genomic amplification, leading to ineffective treatments and unnecessary harm to patients.
Techniques are developed to determine whether a mutation in a gene occurred before or after amplification, enhancing predictive accuracy of treatment effectiveness by identifying patient sub-populations with differential cancer subtypes.
This approach enables the identification of patient sub-populations that may have unrecognized cancer subtypes, preventing administration of ineffective treatments and reducing harmful side-effects.
Smart Images

Figure US2025035547_02012026_PF_FP_ABST
Abstract
Description
DETERMINING RELATIVE TIMING OF MUTATION AND AMPLIFICATIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional App. No. 63 / 665,829, which was filed on June 28, 2024 and is incorporated by reference herein in its entirety.BACKGROUND
[0002] Copy number variants are sequences that are repeated in the genome of a subject where the number and type of repeated sequences vary across different individuals of the same species. In some cases, the number of copies of a given sequence in a subject's genome can increase in somatic cells during a process referred to as "amplification.” Amplification of a gene in a subject's genome can impact how much the gene is expressed, in some cases.
[0003] Copy number variants, as well as amplification variants, are common in cancer cells. These variants can have a significant impact on cancer type, progression, and treatment options. For example, cancer cells having particular variants (e.g., amplification variants) in an androgen receptor (AR) gene may be resistant to some treatment options, and susceptible to other treatment options. If the AR gene contains one or more point mutations that impact AR gene expression, and the mutated AR gene is amplified in the genome of the cancer cells and overexpressed due to the amplification, then the copy number variant, as well as the one or more point mutations, may be particularly relevant for selecting potential treatments.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:
[0005] FIG. 1 illustrates an example environment for determining the timing of genomic amplification of a mutation.
[0006] FIG. 2 illustrates an example report summarizing predicted categories of a cancer of a subject.
[0007] FIG. 3 illustrates a process for determining whether a mutation occurred prior to genomic amplification.
[0008] FIG. 4 illustrates an example environment for sequencing various nucleic acid molecules.
[0009] FIG. 5 illustrates one or more devices configured to perform various operations described herein.
[0010] FIG. 6 illustrates an experimentally derived example of ratios of expected and observed AR-related variant allele frequencies for various patients with prostate cancer.
[0011] FIGS. 7A to 7D show results of an experimental example related to implementations of the present disclosure.
[0012] Some of the materials submitted herewith may be better understood in color. Applicant considers the color versions of the disclosure as part of the original submission and reserves the right to present color images of the materials of this disclosure in later proceedings.DETAILED DESCRIPTION
[0013] Various implementations of the present disclosure relate to techniques for inferring whether a mutation in an amplified gene occurred prior to, or after, the gene was amplified. For example, the genome of a cancer cell may include multiple copies of a particular gene-of-interest. Further, the gene-of-interest may include one or more additionalvariants (e.g., mutations), which may have clinical significance. Various techniques described herein may be utilized to determine whether the variant(s) in the gene-of-interest were present before the gene was amplified. The relative timing between mutations and amplification of genomic sequences expressed by cancer cells can have a significant impact on the prognosis, diagnosis, and treatment options for individuals with cancer.
[0014] Implementations of the present disclosure provide significant improvements to the technical field of cancer diagnosis and treatment. A plethora of treatment options, such as chemotherapies, radiotherapy, immunotherapies, and surgeries, are available for individuals with various types of cancers. However, many of these options have significant drawbacks. Some treatments can result in significant discomfort and harm to patients. Further, some treatments are associated with a significant expense. For treatments that can result in a significant improvement in a patient's disease, these drawbacks may be acceptable to patients. However, treatments that fail to improve a patient's condition cause unnecessary harm.
[0015] In various implementations of the present disclosure, determining the relative timing between a mutation in a gene and amplification of that gene can greatly enhance the predictive accuracy that a treatment targeting expression of that gene will be effective. For instance, some previous techniques may infer that a particular treatment option would be effective on all patients whose cancer cells have a particular point mutation in a particular gene. However, if the point mutation was initiated after amplification, then the cancer cells may not respond to the particular treatment option. Therefore, a patient may be prevented from being administered the particular treatment option, with its potential harmful side-effects, if a care provider is able to recognize that the point mutation arose after amplification. By determining whether mutations occurred before or after genomic amplification, implementations of the present disclosure enable the identification of patient sub-populations that may have differential cancer subtypes that were previously unrecognized.
[0016] Various analyses described herein cannot be performed in the human mind, or by pen and paper. For example, techniques described herein include performing various physical and chemical manipulations of physical nucleic acid molecules that cannot be performed mentally. Furthermore, some implementations of the present disclosure include determining a copy number based on numerous sequence reads that are obtained by sequencing nucleic acid molecules in a sample. The complexity of the process of determining copy number, as well as various allele fractions described herein, prevents various implementations of the present disclosure from being calculated mentally.Example Definitions
[0017] As used herein, the terms "deoxyribonucleic acid,” "DNA,” "DNA molecule,” and their equivalents, may refer to a polymer of nucleotides (also referred to as "nucleobases”) containing deoxyribose. The nucleotides in DNA include cytosine (C), guanine (G), adenine (A), and thymine (T). Each DNA nucleotide includes a deoxyribose and a phosphate group. An example single-stranded DNA (ssDNA) molecule includes a chain of covalently bonded DNA nucleotides. In the example ssDNA molecule, the phosphate group of the mth nucleotide is covalently bonded to the deoxyribose of the (m-1)th nucleotide, wherein m is a positive integer greater than 2 and less than or equal to the number of DNA nucleotides in the chain. In various examples, DNA is double-stranded and includes two ssDNA molecules that are complementary to one another and coiled around each other in a double helix form. The nucleotides of one ssDNA molecule are hydrogen bonded to the nucleotides of the other ssDNA molecule. In particular, the pyrimidines (A and T) hydrogen bond to each other, and the purines (C and G) hydrogen bond to each other.
[0018] As used herein, the terms "ribonucleic acid,” "RNA,” "RNA molecule,” and their equivalents, may refer to a polymer of nucleotides containing ribose. The nucleotides in RNA include cytosine (C), guanine (G), adenine (A), and uracil (U). Each RNA nucleotide includes a ribose and a phosphate group. In an example RNA molecule, the phosphate group of the nth nucleotide is covalently bonded to the ribose of the (n-1)th nucleotide, wherein n is a positive integer greater than 2 and less than or equal to the number of RNA nucleotides in the chain. Messenger RNA (mRNA) is a type of RNA molecule that is synthesized (or "transcribed”) by RNA polymerase (an enzyme) to be complementary to a gene encoded in a DNA sequence, and is also used by a ribosome to synthesize a polypeptide or protein. An mRNA is therefore an example of a "coding RNA.” In various cases, intron sequences are removed from an mRNA via a process known as "RNA splicing.” MicroRNA ("miRNA”) are single-stranded RNA molecules that perform post-transcriptional gene expression regulation. For instance, a miRNA may bind to a complementary mRNA molecule, thereby cleaving, destabilizing, or otherwise preventing the mRNA molecule from being translated into a polypeptide or protein by a ribosome. In various examples, a miRNA has a length in a range of 21 to 23 RNA nucleotides. As used herein, the terms "non-coding RNA” may refer to a type of RNA that is not translated into a protein. Examples of non-coding RNA include miRNA, transfer RNA (tRNA), and ribosomal RNA (rRNA). The term "functional RNA,” and its equivalents, may refer to any RNA molecule that impacts a biological process. For instance, functional RNA may include mRNA, miRNA, tRNA, rRNA, and the like.
[0019] As used herein, the term "base,” and its equivalents, may refer to a monomer of a polymer. For example, a base of DNA or RNA is a nucleotide.
[0020] As used herein, the term "base pair,” and its equivalents, may refer to a pair of complementary DNA nucleotides, which are hydrogen-bonded to one another in a double-stranded DNA molecule. For example, a base pair includes a first base in a first ssDNA and a second base in a second ssDNA, wherein the first and second bases are complementary and hydrogen-bonded to one another.
[0021] As used herein, the terms "nucleotide,” "nucleobase,” "nucleic acid,” "nucleic acid molecule,” and their equivalents, may refer to an organic molecule that includes a nitrogenous base, a sugar, and a phosphate group. In various cases, a nucleotide is a monomer of DNA or RNA. A nucleotide, for instance, is a chemical structure.
[0022] As used herein, the terms "3' end,” "3-prime end,” and their equivalents, may refer to a terminus of a singlestranded nucleotide polymer that includes a base whose third carbon in its deoxyribose or ribose is bound to a hydroxyl group while being unbound to another base.
[0023] As used herein, the terms "5' end,” "5-prime end,” and their equivalents, may refer to a terminus of a singlestranded nucleotide polymer that includes a base whose fifth carbon in its deoxyribose or ribose ring is unbound to another base. In some cases, the fifth carbon is bound to a phosphate group.
[0024] As used herein, the "length” of a polymer refers to a number of covalently bonded monomers that are included in the polymer. For instance, the length of a DNA molecule may be the number of covalently bonded nucleotides in at least one strand of the DNA molecule and / or the number of base pairs in the DNA molecule. In various examples, the length of an RNA molecule may be the number of covalently bonded nucleotides in the RNA molecule.
[0025] As used herein, the term "gene,” and its equivalents, refers to a sequence of DNA nucleotides that is transcribed into a functional RNA. The functional RNA, for instance, is RNA that is translated into a polypeptide or protein (e.g., mRNA) or that has some other biological function (e.g., miRNA, tRNA, etc.). A gene is "expressed” when it is used asa template to generate a functional RNA. A subject, for instance, has numerous genes contained in the subject's genome. A gene may include both introns and exons. As used herein, the term "intron,” and its equivalents, may refer to a subset of DNA nucleotides in a gene that is not used to code for any functional RNA that is expressed by the organism. As used herein, the term "exon,” and its equivalents, may refer to a subset of DNA nucleotides in a gene that is used to code for a functional RNA. For instance, an exon may encode a polypeptide or protein that is expressed by the organism. In various examples, a gene can be represented in data (e.g., as data representative of the sequence of DNA nucleotides in the gene) or as a chemical structure (e.g., as the sequence of DNA nucleotides itself).
[0026] As used herein, the term "genome,” and its equivalents, refers to the aggregate of genes of a subject. In various cases, a genome represents the sequences of several linear DNA molecules that are present in a subject's chromosomes. A "reference genome” refers to an aggregation of genes of one or more reference subjects. In various cases, a genome is represented in data.
[0027] As used herein, the terms "pangenome,” "pan-genome,” "supragenome,” and their equivalents, refers to an aggregate set of genes from multiple subgroups (e.g., strains) within a population (e.g., a clade) of subjects. A pangenome, for example, indicates genes that are present in all subjects within the population, as well as genes that are present in some of the subjects of the population. A pangenome is represented in data, for instance.
[0028] As used herein, the term "transcriptome,” and its equivalents, refers to the aggregate of RNA sequences of a subject. In some cases, a transcriptome is limited to mRNA sequences. In various examples, a transcriptome is represented in data.
[0029] As used herein, the term "genomic DNA,” "gDNA,” "chromosomal DNA,” and their equivalents, may refer to DNA molecules that are obtained from a chromosome and / or nucleus of a cell.
[0030] As used herein, the terms "DNA fragment,” "fragment,” and their equivalents, may refer to DNA molecules that are excised and / or broken off from a larger DNA molecule.
[0031] As used herein, the terms "cell-free DNA,” "cfDNA,” and their equivalents, may refer to DNA fragments that are non-encapsulated and obtained outside of cells within a sample (e.g., a liquid biopsy sample).
[0032] As used herein, the terms "circulating tumor DNA,” "ctDNA,” and their equivalents, may refer to a cfDNA molecule that originates from a cancer cell.
[0033] As used herein, the term "promoter,” and its equivalents, may refer to a portion of a DNA molecule that binds one or more proteins in order to initiate transcription of a gene. For example, the promotor is located "upstream” of the gene. For example, the promotor is located between the 5' end of the DNA molecule and the gene. A promotor may include one or more binding sites for RNA polymerase, and / or one or more transcription factor binding sites. In some examples, a promotor includes one or more CpG islands. A promoter, for instance, includes a transcription start site.
[0034] As used herein, the terms "CpG island,” "CGI,” "CpG site,” and their equivalents, may refer to a continuous portion of a DNA molecule whose sequence includes greater than a threshold amount (e.g., greater than 50%) of G-C base pairs.
[0035] As used herein, the term "enhancer,” and its equivalents, may refer to a portion of a DNA molecule that binds one or more proteins in order to increase the chance that a gene will be transcribed. For instance, an enhancer includes one or more transcription factor binding sites. In various cases, an enhancer includes one or more CpG islands.
[0036] As used herein, the term "cancer,” and its equivalents, may refer to a condition of a subject in which particular cells (referred to as "cancer cells”) divide uncontrollably in the subject's body. In some cases, a cancer is characterized by a location or tissue type from which the cancer cells originated. In some examples, a cancer is characterized by a location or tissue type in which the cancer cells are located.
[0037] As used herein, the terms "tumor,” "neoplasm,” and their equivalents, may refer to a mass of tissue including cancer cells.
[0038] As used herein, the terms "tissue of origin,” "tissue origin,” and their equivalents, refers to a differentiated type of tissue from which cancer cells in the body of a subject began dividing uncontrollably in the subject's body.
[0039] As used herein, the terms "liquid biopsy,” "fluid biopsy,” and their equivalents, may refer to a process of obtaining a fluid sample from a subject's body. The sample, for instance, can be referred to as a "liquid biopsy sample.” Examples of fluids that are sampled from the body include blood, plasma, cerebrospinal fluid, sputum, stool, urine, lymphatic fluid, and saliva.
[0040] As used herein, the term "tissue biopsy,” and its equivalents, may refer to a process of obtaining a sample of cells from a subject's body. A tissue biopsy, in various cases, is performed by cutting a mass of cells from the subject's body. For instance, a tissue biopsy is a procedure performed by a surgeon, interventional radiologist, interventional cardiologist, or other specialized clinician. The term "tissue” or "tissue biopsy sample” can be used to refer to the sample of cells obtained using a tissue biopsy.
[0041] As used herein, the term "subject,” and its equivalents, may refer to a human or non-human animal. A subject that is receiving care from at least one care provider may be referred to as a "patient.”
[0042] As used herein, the terms "machine learning,” "ML,” "computer learning,” "artificial intelligence,” and their equivalents, may refer to the use of a computing devices to learn patterns in training data. The process of learning these patterns may be referred to as "training.” In particular cases, one or more computing devices may perform machine learning by executing a machine learning model. As used herein, the terms "machine learning model,” "ML model,” and their equivalents, may refer to data encoding instructions that, when executed by at least one computing device, causes the at least one computing device to learn patterns in training data by optimizing one or more metrics, values, or other types of parameters. After training, an ML model, when executed by at least one computing device, causes the at least one computing device to utilize the optimized parameters in order to perform one or more tasks.
[0043] As used herein, the term "variant,” and its equivalents, may refer to a difference between a subject genetic sequence and a reference sequence. For instance, a variant may correspond to a difference between one or more nucleotides in a genome of a subject and one or more corresponding nucleotides in at least one reference genome or pangenome. A variant may be characterized by its identity (e.g., what nucleotides are different), its position (e.g., where are the nucleotides located in the genome, what chromosome contains the nucleotides, what gene contains the nucleotides, etc.), its length (e.g., how many nucleotides are different from the reference sequence), its type (e.g., substitution, insertion, deletion, copy number alternation, rearrangement of fusion, etc.), and other features that indicates its significance and / or relevance. In some cases, a variant represents any apparent alteration in a sequence that has been read from a nucleic acid molecule with respect to the reference sequence, such as reads cleaved by restriction enzymes (RE). In various examples, a variant can be represented in data (e.g., by data characterizing the variant) or as a chemical structure (e.g., the nucleotides themselves). As used herein, the term "mutation,” and its equivalents, may refer to a change in a gene.
[0044] As used herein, the term "allele,” and its equivalents, may refer to a sequence of one or more nucleotides at a locus and / or gene. In many cases, an individual inherits one or two alleles for each gene, from each parent. In the case of some sex-linked genes, an individual has a single allele. The most common allele in a sample may be referred to as a "major allele.” Alleles present in a sample that are not a major allele can be referred to as a "minor allele.” In some cases, the second-most-prevalent allele in a sample is referred to as a "minor allele.”
[0045] As used herein, the term "sex-linked,” and its equivalents, may refer to a characteristic (e.g., a gene) present on an X or Y chromosome. For instance, an “X-linked” gene is present on an X chromosome, and a "Y-linked” gene is present on a Y chromosome.
[0046] As used herein, the terms "allele fraction,” "allele frequency,” and their equivalents, may refer to a relative prevalence of nucleic acid molecules (e.g., DNA) in a sample that exhibit one or more alleles-of-interest. For instance, an allele fraction for a sample containing a plurality of nucleic acid molecules can be calculated by identifying sequence reads representing the nucleic acid molecules, aligning the sequence reads in a particular region (e.g., gene) of interest, determining a number of copies of a particular allele (e.g., among multiple alleles) in the sequence reads, and calculating the allele fraction by dividing the number of copies of the particular allele by the total number of all alleles in the sequence reads.
[0047] As used herein, the term "major allele frequency,” and its equivalents, may refer to an allele frequency of a major allele.
[0048] As used herein, the term "minor allele frequency,” and its equivalents, may refer to an allele frequency of one or more minor alleles.
[0049] As used herein, the term "substitution,” and its equivalents, can refer to a nucleotide in a subject sequence that is different than an equivalent nucleotide (e.g., a nucleotide at the same position) in a reference sequence.
[0050] As used herein, the term "insertion,” and its equivalents, can refer to a nucleotide in a subject sequence that is added with respect to a reference sequence.
[0051] As used herein, the term "deletion,” and its equivalents, can refer to the removal of a nucleotide from a nucleotide sequence.
[0052] As used herein, the terms "copy number alternation,” "CAN,” "copy number variation,” "CNV,” and their equivalents, can refer to a portion of a reference sequence that is repeated.
[0053] As used herein, the term "copy number,” and its equivalents, can refer to a number of copies of a sequence, allele, or gene present in a genome.
[0054] As used herein, the terms "rearrangement of fusion,” "fusion rearrangement,” "translocation,” and their equivalents, can refer to a change in the relative position of one or more portions of a reference sequence, thereby generating a gene that was not present in the reference sequence.
[0055] As used herein, the terms "gene amplification,” "amplification,” "genomic amplification,” and their equivalents, may refer to an increase in a number of copies of a gene or allele in a somatic cell. In various cases, amplification refers to a copy number increase of a region in a genome, such as a region of a chromosome arm. Gene amplification, in some cases, occurs in cancer cells.
[0056] As used herein, the term "sequencing,” and its equivalents, may refer to a process of identifying the order and identity of monomers in a polymer chain, such as the order and identity of nucleotides in a DNA or RNA molecule. Theterms "whole genome sequencing,” "WGS,” and their equivalents, may refer to the process of sequencing an entire genome of a subject, including the introns and exons of the genes of the subject. The terms "whole exome sequencing,” "WES,” and their equivalents, may refer to the process of sequencing all exomes of a subject. The term "targeted sequencing,” and its equivalents, may refer to the process of sequencing a portion of the genome of a subject, such as sequencing a single gene of the subject. Various techniques can be utilized to sequence a DNA or RNA molecule, such as massively parallel sequencing (MPS), nanopore sequencing, direct sequencing, Sanger sequencing, or next-generation sequencing. In various cases, sequencing is performed on physical molecules (e.g., RNA or DNA) and is used to generate data.
[0057] As used herein, the terms "massive parallel sequencing,” "massively parallel sequencing,” "MPS,” and their equivalents, may refer to a technique for simultaneously performing multiple reactions that can be used to identify the order and identity of monomers in multiple polymer chains. In particular cases, massive parallel sequencing can be performed using sequencing-by-synthesis on clonally amplified DNA molecules that are located in spatially separated regions, which are individually monitored by sensors.
[0058] As used herein, the term "nanopore sequencing,” and its equivalents, may refer to a technique for identifying the order and identity of monomers in a polymer chain by transporting the polymer chain from a first space to a second space, wherein the first space and the second space are separated by a substrate, by directing the polymer chain through a small hole (known as a "nanopore”) embedded in the substrate, and monitoring a relative electrical signal (e.g., a voltage or current) between the first space and the second space.
[0059] As used herein, the term "sensor,” and its equivalents, may refer to a physical device or other apparatus that is configured to detect one or more detection signals.
[0060] As used herein, the term "detection signal,” and its equivalents, may refer to a physical signal that can be identified, characterized, or otherwise perceived by a sensor.
[0061] As used herein, the term "sequence read data,” and its equivalents, may refer to data that is indicative of an order and identity of monomers in a polymer, such as the order and identity of nucleotides in a DNA or RNA sequence. In various implementations, sequence read data is generated via a sequencing operation.
[0062] As used herein, the term "image,” and its equivalents, may refer to 2D or 3D array of data indicative of an array of pixels or voxels.
[0063] As used herein, the term "ligating,” and its equivalents, may refer to a process of joining two molecules together, for example, with a chemical bond.
[0064] As used herein, the term "adapter,” and its equivalents, may refer to an oligonucleotide that can be ligated to a target nucleic acid molecule. In various cases, an adapter prepares the target nucleic acid molecule for sequencing.
[0065] As used herein, the term "bait molecule,” and its equivalents, may refer to a nucleic acid molecule having a region that is complementary to a region of a target molecule (e.g., cfDNA). A bait molecule includes, for instance, a nucleic acid molecule that can hybridize to ( / .e., is complementary to) a target molecule can be used to capture the target molecule. In some instances, the bait molecule is a capture oligonucleotide (or capture probe). In some instances, the bait molecule is suitable for solution phase hybridization to the target molecule. In some instances, the bait molecule is suitable for solid phase hybridization to the target molecule. In some instances, the bait molecule is suitable for both solution-phase and solid-phase hybridization to the target molecule. The design and construction of bait molecules is described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941 .
[0066] As used herein, the term "amplifying,” and its equivalents, may refer to a process of generating copies of a target molecule, such as a nucleic acid molecule.
[0067] As used herein, the term "hybridization,” and its equivalents, may refer to a process by which to complementary single-stranded nucleic acid molecules bind to one another, thereby forming a double-stranded nucleic acid molecule. In certain examples, the double-stranded nature of the nucleic acid molecule is maintained under stringent hybridization conditions. Exemplary stringent hybridization conditions include an overnight incubation at 42 °C in a solution including 50% formamide, 5XSSC (750 mM NaCI, 75 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5XDenhardt's solution, 10% dextran sulfate, and 20 pig / ml denatured, sheared salmon sperm DNA, followed by washing the filters in 0.1XSSC at 50 °C.
[0068] As used herein, the term "complementary,” and its equivalents, may refer to a state of two single-stranded nucleic acid molecules with respective sequences that cause the nucleic acid molecules to spontaneously hybridize to one another. One nucleic acid molecule, for instance, may have a sequence that causes each nucleic acid to hydrogen bond to a respective nucleic acid in the other nucleic acid molecule.
[0069] As used herein, the term "coverage,” "depth,” and their equivalents, may refer to a number of reads in sequence read data that indicate one or more nucleotides of interest (e.g., at a given genomic position).
[0070] As used herein, the term "minor allele coverage ratio,” and its equivalents, may refer to a haploid coverage ratio that is proportional to a minor allele frequency and a total coverage ratio of a sample. In some cases, the minor allele coverage ratio is equal to product of the minor allele frequency, the total coverage ratio, and a scaling factor (e.g., 2).
[0071] As used herein, the term "major allele coverage ratio,” and its equivalents, may refer to a haploid coverage ratio that is proportional to a major allele frequency and a total coverage ratio of a sample. In some cases, the major allele coverage ratio is equal to product of the major allele frequency, the total coverage ratio, and a scaling factor (e.g., 2).
[0072] As used herein, the terms "therapy,” "treatment,” and their equivalents, may refer to a composition or process that can be used to remediate a health problem. Cancer therapies, for instance, include surgery, radiotherapy, chemotherapy, immunotherapy, cell-based therapies, and the like. Examples of cancer therapies include abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Haris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak),denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane 1131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), Lenvatinib mesylate (Lenvima), letrozole (Femara), lisocabtagene maraleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu 177-dotatate (Lutathera), margetuximabcmkb (Margenza), midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasudotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib camsylate (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecanhziy (Trodelvy), seliciclib, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-tebn (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), and combinations thereof. Examples of cancer therapies also include targeted antibody-based therapies (antibody-drug conjugates, antibody-radioisotope conjugates, and targeted immune cell therapies (e.g., immune effector cells genetically modified to express a chimeric antigen receptor (CAR).
[0073] As used herein, the term "treatment-responsive,” and its equivalents, may refer to a type of cancer cells that can be substantially killed using a predetermined type of therapy. For example, cancer cells of a subject may be responsive to a particular treatment if, after the subject is administered the treatment, the cancer cells are diminished by a particular progression level (e.g., radiographic progression level, marker-based progression level, such as prostate-specific antigen (PSA) progression, etc.). Accordingly, the responsiveness of the cells to the type of therapy may indicate the effectiveness of that therapy.
[0074] As used herein, the term "treatment-resistant,” and its equivalents, may refer to a type of cancer that cannot be substantially killed using a predetermined type of therapy.
[0075] As used herein, the term "metastasis profile,” and its equivalents, may refer to a propensity of a type of cancer to metastasize into one or more differentiated tumor types besides the cancer's tissue origin. In some implementations, the metastasis profile can further indicate the type of tissue in which the cancer can or is likely to metastasize.
[0076] As used herein, the term "clinical trial,” and its equivalents, may refer to a research study used to evaluate a hypothesis based on participation by one or more subjects. In various examples, a clinical trial can be used to assess the efficacy and / or safety of a proposed therapy. A clinical trial may be performed in furtherance of approval of a treatment by a regulatory authority (e.g., the United States Food & Drug Administration (FDA)).Description of Example Implementations
[0077] Various implementations of the present disclosure will now be described with reference to the accompanying Figures.
[0078] FIG. 1 illustrates an example environment 100 for determining the timing of amplification of a mutation. In various cases, a subject 102 presents with one or more symptoms of a condition. The condition, for instance, is a pathological condition.
[0079] The subject 102, for instance, may present to the clinical environment with a lesion 104. In various cases, the lesion 104 may be a tumor that includes cancer cells. According to various examples, the subject 102 has one or more types of cancer, such as adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, a neuroblastoma, non-Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, a teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, a vascular tumor, or combinations or metastases thereof.
[0080] In some embodiments, the subject 102 has a B cell cancer (multiple myeloma), a melanoma, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, cancer of an oral cavity, cancer of a pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel cancer, appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, a cancer of hematological tissue, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma,Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancer, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, or a carcinoid tumor.
[0081] In some embodiments, the subject 102 has acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B-cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with an IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpressed / amplified), breast cancer (HER2+), breast cancer (HR+, HER2- ), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin-associated periodic syndrome, a cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, a diffuse large B-cell lymphoma, fallopian tube cancer, a follicular B-cell non-Hodgkin lymphoma, a follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, a gastrointestinal stromal tumor, a gastrointestinal stromal tumor (KIT+), a giant cell tumor of the bone, a glioblastoma, granulomatosis with polyangiitis, a head and neck squamous cell carcinoma, a hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, a mantle cell lymphoma, medullary thyroid cancer, melanoma, a melanoma with a BRAF V600 mutation, a melanoma with a BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman's disease, multiple hematologic malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, a non-Hodgkin's lymphoma, a nonresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, a non-small cell lung cancer, a non-small cell lung cancer (ALK+), a non-small cell lung cancer (PD-L1 +), a non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), a non-small cell lung cancer (with BRAF V600E mutation), a non-small cell lung cancer (with an EGFR exon 19 deletion or exon 21 substitution (L858R) mutations), a non-small cell lung cancer (with an EGFR T790M mutation), a non-small cell lung cancer KRAS (+ / - G12C), a non-small cell lung cancer TMB-H, a non-small cell lung cancer MET exon 14 skipping, a non-small cell lung cancer ERBB2 inframe indel, a non-small cell lung cancer EGFR exon 20 indel, a neurotrophic tyrosine receptor kinase (NTRK)-positive cancer, ovarian cancer, ovarian cancer (with a BRCA mutation), pancreatic cancer, a pancreatic, gastrointestinal, or lung origin neuroendocrine tumor, a pediatric neuroblastoma, a peripheral T- cell lymphoma, peritoneal cancer, prostate cancer, a renal cell carcinoma, a small lymphocytic lymphoma, a soft tissue sarcoma, a solid tumor (MSI-H / dMMR), a squamous cell cancer of the head and neck, a squamous non-small cell lung cancer, thyroid cancer, a thyroid carcinoma, urothelial cancer, a urothelial carcinoma, or Waldenstrom's macroglobulinemia.
[0082] In particular examples, the subject 102 has prostate cancer. In some cases, the subject 102 has metastatic hormone-sensitive prostate cancer (mHSPC). In some examples, the subject 102 has metastatic castration-resistant prostate cancer (mCRPC). For instance, the subject 102 may have previously received one or more treatments for the prostate cancer, such as androgen receptor (AR) signaling inhibitors (ASIs) and / or taxanes. These previous treatments may have already been administered in one or more rounds.
[0083] In various cases, a care provider 106 (also referred to as a "healthcare provider”) is responsible for diagnosing and / or treating the subject 102. According to some implementations, the lesion 104 may be initially identified using a noninvasive technique. For example, the lesion 104 may be visualized using an imaging modality, such as ultrasound, x-ray, computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), single photon emission CT (SPECT), or any combination thereof. Using the noninvasive technique, the care provider 106 may identify the presence of the lesion 104 but may be unable to determine whether the lesion 104 is a cancerous tumor using noninvasive diagnostic methodologies.
[0084] In various implementations, the care provider 106 is unable to accurately identify a condition of the subject 102 based solely on noninvasive diagnostic techniques. In various cases, the care provider 106 cannot conclusively determine whether the subject 102 has a type of cancer based on noninvasive diagnostic techniques. For example, the care provider 106 is unable to identify a type of the lesion 104 (e.g., a tumor) using imaging techniques. The care provider 106 may be unable to identify a characteristic of the subject 102 presenting with a disease (e.g., cancer), wherein the characteristic is determinative of, or at least correlated with, an effectiveness of at least one therapy at treating the disease, an ineffectiveness of at least one therapy at treating the disease, a survivability (e.g., a likelihood that the subject will survive by a predetermined date or time), an expected quality of life, at least one predetermined symptom, at least one comorbidity, another factor relevant to the prognosis associated with the disease, or any combination thereof.
[0085] To further assess the condition of the subject 102, a sample 108 is obtained from the subject 102. In some examples, the sample 108 includes a tissue biopsy sample. For instance, the sample 108 is obtained by removing cells from the lesion 104 and from the subject 102. In some cases, the tissue biopsy sample is surgically excised from the subject 102. The care provider 106 could identify the condition of the subject 102 using histochemistry and / or immunohistochemistry. For instance, the care provider 106 could surgically remove a tissue sample from the lesion 104 and / or review the tissue sample using histochemistry and / or immunohistochemistry.
[0086] In some cases, the sample includes a liquid biopsy sample. The liquid biopsy sample 108, for instance, includes blood, plasma, cerebrospinal fluid, sputum, stool, urine, lymphatic fluid, saliva, or some other fluid obtained from the body of the subject 102. In some cases, a blood sample is obtained intravenously from the subject 102. The liquid biopsy sample 108, according to various examples, is a plasma sample obtained from the blood of the subject 102. The liquid biopsy sample 108, for instance, can be obtained in a minimally invasive procedure, which could be performed by a medical technician rather than a surgeon.
[0087] The sample 108 includes nucleic acid molecules 110. According to some examples, the nucleic acid molecules 110 include genomic DNA (gDNA). For instance, the nucleic acid molecules 110 include chromosomal DNA that is located in, or extracted from, cells in the sample 108. According to some cases, the DNA is extracted from nuclei and the cells in the sample 108 using mechanical shearing and / or the introduction of a chemical (e.g., a detergent). The DNA may be subsequently isolated from proteins and other cellular materials. In some implementations, the nucleic acid molecules 110 indicate an entire genome of the subject 102 and / or the lesion 104. Thus, a genome of the subject 102 and / or the lesion 104 can be determined by sequencing the DNA in the nucleic acid molecules 110.
[0088] In some examples, the nucleic acid molecules 110 include RNA. In some implementations, the nucleic acid molecules 110 include messenger RNA (mRNA), microRNA, non-coding RNA, functional RNA, or any combinationFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT thereof. Various RNA in the nucleic acid molecules 110 may be indicative of proteins expressed in the cells of the subject 102 and / or the lesion 104.
[0089] In some cases, the nucleic acid molecules cell-free DNA (cfDNA). In examples in which the subject 102 has cancer (e.g., the lesion 104 is a cancerous tumor), the cfDNA, for instance, includes circulating tumor DNA (ctDNA) and / or non-ctDNA. In cases wherein the lesion 104 is a tumor, cancer cells within the lesion 104 will lyse and release the ctDNA into the bloodstream of the subject 102. In some cases, the ctDNA is released from circulating tumor cells (CTCs). Further, other cells additionally release non-ctDNA into the bloodstream of the subject. In general, the cfDNA includes fragments with lengths that are in a range of 1 to 500, 3 to 500, or 100 to 500 bases long. For instance, the cfDNA includes fragments that are 170 bases long and / or fragments that are 340 bases long. For example, the cfDNA includes fragments that are 100 to 240 bases long and / or fragments that are 270 to 410 bases long.
[0090] In various cases, the sample 108 is transported to a location that is remote from the subject 102 for further processing. For example, the sample 108 is removed from the subject 102 in a clinical environment (e.g., a hospital) and is then transported to a remote laboratory for further testing and analysis.
[0091] A sequencer 112 is configured to generate sequence read data 114 indicating the sequences of the nucleic acid molecules 110. The sequencer 112, for instance, includes one or more devices that are configured to generate the sequence read data 114 by processing at least a portion of the sample 108. In some cases, the nucleic acid molecules 110 are extracted from the sample 108. The extraction can be performed by the sequencer 112, by another device, manually (e.g., by a laboratory technician), or any combination thereof. Any appropriate extraction method known to those of ordinary skill in the art can be utilized.
[0092] In various cases, the sequencer 112 is configured to perform one or more processes (e.g., chemical reactions) on the nucleic acid molecules 110 in order to prepare the nucleic acid molecules 110 for sequencing. For instance, the sequencer 112 may ligate adapters onto the nucleic acid molecules 110 and / or amplify the nucleic acid molecules 110, such that numerous copies of the ligated nucleic acid molecules 110 are available for sequencing. Examples of the adapters include, for example, amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. The nucleic acid molecules 110 (e.g., the ligated nucleic acid molecules 110) may be amplified by generating multiple copies of the nucleic acid molecules 110 using one or more techniques such as polymerase chain reaction (PCR), a non-PCR amplification technique, or an isothermal amplification technique.
[0093] The sequencer 112 may identify the length, position, and identity of the bases in the nucleic acid molecules 110 by sequencing the nucleic acid molecules 110 (e.g., the amplified and / or ligated nucleic acid molecules 110). In various implementations, the sequencer 112 utilizes first-generation sequencing (e.g., Sanger sequencing), second- generation sequencing (e.g., massive parallel sequencing), third-generation sequencing (e.g., nanopore sequencing), or a combination thereof. In some cases, the sequencer 112 is configured to sequence substantially all of the nucleotides of all of the nucleic acid molecules 110 fragments obtained from the sample 108. In some examples, the sequencer 112 is configured to perform targeted sequencing. For instance, the sequencer 112 may determine whether the nucleic acid molecules 110 fragments contain one or more predetermined sequences at one or more genomic locations.
[0094] In various cases, the sequencer 112 includes one or more sensors that are configured to detect physical signals (also referred to as “detection signals”) that are indicative of the nucleotide sequences of the nucleic acidFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT molecules 110. The sequencer 112 may perform sequencing-by-synthesis. For example, the sequencer 112 may include one or more optical sensors configured to detect optical signals emitted from fluorescently tagged nucleotide triphosphates (NTPs) that are joined together in a synthesized DNA strand using the ligated nucleic acid molecules 110 as templates. The optical signals detected by the optical sensor(s), for instance, are indicative of the sequences of the nucleic acid molecules 110. The sequencer 112 may perform nanopore sequencing. In various cases, the sequencer 112 includes one or more electrical sensors configured to measure an electrical signal (e.g., an electrical current) across a substrate as the ligated nucleic acid molecules 110 are directed through a nanopore extending through the substrate. The electrical signal over time, in various cases, is indicative of the sequences of the nucleic acid molecules 110 in the sample 108. The sequencer 112, in various implementations, is configured to generate the sequence read data 114 as digital data based on the analog signals detected by the sensor(s). For instance, the sequencer 112 includes one or more analog to digital converters (ADCs). In various cases, the sequencer 112 includes at least one processor configured to generate the sequence read data 114.
[0095] In some implementations, the sequencer 112 performs RNA sequencing (RNA-seq) on the nucleic acid molecules 110. For example, the nucleic acid molecules 110 include RNA that is extracted from the sample 108. In some examples, the RNA in the nucleic acid molecules 110 is fragmented. In various implementations, complementary DNA (cDNA) is generated using reverse transcriptase, such that the cDNA includes sequences that are complementary to the RNA in the nucleic acid molecules 110 from the sample 108. The cDNA, according to various cases, can be sequenced using the DNA sequencing techniques described above. Accordingly, in some cases, the sequence read data 114 indicates sequences of RNA present in the sample 108, which may be indicative of the transcriptome of the subject 102 and / or the lesion 104.
[0096] In various cases, the sequencer 112 performs sequencing on a subset of the nucleic acid molecules 110. For instance, the sequencer 112 may perform targeted sequencing on one or more predetermined genes, such as any of the genes described herein. The sequencer 112, in some cases, may refrain from sequencing at least a portion of the nucleic acid molecules 110 that do not correspond to the subset.
[0097] In various implementations of the present disclosure, the environment 100 can be utilized to assess whether a particular gene mutation occurred before, or after, the gene is amplified. This determination can have important consequences for care, diagnosis, and prognosis of the subject 102.
[0098] During a cancer disease process, a particular mutation may occur within the genome of cells that develop into cancer cells and / or are cancer cells. This mutation may characterize the type of cancer that the subject 102 is experiencing, a subtype of the cancer, an expected prognostic outcome of the subject 102, and so on. In particular cases, the mutation may impact the susceptibility of the cancer cells to a particular type of treatment. In some examples, the mutation may cause the cancer cells to be resistant to a type of treatment.
[0099] In some cases, the mutation is amplified in the genome of the cancer cells. That is, multiple copies of the mutation may develop in various generations of the cancer cells. The presence of multiple copies of the mutation can further impact its pertinence to diagnostic, therapeutic, and prognostic characteristics of the disease of the subject 102. For example, amplification of the mutation may, in some cases, cause greater expression of the mutation. However, if the mutation itself occurred after amplification, such that the mutation is presence in an insignificant fraction of the copies of an amplified gene, then the presence of the mutation may have limited clinical relevance. It may be clinicallyFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT important to assess whether a particular mutation-of-interest, which can be represented by a sequence-of-interest, occurred before or after amplification of the sequence-of-interest in the cancer disease process of the subject 102.
[0100] In various implementations of the present disclosure, the environment 100 can be utilized to determine whether a given mutation (e.g., resulting in a sequence-of-interest) occurred before or after amplification. In various cases, the sequence read data 114 is output to an allele fraction calculator 116. The allele fraction calculator 116 is configured to generate an expected fraction 118 of one or more genomic sequences-of-interest (e.g., one or more sequences associated with particular alleles) of the sample 108 as indicated in the sequence read data 114. The expected fraction 118, in various cases, is represented as a number that is equal to 1.0 or 2.0, for instance. The expected fraction 118 represents an expected allelic fraction for a given sequence-of-interest (e.g., a gene) if it were fully clonal and present on every allele of an amplification of that region.
[0101] Fractions of various types of sequences-of-interest may be determined by the allele fraction calculator 116, such as a sequence representing an allele associated with at least one of ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7,FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT VEGF, VEGFA, or VEGFB. In particular examples, the expected fraction 118 is calculated by the allele fraction calculator 116 for a sequence (e.g., a sequence corresponding to an allele and / or mutation) corresponding to at least one of an AR gene, ATM, TET2, DNMT3A, ASXL1, LYN, SF3B1, RB1, TP53, or SPOP.
[0102] In particular cases, the allele fraction calculator 116 calculates the expected fraction 118 based on a sequence (e.g., a sequence corresponding to an allele and / or mutation) associated with a sex-linked gene. Examples of x-linked genes include AR gene, ATRX, CD99, DKC1, EDA2R, ELF4, FAM123B, FOXP3, LDOC1, RBBP7, RPS6KA6, PHF6, RPL10, and WTX. Examples of Y-linked genes include KDM5D and UTY. In some cases, the sequence is associated with an autosomal gene.
[0103] In various cases, the allele fraction calculator 116 aligns the sequence reads onto a reference genome in order to identify the genomic position of the sequence reads. Accordingly, the allele fraction calculator 116 may determine a number of the sequence reads that correspond to the sequence-of-interest. The expected fraction 118, in various cases, represents a quotient of the number of the sequence reads that correspond to the allele over a number of sequence reads observed at the sequence-of-interest (e.g., the gene-of-interest).
[0104] In various cases, the expected fraction 118 is calculated based on a copy number associated with the sequence-of-interest. For example, the copy number corresponds to a number of copies of the sequence-of-interest present in the genome of cells of the subject 102, such as cells in the lesion 104 itself. The genomes of cancer cells, for instance, can include various types of copy number variants. In some cases, these copy number variants can have a significant impact on the prognostic outcomes of host subjects, treatment resistance, treatment efficacy, and other characteristics related to expression. In some cases, if the cells of the lesion 104 include a significant number of copies of a gene, the cells may overexpress the gene.
[0105] Copy number can be calculated, for instance, by analyzing the sequence read data 114. In various cases, the copy number represents a modeled copy number, such as a copy number determined according to the techniques described in International Publication No. WO 2023 / 060250, which is incorporated by reference herein in its entirety. For instance, the modeled copy number is determined by fitting a model (e.g., a copy number grid model corresponding to the modeled copy number) to the sequence read data 114. In various cases, the modeled copy number can be calculated by determining a minor allele coverage ratio (e.g., for a gene-of-interest) by analyzing the sequence read data 114 and determining a major allele coverage ratio by analyzing the sequence read data 114. The minor and major allele coverage ratios, for instance, logically relate to the copy number of the gene associated with the minor and major alleles, as well as the tumor purity of the sample and the ploidy of the gene. The modeled copy number can further be determined by segmenting the genome into genomic segments and generating input data by determining a difference and / or sum between the major allele coverage ratio and the minor allele coverage ratio for genetic loci in the genome. If the difference between the major allele coverage ratio and the minor allele coverage ratio is plotted against the sum of the major allele coverage ratio and the minor allele coverage ratio, each genetic locus is expected to lie on one of a set of evenly spaced grid points. Copy numbers, for instance, are necessarily integer values. Various copy number grid models can be generated that represent the copy number space scaled and translated as a function of ploidy and tumor purity values. For instance, the grid models can correspond to different tumor purity and tumor ploidy estimates.
[0106] In some cases, the modeled copy number is determined by further fitting different copy number grid models (corresponding to allowed copy number states) to the input data, and selecting one of the copy number grid modelsFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT that fits the input data. The modeled copy number, for instance, corresponds to the selected copy number grid model. In various cases, the modeled copy number is an integer value.
[0107] In some cases, the expected fraction 118 is based, at least in part, on a tumor purity of the sample 108. The tumor purity, in various cases, represents the amount of nucleic acid molecules 110 in the sample 108 and / or the amount of sequence reads in the sequence read data 114 corresponding to nucleic acid molecules of a tumor, such as the lesion 104. For instance, the tumor purity may be determined based on an amount of somatic copy-number alterations (SCNA), an amount of single-nucleotide variants (SNVs), a minor allele frequency (MAF), or any combination thereof, within the sequence read data 114. In cases where the sample 108 is a tissue biopsy sample, the tumor purity may be relatively high. However, in cases where the sample 108 is a liquid biopsy sample, the tumor purity may be relatively low.
[0108] Tumor purity can be calculated by analyzing the sequence read data 114. In particular cases, the tumor purity can be calculated by identifying a fraction of the sequence read data 114 that corresponds to tumor characteristics. These characteristics may include at least one of the presence of one or more variants associated with cancer, a length of DNA fragments associated with cancer cells, copy number patterns, and the like. In various cases, tumor purity can be represented as a fraction, percentage, or decimal number that is less than or equal to one.
[0109] According to some examples, the allele fraction calculator 116 determines the expected fraction 118 using the following Equation 1:wherein CN is the copy number of the sequence-of-interest (e.g., gene-of-interest) in the sample 108 and TP is the tumor purity of the nucleic acid molecules 110 in the sample 108. CN, for instance, can be representative of a modeled copy number. Equation 1 can be used to determine the expected fraction 118 for a sex linked gene in a male subject, for instance. As used herein, the term “male subject” may refer to a subject having a genome with a single y chromosome and a single x chromosome.
[0110] In various implementations, the allele fraction calculator 116 determines the expected fraction 118 using the following Equation 2:wherein CNN is an expected number of copies of the sequence-of-interest (e.g., the allele sequence) in a non- cancerous sample. For instance, if the subject 102 is male, the CNN for X- or Y-linked genes (of which the male subject 102 has a single copy) is 1.0. If the subject 102 is not male and / or the gene is an autosomal gene, the CNN, for instance, is 2.0.
[0111] In various implementations, a ratio calculator 120 receives the expected fraction 118. The ratio calculator 120 calculates a ratio 122 based on the expected fraction 118 and an observed fraction 124 associated with the sequence-of-interest. The ratio may be in a range of 0.0 to 1.0, for example. For instance, the ratio 122 is calculated using the following Equation 3:wherein FO is the observed fraction 124 and FE is the expected fraction 118.FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT
[0112] In various cases, the observed fraction 124 corresponds to an observed allelic fraction of the sequence-of- interest. In various cases, the allele fraction calculator 116 determines the observed fraction 124 by analyzing the sequence read data 114. According to some examples, the observed allelic fraction is determined by determining the number of reads that have the sequence-of-interest (e.g., a variant or mutation) in a locus divided by the number of reads spanning the locus.
[0113] A timing determiner 126 is configured to determine whether a mutation associated with the sequence-of- interest (e.g., gene) occurred prior to amplification of the sequence-of-interest based, at least in part, on the ratio 122. In some examples, the timing determiner 126 compares the ratio 122 to one or more thresholds. In various examples, the threshold(s) are in a range of 0.1 to 0.9, 0.2 to 0.8, 0.3 to 0.7, 0.4 to 0.6, or 0.45 to 0.55. For instance, at least one of the threshold(s) is 0.5. In various implementations, the timing determiner 126 infers that the mutation occurred prior to amplification if the ratio 122 is above at least one of the threshold(s). In contrast, the timing determiner 126 infers that the mutation occurred after amplification if the ratio 122 is below at least one of the threshold(s). In some implementation, a numeric discrepancy between the threshold(s) and the ratio 122 is indicative of (e.g., proportional to) a certainty of the inference of whether the mutation occurred before or after amplification.
[0114] The timing determiner 126 provides a timing indicator 130 to a report generator 132. The timing indicator 130 is generated based on the comparison of the ratio 122 to the one or more thresholds. For example, the timing indicator 130 may include an indication of whether the timing determiner 126 has inferred that the mutation occurred prior to amplification, an indication of whether the timing determiner 126 has inferred that the mutation occurred after amplification, or a combination thereof. The timing indicator 130, in some cases, includes a certainty of the indication that the mutation occurred before or after amplification.
[0115] The report generator 132 is configured to generate a report 134 based, at least in part, on the timing indicator 130. The report 134, for example, includes consumable data that can inform the care provider 106 about a condition of the subject 102. In various implementations, the report 134 may indicate the results of additional analyses, such as the results of a histological study, whole transcriptome sequencing, cfRNA sequencing, whole exome sequencing, whole genome sequencing, a cancer (e.g., DNA) hotspot panel test, a DNA methylation test, a tumor mutational burden (TMB) test, a DNA fragmentation test, an RNA fragmentation test, a microsatellite instability (MSI) test, or a viral status test. The report 134, for example, may include a genomic profile of the subject 102 based on various combinations of the above analyses and tests.
[0116] Optionally, the report generator 132 includes a classifier that classifies a condition of the subject 102 based, at least in part, on the timing indicator 130. For instance, the report generator 132 may include a pretrained machine learning (ML) model that is configured to classify the condition of the subject 102 based on input data that includes the timing indicator 130. The input data may also include other characteristics of the subject 102 and / or the sample 108, for instance.
[0117] The classifier, in various cases, includes one or more model types. For instance, the classifier includes at least one of an artificial neural network (e.g., feedforward neural networks, multi-layer perceptrons (MLPs), convolutional neural networks (CNNs), and backpropagation models), a nearest-neighbor model, a regression analysis model, a clustering model (e.g., k-means clustering, mean-shift clustering, expectation-maximization (EM) clustering, and agglomerative hierarchical clustering), a principal component analysis model, a gradient boosting model, a random forest, or the like.FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT
[0118] The classifier, in various cases, may be defined by various parameters that enable the classifier to identify predictive attributes of the input data that are correlated to or otherwise associated with example categories, such as cancer types, treatment resistance, effective treatments, prognostic indicators, or other clinically relevant classifications.
[0119] In some implementations, the report 134 indicates that a follow-up test of the subject 102 is recommended and / or needed. For instance, in response to determining that the categorization of the condition of the subject 102 is inconclusive, the report generator 128 may generate the report 134 to indicate that one or more additional tests (e.g., a histological study, genome sequencing, exome sequencing, additional DNA sequencing, RNA sequencing, transcriptome sequencing, etc.) should be performed in order to accurately identify the condition of the subject 102.
[0120] In various cases, the report 134 is output to a clinical device 136. For example, the report generator 132 transmits the report 134 to the clinical device 136. In various implementations, the clinical device 136 is a computing device that is operated by, owned by, or otherwise associated with the care provider 106. For instance, the clinical device 136 may be a desktop computer, a laptop computer, a smart phone, or some other computing device associated with the care provider 106. The clinical device 136, in various cases, outputs the report 134 to the care provider 106. In some cases, the clinical device 136 includes a display (e.g., a screen) that visually presents the report 134. In various cases, the clinical device 136 includes a speaker that outputs a sound indicative of the report 134. The clinical device 136, in various cases, may output the information in the report 134 using one or more output mechanisms or devices.
[0121] The care provider 106 may review the report 134 by interacting with the clinical device 136. The report 134, in various cases, may enhance the clinical decision-making of the care provider 106. For instance, the care provider 106 may prepare and / or administer a therapy to the subject 102 based on the report 134. According to various implementations, the care provider 106 may initiate the therapy and / or refer the subject 102 to another care provider to receive the therapy. In various cases, if the predicted condition of the subject 102 is a disease (e.g., cancer), the care provider 106 may prescribe, recommend, or administer an agent in order to treat the disease the subject 102. For instance, if the predicted condition of the subject 102 is a type of cancer, the care provider 106 may administer an anticancer therapy to the subject 102.
[0122] In various implementations, the care provider 106 may develop a diagnosis and / or prognosis of the subject 102 based on the report 134. For instance, the care provider 106 may determine to administer a treatment including at least one of a chemotherapy, radiation therapy, immunotherapy, or surgery to the patient. In some cases, the care provider 106 determines a dosage of a treatment based on the report 134. In various implementations, the care provider 106 may communicate information in the report 134 to the subject 102. According to some cases, the care provider 106 may determine that the subject 102 qualifies for a clinical trial based on the report 134.
[0123] FIG.1 illustrates various elements that can be embodied in one or more computing devices. For example, at least a portion of the functions of the sequencer 112, the allele fraction calculator 116, the ratio calculator 120, the timing determiner 126, the report generator 132, and the clinical device 136 are performed by one or more processors in at least one computing device. Examples of computing devices include server computers, desktop computers, laptop computers, tablet computers, mobile phones, wearable devices, Internet of Things (IoT) devices, and the like. In various cases, instructions for performing at least a portion of the functions of these elements are stored in memory and / or in a non-transitory computer readable medium. The instructions, for instance, are executed by the processor(s).FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT
[0124] FIG.1 also illustrates various types of data. For example, the sequence read data 114, the expected fraction 118, the ratio 122, the observed fraction 124, the timing indicator 130, the report 134, or any combination thereof, includes data. The various types of data illustrated in FIG.1 may be stored, such as in memory or in non-transitory computer readable media. In various implementations, at least a portion of the data is transmitted or otherwise output by one or more computing devices. For example, a computing device may transmit one or more communication signals to another computing device, wherein the communication signal(s) encode at least a portion of the data. Examples of communication signals include electromagnetic signals, optical signals, ultrasonic signals, optical signals, and electrical signals. For example, communication signals can be transmitted wirelessly and / or in a wired fashion. The communication signals, for instance, are transmitted over one or more wireless channels and / or one or more wired channels (e.g., optical cabling, electrical cabling, etc.). In various cases, the communication signal(s) are transmitted over one or more communication networks. A communication network, for instance, may be defined according to one or more physical channels, such as one or more frequency spectra. In some cases, a communication network is defined according to one or more communication protocols and / or standards. Examples of communication networks include fiber optic networks, Institute of Electrical and Electronics Engineers (IEEE) networks (e.g., WI-FI™ networks, WiMAX networks, BLUETOOTH™ networks, etc.), cellular networks (e.g., a 3rdGeneration Partnership Project (3GPP) radio network, such as a Long Term Evolution (LTE) network, a New Radio (NR) network; or a cellular core network such as a 3rdGeneration (3G) core, a 4thGeneration (4G) core, a 5thGeneration (5G) core, etc.), ultrasonic networks, and the like. In some cases, the data is broadcasted from one device to multiple other devices. In some cases, the data is unicasted from one device to another device. For instance, various forms of data described herein may be transmitted via a peer-to-peer (P2P) connection.
[0125] A particular example will now be described with reference to FIG.1. For instance, the subject 102 may be a male subject having a single X chromosome. The subject 102 may have prostate cancer. According to some examples, the prostate cancer is metastatic. For instance, the prostate cancer may be hormone-sensitive and / or castration-resistant. The care provider 106 may obtain the sample 108 by performing a tissue biopsy procedure on the lesion 104. In various cases, the sequence read data 114 indicates that cancer cells in the lesion 104 exhibit one or more variants in an AR gene, which is present on the X chromosome. The variant(s), for instance, are associated AR ligand binding mutations, which may represent an acquired mechanism of resistance to ASIs. The AR gene, furthermore, has been amplified such that the genome of the cancer cells have numerous copies of the AR gene.
[0126] In various cases, a sequence-of-interest including at least a portion of the mutated AR gene is identified. The allele fraction calculator 116 determines the expected fraction 118 based on the sequence-of-interest representing the mutated AR gene and Equation 1 and / or 2. For instance, because the AR gene is present on the X chromosome of the male subject 102, the observed fraction 124 may be 1.0. In various cases, the ratio 122 of the expected fraction 118 and the observed fraction 124 is compared to a threshold of 0.5, which may cause the timing determiner 126 to infer whether the mutations of the AR gene occurred prior to amplification. The timing indicator 130 is included in the report 134 provided by the care provider 106.
[0127] The report 134, for instance, may impact a treatment that the care provider 106 recommends and / or provides to the subject 102. In some examples, the timing of the AR gene mutations relative to amplification is at least one indicator that the subject 102 has a particular subtype of prostate cancer. For instance, the care provider 106 mayFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT determine to administer an AR-targeted therapy to the subject 102 based on the timing indicator 130. In particular cases, the care provider 106 may determine whether cancer cells of the subject 102 are susceptible and / or resistant to one or more therapies based on the timing indicator 140. These therapies, for instance, may include administration of at least one of an AR inhibitor, an AR degrader, an androgen deprivation therapy (ADT), an ASI, or systemic taxane.
[0128] FIG.2 illustrates an example report 200 summarizing predicted categories of a cancer of a subject. In various cases, the report 200 is the report 134 described above with reference to FIG.1. The report 200, for instance, may be displayed to a patient and / or care provider. In some cases, the report 200 is generated based on features of a sample (e.g., a liquid biopsy sample) obtained from the subject.
[0129] The report 200 includes a tissue origin 202 of the cancer. The tissue origin 202, for instance, indicates a histological tissue type 204, a primary site 206, cell subtype 207, or any combination, of the cancer.
[0130] In various cases, the report 200 includes one or more therapy indicators 208. For instance, the therapy indicator(s) 208 convey whether the cancer is predicted to be resistant to one or more predetermined therapies and / or whether the cancer is predicted to be responsive to one or more predetermined therapies. In some cases, the therapy indicator(s) 208 convey at least one dosage of a therapy that is predicted to effectively treat the cancer of the subject.
[0131] In some examples, the report 200 includes one or more prognostic indicators 210. The prognostic indicator(s) 210, for instance, indicate a prognosis of the subject in view of the categorized cancer. For example, the prognostic indicator(s) 210 may indicate a survivability, a recoverability, a quality of life indicator, or other information indicative of the prognosis of the subject.
[0132] The report 200 may include a trial qualification 212 of the subject. The trial qualification 212, for instance, indicates whether the subject is predicted to qualify for a predetermined clinical trial.
[0133] The report 200, in various implementations, includes a metastasis profile 214 of the subject. The metastasis profile 214, for instance, indicates a likelihood that the cancer will metastasize (e.g., at a particular point in time), one or more tissues in which the cancer is predicted to metastasize, or the like.
[0134] In various cases, the report 200 includes recommended follow-up tests 216. For example, the report 200 may include a recommendation to perform whole genome sequencing on the subject, particularly in cases if the cancer cannot be categorized above a threshold certainty.
[0135] The report 200 may include a genomic profile 218 of the subject. In various cases, the genomic profile 218 includes or is generated based on the results of one or more genomic analyses of the subject.
[0136] In various examples, the tissue origin 202, the histological tissue type 204, the primary site 206, the cell subtype 207, the therapy indicator(s) 208, the prognostic indicator(s) 210, the trial qualification 212, the metastasis profile 214, the follow-up tests 216, the genomic profile 218, or any combination thereof, may be determined based, at least in part, on the timing indicator 130. In some implementations, the timing indicator 130 itself is included in the report 200.
[0137] FIG.3 illustrates a process 300 for determining whether a mutation occurred prior to amplification. In various cases, the process 300 may be performed by an entity including at least one processor, a computing device, a medical device, or a combination thereof.
[0138] At 302, the entity determines an observed allele fraction of a sequence in a sample. The observed allele fraction may be the observed fraction 124 discussed above with respect to FIG.1. In various cases, the entity identifiesFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT data indicative of nucleic acid molecules sequences of the sample. For example, the data is sequence read data. In various cases, the sequence read data is representative of DNA sequences in the sample. In various cases, the entity determines the observed allele fraction of a predetermined sequence present in the sequence read data. For instance, the entity determines the observed allele fraction of an allele associated with a gene and / or the gene itself. In various examples, the gene is a sex-linked gene (e.g., a y-linked gene or an x-linked gene). In some cases, the sample is obtained from a male subject. For instance, the gene is an AR gene. According to some examples, the observed allele fraction is an AR short variant allelic fraction. For instance, the AR short variant allelic fraction is the observed allele fraction for a sequence-of-interest associated with one or more variants in the AR gene. In some cases, the subject has prostate cancer.
[0139] At 304, the entity determines an expected allele fraction of the sequence. The expected allele fraction may be the expected fraction 118 discussed above with respect to FIG.1. The expected allele fraction, for instance, represents an allelic fraction for the sequence if it were fully clonal and present on every allele of an amplification of the sequence in the genome. That is, the expected allele fraction represents the allelic fraction for the sequence if it was present on every copy of the sequence in the genome of the sample. In various cases, the expected allele fraction is calculated based on a tumor purity of the sample and / or a copy number of the sequence in the genome of the sample. The copy number, for instance, is a modeled copy number.
[0140] At 306, the entity determines a ratio of the observed allele fraction to the expected allele fraction. The ratio may be the ratio 122 discussed above with respect to FIG.1. At 308, the entity determines whether a mutation associated with the sequence occurred prior to amplification by comparing the ratio to at least one threshold. For example, the threshold(s) may be in a range of 0.2 to 0.8. In some cases, the ratio is compared to a single threshold defined as 0.5. In various cases, the entity infers that the mutation occurred prior to amplification if the ratio exceeds the threshold. In some examples, the entity infers that the mutation occurred after amplification of the ratio is lower than the threshold.
[0141] In various cases, the relative timing between the mutation and amplification can inform treatment decisions for the subject. For example, in some cases, if cancer cells of the subject are determined to have one or more mutations to the AR gene that occurred prior to amplification, then the cancer cells of the subject may be differentially susceptible to AR inhibitors and / or AR degraders, relative to subjects that have mutations to the AR gene that occurred after amplification. According to some cases, the relative timing can be used to classify the cancer cells of the subject.
[0142] FIG.4 illustrates an example environment 400 for sequencing various nucleic acid molecules 402. Reference will be made to “amplification” in the description of FIG.4, however, it should be noted that the amplification discussed with respect to FIG.4 may be distinct from genomic amplification discussed elsewhere in the disclosure.
[0143] In various implementations, the nucleic acid molecules 402 include cfDNA and / or gDNA. For instance, the nucleic acid molecules 402 may include ctDNA. The nucleic acid molecules 402, in various cases, are extracted from a sample, such as a biological sample obtained from a subject. In some implementations, the nucleic acid molecules 402 include DNA that is complementary to RNA present in the sample.
[0144] The nucleic acid molecules 402, in various cases, are ligated with adapters 404. For examples, the adapters 404 are hybridized to the nucleic acid molecules 402. The adapters 404, for example, include additional nucleic acid molecules. In various implementations, the adapters 404 have a shorter length than the nucleic acid molecules 402FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT being sequenced. For instance, the adapters 404 include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. Although FIG.4 illustrates adapters 404 being ligated to one end of each of the nucleic acid molecules 402, implementations are not so limited. For example, the adapters 404 may be ligated to both ends of each of the nucleic acid molecules 402.
[0145] In various examples, the nucleic acid molecules 402 ligated with the adapters 404 are amplified in order to generate amplified molecules 406. Various amplification techniques can be performed. For instance, the amplified molecules 406 are generated using PCR, a non-PCR amplification technique, an isothermal amplification technique, or any combination thereof.
[0146] Amplified molecules 406 may be captured by bait molecules 410 and sequenced. In some implementations, the amplified molecules 406 are sequenced via sequencing-by-synthesis. In various cases, fluorescently tagged deoxyribonucleotide triphosphates (dNTP) 412 are utilized to synthesize a strand that is complementary to DNA strands bound to the substrate 408. When a dNTP 412 is added to the strand (e.g., by an enzyme), the dNTP 412 emits an optical signal 414. In various implementations, the frequency of the optical signal 414 is dependent on the type of dNTP 412 from which the optical signal 414 is emitted. By detecting the optical signals 414 as the strand is being synthesized, the sequence of the original nucleic acid molecules 402 can be derived.
[0147] In some implementations, the amplified molecules 406 are sequenced via nanopore sequencing. For instance, the amplified molecules 406 are directed through a nanopore 416 extending through a substrate 418. In various cases, the amplified molecules 406 are negatively charged, such that they can be directed through the nanopore 416 by imposing an electrical field across the substrate 418. In various cases, the amplified molecules 406 and the nanopore 416 are in the presence of a charged solution. Thus, charged solutes traveling through the nanopore 416 can be monitored by reviewing an electrical signal (e.g., a current) sensed between electrodes 420 on either side of the substrate 418. As an amplified molecule 406 is directed through the nanopore 416, the individual bases within the amplified molecule 406 will block the nanopore 416, which may decrease the amount of charged solutes traveling through the nanopore 416 and consequently, the magnitude of the electrical signal detected by the electrodes 420. Each of the four types of bases within the amplified molecules 406, may block the nanopore 416 to a different extent. Therefore, the sequence of the nucleic acid molecules 402 can be derived by analyzing the measured electrical signal with respect to time as the amplified molecules 406 are directed through the nanopore 416.
[0148] FIG.5 illustrates one or more devices 500 configured to perform various operations described herein. The device(s) 500 include one or more processor(s) 502. In some implementations, the processor(s) 502 includes a central processing unit (CPU), a graphics processing unit (GPU), both CPU and GPU, or other processing unit or component known in the art.
[0149] The processor(s) 502 is operably connected to memory 504. In various implementations, the memory 504 is volatile (such as random access memory (RAM)), non-volatile (such as read only memory (ROM), flash memory, etc.) or some combination of the two. The memory 504 stores instructions that, when executed by the processor(s) 502, causes the processor(s) 502 to perform various operations. In various examples, the memory 504 stores methods, threads, processes, applications, objects, modules, any other sort of executable instruction, or a combination thereof. In some cases, the memory 504 stores files, databases, or a combination thereof. In some examples, the memory 504 includes, but is not limited to, RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flashFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT memory, or any other memory technology. In some examples, the memory 504 includes one or more of CD-ROMs, digital versatile discs (DVDs), content-addressable memory (CAM), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the processor(s) 502. For instance, the memory 504 stores instructions that, when executed by the processor(s) 502, causes the processor(s) 502 to perform operations of the allele fraction calculator 116, the ratio calculator 120, the timing determiner 126, the report generator 132, or any combination thereof.
[0150] The processor(s) 502 is operably connected to one or more input devices 506 and one or more output devices 508. Collectively, the input device(s) 506 and the output device(s) 508 function as an interface between at least one user and the device(s) 500. The input device(s) 506 is configured to receive an input from a user and includes at least one of a keypad, a cursor control, a touch-sensitive display, a voice input device (e.g., a microphone), a haptic feedback device (e.g., a gyroscope), or any combination thereof. The output device(s) 508 includes at least one of a display, a speaker, a haptic output device, a printer, or any combination thereof. In various examples, the processor(s) 502 causes a display among the input device(s) 506 to visually output various data described herein. In some implementations, the input device(s) 506 includes one or more touch sensors, the output device(s) 508 includes a display screen, and the touch sensor(s) are integrated with the display screen.
[0151] In various implementations, the processor(s) 502 is operably connected to one or more transceivers 510 that transmit and / or receive data over one or more communication networks 512. For example, the transceiver(s) 510 includes a network interface card (NIC), a network adapter, a local area network (LAN) adapter, or a physical, virtual, or logical address to connect to the various external devices and / or systems. In various examples, the transceiver(s) 510 includes any sort of wireless transceivers capable of engaging in wireless communication (e.g., radio frequency (RF) communication). For example, the communication network(s) 512 includes one or more wireless networks that include a 3rdGeneration Partnership Project (3GPP) network, such as a Long Term Evolution (LTE) radio access network (RAN) (e.g., over one or more LTE bands), a New Radio (NR) RAN (e.g., over one or more NR bands), or a combination thereof. In some cases, the transceiver(s) 510 includes other wireless modems, such as a modem for engaging in WI-FI®, WIGIG®, WIMAX®, BLUETOOTH®, or infrared communication over the communication network(s) 512.
[0152] The device(s) 500 may further include the sequencer 112. In various implementations, the sequencer 112 includes one or more fluidic circuits 514 configured to receive a sample 516 derived from a subject 517. The sequencer 112, in various cases, may be configured to generate data indicative of one or more sequences of nucleic acid molecules (e.g., DNA and / or RNA) present in the sample 516. In various cases, the sequencer 112 introduces one or more reagents 518 to the fluidic circuit(s) 514 in order to prepare for and perform sequencing of the nucleic acid molecules. Further, the sequencer 112 may include one or more sensors 520 configured to measure or otherwise detect detection signals from the fluidic circuit(s) 514, which may be indicative of the sequences of the nucleic acid molecules. According to various implementations, the sensor(s) 520 may further include one or more ADCs. The sequencer 112, in various cases, outputs sequence read data to the processor(s) 502 for additional processing. First Experimental Example
[0153] FIG.6 illustrates an experimentally derived example of ratios of expected and observed AR-related variant allele frequencies for various patients with prostate cancer. As shown, most patients had AR-related mutations thatFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT occurred after amplification. However, some patients had AR-related mutations corresponding to the variant allele after amplification. The latter patient group may be particularly sensitive to next-generation AR agents, such as AR degraders (e.g., ARV-110). Second Experimental Example
[0154] Numerous preclinical hypotheses, such as the interaction between androgen receptor (AR) signaling and the PI3K-Akt pathway in patients with PTEN loss or the induction of a ‘BRCAness’ phenotype by AR signaling inhibitors (ASIs) regardless of BRCA1 / 2 status, have failed to translate effectively into clinical benefits for patients. This gap in clinical translatability, particularly in metastatic castration-resistant prostate cancer (mCRPC) research, may arise from a misalignment between preclinical models, where these hypotheses originate, and the mCRPC disease that is often targeted first in clinical trials. A major problem is that it is not fully understand if the genetic and biological mechanisms identified in early-stage tumors are relevant in mCRPC, which shows a lot of tumor variability. By studying the clinical and genomic characteristics of tumors in both metastatic hormone-sensitive prostate cancer (mHSPC) and mCRPC stages, and examining how genetic changes or theoretical concepts relate to treatment responses over time, insights that help close the gap between lab research and real-world treatments could be gained. This knowledge could improve how clinical trials are designed, making them more effective.
[0155] Leveraging a comprehensive dataset from electronic health records (EHR) and genomic profiling, two unique patient subsets with distinct clinical and genomic features were identified. A subset of mHSPC patients who are at a high-risk of progression or death on ASIs with tumors were more likely to have AR alterations in comparison to broader mHSPC patients is reported. Another subset of patients had a unique sequence for gain of AR alteration. These patients harbored tumors where AR with the ligand binding domain mutation was amplified, in contrast to the most of patients where there is an amplification of wild-type AR. 1. Introduction
[0156] The journey towards novel treatments for metastatic prostate cancer, particularly in the transition from preclinical models to clinical efficacy, is fraught with both significant challenges and untapped opportunities. Despite insightful preclinical hypotheses, such as exploring the interplay between androgen receptor (AR) signaling and the PI3K-Akt pathway in PTEN-loss tumors (Carver, B. S. et al., Cancer Cell 19, 575–586 (2011); Mulholland, D. J. et al., Cancer Cell 19, 792–804 (2011)) or investigating the induction of ‘a 'BRCAness' phenotype by ASIs (Sweeney, et al., Lancet 398 (10295), p131-142 (2021)), clinical translations have been less straightforward. This is exemplified in phase 3 trials like IPATential150 (Sweeney, et al., Lancet 398 (10295), p131-142 (2021)), PROpel (Saad, et al., Lancet Oncol. 24, 1094–1108 (2023)), and MAGNITUDE (Chi, et al., J. Clin. Oncol.40, 12–12 (2022)).
[0157] Clinical trials investigating the activity of targeted investigational drugs typically begin at the metastatic castration-resistant prostate cancer (mCRPC) stage, where patients have already undergone multiple rounds of effective therapies, including ASIs and taxanes. These trials, especially those targeting specific genetic alterations identified at diagnosis, are predicated on the assumption that such alterations will continue to significantly drive tumor growth and that these genetic changes will predict clinical responses in the castration-resistant stage. Similarly, it is assumed that mechanistic models, such as the interplay between AR and PI3K-Akt signaling developed from studies on homogeneous prostate cell lines, might be applicable in the heterogeneous setting of mCRPC7–9. However, the impact of tumor evolution and the accompanying heterogeneity in mCRPC on the genetic alterations identified atFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT diagnosis or on the mechanistic models remains unclear. This study aims to assess whether genetic alterations or mechanistic models, previously identified as associated with treatment outcomes, retain their relevance across both hormone-sensitive and castration-resistant stages of prostate cancer.
[0158] A comprehensive database that integrates electronic health records (EHR) with extensive genomic profiling data to clinically and genomically characterize tumors in patients with mHSPC and mCRPC was utilized. Subsequently, whether genetic alterations common to both settings exert a consistent influence on treatment outcomes with ASIs or taxanes across these stages was examined.
[0159] The objective of this study was to test whether the impact of genetic events identified at diagnosis on predicted clinical responses to ASIs / taxanes or reciprocal feedback regulation between PI3K-Akt and AR signaling may shift as the disease progresses from hormone-sensitive to the castration resistant stage. Consequently, further efforts are s needed for designing informed clinical trials and ensuring that therapeutic strategies are tailored to the evolving nature of the disease. 2. Patients and Methods
[0160] This study leveraged the nationwide (US-based) de-identified Flatiron Health-Foundation Medicine, Inc. (FH- FMI) clinico-genomic database (CGDB), encompassing data from around 800 care sites from 280 U.S. cancer clinics. The study employed a retrospective longitudinal analysis, extracting clinical data from electronic health records (EHRs). These records comprised both structured and unstructured patient-level information, which was meticulously curated through technology-enabled abstraction. Genomic data, obtained from comprehensive genomic profiling (CGP) tests by Foundation Medicine, Inc. (FMI) (of Cambridge, MA), were linked to the clinical records through de-identified deterministic matching, as described elsewhere (Singal, G. et al., JAMA 321, 1391–99 (2019)). The genomic analysis involved the identification of genomic alterations in over 300 cancer-associated genes, utilizing the next-generation sequencing (NGS)-based FoundationOne® or FoundationOne® CDx panels (Frampton et al., Nature Biotechnology 31, 1023-31 (2013); Woodhouse et al., PLOS ONE 15, e0237802 (2020)).
[0161] The FH-FMI metastatic prostate cancer database encompassed patients diagnosed with metastatic prostate cancer, as confirmed through chart review. Inclusion criteria stipulated at least two visits with the EHR network after 2011 and CGP testing by FMI following the diagnosis of prostate cancer. Critically, patients were also required to exhibit evidence of either castration-resistant or hormone-sensitive status, which was established by FH through abstraction of information available in patient charts and associated documents.
[0162] Given the study's dual focus on both prevalence assessment and the impact of specific biomarkers on treatment outcomes, two distinct cohort types were formed. The first one centered on mHSPC or mCRPC, facilitating the evaluation of prevalence rates. The second cohort comprised patients who initiated treatment within the respective clinical scenario, enabling the investigation of treatment outcomes. The first cohort was designated as the “main” cohort and the second cohort, which is a subset of the first, was designated as the “treatment” cohort.
[0163] Main Cohort: Two primary cohorts were established, each corresponding to mHSPC and mCRPC settings. To construct these cohorts, individuals within the CGDB who appeared in both the June 30, 2023 release and the March 31, 2023 release were initially identified (to ensure sufficient time to ascertain overall survival). Subsequently, only patients with FMI genomic baitsets (designed for liquid or tumor samples) and a CGP specimen that had undergone successful quality control were retained. The collection of these specimens occurred within a time windowFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT spanning 30 days before to 100 days after the respective setting's start date, whether it was mHSPC or mCRPC. Patients for whom the record showed a positive CRPC status but who were missing a CRPC date were excluded. Patients could be part of both the mHSPC and mCRPC cohort, with the same sample potentially meeting eligibility criteria for both settings. mHSPC patients who either progressed to mCRPC or succumbed to the disease within 100 days of the mHSPC setting date were defined as rapid progressors.
[0164] Treatment Cohort: From both of the main cohorts, treatment cohorts of patients who were alive, uncensored, and treated with systemic antineoplastic therapy within 90 days of the setting start (to facilitate a landmark analysis) were created. While Flatiron Health publishes algorithms to determine patient lines of therapy, these do not incorporate ADT. Therefore, treatment was assigned with a custom definition using structured fields within the EHR or from abstracted oral treatment episodes. All non-ADT therapies received within the 0-90 days after the setting start date were evaluated and classified into ADT (ADT was also assumed when no therapies were received), ASIs, and taxane. The non-ADT treatments received within the 0-90d after the mHSPC setting start date and prior to mCRPC setting start date were used in treatment classification in mHSPC patients. ADT could be given alongside ASIs and taxanes, but otherwise no combination therapies were allowed. Patients who did not receive one of these therapy classes were excluded from the treatment cohorts.
[0165] Relative AR variant Allele Frequency: Relative AR variant allelic frequency was calculated by first determining an expected maximal variant allelic frequency, to account for tumor purity, using Equation 2. Subsequently, the observed variant allele frequency was normalized against this expected variant allelic frequency for AR. Samples with a relative AR variant allele frequency greater than 0.5, whereby the majority of the tumor reads were from the variant, were classified as having the mutation occurring prior to the amplification. 3. Results
[0166] Patients whose tumors gain ligand binding mutations within AR before an amplification event may represent a distinct subset of metastatic prostate cancer. AR ligand binding (LBD) mutations are an acquired mechanism of resistance to ASIs20. The investigation of RWD genomic data identified three of five mHSPC patients with rapid progression had AR point mutations prior to ASI exposure, possibly in response to earlier generation of ASIs or ADT. This raises an interesting possibility that a subsequent amplification event could propagate these pre-existing mutations across all amplified AR gene copies. To investigate whether this event occurred, the occurrence of AR LBD mutations before the amplification event was explored.
[0167] In a comprehensive analysis of a research database encompassing 20,000 prostate cancer cases, genomic profiling was performed. The dataset included 11,378 local and 4,727 distant metastatic biopsies. A range of AR alterations was observed, with a notable clustering of mutations within the ligand-binding domain (FIG.7A). AR short variant mutations were present in 4.7% of cases, while amplifications were found in 15.1%. Interestingly, 1.1% of cases showed both short variant mutations and amplifications concurrently in the AR (FIG.7B).
[0168] To further understand the temporal sequence of these genomic events, the proportion of tumor AR reads containing mutations was assessed (FIG.7C). The analysis revealed that in the majority of instances, AR mutations were present in a minority of the total AR reads, suggesting that these mutations most likely occurred after amplification. However, data from 38.3% of cases indicate that the AR mutation may have arisen before the amplification (FIG.7D). In conclusion, these findings have delineated novel patient subsets with AR alterations where mutations precedeFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT amplification, potentially influencing differential responses to AR-targeted therapies compared to cases where amplification occurs first. 4. Conclusions
[0169] In conclusion, the study supports the notion that the association between genomic biomarkers, mechanistic hypotheses, and their impact on therapy response or resistance varies across different stages of prostate cancer. This variation may stem from increased tumor heterogeneity due to extensive anti-cancer treatments. Conversely, in earlier stages where tumors are more homogeneous, there might be a clearer link between genomic alterations and tumor behavior. A subset of patients in early stages, characterized by rapid progression or mortality could potentially benefit from more aggressive treatment strategies, such as the use of AR degraders (e.g., ARV-110). Example Clauses
[0170] The following clauses include various implementations of the present disclosure: 1. A method, including: determining, by analyzing sequence read data of a sample obtained from a subject, an observed allele fraction of a gene in the sample; determining an expected allele fraction based on: a modeled copy number of the gene in the sample, and a tumor purity of the sample; determining a ratio of the observed allele fraction to the expected allele fraction; and determining whether a mutation associated with the gene occurred prior to amplification by comparing the ratio to a threshold. 2. The method of clause 1, wherein the subject has prostate cancer. 3. The method of clause 1 or 2, wherein the subject has at least one of adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, a neuroblastoma, non-Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, a teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, or a vascular tumor. 4. The method of any of clauses 1 to 3, wherein the sample includes a tissue biopsy sample, a liquid biopsy sample, or a normal control. 5. The method of clause 4, wherein the sample is a liquid biopsy sample and includes blood, plasma, cerebrospinal fluid, sputum, stool, urine, lymphatic fluid, or saliva. 6. The method of clause 4 or 5, wherein the sample is a liquid biopsy sample and includes circulating tumor cells (CTCs). 7. The method of any of clauses 4 to 6, wherein the sample is a liquid biopsy sample and includes at least one of cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). 8. The method of any of clauses 1 to 7, wherein the gene is a x-linked gene 9. The method of any of clauses 1 to 8, wherein the gene is a y-linked gene. 10. The method of any of clauses 1 to 9, wherein the gene is an autosomal gene. 11. The method of any of clauses 1 to 10, wherein the gene includes at least one of an AR gene, ATM, TET2, DNMT3A, ASXL1, LYN, SF3B1, RB1, TP53, or SPOP.FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT 12. The method of any of clauses 1 to 11, wherein the observed allele fraction is an AR short variant allelic fraction. 13. The method of any of clauses 1 to 12, wherein determining, by analyzing the sequence read data, the observed allele fraction of the gene in the sample includes: determining an allelic fraction of the gene in the sample by analyzing the sequence read data. 14. The method of any of clauses 1 to 13, wherein determining the expected allele fraction based on the modeled copy number of the gene in the sample and the tumor purity of the sample includes calculating the expected allele fraction using the following equation:wherein CN is the modeled copy number of the gene in the sample, and TP is the tumor purity of the sample. 15. The method of any of clauses 1 to 14, wherein determining the expected allele fraction based on the modeled copy number of the gene in the sample and the tumor purity of the sample includes calculating the expected allele fraction using the following equation:wherein CNNis an expected number of copies of the gene in a non-cancerous sample, CN is the modeled copy number of the gene in the sample, and TP is the tumor purity of the sample. 16. The method of clause 15, wherein the gene is a sex-linked gene, and wherein the expected number of copies of the gene in the non-cancerous sample is 1.0. 17. The method of clause 15 or 16, wherein the gene is an autosomal gene, and wherein the expected number of copies of the gene in the non-cancerous sample is 2.0. 18. The method of any of clauses 1 to 17, further including: receiving, from a sequencer, the sequence read data of the sample. 19. The method of any of clauses 1 to 18, further including: determining the modeled copy number of the gene in the sample by: generating, based on the sequence read data, a major allele coverage ratio and a minor allele coverage ratio; segmenting one or more nucleic acid sequences associated with the sequence read data into segments; generating copy number grid model input features including: a sum of the major allele coverage ratio and the minor allele coverage ratio; and a difference of the major allele coverage ratio and the minor allele coverage ratio; fitting copy number grid models including allowed copy number states to the copy number grid model input features; selecting a copy number grid model among the copy number grid models; and assigning the modeled copy number for at least a portion of the one or more nucleic acid sequences based on the selected copy number grid model. 20. The method of any of clauses 1 to 19, further including: determining the tumor purity of the sample based on characteristics of nucleic acid molecules indicated by the sequence read data, the characteristics including at least one of an amount of somatic copy-number alterations (SCNA), an amount of single-nucleotide variants (SNVs), or a minor allele frequency (MAF). 21. The method of any of clauses 1 to 20, wherein the ratio is in a range of about 0.0 to about 1.0. 22. The method of any of clauses 1 to 21, wherein the ratio is in a range of about 0.2 to about 0.8.FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT 23. The method of any of clauses 1 to 22, further including: generating, based on whether the mutation associated with the gene occurred prior to the amplification, a genomic profile of the subject. 24. The method of clause 23, wherein the genomic profile includes results from at least one of: a comprehensive genomic profiling test; a whole genome sequencing (WGS) test; a whole exome sequencing (WES) test; a gene expression profiling test; a cancer hotspot panel test; a DNA methylation test; a DNA fragmentation test; or an RNA fragmentation test. 25. The method of any of clauses 1 to 24, further including: predicting that a pathological condition of the subject is resistant to a therapy based on whether the mutation associated with the gene occurred prior to the amplification; or predicting that the pathological condition of the subject is susceptible the therapy based on whether the mutation associated with the gene occurred prior to the amplification. 26. The method of clause 25, wherein the therapy includes at least one of an AR inhibitor, an AR degrader, an androgen deprivation therapy (ADT), an androgen signaling inhibitor (ASI), or systemic taxane. 27. The method of clause 25 or 26, wherein the therapy includes an AR-targeted therapy. 28. The method of any of clauses 25 to 27, wherein the therapy includes at least one of chemotherapy, radiation therapy, immunotherapy, or surgery. 29. The method of any of clauses 25 to 28, wherein the therapy is an anticancer therapy. 30. The method of any of clauses 25 to 29, further including: determining a dosage of the therapy based on whether the mutation associated with the gene occurred prior to the amplification. 31. The method of any of clauses 1 to 30, further including: receiving a plurality of nucleic acid molecules obtained from the sample; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules; capturing all or a subset of the amplified nucleic acid molecules; and sequencing, by a sequencer, the captured nucleic acid molecules to obtain a plurality of sequence reads that represent the captured nucleic acid molecules, thereby generating the sequence read data for a genome of the sample. 32. The method of clause 31, wherein the one or more adapters include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. 33. The method of clause 31 or 32, wherein the captured nucleic acid molecules are captured from the amplified nucleic acid molecules by hybridization to one or more bait molecules. 34. The method of clause 33, wherein the one or more bait molecules include one or more additional nucleic acid molecules, each of the one or more additional nucleic acid molecules including a region that is complementary to a region of a captured nucleic acid molecule. 35. The method of any of clauses 31 to 34, wherein amplifying the one or more ligated nucleic acid molecules includes performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique. 36. The method of any of clauses 31 to 35, wherein sequencing the captured nucleic acid molecules includes use of a massively parallel sequencing (MPS) technique, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing.FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT 37. The method of any of clauses 31 to 36, wherein sequencing the captured nucleic acid molecules includes next-generation sequencing (NGS). 38. The method of any of clauses 31 to 37, wherein sequencing the captured nucleic acid molecules includes sequencing-by-synthesis or nanopore sequencing. 39. The method of any of clauses 1 to 38, further including: generating ligated molecules by ligating adapters onto nucleic acid molecules of the sample; generating amplified ligated molecules by amplifying the ligated molecules; generating, using the amplified ligated molecules, detection signals; detecting, by at least one sensor, the detection signals; and generating the sequence read data based on the detection signals. 40. The method of clause 39, wherein generating, using the amplified ligated molecules, the detection signals includes: synthesizing, by a polymerase using fluorescently tagged nucleotide triphosphates (NTPs), a synthesized nucleic acid molecule that is complementary to one of the amplified ligated molecules, and wherein detecting, by the at least one sensor, the detection signals includes: detecting, by at least one optical sensor, optical signals emitted by the fluorescently tagged NTPs upon binding to the synthesized nucleic acid molecule, the optical signals being indicative of at least one sequence of the nucleic acid molecules of the sample. 41. The method of clause 39 or 40, wherein generating, using the amplified ligated molecules, the detection signals includes: directing the amplified ligated molecules through a nanopore extending from a first space to a second space through a substrate, and wherein detecting, by the at least one sensor, the detection signals includes: detecting, by sensors disposed in the first space and the second space, an electrical signal over time, the electrical signal being indicative of at least one sequence of the nucleic acid molecules of the sample. 42. The method of any of clauses 39 to 41, wherein the sequence read data indicates a predetermined panel of genes of the sample. 43. The method of clause 42, wherein the predetermined panel includes one or more of ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1,FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, or VEGFB. 44. The method of any of clauses 1 to 43, further including: determining, based on whether the mutation associated with the gene occurred prior to the amplification, whether the subject is eligible for a clinical trial. 45. A system, including: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including: determining, by analyzing sequence read data of a sample obtained from a subject, an observed allele fraction of a gene in the sample; determining an expected allele fraction based on: a modeled copy number of the gene in the sample, and a tumor purity of the sample; determining a ratio of the observed allele fraction to the expected allele fraction; and determining whether a mutation associated with the gene occurred prior to amplification by comparing the ratio to a threshold. 46. The system of clause 45, further including: a sequencer configured to generate the sequence read data by sequencing a plurality of nucleic acid molecules in the sample. 47. The system of clause 45 or 46, further including: a transceiver configured to transmit data indicating whether the mutation associated with the gene occurred prior to the amplification. 48. The system of any of clauses 45 to 47, further including: an output device configured to output an indication of whether the mutation associated with the gene occurred prior to the amplification. 49. A non-transitory computer readable medium storing instructions for performing operations including: determining, by analyzing sequence read data of a sample obtained from a subject, an observed allele fraction of a gene in the sample; determining an expected allele fraction based on: a modeled copy number of the gene in the sample, and a tumor purity of the sample; determining a ratio of the observed allele fraction to the expected allele fraction; and determining whether a mutation associated with the gene occurred prior to amplification by comparing the ratio to a threshold. 50. A method, including: providing a plurality of nucleic acid molecules obtained from a tumor sample of a male subject; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing, by a sequencer, all or a subset of the captured amplified nucleic acid molecules to obtain a plurality of sequence reads that represent the sequenced amplified nucleic acid molecules thereby generating sequence read data; receiving, at one or more processors, the sequence read data for the plurality of sequence reads; determining, by analyzing the sequence readFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT data using the one or more processors, an observed allele fraction of a x-linked gene in the tumor sample; determining, using the one or more processors, an expected allele fraction based on: a modeled copy number of the x-linked gene in the tumor sample, and a tumor purity of the tumor sample; determining, using the one or more processors, a ratio of the observed allele fraction to the expected allele fraction; determining, using the one or more processors, that a mutation associated with the x-linked gene occurred prior to amplification by determining that the ratio exceeds a threshold; and based on determining that the mutation associated with the x-linked gene occurred prior to the amplification, predicting that a cancer of the male subject is responsive to a predetermined therapy. 51. The method of clause 50, wherein the x-linked gene includes an androgen receptor (AR) gene. 52. The method of clause 51, wherein the observed allele fraction is an AR short variant allelic fraction. 53. The method of any of clauses 50 to 52, wherein determining the expected allele fraction based on the modeled copy number of the x-linked gene in the tumor sample and the tumor purity of the tumor sample includes calculating the expected allele fraction using the following equation:wherein CN is the modeled copy number of the x-linked gene in the tumor sample, and TP is the tumor purity of the tumor sample. 54. The method of any of clauses 50 to 53, wherein determining the expected allele fraction based on the modeled copy number of the gene in the sample and the tumor purity of the sample includes calculating the expected allele fraction using the following equation:wherein CNNis an expected number of copies of the gene in a non-cancerous sample, CN is the modeled copy number of the gene in the sample, and TP is the tumor purity of the sample. 55. The method of any of clauses 50 to 54, wherein the cancer of the male subject is prostate cancer. 56. The method of any of clauses 50 to 55, wherein the predetermined therapy includes an AR inhibitor and / or an AR degrader. Conclusion
[0171] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.
[0172] The features disclosed in the foregoing description, or the following claims, or the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for attaining the disclosed result, as appropriate, may, separately, or in any combination of such features, be used for realizing implementations of the disclosure in diverse forms thereof.
[0173] As will be understood by one of ordinary skill in the art, each implementation disclosed herein can comprise, consist essentially of or consist of its particular stated element, step, or component. Thus, the terms “include” or “including” should be interpreted to recite: “comprise, consist of, or consist essentially of.” The transition term “comprise” or “comprises” means has, but is not limited to, and allows for the inclusion of unspecified elements, steps, ingredients,FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT or components, even in major amounts. The transitional phrase “consisting of” excludes any element, step, ingredient or component not specified. The transition phrase “consisting essentially of” limits the scope of the implementation to the specified elements, steps, ingredients or components and to those that do not materially affect the implementation. As used herein, the term “based on” is equivalent to “based at least partly on,” unless otherwise specified.
[0174] Unless otherwise indicated, all numbers expressing quantities, properties, conditions, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. When further clarity is required, the term “about” has the meaning reasonably ascribed to it by a person skilled in the art when used in conjunction with a stated numerical value or range, i.e., denoting somewhat more or somewhat less than the stated value or range, to within a range of ±20% of the stated value; ±19% of the stated value; ±18% of the stated value; ±17% of the stated value; ±16% of the stated value; ±15% of the stated value; ±14% of the stated value; ±13% of the stated value; ±12% of the stated value; ±11% of the stated value; ±10% of the stated value; ±9% of the stated value; ±8% of the stated value; ±7% of the stated value; ±6% of the stated value; ±5% of the stated value; ±4% of the stated value; ±3% of the stated value; ±2% of the stated value; or ±1% of the stated value.
[0175] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. Any numerical value, however, inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements.
[0176] The terms “a,” “an,” “the,” and similar referents used in the context of describing implementations (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein is intended merely to better illuminate implementations of the disclosure and does not pose a limitation on the scope of the disclosure. No language in the specification should be construed as indicating any non-claimed element essential to the practice of implementations of the disclosure.
[0177] Groupings of alternative elements or implementations disclosed herein are not to be construed as limitations. Each group member may be referred to and claimed individually or in any combination with other members of the group or other elements found herein. It is anticipated that one or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT
[0178] Unless otherwise indicated, the practice of the present disclosure can employ conventional techniques of immunology, molecular biology, microbiology, cell biology and recombinant DNA. These methods are described in the following publications. See, e.g., Sambrook, et al. Molecular Cloning: A LaboratoryManual, 2nd Edition (1989); F. M. Ausubel, et al. eds., Current Protocols in Molecular Biology, (1987); the series Methods IN Enzymology (Academic Press, Inc.); M. MacPherson, et al., PCR: A Practical Approach, IRL Press at Oxford University Press (1991); MacPherson et al., eds. PCR 2: Practical Approach, (1995); Harlow and Lane, eds. Antibodies, A Laboratory Manual, (1988); and R. I. Freshney, ed. Animal Cell Culture (1987).
[0179] Tumor mutational burden (TMB) is a measure of the number of mutations carried by tumor cells. By comparing DNA sequences from a patient’s healthy tissues and tumor cells, the number of acquired somatic mutations present in tumors, but not in normal tissues, may be determined. In some instances, driver mutations may be excluded from a TMB calculation.
[0180] In certain examples, "tumor mutational burden" or “TMB” refers to the number of somatic mutations in a tumor's genome and / or the number of somatic mutations per area of the tumor's genome. In some embodiments, TMB, as used herein, refers to the number of somatic mutations per megabase (Mb) of DNA sequenced. In some embodiments, germline (inherited) variants are excluded when determining TMB, given that the immune system has a higher likelihood of recognizing these as self. In various cases, driver mutations are excluded from a TMB calculation.
[0181] Microsatellites are highly polymorphic DNA-repeat regions. In certain examples, “microsatellite” refers to a repetitive nucleic acid having repeat units of less than 10 base pairs or nucleotides in length. In certain examples, a microsatellite refers to a tract of tandemly repeated (i.e. adjacent) DNA motifs ranging from one to six or up to ten nucleotides, with each motif repeated 5 to 50 repeated times. “Microsatellite instability” refers to genetic instability in the microsatellite regions. Cancer patients with microsatellite instability classified as being high (MSI-H or MSI-High) frequently exhibit an accumulation of somatic mutations in tumor cells that leads to a range of molecular and biological changes including high tumor mutational burden, increased expression of neoantigens and abundant tumor-infiltrating lymphocytes. Chang et al. “Microsatellite Instability: A Predictive Biomarker for Cancer Immunotherapy,” Appl Immunohistochem Mol Morphol, 26(2):e15-e21 (2018). These changes have been linked to increased sensitivity to checkpoint inhibitor drugs, such as pembrolizumab, which is used to treat advanced melanoma, head and neck squamous cell carcinoma, non-small cell lung cancer (NSCLC), and classical Hodgkin lymphoma.
[0182] A viral status test refers to a test that identifies the presence of viral RNA or DNA in a subject. The test can identify viral load and / or viral identity. For example, the viral status test can identify the presence of viral RNA or DNA associated with the occurrence of certain cancers. Examples of such viruses include Hepatitis B Virus (HBV) and Hepatitis C Virus (HCV), Kaposi Sarcoma-Associated Herpesvirus (KSHV), Merkel Cell Polyomavirus (MCV), Human Papillomavirus (HPV), Human Immunodeficiency Virus Type 1 (HIV-1, or HIV), Human T-Cell Lymphotropic Virus Type 1 (HTLV-1), and Epstein-Barr Virus (EBV).
[0183] Cancer “hotspot” mutations give rise to oncological outcomes. PhyloP, SIFT, Grantham, COSMIC and PolyPhen-2 are in silico tools that can be used to assess pathogenicity of identified variants. Exemplary hotspot genes and mutations include EGFR exon 19 activating mutation, EGFR exon 19 deletion, EGFR exon 19 insertion, EGFR exon 19 sensitizing mutation, EGFR exon 20 activation mutation, EGFR exon 20 insertion, EGFR G719 mutation, EGFR L858R mutation, EGFR L861 mutation, EGFR S768 mutation, EGFR T790M mutation, C797 mutation, KITFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT activating mutation, KRAS activating mutation, MET activating mutation, NRAS activating mutation, PMS2 promoter mutations, among many others. Hotspot mutations also occur in the following genes: AKT2, BRCA1, BRCA2, ERC1, NSD1, POLH, PPM1G, PTEN, RAD18, RAD51, RAD51B, RB1, TERT, TP53, TP53Bp1, ALK, ARMT1, ATAD5, ATG7, ATIC, AXL, BIRC6, BRD3, BRD4, CAPRIN1, CCAR2, CCDC6, CDK5RAP2, CHD9, CIT, CTNNB1, CUL1, EBF1, EIF3E, HIP1, HMGA2, IRF2BP2, NOTCH1, NOTCH4, NPM1, OFD1, TACC1, TACC3, TERF2, TMEM106B, UBE2L3, USP10, WRDR48, YAP1, ZEB2, and ZMYND8.
[0184] A “DNA methylation test” refers to an assay, which can be commercially available, for distinguishing methylated versus unmethylated cytosine loci in DNA. Techniques for measuring cytosine methylation include bisulfite- based methylation assays. The addition of bisulfite to DNA results in the methylation of unmethylated cytosine and its ultimate conversion to the nucleotide uracil. Uracil has similar binding properties to thiamine in the DNA sequence. Previously methylated cytosine does not undergo similar chemical conversion on exposure to bisulfite. Bisulfite assays can thus be used to discriminate previously methylated versus unmethylated cytosine.
[0185] An exemplary quantitative methylation detection assay combines bisulfite treatment and restriction analysis COBRA, which uses methylation sensitive restriction endonucleases, gel electrophoresis, and detection based on labeled hybridization probes. (Ziong and Laird, Nucleic Acid Res.199725; 2532-4). Another exemplary detection assay is the methylation specific polymerase chain reaction PCR (MSPCR) for amplification of DNA segments of interest. This assay can be performed after sodium bisulfite conversion of cytosine and uses methylation sensitive probes. Other detection assays include the Quantitative Methylation (QM) assay, which combines PCR amplification with fluorescent probes designed to bind to putative methylation sites; MethyLightTM(Qiagen, Redwood City, CA) a quantitative methylation detection assay that uses fluorescence-based PCR (Eads, et al., Cancer Res.1999; 59:2302- 2306); and Ms-SNuPE, a quantitative technique for determining differences in methylation levels in CpG sites. As with other techniques, Ms-SNuPE also requires bisulfite treatment to be performed first, leading to the conversion of unmethylated cytosine to uracil while methyl cytosine is unaffected. PCR primers specific for bisulfite converted DNA are then used to amplify the target sequence of interest. The amplified PCR product is isolated and used to quantitate the methylation status of the CpG site of interest. (Gonzalgo and Jones Nuclei Acids Res1997; 25:252-31).
[0186] In particular embodiments, pyrosequencing can be used to detect marker methylation. Pyrosequencing is a method of DNA sequencing that relies on detection of the release of pyrophosphates as DNA is synthesized (and is therefore a “sequencing by synthesis” technique). To assess methylation by pyrosequencing, a DNA sample can be incubated with sodium bisulfite, converting unmethylated cytosine to uracil. The presence of uracil will result in thymine incorporation during PCR amplification. Therefore, sequencing results that include thymine at a nucleotide position that is known to encode cytosine can be interpreted as unmethylated sites. In contrast cytosines present in the sequencing results indicate that the site was methylated in the original DNA sample, because methylation protects cytosine from conversion to uracil upon treatment. Bisulfite treatment can also be performed on control samples with known methylation patterns, to reduce or eliminate false positive results. Commercially available pyrosequencing machines include Pyro Mark Q96 (Qiagen, Hilden, Germany). For more details on methods to use pyrosequencing for measurement of methylation, see Delaney et al. Methods Mol Biol.20151343: 249-264. Pyrosequencing is especially useful for detecting methylation in the CpG sites within genes.FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT
[0187] In particular embodiments, a protein marker is detected by contacting a sample with reagents (e.g., antibodies), generating complexes of reagent and marker(s), and detecting the complexes. Particular embodiments for detecting and measuring protein levels can use methods including agglutination, chemiluminescence, electro- chemiluminescence (ECL), enzyme-linked immunoassays (ELISA), immunoassay, immunoblotting, immunodiffusion, immunoelectrophoresis, immunofluorescence, immunohistochemistry, immunoprecipitation, mass-spectrometry, and western blot. See also, e.g., E. Maggio, Enzyme-Immunoassay (1980), CRC Press, Inc., Boca Raton, Fla; and U.S. Pat. Nos.4,727,022; 4,659,678; 4,376,110; 4,275,149; 4,233,402; and 4,230,797.
[0188] Read depth refers to the number of times that a specific genomic site is sequenced during a sequencing run.
[0189] Certain implementations are described herein, including the best mode known to the inventors for carrying out implementations of the disclosure. Of course, variations on these described implementations will become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventor expects skilled artisans to employ such variations as appropriate, and the inventors intend for implementations to be practiced otherwise than specifically described herein. Accordingly, the scope of this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by implementations of the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
Claims
FMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT CLAIMS What is claimed is:
1. A method, comprising: determining, by analyzing sequence read data of a sample obtained from a subject, an observed allele fraction of a gene in the sample; determining an expected allele fraction based on: a modeled copy number of the gene in the sample, and a tumor purity of the sample; determining a ratio of the observed allele fraction to the expected allele fraction; and determining whether a mutation associated with the gene occurred prior to amplification based on comparing the ratio to a threshold.
2. The method of claim 1, wherein the subject has prostate cancer.
3. The method of claim 1, wherein the gene is a x-linked gene 4. The method of claim 1, wherein the gene is a y-linked gene.
5. The method of claim 1, wherein the gene comprises at least one of an androgen receptor (AR) gene, ATM, TET2, DNMT3A, ASXL1, LYN, SF3B1, RB1, TP53, or SPOP.
6. The method of claim 1, wherein the observed allele fraction is an AR short variant allelic fraction.
7. The method of claim 1, wherein determining the expected allele fraction based on the modeled copy number of the gene in the sample and the tumor purity of the sample comprises calculating the expected allele fraction using the following equation:wherein CN is the modeled copy number of the gene in the sample, and TP is the tumor purity of the sample.
8. The method of claim 1, wherein determining the expected allele fraction based on the modeled copy number of the gene in the sample and the tumor purity of the sample comprises calculating the expected allele fraction using the following equation:wherein CNN is an expected number of copies of the gene in a non-cancerous sample, CN is the modeled copy number of the gene in the sample, and TP is the tumor purity of the sample.
9. The method of claim 8, wherein the gene is a sex-linked gene, and wherein the expected number of copies of the gene in the non-cancerous sample is 1.
0.
10. The method of claim 8, wherein the gene is an autosomal gene, and wherein the expected number of copies of the gene in the non-cancerous sample is 2.
0.
11. The method of claim 1, wherein the ratio is in a range of about 0.2 to about 0.
8.
12. The method of claim 1, further comprising: predicting that a pathological condition of the subject is resistant to a therapy based on whether the mutation associated with the gene occurred prior to the amplification; orFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT predicting that the pathological condition of the subject is susceptible the therapy based on whether the mutation associated with the gene occurred prior to the amplification, wherein the therapy comprises at least one of an AR inhibitor, an AR degrader, an androgen deprivation therapy (ADT), an androgen signaling inhibitor (ASI), or systemic taxane.
13. A system, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: determining, by analyzing sequence read data of a sample obtained from a subject, an observed allele fraction of a gene in the sample; determining an expected allele fraction based on: a modeled copy number of the gene in the sample, and a tumor purity of the sample; determining a ratio of the observed allele fraction to the expected allele fraction; and determining whether a mutation associated with the gene occurred prior to amplification by comparing the ratio to a threshold.
14. The system of claim 13, further comprising: a sequencer configured to generate the sequence read data by sequencing a plurality of nucleic acid molecules in the sample.
15. A method, comprising: providing a plurality of nucleic acid molecules obtained from a tumor sample of a male subject; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing, by a sequencer, all or a subset of the captured amplified nucleic acid molecules to obtain a plurality of sequence reads that represent the sequenced amplified nucleic acid molecules thereby generating sequence read data; receiving, at one or more processors, the sequence read data for the plurality of sequence reads; determining, by analyzing the sequence read data using the one or more processors, an observed allele fraction of a x-linked gene in the tumor sample; determining, using the one or more processors, an expected allele fraction based on: a modeled copy number of the x-linked gene in the tumor sample, and a tumor purity of the tumor sample; determining, using the one or more processors, a ratio of the observed allele fraction to the expected allele fraction; determining, using the one or more processors, that a mutation associated with the x-linked gene occurred prior to amplification by determining that the ratio exceeds a threshold; andFMI Docket No.: 0186-WO / 0139-CG L&H Docket No.: F171-6004PCT based on determining that the mutation associated with the x-linked gene occurred prior to the amplification, predicting that a cancer of the male subject is responsive to a predetermined therapy.
16. The method of claim 15, wherein the x-linked gene comprises an AR gene, and wherein the observed allele fraction is an AR short variant allelic fraction.
17. The method of claim 15, wherein determining the expected allele fraction based on the modeled copy number of the x-linked gene in the tumor sample and the tumor purity of the tumor sample comprises calculating the expected allele fraction using the following equation:wherein CN is the modeled copy number of the x-linked gene in the tumor sample, and TP is the tumor purity of the tumor sample.
18. The method of claim 15, wherein determining the expected allele fraction based on the modeled copy number of the gene in the sample and the tumor purity of the sample comprises calculating the expected allele fraction using the following equation:wherein CNNis an expected number of copies of the gene in a non-cancerous sample, CN is the modeled copy number of the gene in the sample, and TP is the tumor purity of the sample.
19. The method of claim 15, wherein the cancer of the male subject is prostate cancer.
20. The method of claim 15, wherein the predetermined therapy comprises an AR inhibitor and / or an AR degrader.
Citation Information
Patent Citations
Biomarkers For Prostate Cancer
US20080181850A1
Biomarkers indicative of prostate cancer and treatment thereof
US20230133972A1
Systems and methods for evaluating tumor fraction
WO2022271159A1