Loss-of-function calculation model based on allele frequency
Competitive probability models are generated through computer systems and germline single nucleotide polymorphic site modeling is used to solve the diagnostic problems of homozygous and heterozygous deletion of somatic cells in humoral fluids, achieving higher cancer detection accuracy and personalized therapeutic support.
Patent Information
- Application Number
- CN202080031940.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-25
- Filing Date
- 2020-02-27
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-02-27
AI Technical Summary
The prior art is difficult to efficiently distinguish and diagnose somatic homozygous deletion and somatic heterozygous deletion in bodily fluids, resulting in insufficient accuracy and sensitivity of cancer detection.
A competitive probability model was generated using a computer system, and the probability of somatic homozygous deletion and somatic heterozygous deletion were calculated based on germline single nucleotide polymorphism sites, and the log-likelihood ratio was used to determine the sample state.
It improves the diagnostic accuracy and sensitivity of somatic cell deletion status, can more accurately identify cancer-related gene status, and supports personalized treatment decisions.
Smart Images

Figure CN113748467B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. Provisional Application No. 62 / 811,159, filed February 27, 2019, and U.S. Provisional Application No. 62 / 823,585, filed March 25, 2019, which are incorporated herein by reference for all purposes.
[0003] background
[0004] A tumor is an abnormal growth of cells. When cells (such as tumor cells) die, fragmented DNA is often released into body fluids. Therefore, some of the cell-free DNA in body fluids is tumor DNA. Tumors can be benign or malignant. Malignant tumors are often called cancers.
[0005] Cancer is a leading cause of disease worldwide. Every year, tens of millions of people are diagnosed with cancer worldwide, and more than half of those diagnosed die from it. In many countries, cancer ranks as the second most common cause of death after cardiovascular disease. Early detection is associated with improved outcomes for many cancers.
[0006] Cancer is caused by the accumulation of mutations and / or epigenetic variations within an individual's normal cells, at least some of which lead to misregulation of cell division. Such mutations, or states of genetic material, typically include copy number variations (CNVs), copy number aberrations (CNAs), single nucleotide variants (SNVs), gene fusions, and indels, and epigenetic variations include modifications to the fifth atom of the cytosine 6-atom ring and the association of DNA with chromatin and transcription factors.
[0007] In specific instances, loss of heterozygosity (LOH) and biallelic copy number loss of homologous recombination repair (HRR) genes (BRCA1 / 2) are associated with loss of function in tumor suppressors, leading to cancer. In many cases, the specific state of a gene of interest can inform the type of treatment. For example, one state of a gene may be responsive to one group of drugs, while another state of the gene may not be responsive. Therefore, it is increasingly important to be able to not only diagnose cancer and other diseases, but also to characterize the underlying causes of the disease.
[0008] Cancer is typically detected by biopsy of the tumor and then analyzing cells, markers, or DNA extracted from the cells. Research into cancer detection based on analysis of body fluids is ongoing. If successful, these tests have the advantage that they are non-invasive and can be performed without identifying suspected cancer cells through a biopsy. However, the fact that the amount of nucleic acid in body fluids is very low makes it complicated to successfully perform these types of tests. In addition, the amount of detectable tumor-associated cell-free nucleic acid in body fluids may also make the analysis and detection of cancer in cell-free DNA difficult. In other words, tumor DNA in body fluids may be contaminated with normal DNA, making it difficult to computationally analyze and detect the specific cause of the tumor in the cell-free DNA sample.
[0009] Overview
[0010] The present disclosure relates to computer technology for providing accurate diagnosis of various states of genetic material (such as genes sequenced from cell-free DNA in a sample). The state may include the mutation state of the gene, such as but not limited to somatic homozygous deletion, somatic heterozygous deletion, copy number variation ("CNV") (including specific copy number wild type, amplification or loss) and / or other states. Accurate diagnosis can be based on one or more probability models of the state. For example, a computer system can generate competitive models, and each model outputs the probability that the genetic material is in a certain state.
[0011] Each model can be trained on a sample training set to output the probability that the genetic material is in a corresponding state. For example, a first model can be related to and output a first probability that the genetic material contains a somatic homozygous deletion of alleles of a specific gene. A second model can be related to and output a second probability that the genetic material contains a somatic heterozygous deletion of alleles of a specific gene. Other models can be related to and output the probability of other types of states of the genetic material, such as CNV. The computer system can compare the outputs of each competing model to determine which is more likely. For example, the computer system can use the log-likelihood ratio of the competing first and second probabilities to determine whether the genetic material includes a somatic homozygous deletion or a somatic heterozygous deletion.
[0012] In some embodiments, the computer system can use various probability distributions to generate models.For example, the computer system can use beta-binomial distribution, binomial distribution, normal (also referred to as " Gaussian ") distribution and / or other types of probability modeling techniques.The computer system can model a state (such as the allele count supporting a particular state) based on a training data set to set a baseline expectation for an abnormal or tumor state.For example, the computer system can identify germline single nucleotide polymorphisms (SNPs) positions observed in "normal" or non-tumor samples, such as samples in which somatic variants are not observed.These samples will also be referred to as tumor not detected (TND) samples.
[0013] Because TND samples are normal, the computer system can assume that the germline SNP position will not cause an abnormal state. Like this, the computer system can utilize these SNP sites as reference expected values for modeling allele counts to determine the probability of a state. For example, the deviation from the nucleotide calls observed at each SNP position can indicate the probability that this deviation causes a specific state such as a tumor or other abnormal state. The computer system can accordingly train the model based on the expected value calculated from the data of the germline SNPs derived from the TND sample. For each SNP site, such calculated data can include: the prevalence rate of heterozygosity, MAF standard deviation, genotype, germline prevalence rate (a priori) and / or other data that can inform individual sample analysis.
[0014] Utilizing the expected value calculated, the computer system can model state based on the sequence reads of the tested individual sample aligned with the region of interest (such as upstream, downstream and the region including gene of interest).In some embodiments, the molecular sequence reads produced from the individual sample can be compared with the reference genome to identify the allele (mutant or wild type) supported by potential molecules.Based on the comparison of the sequence reads produced from the individual sample, the computer system can identify the number of molecules supporting alternative alleles, and calculate the total number of molecules.The computer system can model these and / or other data from the individual sample with the expected value data calculated from each germline SNP in the region of interest.In some instances, order-checking can be the targeted order-checking based on plasma cell-free DNA (cfDNA).
[0015] In one aspect, the present disclosure relates to a computer system that is improved to distinguish somatic homozygous deletions and somatic heterozygous deletions of genes in samples that do not show germline deletions of genes. The computer system may include a processor that is programmed to: via a first probability distribution, a first model of allele counts is generated based on one or more germline single nucleotide polymorphisms (SNPs) positions related to the gene, and the first model represents somatic homozygous deletions. The processor can also, via a second probability distribution, generate a second model of allele counts in the sample based on one or more germline SNP positions, and the second model represents somatic heterozygous deletions. The processor can compare the first output of the first model and the second output of the second model. Based on the comparison, the processor can generate a prediction of the somatic homozygous deletion of genes in the sample.
[0016] In some embodiments, the first model can represent a first probability that a sample includes a somatic homozygous deletion, and the second model represents a second probability that a sample includes a somatic heterozygous deletion.
[0017] In some embodiments, the first probability distribution is the same type of probability distribution as the second probability distribution.
[0018] In some embodiments, to generate the first model, the processor is programmed to determine one or more parameters that are input to the first probability distribution.
[0019] In some embodiments, the first probability distribution comprises a beta-binomial distribution, a binomial distribution, or a normal distribution.
[0020] In some embodiments, to generate the first model of allele counts, the processor may also determine the prevalence of heterozygosity for one or more germline SNPs in the training set of samples for input into the first probability distribution.
[0021] In some embodiments, the sample training set may include more than one sample in which tumor was not detected (TND).
[0022] In some embodiments, to generate the first model of allele counts, the processor may also determine the standard deviation of the minor allele frequency (MAF) associated with each of the one or more germline SNPs in the sample training set for input to the first probability distribution.
[0023] In some embodiments, to generate the first model, the processor may also determine the number of molecules in the sample that support the mutant allele for input to the first probability distribution.
[0024] In some embodiments, to generate the first model, the processor may also determine the total number of molecules in the sample for input to the first probability distribution.
[0025] In some embodiments, to generate the first model, the processor may further calculate a first likelihood of allele counts for one or more germline SNP positions in the sample assuming a somatic homozygous deletion based on sequence read coverage associated with the somatic homozygous deletion.
[0026] In some embodiments, to generate the second model, the processor may further calculate a second likelihood of allele counts for one or more germline SNP positions in the sample hypothesized somatic heterozygous deletion based on sequence read coverage associated with the somatic heterozygous deletion.
[0027] In some embodiments, to generate the second model, the processor may also determine a mean of the tumor fractions estimated from the samples for input to the second probability distribution of the second model.
[0028] In some embodiments, the tumor fraction can be estimated based on sequence coverage information.
[0029] In some embodiments, to generate the second model, the processor may also determine a standard deviation of the tumor fraction estimated from the sample for input into the second probability distribution of the second model.
[0030] In some embodiments, the processor can also read more than one sample, identify a sample group that includes a germline deletion from the more than one sample, filter out the sample group from the more than one sample, and identify the presence of a somatic homozygous deletion or a somatic heterozygous deletion from the filtered more than one sample.
[0031] In some embodiments, the first output may include a first probability that a somatic homozygous deletion is present, and the second output may include a second probability that a somatic heterozygous deletion is present.
[0032] In some embodiments, to compare the first output of the first model and the second output of the second model, the processor may further perform a log-likelihood function based on the first output and the second output.
[0033] In some embodiments, the gene may include BRCA1, BRCA2, or ATM.
[0034] In another aspect, the present disclosure relates to a system that can include a processor programmed to generate a first probability that a gene in a sample contains a somatic homozygous deletion, generate a second probability that the gene in the sample contains a somatic heterozygous deletion, compare the first probability and the second probability, and generate a prediction of whether the sample contains a somatic homozygous deletion or a somatic heterozygous deletion.
[0035] In another aspect, the present disclosure relates to a system that can include a processor programmed to generate a first probability that genetic material in a sample contains a first state, generate a second probability that genetic material in the sample contains a second state, compare the first and second probabilities, and generate a prediction of whether the sample contains the first state or the second state.
[0036] In some embodiments, the first state comprises a somatic homozygous deletion and the second state comprises a somatic heterozygous deletion.
[0037] In some embodiments, the first state may include a first copy number variation (CNV), and the second state may include a second CNV that is different from the first CNV.
[0038] In some embodiments, the first CNV and / or the second CNV can be associated with a deleterious condition.
[0039] In some embodiments, to generate the first probability, the processor can also read one or more germline single nucleotide polymorphism (SNP) positions associated with the gene and determine the standard deviation of the minor allele frequency (MAF) associated with each of the one or more germline SNPs in the sample training set.
[0040] In some embodiments, to generate the first probability, the processor may also determine the standard deviation of the minor allele frequency (MAF) associated with each of the one or more germline SNPs in the sample training set for input to the probability distribution.
[0041] On the other hand, the present disclosure relates to a method implemented by a processor. The method may include, by the processor, via a first probability distribution, generating a first model of allele counts based on one or more germline single nucleotide polymorphism (SNP) positions associated with the gene, the first model representing somatic homozygous deletions. The method may also include, by the processor, via a second probability distribution, generating a second model of allele counts in the sample based on one or more germline SNP positions, the second model representing somatic heterozygous deletions. The method may include, by the processor, comparing the first output of the first model and the second output of the second model. The method may also include, by the processor, based on the comparison, generating a prediction of the presence of somatic homozygous deletions of the gene in the sample.
[0042] In another aspect, the present disclosure relates to another method implemented by a processor. The method may include generating, by the processor, a first probability that a gene in a sample contains a somatic homozygous deletion. The method may also include generating, by the processor, a second probability that the gene in the sample contains a somatic heterozygous deletion. The method may also include comparing, by the processor, the first probability and the second probability. The method may also include generating, by the processor, a prediction of whether the sample contains a somatic homozygous deletion or a somatic heterozygous deletion.
[0043] In another aspect, the present disclosure relates to another method implemented by a processor.
[0044] The method may include generating, by the processor, a first probability that the genetic material in the sample contains a first state. The method may also include generating, by the processor, a second probability that the genetic material in the sample contains a second state. The method may also include comparing, by the processor, the first probability and the second probability. The method may also include generating, by the processor, a prediction of whether the sample contains the first state or the second state.
[0045] In another aspect, the present disclosure relates to a method for administering to a subject determined to have a somatic homozygous deletion based on the disclosure herein a therapeutic intervention effective for treating a cancer associated with the somatic homozygous deletion.
[0046] In some embodiments, therapeutic intervention may include poly ADP ribose polymerase (PARP) inhibitors. Examples of PARP inhibitors include OLAPARIB, TALAZOPARIB, RUCAPARIB, NIRAPARIB (trade name ZEJULA), and the like.
[0047] In some embodiments, therapeutic intervention can include a base excision repair (BER) inhibitor. For example, OLAPARIB can inhibit BER.
[0048] In another aspect, the present disclosure relates to a method for administering to a subject determined to have a particular state of genetic material based on the present disclosure a therapeutic intervention effective for treating a disease associated with that state of genetic material.
[0049] In another aspect, the present disclosure relates to a method for administering a therapeutic intervention other than a PARP inhibitor to a subject determined based on the present disclosure to not have a somatic homozygous deletion.
[0050] In some embodiments of various aspects of the present disclosure, the result of system and / or method disclosed herein is used as input to generate a report.The report can be in paper or electronic format.For example, the information about the disappearance or other state of gene and / or genetic material determined by method disclosed herein or system and / or the information derived therefrom can be presented in such a report.Method disclosed herein or system can also include that the report is transmitted to a third party, such as an experimenter or a health care practitioner who obtains a sample from it.
[0051] The various operations of the methods disclosed herein, or the operations performed by the systems disclosed herein, can be performed at the same time or at different times, and / or in the same geographical location or in different geographical locations (e.g., countries). The various steps of the methods disclosed herein can be performed by the same person or by different people. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 An example of a system for training a model to predict the state of genetic material based on the probability of each state is shown, according to an embodiment of the present disclosure.
[0054] Figure 2 Schematic diagram showing determination of allele counts of germline SNPs to predict genetic status according to embodiments of the present disclosure.
[0055] Figure 3 A process of predicting somatic homozygous deletions or somatic heterozygous deletions based on a trained model according to an embodiment of the present disclosure is shown.
[0056] Figure 4A process for predicting the state of genetic material based on a trained model according to an embodiment of the present disclosure is shown.
[0057] Figure 5 Shown are types of somatic deletions according to embodiments of the present disclosure.
[0058] Figure 6A An example diagram showing a homozygous deletion of BRCA1 according to an embodiment of the present disclosure.
[0059] Figure 6B An example diagram showing a BRCA2 heterozygous deletion according to an embodiment of the present disclosure.
[0060] Figure 7A An example graph showing the prevalence of hets in TND samples according to an embodiment of the present disclosure.
[0061] Figure 7B An example graph showing MAF across TND samples according to an embodiment of the present disclosure.
[0062] Figure 8A An example graph showing MAF values for BRCA1 according to embodiments of the present disclosure.
[0063] Figure 8B An example graph showing MAF values for BRCA2 according to embodiments of the present disclosure.
[0064] Figure 9A An example graph showing a comparison of scores between the beta-binomial model and the binomial model for the BRCA2 panel, according to embodiments of the present disclosure.
[0065] Figure 9B An example graph showing a comparison of scores between a beta-binomial model and a Gaussian model for the BRCA2 panel, according to embodiments of the present disclosure.
[0066] Figure 10A An example graph showing the distribution of LLR scores for BRCA1 negative samples according to embodiments of the present disclosure.
[0067] Figure 10B An example graph showing the distribution of LLR scores for BRCA2-negative samples according to embodiments of the present disclosure.
[0068] Figure 11A An example graph showing the limit of detection (LoD) loss for BRCA1 according to embodiments of the present disclosure.
[0069] Figure 11B An example diagram showing a homozygous deletion of the LoD HRR of BRCA1 according to an embodiment of the present disclosure.
[0070] Figure 12A An example diagram showing LoD deletions for BRCA2 according to embodiments of the present disclosure.
[0071] Figure 12B An example diagram showing a homozygous deletion of the LoD HRR of BRCA2 according to an embodiment of the present disclosure.
[0072] Figure 13 An example graph showing TF prevalence versus cancer type according to embodiments of the present disclosure is shown.
[0073] Figure 14 An example graph showing LLR score density for BRCA1 and BRCA2 according to embodiments of the present disclosure.
[0074] Figure 15 An example graph showing the prevalence of BRCA2 homozygous deletions according to embodiments of the present disclosure.
[0075] Figure 16 An example graph showing the prevalence of BRCA1 homozygous deletions according to embodiments of the present disclosure.
[0076] Figure 17 An example of a BRCA2 homozygous deletion and potential clinical actionability according to embodiments of the present disclosure is shown.
[0077] Figure 18A An example diagram showing a homozygous deletion of BRCA1 according to an embodiment of the present disclosure.
[0078] Figure 18B An example diagram showing a homozygous deletion of BRCA1 according to an embodiment of the present disclosure.
[0079] Figure 19A An example diagram showing a BRCA2 homozygous deletion according to an embodiment of the present disclosure.
[0080] Figure 19B An example diagram showing a BRCA2 homozygous deletion according to an embodiment of the present disclosure.
[0081] Figure 20A An example diagram illustrating biallelic somatic copy number loss of BRCA1 according to embodiments of the present disclosure.
[0082] Figure 20B An example diagram showing BRCA1 LOH according to an embodiment of the present disclosure is shown.
[0083] Figure 21A An example diagram illustrating BRCA2 biallelic somatic copy number loss according to embodiments of the present disclosure.
[0084] Figure 21B An example diagram of BRCA2 LOH is shown, according to an embodiment of the present disclosure.
[0085] Figure 22 Graph showing the prevalence of BRCA1 and BRCA2 somatic deletions according to embodiments of the present disclosure.
[0086] definition
[0087] A subject refers to an animal, such as a mammalian species (preferably a human) or an avian (e.g., bird) species, or other organisms, such as a plant. More specifically, a subject can be a vertebrate, for example a mammal, such as a mouse, a primate, a monkey, or a human. Animals include farm animals, sport animals, and pets. A subject can be a healthy individual, an individual with symptoms or signs, or an individual suspected of having a disease or disease predisposition, or an individual in need of treatment or suspected of needing treatment.
[0088] Genetic variant refers to the change, variant or polymorphism in the nucleic acid sample or genome of the experimenter.Such change, variant or polymorphism can be relative to a reference genome, and the reference genome can be a species, experimenter or other individual reference genome (for example, hG19 or hG38 for the mankind).Variation includes one or more single nucleotide variation (SNV), insertion, deletion, repetition, small insertion, small deletion, small repetition, structural variant connection, variable length tandem repeat and / or flanking sequence, copy number variation (CNV), transversion, gene fusion and other rearrangements are also forms of genetic variation.Variation can be base change, insertion, deletion, repetition, copy number variation, transversion or its combination.
[0089] Cancer markers are genetic variants that are associated with the presence of cancer or the risk of developing cancer. Cancer markers can provide an indication that a subject has cancer or is at a higher risk of developing cancer than an age- and sex-matched subject of the same species who does not have the cancer marker. A cancer marker may or may not be the cause of the cancer.
[0090] As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides, about 100 nucleotides, about 50 nucleotides, or about 10 nucleotides in length) that is used to distinguish nucleic acids from different samples (e.g., representing a sample index) or different nucleic acid molecules of different types or that have been differently treated in the same sample (e.g., representing a molecular barcode). Nucleic acid tags include predetermined, fixed, non-random, random, or semi-random oligonucleotide sequences. These nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags optionally have the same length or different lengths. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, including 5' or 3' single-stranded regions (e.g., overhangs), and / or include one or more other single-stranded regions at other positions within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information, such as the sample source, form, or treatment of a given nucleic acid. For example, nucleic acid tags can also be used to enable the collection and / or parallel processing of more than one sample comprising nucleic acids with different molecular barcodes and / or sample indices, wherein the nucleic acids are subsequently deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish amplicons of different molecules or different parent molecules in the same sample or subsample). This includes, for example, uniquely labeling different nucleic acid molecules in a given sample, or non-uniquely labeling these molecules. In the case of non-unique labeling applications, a limited number of tags (i.e., molecular barcodes) can be used to label each nucleic acid molecule so that different molecules can be distinguished based on their endogenous sequence information (e.g., their mapping to the start and / or end position of a selected reference genome, subsequences at one or both ends of the sequence, and / or sequence length) in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used such that any two molecules can have identical endogenous sequence information (e.g., start and / or end position, subsequence at one or both ends of the sequence, and / or length) and also have a low probability (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% probability) of having the same molecular barcode.
[0091] Adaptor is a short nucleic acid (for example, a length less than 500, 100 or 50 nucleotides), usually at least partially double-stranded, for being connected to one end or both ends (an adaptor at every end) of a sample nucleic acid molecule. Adaptor can include a primer binding site to allow the amplification of the nucleic acid molecules with both ends being adaptors, and / or sequencing primer binding sites, including primer binding sites for next generation sequencing (NGS). Adaptor can also include a binding site for capturing a probe, such as an oligonucleotide attached to a flow cell support. Adaptor can also include a barcode as described above. Barcode is preferably positioned relative to primer and sequencing primer binding site so that barcode is included in the amplicon and sequencing read of nucleic acid molecules. The adaptor of identical or different sequence can be connected to the corresponding end of nucleic acid molecules. Sometimes, except that barcode is different, identical adaptor is connected to the corresponding end. A preferred adapter is a Y-shaped adapter, wherein one end is blunt-ended or tailed as described herein for ligation to a nucleic acid molecule also blunt-ended or tailed having one or more complementary nucleotides, and the other end of the Y-shaped adapter comprises a non-complementary sequence that does not hybridize to form a double strand. Another preferred adapter is a bell-shaped adapter, which also has a blunt-ended or tailed end for ligation to the nucleic acid to be analyzed.
[0092] As used herein, the term "sequencing" refers to any of several techniques for determining the sequence of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Exemplary sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exome sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid phase sequencing, high throughput sequencing, massively parallel signature sequencing, emulsion PCR, low denaturing temperature co-amplification-PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD TM Sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as, for example, a genetic analyzer commercially available from Illumina or Applied Biosystems.
[0093] The phrase "next generation sequencing" or NGS refers to sequencing technologies that have increased throughput, e.g., the ability to generate hundreds of thousands of relatively small sequence reads at a time, as compared to traditional Sanger and capillary electrophoresis-based methods. Some examples of next generation sequencing technologies include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.
[0094] DNA (deoxyribonucleic acid) is a nucleotide chain that includes four types of nucleotides: adenine (A), thymine (T), cytosine (C), and guanine (G). RNA (ribonucleic acid) is a nucleotide chain that includes four types of nucleotides: A, uracil (U), G, and C. Certain pairs of nucleotides specifically bind to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid chain is combined with a second nucleic acid chain composed of nucleotides complementary to those in the first chain, the two chains combine to form a double strand. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "genetic sequence," or "fragment sequence," or "nucleic acid sequencing read" means any information or data indicating the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule (e.g., a whole genome, a whole transcriptome, an exome, an oligonucleotide, a polynucleotide, or a fragment) of a nucleic acid such as DNA or RNA. It should be understood that the present teachings contemplate sequence information obtained using all available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.
[0095] "Polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogs) linked by internucleosidic linkages. Typically, a polynucleotide comprises at least three nucleosides. Oligonucleotides typically range in size from a few monomeric units, e.g., 3-4, to hundreds of monomeric units. Unless otherwise indicated, whenever a polynucleotide is represented by a sequence of letters, such as "ATGCCTG," it will be understood that the nucleotides are in a 5'→3' order from left to right, and that "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents thymidine. As is standard in the art, the letters A, C, G, and T may be used to refer to a base itself, a nucleoside, or a nucleotide comprising a base.
[0096] The phrase "sequence read coverage" refers to the number of sequence reads that align to a site of a reference sequence. "Sequence coverage information" refers to information indicating the sequence read coverage of a given site of a reference sequence. Sequence coverage information can include the number or identity of sequence reads that align to that site and / or other information indicating the sequence read coverage at that site.
[0097] The phrase "molecular coverage" refers to the number of molecules that cover a site of a reference sequence. Molecules can be identified based on the sequence reads and molecular barcodes described herein. Similarly, whether a molecule covers a site of a reference sequence can be determined based on the sequence reads generated by the molecule aligned to the site.
[0098] Reference sequence is a known sequence for the purpose of comparing with a sequence determined by experiment. For example, a known sequence can be a whole genome, a chromosome or any section thereof. A reference sequence generally includes at least 20, 50, 100, 200, 250, 300, 350, 400, 450, 500, 1,000, 10,000, 100,000, 1,000,000, 10,000,000, 100,000,000, 1,000,000,000 or more nucleotides. A reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can include a non-continuous segment (segment) compared with different regions of a genome or chromosome. Reference to the human genome includes, for example, hG19 and hG38.
[0099] The term "specified position" in a reference sequence refers to the genomic coordinates in the reference sequence.
[0100] If the first single-stranded nucleic acid sequence or its complement and the second single-stranded sequence or its complement are aligned with overlapping but non-identical segments of a continuous reference sequence, such as the sequence of a human chromosome, then the first nucleic acid sequence overlaps with the second nucleic acid sequence. If either strand of a fully or partially double-stranded nucleic acid overlaps with either strand of another fully or partially double-stranded nucleic acid, then a fully or partially double-stranded nucleic acid overlaps with another fully or partially double-stranded nucleic acid.
[0101] A "C" to "T" variant or conversion refers to the presence of a "T" base in the sequenced polynucleotide at the coordinate position occupied by a "C" base in the reference sequence. A "G" to "A" variant or conversion refers to the presence of an "A" base in the sequenced polynucleotide at the coordinate position occupied by a "G" base in the reference sequence.
[0102] Nucleic acid molecules can be conceptually divided into a 5' terminal end, an internal portion, and a 3' terminal end. The terminal end can be specified based on a predetermined number of nucleotides from the end. For example, the 5' terminal end is represented by, for example, 20 terminal nucleotides from the 5' end. The 3' terminal end is represented by, for example, 20 terminal nucleotides from the 3' end. Alternatively, a nucleic acid molecule can be divided into a terminal portion and a remainder as described above.
[0103] The term "minor allele frequency" ("MAF") refers to the frequency with which a minor allele (eg, an allele that is not the most common) occurs in a given population of nucleic acids, such as a sample.
[0104] "Tumor fraction" (TF) refers to the fraction of DNA molecules associated with a tumor in a given sample. TF can be obtained based on detecting a reduction in the coverage of variant alleles in tumor cells. A lower TF in a given sample may affect the MAF of a given variant allele in a given sample, and therefore affect the detectability of a given variant allele.
[0105] The term "tumor not detected" or "TND" refers to a sample in which no somatic single nucleotide variants, indels, copy number variants, or fusions are detected.
[0106] The terms "process," "calculate," and "compare" can be used interchangeably. The terms can refer to determining differences, such as differences in quantity or sequence. For example, gene expression, copy number variation (CNV), insertion / deletion, and / or single nucleotide variant (SNV) values or sequences can be processed.
[0107] Adaptor is an artificially synthesized sequence that can be coupled to nucleic acid molecules or polynucleotide sequences by any method (including connection, hybridization and / or amplification). Adaptor is a short nucleic acid (for example, less than 500, 100 or 50 nucleotides long) for being connected to an end or two ends of sample nucleic acid molecules, typically at least partially double-stranded. Adaptor can include primer binding sites to allow the amplification of nucleic acid molecules flanked by two end flanks of adaptors, and / or can include sequencing primer binding sites, including primer binding sites for next generation sequencing (NGS). Adaptor can also include binding sites for capture probes, such as oligonucleotides attached to flow cell supports. Adaptor can also include barcodes as described above. Label is preferably positioned relative to primer and sequencing primer binding site so that label is included in amplicon and sequencing reads of nucleic acid molecules. Identical or different adaptors can be connected to the corresponding ends of nucleic acid molecules. Sometimes, except that label is different, identical adaptors are connected to the corresponding ends. A preferred adapter is a Y-shaped adapter, wherein one end is blunt-ended or tailed as described herein for ligation to a nucleic acid molecule that is also blunt-ended or tailed and has one or more complementary nucleotides. Another preferred adapter is a bell-shaped adapter, which also has a blunt-ended or tailed end for ligation to the nucleic acid to be analyzed.
[0108] Detailed description
[0109] Figure 1 An example of a system 100 for training and using a computer model to predict the state of genetic material based on the probability of each state, according to an embodiment of the present disclosure, is shown. The system can process a sample 101 to train one or more models 140 (illustrated as models 140a...n), each of which outputs a probability that genetic material of a sample, such as from an individual under test (IUT) 111, is in a particular state. In some examples, sample 101 can include a set of various genes of interest being studied.
[0110] For example, the system can use model 140a to determine that the sample from IUT 111 contains the probability of somatic homozygous deletions related to the gene. The system can use another model 140b to determine that the sample from IUT 111 contains the probability of somatic heterozygous deletions related to the gene. Then, the system can compare the probabilities with each other to determine which of somatic homozygous deletions or somatic heterozygous deletions is more likely. The system can also provide other types of accurate diagnosis based on competing probabilities. For example, the system can model CNVs by modeling the probabilities of different copy numbers. Based on the comparison of the output probabilities of each model (each model can correspond to different copy number predictions), the system can determine the CNVs in the sample of IUT 111.
[0111] The system 100 may include a sequencing system 102, a computer system 110, and / or other components. It should be noted that the sequencing system 102 and the computer system 110 may be remote from each other and connected to each other via a computer network (not shown). The sequencing system 102 may include a sample collection and preparation pipeline 103, a sequencing pipeline 105, and a sequence read data storage 109 and / or other components. The sequencing pipeline 105 may include one or more sequencing devices 107 (in Figure 1 shown as sequencing devices 107a…n).
[0112] Computer system 110 may include a sequence analysis pipeline 112 , a processor 120 , a storage device 122 , a data pre-processing subsystem 124 , a classifier 130 , a model validator 132 , and / or other components.
[0113] The sequence analysis pipeline 112 may include a sequence quality control (QC) component 113, an alignment component 114, other analysis components 115, and an analysis QC component 116. Output from the sequence analysis pipeline 112 may be stored in an analysis data store 117. A data preprocessing subsystem 124 may preprocess data from the sequence analysis pipeline 112 to generate a training dataset 125. For example, the training dataset 125 may include data from a sample 101 in which no tumor was detected ("TND") (when cancer is to be diagnosed) or other normal samples (when other types of diseases or conditions are to be diagnosed). Examples disclosed throughout the disclosure may be described with reference to TND samples.
[0114] In some embodiments, the training data set 125 may be stored in a training data store 126. Figure 2 1 to illustrate example operations of the processor 120 . Figure 2A schematic diagram 200 is shown of determining allele counts of germline SNPs to predict the status of a gene 201 according to an embodiment of the present disclosure. In some examples, for a region of interest 201 around a gene 201, the processor 120 can identify germline SNPs for TND samples. In one training example, germline SNPs were selected from 28,199 samples. Of these samples, 5105 samples (18%) were identified as having TND and applied to the population allele / genotype frequencies. The germline SNPs were selected to meet the following criteria: (1) they were within 3 Mb of the selected gene (such as BRCA1, BRCA2, ATM), (2) the frequency of heterozygote calls (MAF>25% and MAF<75%) in the 5105 TND samples was between 5% and 95%, and (3) the variant was not called as somatic in all 28,199 samples. Region of interest 203 may include N bases upstream of the start of gene 201 and M bases downstream of the end of gene 201. The values of N and M may be the same or different. In some examples, each of N and M may be 3,000,000 nucleotides (3 Mb).
[0115] exist Figure 2 In the example shown, the reference wild-type nucleotide at SNP site (i) (illustrated as SNP (i)) can be "G". The nucleotides called at this position can be different from each other throughout the TND sample. Because the TND sample is normal, the processor 120 can assume that SNP (i) and other SNP sites of the TND sample will not lead to tumors or abnormal states. Therefore, these SNP sites can each serve as a reference expected value for modeling allele counts for probability determination of gene states. For example, the deviation of nucleotide calls observed at each SNP position can indicate the probability that such deviation leads to a specific state (such as a tumor or other abnormal state of gene 201). The processor 120 can accordingly train the model 140 based on the expected value derived from the calculation of data on germline SNPs from the TND sample. For each SNP site, such calculated data can include: the prevalence of heterozygosity, the standard deviation of the minor allele frequency (MAF), the genotype, the germline prevalence (historical) and / or other data.
[0116] Using the calculated expected value, the processor 120 can model the state of the gene 201 based on the sequence reads of the sample of the IUT 111 aligned with the region of interest 203. For example, the processor 120 can generate competitive models 140, each of which outputs a corresponding score representing the probability that the gene 201 is in a specific state. The processor 120 can compare the corresponding scores to calculate a predicted score, which can be compared with a threshold score to determine the state of the gene 201. The processor 120 can calculate the threshold score based on observed data from the training sample, which will be further described below.
[0117] In some embodiments, the molecular sequence reads generated from the sample of IUT 111 can be aligned with the reference genome to identify the alleles (mutants or wild types) supported by the potential molecules. The sample of IUT 111 can be prepared in the sample collection and preparation pipeline 103 and sequenced in the sequencing pipeline 105. Each molecule can be associated with a sequence read. Multiple sequence reads from multiple molecules of the sample of IUT 111 can cover a given germline SNP site.
[0118] Based on the alignment of sequence reads generated from the sample of IUT 111, processor 120 can identify the number of molecules that support the SNP allele and calculate the total number of molecules. Processor 120 can use the expected value data calculated from each germline SNP in region of interest 203 to model these and / or other data for the sample of IUT 111. For example, processor 120 can generate a first output of allele count model 140a representing the probability of a first state of gene 201 and a second output of allele count model 140b representing the probability of a second state of gene 201.
[0119] Processor 120 can implement different types of probability distributions to generate model 140. In addition, model 140 can model various types of states of gene 201, or more generally, various types of states of genetic material. Attention will now be directed to the modeling and examples of the types of states modeled by processor 120.
[0120] Generally speaking, the processor 120 may implement the classifier 130 (by being programmed by the classifier 130). Alternatively, it should be noted that the classifier 130 may comprise a hardware module. In any case, the classifier 130 (which may program the processor 120) may classify the gene (such as a gene) based on the alleles detected in the region of interest associated with the gene. Figure 2 More specifically, based on the training data set 125, the classifier 130 can be based on the region of interest (such as Figure 2The method of claim 1 , wherein one or more probability models 140 (illustrated as models 140a, 140b, ..., 140n) are used to determine a particular state of a gene based on the germline single nucleotide polymorphism (SNP) position in the region of interest 203 (shown in FIG. 1 ). Each model 140 can correspond to a corresponding probability of a gene state. The SNP position can be based on the sequencing reads generated from various samples 101 by the sequencing system 102. The state can include the mutation state of the gene, such as, but not limited to, somatic homozygous deletion, somatic heterozygous deletion, copy number variation ("CNV") (including specific copy number wild type, gain or loss) and / or other states of the gene.
[0121] In various embodiments, the classifier 130 can apply a model that can be generated based on the training data set 125 to determine the status of genes in an individual's sample. For example, the classifier 130 can determine the probability that genes in a cfDNA molecular sample from an individual include somatic homozygous deletions, somatic heterozygous deletions, and / or other states that may be associated with a disease such as cancer or other health conditions. Based on this prediction, precise treatment can be customized for the individual. Similarly, the computer system 110 can be improved to provide advanced diagnostic capabilities based on non-invasive analysis of genetic material such as cfDNA.
[0122] It is noted that although the examples described herein may relate to determining gene states, the states of other genetic materials such as chromosomes, exomes, and / or other genetic materials may also be determined. For example, CNVs of chromosomes, exomes, and / or other genetic materials may be determined. Having provided a functional description of the classifier 130, attention will now be turned to a more detailed example of determining gene states by training various models 140 and using the models 140 to predict the probability that a particular sample under examination exhibits a particular gene state.
[0123] Model training based on TND samples
[0124] In some embodiments, the classifier 130 uses data from the samples 101. The data may include a set of samples in which no tumor was detected ("TND" samples). The classifier 130 may use the TND samples to determine the prevalence of heterozygosity for each germline SNP in the TND samples and the standard deviation of the minor allele frequency (MAF) for each germline SNP in the TND samples. It should be noted that in the formulas and calculations described throughout this disclosure, variance may be used instead of standard deviation, as long as such calculations are appropriately adjusted to use variance instead of standard deviation. The prevalence and standard deviation of heterozygosity can provide a baseline expectation for a "normal" sample, i.e., a sample that does not exhibit a disease state. The classifier 130 may also estimate the germline prevalence (a priori) g for each site i i An example calculation of the prevalence of heterozygosity for each germline SNP can be given by equation (1):
[0125] p a (g i )=P(b ij =a|g i ) (1),
[0126] in:
[0127] p a (g i ) represents the prevalence of heterozygosity for each germline SNP,
[0128] b ij represents the set of bases observed at SNP site i, and
[0129] g i represents the genotype at SNP site i (AA / Aa / aa).
[0130] Modeling of somatic homozygous deletions
[0131] Sorter 130 can generate the first model 140a of allele count based on one or more germline single nucleotide polymorphism (SNP) positions associated with gene via probability distribution.The first model can represent the somatic homozygous deletion of (such as modeling) gene.For example, given the prevalence rate (equation (1)) of each germline SNP heterozygosity in TND sample and the standard deviation of the MAF of each germline SNP in TND sample, sorter 130 can be to the probability modeling that the gene of the specific sample from individual is associated with somatic homozygous deletion.For this reason, sorter 130 can read the number of the molecule that has somatic homozygous deletion in the gene of the sample of IUT 111, and the total number of molecules of the sample of IUT 111.For example, sorter 130 can generate the model of the probability that gene has somatic homozygous deletion in representative sample, such as model 140a. In some embodiments, the classifier 130 may use a beta-binomial probability distribution to generate the model 140a, although other probability distributions may be used, such as a binomial probability distribution, a normal (Gaussian) distribution, and / or other probability modeling.
[0132] The beta-binomial distribution is a binomial distribution of n Bernoulli trials, where the probability of success on each trial is fixed but drawn randomly from a beta distribution. The beta-binomial distribution can be expressed using two parameters: α and β (which are uniquely determined by the mean and standard deviation of the distribution). When n = 1, the distribution simplifies to a Bernoulli distribution. For α = β = 1, it is a discrete uniform distribution from 0 to n.
[0133] The binomial distribution is the probability distribution of a binomial random variable. The binomial random variable is the number of successes in N repeated trials of a binomial experiment. The binomial distribution has the following properties: the mean (μx) of the distribution is equal to n*P; the variance is given by: n*P*(1-P); and the standard deviation (σx) is given by equation (2):
[0134]
[0135] The normal distribution can be defined using the normal equation:
[0136] Y={1 / [σ*sqrt(2π)]}*e-(x-μ)2 / 2σ2 (3),
[0137] in:
[0138] X is a normal random variable,
[0139] μ is the mean value,
[0140] σ is the standard deviation,
[0141] π is approximately 3.14159, and
[0142] e is approximately 2.71828.
[0143] For illustrative purposes, an example using a β-binomial probability distribution will now be described. Those skilled in the art will appreciate that, based on the disclosure herein, binomial, normal, and / or other probability distributions may also be used. For the β-binomial probability distribution, the classifier 130 may use the dbetabinom function in the VGAM package of the R project according to equation (4):
[0144] P(b ij |g i |TF)=P(b ij |g i )=VGAM::dbetabinom(m i , R i , p a (g i ), sd(g i )) (4),
[0145] in:
[0146] m i represents the number of molecules supporting the SNP allele at SNP site i,
[0147] R i represents the total number of molecules,
[0148] p a (g i) represents the prevalence of heterozygosity at SNP site i, and
[0149] sd(g i ) represents the standard deviation of MAF.
[0150] The classifier 130 may generate a first probability output (L1) of the first model 140a according to equation (5):
[0151] L1=prod(sum gi (prior gi *dbtabinom(m i , R i , p a (g i ), sd(g i ))) (5).
[0152] Somatic loss of heterozygosity modeling
[0153] The classifier 130 can generate a second model 140b of allele counts in the sample based on one or more germline SNP positions via a probability distribution. The second model 140b can represent (such as modeling) somatic heterozygous deletions of genes. Because the detection of heterozygous deletions can be affected by TF, the classifier 130 can determine the mean mu.tf (also denoted as μ.tf) and standard deviation sd.tf (also denoted as σ.tf) of TF based on the read coverage (sequence read coverage) in the sample of IUT 111.
[0154] In some embodiments, the classifier 130 may use a beta-binomial probability distribution to generate the model 140b, although other probability distributions such as a binomial probability distribution, a Gaussian distribution, and / or other probability modeling may be used.
[0155] For the β-binomial probability distribution, the classifier 130 may use the dbetabinom function in the VGAM package of the R project according to equation (6):
[0156] P(b ij |g i |TF)=P(b ij |g i )=VGAM::dbetabinom(m i , R i ,mu i , sd i ) (6),
[0157] in:
[0158] m irepresents the number of molecules supporting the SNP allele at SNP site i,
[0159] R i represents the total number of molecules,
[0160] mui i represents the mean value of TF calculated for IUT 111 samples, and
[0161] sd i represents the standard deviation of TF calculated for IUT 111 samples.
[0162] The classifier 130 may generate a second probability output L0 of the second model 140b according to equation (7):
[0163] L0=prod(sum gi (prior gi *dbtabinom(m i , R i , m i (g i ), sd i (g i )))(7),
[0164] in:
[0165] If g i =AA or aa, then m i (g i )=pa(g i );
[0166] 1-m.tf*p a (g i –Aa)+m.tf*max(p a (g i =aa),p a (g i =AA)),
[0167] If g i =AA or aa, then sd i (g i )=sd(g i )
[0168] var_prod(sd.tf,1-mf.tf,sd(g i =Aa),p a (g i =Aa))+
[0169] var_prof(sd.tf,m.tf,sd(g i =aa),p a(g i =aa))
[0170] Where var_prod(va,ma,vb,mb)=va*vb+va*mb 2 +vb*ma 2
[0171] It should be noted that the first model 140a and the second model 140b do not need to use the same probability distribution, as long as the first model 140a and the second model 140b each output a probability.
[0172] The classifier 130 can compare the first probability output of the first model 140a and the second probability output of the second model 140b to determine which probability output is more likely. For example, the classifier 130 can determine whether somatic homozygous deletion or somatic heterozygous deletion is more likely. In a specific example, the classifier 130 can use the log-likelihood ratio ("LLR") to generate an LLR score based on the first probability output (probability of somatic homozygous deletion) and the second probability output (probability of somatic heterozygous deletion). In some embodiments, one of the first probability output or the second probability output can be used as a zero probability, so that if the LLR score does not exceed the threshold cutoff score, the zero probability is rejected. For example, the classifier 130 can compare the LLR score with the threshold cutoff score to determine whether the second probability output should be rejected. In other words, if the LLR score exceeds the threshold cutoff score, the classifier 130 can determine that the first probability output should be selected. In this example, the classifier 130 can generate a prediction of the presence of somatic homozygous deletion of the gene in the sample of the IUT 111 based on this comparison.
[0173] In some examples, to reduce error, model 140A or 140B can be used to determine the sample genotype for each SNP that overlaps with a given gene. If no germline SNP is determined to be heterozygous, the given gene can be labeled as "no call" and no somatic homozygous deletions or somatic heterozygous deletions are associated with the given gene.
[0174] Learning threshold score cutoff
[0175] In some embodiments, the threshold cutoff scores can be customized based on the different genes or other genetic material being analyzed. For example, the BRCA1 gene may be associated with a different threshold cutoff score than the BRCA2 gene. Other genes may also be associated with customized threshold cutoff scores. In these embodiments, the classifier 130 can be trained to determine the threshold cutoff scores. In some of these embodiments, the classifier 130 can be trained to determine the threshold cutoff scores for specific genes. For example, the classifier 130 can be trained using simulations of somatic heterozygous deletions starting from TND samples. Figure 10A and Figure 10B For example, when a BRCA1 and BRCA2 negative sample does not have a homozygous deletion, a limit of blank (LoB) or the highest LLR score would be expected. Figure 10A and Figure 10B , 100,000 somatic heterozygous deletions starting from TND samples were simulated. The TF distribution observed in 28,000 samples was used as the TF for determining LoB for BRCA1 and BRCA2. As shown in the figure, the threshold cutoff scores for BRCA1 and BRCA2, compared to the LLR scores, were 20.1 and 0, respectively. Therefore, if a somatic deletion of BRCA1 is observed in a sample from IUT 111 and the LLR score of BRCA1 in the sample from IUT 111 is >20.1, then the classifier 130 can predict that the somatic deletion is a somatic homozygous deletion. Similarly, if a somatic deletion of BRCA2 is observed in a sample from IUT 111 and the LLR score of BRCA2 in the sample from IUT 111 is >0, then the classifier 130 can predict that the somatic deletion is a somatic homozygous deletion. It should be noted that other genes can be similarly simulated to determine the threshold cutoff scores.
[0176] In some embodiments, the model validator 132 can use simulation data and / or clinical data to validate the results of the model 140. For example, the model validator 132 can consult the diagnostic result data store 150 and / or the clinical result data store 160 to validate the prediction. For the simulation results, a panel of known samples can be modeled to generate predictions of the genetic material state of these samples. These results can be used to validate the results of previous predictions and / or future predictions.
[0177] Figure 3 FIG300 is a diagram illustrating a process 300 for predicting somatic homozygous deletions or somatic heterozygous deletions based on a trained model according to an embodiment of the present disclosure. The process 300 is provided by way of example, as there are many ways to perform the methods described herein. Although the process 300 is primarily described as consisting of Figure 1Computer system 110 (via processor 120 ) is shown as being performed, but process 300 may be performed by other systems or combinations of systems or performed in other ways. Figure 3 Each block shown in the may also represent one or more processes, methods, or subroutines, and one or more blocks may include machine-readable instructions stored on a non-transitory computer-readable medium and executed by a processor or other type of processing circuitry to perform one or more operations described herein. The various operations of the process 300 disclosed herein, or blocks performed by the system disclosed herein, may be performed at the same or different times, in the same or different geographic locations, such as countries, and / or by the same or different persons.
[0178] In operation 302, the processor 120 can read the germline SNP data from a set of samples (including TND samples). In operation 304, the processor 120 can determine the prevalence and SD of heterozygosity of MAF based on the germline SNP data. In operation 306, the processor 120 can determine a first number of reads for each germline SNP site that supports determining that a sample from an individual comprises a somatic homozygous deletion at a gene.
[0179] In operation 308, processor 120 can generate the first output of the first model of the probability that gene is associated with somatic homozygous deletion based on the prevalence rate of heterozygosity, the standard deviation (sd) of MAF, germline SNP data and the number of reads. In operation 310, processor 120 can determine the mean value and standard deviation of TF based on the sample from individual. In operation 312, processor 120 can determine the second number of reads of each germline SNP site that supports determining that the sample from individual comprises somatic heterozygous deletion at gene place. In operation 314, processor 120 can generate the second output of the second model of the probability that gene is associated with somatic heterozygous deletion based on the mean value and SD of TF, germline SNP data and the second number of reads. In operation 316, processor 120 can compare this first output and this second output. In operation 318, processor 120 can determine whether to select the first output based on comparison. In operation 320, processor 120 can generate the prediction that gene comprises somatic homozygous deletion based on determining whether to select the first output.
[0180] Classifier 130 can apply various modeling techniques to predict the status of genes. Classifier 130 can also use other modeling techniques. For example, Figure 9A and Figure 9B A comparison of the results of different modeling techniques is shown. Other probabilistic techniques can also be used.
[0181] Modeling other types of states of genetic material
[0182] The classifier 130 can model other types of states of genetic material. For example, the classifier 130 can predict various types of states of genetic material, such as CNV. Figure 4 , which shows a process 400 for predicting the state of genetic material. Process 400 is provided by way of example, as there are many ways to perform the methods described herein. Although method 400 is primarily described as consisting of Figure 1 Computer system 110 (via processor 120 ) is shown as being performed, but process 400 may be performed by other systems or combinations of systems or performed in other ways. Figure 4 Each block shown in the 400 may also represent one or more processes, methods, or subroutines, and one or more blocks may include machine-readable instructions stored on a non-transitory computer-readable medium and executed by a processor or other type of processing circuitry to perform one or more operations described herein. The various operations of the process 400 disclosed herein, or blocks performed by the system disclosed herein, may be performed at the same or different times, in the same or different geographic locations, such as countries, and / or by the same or different persons.
[0183] about Figure 4 The examples described include determining CNVs in a sample from an IUT 111. More specifically, the examples can be used to determine copy number variations (such as amplifications) in genetic material from a sample from an IUT 111. However, other types of states of genetic material can be determined in a similar manner, using alternative (competing) probabilities for different states and selecting the most likely probability.
[0184] In operation 402, processor 120 may generate a first model that models a first state of the genetic material. The first state may include a first CNV or other state. In operation 404, processor 120 may generate a second model that models a second state of the genetic material. The second state may include a second CNV or other state. In operation 406, processor 120 may generate a first score based on the first model. The first score may indicate the probability that the genetic material is in the first state.
[0185] In operation 408, processor 120 may generate a second score based on the second model. The second score may indicate a probability that the genetic material is in the second state. In operation 410, processor 120 may compare the first score and the second score. In operation 412, processor 120 may generate a prediction based on the comparison of the first score and the second score, indicating whether the genetic material is in the first state or the second state.
[0186] Similar to how classifier 130 uses MAF associated with germline SNPs to generate somatic heterozygous deletion and somatic homozygous deletion probabilities, MAF can be used to interpret CNV probabilities. For example, the MAF of a germline SNP in a sample where no CNV is detected can be used to determine whether the reads in the sample support a specific amplification.
[0187] Figure 5 Types of somatic deletions according to embodiments of the present disclosure are shown. Somatic homozygous deletions can be generated in two ways: (1) The germline has a single copy of the gene and the somatic cell acquires a secondary deletion (similar to the LoD detected by single copy amplification). These can be detected based on coverage + no heterozygous SNP overlap. In some cases, although these are not observed, the germline may not have a copy of the gene. (2) A second way in which somatic homozygous deletions can be generated is that the germline has two copies of the gene and the somatic cell loses both copies (this situation is observed at a higher prevalence). In some embodiments, for biallelic somatic copy number loss, the reference allele frequency for the germline heterozygous SNP in a mixture of germline and somatic cells is 0.5. In the case of somatic LOH, the reference allele frequency is 0.5-0.5*TF (tumor fraction) or 0.5+0.5*TF, depending on whether the reference allele is lost or retained in the cancer cell. In some embodiments, for LOH, the expected allele frequency may depend on the fraction of tumor cells. Therefore, the system can distinguish LOH from biallelic copy number loss based on the calculated allele frequency compared to the expected allele frequency of 0.5.
[0188] Figure 6A Example graphs 600(A)(1) and 600(A)(2) are shown for a homozygous deletion of BRCA1 according to embodiments of the present disclosure. Figure 6B Example Figures 600(A)(1) and 600(A)(2) of BRCA2 heterozygous deletions are shown according to embodiments of the present disclosure. Referring to Figures 600(A)(1) and 600(B)(1), for a given cfDNA sample, normalized molecular coverage (y-axis) is shown by targeted probes sorted by genomic position (x-axis). Chromosome segregation is represented by vertical lines, and identifiers are shown at the bottom of the figure. Regions without somatic copy number variation exhibit molecular coverage close to 2, while somatic deletions can be identified by molecular coverage levels below 2. Referring to Figures 600(B)(1) and 600(B)(2), for the same sample, MAFs (y-axis) with known germline SNPs are shown relative to their genomic position (x-axis). Somatic deletions observed in the coverage plots in the top row exhibit germline variant MAFs close to 50% (see Figure 6A), whereas heterozygous deletions produce unbalanced germline variant MAFs (see Figure 6B ).
[0189] Figure 7A Shown are example graphs of the prevalence of heterozygous genotypes for known germline SNPs overlapping ATM, BRCA1, and BRCA2 genes observed in TND samples, according to embodiments of the present disclosure. Figure 7B An example graph showing MAF across TND samples according to an embodiment of the present disclosure is shown.
[0190] Figure 8A Shown are example graphs of MAF values for BRCA1 according to embodiments of the present disclosure. Figure 8B Shown are example graphs of MAF values for BRCA2 according to embodiments of the present disclosure. Figure 8A and Figure 8B Shown are examples of MAFs (y-axis) for 9 known germline SNVs, for 3 possible genotypes (homozygous alternative allele / heterozygous / homozygous reference allele) for each SNP (x-axis). Figure 9A Shown is an example graph of score comparison between the beta-binomial model and the binomial model for the BRCA2 group, according to embodiments of the present disclosure. Figure 9B An example graph showing a score comparison between a beta-binomial model and a Gaussian model for the BRCA2 group, according to an embodiment of the present disclosure. Figure 10A An example graph showing the distribution of LLR scores for BRCA1 negative samples according to embodiments of the present disclosure is shown. Figure 10B Shown is an example graph of the distribution of LLR scores for BRCA2-negative samples according to embodiments of the present disclosure.
[0191] Figure 11A Shown are example diagrams of LoD deletions for BRCA1, according to embodiments of the present disclosure. Figure 11B An example graph showing the LoD of loss of heterozygosity (LOH) (interchangeably referred to herein as "deletion of heterozygosity") of BRCA1 according to embodiments of the present disclosure is shown. Simulation: 100k homozygous somatic deletions starting from a TND sample.
[0192] TF used = distribution of TFs observed in 28,199 samples. LoD depends on two factors (two-step algorithm): (1) deletion detection sensitivity (based on coverage only): BRCA1 amplification / deletion mean cutoff = 0.05; and (2) ability to distinguish between homozygous and heterozygous somatic deletions (LLR assay).
[0193] Figure 12AShown are example graphs of LoD deletions for BRCA2, according to embodiments of the present disclosure. Figure 12B An example graph showing LoD of LOH for BRCA2 according to embodiments of the present disclosure is shown. Simulation: 100k homozygous somatic deletions starting from a TND sample.
[0194] TF used = distribution of TFs observed in 28,199 samples.
[0195] The LoD depends on two factors (two-step algorithm): (1) deletion detection sensitivity (based on coverage alone): BRCA2 amplification / deletion mean cutoff = 0.09; and (2) the ability to distinguish between homozygous and heterozygous somatic deletions (LLR assay).
[0196] Figure 13 Shown is an example graph of TF prevalence versus cancer type, according to embodiments of the present disclosure.
[0197] Figure 14 An example density plot of LLR scores for BRCA1 and BRCA2 is shown, according to embodiments of the present disclosure. A set of 28,000 training samples was randomly selected with cutoffs of 2.5 and 0 (determined in the LoB section) to call samples with homozygous BRCA1 / 2 deletions. 387 and 994 samples showed somatic deletions of BRCA1 and BRCA2, respectively. Of these samples, 49 and 60 were called as having homozygous BRCA1 and BRCA2 deletions, respectively.
[0198] Figure 15 Shown are example graphs of the prevalence of BRCA2 homozygous deletions observed in more than one cancer type population, according to embodiments of the present disclosure. Figure 16 Shown are example graphs of the prevalence of homozygous deletions in BRCA1 observed in more than one cancer type population, according to embodiments of the present disclosure. Figure 17 An example of a BRCA2 homozygous deletion and potential clinical actionability according to embodiments of the present disclosure is shown. Figure 17The figure shown is from "Integrative clinical genomics of advanced prostate cancer," Cell 161:1215–1228 (2015), Robinson D, Van Allen EM, Wu YM, Schultz N, Lonigro RJ, Mosquera JM, Montgomery B, Taplin ME, Pritchard CC, Attard G, et al. ("Robinson"), the entire contents of which are incorporated herein by reference. Robinson showed that a comprehensive analysis of somatic and pathogenic germline alterations in BRCA2 identified 19 / 150 (12.7%) cases of BRCA2 loss, of which approximately 90% exhibited biallelic loss. This was often the result of somatic point mutations and loss of heterozygosity, as well as homozygous deletions. A clinical trial evaluating poly(ADP-ribose) polymerase (PARP) inhibition in unselected affected individuals with mCRPC showed that more than one affected individual who experienced clinical benefit in the trial had biallelic BRCA2 loss, providing further evidence of clinical actionability.
[0199] Figure 18A Shown is an example diagram of a homozygous deletion of BRCA1 according to an embodiment of the present disclosure. Figure 18B Shown is an example diagram of a homozygous deletion of BRCA1 according to an embodiment of the present disclosure. Figure 19A Shown are example diagrams of BRCA2 homozygous deletions according to embodiments of the present disclosure. Figure 19B Shown are example diagrams of BRCA2 homozygous deletions according to embodiments of the present disclosure. Figure 18A 、 Figure 18B 、 Figure 19A and Figure 19B It is based on a map of the human genome.
[0200] Figure 20A An example diagram of BRCA1 biallelic somatic copy number loss is shown, according to an embodiment of the present disclosure.For the purposes of this disclosure, the term "biallelic somatic copy number loss" will be used interchangeably with "homozygous deletion." Figure 20B An exemplary diagram of BRCA1 LOH according to an embodiment of the present disclosure is shown. For the purposes of this disclosure, the term "LOH" will be used interchangeably with "loss of heterozygous." Figure 21A Shown are example diagrams of BRCA2 biallelic somatic copy number loss, according to embodiments of the present disclosure. Figure 21B Shown are example diagrams of BRCA2 LOH according to embodiments of the present disclosure. Figure 20A、 Figure 20B 、 Figure 21A and Figure 21B is a diagram based on three (human) chromosomes. Figure 22 A graph showing the prevalence of BRCA1 and BRCA2 somatic deletions according to embodiments of the present disclosure.
[0201] Computer Implementation
[0202] The method can be computer-implemented such that any or all of the steps described in the specification or appended claims, except the wet chemistry steps, can be performed on a suitably programmed computer. The computer can be a mainframe computer, a personal computer, a tablet computer, a smartphone, the cloud, an online data store, a remote data store, or the like. The computer can be operated in one or more locations.
[0203] Various operations of the present method may utilize information and / or programs and generate results that are stored on a computer-readable medium (e.g., a hard drive, secondary storage, external storage, a server; a database, a portable storage device (e.g., a CD-R, DVD, ZIP disk, flash memory card), etc.).
[0204] The present disclosure also includes an article of manufacture for analyzing a nucleic acid population, the article of manufacture comprising a machine-readable medium containing one or more programs that, when executed, perform the steps of the present method.
[0205] The present disclosure may be implemented in hardware and / or software. For example, different aspects of the present disclosure may be implemented in client-side logic or server-side logic. The present disclosure or its components may be embodied in a fixed medium program component containing logic instructions and / or data that, when loaded into a suitably configured computing device, causes the device to perform according to the present disclosure. The fixed medium containing the logic instructions may be delivered to an observer on a fixed medium for physical loading into the observer's computer, or the fixed medium containing the logic instructions may reside on a remote server that the observer accesses via a communication medium to download the program component.
[0206] The present disclosure provides a computer control system programmed to implement the methods of the present disclosure. The processor 120 may include a single-core or multi-core processor, or more than one processor for parallel processing. The storage device 122 may include random access memory, read-only memory, flash memory, a hard disk, and / or other types of memory. The computer system 110 may include a communication interface (e.g., a network adapter) for communicating with one or more other systems and peripheral devices, such as a cache, other memory, data storage, and / or an electronic display adapter. The components of the computer system 110 may communicate with each other via an internal communication bus such as a motherboard. The storage device 122 may be a data storage unit (or data repository) for storing data. The computer system 110 may be operably coupled to a computer network ("network") via a communication interface. The network may be the Internet, an intranet and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. In some cases, the network is a telecommunications and / or data network. The network may include a local area network. The network may include one or more computer servers that may support distributed computing, such as cloud computing. In some cases, with the aid of computer system 110 , the network may implement a peer-to-peer network, which enables devices coupled to computer system 110 to operate as either clients or servers.
[0207] The processor 120 can execute a series of machine-readable instructions that can be embodied as a program or software. The instructions can be stored in a memory location, such as a storage device 122. The instructions can be directed to the processor 120, which can then program or otherwise configure the processor 120 to implement the methods of the present disclosure. Examples of operations performed by the processor 120 can include reading, decoding, executing, and writing back.
[0208] Processor 120 may be part of a circuit, such as an integrated circuit. One or more other components of system 100 may be included in the circuit. In some cases, the circuit may include an application-specific integrated circuit (ASIC).
[0209] Storage device 122 can store files such as drivers, libraries, and saved programs. Storage device 122 can store user data, such as user preferences and user projects. In some cases, computer system 110 may include one or more additional data storage units that are external to computer system 110, such as located on a remote server that communicates with computer system 110 via an intranet or the Internet.
[0210] Computer system 110 can communicate with one or more remote computer systems via a network. For example, computer system 110 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., laptop PCs), tablet or tablet PCs (e.g., iPad, Galaxy Tab), phones, smartphones (e.g. iPhone, Android supported devices, ) or a personal digital assistant. Users can access the computer system 110 via a network.
[0211] The methods described herein may be implemented by means of machine (e.g., computer processor) executable code stored in an electronic storage location of the computer system 110, such as, for example, on the storage device 122. The machine executable code or machine readable code may be provided in the form of software. During use, the code may be executed by the processor 120. In some cases, the code may be retrieved from the storage unit 915 and stored on the storage device 122 for immediate access by the processor 120.
[0212] The code may be precompiled and configured for use with a machine having a processor suitable for executing the code, or may be compiled during runtime. The code may be provided in a programming language that may be selected to enable the code to be executed in a precompiled or as-compiled manner.
[0213] Aspects of the systems and methods provided herein, such as computer system 110, can be embodied in programming. Various aspects of the technology can be considered "products" or "articles of manufacture," which are generally in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk.
[0214] "Storage" type media may include any or all tangible memories of computers, processors, etc., or their related modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-temporary storage for software programming at any time. All or part of the software may sometimes be communicated via the Internet or various other telecommunications networks. For example, such communications may enable software to be loaded from one computer or processor to another, for example, from a management server or host to a computer platform of an application server. Therefore, another type of medium that may have software elements includes light waves, radio waves, and electromagnetic waves, such as light waves, radio waves, and electromagnetic waves used across physical interfaces between local devices, through wired and fiber optic landline networks, and through air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered to be media with software. As used herein, unless limited to non-temporary, tangible "storage" media, terms such as computer or machine "readable media" refer to any medium that participates in providing instructions to a processor for execution.
[0215] Thus, a machine-readable medium, such as computer executable code, may take a variety of forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in one or more computers, etc., such as the storage devices shown in the accompanying drawings that can be used to implement databases, etc. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that make up a bus within a computer system. Carrier transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punched card stock tape, any other physical storage medium with a pattern of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can participate in carrying one or more sequences of one or more instructions to a processor for execution.
[0216] Computer system 110 may include or communicate with an electronic display 935 that includes a user interface (UI) for providing, for example, reports. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0217] The method and system of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software after being executed by the processor 120 .
[0218] Sample collection and analysis lines
[0219] Sample 101 can be any biological sample separated from a subject. Sample can include body tissue, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (white blood cell) or white blood cells (leucocyte), endothelial cells, biopsy tissue, cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid, fluid in the intercellular space (including gingival sulcus fluid), bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. Sample is preferably body fluid, particularly blood and its fraction (fraction), and urine. Such sample can include nucleic acid. Such sample can be referred to as nucleic acid sample. In some samples of these samples, nucleic acid may fall off from tumor. Nucleic acid can include DNA and RNA, and can be double-stranded and / or single-stranded form. In instances where the nucleic acid comprises RNA, the systems and methods described herein can determine somatic deletions in a gene of interest encoded by RNA by comparing the gene expression of the gene of interest relative to a reference gene (such as an endogenous control gene like GAPDH) to a training threshold calculated from a normal sample. The sample can be in the form originally isolated from the subject, or can have undergone further processing to remove or add components such as cells, enrich one component relative to another component, or convert one form of nucleic acid into another form of nucleic acid, such as converting RNA into DNA or converting single-stranded nucleic acid into double-stranded nucleic acid. Thus, for example, the body fluid used for analysis is plasma or serum containing cell-free nucleic acids, such as cell-free DNA (cfDNA).
[0220] The volume of plasma can depend on the desired read depth for the region being sequenced. Exemplary volumes are 0.4-40 ml, 5-20 ml, 10-20 ml. For example, the volume can be 0.5 ml, 1 ml, 5 ml, 10 ml, 20 ml, 30 ml, or 40 ml. The volume of plasma sampled can be 5 ml to 20 ml.
[0221] A sample can contain various amounts of nucleic acid comprising genome equivalents. For example, a sample of about 30 ng of DNA can contain about 10,000 (104 ) haploid human genome equivalents, and in the case of cfDNA, can contain approximately 200 billion (2x10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, can contain about 600 billion individual molecules.
[0222] The sample can comprise nucleic acids from different sources, such as from cells and free cells. The sample can comprise nucleic acids carrying mutations. For example, the sample can comprise DNA carrying germline mutations and / or somatic mutations. The sample can comprise DNA carrying cancer-associated mutations (e.g., cancer-associated somatic mutations).
[0223] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification range from about 1 fg to about 1 μg, such as 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, the amount can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The amount can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method can include obtaining 1 femtogram (fg) to 200 ng.
[0224] Cell-free nucleic acid sample refers to a sample comprising cell-free nucleic acid from a subject. Cell-free nucleic acid is a nucleic acid that is not contained in a cell or is not otherwise bound to a cell. For example, a cell-free nucleic acid sample may include nucleic acids retained in a sample after removal of intact cells. Cell-free nucleic acid may refer to all unencapsulated nucleic acids derived from a subject's body fluids (e.g., blood, urine, CSF, etc.). Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA) and its hybrids, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA) or a fragment of any of these. Cell-free nucleic acid may be double-stranded, single-stranded or its hybrids. Cell-free nucleic acid can be released into body fluids by secretion or cell death processes such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, e.g., circulating tumor DNA (ctDNA). Others are released from healthy cells. ctDNA can be unencapsulated tumor-derived fragmented DNA.Cell-free fetal DNA (cffDNA) is fetal DNA that circulates freely in maternal blood.
[0225] The cell-free nucleic acid or a protein associated therewith can have one or more epigenetic modifications, for example, the cell-free nucleic acid can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.
[0226] The cell-free nucleic acids have an exemplary size distribution of about 100-500 nucleotides, with molecules of 110 to about 230 nucleotides representing about 90% of the molecules, a mode in humans of about 168 nucleotides, and a second, smaller peak in the range of 240 to 440 nucleotides. The cell-free nucleic acids can be about 160 to about 180 nucleotides, or about 320 to about 360 nucleotides, or about 440 to about 480 nucleotides.
[0227] Cell-free nucleic acid can be separated from body fluid by a partitioning step, in which the cell-free nucleic acid as found in the solution is separated from intact cells and other insoluble components of the body fluid. Partitioning can include techniques such as centrifugation or filtration. Alternatively, cells in the body fluid can be lysed, and the cell-free nucleic acid and cellular nucleic acid are processed together. Typically, after adding a buffer and washing steps, the cell-free nucleic acid can be precipitated with alcohol. Further purification steps such as silica-based columns can be used to remove contaminants or salts. For example, non-specific bulk carrier nucleic acid can be added throughout the reaction to optimize certain aspects of the procedure such as yield.
[0228] After such treatment, the sample may include a variety of forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. Optionally, single-stranded DNA and RNA may be converted to double-stranded forms so that they are included in subsequent processing and analysis steps.
[0229] Label
[0230] In some embodiments, nucleic acid molecules (from polynucleotide samples) can be tagged with sample indexes and / or molecular barcodes (commonly referred to as "tags"). Tags can be incorporated into or otherwise connected to adapters by methods such as chemical synthesis, ligation (e.g., blunt end ligation or sticky end ligation) or overlap extension polymerase chain reaction (PCR). These adapters can ultimately be connected to target nucleic acid molecules. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are typically applied to introduce sample indexes into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., more than one microwell in an array). Molecular barcodes and sample indexes can be introduced simultaneously or in any order. In some embodiments, molecular barcodes and / or sample indexes are introduced before and / or after the sequence capture step is performed. In some embodiments, only molecular barcodes are introduced before probe capture, and sample indexes are introduced after the sequence capture step is performed. In some embodiments, both molecular barcodes and sample indexes are introduced before the probe-based capture step is performed. In some embodiments, sample indexes are introduced after the sequence capture step is performed. In some embodiments, the molecular barcode is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample via an adapter via ligation (e.g., blunt end ligation or sticky end ligation). In some embodiments, the sample index is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample by overlap extension polymerase chain reaction (PCR). Typically, sequence capture schemes involve the introduction of a single-stranded nucleic acid molecule complementary to a target nucleic acid sequence, such as a coding sequence of a genomic region, and mutations in such a region are associated with cancer types.
[0231] In some embodiments, the tag can be located at one or both ends of the sample nucleic acid molecule. In some embodiments, the tag is an oligonucleotide of a predetermined sequence, a random sequence, or a semi-random sequence. In some embodiments, the length of the tag can be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotides. The tag can be randomly or non-randomly attached to the sample nucleic acid.
[0232] In some embodiments, each sample is uniquely labeled with a sample index or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or subsample is uniquely labeled with a molecular barcode or a combination of molecular barcodes. In other embodiments, more than one molecular barcode can be used so that the molecular barcodes do not have to be unique to each other (e.g., non-unique molecular barcodes) in more than one. In these embodiments, molecular barcodes are usually attached (e.g., by connection) to individual molecules so that the combination of molecular barcodes and sequences to which they are attached produces a unique sequence that can be traced back individually. The detection of non-unique marker molecular barcodes in conjunction with endogenous sequence information (e.g., corresponding to the beginning (start) and / or end (termination) portion of the sequence of the original nucleic acid molecule in the sample, the subsequence of the sequence read at one or both ends, the length of the sequence read, and / or the length of the original nucleic acid molecule in the sample) is generally allowed to assign a unique identity to a specific molecule. The length or base pair number of the individual sequence reads are also optionally used to assign a unique identity to a given molecule. As described herein, fragments from a single nucleic acid strand have been assigned a unique identity, thereby allowing subsequent identification of fragments from that parental strand and / or the complementary strand.
[0233] In some embodiments, molecular barcodes are introduced into molecules in a sample at an expected ratio of a set of identifiers (e.g., a combination of unique or non-unique molecular barcodes). An example format uses about 2 to about 1,000,000 different molecular barcodes, or about 5 to about 150 different molecular barcodes, or about 20 to about 50 different molecular barcodes, connected to both ends of the target molecule. Alternatively, about 25 to about 1,000,000 different molecular barcodes can be used. For example, 20-50x20-50 molecular barcodes can be used. These number of identifiers are generally sufficient to provide different molecules with the same starting and ending points with a high probability (e.g., at least 94%, 99.5%, 99.99% or 99.999%) of receiving different identifier combinations. In some embodiments, about 80%, about 90%, about 95% or about 99% of molecules have the same molecular barcode combination.
[0234] In some embodiments, the assignment of unique or non-unique molecular barcodes in a reaction is performed using methods and systems described in, for example, U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, only endogenous sequence information (e.g., start and / or end position, subsequences at one or both ends of the sequence, and / or length) can be used to identify different nucleic acid molecules in a sample.
[0235] Amplification
[0236] Sample nucleic acids flanked by adapters can be amplified by PCR and other amplification methods, typically primed from primers that bind to primer binding sites in the adapters flanking the DNA molecule to be amplified. Amplification methods can include cycles of extension, denaturation, and annealing produced by thermal cycling, or can be isothermal cycling, as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication.
[0237] One or more amplifications can be applied to introduce barcodes into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures. Molecular tags and sample indexes / tags can be introduced simultaneously or in any order. Molecular tags and sample indexes / tags can be introduced before and / or after sequence capture. In some cases, only the molecular tags are introduced before probe capture, while the sample indexes / tags are introduced after sequence capture. In some cases, both the molecular tags and sample indexes / tags are introduced before probe capture. In some cases, the sample indexes / tags are introduced after sequence capture. Typically, sequence capture involves introducing a single-stranded nucleic acid molecule complementary to a target sequence, such as a coding sequence of a genomic region, and mutations in such a region are associated with cancer types. Typically, amplification produces more than one non-uniquely or uniquely tagged nucleic acid amplicon, and the molecular tags and sample indexes / tags range in size from 200nt to 700nt, 250nt to 350nt, or 320nt to 550nt. In some embodiments, the amplicon has a size of about 300nt. In some embodiments, the amplicon has a size of about 500nt.
[0238] Enrichment
[0239] In some embodiments, the enrichment sequence is before nucleic acid sequencing. Optionally, specific target regions or non-specifically ("target sequences") are enriched. In some embodiments, the target region of interest can be enriched with nucleic acid capture probes ("bait") selected for one or more bait groups using differential tiling and capture schemes. Differential tiling and capture schemes typically use bait groups of different relative concentrations to differentially tile (e.g., with different "resolutions") on the genomic regions associated with the bait, obey a set of constraints (e.g., sequencer constraints, such as sequencing capacity, the utility of each bait, etc.), and capture target nucleic acids at the level required for downstream sequencing. These target genomic regions of interest optionally include the natural or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, biotin-labeled beads with probes of one or more regions of interest can be used to capture target sequences, and optionally subsequently amplify these regions to enrich the regions of interest.
[0240] Sequence capture generally includes the use of oligonucleotide probes that hybridize with the target nucleic acid sequence.In certain embodiments, the probe set strategy includes tiling the probe on the region of interest.The length of these probes can be, for example, approximately 60 to approximately 120 Nucleotide.The degree of depth of this group can be approximately 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x or more.The effectiveness of sequence capture generally depends in part on the sequence length of the target molecule that is complementary to the probe sequence (or close to complementary).In some embodiments, the colony of enrichment can be amplified before order-checking.
[0241] Sequencing pipeline
[0242] The sample nucleic acid flanked by adapters can be sequenced with or without prior amplification, such as by one or more sequencing devices 107. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrophosphate sequencing, synthetic sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, connection sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, single-molecule synthetic sequencing (SMSS) (Helicos), massively parallel sequencing, cloning single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford nanopore, RocheGenia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent or nanopore platforms. Sequencing reactions can be carried out in a variety of sample processing units, which can be multi-lane, multichannel, multiporous or other devices that substantially process more than one sample group simultaneously. The sample processing unit can also include more than one sample chamber to be able to process more than one run simultaneously.
[0243] Sequencing reactions can be performed on one or more fragment types known to contain markers for cancer or other diseases. Sequencing reactions can also be performed on any nucleic acid fragment present in the sample. The sequence reaction can provide at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% sequence coverage of the genome. In other cases, the sequence coverage of the genome can be less than 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100%.
[0244] Multiple sequencing reactions can be used to perform simultaneous sequencing reactions. In some cases, the cell-free polynucleotides can be sequenced using at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, the cell-free polynucleotides can be sequenced using less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. The sequencing reactions can be performed sequentially or simultaneously. Subsequent data analysis can be performed on all or part of the sequencing reactions. In some cases, data analysis can be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. In other cases, data analysis can be performed on less than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100,000 sequencing reactions. An exemplary read depth is 1000-80000 reads / locus (bases).
[0245] Sequence analysis pipeline
[0246] In some embodiments, the nucleic acids in the sample can be contacted with a sufficient number of adapters comprising molecular barcodes so that the probability that any two copies of the same nucleic acid molecule receive the same combination of molecular barcodes from the adapters connected at both ends is low (e.g., <1% or 0.1%). Using adapters in this manner allows the identification of families of nucleic acid sequences (sequence reads) generated by a given nucleic acid molecule. For example, nucleic acid sequences having the same start and end points on a reference sequence and connected to the same combination of molecular barcodes can be considered to be part of a family. Thus, a family represents the sequence of amplified products of a given nucleic acid molecule in a sample, where family members are sequence reads generated from the amplified products. The sequences of the family members can be compiled to obtain one or more common nucleotides or a complete common sequence of the nucleic acid molecules in the original sample, as modified by blunt end formation and adapter connection. In other words, the nucleotides occupying a specified position of the nucleic acid in the sample are determined to be the common nucleotides occupying the corresponding position in the family member sequence. A family can include the sequence of one or both chains of a double-stranded nucleic acid. If the members of a family include sequences from both chains of a double-stranded nucleic acid, the sequences of one chain are converted into their complementary sequences for the purpose of compiling all sequences to obtain one or more common nucleotides or sequences. Some families only include a single member sequence. In this case, the sequence can be obtained as the sequence of the nucleic acid in the sample before amplification. Alternatively, families with only a single member sequence can be eliminated from subsequent analysis.
[0247] Nucleotide variation in sequenced nucleic acid can be determined by comparing sequenced nucleic acid with reference sequence. Reference sequence is typically a known sequence, for example, a known full genome or partial genome sequence from a subject (for example, a full genome sequence of a human subject). Reference sequence can be, for example, hG19 or hG38. As described above, sequenced nucleic acid can represent a sequence directly determined for the nucleic acid in the sample, or a consensus sequence of an amplification product of such nucleic acid. Comparisons can be made at one or more designated positions on the reference sequence. When the corresponding sequences are aligned to the greatest extent, a subset of sequenced nucleic acid can be identified, including a position corresponding to the designated position of the reference sequence. Within such a subset, it can be determined which (if any) sequenced nucleic acid comprises nucleotide variation at the designated position, and optionally which (if any) comprises reference nucleotides (i.e., the same as in the reference sequence). If the number of sequenced nucleic acids comprising nucleotide variants in the subset exceeds a selected threshold value, variant nucleotides can be called at the designated position. The threshold value can be a simple numerical value, such as at least 1, 2, 3, 4, 5, 6, 7, 9 or 10 nucleic acids through sequencing in the subset that includes nucleotide variants, or it can be a ratio, such as at least 0.5, 1, 2, 3, 4, 5, 10, 15 or 20 of the nucleic acids through sequencing in the subset that includes nucleotide variants, and other possibilities. The comparison can be repeated for any interested designated position in the reference sequence. Sometimes, a designated position that occupies at least about 20, 100, 200 or 300 consecutive positions on the reference sequence, for example, about 20-500 or about 50-300 consecutive positions, can be compared.
[0248] The present methods can be used to identify the presence or absence of genetic events in a subject that may cause a condition, particularly cancer, to characterize the condition (e.g., to stage a cancer or determine the heterogeneity of a cancer), to monitor the response of the condition to treatment, and to provide a prognostic risk of progression of the condition or subsequent course of the condition.
[0249] A variety of cancers can be detected using this method. Cancer cells, like most cells, can be characterized by a turnover rate, in which old cells die and are replaced by newer cells. Typically, dying cells that come into contact with the vascular system in a given subject can release DNA or fragments of DNA into the bloodstream. This is also true for cancer cells during the various stages of the disease. Cancer cells can also be characterized by a variety of genetic aberrations, such as copy number variations and rare mutations, depending on the stage of the disease. This phenomenon can be used to detect the presence or absence of cancer in an individual using the methods and systems described herein.
[0250] The types and numbers of cancers that can be detected may include blood cancer, brain cancer, lung cancer, skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, homogeneous tumors, etc.
[0251] Cancer can be detected from genetic variations including mutations, rare mutations, insertions / deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural changes, gene fusions, chromosome fusions, gene truncations, gene amplifications, gene duplications, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, and abnormal changes in epigenetic patterns.
[0252] Genetic data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and staging. Genetic profile data can allow for the characterization of specific subtypes of cancer, which can be important in the diagnosis or treatment of that specific subtype. This information can also provide a subject or practitioner with prognostic clues about a specific type of cancer and allow the subject or practitioner to adjust treatment options based on the progression of the disease. Some cancers progress, becoming more aggressive and genetically unstable. Other cancers can remain benign, inactive, or dormant. The systems and methods of the present disclosure can be used to determine disease progression.
[0253] This analysis can also be used to determine the effectiveness of a specific treatment option. If the treatment is successful, the successful treatment option can increase the amount of copy number variation or rare mutation detected in the experimenter's blood as more cancer may die and DNA is shed. In other examples, this may not happen. In another example, perhaps some treatment options may be related to the genetic profile of the cancer over time. This correlation can be used to select therapy. In addition, if it is observed that the cancer is in remission after treatment, this method can be used to monitor residual disease or the recurrence of the disease.
[0254] This method can also be used to detect genetic variations in conditions other than cancer. After the presence of certain diseases, immune cells, such as B cells, can undergo rapid clonal expansion. Use copy number variation detection to monitor clonal expansion and to monitor certain immune states. In this example, copy number variation analysis can be carried out over time to produce a spectrum of how a particular disease may progress. Copy number variation or even rare mutation detection can be used to determine how pathogen colonies change during the course of infection. This can be particularly important during chronic infections such as HIV / AIDs or hepatitis infections, where viruses can change life cycle states and / or mutate into more virulent forms during the course of infection. When immune cells attempt to destroy transplanted tissue, this method can be used to determine or analyze the rejection activity of the host body to monitor the state of transplanted tissue and change the course of treatment or prevent rejection.
[0255] In addition, the methods of the present disclosure can be used to characterize the heterogeneity of abnormal conditions in a subject, the method comprising generating a genetic profile of extracellular polynucleotides in the subject, wherein the genetic profile comprises more than one data obtained from copy number variation and rare mutation analysis. In some cases, including but not limited to cancer, the disease can be heterogeneous. The diseased cells can be different. In the example of cancer, some tumors are known to contain different types of tumor cells, some cells are at different stages of cancer. In other examples, heterogeneity can include multiple lesions of the disease. Again, in the example of cancer, there can be multiple tumor lesions, perhaps one or more of which are the result of metastasis that has spread from the primary site.
[0256] The present method can be used to generate or analyze a fingerprint or dataset that is the sum of genetic information from different cells in a heterogeneous disease. The dataset can include copy number variation and rare mutation analysis alone or in combination.
[0257] The present methods can be used to diagnose, prognose, monitor or observe cancer or other diseases of fetal origin. That is, these methods can be used in pregnant subjects to diagnose, prognose, monitor or observe cancer or other diseases in unborn subjects whose DNA and other polynucleotides can co-circulate with maternal molecules.
[0258] Precision treatment examples
[0259] The precise diagnosis provided by the improved computer system 110 can lead to a precise treatment plan, which can be identified by the computer system 110 (and / or planned by a health professional). For example, one type of precise diagnosis and treatment may be related to genes in the homologous recombination repair (HRR) pathway.
[0260] Homologous recombination is a type of genetic recombination in which nucleotide sequences are exchanged between two similar or identical DNA molecules. It is most widely used by cells to precisely repair harmful breaks in both strands of DNA, called double-strand breaks (DSBs). HRR provides a mechanism for error-free removal of damage present in DNA that has already been replicated (S and G2 phases), eliminating chromosome breaks before cell division occurs. The leading models for how homologous recombination repairs double-strand breaks in DNA are the homologous recombination repair pathways, which mediate the double-strand break repair (DSBR) pathway and the synthesis-dependent strand annealing (SDSA) pathway. Germline and somatic defects in homologous recombination genes are strongly associated with breast, ovarian, and prostate cancers.
[0261] The number and type of variant nucleotides in a sample can provide an indication of the suitability of the subject providing the sample for treatment, i.e., therapeutic intervention. For example, various poly ADP ribose polymerase (PARP) inhibitors have been shown to prevent the growth of breast, ovarian, and prostate cancer tumors caused by inherited mutations in the BRCA1 or BRCA2 genes. Some of these therapeutic agents can inhibit base excision repair (BER), which can compensate for the deficiency of HRR.
[0262] On the other hand, some BRCA and HRR wild-type patients may not derive clinical benefit from PARP inhibitor treatment. In addition, not all ovarian cancer patients with BRCA mutations will respond to PARP inhibitors. In addition, different types of mutations may indicate different therapies. For example, a somatic heterozygous deletion of an HRR gene may indicate a different therapy than a somatic homozygous deletion. Therefore, the state of the genetic material may affect the treatment. In one example, a PARP inhibitor can be administered to an individual who contains a somatic homozygous deletion in the HRR gene, but not to an individual who contains a wild-type allele or a somatic heterozygous deletion in the HRR gene.
[0263] Nucleotide variations in sequenced nucleic acids can be determined by comparing the sequenced nucleic acids with a reference sequence. A reference sequence is typically a known sequence, for example, a known full genome or partial genome sequence from a subject, a full genome sequence of a human subject. The reference sequence can be hG19. As described above, sequenced nucleic acids can represent sequences directly determined for nucleic acids in a sample, or a consensus sequence of amplified products of such nucleic acids. Comparisons can be made at one or more designated positions on the reference sequence. When the corresponding sequences are aligned to the greatest extent possible, a subset of sequenced nucleic acids can be identified, including positions corresponding to designated positions of the reference sequence. Within such a subset, it can be determined which (if any) sequenced nucleic acids contain nucleotide variations at designated positions, and optionally which (if any) contain reference nucleotides (i.e., identical to those in the reference sequence). If the number of sequenced nucleic acids for the nucleotide variants contained in the subset exceeds a threshold, the variant nucleotide can be called at the designated position. The threshold value can be a simple numerical value, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleic acids sequenced in the subset comprising the nucleotide variant, or it can be a ratio, such as at least 0.5%, 1%, 2%, 3%, 4%, 5%, 10%, 15% or 20% of the nucleic acids sequenced in the subset comprising the nucleotide variant, as well as other possibilities. The comparison can be repeated for any designated position of interest in the reference sequence. Sometimes, a designated position occupying at least 20, 100, 200 or 300 consecutive positions on the reference sequence, for example, 20-500 or 50-300 consecutive positions, can be compared. Example
[0264] The modeling described herein was applied to plasma samples from 28,199 patients with advanced solid tumors sequenced using Guardant Health, Inc.'s 73-gene next-generation sequencing ctDNA panel.
[0265] Examples of the results show that the sensitivity for detecting BRCA1 / 2 gene deletions is 95% for samples showing a tumor fraction of 9%-11%. The limit of detection for LOH and biallelic copy number loss is 11%-13%. The prevalence of BRCA1 somatic deletions observed in breast, colorectal, prostate, and endometrial cancers is greater than 3%. The prevalence of BRCA2 somatic deletions observed in breast, lung, prostate, head and neck cancer (HNSCC), and hepatocellular carcinoma is greater than 6%.
[0266] In a cohort of 5,568 patients with typical HRD-associated cancers, somatic LOH and biallelic somatic copy number loss were detected in BRCA1 in 2.7% of samples and in BRCA2 in 8.0%, which is consistent with previously reported tissue prevalence. BRCA1 LOH and BRCA2 LOH were observed in 2.4% (134 / 5568) and 7.4% (415 / 5568) of typical homologous recombination-deficient (HRD) cancers, including breast, ovarian, prostate, and pancreatic cancers. In this same cohort of HRD cancers, biallelic somatic copy number loss was observed in 0.3% (19 / 5568) and 0.5% (31 / 5568) of BRCA1 and BRCA2, respectively. Based on the application of the model described here, BRCA1 / 2 somatic LOH and biallelic somatic copy number loss can be accurately detected in ctDNA. The ability to identify such therapeutically targetable genomic alterations through noninvasive ctDNA assessment has important clinical implications, particularly for patients with diseases that are challenging to detect by tissue, such as breast and prostate cancers, which primarily metastasize to bone and brain due to their deep visceral location.
[0267] All patent applications, websites, other publications, registration numbers, etc. cited above or below are incorporated by reference in their entirety for all purposes, as if each individual item was specifically and individually indicated to be incorporated by reference in this manner. If different versions of a sequence are associated with an accession number at different times, it is meant the version associated with that accession number at the time of the effective filing date of this application. The effective filing date means the earlier of the actual filing date or the filing date of the priority application (if applicable) citing that accession number. Similarly, if different versions of a publication, website, etc. are published at different times, it is meant the most recently published version at the time of the effective filing date of the application, unless otherwise indicated. Any feature, step, element, embodiment, or aspect of the present disclosure may be used in combination with any other, unless otherwise specifically indicated. Although the present disclosure has been described in considerable detail by way of illustration and example for the purposes of clarity and understanding, it will be apparent that certain changes and modifications may be made within the scope of the appended claims.
Claims
1. A computer system for distinguishing between somatic homozygous and somatic heterozygous deletions of a gene in a sample that does not exhibit a germline deletion of the gene, the computer system comprising: A processor programmed to: generating a first model of allele counts based on one or more germline single nucleotide polymorphism (SNP) positions associated with the gene via a first probability distribution, the first model representing the somatic homozygous deletion; generating, via a second probability distribution, a second model of allele counts in the sample based on the one or more germline SNP positions, the second model representing the somatic heterozygous deletion; comparing a first output of the first model and a second output of the second model; and generating a prediction of the presence of a somatic homozygous deletion of the gene in the sample based on the comparison, wherein the first model represents a first probability that the sample contains the somatic homozygous deletion, and the second model represents a second probability that the sample contains the somatic heterozygous deletion, wherein to generate the first model of allele counts, the processor is further programmed to: determining the prevalence of heterozygosity for the one or more germline SNPs in the sample training set for input to the first probability distribution, and In order to generate the second model, the processor is further programmed to: A mean of the tumor fractions estimated from the samples is determined for input to a second probability distribution of the second model.
2. The computer system of claim 1, wherein the sample is a cell-free DNA sample.
3. The computer system of claim 1 or 2, wherein the first probability distribution is the same type of probability distribution as the second probability distribution.
4. The computer system of claim 1 or 2, wherein to generate the first model, the processor is programmed to determine one or more parameters for input to the first probability distribution. 5 . The computer system of claim 4 , wherein the first probability distribution comprises a probability distribution type comprising one of: a beta-binomial distribution, a binomial distribution, or a normal distribution.
6. The computer system of claim 4, wherein the sample training set comprises more than one sample in which tumor was not detected (TND).
7. The computer system of claim 4, wherein to generate the first model of allele counts, the processor is further programmed to: A standard deviation of the minor allele frequencies (MAFs) associated with the one or more germline SNPs in the sample training set is determined for input to the first probability distribution.
8. The computer system of claim 7, wherein to generate the first model, the processor is further programmed to: The number of molecules supporting the mutant allele in the sample is determined for input to the first probability distribution.
9. The computer system of claim 8, wherein to generate the first model, the processor is further programmed to: A total number of molecules in the sample is determined for input to the first probability distribution.
10. The computer system of claim 9, wherein to generate the first model, the processor is further programmed to: A first likelihood of allele counts for the one or more germline SNP positions in the sample assuming a somatic homozygous deletion is calculated based on the molecular coverage associated with the somatic homozygous deletion.
11. The computer system of claim 10, wherein to generate the second model, the processor is further programmed to: A second likelihood of allele counts for the one or more germline SNP positions in the sample assuming a somatic heterozygous deletion is calculated based on the molecular coverage associated with the somatic heterozygous deletion.
12. The computer system of claim 4, wherein the tumor score is estimated based on sequence coverage information.
13. The computer system of claim 4, wherein to generate the second model, the processor is further programmed to: A standard deviation of the tumor fraction estimated from the sample is determined for input to the second probability distribution of the second model.
14. The computer system of claim 1 or 2, wherein the processor is further programmed to: Read more than one sample; identifying a group of samples comprising a germline deletion from the more than one samples; and filtering the set of samples from the more than one samples; and The presence of the somatic homozygous deletion or the somatic heterozygous deletion is identified from the filtered more than one sample.
15. The computer system of claim 1, wherein the first output comprises a first probability that the somatic homozygous deletion is present, and the second output comprises a second probability that the somatic heterozygous deletion is present.
16. The computer system of claim 12, wherein to compare the first output of the first model and the second output of the second model, the processor is further programmed to: A log-likelihood function is performed based on the first output and the second output.
17. The computer system of claim 1 or 2, wherein the gene comprises one of the following: BRCA1, BRCA2, and ATM.
18. A method implemented by a processor, the method comprising: generating, by the processor, a first model of allele counts based on one or more germline single nucleotide polymorphism (SNP) positions associated with a gene via a first probability distribution, the first model representing somatic homozygous deletions; generating, by the processor, a second model of allele counts in the sample based on the one or more germline SNP positions via a second probability distribution, the second model representing somatic loss of heterozygosity; comparing, by the processor, a first output of the first model and a second output of the second model; and generating, by the processor, a prediction that the somatic homozygous deletion of the gene is present in the sample based on the comparison, wherein the first model represents a first probability that the sample contains the somatic homozygous deletion, and the second model represents a second probability that the sample contains the somatic heterozygous deletion, wherein to generate the first model of allele counts, the processor is further programmed to: determining the prevalence of heterozygosity for the one or more germline SNPs in the sample training set for input to the first probability distribution, and In order to generate the second model, the processor is further programmed to: A mean of the tumor fractions estimated from the samples is determined for input to a second probability distribution of the second model.
19. The system of any one of claims 1 to 17 or the method of claim 18, further comprising generating a report.
20. The method or system of claim 19, wherein the report comprises information about the status of the genes and / or genetic material in the sample and / or information derived from the status of the genes and / or genetic material in the sample.
21. The method or system of claim 19, further comprising transmitting the report to a third party.
22. The method or system of claim 20, further comprising transmitting the report to a third party.
23. The method or system of claim 21 or 22, wherein the third party is a subject or a healthcare practitioner from whom the sample was obtained.
24. A non-transitory computer-readable medium storing machine-readable instructions that, when executed by a processor, perform the following steps: generating, by the processor, a first model of allele counts based on one or more germline single nucleotide polymorphism (SNP) positions associated with a gene via a first probability distribution, the first model representing somatic homozygous deletions; generating, by the processor, a second model of allele counts in the sample based on the one or more germline SNP positions via a second probability distribution, the second model representing somatic loss of heterozygosity; comparing, by the processor, a first output of the first model and a second output of the second model; and generating, by the processor, a prediction that the somatic homozygous deletion of the gene is present in the sample based on the comparison, wherein the first model represents a first probability that the sample contains the somatic homozygous deletion, and the second model represents a second probability that the sample contains the somatic heterozygous deletion, wherein to generate the first model of allele counts, the processor is further programmed to: determining the prevalence of heterozygosity for the one or more germline SNPs in the sample training set for input to the first probability distribution, and In order to generate the second model, the processor is further programmed to: A mean of the tumor fractions estimated from the samples is determined for input to a second probability distribution of the second model.
Citation Information
Patent Citations
Oligonucleotides
US20010053519A1
Method and apparatus for imaging a sample on a device
US20030152490A1
Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels
US20110160078A1
Oligonucleotides
US6582908B2
Compositions and methods of selective nucleic acid isolation
US7537898B2