Detection of homologous recombination defects based on the methylation status of cell-free nucleic acid molecules
A computational method using the methylation status of cell-free nucleic acids improves cancer detection by predicting homologous recombination deficiencies, enhancing sensitivity and enabling targeted treatments.
Patent Information
- Application Number
- JP2025536227
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-21
- Filing Date
- 2023-12-19
- Publication Date
- 2026-01-14
AI Technical Summary
Existing cancer detection methods using liquid biopsies face challenges due to the low amount and heterogeneous forms of nucleic acids in bodily fluids, making it difficult to accurately classify samples containing tumor-derived DNA.
A computational approach that utilizes the methylation status of cell-free nucleic acid molecules to determine homologous recombination deficiency, employing a predictive model trained on nucleotide sequences with methylated cytosines to classify samples and predict the presence of homology-directed repair deficiencies.
Enhances the sensitivity and accuracy of cancer detection by classifying samples with improved sensitivity, allowing for targeted treatments such as PARP inhibitors based on the presence of homologous recombination repair deficiencies.
Smart Images

Figure 2026501232000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 476,614, filed December 21, 2022, which is incorporated by reference for all purposes. [Background technology]
[0002] background Cancer is a leading cause of disease worldwide. Each year, tens of millions of people worldwide are diagnosed with cancer, and more than half ultimately die from it. In many countries, cancer is the second leading cause of death after cardiovascular disease. Early detection is often associated with improved cancer outcomes.
[0003] Cancer can be caused by the accumulation of genetic variations in an individual's normal cells, resulting in improperly regulated cell division in at least some of them. Such variations generally include copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), epigenetic variations including 5-methylation of cytosine (5-methylcytosine), and association of DNA with chromatin and transcription factors.
[0004] Cancer is often detected by tumor biopsy followed by analysis of cells, markers, or DNA extracted from the cells. However, it has recently been proposed that cancer can also be detected from cell-free nucleic acids in bodily fluids such as blood or urine. Such tests have the advantage of being non-invasive and can be performed without the need to identify suspected cancer cells through a biopsy. However, such tests are complicated by the fact that the amount of nucleic acid in bodily fluids is very low and that nucleic acids exist in heterogeneous forms (e.g., RNA and DNA, single-stranded and double-stranded, and various states of post-replicative modification and association with proteins such as histones).
[0005] Thus, there is a need for improved systems and methods for improved cancer detection using liquid biopsy assays. It is therefore an object of the present disclosure to provide computer-implemented systems and methods with increased sensitivity that have an improved ability to classify samples as containing tumor-derived DNA.
[0006] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain implementations and, together with the written description, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings, which are included by way of example and not by way of limitation. It will be understood that like reference numerals identify like components throughout the drawings, unless the context indicates otherwise. It will also be understood that some or all of the figures may be schematic for purposes of illustration and do not necessarily indicate the actual relative size or location of the depicted elements. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 is a schematic diagram of an example architecture for partitioning nucleic acids based on cytosine methylation in one or more genomic regions of a reference sequence, according to one or more implementations.
[0008] [Figure 2] FIG. 2 is a schematic diagram of an example framework for determining homologous recombination deficiency status based on the methylation status of cell-free nucleic acid molecules using one or more computational models, according to one or more implementations.
[0009] [Figure 3]FIG. 3 is a schematic diagram of an example framework for generating a computational model for determining a subject's homologous recombination deficiency status, according to one or more implementations.
[0010] [Figure 4] FIG. 4 is a flow diagram of an example process, according to one or more implementations, for determining the probability that one or more subjects have a homologous recombination repair deficiency based on methylation data derived from samples obtained from the one or more subjects.
[0011] [Figure 5] FIG. 5 is a block diagram illustrating components of a machine in the form of a computer system that can read and execute instructions from one or more machine-readable media to perform any one or more methodologies described herein, according to one or more example implementations.
[0012] [Figure 6] FIG. 6 is a block diagram illustrating a representative software architecture that may be used in conjunction with one or more hardware architectures described herein, according to one or more example implementations.
[0013] [Figure 7] FIG. 7 is a chart showing genomic regions identified as predictor regions for homology-directed repair deficiency using the techniques described herein.
[0014] [Figure 8] FIG. 8 includes several graphs showing probit values for determining whether a subject is positive for homologous repair deficiency (HRD) or negative for HRD using the techniques described herein for several forms of cancer. DETAILED DESCRIPTION OF THE INVENTION
[0015] overview In one aspect, the method includes obtaining, by a computing system having one or more hardware processors and memory, training sequence data including training sequence representations from a plurality of samples, where each training sequence representation includes a nucleotide sequence corresponding to a fragment of a nucleic acid included in one sample of the plurality of samples, and where each sample of the plurality of samples corresponds to a subject to be classified as having a homologous recombination repair deficiency. The method also includes determining, by the computing system, a subset of the training sequence representations that correspond to nucleic acids having at least a threshold amount of methylated cytosines within one or more regions of the nucleotide sequence. The method further includes analyzing, by the computing system, the subset of training sequence representations to determine quantitative measures derived from the subset of training sequence representations, where each quantitative measure corresponds to a classification region among a plurality of classification regions of a reference genome, and where each classification region of the plurality of classification regions has a threshold amount of methylated cytosines in the subject in which cancer is to be detected. The method further includes analyzing, by a computing system and using one or more computer-based techniques, the quantitative measures of the plurality of classification regions to determine a subset of the plurality of classification regions having at least a threshold likelihood of indicating a homology directed repair deficiency. The method includes generating, by the computing system, a predictive model for determining the probability of a homology directed repair deficiency being present in one or more additional subjects, the predictive model including a plurality of variables and a plurality of weights, each weight of the plurality of weights corresponding to each variable of the plurality of variables, each variable of the plurality of variables corresponding to each classification region of the subset of the plurality of classification regions, and each weight corresponding to each variable indicating the likelihood that the each classification region exhibits a homology directed repair deficiency. The method also includes administering to the subject a treatment suitable for treating the homology directed repair deficiency based on the classification of the subject as having a homology directed repair deficiency.
[0016] In one or more examples, the method may include: analyzing, by a computing system, a subset of the training sequence representations to determine additional quantitative measures derived from the subset of training sequence reads, each quantitative measure corresponding to one control region of a plurality of control regions of the reference genome, each control region of the plurality of control regions having a threshold amount of methylated cytosines in a subject in which cancer is detected and in a further subject in which cancer is not detected; and determining, by the computing system, normalized quantitative measures corresponding to a subset of the plurality of classification regions, each normalized quantitative measure being determined according to the quantitative measure corresponding to the one classification region of the subset of the plurality of classification regions and the additional quantitative measure.
[0017] In various examples, the method may include implementing, by a computing system, a predictive model to determine an individual probability that a homology-directed repair deficiency is present in each sample of the plurality of samples based on a normalized quantitative measure corresponding to the individual sample, and determining, by the computing system, a threshold probability based on the individual probabilities to indicate the presence of a homology-directed repair deficiency for a given subject.
[0018] Further, the method may include determining, by a computing system, responsiveness to a treatment for a group of subjects, wherein cancer is detected in the group of subjects and a treatment is provided to treat the cancer; and determining, by the computing system, a plurality of samples corresponding to subjects having a homologous recombination repair deficiency based on the responsiveness to the treatment of a portion of the group of subjects being at least a threshold level of responsiveness.
[0019] Further, the method may include analyzing, by a computing system, additional sequence reads from samples of a group of subjects in which cancer is detected to determine whether one or more genomic mutations are present with respect to one or more genomic regions, wherein the one or more genomic mutations correspond to a homologous recombination repair pathway; and determining, by the computing system, a plurality of samples to be used to generate training sequence representations by identifying a portion of samples from the group of subjects in which the one or more genomic mutations are present.
[0020] In at least some examples, the one or more computational techniques include implementing one or more logistic regression models with elastic regularization.
[0021] In one or more additional examples, the method may include implementing, by a computing system, the predictive model to determine the probability of the presence of a homologous recombination repair deficiency in a plurality of additional samples, wherein the plurality of additional samples are derived from additional subjects, and wherein a first form of cancer is detected in a first portion of the additional subjects, and a second form of cancer is detected in a second portion of the additional subjects.
[0022] In one or more further examples, the method may include implementing, by a computing system, the predictive model to determine the probability of the presence of a homologous recombination repair deficiency in a plurality of additional samples, wherein the plurality of additional samples are derived from additional subjects presenting with a single form of cancer.
[0023] In one or more examples, the method may include analyzing, by a computing system, a subset of the training sequence reads to determine groups of training sequence reads that correspond to a plurality of genomic regions associated with a homologous recombination repair pathway; and determining, by the computing system, one or more additional quantitative measures based on the number of groups of training sequence representations that correspond to at least a portion of the plurality of genomic regions.
[0024] Additionally, a plurality of the classification regions may have a cytosine-guanine content of at least a threshold amount.
[0025] The method may also include determining, by a computing system, tumor fraction estimates for a number of samples, the number of samples corresponding to subjects in which cancer is detected; analyzing, by the computing system, the tumor fraction estimates with respect to a threshold tumor fraction estimate; and determining, by the computing system, a number of samples to use to obtain training sequence reads based on identification, by the computing system, of at least a portion of the number of samples having tumor fraction estimates that correspond to at least the threshold tumor fraction estimate.
[0026] In various examples, the method may include obtaining, by a computing system, test sequence data from an additional subject not included in the plurality of subjects, wherein the test sequence data includes test sequencing representations from a sample of the additional subject, wherein each test sequencing representation includes a nucleotide sequence corresponding to a fragment of nucleic acid included in the additional sample, and each test sequencing read corresponds to a molecule having a threshold amount of methylated cytosines included within a region of nucleotides; and determining a probability of the presence of a homologous recombination repair deficiency in the additional subject using the predictive model and the additional sequence data.
[0027] Furthermore, the method may include combining a plurality of nucleic acids derived from at least one of a subject's blood or tissue with a solution containing an amount of a methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution, and performing multiple washes of the nucleic acid-MBD protein solution with a salt solution to produce several nucleic acid fractions, each nucleic acid fraction having a threshold number of methylated cytosines within a region of the plurality of nucleic acids having at least a threshold cytosine-guanine content. Other technical features will be readily apparent to those skilled in the art from the following figures, descriptions, and claims.
[0028] In at least some examples, the step of determining a subset of the plurality of classification regions having at least a threshold likelihood of indicating a homologous recombination repair deficiency may include: determining, by a computing system, for each classification region of the plurality of classification regions, a difference between a first portion of a normalized quantitative measure derived from a sample corresponding to a subject in which a homologous recombination repair deficiency is present and a second portion of a normalized quantitative measure derived from a sample corresponding to an additional subject in which a homologous recombination repair deficiency is not present; and determining, by the computing system, that the individual classification region is to be included in the subset of the plurality of classification regions based on the difference between the first portion of the normalized quantitative measure for the individual classification region and the second portion of the normalized quantitative measure for the individual classification region being at least a threshold difference.
[0029] In one or more additional examples, the treatment may include a polyadenosine diphosphate (ADP) ribose polymerase (PARP) inhibitor. The method also includes administering to the additional subject, based on the classification as having a homology directed repair deficiency, a treatment suitable for treating the homology directed repair deficiency.
[0030] In one or more further examples, the method may include determining, by the computing system, an additional subset of training sequence representations corresponding to additional nucleic acids having methylation below an additional threshold amount; analyzing, by the computing system, an additional subset of the training sequence reads to determine an additional group of training sequence representations corresponding to a plurality of genomic regions associated with a homologous recombination repair pathway; and determining, by the computing system, one or more additional quantitative measures based on an additional number of the additional group of training sequence representations corresponding to at least a portion of the plurality of genomic regions.
[0031] The method may also include analyzing, by the computing system, the difference between the one or more additional quantitative measures and the one or more further quantitative measures to determine one or more additional variables for the predictive model.
[0032] The method may further include the steps of: analyzing, by a computing system, the test sequencing reads to determine a first additional quantitative measure corresponding to each classification region of the plurality of classification regions; analyzing, by a computing system, the test sequencing reads to determine a second additional quantitative measure from the test sequencing reads corresponding to each control region of the plurality of control regions, wherein each control region of the plurality of control regions has a threshold amount of methylated cytosines in the subject in which cancer is detected and in an additional subject in which cancer is not detected; determining, by the computing system, additional normalized quantitative measures corresponding to a subset of the plurality of classification regions, wherein each additional normalized quantitative measure is determined according to the first additional quantitative measure and the second additional quantitative measure; and generating, by the computing system, an input vector including the normalized quantitative measures, wherein the predictive model uses the input vector to determine the probability of a homologous recombination repair deficiency being present in the additional subject.
[0033] In one or more examples, one of the washes is performed with a solution having one concentration of sodium chloride (NaCl) to produce one of several nucleic acid fractions having a range of binding energies for the MBD protein.
[0034] In at least some examples, the method may include determining that a first nucleic acid fraction is associated with a first partition of a plurality of partitions of nucleic acids, the first partition corresponding to a first range of binding energies for MBD proteins; causing binding of a first molecular barcode to the nucleic acids of the first nucleic acid fraction, the first molecular barcode being associated with the first partition; determining that a second nucleic acid fraction is associated with a second partition of the plurality of partitions of nucleic acids, the second partition corresponding to a second range of binding energies for MBD proteins that is different from the first range of binding energies for MBD proteins; and causing binding of a second molecular barcode to the nucleic acids of the second nucleic acid fraction, the second molecular barcode being associated with the second partition.
[0035] In various examples, the method may include combining at least a portion of some nucleic acid fractions with an amount of a restriction enzyme that cleaves molecules having one or more unmethylated cytosines to produce at least a portion of a plurality of samples used to generate training sequence representations, wherein a threshold amount of methylated cytosines corresponds to a minimum frequency of methylated cytosines within a region having at least a threshold cytosine-guanine content.
[0036] The method may also include combining at least a portion of some nucleic acid fractions with an amount of a restriction enzyme that cleaves molecules with one or more methylated cytosines to produce at least a portion of a plurality of samples used to generate training sequence representations, wherein the threshold amount of unmethylated cytosines corresponds to the highest frequency of uncleaved methylated cytosines within a region having at least a threshold cytosine-guanine content. The method also includes administering to the subject a treatment suitable for treating the homology-directed repair deficiency based on the classification of the subject as having a homology-directed repair deficiency. The method also includes administering to the subject a treatment suitable for treating the homology-directed repair deficiency based on the classification of the subject as having a homology-directed repair deficiency. Other technical features will be readily apparent to those skilled in the art from the following figures, descriptions, and claims.
[0037] In one aspect, the computing device includes a processor, which, when executed by the processor, performs the following steps: obtaining training sequence data including training sequence representations from a plurality of samples, wherein each training sequence representation comprises a nucleotide sequence corresponding to a fragment of a nucleic acid included in one sample of the plurality of samples, wherein each sample of the plurality of samples corresponds to a subject classified as having a homologous recombination repair deficiency; determining a subset of the training sequence representations corresponding to nucleic acids having at least a threshold amount of methylated cytosines within one or more regions of the nucleotide sequence; and analyzing the subset of training sequence representations to determine quantitative measures derived from the subset of training sequence representations, wherein each quantitative measure corresponds to a classification region of a plurality of classification regions of a reference genome, wherein each classification region of the plurality of classification regions corresponds to a region where cancer is detected. The apparatus also includes a memory having stored thereon instructions that configure the apparatus to: determine that a subset of the plurality of classification regions has a threshold amount of methylated cytosines in the selected subject; analyze the quantitative measures of the plurality of classification regions using one or more computer-based techniques to determine a subset of the plurality of classification regions having at least a threshold likelihood of being indicative of a homology-directed repair deficiency; and generate a predictive model for determining the probability of a homology-directed repair deficiency being present in one or more additional subjects, the predictive model including a plurality of variables and a plurality of weights, each weight of the plurality of weights corresponding to a respective variable of the plurality of variables, each variable of the plurality of variables corresponding to a respective classification region of the subset of the plurality of classification regions, and each weight corresponding to a respective variable indicating a likelihood that the respective classification region exhibits a homology-directed repair deficiency.
[0038] The computing device may also include additional instructions that, when executed by the processor, configure the device to: analyze a subset of the training sequence representations to determine additional quantitative measures derived from the subset of training sequence reads, each quantitative measure corresponding to a control region of a plurality of control regions of the reference genome, each control region of the plurality of control regions having a threshold amount of methylated cytosines in a subject in which cancer is detected and in a further subject in which cancer is not detected; and determine normalized quantitative measures corresponding to a subset of the plurality of classification regions, each normalized quantitative measure determined according to the quantitative measure corresponding to the one classification region of the subset of the plurality of classification regions and the additional quantitative measure.
[0039] Additionally, the computing device may include additional instructions that, when executed by the processor, configure the device to implement the predictive model to determine individual probabilities of a homology-directed repair deficiency being present in each sample of the plurality of samples based on the normalized quantitative measure corresponding to the individual sample, and to determine, by the computing system, a threshold probability based on the individual probabilities to indicate the presence of a homology-directed repair deficiency for a given subject.
[0040] Additionally, the computing device may include additional instructions that, when executed by the processor, configure the device to: determine responsiveness to a treatment for a group of subjects, wherein cancer is detected in the group of subjects and a treatment is provided to treat the cancer; and determine a plurality of samples corresponding to subjects having a homologous recombination repair deficiency based on a portion of the group of subjects having at least a threshold level of responsiveness to the treatment.
[0041] In one or more examples, the computing device may include additional instructions that, when executed by the processor, configure the device to: analyze additional sequence reads from samples of a group of subjects in which cancer is detected to determine whether one or more genomic mutations are present with respect to one or more genomic regions, wherein the one or more genomic mutations correspond to a homologous recombination repair pathway; and determine a plurality of samples to be used to generate training sequence representations by identifying a portion of the samples from the group of subjects in which the one or more genomic mutations are present.
[0042] In various examples, the one or more computational techniques include implementing one or more logistic regression models with elastic regularization.
[0043] In at least some examples, the computing device may include additional instructions that, when executed by the processor, configure the device to implement the predictive model to determine the probability of the presence of a homologous recombination repair deficiency in a plurality of additional samples, wherein the plurality of additional samples are derived from additional subjects having a first form of cancer detected in a first portion of the additional subjects and a second form of cancer detected in a second portion of the additional subjects.
[0044] In one or more additional examples, the computing device may include additional instructions that, when executed by the processor, configure the device to implement the predictive model to determine the probability of a homologous recombination repair deficiency being present in a plurality of additional samples, wherein the plurality of additional samples are derived from additional subjects presenting with a single form of cancer.
[0045] In one or more further examples, the computing device may include additional instructions that, when executed by the processor, configure the device to analyze a subset of the training sequence reads to determine groups of training sequence reads that correspond to a plurality of genomic regions associated with a homologous recombination repair pathway, and to determine one or more additional quantitative measures based on the number of groups of training sequence reads that correspond to at least a portion of the plurality of genomic regions.
[0046] In various examples, the plurality of classification regions have a cytosine-guanine content of at least a threshold amount.
[0047] The computing device may also include additional instructions that, when executed by the processor, configure the device to: determine tumor fraction estimates for a number of samples, where a number of samples correspond to subjects in which cancer is detected; analyze the tumor fraction estimates with respect to a threshold tumor fraction estimate; and determine, by the computing system, a number of samples to use for obtaining training sequence reads based on identification, by the computing system, of at least a portion of the number of samples having tumor fraction estimates that correspond to at least the threshold tumor fraction estimate.
[0048] Additionally, the computing device may include additional instructions that, when executed by the processor, configure the device to obtain test sequence data from an additional subject not included in the plurality of subjects, wherein the test sequence data includes test sequencing indications from a sample of the additional subject, wherein each test sequencing indication includes a nucleotide sequence corresponding to a fragment of nucleic acid included in the additional sample, and each test sequencing read corresponds to a molecule having a threshold amount of methylated cytosines included within a region of nucleotides; and use the predictive model and the additional sequence data to determine the probability of a homologous recombination repair deficiency being present in the additional subject.
[0049] The computing device may also include additional instructions that, when executed by the processor, configure the device to: determine, for each classification region of the plurality of classification regions, a difference between a first portion of the normalized quantitative measure derived from a sample corresponding to a subject in which a homologous recombination repair deficiency is present and a second portion of the normalized quantitative measure derived from a sample corresponding to an additional subject in which a homologous recombination repair deficiency is not present; and determine that the each classification region is to be included in a subset of the plurality of classification regions based on the difference between the first portion of the normalized quantitative measure for the each classification region and the second portion of the normalized quantitative measure for the each classification region being at least a threshold difference.
[0050] In one or more examples, the treatment is a polyadenosine diphosphate (ADP) ribose polymerase (PARP) inhibitor.
[0051] Additionally, the computing device may include additional instructions that, when executed by the processor, configure the device to: determine additional subsets of training sequence representations corresponding to additional nucleic acids having methylation below an additional threshold amount; analyze the additional subsets of training sequence reads to determine additional groups of training sequence representations corresponding to a plurality of genomic regions associated with a homology-directed repair pathway; and determine one or more additional quantitative measures based on additional numbers of the additional groups of training sequence representations corresponding to at least a portion of the plurality of genomic regions.
[0052] Additionally, the computing device may include additional instructions that, when executed by the processor, configure the device to analyze differences between the one or more additional quantitative measures and the one or more further quantitative measures to determine one or more additional variables for the predictive model.
[0053] In various examples, the computing device may include additional instructions that, when executed by the processor, configure the device to: analyze the test sequencing reads to determine a first additional quantitative measure corresponding to each classification region of the plurality of classification regions; analyze the test sequencing reads to determine a second additional quantitative measure from the test sequencing reads corresponding to each control region of the plurality of control regions, wherein each control region of the plurality of control regions has a threshold amount of methylated cytosines in a subject in which cancer is detected and in an additional subject in which cancer is not detected; determine additional normalized quantitative measures corresponding to a subset of the plurality of classification regions, wherein each additional normalized quantitative measure is determined according to the first additional quantitative measure and the second additional quantitative measure; and generate an input vector including the normalized quantitative measures, wherein a predictive model uses the input vector to determine the probability that a homologous recombination repair deficiency is present in the additional subject.
[0054] In one or more examples, the computing device includes additional instructions that, when executed by the hardware processor, cause the hardware processor to: determine, for each classification region of the plurality of classification regions, a difference between a first portion of the normalized quantitative measure derived from a sample corresponding to a subject in which a homologous recombination repair deficiency is present and a second portion of the normalized quantitative measure derived from a sample corresponding to an additional subject in which a homologous recombination repair deficiency is not present; and determine that the each classification region is to be included in a subset of the plurality of classification regions based on the difference between the first portion of the normalized quantitative measure for the each classification region and the second portion of the normalized quantitative measure for the each classification region being at least a threshold difference.
[0055] In one aspect, the one or more non-transitory computer readable storage media, when executed by a computer, cause the computer to: obtain training sequence data comprising training sequence representations from a plurality of samples, wherein each training sequence representation comprises a nucleotide sequence corresponding to a fragment of a nucleic acid included in one sample of the plurality of samples, wherein each individual sample of the plurality of samples corresponds to a subject classified as having a homology directed repair deficiency; determine a subset of the training sequence representations that correspond to nucleic acids having at least a threshold amount of methylated cytosines within one or more regions of the nucleotide sequence; and analyze the subset of training sequence representations to determine quantitative measures derived from the subset of training sequence representations, wherein each quantitative measure corresponds to a classification region of a plurality of classification regions of a reference genome. In response, the method includes instructions to: determine that each classification region of the plurality of classification regions has a threshold amount of methylated cytosines in a subject in which cancer is detected; analyze the quantitative measures of the plurality of classification regions using one or more computer-based techniques to determine a subset of the plurality of classification regions having at least a threshold likelihood of being indicative of a homology-directed repair deficiency; and generate a predictive model for determining the probability of a homology-directed repair deficiency being present in one or more additional subjects, the predictive model comprising a plurality of variables and a plurality of weights, each weight of the plurality of weights corresponding to a respective variable of the plurality of variables, each variable of the plurality of variables corresponding to a respective classification region of the subset of the plurality of classification regions, and each weight corresponding to a respective variable indicating a likelihood that the respective classification region will exhibit a homology-directed repair deficiency.
[0056] The one or more non-transitory computer-readable storage media may also include instructions that, when executed by a computer, cause the computer to: analyze a subset of the training sequence representations to determine additional quantitative measures derived from the subset of training sequence reads, wherein each quantitative measure corresponds to a control region of a plurality of control regions of a reference genome, each control region of the plurality of control regions having a threshold amount of methylated cytosines in a subject in which cancer is detected and in an additional subject in which cancer is not detected; and determine normalized quantitative measures corresponding to a subset of the plurality of classification regions, wherein each normalized quantitative measure is determined according to the quantitative measure corresponding to the one classification region of the subset of the plurality of classification regions and the additional quantitative measure.
[0057] Further, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to implement a predictive model to determine individual probabilities of the presence of a homology-directed repair deficiency in each sample of the plurality of samples based on normalized quantitative measures corresponding to the individual samples, and to determine a threshold probability based on the individual probabilities to indicate the presence of a homology-directed repair deficiency for a given subject.
[0058] Further, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to determine responsiveness to a treatment for a group of subjects, wherein cancer is detected in the group of subjects and a treatment is provided to treat the cancer, and determine a plurality of samples that correspond to subjects having a homologous recombination repair deficiency based on the responsiveness of a portion of the group of subjects to the treatment being at least a threshold level of responsiveness.
[0059] In one or more examples, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to analyze additional sequence reads from samples of a group of subjects in which cancer is detected to determine whether one or more genomic mutations are present with respect to one or more genomic regions, wherein the one or more genomic mutations correspond to a homologous recombination repair pathway, and determine a plurality of samples to be used to generate training sequence representations by identifying a portion of the samples from the group of subjects in which the one or more genomic mutations are present.
[0060] The one or more computational techniques include implementing one or more logistic regression models with elastic regularization.
[0061] In one or more additional examples, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to implement a predictive model to determine the probability of the presence of a homologous recombination repair deficiency in a plurality of additional samples, wherein the plurality of additional samples are derived from additional subjects having a first form of cancer detected in a first portion of the additional subjects and a second form of cancer detected in a second portion of the additional subjects.
[0062] In one or more further examples, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to implement a predictive model to determine the probability of a homologous recombination repair deficiency being present in a plurality of additional samples, wherein the plurality of additional samples are derived from additional subjects presenting with a single form of cancer.
[0063] In at least some examples, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to analyze a subset of the training sequence reads to determine groups of training sequence reads that correspond to a plurality of genomic regions associated with a homology-directed repair pathway, and determine one or more additional quantitative measures based on the number of groups of training sequence representations that correspond to at least a portion of the plurality of genomic regions.
[0064] In various examples, the plurality of classification regions have a cytosine-guanine content of at least a threshold amount.
[0065] The one or more non-transitory computer-readable storage media may also include instructions that, when executed by a computer, cause the computer to determine tumor fraction estimates for a number of samples, where the number of samples correspond to subjects in which cancer is detected; analyze the tumor fraction estimates with respect to a threshold tumor fraction estimate; and determine a number of samples to use to obtain training sequence reads based on identifying at least a portion of the number of samples having tumor fraction estimates that correspond to at least the threshold tumor fraction estimate.
[0066] Further, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to obtain test sequence data from an additional subject not included in the plurality of subjects, wherein the test sequence data includes test sequencing representations from samples of the additional subject, each test sequencing representation including a nucleotide sequence corresponding to a fragment of nucleic acid included in the additional sample, and each test sequencing read corresponding to a molecule having a threshold amount of methylated cytosines included within a region of nucleotides; and use the predictive model and the additional sequence data to determine the probability that a homologous recombination repair deficiency is present in the additional subject.
[0067] Treatment may include polyadenosine diphosphate (ADP) ribose polymerase (PARP) inhibitors.
[0068] The one or more non-transitory computer-readable storage media may also include instructions that, when executed by the computer, cause the computer to determine additional subsets of training sequence representations corresponding to additional nucleic acids having methylation below an additional threshold amount; analyze, by the computing system, the additional subsets of training sequence reads to determine additional groups of training sequence representations corresponding to a plurality of genomic regions associated with a homology-directed repair pathway; and determine one or more further quantitative measures based on additional numbers of the additional groups of training sequence representations corresponding to at least a portion of the plurality of genomic regions.
[0069] Additionally, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to analyze differences between the one or more additional quantitative measures and the one or more further quantitative measures to determine one or more additional variables for the predictive model.
[0070] Further, the one or more non-transitory computer-readable storage media may include instructions that, when executed by a computer, cause the computer to: analyze the test sequencing reads to determine a first additional quantitative measure corresponding to each classification region of the plurality of classification regions; analyze, by the computing system, the test sequencing reads to determine a second additional quantitative measure from the test sequencing reads corresponding to each control region of the plurality of control regions, wherein each control region of the plurality of control regions has a threshold amount of methylated cytosines in the subject in which cancer is detected and in an additional subject in which cancer is not detected; determine additional normalized quantitative measures corresponding to a subset of the plurality of classification regions, wherein each additional normalized quantitative measure is determined according to the first additional quantitative measure and the second additional quantitative measure; and generate an input vector including the normalized quantitative measures, wherein the predictive model uses the input vector to determine the probability that a homologous recombination repair deficiency is present in the additional subject.
[0071] definition In order to more readily understand this disclosure, certain terms are first defined below. Additional definitions for these and other terms may be found throughout the specification. In the event that a definition of a term below conflicts with a definition in an application or patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.
[0072] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or that will become apparent to those skilled in the art upon reading this disclosure, and so forth.
[0073] It should also be understood that the terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting. Furthermore, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In describing and claiming the methods, computer-readable media, and systems, the following terminology and grammatical variations thereof will be used in accordance with the definitions set forth below.
[0074] About: As used herein, "about" or "approximately," when applied to one or more values or elements of interest, refers to a value or element that is similar to the specified reference value or element. In certain implementations, the term "about" or "approximately," unless otherwise specified or otherwise clear from the context, refers to a set of values or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or fewer percent in either direction (greater or less) of the specified reference value or element (except where such number exceeds 100% of the possible values or elements).
[0075] Administer: As used herein, "administering" or "administering" a therapeutic agent (e.g., an immunological therapeutic agent) to a subject means giving, applying, or contacting the composition to the subject. Administration may be accomplished by any of several routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal.
[0076] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that can be used to ligate one or both ends of a given sample nucleic acid molecule that is at least partially double-stranded. The adapter may include nucleic acid primer binding sites to enable amplification of the nucleic acid molecule flanked by the adapters at both ends, and / or sequencing primer binding sites, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. The adapter may also include a binding site for a capture probe, such as an oligonucleotide attached to a flow cell support or the like. The adapter may also include a nucleic acid tag, as described herein. The nucleic acid tag can be positioned relative to the amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequence read of a given nucleic acid molecule. The same or different adapters can be ligated to each end of a nucleic acid molecule. In some implementations, the same adapter is ligated to each end of a nucleic acid molecule, except for the different nucleic acid tags. In some implementations, the adapters are Y-shaped adapters with one end blunt or tailed for joining to a similarly blunt or tailed nucleic acid molecule with one or more complementary nucleotides, as described herein. In yet another exemplary implementation, the adapters are bell-shaped adapters with a blunt or tailed end for joining to a nucleic acid molecule to be analyzed. Other examples of adapters include T-tailed adapters and C-tailed adapters.
[0077] Alignment: As used herein, "alignment" or "aligning" refers to determining whether at least two sequence representations have at least a threshold amount of homology. In one or more examples, the threshold amount of homology can be at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9%. In situations where two sequence representations have at least a threshold amount of homology, the two sequence representations can be referred to as "aligned."
[0078] Amplify: As used herein, "amplify" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or a portion of a polynucleotide, starting from a small amount of polynucleotide (e.g., a single polynucleotide molecule), where an amplification product or amplicon is generally detectable. Polynucleotide amplification encompasses a variety of chemical and enzymatic processes.
[0079] Barcode: As used herein, "barcode" or "molecular barcode" in the context of nucleic acids refers to a nucleic acid molecule that contains a sequence that can serve as a molecular identifier. For example, an individual "barcode" sequence can be added to each DNA fragment during next-generation sequencing (NGS) library preparation, thus allowing each read to be identified and sorted prior to final data analysis.
[0080] Cancer type: As used herein, "cancer type" refers to the type or subtype of cancer as defined, for example, by histopathology. Cancer type can be determined by any conventional criteria, for example, by occurrence in a given tissue (e.g., blood cancer, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, intestinal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, etc. Cancers can be defined as cancers of unknown primary origin, etc., and / or based on the same cellular lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), and / or exhibiting cancer markers such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors, and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether primary or secondary.
[0081] Carrier Signal: As used herein, "carrier signal" refers to any intangible medium capable of storing, encoding, or carrying transitory or non-transitory instructions 502 for execution by machine 500, including digital or analog communication signals or other intangible media for facilitating communication of such instructions 502. Instructions 502 may be transmitted or received over network 534 using transitory or non-transitory transmission media, via network interface devices, and using any one of several well-known transfer protocols.
[0082] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acid that is not contained within or otherwise associated with cells, or, in some implementations, nucleic acid that remains in a sample after removal of intact cells. Cell-free nucleic acid can include, for example, all unencapsulated nucleic acids originating from a subject's bodily fluid (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acid can be double-stranded, single-stranded, or a hybrid thereof. Cell-free nucleic acid can be released into bodily fluids by secretion or cell death processes, such as cell necrosis, apoptosis, etc. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released from cancer cells into body fluids. Others are released from healthy cells. ctDNA can be unencapsulated tumor-derived fragmented DNA. Cell-free nucleic acids can have one or more epigenetic modifications, for example, cell-free nucleic acids can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.
[0083] Cellular nucleic acid: As used herein, "cellular nucleic acid" refers to nucleic acid that is located within one or more cells, at least at the time the sample is taken or collected from a subject, even if the nucleic acid is later removed as part of a given analytical process.
[0084] Classification region: As used herein, " classification region " refers to a genomic region that can show sequence-independent changes in neoplastic cells (such as tumor cells and cancer cells), or that can show sequence-independent changes in the cfDNA of subjects with cancer compared with the cfDNA of subjects without cancer. " Classification region " can also refer to the genomic region that is related to homologous recombination pathway. Examples of sequence-independent changes include, but are not limited to, changes in methylation rates (increase or decrease), nucleosome distribution, CTCF binding, transcription start sites, and regulatory protein binding regions. Classification regions can be enriched by one or more probes. Furthermore, classification regions can be defined by pairs of primer binding sites. Furthermore, classification regions can be defined by a predetermined start genomic locus and a predetermined end genomic locus. Classification regions can contain about 25 nucleotides to about 250 nucleotides, about 50 nucleotides to about 200 nucleotides, or about 75 nucleotides to about 150 nucleotides.
[0085] Communications Network: As used herein, a "communications network" refers to one or more portions of a network 114, 1034, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a plain old telephone service (POTS) network, a cellular network, a wireless network, a Wi-Fi network, another type of network, or a combination of two or more such networks. For example, the network 114, 1034 or a portion of a network may include a wireless or cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, coupling is performed to connect the third generation partnership project, including Single Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, and 3G. The wireless communication network may implement any of various types of data transfer technologies, such as the Third Generation Partnership Project (3GPP®), Fourth Generation Wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE) standards, other technologies defined by various standards-setting organizations, other long-range protocols, or other data transfer technologies.
[0086] Coverage: As used herein, "coverage" or "coverage metrics" refers to the number of nucleic acid molecules or sequencing reads that correspond to a particular genomic region of a reference sequence.
[0087] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to natural or modified nucleotides that have a hydrogen group at the 2' position of the sugar moiety. DNA can contain a chain of nucleotides containing four types of nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to natural or modified nucleotides that have a hydroxyl group at the 2' position of the sugar moiety. RNA can contain a chain of nucleotides containing four types of nucleotides: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to natural or modified nucleotides. Certain pairs of nucleotides specifically bind to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides complementary to those of the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "sequence representation," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," "fragment sequence," "sequencing read," or "nucleic acid sequencing read" refers to any information or data that indicates the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule of nucleic acid, such as DNA or RNA (e.g., a whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment).It should be understood that the present teachings contemplate sequence information obtained using all available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.
[0088] Driver mutation: As used herein, "driver mutation" means a mutation that drives cancer progression.
[0089] Homologous recombination deficiency: As used herein, "homologous recombination deficiency" or HRD or "homologous recombination repair deficiency" refers to the inability of a cell to efficiently repair double-stranded DNA breaks using the homologous recombination repair pathway. HRD is typically characterized by mutations in one or more genomic regions that regulate the homologous recombination repair pathway.
[0090] Homologous recombination repair pathway: As used herein, "homologous recombination repair pathway" refers to one or more processes that use a group of proteins to repair damage to DNA caused by double-strand breaks.
[0091] Hypermethylated: As used herein, "hypermethylated" refers to an increased level or degree of methylation of a nucleic acid molecule(s) relative to other nucleic acid molecules within a population (e.g., a sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypermethylated DNA may include DNA molecules that contain at least one methylated cytosine, at least two methylated cytosines, at least three methylated cytosines, at least five methylated cytosines, or at least ten methylated cytosines.
[0092] Hypomethylated: As used herein, "hypomethylated" refers to a reduced level or degree of methylation of a nucleic acid molecule(s) relative to other nucleic acid molecules within a population (e.g., a sample) of nucleic acid molecules from the same genomic locus. In some embodiments, hypomethylated DNA includes unmethylated DNA molecules. In some embodiments, hypomethylated DNA may include DNA molecules containing zero methylated cytosines, at most one methylated cytosine, at most two methylated cytosines, at most three methylated cytosines, at most four methylated cytosines, or at most five methylated cytosines.
[0093] Immunotherapy: As used herein, "immunotherapy" refers to treatment with one or more agents that stimulate the immune system to kill or at least inhibit the growth of cancer cells, preferably to reduce further cancer growth, shrink the size of cancer, and / or eliminate cancer. Some such agents bind to targets present in cancer cells, some bind to targets present in immune cells but not in cancer cells, and some bind to targets present in both cancer cells and immune cells. Such agents include, but are not limited to, checkpoint inhibitors and / or antibodies. Checkpoint inhibitors are inhibitors of immune system pathways that maintain self-tolerance and modulate the duration and magnitude of physiological immune responses in peripheral tissues, minimizing secondary tissue damage (see, e.g., Pardoll, Nature Reviews Cancer 12, 252-264 (2012)). Examples of agents include antibodies against PD-1, PD-2, PD-L1, PD-L2, CTLA-40, OX40, B7.1, B7He, LAG3, CD137, KIR, CCR5, CD27, or CD40. Other examples of agents include pro-inflammatory cytokines such as IL-1β, IL-6, and TNF-α. Another example of an agent is T cells activated against tumors, for example, T cells activated by expressing a chimeric antigen that targets a tumor antigen recognized by the T cell.
[0094] Indel: As used herein, "indel" refers to a mutation involving the insertion or deletion of nucleotides in the genome of a subject.
[0095] Machine-readable medium: As used herein, "machine-readable medium" refers to a component, device, or other tangible medium capable of temporarily or permanently storing instructions 502 and data, which may include, but are not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage (e.g., erasable programmable read-only memory (EEPROM)), and / or any suitable combination thereof. The term "machine-readable medium" may be considered to include a single medium or multiple media (e.g., centralized or distributed databases, or associated caches and servers) capable of storing instructions 502. The term "machine-readable medium" shall also be considered to include any medium or combination of media capable of storing instructions 502 (e.g., code) for execution by a machine 500 such that the instructions 502, when executed by one or more processors 504 of the machine 500, cause the machine 500 to perform any one or more of the methodologies described herein. Thus, a "machine-readable medium" refers to a single storage device or apparatus, as well as a "cloud-based" storage system or storage network that includes multiple storage devices or apparatuses. Signals themselves are excluded from the term "machine-readable medium."
[0096] Maximum MAF: As used herein, "maximum MAF" or "max MAF" refers to the maximum MAF of all somatic variants in a sample.
[0097] Methylation: As used herein, the term "methylation" or "DNA methylation" refers to the addition of a methyl group to a nucleotide base in a nucleic acid molecule. In some embodiments, methylation refers to the addition of a methyl group to a cytosine at a CpG site (i.e., a cytosine followed by a guanine in the 5' to 3' direction of a nucleic acid sequence). In some embodiments, DNA methylation refers to the addition of a methyl group to an adenine, e.g., N 6 -methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the six-carbon ring of cytosine). In some embodiments, 5-methylation refers to the addition of a methyl group to the 5C position of cytosine to create 5-methylcytosine (5mC). In some embodiments, methylation includes derivatives of 5mC. Derivatives of 5mC include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the six-carbon ring of cytosine). In some embodiments, 3C methylation includes the addition of a methyl group to the 3C position of cytosine to create 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA within a promoter region is methylated, gene transcription can be suppressed. DNA methylation is crucial for normal development, and abnormal methylation can disrupt epigenetic regulation. Disruption, e.g., suppression, in epigenetic regulation can cause diseases, such as cancer. Promoter methylation in DNA can indicate cancer.
[0098] Methylation-dependent nuclease: As used herein, the term "methylation-dependent nuclease" refers to a nuclease that preferentially cleaves methylated DNA compared to unmethylated DNA. For example, a methylation-dependent nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the methylation of at least one of the nucleic acid bases, e.g., cytosine, within the recognition sequence. In some embodiments, the nucleolytic activity of a methylation-dependent nuclease is at least 10-fold, 20-fold, 50-fold, or 100-fold higher for a methylated recognition site compared to an unmethylated control in a standard nucleolytic assay. Methylation-dependent nucleases include methylation-dependent restriction enzymes.
[0099] Methylation-dependent restriction enzyme: As used herein, "methylation-dependent restriction enzyme" or "MDRE" refers to a restriction enzyme that depends on DNA methylation (e.g., cytosine methylation), i.e., the presence or absence of methyl groups in nucleotide bases alters the rate at which the enzyme cleaves target DNA. In some embodiments, a methylation-dependent restriction enzyme does not cleave DNA if a particular nucleotide base is unmethylated in the recognition sequence. For example, MspJI is a methylation-dependent restriction enzyme with the recognition sequence "mCNNR(N9)," and does not cleave DNA if methylated cytosine (mC) is absent within the recognition sequence.
[0100] Methylation-sensitive nuclease: As used herein, the term "methylation-sensitive nuclease" refers to a nuclease that preferentially cleaves unmethylated DNA compared to methylated DNA. For example, a methylation-sensitive nuclease can cleave at or near a recognition sequence, such as a restriction site, in a manner that depends on the lack of methylation of at least one of the nucleic acid bases, e.g., cytosine, within the recognition sequence. In some embodiments, the nucleolytic activity of a methylation-sensitive nuclease is at least 10-fold, 20-fold, 50-fold, or 100-fold higher for an unmethylated recognition sequence than for a methylated control in a standard nucleolytic assay. Methylation-sensitive nucleases include methylation-sensitive restriction enzymes.
[0101] Methylation-sensitive restriction enzyme: As used herein, "methylation-sensitive restriction enzyme" or "MSRE" refers to a restriction enzyme that is sensitive to the methylation state of DNA (e.g., cytosine methylation), i.e., the presence or absence of a methyl group at a nucleotide base alters the rate at which the enzyme cleaves a target DNA. In some embodiments, a methylation-sensitive restriction enzyme does not cleave DNA if a particular nucleotide base is methylated in the recognition sequence. For example, HpaII is a methylation-sensitive restriction enzyme with the recognition sequence "CCGG," and does not cleave DNA if the second cytosine in the recognition sequence is methylated.
[0102] Methylation rate: As used herein, "methylation rate" refers to the probability, likelihood, or percentage that a given base (e.g., a cytosine residue of a CpG) on a DNA molecule in a particular genomic region analyzed in a sample is methylated. In some embodiments, the methylation rate can be applied to a defined region containing one or more potentially methylated bases. In some embodiments, the methylation rate refers to the percentage of methylated CpG residues in a DNA molecule. In some embodiments, the methylation rate refers to the percentage of methylated CpG residues in molecules aligned to a particular genomic location or genomic region. Methylation rates can be measured by various methods, including, but not limited to, using either bisulfite sequencing (any single-base resolution, such as TAPS, EM-SEQ, etc.) or partitioning (DNA molecule resolution). Methylation rates can be measured in various ways. One estimate can be by counting how many DNA fragments end up in each methylation-dependent partition, or, in the case of bisulfite sequencing, by counting the number of converted CpGs per fragment. Furthermore, in the case of methylation-dependent partitioning, rate calculations can be normalized using a set of predefined regions with known methylation states or spike-in synthetic DNA with known methylation states to obtain a rate-parameterized partition distribution and estimate rates using a maximum likelihood approach.
[0103] Methylation state: As used herein, "methylation state" can refer to the presence or absence of a methyl group on a DNA base (e.g., cytosine) at a specific genomic position in a nucleic acid molecule. It can also refer to the degree of methylation in a nucleic acid sequence (e.g., hypermethylated, hypomethylated, intermediately methylated, or unmethylated nucleic acid molecule). Methylation state can also refer to the number of methylated nucleotides in a particular nucleic acid molecule.
[0104] Modified nucleotide-specific binding reagent: As used herein, refers to a binding reagent that is specific for or targets a modified nucleotide. For example, the modified nucleotide may be a methylated nucleotide, and thus the binding reagent may be specific for methylated nucleotides. Examples of binding reagents include, but are not limited to, the methyl-binding domain (MBD) of a methylation-binding protein ("MBP") or variants thereof, antibodies (and antibody variants, e.g., single-chain antibodies), aptamers, or combinations thereof. Thus, as disclosed throughout, the use of an MBD can be interchanged with any other modified nucleotide-specific binding reagent, provided that the modified nucleotide-specific binding reagent has the desired specificity and affinity for the particular modified base of interest in the selected implementation.
[0105] Variant allele fraction: As used herein, "variant allele fraction," "mutation dose," or "MAF" refers to the fraction of nucleic acid molecules that have an allelic alteration or mutation at a given genomic location in a given sample. The MAF is generally expressed as a fraction or percentage. For example, the MAF can be less than about 0.5, less than 0.1, less than 0.05, or less than 0.01 (i.e., less than about 50%, less than 10%, less than 5%, or less than 1%) of all somatic variants or alleles present at a given locus.
[0106] Mutation: As used herein, "mutation" refers to a variation from a known reference sequence, including, for example, single nucleotide variants (SNVs), copy number variants or variations (CNVs) / abnormalities, insertions or deletions (indels), gene fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants.Mutation can be a germline mutation or a somatic mutation.In some cases, the reference sequence for comparison is the wild-type genome sequence of the species of the subject that provides the test sample, typically the human genome.
[0107] Variation caller: As used herein, "variation caller" refers to an algorithm (embodied in software or otherwise computer-implemented) used to identify variants in test sample data (e.g., sequence information obtained from a subject).
[0108] Mutation count: As used herein, "mutation count" or "mutational count" refers to the number of somatic mutations in the whole genome or exome or targeted region of a nucleic acid sample.
[0109] Negative control region: As used herein, "negative control region" refers to a genomic region that contains nucleic acids that have below a threshold number of cytosines that are methylated in cells derived from subjects without cancer and also in cells derived from subjects without cancer.
[0110] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to an abnormal growth of cells in a subject. A neoplasm or tumor can be benign, presumably malignant, or malignant. A malignant tumor is called a cancer or cancerous tumor.
[0111] Next-generation sequencing: As used herein, "next-generation sequencing" or "NGS" refers to a sequencing technology that has increased throughput compared to traditional Sanger and capillary electrophoresis-based approaches, e.g., the ability to generate hundreds of thousands of relatively small sequencing reads at a time. Some examples of next-generation sequencing technologies include, but are not limited to, sequencing-by-synthesis, sequencing-by-ligation, and sequencing-by-hybridization.
[0112] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, less than about 50 nucleotides, or less than about 10 nucleotides in length) used to identify nucleic acids from different samples (e.g., representing a sample index) or to identify different nucleic acid molecules in the same sample that have been differently typed or processed (e.g., representing a molecular barcode). Nucleic acid tags comprise predetermined, fixed, non-random, random, or semi-random oligonucleotide sequences. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or various lengths, as desired. Nucleic acid tags can also include double-stranded molecules with one or more blunt ends, 5' or 3' single-stranded regions (e.g., overhangs), and / or one or more other single-stranded regions elsewhere within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the sample of origin, form, or processing of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indices, where the nucleic acids are subsequently deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between different molecules or amplicons of different parent molecules in the same sample or subsample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample or non-uniquely tagging such molecules.For non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule such that different molecules can be distinguished based on their intrinsic sequence information (e.g., start and / or end positions where they map to a selected reference sequence, subsequences at one or both ends of the sequence, and / or sequence length) in combination with at least one molecular barcode. A sufficient number of different molecular barcodes are used so that the probability that any two molecules will have the same intrinsic sequence information (e.g., start and / or end positions, subsequences at one or both ends of the sequence, and / or length), as well as the same molecular barcode, is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% likelihood).
[0113] Partitioning: As used herein, "partitioning" refers to physically separating or fractionating a mixture of nucleic acid molecules in a sample based on characteristics of the nucleic acid molecules. Partitioning can be a physical partitioning of molecules. Partitioning can include separating nucleic acid molecules into groups or sets based on the level of epigenetic traits (e.g., related to methylation). For example, nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In some embodiments, methods and systems used for partitioning can be found in PCT Patent Application No. PCT / US2017 / 068329, which is incorporated herein by reference in its entirety.
[0114] Distribution set: As used herein, "distribution set" or "distribution" refers to a set of nucleic acid molecules distributed into sets or groups based on the differential binding affinity of the nucleic acid molecules or proteins associated with the nucleic acid molecules to a binder. A distribution set may also be referred to as a subsample. A binder preferentially binds to nucleic acid molecules containing nucleotides with epigenetic modifications. For example, if the epigenetic modification is methylation, the binder may be a methyl-binding domain (MBD) protein. In some embodiments, a distribution set may include nucleic acid molecules belonging to a particular level or degree of epigenetic trait (e.g., methylation). For example, nucleic acid molecules may be distributed into three sets: one set of highly methylated nucleic acid molecules (first subsample, high distribution, high distribution set, or high methylation distribution set), a second set of hypomethylated nucleic acid molecules (second subsample, low distribution, low distribution set, or hypomethylation distribution set), and a third set of intermediately methylated nucleic acid molecules (third subsample, intermediate distribution set, intermediate methylation distribution set, residual distribution, or residual distribution set). In another example, the nucleic acid molecules may be distributed based on the number of methylated nucleotides, with one distribution set having nucleic acid molecules with 9 methylated nucleotides and another distribution set having unmethylated nucleic acid molecules (0 methylated nucleotides).
[0115] Polynucleotide: As used herein, "polynucleotide," "nucleic acid," "nucleic acid molecule," "polynucleotide molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleoside linkages. A polynucleotide can contain at least three nucleosides. Oligonucleotides often range in size from a small number of monomeric units, e.g., 3-4, to hundreds of monomeric units. Whenever a polynucleotide is represented by a sequence of letters, e.g., "ATGCCTG," it is understood that the nucleotides are in 5'→3' order from left to right, and that, in the case of DNA, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine, unless otherwise specified. The letters A, C, G, and T may be used to refer to the base itself, a nucleoside, or a nucleotide that includes the base, as is standard in the art.
[0116] Positive control region: As used herein, a "positive control region" refers to a genomic region that is methylated in cells and contains at least a threshold number of nucleic acids with cytosines from both subjects without a homologous recombination repair deficiency and subjects in which a recombination repair deficiency is present.
[0117] Probe: As used herein, "probe" refers to a polynucleotide that includes a function. The function may be a detectable label (fluorescence), a binding moiety (biotin), or a solid support (magnetically attractable particle or chip). A probe may include a single-stranded DNA / RNA polynucleotide or a double-stranded DNA polynucleotide (e.g., SureSelect® Probe, Agilent Technologies) that hybridizes to a target nucleic acid sequence. Sequence capture using a probe generally depends in part on the number of consecutive nucleotides in at least a portion of the target nucleic acid sequence that are complementary (or nearly complementary) to the sequence of the probe. In some examples, the probe may correspond to a driver mutation.
[0118] Processing: As used herein, "processing," "calculating" The terms "comparing" and "comparing" may be used interchangeably. In application, the term refers to determining differences, e.g., differences in number or sequence. For example, gene expression, copy number variation (CNV), indel, and / or single nucleotide variant (SNV) values or sequences can be processed.
[0119] Processor: As used herein, "processor" refers to any circuit or virtual circuit (a physical circuit emulated by logic running on an actual processor) that manipulates data values in accordance with control signals (e.g., "commands," "op codes," "machine code," etc.) to produce corresponding output signals that are applied to operate a machine. A processor may be, for example, a CPU, a RISC processor, a CISC processor, a GPU, a DSP, an ASIC, an RFIC, or any combination thereof. A processor may also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") capable of simultaneously executing instructions.
[0120] Promoter region: As used herein, "promoter region" refers to a DNA sequence recognized by the synthetic machinery of the cell, or introduced synthetic machinery, necessary to initiate the specific transcription of a gene.
[0121] As used herein, "quantitative measurement" refers to an absolute or relative measurement. A quantitative measurement can be, but is not limited to, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation, or quantile), or a degree or relative quantity (e.g., high, medium, and low). A quantitative measurement can be a ratio of two quantitative measurements. A quantitative measurement can be a linear combination of quantitative measurements. A quantitative measurement can be a normalized measurement.
[0122] As used herein, " reference sequence " refers to the known sequence used for the purpose of comparison with experimentally determined sequences.For example, known sequence can be the whole genome, chromosome, or any segment thereof.Reference sequence can comprise at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000 or more nucleotides.Reference sequence can be aligned with a single continuous sequence of genome or chromosome, or can comprise discontinuous segments that are aligned with different regions of genome or chromosome.Examples of reference sequence include, for example, human genome reference sequence, such as hG19 and hG38.
[0123] As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.
[0124] Sensitivity: As used herein, "sensitivity" means the probability of detecting the presence of single-base variants, insertions, and deletions at a given MAF and coverage, and the probability of detecting the presence of copy number variants at a given tumor fraction and coverage.
[0125] As used herein, "sequencing" refers to any of several technologies used to determine the sequence (e.g., identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, and the like. Examples of sequencing techniques include high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperatures (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some implementations, sequencing can be performed by a genetic analyzer, such as a commercially available genetic analyzer from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among others.
[0126] Single nucleotide variant: As used herein, "single nucleotide variant" or "SNV" refers to a mutation or variation in a single base that occurs at a specific position in the genome.
[0127] Somatic mutation: As used herein, the terms "somatic mutation" or "somatic mutation" are used interchangeably. They refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.
[0128] Specific binding: As used herein, "specific binding," in the context of a probe or other oligonucleotide and a target sequence, means that, under appropriate hybridization conditions, the oligonucleotide or probe hybridizes to its target sequence or a copy thereof to form a stable probe:target hybrid, while at the same time minimizing the formation of stable probe:non-target hybrids. Thus, the probe hybridizes to the target sequence or a copy thereof to a sufficiently greater extent than to non-target sequences, allowing for capture or detection of the target sequence. Suitable hybridization conditions are well known in the art and can be predicted based on sequence composition or determined using routine testing methods (see, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2nd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) §§ 1.90-1.91, 7.37-7.57, 9.47-9.51 and 11.47-11.57, especially §§ 9.50-9.51, 11.12-11.13, 11.45-11.47 and 11.55-11.57, which are incorporated herein by reference).
[0129] As used herein, "subject" refers to an animal, such as a mammalian species (e.g., a human), or an avian (e.g., a bird) species, or other organism, such as a plant. More specifically, a subject can be a vertebrate, such as a mammal, such as a mouse, a primate, a monkey, or a human. Animals include farm animals (e.g., beef cattle, dairy cattle, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or service animals). A subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended interchangeably with "subject."
[0130] For example, the subject can be the individual who has been diagnosed with cancer, who is going to receive cancer treatment, and / or who has received at least one cancer treatment.The subject can be in the remission stage of cancer.As another example, the subject can be the individual who has been diagnosed with autoimmune disease.As another example, the subject can be the female individual who is pregnant or planning to become pregnant, who has been diagnosed with or suspected to have disease, such as cancer, autoimmune disease.
[0131] Target region: As used herein, "target region" refers to a genomic region of interest. For example, the genomic region of interest may correspond to one or more mutations consistent with one or more cancer types. Furthermore, the genomic region of interest may be enriched by one or more probes.
[0132] Threshold: As used herein, "threshold" refers to a predetermined value used to characterize experimentally determined values for the same parameter in different samples according to their relationship to the threshold.
[0133] Tumor fraction: As used herein, " tumor fraction " refers to the estimated fraction of nucleic acid molecules derived from tumor in a given sample.For example, the tumor fraction of sample can be the maximum MAF of sample or the pattern of sequencing coverage of sample or the length of the cfDNA fragments in sample or any other selected feature of sample.In some cases, the tumor fraction of sample is equal to the maximum MAF of sample.
[0134] Variant: As used herein, "variant" can be referred to as an allele. Variants are usually present at a frequency of 50% (0.5) or 100% (1), depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and usually have a frequency of 0.5 or 1. However, somatic variants are acquired variants and usually have a frequency of <0.5. The major and minor alleles of a genetic locus refer to nucleic acids having a locus that is occupied by a nucleotide of a reference sequence, and each variant nucleotide is different from the reference sequence. Measurements at a locus can be obtained in the form of an allele fraction (AF), which measures the frequency at which an allele is observed in a sample. (Mode for Carrying Out the Invention)
[0135] Detailed Description Cancer is usually caused by the accumulation of mutations in the genes of an individual's cells, at least in part resulting in improperly regulated cell division. Such mutations can include single nucleotide variations (SNVs), gene fusions, insertions, deletions, transversions, translocations, and inversions. These mutations can also include copy number variations, which correspond to an increase or decrease in the copy number of genes in the tumor genome compared to the individual's non-cancerous cells. The degree of mutation present in cell-free nucleic acid and the amount of mutated cell-free nucleic acid in a sample can be used as biomarkers to determine tumor progression, predict patient outcome, and improve treatment selection. In various examples, the degree of mutation present in cell-free nucleic acid can be indicated by tumor cell copy number and tumor fraction for a given sample.
[0136] Additionally, cancer can be manifested by non-sequence alterations such as methylation. Examples of methylation changes in cancer include localized increases in DNA methylation in CpG islands at the TSSs of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This increased methylation level can be associated with abnormal loss of transcriptional capacity of the associated gene and occurs at least as frequently as point mutations and deletions as a cause of altered gene expression.
[0137] Therefore, DNA methylation profiling can be used to detect the abnormal methylation in the DNA of sample.The DNA is usually hypermethylated or hypomethylated in a given sample type (for example, the cfDNA derived from bloodstream), but for example, because the tissue contribution to this sample type is abnormally increased (for example, due to the increased DNA loss in or around neoplasia or cancer), and / or the degree of genomic methylation is changed during development or is disturbed by disease, for example, cancer or any cancer-related disease, the degree of abnormal methylation can correspond to a certain genomic region (" differentially methylated region " or " DMR "), which can show the degree of abnormal methylation that correlates with neoplasia or cancer.
[0138] Defects in homologous recombination repair pathway can be determined by identifying mutations in the genes involved in the regulation of homologous recombination repair pathway.For example, somatic mutations in genes previously identified as being related to the regulation of homologous recombination repair pathway can be identified by analyzing the mutations present in the cell-free nucleic acid obtained from the subject.The accuracy of existing techniques for detecting somatic mutations in the cell-free nucleic acid of subjects is typically less than desired.In at least some scenarios, homologous recombination repair deficiency can be characterized by the presence of deletions in the genes that regulate homologous recombination repair pathway.The sensitivity of existing techniques for detecting the deletions corresponding to the genes that regulate homologous recombination repair pathway is somewhat limited.
[0139] The method and system described herein are intended to determine whether a subject has a homologous recombination repair deficiency by analyzing methylation data obtained from a sample containing the subject's cell-free nucleic acid. The methylation data can indicate the amount of nucleic acid molecules with methylated cytosine in some genomic regions. For example, the methylation data can correspond to the differentially methylated genomic regions in subjects with cancer. Furthermore, the methylation data can correspond to the genomic regions previously identified as having one or more mutations present in individuals with cancer. By analyzing methylation data to determine whether a subject has a homologous recombination repair deficiency, the accuracy of detecting homologous recombination repair deficiency is improved compared to the accuracy achieved using existing technology.
[0140] In one or more implementations, one or more computational models can be generated to determine the subject's status with respect to homology-directed repair deficiency. The one or more computational models can implement at least one of one or more machine learning techniques or one or more statistical techniques to determine the subject's status with respect to homology-directed repair deficiency. In various examples, the one or more computational models can analyze sequencing data corresponding to a sample obtained from the subject to determine the subject's status with respect to homology-directed repair deficiency. The sequencing data can indicate nucleic acid molecules having at least one of more than expected methylated cytosines or fewer than expected methylated cytosines within some genomic regions. Some genomic regions can indicate one or more mutations in individuals with one or more forms of cancer and / or individuals with homology-directed repair deficiency. Some genomic regions can also include regions that are differentially methylated in individuals with homology-directed repair deficiency.
[0141] 1 is a schematic diagram of an example environment 100 for partitioning nucleic acids based on the methylation of cytosine molecules in several genomic regions of a reference sequence, according to one or more implementations. In one or more examples, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include bile duct cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphocytic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary tumor, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.
[0142] The environment 100 may include a sample 102. The sample 102 may be derived from a fluid obtained from a subject. For example, the sample 102 may be derived from blood obtained from a subject. In one or more additional examples, the sample 102 may be derived from a tissue of the subject. In various examples, the sample 102 may be derived from multiple sources. For example, the sample 102 may be derived from one or more bodily fluids of the subject and a tissue of the subject. In one or more instances, the subject may be a mammal. In one or more additional instances, the subject may be a human. In one or more further instances, the subject may be a non-human mammal.
[0143] The sample 102 may include several nucleic acids 104. Each nucleic acid 104 may include several regions having at least a threshold number of cytosine and guanine molecules. In one or more examples, each nucleic acid 104 may include a region having at least a threshold number of cytosine-guanine pairs. In various examples, at least a portion of the cytosine-guanine pairs contained within a region may be sequentially positioned within the sequence of the nucleic acid 104. In one or more instances, a region of a nucleic acid having at least a threshold amount of cytosine-guanine pairs may be referred to herein as a "CG region" or "CpG region." In one or more instances, a CG region may include at least 200 base pairs. In one or more instances, a CG region may include 200 base pairs to 5000 base pairs, 300 base pairs to 3000 base pairs, 200 base pairs to 2500 base pairs, or 500 base pairs to 1500 base pairs. Furthermore, the CG region may have a GC percentage of at least 50% and a ratio of observed CpGs to predicted CpGs of at least 60%. The ratio of observed CpGs to predicted CpGs can be calculated as follows: observed CpGs is the number of CpGs identified in a given genomic region, and predicted CpGs is the number of cytosines multiplied by the number of guanines divided by the number of bases in the genomic region. The predicted CpGs are: ((number of cytosines + number of guanines) / 2) / length of genomic region For example, CG regions can be determined using the techniques described by Gardiner-Garden M, Frommer M (1987). "CpG islands in vertebrate genomes". Journal of Molecular Biology. 196 (2): 261-282. and / or Saxonov S, Berg P, Brutlag DL (2006). "A genome-wide analysis of CpG dinucleotides in the human genome distinguishes two distinct classes of promoters". Proc Natl Acad Sci USA. 103 (5): 1412-1417.
[0144] 1, a portion of the sequence of an example nucleic acid 104 can include a first CG region 106, a second CG region 108, and a third CG region 110. Although the example of FIG. 1 illustrates a portion of the sequence of a nucleic acid 104 having three CG regions, the nucleic acids 104 included in the sample 102 can have a different number of CG regions. For example, an individual nucleic acid 104 included in the sample 102 can include at least 1 CG region, at least 5 CG regions, at least 10 CG regions, at least 25 CG regions, at least 50 CG regions, at least 100 CG regions, at least 250 CG regions, at least 500 CG regions, or at least 1000 CG regions.
[0145] An individual CG region may include some molecules with methylated cytosines. In the example of FIG. 1, the first CG region 106 may include molecules with methylated cytosines 112. In the example of FIG. 1, the molecules with methylated cytosines 112 are 5-methylcytosines. An individual CG region may also include some unmethylated cytosines. For example, the CG region 106 may include molecules with unmethylated cytosines 116. In various examples, at least a portion of the CG regions of the nucleic acid 104 may correspond to classification regions of a reference genome. In the example of FIG. 1, the first CG region 106, the second CG region 108, and the third classification region 110 may be classification regions. The classification regions may correspond to genomic regions of the reference genome corresponding to one or more mutations consistent with one or more biological conditions, such as one or more forms of cancer. The one or more mutations may include at least one of somatic mutations or germline mutations.
[0146] In at least some examples, the classification region can correspond to the genomic region that is enriched as part of the diagnostic assay.In various examples, the classification region can be the differentially methylated region that is located in the genomic region that corresponds to the defect in the homologous recombination repair pathway.For example, gene 118 can correspond to the genomic region that is related to the homologous recombination repair deficiency.In one or more additional examples, the classification region can be determined by identifying the differentially methylated region that is located in at least a part of the gene that corresponds to the homologous recombination repair deficiency, wherein the differentially methylated region is present in the subject that has the homologous recombination repair deficiency, and the differentially methylated region is absent in the subject that does not have the homologous recombination repair deficiency.
[0147] In addition to several CG regions, each individual nucleic acid 104 may correspond to one or more positive control regions of the reference genome, such as positive control region 120. Positive control region 120 may contain at least a threshold number of methylated cytosines in a genomic region of nucleic acid derived from a cell obtained from a subject without cancer, and is also methylated in a genomic region of nucleic acid derived from a subject with cancer. In various examples, positive control region 120 may be hypermethylated in nucleic acid derived from a subject without cancer and in nucleic acid derived from a subject with cancer. Each individual nucleic acid 104 may also include one or more negative control regions, such as negative control region 122. Negative control region 122 may contain less than a threshold number of methylated cytosines in nucleic acid derived from a subject without cancer and in nucleic acid derived from a subject without cancer. In one or more instances, negative control region 122 may be hypomethylated in both a subject without cancer and a subject with cancer. In various examples, the positive control region and negative control region can be used to perform normalization calculations. Normalization calculations can be performed to generate input data for one or more computational models implemented to determine homology directed repair deficiency for a given sample 102.
[0148] A molecular separation process 124 can be performed. The molecular separation process 124 can separate the nucleic acids 104 contained in the sample 102 based on the amount of cytosine methylation in each individual nucleic acid 104. In one or more examples, the molecular separation process 124 can separate the nucleic acids 104 contained in the sample 102 based on the amount of cytosine methylation in the CG regions of each individual nucleic acid 104. In various examples, the molecular separation process 124 can separate the nucleic acids 104 into multiple groups, each group corresponding to a different amount of cytosine methylation in the nucleic acids 104.
[0149] 1 , the molecular separation process 124 can be performed in relation to a first methylation threshold 126. Performing the molecular separation process 124 in relation to the first methylation threshold 126 can result in a first distribution 128 of nucleic acids. In one or more examples, the first methylation threshold 126 can refer to a first threshold number of methylated cytosines located within CG regions of the nucleic acids 104. The molecular separation process 124 can identify some nucleic acids 104 that have fewer methylated cytosines within the CG regions than the first methylation threshold 126. In various examples, the first methylation threshold 126 can correspond to a first methylation rate.
[0150] The molecular separation process 124 can also be performed with respect to a second methylation threshold 130. The second methylation threshold 130 can indicate an amount of cytosine methylation in one or more genomic regions of the nucleic acids 104 that is greater than the amount of cytosine methylation in one or more regions corresponding to the first methylation threshold 126. The second methylation threshold 130 can indicate a number of methylated cytosines per given number of nucleic acids. In one or more additional examples, the second methylation threshold 130 can correspond to a rate of methylation of the nucleic acids that is greater than the rate of methylation corresponding to the first methylation threshold 126. Performing the molecular separation process 124 with respect to the second methylation threshold 130 can result in a second distribution 132 of the nucleic acids. In one or more examples, the molecular separation process 124 can identify nucleic acids 104 having an amount of cytosine methylation greater than a first methylation threshold 126 and an amount of cytosine methylation less than a second methylation threshold 130, resulting in a second distribution 132 of nucleic acids.
[0151] Additionally, the molecular separation process 124 can be performed with respect to a third methylation threshold 134. The third methylation threshold 134 can indicate an amount of cytosine methylation in one or more genomic regions of the nucleic acids 104 that is greater than the amount of cytosine methylation in one or more regions corresponding to the first methylation threshold 126 and greater than the amount of cytosine methylation in one or more regions corresponding to the second methylation threshold 130. The third methylation threshold 134 can indicate the number of molecules having methylated cytosines per given number of nucleic acids. In one or more additional examples, the third methylation threshold 134 can correspond to a rate of cytosine methylation that is greater than the rate of methylation corresponding to the first methylation threshold 126 and greater than the rate of methylation corresponding to the second methylation threshold 130. Performing the molecular separation process 124 with respect to the third methylation threshold 134 can result in a third distribution 136 of nucleic acids. In one or more examples, the molecular separation process 124 can identify nucleic acids 104 that have a higher amount of cytosine methylation than the nucleic acids 104 contained in the second distribution of nucleic acids 132. As such, the amount of cytosine methylation of the nucleic acids contained in the first distribution 128, the second distribution 132, and the third distribution 136 increases from the first distribution 128 to the second distribution 132 and increases from the second distribution 132 to the third distribution 136. In one or more instances, the first distribution of nucleic acids 128 can be referred to as a low-methylation distribution, the second distribution of nucleic acids 132 can be referred to as an intermediate distribution, and the third distribution of nucleic acids 136 can be referred to as a high-methylation distribution.
[0152] In one or more examples, the amount of cytosine methylation of a nucleic acid can correspond to the strength of binding to a methyl-binding domain (MBD). In these scenarios, the first partitioning 128, second partitioning 132, and third partitioning 134 can occur based on the different strengths of binding to the MBD for nucleotides with different amounts of cytosine methylation. In one or more examples, the molecular separation process 124 can include a series of washes in which the nucleic acid 104 is contacted with solutions having different concentrations of sodium chloride (NaCl).
[0153] Partitioning of nucleic acids can be performed by contacting the nucleic acid with a modified nucleotide-specific binding reagent, such as the MBD of MBP. The modified nucleotide-specific binding reagent can bind to 5-methylcytosine (5mC). The modified nucleotide-specific binding reagent, such as the MBD, can be coupled to paramagnetic beads such as Dynabeads® M-280 streptavidin via a biotin linker. Partitioning into fractions with different degrees of methylation can be performed by increasing the NaCl concentration in a series of washes. Sequences eluted from the modified nucleotide-specific binding reagent are partitioned into two or more fractions (e.g., low, high) depending on which wash (e.g., NaCl concentration) the sequence eluted from. The resulting partition can contain one or more of the following nucleic acid forms: double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments.
[0154] Binding of nucleic acids using modified nucleotide-specific binding reagents can be a function of the number of methylated (or modified) sites per molecule, with molecules with more methylation eluted by increasing salt concentration. A series of elution buffers with increasing NaCl concentrations can be used to elute DNA into distinct populations based on the degree of methylation. In one or more implementations, the salt concentration can range from about 100 mM to about 2500 mM NaCl. In various implementations, the molecular separation process 124 results in three partitions. Molecules are contacted with a solution containing molecules containing a methyl-binding domain at a first salt concentration, allowing the molecules to bind to a capture moiety, such as streptavidin. At the first salt concentration, one population of molecules binds to the MBD, and one population remains unbound. The unbound population can be separated as a "hypomethylated" population (low partition). For example, the first partition 128 can represent hypomethylated forms of DNA that remain unbound at low salt concentrations. In one or more instances, the concentration of NaCl in the solution used to effect the first distribution 128 can be about 100 nM, about 120 nM, about 140 nM, about 160 nM, about 180 nM, about 200 nM, or about 250 nM.
[0155] The second partition 132 may be referred to as an "intermediate partition" and may represent intermediate methylation of CG regions of DNA eluted using an intermediate salt concentration, e.g., a concentration between 100 mM and 2000 mM. In one or more additional examples, the NaCl concentration of the solution used to generate the second partition 132 may be about 100 mM to about 500 mM, about 100 mM to about 1000 mM, about 100 mM to about 1500 mM, about 250 mM to about 1000 mM, about 250 mM to about 1500 mM, about 500 mM to about 1500 mM, about 250 mM to about 2000 mM, about 500 mM to about 2000 mM, or about 1000 mM to about 2000 mM. The third partition 136 may represent highly methylated forms of nucleic acids (high partition) and is eluted using a high salt concentration, e.g., at least about 2000 mM. In one or more further examples, the NaCl concentration of the solution used to generate the third partition 136 may be about 2000 mM to about 5000 mM, about 2000 mM to about 4000 mM, about 2000 mM to about 3500 mM, about 2000 mM to about 3000 mM, or about 2500 mM to about 4000 mM.
[0156] In various examples, the first distribution 128 may correspond to a first range of binding strengths of nucleic acids to MBDs and a first amount of methylated cytosines in CG regions, and the second distribution 132 may correspond to a second range of binding strengths of nucleic acids to MBDs and a second amount of methylated cytosines in CG regions. The first range of binding strengths may be less than the second range of binding strengths. In one or more scenarios, a first solution having a first NaCl concentration can separate a first group of nucleic acids having the first range of binding strengths from the MBD, and a second solution having a second NaCl concentration higher than the first NaCl concentration can separate a second group of nucleic acids having the second range of binding strengths from the MBD. Furthermore, the third distribution 136 may correspond to a third range of binding strengths and a third amount of methylated cytosines in CG regions. The third range of binding strengths may be greater than the first range of binding strengths and greater than the second range of binding strengths. In one or more cases, a third group of nucleic acids having a third range of binding strengths can be separated by a third solution having a third NaCl concentration, which can be greater than the first NaCl concentration and greater than the second NaCl concentration.
[0157] In one or more instances, a plurality of nucleic acids from at least one of a subject's blood or tissue can be combined with a solution containing a quantity of MBD to produce a nucleic acid-MBD solution. A first wash of the nucleic acid-MBD solution can be performed using a first wash solution containing a first NaCl concentration to produce a first nucleic acid fraction and a first residual solution. The first nucleic acid fraction can include a first portion of the plurality of nucleic acids, and the first residual solution can include a second portion of the plurality of nucleic acids. In one or more instances, the first portion of the plurality of nucleic acids can have a first range of binding energies to the MBD that is less than a second range of binding energies to the MBD of the second portion of the plurality of nucleic acids.
[0158] Further, a second wash of the first residual solution can be performed using a second wash solution containing a second NaCl concentration higher than the first NaCl concentration to produce a second nucleic acid fraction and a second residual solution. The second nucleic acid fraction can include a first subset of a second portion of the plurality of nucleic acids, and the second residual solution can include a second subset of a second portion of the plurality of nucleic acids. The first subset of a second portion of the plurality of nucleic acids can have a third range of binding energies to the MBD that is less than the fourth range of binding energies to the MBD of the second subset of a second portion of the plurality of nucleic acids. In various examples, the second range of binding energies can include a third range of binding energies and a fourth range of binding energies. Further, a third wash of the second residual solution can be performed using a third solution containing a third NaCl concentration higher than the second NaCl concentration to produce a third nucleic acid fraction containing a second subset of a second portion of the plurality of nucleic acids.
[0159] After the first wash, the second wash, and the third wash, a determination can be made that the first nucleic acid fraction is associated with the first distribution 128. A first molecular barcode can then be attached to a first portion of the plurality of nucleic acids using a first molecular barcode indicative of the first distribution 128. In this manner, a sequencing read corresponding to the first distribution 128 can be identified based on determining that the sequencing read comprises the first molecular barcode. Further, a determination can be made that the second nucleic acid fraction is associated with a second distribution 132 of the plurality of distributions. In these circumstances, a second molecular barcode can be attached to a first subset of the second portion of the plurality of nucleic acids using a second molecular barcode indicative of the second distribution 132. As a result, a sequencing read corresponding to the second distribution 132 can be identified based on determining that the sequencing read comprises the second molecular barcode. Further, a determination can be made that the third nucleic acid fraction is associated with the third distribution 136. A third molecular barcode can then be attached to a second subset of the second portion of the plurality of nucleic acids using a third molecular barcode that indicates a third distribution 136. In these cases, sequencing reads that correspond to the third distribution 136 can be identified based on determining that the sequencing reads include the third molecular barcode.
[0160] In one or more additional examples, the molecular separation process 124 may include performing one or more bisulfite sequencing processes to determine the amount of methylation of the nucleic acids 102. In one or more illustrative examples, the bisulfite sequencing may be performed according to Li Y, Tollefsbol TO. DNA methylation detection: bisulfite genomic sequencing analysis. Methods Mol Biol. 2011;791:11-21. doi: 10.1007 / 978-1-61779-316-5_2. PMID: 21913068; PMCID: PMC3233226. In one or more further examples, at least one of the first distribution of nucleic acids 128, the second distribution of nucleic acids 132, or the third distribution of nucleic acids 136 may be subjected to an additional separation process. For example, the one or more molecular separation processes 124 may include digestion with a methyl-sensitive restriction enzyme (MSRE) of at least one of the first partition of nucleic acids 128, the second partition of nucleic acids 132, or the third partition of nucleic acids 136. Digestion of the nucleic acids contained in the first partition 128, the second partition 132, and / or the third partition 136 with an MSRE can result in separation of nucleic acids contained in one of the first partition 126, the second partition 132, or the third partition 136 that do not have a level of methylation corresponding to the first methylation threshold 126, the second methylation threshold 130, or the third methylation threshold 134, respectively. Digestion of nucleic acids contained in at least one of the first distribution 128, the second distribution 132, or the third distribution 136 using an MSRE can increase the amount of nucleic acids contained in the first distribution 128 having an amount of methylation below the first methylation threshold 126, increase the amount of nucleic acids contained in the second distribution 132 having an amount of methylation between the second methylation threshold 130 and the first methylation threshold 126, and / or increase the amount of nucleic acids contained in the third distribution 136 having an amount of methylation between the third methylation threshold 134 and the second methylation threshold 132.
[0161] The environment 100 may include a sequencing machine 138. In one or more examples, the sequencing machine 138 may be any of several sequencing machines capable of performing one or more sequencing operations that amplify nucleic acids present in the sample 104. In various examples, the sequencing machine 138 may perform next-generation sequencing operations.
[0162] In the example of FIG. 1 , the separation of nucleic acids 102 into first distribution 128 and second distribution 132 can be optional, as indicated by the dotted line between molecular separation process 124 and first distribution 128 and second distribution 132. That is, in at least some instances, only third methylation threshold 134 is used to generate third distribution 136 corresponding to highly methylated nucleic acids, which is then provided to sequencing machine 138. Any remaining nucleic acids not included in third distribution 136 are not provided to sequencing machine 138. Furthermore, in these scenarios, third distribution 136 is subjected to MSRE digestion, and any remaining nucleic acids not included in the third distribution are not subjected to MSRE digestion. Furthermore, in these scenarios, the nucleic acids included in third distribution 136 are subjected to blunt-end ligation.
[0163] In one or more additional examples, the separation of nucleic acids into the second partition 132 is optional, and the molecular separation process 124 using the first methylation threshold 126 and the third methylation threshold 130 results in a first partition 128 corresponding to low methylated nucleic acids and a third partition 136 corresponding to high methylated nucleic acids. In these scenarios, the nucleic acids contained in the first partition 128 and the third partition 132 are provided to the sequencing machine 136, and the nucleic acids corresponding to the second partition 132 are not provided to the sequencing machine 138. Further, in various examples, the molecular separation process 124 can result in two partitions: a partition combining the nucleic acids from the first partition 128 and the second partition 132, and an additional partition containing the nucleic acids of the third partition 136.
[0164] Prior to sequencing, the extracted polynucleotides and adapters can be blunt-end ligated, and tags (e.g., molecular barcodes) can be added to the extracted polynucleotides. The extracted polynucleotides can also be enriched by hybridizing them with probes corresponding to classification regions of a reference sequence. The enrichment process can identify thousands, hundreds of thousands, or even millions of polynucleotides corresponding to the classification regions associated with the probes. In one or more examples, the enrichment process can be performed on a genomic region that is part of a diagnostic assay. In these cases, the genomic region can correspond to a nucleotide sequence of a reference genome that indicates the presence of one or more forms of cancer. In one or more additional examples, the enrichment process can be performed on a genomic region where mutations can result in defects in the homologous recombination repair mechanism. In these scenarios, the genomic region can contain genes corresponding to the homologous recombination repair pathway. After the enrichment process, there can also be thousands, or even millions, of unenriched polynucleotides corresponding to non-classified regions of the reference sequence.
[0165] After the enrichment process, the enriched polynucleotides can be amplified according to one or more amplification processes. The one or more amplification processes can produce thousands, or even millions, of copies of each enriched polynucleotide. In one or more instances, a portion of the non-enriched polynucleotides can be amplified, but in some cases, not to the same extent as the enriched polynucleotides. The one or more amplification processes can produce amplification products that are subjected to one or more sequencing operations. After performing one or more sequencing operations on the sample 104, sequencing data 140 can be produced by the sequencing machine 138.
[0166] Sequencing data 140 may include alphanumeric representations of the nucleic acids contained in the amplification products produced by sequencing machine 140. For example, sequencing data 140 may include, for each nucleic acid in the amplification product, data corresponding to a character string representing each strand of nucleotides corresponding to the individual nucleic acid.
[0167] The sequencing data 140 may be stored in one or more data files. For example, the sequencing data 140 may be stored in a FASTQ file, which includes a text-based sequencing data file format that stores raw sequence data and quality scores. In one or more additional examples, the sequencing data 142 may be stored in a data file according to the binary base call (BCL) sequence file format. In one or more further examples, the sequencing data 142 may be stored in a BAM file. In one or more examples, the sequencing data 140 may include at least about 1 gigabyte (GB), at least about 2 GB, at least about 3 GB, at least about 4 GB, at least about 5 GB, at least about 8 GB, or at least about 10 GB. An individual sequence representation included in the sequencing data 140 may be referred to herein as a "read" or "sequencing read." In various examples, an individual first nucleic acid included in the sample 102 may correspond to many sequence representations included in the sequencing data 140 as a result of the individual first nucleic acid being amplified. In one or more additional examples, an individual second nucleic acid contained in sample 102 may correspond to a single sequence representation or a small number of sequence representations contained in sequencing data 140 as a result of the individual second nucleic acid not being amplified.
[0168] The environment 100 may also include performing a computational analysis 142 based on the sequencing data 140. The computational analysis 142 may include analyzing sequence reads corresponding to one or more of the distributions 128, 132, 136 to determine an indicator of homologous recombination deficiency (HRD) status 144. The indicator of HRD status 144 may correspond to a probability that a homologous recombination repair deficiency is present in the subject. In various examples, the computational analysis 142 may implement at least one of one or more machine learning techniques or one or more statistical techniques to generate the indicator of HRD status 144.
[0169] In one or more examples, the computational analysis 142 may include determining a quantity of sequencing reads having an amount of methylation corresponding to at least one of the first methylation threshold 126, the second methylation threshold 130, or the third methylation threshold 134 and corresponding to at least one of the first CG region 106, the second CG region 108, or the third CG region 110. In one or more instances, the computational analysis 142 may include determining a number of sequence reads included in the sequencing data 140 that correspond to nucleic acids included in the third distribution 136. In various examples, the computational analysis 142 may implement one or more computational models that include components corresponding to a portion of the classification regions. For example, a training process may be performed to identify a subset of the classification regions that are predictive of HRD status. The subset of classification regions may correspond to one or more components of the one or more computational models implemented as part of the computational analysis 142 to generate the indicator 144 of HRD status.
[0170] 2 is a schematic diagram of an example framework 200 for determining homologous recombination deficiency status based on the methylation status of cell-free nucleic acid molecules using one or more computational models, according to one or more implementations. The framework 200 can include performing one or more library preparation and sequencing processes 202. The one or more library preparation and sequencing processes 202 can be performed on several samples 204. The several samples 204 can be obtained from subjects 206. In one or more instances, a first portion of the subjects 206 can be subjects without a homology-directed repair deficiency. Additionally, a homology-directed repair deficiency can be present in a second portion of the subjects 206. The subjects 206 can be training subjects from which data is generated for training one or more computational models for determining the subject's status with respect to homologous recombination deficiency.
[0171] In one or more examples, a first portion of the subjects 206 can be determined based on the absence of one or more mutations in a genomic region associated with homologous recombination repair. In one or more additional examples, a second portion of the subjects 206 can be determined based on the presence of one or more mutations in a genomic region associated with homologous recombination repair. In one or more instances, the one or more mutations in a genomic region associated with homologous recombination repair can include at least one of a germline deletion, a germline rearrangement, a germline fusion, a somatic deletion, a somatic rearrangement, a somatic fusion, or a homozygous deletion. In various examples, the deletion present in the second portion of the subjects 206 having a homologous recombination deficiency can include a single-base variant or an indel. In one or more further examples, a second portion of the subjects 206 having a homologous recombination deficiency can be determined based on responsiveness to an inhibitor of polyadenosine diphosphate (ADP) ribose polymerase (PARP). Individuals with a homologous recombination deficiency tend to be responsive to treatment with a PARP inhibitor. (See Keung MYT, Wu Y, Vadgama JV. PARP Inhibitors as a Therapeutic Agent for Homologous Recombination Deficiency in Breast Cancers. J Clin Med. 2019 Mar 30;8(4):435. doi: 10.3390 / jcm8040435. PMID:30934991; PMCID: PMC6517993.) As a result, the second portion of subjects 206 can include a portion of subjects 206 that have a tumor and have a reduction in tumor cells in response to treatment of the tumor with one or more PARP inhibitors.
[0172] The library preparation and sequencing process 202 may include extraction of nucleic acid molecules from a sample 204. In one or more implementations, the nucleic acid molecules include cell-free nucleic acids (e.g., cell-free DNA). In various implementations, the sample 204 may include one or more samples selected from one or more of blood, plasma, serum, urine, feces, saliva samples, combinations thereof, and the like. In one or more additional examples, the sample 204 may include one or more samples selected from one or more of whole blood, blood fractions, tissue biopsies, pleural fluid, pericardial fluid, cerebrospinal fluid, and peritoneal fluid.
[0173] Extraction of nucleic acid molecules from sample 204 may include implementing one or more cell lysis techniques to cleave membranes of cells contained in sample 204 and applying one or more proteases to degrade proteins contained in sample 204. Extraction of nucleic acid molecules from sample 204 may also include several washing and / or elution techniques to separate the nucleic acid molecules from other components contained in sample 204. In various examples, thousands, up to millions, or up to billions of nucleic acid molecules may be extracted from sample 204.
[0174] The one or more library preparation and sequencing processes 202 may include one or more separation processes that separate nucleic acid molecules into several partitions based on the characteristics of the nucleic acid molecules. Examples of characteristics that can be used to partition nucleic acid molecules include multiple different nucleotide modifications, methylation levels, nucleosome binding, sequence mismatches, immunoprecipitation, and / or proteins that bind to DNA. In one or more instances, a heterogeneous population of nucleic acid molecules can be partitioned into nucleic acid molecules with one or more epigenetic modifications and nucleic acid molecules without one or more epigenetic modifications. Examples of epigenetic modifications include, but are not limited to, the presence or absence of methylation; the level of methylation, hydroxymethylation, and the type of methylation (5' cytosine or 6 methyladenine).
[0175] In one or more examples, the nucleic acid molecules extracted from the sample 204 may contain nucleic acids with various levels of methylation. The methylation may result from any one or more post-replication or post-transcriptional modifications. Post-replication modifications include, but are not limited to, modifications of the nucleotide cytosine, including 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxylcytosine. One or more library preparation and sequencing processes 202 may separate the nucleic acid molecules extracted from the sample 204 into several partitions, each corresponding to a different level of methylation. For example, one or more library preparation and sequencing processes 202 may produce a first partition of nucleic acid molecules with a first level of methylation, a second partition of nucleic acid molecules with a second level of methylation, and a third partition of nucleic acid molecules with a third level of methylation. In various examples, the second level of methylation may be higher than the first level of methylation, and the third level of methylation may be higher than both the first level of methylation and the second level of methylation. In one or more instances, the one or more library preparation and sequencing processes 202 may include the molecular separation process 124 described with respect to FIG.
[0176] In at least some examples, molecular barcodes can be added to nucleic acids corresponding to one or more methylation distributions. For example, one or more first molecular barcodes can be added to nucleic acids having a first methylation level and included in the first methylation distribution, one or more second molecular barcodes can be added to nucleic acids having a second methylation level and included in the second methylation distribution, and one or more third molecular barcodes can be added to nucleic acids having a third methylation level and included in the third methylation distribution. In scenarios where nucleic acids included in sample 204 are subjected to one or more enrichment processes, molecular barcodes can be added to the enriched nucleic acids.
[0177] The library preparation and sequencing process 202 may also include one or more enrichment processes. The one or more enrichment processes may amplify the number of nucleic acids in the sample 204 that have one or more specific sequences. In various examples, the nucleic acids in the sample 204 may be enriched for a methylation panel region 208. The methylation panel region 208 may correspond to a genomic region of a reference genome that is differentially methylated in subjects with one or more biological conditions. For example, the methylation panel region 208 may include one or more differentially methylated genomic regions involved in subjects with one or more homologous recombination repair deficiencies. In one or more additional examples, the methylation panel region 208 may include one or more differentially methylated genomic regions in individuals with one or more forms of cancer. In at least some examples, the methylation panel region 208 may be the subject of one or more diagnostic assays.
[0178] One or more enrichment processes included in the library preparation and sequencing process 202 can also be performed on genome panel regions 210. The genome panel regions 210 can include one or more portions of a reference genome that are the subject of one or more diagnostic assays. For example, the genome panel regions 210 can correspond to several genomic regions in which at least one somatic mutation or germline mutation is present in individuals with a biological condition. In one or more instances, the genome panel regions 210 can correspond to several genomic regions in which at least one somatic mutation or one or more germline mutations is present in individuals with one or more forms of cancer. In one or more additional instances, the genome panel regions 210 can include driver mutations corresponding to one or more forms of cancer. In various instances, one or more of the methylation panel regions 208 can include or overlap at least a portion of one or more genome panel regions 210.
[0179] In various examples, library preparation and sequencing process 202 may include performing one or more enrichment processes for nucleic acids having one or more amounts of methylation. For example, library preparation and sequencing process 202 may include performing one or more enrichment processes for nucleic acids having genomic regions with at least a threshold amount of methylation. In one or more examples, library preparation and sequencing process 202 may include performing one or more enrichment processes for nucleic acids having at least one hypermethylated genomic region. In one or more instances, library preparation and sequencing process 202 may include performing one or more enrichment processes for nucleic acids having at least a threshold amount of methylation in one or more CG regions corresponding to at least one of methylation panel region 208 or genomic panel region 210.
[0180] The library preparation and sequencing process 202 may include performing one or more amplification processes and one or more sequencing processes to generate sequencing data 212. The sequencing data 212 may include an alphanumeric representation of the nucleic acids contained in the amplification products produced by the one or more library preparation and sequencing processes 202. For example, the sequencing data 212 may include, for each nucleic acid in the amplification product, data corresponding to a string representing each strand of nucleotides corresponding to the individual nucleic acid. The sequencing data 212 may be stored in one or more data files.
[0181] Framework 200 may also include determining sequence reads for one or more methylation distributions in operation 214. In one or more examples, the sequencing data 212 may be analyzed to determine sequence reads corresponding to nucleic acids having at least one CG region corresponding to the at least one methylation distribution. In one or more instances, determining sequence reads for the one or more methylation distributions in operation 214 may include determining sequence reads corresponding to a hypermethylated distribution included in the sequencing data 212. In at least some examples, determining sequence reads corresponding to the at least one methylation distribution may include analyzing the sequencing data 212 to determine sequence reads having one or more molecular barcodes corresponding to the at least one methylation distribution.
[0182] In operation 214, training data 216 can be generated by determining sequence reads included in the sequencing data 212 that correspond to one or more methylation distributions. The training data 216 can be used to train at least one of one or more machine learning models or one or more statistical models to determine a subject's homologous recombination deficiency status. The training data 216 can include methylation panel methylation data 218. The methylation panel methylation data 218 can include sequence reads included in the sequencing data 212 that correspond to nucleic acids included in one or more methylation distributions and that correspond to the methylation panel regions 208. In various examples, the sequence reads in the sequencing data 212 that correspond to at least one methylation distribution can be further analyzed to determine a subset of sequence reads that correspond to at least one methylation distribution and that also correspond to the methylation panel regions 208. In one or more instances, the methylation panel methylation data 218 can include sequence reads that correspond to the methylation panel regions 208 and have at least a threshold amount of methylated cytosines in CG regions of the methylation panel regions 208. For example, in operation 214, the sequencing data 212 can be analyzed to generate methylation panel methylation data 218, where the methylation panel methylation data 218 thus corresponds to a high methylation distribution and includes sequence reads that correspond to the methylation panel regions 208. In one or more examples, the methylation panel methylation data 218 can be determined by analyzing the sequencing data 212 to determine sequence reads that include one or more molecular barcodes that correspond to the methylation panel regions 208. In one or more additional examples, the methylation panel methylation data 218 can be determined by analyzing the sequencing data 212 to determine sequence reads that have at least a threshold amount of homology to the methylation panel regions 208.
[0183] The training data 216 may also include genome panel methylation data 220. The genome panel methylation data 220 may include sequence reads included in the sequencing data 212 that correspond to nucleic acids included in one or more methylation distributions and that correspond to the genome panel regions 210. In one or more examples, the sequence reads in the sequencing data 212 that correspond to at least one methylation distribution can be further analyzed to determine an additional subset of sequence reads that correspond to at least one methylation distribution and that also correspond to the genome panel regions 210. In one or more instances, the genome panel methylation data 220 may include sequence reads that correspond to the genome panel regions 210 and have additional less-than-threshold amounts of methylated cytosines in CG regions of the genome panel regions 210. As an example, in operation 214, the sequencing data 212 can be analyzed to generate screening panel methylation data 220, such that the screening panel methylation data 220 includes sequence reads that correspond to a low methylation distribution and that correspond to the genome panel regions 210. In one or more examples, the methylation data 220 of the genome panel can be determined by analyzing the sequencing data 212 to determine sequence reads that include one or more molecular barcodes that correspond to the genome panel regions 210. In one or more additional examples, the methylation data 220 of the genome panel can be determined by analyzing the sequencing data 212 to determine sequence reads that have at least a threshold amount of homology to the genome panel regions 210.
[0184] One or more alignment processes can be performed to generate training data 216. For example, one or more alignment processes can be performed to determine the amount of homology between sequence reads included in sequencing data 212 and methylation panel regions 208 to determine methylation panel methylation data 218. Further, one or more alignment processes can be performed to determine the amount of homology between sequence reads included in sequencing data 212 and genomic panel regions 210 to determine genomic panel methylation data 220. Further, one or more alignment processes can be performed to determine the amount of homology between sequence reads included in sequencing data 212 and one or more molecular barcodes. The one or more molecular barcodes can correspond to at least one of one or more methylation distributions, methylation panel regions 208, or genomic panel regions 210.
[0185] The amount of homology between a given sequence read and one or more genomic regions of a reference sequence can indicate the number of positions of the reference sequence that have the same nucleotide as the corresponding position of the given sequence read.Based on determining that the genomic region of the sequence read and the genomic region of the reference sequence have at least a threshold amount of homology, the sequence read can be aligned with the genomic region of the reference sequence.In the scenario where the sequence read has at least a threshold amount of homology with multiple genomic regions of the reference sequence, the genomic region of the reference sequence that has the largest amount of homology with the sequence read can be determined to be aligned with the sequence read.
[0186] The amount of homology between a given sequence read and a portion of a reference sequence can be determined using the BLAST program (basic local alignment search tool) and PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), or by using the Gap program (Wisconsin Sequence Analysis Package, Genetics Computer Group, University Research Park, Madison Wis.) with default settings, which uses the Needleman and Wunsch algorithm (J. Mol. Biol. 48; 443-453 (1970)). The amount of homology between a sequence read and a portion of a reference sequence can also be determined using the Burrows-Wheeler aligner (Li, H., & Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics, 25(14), 1754-1760).
[0187] In one or more instances, the sequencing data 212 can be analyzed to determine a first group of sequence reads corresponding to a high methylation distribution. The first group of sequence reads can also be analyzed to determine a first subset of the first group of sequence reads that correspond to the methylation panel regions 208 and a second subset of the first group of sequence reads that correspond to the genomic panel regions 210. In these scenarios, the training data 216 can include sequence reads that correspond to nucleic acids having at least a threshold amount of methylation associated with a high methylation distribution and that also correspond to the methylation panel regions 208 and the genomic panel regions 210.
[0188] In one or more additional instances, the sequencing data 212 can be analyzed to determine a second group of sequence reads corresponding to a low methylation distribution. The second group of sequence reads can also be analyzed to determine a first subset of the second group of sequence reads corresponding to the methylation panel region 208 and a second subset of the second group of sequence reads corresponding to the genomic panel region 210. In these situations, the training data 216 can include sequence reads that correspond to nucleic acids having at least a threshold amount of methylation associated with a high methylation distribution and having an additional threshold amount or less of methylation associated with a low methylation distribution, and that also correspond to the methylation panel region 208 and the genomic panel region 210.
[0189] In at least some instances, training data 216 may include both a first group of sequence reads corresponding to a high methylation distribution and a second group of sequence reads corresponding to a low methylation distribution. In these cases, training data 216 may include sequence reads corresponding to a first nucleic acid having at least a threshold amount of methylation associated with a high methylation distribution and a second nucleic acid having an additional threshold amount or less of methylation associated with a low methylation distribution, and also corresponding to methylation panel region 208 and genomic panel region 210.
[0190] In various examples, at least one of a tumor fraction or a tumor cell copy number for a given sample 204 can be generated based on a portion of the sequencing data 212 corresponding to the given sample 204. In one or more instances, the tumor fraction and / or tumor cell copy number can be calculated as described in U.S. Patent Application No. 17 / 691,049, filed March 9, 2022, which is incorporated herein in its entirety. A subset of samples 204 having at least one of a threshold tumor fraction or a threshold tumor cell copy number can be determined. In various examples, the sequence reads included in the training data 216 can be derived from a subset of samples 204 having at least the threshold tumor fraction and / or the threshold tumor cell copy number.
[0191] The architecture 200 may include an HRD state computing system 222 that obtains training data 216 and analyzes the training data 216 to generate one or more models for determining a subject's HRD state. The HRD state computing system 222 may include one or more computing devices 224. The one or more computing devices 224 may include at least one of one or more desktop computing devices, one or more mobile computing devices, or one or more server computing devices. In various examples, at least a portion of the one or more computing devices 224 may be included in a remote computing environment, such as a cloud computing environment. In one or more examples, the library preparation and sequencing process 202, the determining sequence reads for one or more methylation distributions in operation 214, and the operations performed by the computing system 224 may be performed by a single entity. In one or more additional examples, the library preparation and sequencing process 202, the determining sequence reads for one or more methylation distributions in operation 214, and the operations performed by the computing system 224 may be performed by multiple organizations.
[0192] In operation 226, the HRD state computing system 222 can generate a training quantitative measure 228. The training quantitative measure 228 corresponds to sequence reads included in the training data 216 and may correspond to the number of sequence representations that correspond to individual classification regions of the reference sequence. In one or more examples, prior to determining the training quantitative measure 228, the HRD state computing system 222 can identify one or more groups of sequence representations. For example, individual sequence representations may correspond to individual sequencing reads included in the sequencing data 212. In these scenarios, a sequence representation may include multiple reads that correspond to a single nucleic acid molecule included in the sample 204. In one or more additional examples, a sequence representation may correspond to individual nucleic acid molecules included in the sample 204. In these situations, the HRD state computing system 222 can determine groups of reads included in the training data 216 that correspond to individual nucleic acid molecules included in the sample 204 based on molecular barcodes common to each group of sequencing reads. That is, each individual nucleic acid molecule contained in sample 204 can be encoded with a molecular barcode that uniquely identifies the individual nucleic acid molecule, and in at least some cases, the individual nucleic acid molecule can be represented by multiple sequencing reads included in training data 216. Thus, when multiple sequence representations corresponding to a single nucleic acid molecule contained in sample 204 are present in training data 216, the multiple sequence representations can be grouped together by HRD state computing system 222. In various examples, a group of sequence representations corresponding to a single nucleic acid molecule contained in sample 204 may be referred to herein as a "family." Furthermore, the start and end positions relative to a reference sequence of sequence representations having a common molecular barcode can be used to group the sequence representations corresponding to each individual nucleic acid contained in sample 204. In one or more instances, an individual sequence representation representing a family of sequence representations corresponding to a single nucleic acid molecule contained in sample 204 may be referred to herein as a "consensus sequence representation."
[0193] One or more classification regions may correspond to a genomic region of the reference sequence that has a certain amount of methylation in cfDNA obtained from a subject with a homologous recombination repair deficiency, compared to a certain amount of methylation in a genomic region in cfDNA obtained from a subject with a homologous recombination repair deficiency. One or more classification regions may also include at least a threshold amount of cytosine-guanine content. In various examples, one or more classification regions may include a series of cytosine-guanine (CG) pairs (CpG sites) in the 5'->3' direction, for example, at least 3 CpG sites, at least 5 CpG sites, at least 8 CpG sites, at least 10 CpG sites, at least 12 CpG sites, at least 15 CpG sites, at least 18 CpG sites, or at least 20 CpG sites. In one or more embodiments, the training quantitative measure 228 may indicate the number of sequence representations from the training data 216 corresponding to each genome panel region 210. Further, the one or more classification regions may correspond to genomic regions of the reference sequence that contain at least one of one or more germline mutations or one or more somatic mutations in individuals with a homologous recombination repair deficiency, and / or genomic regions that correspond to differentially methylated regions in individuals with at least one HRD or one or more forms of cancer. In one or more additional examples, the training quantitative measure 228 may indicate a number of sequence representations from the training data 226 that correspond to individual methylation panel regions 208. In at least some examples, the training quantitative measure 228 may indicate a number of sequence representations from the training data 216 that correspond to individual methylation panel regions 208 and a number of sequence representations from the training data 216 that correspond to individual genomic panel regions 210.
[0194] In various examples, the training quantitative measure 228 may include a normalized quantitative measure. The normalized quantitative measure may be determined by analyzing the number of sequence representations from the training data 216 corresponding to a classification region with respect to the number of sequence representations from the training data 216 corresponding to one or more control regions. In one or more examples, the normalized quantitative measure may be determined by analyzing the number of sequence representations from the training data 216 corresponding to a classification region with respect to at least one of a first number of sequence representations from the training data 216 corresponding to one or more positive control regions or a second number of sequence representations from the training data 216 corresponding to one or more negative control regions. The positive control region may have at least a threshold amount of sequence representations corresponding to nucleic acids having methylated cytosines and may include a genomic region of the reference sequence containing at least a threshold number of CpG sites. The positive control region may have at least a threshold amount of sequence representations corresponding to nucleic acids having methylated cytosines in a sample obtained from a subject with a homology-directed repair deficiency and in a sample obtained from a subject without a homology-directed repair deficiency. In one or more examples, the negative control region may comprise a genomic region of the reference sequence having a sequence representation corresponding to a nucleic acid having less than a threshold amount of methylated cytosine and at least a threshold number of CpG sites.The negative control region may have a sequence representation corresponding to a nucleic acid having less than a threshold amount of methylated cytosine in a sample obtained from a subject with a homology-directed repair deficiency and in a sample obtained from a subject without a homology-directed repair deficiency.
[0195] The training quantitative measures 228 can also be normalized with respect to guanine-cytosine (GC) content. For example, for each methylation panel region 208 and / or each genome panel region 210, a GC content indicating the number of guanine nucleotides and the number of cytosine nucleotides in the sequence representation corresponding to the individual methylation panel region 208 and / or each genome panel region 210 can be determined. Furthermore, the frequency of GC content can be determined for one of a plurality of GC content distributions. Each GC content distribution can correspond to a different range of GC content values. In this manner, the frequency of GC content for a given methylation panel region 208 or a given genome panel region 210 can be represented by the GC content distribution for the individual methylation panel region 208 and / or each genome panel region 210. The amount of expected coverage for each methylation panel region 208 and / or each genome panel region 210 can be determined based on the frequency of GC content for the methylation panel region 208 and / or each genome panel region 210. At least a portion of the normalized training quantitative measures may include GC-normalized coverage data determined based on the amount of coverage expected for each HRD genomic region 208 and / or each screening panel region 210.
[0196] In operation 230, the HRD status computing system 222 can analyze the training quantitative measure 228 to generate predictor regions 232. The predictor regions 232 can indicate genomic regions included in at least one of the methylation panel regions 208 or the genomic panel regions 210 that include a methylation pattern indicative of a homologous recombination repair deficiency. For example, the HRD status computing system 222 can analyze the training quantitative measure 228 to determine CG regions that have a different amount of methylation in subjects with a homologous recombination repair deficiency relative to subjects without a homologous recombination repair deficiency. In one or more examples, at least a portion of the predictor regions 232 can include genomic regions corresponding to at least a subset of the methylation panel regions 208 or the genomic panel regions 210 in which the CG regions have a greater amount of methylation in subjects with a homologous recombination repair deficiency than in subjects without a homologous recombination repair deficiency. In one or more additional examples, at least a portion of the predictor region 232 may include genomic regions corresponding to at least a subset of the methylation panel region 208 or genomic panel region 210 in which the CG regions have less methylation in subjects with a homology-directed repair deficiency than in subjects without a homology-directed repair deficiency.
[0197] The HRD state computing system 222 may implement at least one of one or more machine learning techniques or one or more statistical techniques to determine the predictor regions 232. In one or more examples, the HRD state computing system 222 may implement one or more logistic regression models to determine the predictor regions 232. In one or more instances, the HRD state computing system 222 may implement one or more elastic net linear regression algorithms to determine the predictor regions 232. In one or more additional instances, the HRD state computing system 222 may implement one or more lasso regression techniques to determine the predictor regions 232. In various examples, the HRD state computing system 222 may analyze quantitative measures corresponding to hundreds of genomic regions, up to thousands of genomic regions, or up to tens of thousands of genomic regions to determine the predictor regions 232.
[0198] In operation 234, the HRD status computing system 222 can use the predictor regions to generate a predictive model for HRD status, e.g., an HRD status predictor model 236. The HRD status predictor model 236 can include individual components corresponding to individual predictor regions 232. For example, the HRD status predictor model 236 can include individual variables with individual weights corresponding to individual predictor regions 232. The HRD status prediction model 236 can generate a model output 238 having at least two outcomes: HRD positive 240 or HRD negative 242. In one or more examples, the model output 238 generated using the HRD status prediction model 238 can correspond to the probability that a given sample is derived from a subject in whom a homologous recombination repair deficiency is present. The HRD status computing system 222 can analyze the probability that a given sample is derived from a subject in whom a homologous recombination repair deficiency is present relative to a threshold probability to determine whether the model output 238 is HRD positive 240 or HRD negative 242.
[0199] In one or more instances, after the HRD status computing system 222 generates the HRD status prediction model 236 using the training quantitative measures 228, the HRD status prediction model 236 can be used to determine the HRD status of additional subjects 244 not included in the training subjects 206. For example, the subject sequencing data 246 can be generated using one or more samples from the additional subjects 244. By way of example, the subject sequencing data 246 can be generated from one or more cfDNA samples from the additional subjects 244. In various instances, the subject sequencing data 246 can be generated according to a process similar or the same as the one or more library preparation and sequencing processes 202. The subject sequencing data 246 can include sequence representations corresponding to nucleotide sequences of nucleic acids included in the one or more samples from the additional subjects 244. In at least some instances, the sequence representations included in the subject sequencing data 246 can correspond to sequence reads generated by the one or more library preparation and sequencing processes.
[0200] In operation 248, the subject sequencing data 246 can be analyzed to generate a subject quantitative measure. In one or more examples, operation 248 can be performed by the HRD status computing system 222. The subject quantitative measure generated in operation 248 can be included in model input data 250. The model input data 250 can be provided to the HRD status prediction model 236 to determine the HRD status of the additional subject 244. In one or more instances, the model input data 250 can include normalized quantitative metrics corresponding to the predictor region 232. For example, in operation 248, a number of sequence representations corresponding to nucleic acids contained in one or more samples from the additional subject 244 that have at least a threshold amount of methylation can be determined. In at least some examples, in operation 248, sequence representations corresponding to nucleic acids corresponding to a hypermethylated distribution can be identified. Further, in operation 248, at least one of the number of sequence representations or the number of nucleic acids associated with one or more samples obtained from the additional subject 244 and corresponding to the predictor region 232 can be determined. Additionally, one or more normalization procedures can be performed to generate normalized quantitative measures determined based on sequence representations from subject sequencing data 246 corresponding to at least a portion of predictor region 232. In various examples, model input data 250 can include input vectors representing normalized quantitative measures determined using subject sequencing data 246.
[0201] Based on the model input data 250, the HRD status prediction model 236 can determine a model output 238 for the additional subject 244. The model output 238 can indicate a probability of a homologous recombination repair deficiency being present in the additional subject 244. The model output 238 can indicate that the additional subject 244 corresponds to an HRD-positive 240 status or an HRD-negative 242 status. In one or more instances, the model output 238 can be used to determine that the additional subject 244 has an HRD-positive status 240 based on the probability that a homologous recombination repair deficiency is present in the additional subject 244 being at least a threshold probability. In one or more examples, the model output 238 can be used to determine one or more treatment recommendations for the additional subject 244. For example, in a scenario where the model output 238 for the additional subject 244 is HRD-positive 240, the treatment for the additional subject 244 can include one or more PARP inhibitors.
[0202] Figure 3 is a schematic diagram of an example framework 300 for generating one or more computational models for determining a subject's homologous recombination deficiency status, according to one or more implementations. In at least some examples, the operations described in connection with framework 300 can be performed by the HRD status computing system 200 described in connection with Figure 2. Framework 300 can include computational analysis 302 performed on methylation training data 304. Methylation training data 304 can be derived from a first group of training subjects 306 that are not homologous recombination repair deficient and a second group of training subjects 308 that are not homologous recombination repair deficient.
[0203] The methylation training data 304 may include sequence representations of nucleic acids derived from samples obtained from the first group of training subjects 306 and the second group of training subjects 308. The sequence representations may correspond to polynucleotide molecules or sequence reads derived from samples obtained from the first group of training subjects 306 and the second group of training subjects 308. The methylation training data 304 may also include sequence representations corresponding to nucleic acids having at least a threshold amount of methylation in CG regions of one or more classification regions. For example, the methylation training data 304 may include sequence representations corresponding to nucleic acids that fall within a high methylation distribution for one or more CG regions of at least one classification region. In one or more additional examples, the methylation training data 304 may include sequence representations corresponding to nucleic acids that fall within a low methylation distribution for one or more CG regions of at least one classification region. The one or more classification regions may correspond to genomic regions of a reference genome that may include one or more genomic mutations in individuals in which one or more biological conditions exist. In at least some examples, at least a portion of the classification regions may be differentially methylated in individuals in which one or more biological conditions exist.
[0204] The computer analysis 302 may include analyzing the methylation training data 304 to determine a number of predictor regions 310. The predictor regions 310 may include a subset of the classifier regions. In one or more examples, the predictor regions 310 may be differentially methylated in individuals with a homologous recombination repair deficiency compared to individuals without a homologous recombination repair deficiency. In one or more additional examples, the predictor regions 310 may be differentially methylated in individuals with one or more forms of cancer. In various examples, the predictor regions may include C-G regions with a threshold number of methylated cytosines. For example, at least a portion of the predictor regions 310 may fall within a hypermethylated distribution. Furthermore, the predictor regions 310 may include C-G regions with an additional or fewer methylated cytosines. By way of example, at least a portion of the predictor regions 310 may fall within a hypomethylated distribution.
[0205] In various examples, the predictor region 310 may include a genome panel predictor region 312. The genome panel HRD predictor region 312 may include a genomic region that is part of a screening panel. The screening panel may include a diagnostic process for identifying the presence of a genomic mutation that may be indicative of one or more biological conditions. In one or more examples, the screening panel may include a diagnostic process for identifying the presence of a genomic mutation that may be indicative of one or more forms of cancer. In one or more instances, the genome panel predictor region 312 may include a genomic region corresponding to one or more genomic regions containing at least one of one or more somatic mutations or one or more germline mutations that may be indicative of one or more forms of cancer. Furthermore, the genome panel predictor region 312 may predict the presence of a homologous recombination repair deficiency in a subject.
[0206] The predictor regions 310 may also include a methylation panel predictor region 314. The methylation panel predictor region 314 may include a genomic region that is differentially methylated in individuals with a homologous recombination repair deficiency. Additionally, the methylation panel predictor region 314 may include a portion of one or more genes that are differentially methylated in individuals with a homologous recombination repair deficiency one or more forms of cancer. In various examples, the computer analysis 302 includes analyzing a portion of a gene that is differentially methylated in a first group of training subjects 306 versus a second group of training subjects 308.
[0207] The computer analysis 302 may include generating a model for each classification region, wherein the model generates an indicator of the subject's HRD status. In one or more examples, the indicator of the subject's HRD status may include a probability that the one or more subjects have a homologous recombination repair deficiency. In various examples, the computer analysis 302 may generate a logistic regression model for each classification region to determine the indicator of the subject's HRD status. In at least some examples, the model generated for each classification region may have input data corresponding to the individual classification region and including a quantitative measure indicative of a form of cancer, where one or more mutations are present in the individual classification region in subjects with the form of cancer.
[0208] The quantitative measure may include a metric indicating the number of sequence representations corresponding to each classification region and corresponding to at least one methylation distribution. The number of sequence representations may correspond to the number of sequence reads derived from nucleic acids contained in one or more samples obtained from the first group of training subjects 306 and the second group of training subjects 308 corresponding to each classification region, or the number of nucleic acids present in one or more samples obtained from the first group of training subjects 306 and the second group of training subjects 308 corresponding to each classification region. In one or more examples, the quantitative measure may include a normalized quantitative measure. The normalized quantitative measure may include a ratio of the count of sequence representations derived from samples obtained from the first group of training subjects 306 and the second group of training subjects 308 corresponding to each classification region to the count of sequence representations derived from samples obtained from the first group of training subjects 306 and the second group of training subjects 308 corresponding to one or more control regions of the reference sequence. In various examples, the normalized quantitative measure may also be generated using a CG normalization process.
[0209] In one or more instances, the computer analysis 302 may include analyzing quantitative measures corresponding to at least 1,000 genomic regions, at least 5,000 genomic regions, at least 8,000 genomic regions, at least 10,000 genomic regions, at least 12,000 genomic regions, at least 15,000 genomic regions, at least 18,000 genomic regions, at least 20,000 genomic regions, at least 25,000 genomic regions, or at least 30,000 genomic regions to determine predictor regions 310. At least one of the panel HRD predictor regions 312 or genomic HRD predictor regions 314 may include at least 25 genomic regions, at least 50 genomic regions, at least 75 genomic regions, at least 100 genomic regions, at least 150 genomic regions, at least 200 genomic regions, at least 250 genomic regions, at least 300 genomic regions, at least 350 genomic regions, at least 400 genomic regions, at least 450 genomic regions, or at least 500 genomic regions.
[0210] Analysis of the quantitative measures for each classification region can generate a p-value for the individual classification region. The individual p-value can indicate a measure of the significance of the individual classification region in determining whether the indicator of HRD status determined using the logistic regression model for the individual classification region accurately corresponds to the HRD status of individuals included in at least one of the first group of training subjects 306 or the second group of training subjects. In various examples, the samples obtained from the first group of training subjects 306 and the second group of training subjects 308 can be divided so that a first portion of the samples obtained from the first group of training subjects 306 and the second group of training subjects 308 can be used as training samples for generating a logistic regression model for the individual classification region. A second portion of the samples obtained from the first group of training subjects 306 and the second group of training subjects 308 can be used as test samples to determine the p-value for the individual classification region.
[0211] The classification regions can be ordered according to the p-values corresponding to each classification region. In at least some examples, the lower the p-value for an individual classification region, the more significant the classification region in predicting the subject's HRD status. In one or more examples, the classification regions can be ranked from the classification region with the lowest p-value to the classification region with the highest p-value. In various examples, a subset of the classification regions can correspond to the predictor regions 310 according to their p-values and can be selected for display in the computational model 316. The classification regions selected for display in the computational model 316 can include at least one of one or more genome panel predictor regions 312 or one or more methylation panel predictor regions 314. In one or more instances, 50 classification regions with the lowest p-values may be selected for display in the computational model 316, 100 classification regions with the lowest p-values may be selected for display in the computational model 316, 150 classification regions with the lowest p-values may be selected for display in the computational model 316, 200 classification regions with the lowest p-values may be selected for display in the computational model 316, 250 classification regions with the lowest p-values may be selected for display in the computational model 316, 300 classification regions with the lowest p-values may be selected for display in the computational model 316, 350 classification regions with the lowest p-values may be selected. The classification regions with the 400 lowest p-values may be selected for display in the computational model 316, the classification regions with the 450 lowest p-values may be selected for display in the computational model 316, the classification regions with the 500 lowest p-values may be selected for display in the computational model 316, the classification regions with the 750 lowest p-values may be selected for display in the computational model 316, the classification regions with the 1000 lowest p-values may be selected for display in the computational model 316, or the regions with the 1500 lowest p-values may be selected for display in the computational model 316.
[0212] The computational model 316 may include several components. The components of the computational model 316 may include variables that may predict a subject's HRD status. In one or more examples, an individual component of the computational model 316 may correspond to at least one of the one or more predictor regions 310. In various examples, an individual component of the computational model 316 may be determined based on a sequence representation and one or more genome panel predictor regions 312 included in the methylation training data 304 that correspond to a low methylation distribution. Furthermore, an individual component of the computational model 316 may be determined based on a sequence representation and one or more methylation panel predictor regions 314 included in the methylation training data 304 that correspond to a high methylation distribution.
[0213] In the example of FIG. 3 , the computational model 316 may include several components, such as a first model component 318, a second model component 320, and a third model component 322. In one or more examples, the first model component 318 may include one or more first functions, the second model component 320 may include one or more second functions, and the third model component 322 may include one or more third functions. The first model component 318 may correspond to a first predictor region 324, the second model component 320 may correspond to a second predictor region 326, and the third model component 322 may correspond to a third predictor region 328. Additionally, the components of the computational model 316 may correspond to at least one parameter that may indicate a measure of the component's significance in determining a subject's HRD status. For example, the components of the computational model 316 may correspond to one or more weights. As an example, the first model component 318 may have a first weight 330, the second model component 320 may have a second weight 332, and the third model component 322 may have a third weight 334.
[0214] A computational model 316 can be trained using the methylation training data 304 and the quantitative measure determined for the predictor region 310. In one or more examples, sequence representations from a first group of training subjects 306 and a second group of training subjects 308 can be analyzed to determine the amount of methylation in one or more CG regions of each sequence representation. Further, sequence representations generated from samples obtained from the first group of training subjects 306 and the second group of training subjects 308 having genomic regions with at least a threshold amount of methylation in CG regions can be aligned to a reference genome. The aligned sequence representations can then be analyzed to determine the amount of homology between the sequence representation and the predictor region 310. A count of the aligned sequence representations corresponding to the predictor region 310 can be determined. Further, a normalized quantitative measure can be determined based on the count of the aligned sequence representations corresponding to the predictor region 310 and the count of the sequence representations corresponding to at least one of the one or more positive control regions or one or more negative control regions.
[0215] In at least some examples, the computational model 316 may include one or more machine learning models or one or more statistical models generated using normalized quantitative measures of the predictor space 310. In one or more instances, the computational model 316 may include a logistic regression model derived from samples obtained from the first group of training subjects 306 and the second group of training subjects 308 and trained and validated using normalized quantitative measures corresponding to the predictor space 310. In one or more examples, one or more least absolute shrinkage and selection operator (lasso) regression techniques may be used to generate the computational model 316, including the logistic regression model. In one or more additional examples, the training process of the computational model 316 may include performing one or more elastic regularization processes. In various examples, the training process of the computational model 316 may include using one or more validation techniques and optimizing one or more tuning parameters based on the normalized quantitative measures to generate the computational model 316. Optimization of the tuning parameters may be performed to minimize overfitting of the training data to the computational model 316. In one or more additional instances, the training process for the computational model 316 may include a lambda optimization process to generate the computational model 316. In various examples, the lambda optimization process may determine one or more parameters corresponding to the components of the computational model 316. By way of example, the lambda optimization process may determine a first weight 330 for the first model component 318, a second weight 332 for the second model component 320, and a third weight 334 for the third model component 322.
[0216] The computational model 316 can generate a model output 336. For one or more samples obtained from a given subject, the model output 336 can indicate the subject's status with respect to homology-directed repair deficiency. By way of example, the model output 336 can indicate the probability that a homology-directed repair deficiency is present in the individual. In one or more instances, the model output 336 can include a probit value associated with the subject's status with respect to the presence of a homology-directed repair deficiency in the subject.
[0217] In one or more examples, the computational model 316 can generate a model output 336 indicative of the HRD status of an individual, where the individual may have different forms of cancer. For example, the computational model 316 can determine the probability of a homologous recombination repair deficiency being present in a first group of subjects presenting a first form of cancer and the probability of a homologous recombination repair deficiency being present in a second group of subjects presenting a second form of cancer. In one or more additional examples, the computational model 316 can determine the probability of a homologous recombination repair deficiency being present in subjects presenting a particular form of cancer. By way of example, the computational model 316 can determine the probability of a homologous recombination repair deficiency being present in subjects presenting colorectal cancer or the probability of a homologous recombination repair deficiency being present in subjects presenting prostate cancer. In at least some examples, the form of cancer for which the computational model 316 can be used to determine the HRD status of a subject can depend on the form of cancer present in the first group of training subjects 306 and the second group of training subjects 308.
[0218] In various examples, a threshold probability for the presence of a homology repair deficiency can be used in the computational model 316 to determine the HRD status of one or more subjects. The threshold probability can be determined by analyzing the model output 336 for subjects in which a homology repair deficiency is present and subjects in which a homology repair deficiency is not present. The threshold probability can correspond to the probability of displaying a model output 336 that captures the maximum number of subjects in which a homology repair deficiency is present. In at least some examples, the threshold probability used in the computational model 316 to determine whether a homology repair deficiency is present in a subject can be different for different forms of cancer.
[0219] In one or more examples, the computational model 316 may include additional components for generating the model output 336 in addition to the model components corresponding to the predictor regions 310. For example, the computational model 316 may include one or more first additional components corresponding to copy number variations in one or more genomic regions associated with regulation of the homology-directed repair pathway. Furthermore, the computational model 316 may include one or more second additional components corresponding to loss of heterozygosity in one or more genomic regions associated with regulation of the homology-directed repair pathway. At least one of the first additional component or the second additional component of the computational model 316 can be used to generate the model output 336.
[0220] 4 is a flow diagram of an example process 400 for determining the probability of a homologous recombination repair deficiency in one or more subjects based on methylation data from samples obtained from the one or more subjects, according to one or more implementations. In operation 402, process 400 includes obtaining training sequence data including training sequence representations from multiple samples of multiple training subjects. The sequence representations may correspond to sequence reads generated by one or more sequencing machines. In these scenarios, the sequence reads may correspond to nucleic acids extracted from multiple samples obtained from the multiple training subjects. In one or more additional examples, the sequence representations may represent nucleic acids contained in the multiple samples. In various examples, each sequence representation may correspond to a family of sequence reads corresponding to individual nucleic acids contained in the multiple samples.
[0221] Process 400 may also include, in operation 404, determining a subset of training sequence reads corresponding to nucleic acids having at least a threshold amount of methylated cytosines within one or more regions of the nucleotide sequence of the nucleic acid. Each training sequencing representation may correspond to at least a portion of a nucleic acid from one sample of a plurality of samples having a CG region with a threshold amount of methylated cytosines. In one or more instances, the plurality of samples may include cell-free nucleic acids. In one or more instances, the methylated cytosines may be determined using at least one of bisulfite conversion and sequencing, Tet-assisted bisulfite sequencing (TAB-Seq), differential enzymatic cleavage, treatment with MSRE, or MBD partitioning. In one or more additional instances, the methylated cytosines may be determined using one or more single-molecule sequencing methods, such as nanopore DNA sequencing or those described in Eid, J., et al. (2009) "Real-time DNA sequencing from single polymerase molecules". Science, 323(5910), 133-138.
[0222] In operation 406, process 400 may include analyzing a subset of the training sequence representations to determine a quantitative measure derived from the subset of training sequence representations. The quantitative measure may represent a quantity of sequence representations corresponding to one or more genomic regions of the reference sequence. In one or more examples, the quantitative measure may represent a quantity of sequence representations corresponding to a classification region of the reference genome. The classification region may include a promoter region corresponding to a genomic region containing one or more mutations present in individuals with one or more forms of cancer. The classification region may also include a differentially methylated region, which includes a genomic region having a different amount of cytosine methylation in a CG region of an individual with one or more forms of cancer compared to the amount of additional cytosine methylation in a CG region of an individual without cancer. Furthermore, the classification region may include one or more genomic regions enriched as part of a screening panel used to identify individuals with one or more forms of cancer. Furthermore, the classification region may include one or more genomic regions containing one or more mutations present in individuals with a homologous recombination repair deficiency. In these scenarios, the classification region may include at least one of at least a portion of the ATM gene, at least a portion of the BRCA1 gene, at least a portion of the BRCA2 gene, at least a portion of the CDK12 gene, at least a portion of the CHEK2 gene, at least a portion of the PALB2 gene, or at least a portion of the RAD51D gene.
[0223] The quantitative measure may also include a normalized quantitative measure. The normalized quantitative measure may correspond to a subset of training sequence representations having at least a threshold amount of methylated cytosines in CG regions included in a classification region, compared to a subset of training sequence representations having at least a threshold amount of methylated cytosines in CG regions included in one or more control regions. For example, the normalized quantitative measure may indicate the ratio of the count of the subset of training sequence representations having at least a threshold amount of methylated cytosines in CG regions included in a classification region to the count of the subset of training sequence representations having at least a threshold amount of methylated cytosines in CG regions included in one or more control regions. The one or more control regions may include genomic regions of the reference sequence having at least a threshold amount of methylated cytosines in CG regions of individuals with one or more forms of cancer and in additional individuals without cancer.
[0224] Furthermore, in operation 408, process 400 may include analyzing the quantitative measures to determine a subset of the plurality of classification regions having at least a threshold likelihood of indicating a homology-directed repair deficiency. In one or more examples, one or more models can be generated for each classification region, and the models can then be implemented to determine the subject's status with respect to homology-directed repair deficiency. For example, for each classification region, at least one of one or more machine learning regression models or statistical regression models can be generated based on the quantitative measures for each classification region. A measure of significance for each classification region can be determined based on the one or more models corresponding to each classification region. In one or more instances, a p-value can be calculated for the one or more models for each classification region. The p-value can be used to rank the classification regions to indicate a measure of significance for the classification region in identifying individuals with a homology-directed repair deficiency. The subset of classification regions can be selected according to the number of classification regions having at least a threshold amount of significance in determining individuals with a homology-directed repair deficiency.
[0225] Further, at operation 410, process 400 may include generating a predictive computational model for determining the probability of a homologous recombination repair deficiency in one or more additional subjects. The predictive computational model may include several components, with each component corresponding to at least one classification region of a subset of the multiple classification regions. In one or more examples, the predictive computational model may be generated using one or more elastic regularization techniques to minimize overfitting of the predictive computational model to the data used to train the predictive model. In at least some examples, the predictive computational model may include a machine learning-based regression model or a statistically-based regression model. In one or more instances, the predictive computational model may include a logistic regression model. The predictive model may determine an indicator of HRD status for the subject based on the probability of a homologous recombination deficiency being present in the subject. In one or more instances, the predictive computational model may generate an output indicating a positive HRD status for the subject by determining that the probability of a homologous recombination repair deficiency being present is at least a threshold probability. In one or more additional instances, the predictive computational model may generate an output indicating a negative HRD status for the subject by determining that the probability of a homologous recombination repair deficiency being present is less than a threshold probability.
[0226] In various examples, sequence representations provided to a predictive computational model during or after a training process have at least a threshold amount of cytosine methylation within classification regions. Sequence representations that meet the methylation levels can be produced, at least in part, using one or more molecular separation processes. The molecular separation process can include combining a plurality of nucleic acids from at least one of a subject's blood or tissue with a solution containing a certain amount of methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution. Multiple washes of the nucleic acid-MBD protein solution can then be performed with a salt solution to produce several nucleic acid fractions. Each nucleic acid fraction can have a threshold number of molecules with methylated cytosines within a region of the plurality of nucleic acids that has at least a threshold cytosine-guanine content. In one or more instances, one of the multiple washes can be performed with a solution having a sodium chloride (NaCl) concentration to produce one of the several nucleic acid fractions with a range of binding energies for the MBD protein.
[0227] In one or more examples, a first nucleic acid fraction can be determined to be associated with a first distribution of a plurality of distributions of nucleic acids. The first distribution corresponds to a first range of binding energies for MBD proteins. Further, a first molecular barcode can be bound to nucleic acids of the first nucleic acid fraction. The first molecular barcode can be associated with the first distribution. Further, a second nucleic acid fraction can be determined to be associated with a second distribution of the plurality of distributions of nucleic acids. The second distribution can correspond to a second range of binding energies for MBD proteins that is different from the first range of binding energies for MBD proteins. A second molecular barcode can be bound to nucleic acids of the second nucleic acid fraction. The second molecular barcode is associated with the second distribution. sample
[0228] Isolation and extraction of cell-free polynucleotides can be performed by collecting samples using various techniques. The sample can be any biological sample isolated from a subject. Samples can include body tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies (e.g., biopsies from known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid (e.g., fluid from the spaces between cells), gingival exudate, gingival crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Samples preferably include bodily fluids, particularly blood and its fractions, and urine. Such samples include nucleic acids excreted from tumors. Nucleic acids can include DNA and RNA, and can be in double-stranded or single-stranded forms. Sample can be the form that is originally separated from subject, or can be further processed to remove or add components, for example, cell, or enrich one component with another component, or convert one form of nucleic acid into another form of nucleic acid, for example, RNA into DNA, or single-stranded nucleic acid into double-stranded nucleic acid.Therefore, for example, the body fluid sample for analysis is the plasma or serum that contains cell-free nucleic acid, for example, cell-free DNA (cfDNA).
[0229] In some implementations, the sample volume of bodily fluid obtained from a subject depends on the desired read depth of the region to be sequenced. Example volumes are about 0.4-40 ml, about 5-20 ml, and about 10-20 ml. For example, the volume can be about 0.5 ml, about 1 ml, about 5 ml, about 10 ml, about 20 ml, about 30 ml, about 40 ml, or more milliliters. The volume of blood sampled can be between about 5 ml and about 20 ml.
[0230] Samples can contain varying amounts of nucleic acid. The amount of nucleic acid in a given sample can be equivalent to multiple genome equivalents. For example, a sample of about 30 ng of DNA can contain approximately 10,000 (10 4) haploid human genome equivalents, and in the case of cfDNA, approximately 200 billion (2 x 10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, or in the case of cfDNA, about 600 billion individual molecules.
[0231] In some implementations, the sample contains nucleic acids from different sources, for example, from cells and from cell-free sources (e.g., blood samples, etc.). Typically, the sample contains nucleic acids having mutations. For example, the sample optionally contains DNA having germline mutations and / or somatic mutations. Typically, the sample contains DNA having cancer-associated mutations (e.g., cancer-associated somatic mutations). In some implementations of the present disclosure, the cell-free nucleic acids in a subject may be derived from a tumor. For example, the cell-free DNA isolated from a subject may include ctDNA.
[0232] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), e.g., from about 1 picogram (pg) to about 200 nanograms (ng), from about 1 ng to about 100 ng, or from about 10 ng to about 1000 ng. In some implementations, the sample contains up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain implementations, the amount is up to about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng of cell-free nucleic acid molecules. In some implementations, the method includes obtaining from about 1 fg to about 200 ng of cell-free nucleic acid molecules from the sample.
[0233] Cell-free nucleic acids typically have a size distribution between about 100 and about 500 nucleotides in length, with molecules between about 110 and about 230 nucleotides in length representing about 90% of the molecules in a sample, a mode at about 168 nucleotides in length, and a second minor peak in the range between about 240 and about 440 nucleotides in length. In certain implementations, the cell-free nucleic acids are between about 160 and about 180 nucleotides in length, or between about 320 and about 360 nucleotides in length, or between about 440 and about 480 nucleotides in length.
[0234] In some implementations, cell-free nucleic acids are isolated from bodily fluids by a partitioning step, which separates the cell-free nucleic acids found in solution from intact cells and other insoluble components of the bodily fluid. In some of these implementations, the partitioning step includes techniques such as centrifugation or filtration. Alternatively, cells in the bodily fluid are lysed, and the cell-free and cellular nucleic acids are processed together. Generally, after the addition of a buffer and a washing step, the cell-free nucleic acids are precipitated, for example, using alcohol. In certain implementations, an additional washing step, such as a silica-based column, is used to remove contaminants or salts. For example, nonspecific bulk carrier nucleic acids are optionally added throughout the reaction to optimize certain aspects of the exemplary procedure, such as yield. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. If necessary, the single-stranded DNA and / or single-stranded RNA are converted to double-stranded form so that they can be included in subsequent processing and analysis steps. Further details regarding analysis of cfDNA distribution and associated epigenetic modifications, as optionally adapted for use in practicing the methods disclosed herein, are described, for example, in WO2018 / 119452, filed December 22, 2017, which is incorporated by reference. Nucleic Acid Tags
[0235] In certain implementations, tags providing molecular identifiers or barcodes are incorporated or otherwise attached to the adapters by chemical synthesis, ligation, or overlap extension PCR, among other methods. In some implementations, assignment of unique or non-unique identifiers or molecular barcodes in reactions utilizes systems according to methods described, for example, in U.S. Patent Application Nos. 20010053519, 20030152490, 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, and 9,598,731, each of which is incorporated by reference.
[0236] The tags are randomly or non-randomly linked (e.g., ligated) to the sample nucleic acid. In some implementations, the tags are introduced with a predicted ratio of identifiers (e.g., a combination of unique barcodes and / or non-unique barcodes) to the microwells. For example, the identifiers can be loaded so that about more than 1, more than 2, more than 3, more than 4, more than 5, more than 6, more than 7, more than 8, more than 9, more than 10, more than 20, more than 50, more than 100, more than 500, more than 1000, more than 5000, more than 10000, more than 50,000, more than 100,000, more than 500,000, more than 1,000,000, more than 10,000,000, more than 50,000,000, or more than 1,000,000,000 identifiers are loaded per genome sample. In some implementations, identifiers are loaded such that less than about 2, less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 20, less than 50, less than 100, less than 500, less than 1000, less than 5000, less than 10000, less than 50,000, less than 100,000, less than 500,000, less than 1,000,000, less than 10,000,000, less than 50,000,000, or less than 1,000,000,000 identifiers are loaded per genomic sample.In certain implementations, the average number of identifiers loaded per sample genome is less than about 1 or more than about 1, less than 2 or more, less than 3 or more, less than 4 or more, less than 5 or more, less than 6 or more, less than 7 or more, less than 8 or more, less than 9 or more, less than 10 or more, less than 20 or more, less than 50 or more, less than 100 or more, less than 500 or more, less than 1000, or more than 1,000, less than or more than 5,000, less than or more than 10,000, less than or more than 50,000, less than or more than 100,000, less than or more than 500,000, less than or more than 100,000, less than or more than 500,000, less than or more than 1,000,000, less than or more than 10,000,000, less than or more than 50,000,000, or less than or more than 1,000,000,000 identifiers. Identifiers are generally unique or non-unique.
[0237] One example format uses about 2 to about 1,000,000 different tags, or about 5 to about 150 different tags, or about 20 to about 50 different tags ligated to both ends of a target nucleic acid molecule. 20 to 50 x 20 to 50 tags are generated, for a total of 400 to 2500 tags. Such a number of tags is typically sufficient to ensure that different molecules with the same start and end points receive different combinations of tags with a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%).
[0238] In some implementations, the identifier is an oligonucleotide of predetermined, random, or semi-random sequence. In other implementations, multiple barcodes can be used, where the barcodes are not necessarily unique to each other. In these implementations, the barcodes are typically attached to individual molecules (e.g., by ligation or PCR amplification) so that the combination of the barcode and the sequence to which it can be attached creates a unique sequence that can be individually tracked. As described herein, detection of non-uniquely tagged barcodes in combination with sequence data from the beginning (start) and end (end) of the sequence read typically allows for the assignment of a unique identity to a particular molecule. The length or number of base pairs of each sequence read is also optionally used to assign a unique identity to a given molecule. As described herein, fragments from a single strand of nucleic acid that have been assigned a unique identity can thereby allow for the identification of subsequent fragments from the parental and / or complementary strands. Nucleic Acid Amplification
[0239] The adaptor-flanked sample nucleic acid is typically amplified by PCR and other amplification methods using nucleic acid primers that bind to primer binding sites in the adaptors adjacent to the DNA molecule to be amplified. In some implementations, the amplification method involves cycles of extension, denaturation, and annealing as a result of thermocycling, or may be isothermal, such as in transcription-mediated amplification. Examples of other amplification methods that may be used as needed include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustaining sequence-based replication, among other approaches.
[0240] To introduce a sample index / tag into a nucleic acid molecule using conventional nucleic acid amplification methods, one or more rounds of amplification cycles are generally applied. Amplification is typically performed in one or more reaction mixtures. In some implementations, molecular tags and sample index / tags are introduced before and / or after the sequence capture step. In some implementations, only molecular tags are introduced before probe capture, and sample index / tags are introduced after the sequence capture step. In certain implementations, both molecular tags and sample index / tags are introduced before the probe-based capture step. In some implementations, sample index / tags are introduced after the sequence capture step (i.e., nucleic acid enrichment). Typically, sequence capture protocols involve introducing a targeted nucleic acid sequence, for example, a single-stranded nucleic acid molecule complementary to a coding sequence in a genomic region and a mutation in such a region associated with a cancer type. Typically, the amplification reaction generates multiple non-uniquely or uniquely tagged nucleic acid amplicons having molecular tags and sample indexes / tags ranging in size from about 200 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some implementations, the amplicons have a size of about 300 nt. In some implementations, the amplicons have a size of about 500 nt. Nucleic acid enrichment
[0241] In some implementations, sequences are enriched prior to sequencing of nucleic acids. Enrichment can be performed for specific target regions or nonspecifically ("target sequences") as needed. In some implementations, a differential tiling and capture scheme can be used to enrich targeted regions of interest using nucleic acid capture probes ("baits") selected for one or more bait set panels. Differential tiling and capture schemes generally use bait sets with different relative concentrations to differentially tile (e.g., at different "resolutions") across the genomic intervals associated with the baits, subject to a set of constraints (e.g., sequencer constraints such as sequencing load, availability of each bait, etc.) to capture targeted nucleic acids at a desired level for downstream sequencing. These targeted genomic intervals of interest optionally include natural or synthetic nucleotide sequences of nucleic acid constructs. In some implementations, biotin-labeled beads can be used with probes for one or more intervals of interest to capture target sequences, and optionally, the intervals can then be amplified to enrich the regions of interest.
[0242] Sequence capture typically involves the use of oligonucleotide probes that hybridize to target nucleic acid sequences. In certain implementations, a probe set strategy involves tiling probes across a desired interval. Such probes can be, for example, about 60 to about 120 nucleotides in length. Sets can have depths of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x, or greater. The effectiveness of sequence capture generally depends in part on the length of the target molecule sequence that is complementary (or nearly complementary) to the probe sequence. Nucleic acid sequencing
[0243] After extraction and isolation of cfDNA from the sample, the cfDNA can be sequenced in steps 103 and 104. Generally, the sample nucleic acid, optionally flanked by adapters, with or without prior amplification, is subjected to sequencing. Sequencing methods or commercially available formats optionally utilized include, for example, Sanger sequencing, high-throughput sequencing, bisulfite sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore-based sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing (NGS), single-molecule sequencing-by-synthesis (SMSS) (Helicos), massively parallel sequencing, clonal single molecule arrays (Solexa), shotgun sequencing, sequencing using Ion Torrent, Oxford Nanopore, Roche Genia, primer walking, PacBio, SOLiD, Ion Torrent, or nanopore platforms. Sequencing reactions can be performed in a variety of sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means of processing multiple sample sets substantially simultaneously. Sample processing units may also include multiple sample chambers, allowing for the processing of multiple runs simultaneously.
[0244] The sequencing reaction can be performed on one or more nucleic acid fragment types or intervals known to contain markers for cancer or other diseases. The sequencing reaction can also be performed on any nucleic acid fragment present in the sample. The sequence reaction can provide sequence coverage of at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome. In other cases, sequence coverage of the genome may be less than about 5%, less than 10%, less than 15%, less than 20%, less than 25%, less than 30%, less than 40%, less than 50%, less than 60%, less than 70%, less than 80%, less than 90%, less than 95%, less than 99%, less than 99.9%, or less than 100% of the genome.
[0245] Multiplex sequencing technology can be used to carry out simultaneous sequencing reactions.In some implementations, cell-free polynucleotide is sequenced at least about 1000 times, 2000 times, 3000 times, 4000 times, 5000 times, 6000 times, 7000 times, 8000 times, 9000 times, 10000 times, 50000 times or 100,000 times sequencing reactions.In other implementations, cell-free polynucleotide is sequenced at less than about 1000 times, less than 2000 times, less than 3000 times, less than 4000 times, less than 5000 times, less than 6000 times, less than 7000 times, less than 8000 times, less than 9000 times, less than 10000 times, less than 50000 times or less than 100,000 times sequencing reactions. Sequencing reactions are typically carried out sequentially or simultaneously.Subsequent data analysis is generally carried out on all or part of sequencing reactions.In some implementations, data analysis is carried out on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions. In other implementations, data analysis can be performed on less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. Exemplary read depths are about 1000 to about 50,000 reads per locus (base position).
[0246] In some implementations, a nucleic acid population is prepared for sequencing by enzymatically creating blunt ends on double-stranded nucleic acids having single-stranded overhangs at one or both ends. In these implementations, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Exemplary enzymes, or catalytic fragments thereof as needed, include Klenow large fragment and T4 polymerase. In the 5' overhang, the enzyme typically extends the recessed 3' end on the opposite strand until it overlaps the 5' end, creating a blunt end. In the 3' overhang, the enzyme generally digests from the 3' end to the 5' end of the opposite strand, sometimes beyond the 5' end. If this digestion proceeds beyond the 5' end of the opposite strand, the gap can be filled with the same enzyme with the same polymerase activity used for the 5' overhang. The formation of blunt ends in double-stranded nucleic acids facilitates, for example, adapter binding and subsequent amplification.
[0247] In some implementations, the nucleic acid population is subjected to additional processing, such as converting single-stranded nucleic acids to double-stranded nucleic acids and / or converting RNA to DNA. These forms of nucleic acids can also be ligated with adapters and amplified, if necessary.
[0248] Regardless of whether or not there is previous amplification, the nucleic acid that is subjected to the above-mentioned blunt-end forming process, and optionally other nucleic acids in the sample, can be sequenced to produce sequenced nucleic acids.Sequenced nucleic acids can refer to either the sequence of nucleic acid (i.e., sequence information) or the nucleic acid whose sequence has been determined.Sequencing can be performed to obtain the sequence data of each nucleic acid molecule in the sample directly or indirectly from the consensus sequence of the amplification products of each nucleic acid molecule in the sample.
[0249] In some implementations, double-stranded nucleic acids with single-stranded overhangs in the sample after blunt-end formation are ligated at both ends to adapters containing barcodes, and the nucleic acid sequence and the in-line barcodes introduced by the adapters are determined by sequencing. Blunt-ended DNA molecules are optionally ligated to the blunt ends of at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters). Alternatively, the blunt ends of the sample nucleic acid and adapter may have complementary nucleotides at their tails to facilitate ligation (e.g., cohesive end ligation).
[0250] Typically, a nucleic acid sample is contacted with a sufficient number of adapters so that the probability that either of two copies of the same nucleic acid will have the same combination of adapter barcodes derived from adapters ligated to both ends is low (e.g., less than 1% or less than 0.1%). This use of adapters allows the identification of families of nucleic acid sequences that have the same start and end points on the reference nucleic acid and are ligated to the same combination of barcodes. Such families represent the sequences of the amplification products of the template / parent nucleic acid in the sample before amplification. The sequences of family members can be compiled to obtain consensus nucleotide(s) or a complete consensus sequence for the nucleic acid molecules in the original sample that have been modified by blunt-end formation and adapter ligation. In other words, a nucleotide occupying a particular position in a nucleic acid in the sample is determined to be the consensus of the nucleotides occupying the corresponding positions in the family member sequences. A family can include sequences from one or both strands of a double-stranded nucleic acid. When a family member comprises the sequences of both strands from double-stranded nucleic acid, the sequence of one strand is converted to its complement, so as to compile all sequences and obtain consensus nucleotide(s) or sequence.Some families only comprise a single member sequence.In this case, this sequence can be considered as the sequence of the nucleic acid in the sample before amplification.Alternatively, families that only have a single member sequence can be excluded from subsequent analysis.
[0251] Nucleotide variations in sequenced nucleic acids can be determined by comparing the sequenced nucleic acids with a reference sequence.The reference sequence is often a known sequence, for example, a known whole or partial genome sequence from a subject (for example, the whole genome sequence of a human subject).The reference sequence can be, for example, hG19 or hG38.The sequenced nucleic acid can represent the sequence directly determined for the nucleic acid in the sample, or the consensus of the sequence of the amplification product of such nucleic acid as described above.Comparison can be performed with respect to one or more designated positions on the reference sequence.A subset of sequenced nucleic acids can be identified that includes a position corresponding to the designated position of the reference sequence when each sequence is maximally aligned. Within such a subset, it can be determined that the sequenced nucleic acids, if any, contain a nucleotide variation at a specified position, the length of the given cfDNA fragment based on where its endpoints (i.e., the 5'-end and 3'-end nucleotides) map to the reference sequence, the offset of the midpoint of the given cfDNA fragment from the midpoint of the genomic region in the cfDNA fragment, and, if necessary, that the sequenced nucleic acids contain a reference nucleotide (i.e., the same as in the reference sequence). If the number of sequenced nucleic acids containing a nucleotide variant within the subset exceeds a selected threshold, a variant nucleotide can be called at the specified position. The threshold can be, among other possibilities, a simple number within the subset, for example, at least 1, 2, 3, 4, 5, 6, 7, 9, or 10 sequenced nucleic acids containing a nucleotide variant, or a ratio, for example, at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20 of the sequenced nucleic acids within the subset that contain a nucleotide variant. The comparison can be repeated for any specified position of interest within the reference sequence. Sometimes, comparison can be performed for designated positions occupying at least about 20, 100, 200, or 300 contiguous positions on the reference sequence, e.g., about 20-500, or about 50-300 contiguous positions.
[0252] Further details regarding nucleic acid sequencing, including the formats and applications described herein, can also be found in, for example, Levy et al., Annual Review of Genomics and Human Genetics, 17: 95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364: 1-11 (2012); Voelkerding et al., Clinical Chem., 55: 641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009); Astier et al., J Am Chem Soc., 128(5): 1705-1000 (2009), each of which is incorporated by reference in its entirety. (2006), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Patent No. 7,501,245, U.S. Patent No. 6,818,395, U.S. Patent No. 6,911,345, U.S. Patent No. 7,501,245, U.S. Patent No. 7,329,492, U.S. Patent No. 7,170,050, U.S. Patent No. 7,302,146, U.S. Patent No. 7,313,308, and U.S. Patent No. 7,476,503. Sequencing Panel
[0253] To improve the likelihood of detecting the genomic region of interest, and optionally, mutations indicative of tumors, the DNA section to be sequenced can include a panel of genomic sections containing genes or known genomic regions. Selecting a limited section for sequencing (e.g., a limited panel) can reduce the total amount of sequencing required (e.g., the total amount of nucleotides sequenced). A sequencing panel can target multiple different genes or regions, for example, to detect a single cancer, a set of cancers, or all cancers. Alternatively, DNA can be sequenced by whole genome sequencing (WGS) or other unbiased sequencing methods without using a sequencing panel. Examples of suitable panels and targets for use in panels can be found in the epigenetic targets described in U.S. Provisional Patent Application No. 62 / 799,637, filed January 31, 2019, the entire contents of which are incorporated by reference.
[0254] In some embodiments, a panel targeting multiple different genes or genomic regions (e.g., transcription factor binding regions, distal regulatory elements (DREs), repetitive elements, intron-exon junctions, transcription start sites (TSSs), etc.) is selected such that a determined percentage of subjects with cancer exhibit genetic variants or tumor markers in one or more different genes in the panel. The panel may be selected such that the sequencing region is limited to a fixed number of base pairs. The panel may be selected such that a desired amount of DNA is sequenced. The panel can further be selected to achieve a desired sequence read depth. The panel may be selected such that a desired sequence read depth or sequence read coverage is achieved for a certain amount of sequenced base pairs. The panel may be selected such that a theoretical sensitivity, specificity, and / or accuracy is achieved for detecting one or more genetic variants in a sample.
[0255] Probes for detecting a panel of regions can include probes for detecting genomic regions of interest (hotspot regions) as well as nucleosome-aware probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations affected by nucleosome binding patterns and GC sequence composition. As used herein, regions can also include non-hotspot regions optimized based on nucleosome position and GC model. Panels may include multiple subpanels, including subpanels for identifying tissue of origin (e.g., using published literature to define 50-100 baits representing genes (not necessarily promoters) with the most diverse transcriptional profiles across tissues), whole-genome scaffolding (e.g., to tiling sparsely across chromosomes using only a small number of probes to identify ultraconserved genomic content and establish a copy number baseline), and transcription start site (TSS) / CpG island (e.g., to capture differentially methylated regions (DMRs) in promoters of tumor suppressor genes (e.g., SEPT9 / VIM in colorectal cancer)). In some implementations, the markers for tissue of origin are tissue-specific epigenetic markers.
[0256] Some example lists of genomic locations of interest can be found in Tables 1 and 2. In some implementations, the genomic locations used in the methods of the present disclosure include at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or 97 of the genes in Table 1. In some implementations, the genomic locations used in the methods of the present disclosure include at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 1. In some implementations, the genomic locations used in the methods of the present disclosure include at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the CNVs in Table 1. In some implementations, the genomic locations used in the methods of the present disclosure include at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 1. In some implementations, the genomic locations used in the methods of the present disclosure include at least a portion of at least one, at least two, or three of the indels in Table 1.In some implementations, the genomic locations used in the methods of the present disclosure include at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, or 115 of the genes in Table 2. In some implementations, the genomic locations used in the methods of the present disclosure include at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 2. In some implementations, the genomic locations used in the methods of the present disclosure include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the CNVs in Table 2. In some implementations, the genomic locations used in the methods of the present disclosure include at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 2. In some implementations, the genomic locations used in the methods of the present disclosure comprise at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 2. Each of these genomic locations of interest can be identified as a scaffold region or a hotspot region for a given bait set panel.In one or more examples, the methods of the present disclosure can be implemented using all of the mutations included in Table 1 and / or Table 2. [Table 1] Table 2 [Table 2-1] [Table 2-2]
[0257] In some implementations, one or more regions in the panel include one or more loci from one or more genes for detecting residual cancer after surgery. This detection can occur earlier than is possible with existing cancer detection methods. In some implementations, one or more genomic locations in the panel include one or more loci from one or more genes for detecting cancer in high-risk patient populations. For example, smokers have a much higher rate of lung cancer than the general population. In addition, smokers may also develop other lung conditions that make cancer detection more difficult, such as the development of irregular nodules in the lungs. In some implementations, the methods described herein detect a patient's response to cancer treatment (particularly in high-risk patients) earlier than is possible with existing cancer detection methods.
[0258] The location of genome to be included in sequencing panel can be selected based on the number of subjects with cancer that have tumor marker in that gene or region.The location of genome to be included in sequencing panel can be selected based on the prevalence of subjects with cancer that have tumor marker in that gene.The presence of tumor marker in a region can indicate that the subject has cancer.
[0259] In some cases, the panel can be selected using information from one or more databases. Information about cancer can be derived from cancer tumor biopsies or cfDNA assays. The database can include information describing the population of tumor samples to be sequenced. The database can include information about mRNA expression in tumor samples. The database can include information about regulatory elements or genomic regions in tumor samples. The information about the tumor samples to be sequenced can include the frequency of various genetic variants, describing the genes or regions where the genetic variants are present. The genetic variants can be tumor markers. A non-limiting example of such a database is COSMIC. COSMIC is a list of somatic mutations found in various cancers. In COSMIC, genes are ranked based on the frequency of mutations for specific cancers. Genes can be selected for inclusion in a panel based on the frequency of mutations in a given gene. For example, COSMIC shows that 33% of the population of breast cancer samples to be sequenced have mutations in TP53, and 22% of the population of breast cancer samples sampled have mutations in KRAS. Other ranked genes, including APC, are mutated in only approximately 4% of breast cancer samples sequenced. TP53 and KRAS can be included in a sequencing panel based on their relatively high frequency in sampled breast cancers (e.g., compared to APC, which is present at a frequency of approximately 4%). While COSMIC is presented as a non-limiting example, any database or set of information linking cancer to tumor markers located in genes or genetic regions can be used. In another example, as presented by COSMIC, of 1,156 cholangiocarcinoma samples, 380 samples (33%) had mutations in TP53. Some other genes, such as APC, have mutations in 4-8% of all samples. Therefore, TP53 can be selected for inclusion in a panel based on its relatively high frequency in a cholangiocarcinoma sample population.
[0260] A gene or genomic interval can be selected for a panel if the frequency of the tumor marker in sampled tumor tissue or circulating tumor DNA is significantly higher than that found in a given background population. A combination of genomic locations can be selected for inclusion in a panel so that at least a majority of subjects with cancer can have a tumor marker or genomic region present in at least one of the genomic locations or genes in the panel. A combination of genomic locations can be selected based on data showing that for a particular cancer or set of cancers, a majority of subjects have one or more tumor markers in one or more of the selected regions. For example, to detect cancer 1, a panel including regions A, B, C, and / or D can be selected based on data showing that 90% of subjects with cancer 1 have tumor markers in regions A, B, C, and / or D of the panel. Alternatively, a tumor marker may be shown to be present independently in two or more regions in subjects with cancer, so that, when combined, a majority of the population of subjects with cancer has tumor markers in two or more regions. For example, to detect cancer 2, a panel including regions X, Y, and Z can be selected based on data showing that 90% of subjects have tumor markers in one or more regions, and that in 30% of such subjects, the tumor marker is detected only in region X, while in the remaining subjects in which the tumor marker is detected, the tumor marker is detected only in regions Y and / or Z. A tumor marker present in one or more genomic locations previously shown to be associated with one or more cancers can indicate or predict that a subject has cancer if the tumor marker is detected in 50% or more of those regions. Computational approaches, such as models using conditional probabilities of cancer detection that consider the frequency of cancer for a set of tumor markers in one or more regions, can be used to predict which regions, alone or in combination, may be predictive of cancer.Other approaches to panel selection involve the use of databases containing information from studies using large panels and / or comprehensive genomic profiling of tumors using whole genome sequencing (WGS, RNA-seq, Chip-seq, ATAC-seq, and others). Information gleaned from the literature may also describe pathways commonly affected and mutated in certain cancers. Panel selection can be further informed by the use of genetic ontologies.
[0261] The gene included in the sequencing panel can include the complete transcription region, promoter region, enhancer region, regulatory element, and / or downstream sequence.To further increase the likelihood of detecting mutations that indicate tumors, only exons can be included in the panel.The panel can include all exons of the selected gene, or can include only one or more exons of the selected gene.The panel can include exons from each of multiple different genes.The panel can include at least one exon from each of multiple different genes.
[0262] In some embodiments, a panel of exons from each of a plurality of different genes is selected such that a determined proportion of subjects with cancer exhibit a genetic variant in at least one exon within the panel of exons.
[0263] At least one complete exon from each different gene in the panel of genes can be sequenced. The panel to be sequenced can consist of exons from multiple genes. The panel can consist of exons from 2-100 different genes, 2-70 genes, 2-50 genes, 2-30 genes, 2-15 genes, or 2-10 genes.
[0264] The selected panel can consist of various numbers of exons. The panel can consist of 2-3000 exons. The panel can consist of 2-1000 exons. The panel can consist of 2-500 exons. The panel can consist of 2-100 exons. The panel can consist of 2-50 exons. The panel can consist of 300 or fewer exons. The panel can consist of 200 or fewer exons. The panel can consist of 100 or fewer exons. The panel can consist of 50 or fewer exons. The panel can consist of 40 or fewer exons. The panel can consist of 30 or fewer exons. The panel can consist of 25 or fewer exons. The panel can consist of 20 or fewer exons. The panel can consist of 15 or fewer exons. The panel can consist of 10 or fewer exons. The panel can consist of 9 or fewer exons. The panel may consist of no more than 8 exons. The panel may consist of no more than 7 exons.
[0265] A panel may consist of one or more exons from a plurality of different genes. A panel may consist of one or more exons from each of a proportion of a plurality of different genes. A panel may consist of at least two exons from each of at least 25%, 50%, 75%, or 90% of the different genes. A panel may consist of at least three exons from each of at least 25%, 50%, 75%, or 90% of the different genes. A panel may consist of at least four exons from each of at least 25%, 50%, 75%, or 90% of the different genes.
[0266] The size of a sequencing panel can vary. A sequencing panel can be larger or smaller (in terms of nucleotide size) depending on several factors, including, for example, the total amount of nucleotides to be sequenced or the number of unique molecules to be sequenced for a particular region within the panel. A sequencing panel can be 5 kb to 50 kb in size. A sequencing panel can be 10 kb to 30 kb in size. A sequencing panel can be 12 kb to 20 kb in size. A sequencing panel can be 12 kb to 60 kb in size. A sequencing panel can be at least 10 kb, 12 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 45 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, or 150 kb in size. Sequencing panels can be less than 100 kb, less than 90 kb, less than 80 kb, less than 70 kb, less than 60 kb, or less than 50 kb in size.
[0267] The panel selected for sequencing can consist of at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 genomic locations (e.g., each containing a genomic region of interest).In some cases, the genomic locations in the panel are selected so that the location size is relatively small.In some cases, the region in the panel has a size of about 10 kb or less, about 8 kb or less, about 6 kb or less, about 5 kb or less, about 4 kb or less, about 3 kb or less, about 2.5 kb or less, about 2 kb or less, about 1.5 kb or less, or about 1 kb or less. In some cases, genomic locations within a panel have a size of about 0.5 kb to about 10 kb, about 0.5 kb to about 6 kb, about 1 kb to about 11 kb, about 1 kb to about 15 kb, about 1 kb to about 20 kb, about 0.1 kb to about 10 kb, or about 0.2 kb to about 1 kb. For example, a region within a panel can have a size of about 0.1 kb to about 5 kb.
[0268] The panel selected herein can enable deep sequencing sufficient to detect low-frequency genetic variants (for example, in cell-free nucleic acid molecules obtained from a sample). A certain amount of genetic variants in a sample can be represented in terms of the variant allele fraction for a given genetic variant. Variant allele fraction can refer to the frequency at which a variant allele exists in a given group of nucleic acids, for example, a sample. A genetic variant with a low variant allele fraction can have a relatively low frequency of existence in a sample. In some cases, the panel allows the detection of genetic variants with a variant allele fraction of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, or 0.5%. The panel can allow the detection of genetic variants with a variant allele fraction of 0.001% or higher. The panel can allow the detection of genetic variants with a variant allele fraction of 0.01% or higher. A panel can allow for the detection of genetic variants present in a sample at frequencies as low as 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. A panel can allow for the detection of tumor markers present in a sample at frequencies of at least 0.0001%, 0.001%, 0.005%, 0.01%, 0.025%, 0.05%, 0.075%, 0.1%, 0.25%, 0.5%, 0.75%, or 1.0%. A panel can allow for the detection of tumor markers present in a sample at frequencies as low as 1.0%. A panel can allow for the detection of tumor markers present in a sample at frequencies as low as 0.75%. The panel may allow for the detection of tumor markers at frequencies as low as 0.5% in a sample. The panel may allow for the detection of tumor markers at frequencies as low as 0.25% in a sample. The panel may allow for the detection of tumor markers at frequencies as low as 0.1% in a sample. The panel may allow for the detection of tumor markers at frequencies as low as 0.075% in a sample. The panel may allow for the detection of tumor markers at frequencies as low as 0.05% in a sample.The panel may enable detection of tumor markers at frequencies as low as 0.025% in a sample. The panel may enable detection of tumor markers at frequencies as low as 0.01% in a sample. The panel may enable detection of tumor markers at frequencies as low as 0.005% in a sample. The panel may enable detection of tumor markers at frequencies as low as 0.001% in a sample. The panel may enable detection of tumor markers at frequencies as low as 0.0001% in a sample. The panel may enable detection of tumor markers in sequenced cfDNA at frequencies as low as 1.0% to 0.0001% in a sample. The panel may enable detection of tumor markers in sequenced cfDNA at frequencies as low as 0.01% to 0.0001% in a sample.
[0269] Genetic variants can be represented by the percentage of the population of subjects with disease (for example, cancer).In some cases, at least 1%, 2%, 3%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95% or 99% of the population with cancer show one or more genetic variants in at least one of the regions in panel.For example, at least 80% of the population with cancer can show one or more genetic variants in at least one of the genome positions in panel.
[0270] A panel can be comprised of one or more locations that comprise the genomic region of interest from one or more genes.In some cases, a panel can be comprised of one or more locations that comprise the genomic region of interest from at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 80 genes.In some cases, a panel can be comprised of one or more locations that comprise the genomic region of interest from at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 80 genes. In some cases, the panel may consist of one or more loci containing genomic regions of interest from each of about 1 to about 80, 1 to about 50, about 3 to about 40, 5 to about 30, or 10 to about 20 different genes.
[0271] The location of the genome region in the panel can be selected to detect one or more epigenetically modified regions.One or more epigenetically modified regions can be acetylated, methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.For example, the region in the panel can be selected to detect one or more methylated regions.
[0272] The region in panel can be selected to include the sequence that is differentially transcribed across one or more tissues.In some cases, the location that includes genomic region can include the sequence that is transcribed at a high level in a certain tissue compared with other tissues.For example, the location that includes genomic region can include the sequence that is transcribed in a certain tissue but not in other tissues.
[0273] The genome locations within the panel may include coding and / or non-coding sequences. For example, the genome locations within the panel may include one or more sequences within exons, introns, promoters, 3' untranslated regions, 5' untranslated regions, regulatory elements, transcription start sites, and / or splice sites. In some cases, the regions within the panel may include other non-coding sequences, including pseudogenes, repeat sequences, transposons, viral elements, and telomeres. In some cases, the genome locations within the panel may include sequences within non-coding RNAs, such as ribosomal RNAs, transfer RNAs, Piwi-interacting RNAs, and microRNAs.
[0274] Genomic locations within the panel can be selected such that cancer is detected (diagnosed) with a desired level of sensitivity (e.g., by detection of one or more genetic variants). For example, regions within the panel can be selected such that cancer is detected with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% sensitivity (e.g., by detection of one or more genetic variants). Genomic locations within the panel can be selected such that cancer is detected with 100% sensitivity.
[0275] The genomic locations within the panel can be selected such that cancer is detected (diagnosed) with a desired level of specificity (e.g., by detection of one or more genetic variants). For example, the genomic locations within the panel can be selected such that cancer is detected (e.g., by detection of one or more genetic variants) with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% specificity. The genomic locations within the panel can be selected such that one or more genetic variants are detected with 100% specificity.
[0276] Genomic locations within the panel can be selected to detect (diagnose) cancer with a desired positive predictive value. Positive predictive value can be increased by increasing sensitivity (e.g., the probability that an actual positive will be detected) and / or specificity (e.g., the probability that an actual negative will not be falsely identified as a positive). As a non-limiting example, genomic locations within the panel can be selected to detect one or more genetic variants with a positive predictive value of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. Regions within the panel can be selected to detect one or more genetic variants with a positive predictive value of 100%.
[0277] The genomic locations within the panel can be selected so that cancer is detected (diagnosed) with a desired accuracy. As used herein, the term "accuracy" can refer to the ability of a test to distinguish between a disease state (e.g., cancer) and a healthy state. Accuracy can be quantified using measures such as sensitivity and specificity, predictive value, likelihood ratio, area under the ROC curve, Youden index and / or diagnostic odds ratio.
[0278] Accuracy can be presented as a percentage, referring to the ratio of the number of tests that produced correct results to the total number of tests performed. Regions within the panel can be selected to detect cancer with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. Genomic locations within the panel can be selected to detect cancer with 100% accuracy.
[0279] The panel can be selected to be highly sensitive and detect low frequency genetic variants.For example, the panel can be selected to detect genetic variants or tumor markers that exist in samples at a frequency of 0.01%, 0.05% or 0.001% with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9% sensitivity.The genome location within the panel can be selected to detect tumor markers that exist in samples at a frequency of 1% or less with 70% or higher sensitivity. The panel can be selected to detect tumor markers at frequencies as low as 0.1% in samples with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers at frequencies as low as 0.01% in samples with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers with frequencies as low as 0.001% in a sample with a sensitivity of at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0280] Panel can be selected to have high specificity and detect low frequency genetic variants.For example, panel can be selected to detect genetic variants or tumor markers that exist in samples with a frequency as low as 0.01%, 0.05% or 0.001% with at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5% or 99.9% specificity.Genomic location in panel can be selected to detect tumor markers that exist in samples with a frequency of 1% or less with 70% or higher specificity. The panel can be selected to detect tumor markers with frequencies as low as 0.1% in a sample with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers at frequencies as low as 0.01% in samples with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%. The panel can be selected to detect tumor markers at frequencies as low as 0.001% in samples with a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0281] The panel can be selected to have high accuracy and detect low-frequency genetic variants.The panel can be selected to detect genetic variants or tumor markers that are present in samples at frequencies as low as 0.01%, 0.05%, or 0.001% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy.The genomic locations within the panel can be selected to detect tumor markers that are present in samples at frequencies of 1% or less with 70% or higher accuracy.The panel can be selected to detect tumor markers that are present in samples at frequencies as low as 0.1% with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The panel can be selected to detect tumor markers at frequencies as low as 0.01% in samples with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy. The panel can be selected to detect tumor markers at frequencies as low as 0.001% in samples with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% accuracy.
[0282] The panel can be selected to be highly predictive and detect low frequency genetic variants. The panel can be selected so that a genetic variant or tumor marker present in a sample at a frequency as low as 0.01%, 0.05%, or 0.001% can have a positive predictive value of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%.
[0283] The concentration of the probe or bait used in the panel can be increased (2-6 ng / μL) to capture more nucleic acid molecules in the sample. The concentration of the probe or bait used in the panel can be at least 2 ng / μL, 3 ng / μL, 4 ng / μL, 5 ng / μL, 6 ng / μL, or higher. The probe concentration can be about 2 ng / μL to about 3 ng / μL, about 2 ng / μL to about 4 ng / μL, about 2 ng / μL to about 5 ng / μL, or about 2 ng / μL to about 6 ng / μL. The concentration of the probe or bait used in the panel can be 2 ng / μL or higher to 6 ng / μL or lower. In some cases, this can enable the analysis of more molecules in a biological sample, thus enabling the detection of less frequent alleles.
[0284] In some implementations, after sequencing, sequence reads can be assigned quality scores. The quality scores can be a representation of the sequence reads that indicates whether the sequence reads may be useful in subsequent analysis based on a threshold. In some cases, some sequence reads are not of sufficient quality or length to perform the subsequent mapping step. Sequence reads with a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out from the dataset of sequence reads. In other cases, sequence reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out from the dataset. Sequence reads that meet a certain quality score threshold can be mapped to a reference genome. After mapping alignment, sequence reads can be assigned a mapping score. The mapping score can be a representation of the sequence reads mapped back to the reference sequence, indicating whether each position is uniquely mappable. Sequence reads with a mapping score of at least 90%, 95%, 99%, 99.9%, 99.99%, or 99.999% can be filtered out from the dataset. In other cases, sequencing reads assigned a mapping score of less than 90%, less than 95%, less than 99%, less than 99.9%, less than 99.99%, or less than 99.999% can be filtered out from the dataset. Accurate treatment
[0285] The accurate diagnosis provided by the improved computer system 110 can result in an accurate treatment plan, which can be identified by (and / or curated by) the computer system 110. For example, one type of accurate diagnosis and treatment can relate to genes in the homologous recombination repair (HRR) pathway.
[0286] Homologous recombination is a type of genetic recombination in which nucleotide sequences are exchanged between two similar or identical DNA molecules. Homologous recombination is most widely used by cells to precisely repair harmful breaks that occur in both strands of DNA, known as double-strand breaks (DSBs). HRR provides a mechanism for error-free removal of damage present in replicated DNA (S and G2 phases) to eliminate chromosomal breaks before cell division occurs. The leading model for how DNA double-strand breaks are repaired by homologous recombination is the homology-directed repair pathway, which mediates the double-strand break repair (DSBR) pathway and the synthesis-dependent strand annealing (SDSA) pathway. Germline and somatic defects in homologous recombination genes have been strongly associated with breast, ovarian, and prostate cancer.
[0287] The number and type of variant nucleotides in sample can provide the indication of the suitability of the subject who provides sample for treatment, i.e., therapeutic intervention.For example, various poly ADP-ribose polymerase (PARP) inhibitors have been shown to stop the growth of tumors that originate from breast cancer, ovarian cancer and prostate cancer caused by inherited mutations in BRCA1 or BRCA2 gene.Some of these therapeutic agents can inhibit base excision repair (BER), which can compensate for the defect of HRR.
[0288] On the other hand, certain BRCA and HRR wild-type patients may not realize clinical benefits from treatment with PARP inhibitors. Moreover, not all ovarian cancer patients with BRCA mutations respond to PARP inhibitors. Furthermore, different types of mutations may indicate different treatments. For example, somatic heterozygous deletion in the HRR gene may indicate different treatments than somatic homozygous deletion. Therefore, the state of genetic material may affect treatment. In one example, PARP inhibitors can be administered to individuals with somatic homozygous deletion in the HRR gene, but cannot be administered to individuals with wild-type alleles or somatic heterozygous deletion in the HRR gene.
[0289] In some implementations, a targeted therapy can be administered to a subject with HRD determined by any of the disclosed methods. The targeted therapy can include a PARP inhibitor. Examples of PARP inhibitors that can be administered include one or more of veliparib, olaparib, talazoparib, rucaparib, niraparib, pamiparib, CEP 9722 (Cephalon), E7016 (Eisai), E7449 (Eisai, a PARP 1 / 2 and tankyrase 1 / 2 inhibitor), or 3-aminobenzamide. In some implementations, the targeted therapy can include at least one base excision repair (BER) inhibitor. For example, olaparib can inhibit BER. In certain implementations, the targeted therapy can include a combination of a PARP inhibitor and radiation therapy. In some implementations, the combination of a PARP inhibitor with radiation therapy may allow the PARP inhibitor to convert single-strand breaks generated by radiation therapy into double-strand breaks in tumor tissue (e.g., tissue with BRCA1 / BRCA2 mutations), resulting in more potent treatment per radiation dose. Customized treatment and related administration
[0290] In some implementations, the methods disclosed herein relate to identifying and administering a treatment to a patient with a given disease, disorder, or condition. Essentially any cancer treatment (e.g., surgery, radiation therapy, chemotherapy, etc.) can be included as part of these methods. This includes, for example, the diseases, disorders, or conditions found in Table 3. Table 3. Diseases and Treatments [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4] Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 3-10 Table 3-11 Table 3-12 Table 3-13
[0291] Patients with germline or somatic HRD may be candidates for targeted therapies, including DNA damage response (DDR) inhibitors, such as poly(ADP-ribose) polymerase (PARP) inhibitors (PARPi) [Fong et al., "Poly(ADP)-ribose polymerase inhibition: frequent durable responses in BRCA carrier ovarian cancer correlating with platinum-free interval," J Clin Oncol, 28:2512-9 (2010); Audeh et al., "Oral poly(ADP-ribose) polymerase inhibitor olaparib in patients with BRCA1 or BRCA2 mutations and recurrent ovarian cancer: a proof-of-concept trial," Lancet, 376:245-51 (2010)].
[0292] In some embodiments, a targeted therapy can be administered to a subject with HRD determined by any of the disclosed methods. The targeted therapy can include a PARP inhibitor. Examples of PARP inhibitors that can be administered include one or more of veliparib, olaparib, talazoparib, rucaparib, niraparib, pamiparib, CEP 9722 (Cephalon), E7016 (Eisai), E7449 (Eisai, a PARP 1 / 2 and tankyrase 1 / 2 inhibitor), or 3-aminobenzamide. In some embodiments, the targeted therapy can include at least one base excision repair (BER) inhibitor. For example, olaparib can inhibit BER. In certain embodiments, the targeted therapy can include a combination of a PARP inhibitor and radiation therapy. In some embodiments, the combination of a PARP inhibitor with radiation therapy can allow the PARP inhibitor to convert single-strand breaks caused by radiation therapy into double-strand breaks in tumor tissue (e.g., tissue with a BRCA1 / BRCA2 mutation), resulting in more potent treatment per radiation dose.
[0293] In some embodiments, the therapy is a PARP inhibitor, such as olaparib (Lynparza), rucaparib (Rubraca), niraparib (Zejula), and talazoparib (Talzenna), which can be used to treat mutations in BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, and RAD54L alterations and / or related homologous recombination repair (HRR) genes.
[0294] In certain implementations, the treatment administered to the subject may include at least one chemotherapeutic agent. In some implementations, the chemotherapeutic agent may include alkylating agents (e.g., but not limited to, chlorambucil, cyclophosphamide, cisplatin, and carboplatin), nitrosoureas (e.g., but not limited to, carmustine and lomustine), antimetabolites (e.g., but not limited to, fluorouracil, methotrexate, and fludarabine), plant alkaloids and natural products (e.g., but not limited to, vincristine, paclitaxel, and topotecan), antitumor antibiotics (e.g., but not limited to, bleomycin, doxorubicin, and mitoxantrone), hormonal agents (e.g., but not limited to, prednisone, dexamethasone, tamoxifen, and leuprolide), and biological response modifiers (e.g., but not limited to, Herceptin, Avastin, Erbitux, and Rituxan). In some implementations, the chemotherapy administered to the subject may include FOLFOX or FOLFIRI. Typically, the treatment includes at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to a method of enhancing the immune response to a given cancer type. In certain implementations, immunotherapy refers to a method of enhancing the T cell response to tumors or cancer.
[0295] In some implementations, immunotherapy or immunotherapeutic agents target immune checkpoint molecules. Certain tumors can evade the immune system by exploiting immune checkpoint pathways. Therefore, targeting immune checkpoints has emerged as an effective approach to block tumors' ability to evade the immune system and to activate antitumor immunity against certain cancers. Pardoll, Nature Reviews Cancer, 2012, 12:252-264.
[0296] In certain implementations, the immune checkpoint molecule is an inhibitory molecule that reduces signals involved in T cell responses to antigens. For example, CTLA4 is expressed on T cells and plays a role in downregulating T cell activation by binding to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells. PD-1 is another inhibitory checkpoint molecule expressed on T cells. PD-1 limits T cell activity in peripheral tissues during inflammatory responses. Furthermore, PD-1 ligands (PD-L1 or PD-L2) are commonly upregulated on the surface of many different tumors, resulting in downregulation of anti-tumor immune responses in the tumor microenvironment. In certain implementations, the inhibitory immune checkpoint molecule is CTLA4 or PD-1. In other implementations, the inhibitory immune checkpoint molecule is a PD-1 ligand, e.g., PD-L1 or PD-L2. In other implementations, the inhibitory immune checkpoint molecule is a CTLA4 ligand, e.g., CD80 or CD86. In other implementations, the inhibitory immune checkpoint molecule is lymphocyte activation gene 3 (LAG3), killer cell immunoglobulin-like receptor (KIR), T cell membrane protein 3 (TIM3), galectin 9 (GAL9), or adenosine A2a receptor (A2aR).
[0297] Antagonists that target these immune checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Thus, in certain implementations, the immunotherapy or immunotherapeutic agent is an antagonist of an inhibitory immune checkpoint molecule. In certain implementations, the inhibitory immune checkpoint molecule is PD-1. In certain implementations, the inhibitory immune checkpoint molecule is PD-L1. In certain implementations, the antagonist of an inhibitory immune checkpoint molecule is an antibody (e.g., a monoclonal antibody). In certain implementations, the antibody or monoclonal antibody is an anti-CTLA4, anti-PD-1, anti-PD-L1, or anti-PD-L2 antibody. In certain implementations, the antibody is a monoclonal anti-PD-1 antibody. In some implementations, the antibody is a monoclonal anti-PD-L1 antibody. In certain implementations, the monoclonal antibody is a combination of an anti-CTLA4 antibody and an anti-PD-1 antibody, an anti-CTLA4 antibody and an anti-PD-L1 antibody, or an anti-PD-L1 antibody and an anti-PD-1 antibody. In certain implementations, the anti-PD-1 antibody is one or more of pembrolizumab (Keytruda®) or nivolumab (Opdivo®). In certain implementations, the anti-CTLA4 antibody is ipilimumab (Yervoy®). In certain implementations, the anti-PD-L1 antibody is one or more of atezolizumab (Tecentriq®), avelumab (Bavencio®), or durvalumab (Imfinzi®).
[0298] In certain implementations, the immunotherapy or immunotherapeutic agent is an antagonist (e.g., an antibody) against CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In other implementations, the antagonist is a soluble form of an inhibitory immune checkpoint molecule, such as a soluble fusion protein comprising the extracellular domain of an inhibitory immune checkpoint molecule and the Fc domain of an antibody. In certain implementations, the soluble fusion protein comprises the extracellular domain of CTLA4, PD-1, PD-L1, or PD-L2. In some implementations, the soluble fusion protein comprises the extracellular domain of CD80, CD86, LAG3, KIR, TIM3, GAL9, or A2aR. In one implementation, the soluble fusion protein comprises the extracellular domain of PD-L2 or LAG3.
[0299] In certain implementations, the immune checkpoint molecule is a costimulatory molecule that amplifies signals involved in T cell responses to antigens. For example, CD28 is a costimulatory receptor expressed on T cells. When a T cell binds to an antigen through its T cell receptor, CD28 binds to CD80 (also known as B7.1) or CD86 (also known as B7.2) on antigen-presenting cells, amplifying T cell receptor signaling and promoting T cell activation. Because CD28 binds to the same ligands (CD80 and CD86) as CTLA4, CTLA4 can counteract or modulate costimulatory signaling mediated by CD28. In certain implementations, the immune checkpoint molecule is a costimulatory molecule selected from CD28, inducible T cell costimulator (ICOS), CD137, OX40, or CD27. In other implementations, the immune checkpoint molecule is a ligand for a costimulatory molecule, including, for example, CD80, CD86, B7RP1, B7-H3, B7-H4, CD137L, OX40L, or CD70.
[0300] Agonists targeting these costimulatory checkpoint molecules can be used to enhance antigen-specific T cell responses against certain cancers. Thus, in certain implementations, the immunotherapy or immunotherapeutic agent is an agonist of a costimulatory checkpoint molecule. In certain implementations, the agonist of a costimulatory checkpoint molecule is an agonist antibody, preferably a monoclonal antibody. In certain implementations, the agonist antibody or monoclonal antibody is an anti-CD28 antibody. In other implementations, the agonist antibody or monoclonal antibody is an anti-ICOS, anti-CD137, anti-OX40, or anti-CD27 antibody. In other implementations, the agonist antibody or monoclonal antibody is an anti-CD80, anti-CD86, anti-B7RP1, anti-B7-H3, anti-B7-H4, anti-CD137L, anti-OX40L, or anti-CD70 antibody.
[0301] Therapeutic options for treating particular genetically based diseases, disorders, or conditions other than cancer are generally well known to those of skill in the art and will be apparent given the particular disease, disorder, or condition being considered.
[0302] In certain implementations, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by any method known in the art, including, for example, buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intraauricular, and administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, etc.
[0303] FIG. 5 is a block diagram illustrating components of a machine 500 capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methodologies discussed herein, according to some example implementations. Specifically, FIG. 5 shows a schematic diagram of an example machine 500 in the form of a computer system that can execute instructions 502 (e.g., software, programs, applications, applets, apps, or other executable code) that cause the machine 500 to perform any one or more of the methodologies discussed herein. Thus, the instructions 502 can be used to implement modules or components described herein. The instructions 502 transform a general, unprogrammed machine 500 into a specific machine 500 that is programmed to perform the described and illustrated functions in the described manner. In alternative implementations, the machine 500 may operate as a stand-alone device or be coupled (e.g., networked) to other machines. In a networked deployment, machine 500 may operate as a server or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 500 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a mobile phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing, serially or otherwise, instructions 502 that specify operations to be performed by machine 500.Additionally, although only a single machine 500 is illustrated, the term "machine" shall also be considered to include a collection of machines that individually or jointly execute instructions 502 to perform any one or more of the methodologies discussed herein.
[0304] Machine 500 may include a processor 504, memory / storage 506, and I / O components 508, which may be configured to communicate with each other, for example, via a bus 510. In one example implementation, processor 504 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 512 and processor 514 capable of executing instructions 502. The term "processor" is intended to include a multi-core processor 504, which may consist of two or more independent processors (sometimes referred to as "cores") capable of simultaneously executing instructions 502. Although multiple processors 504 are shown in FIG. 5, the machine 500 may include a single processor 512 with a single core, a single processor 512 with multiple cores (e.g., a multi-core processor), multiple processors 512, 514 with a single core, multiple processors 512, 514 with multiple cores, or any combination thereof.
[0305] Memory / storage 506 may include memory, such as main memory 516 or other memory storage, and storage unit 518, both of which are accessible to processor 504, for example, via bus 510. Storage unit 518 and main memory 516 store instructions 502 that embody any one or more of the methodologies or functions described herein. During execution by machine 500, instructions 502 may reside, completely or partially, within main memory 516, within storage unit 518, within at least one of processors 504 (e.g., within a processor's cache memory), or any suitable combination thereof. Thus, main memory 516, storage unit 518, and memory of processor 504 are examples of machine-readable media.
[0306] The I / O components 508 may include a wide variety of components for receiving input, providing output, generating output, communicating information, exchanging information, capturing measurements, etc. The particular I / O components 508 included in a particular machine 500 will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It will be understood that the I / O components 508 may include many other components not shown in FIG. 5 . The I / O components 508 are grouped according to function solely to simplify the following discussion, and the grouping is in no way limiting. In various example implementations, the I / O components 508 may include a user output component 520 and a user input component 522. User output components 520 may include visual components (e.g., a display such as a plasma display panel (PDP), light-emitting diode (LED) display, liquid crystal display (LCD), projector, or cathode ray tube (CRT)), auditory components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generating devices, etc. User input components 522 may include alphanumeric input components (e.g., a keyboard, a touchscreen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input component), point-based input components (e.g., a mouse, touchpad, trackball, joystick, motion sensor, or other pointing device), tactile input components (e.g., physical buttons, a touchscreen that provides the location or force of a touch or touch gesture, or other tactile input component), audio input components (e.g., a microphone), etc.
[0307] In further example implementations, the I / O component 508 may include a biometrics component 524, a motion component 526, an environmental component 528, or a position component 530, among a wide range of other components. For example, the biometrics component 524 may include components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, gestures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or brain wave-based identification), etc. The motion component 526 may include an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environmental components 528 may include, for example, a light sensor component (e.g., a light meter), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor for detecting concentrations of harmful gases for safety purposes or for measuring pollutants in the air), or other components that may provide indicators, measurements, or signals corresponding to the surrounding physical environment. The location component 530 may include a location sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure from which altitude can be calculated), an orientation sensor component (e.g., a magnetometer), etc.
[0308] Communications can be implemented using a wide variety of technologies. The I / O component 508 may include a communications component 532 operable to couple the machine 500 to a network 534 or a device 536. For example, the communications component 532 may include a network interface component or other suitable device for interfacing with the network 534. In further examples, the communications component 532 may include a wired communications component, a wireless communications component, a cellular communications component, a near field communications (NFC) component, a Bluetooth® component (e.g., Bluetooth® Low Energy), a Wi-Fi® component, and other communications components providing communications via other modalities. The device 536 may be another machine 500 or any of a wide variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0309] Additionally, the communication component 532 may include a component that detects an identifier or is operable to detect an identifier. For example, the communication component 532 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multidimensional barcodes such as Quick Response (QR) codes, Aztec Code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). Additionally, various information can be obtained via the communication component 532, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi signal triangulation, and location via detection of NFC beacon signals that may indicate a specific location.
[0310] As used herein, a "component" refers to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that result in the division or modularization of specific processing or control functions. Components can be combined with other components through their interfaces to perform machine processing. A component may also be a packaged functional hardware unit designed for use with portions of a program that typically perform specific functions of other components and associated functions. A component may constitute either a software component (e.g., code embodied on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit that can perform certain operations and can be configured or arranged in a certain physical manner. In various examples of implementation, one or more computer systems (e.g., stand-alone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or group of processors) may be configured as hardware components that operate by software (e.g., an application or application portion) to perform certain operations described herein.
[0311] A hardware component may be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic permanently configured to perform certain operations. A hardware component may be a special-purpose processor, such as a field-programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor 504 or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of machine 500) uniquely tailored to perform the configured function, and is no longer a general-purpose processor 504. It will be appreciated that the decision to implement a hardware component mechanically in dedicated, permanently configured circuitry or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations. Thus, the phrase "hardware component" (or "hardware-implemented component") should be understood to encompass a tangible entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular manner or to perform certain operations described herein. Given an implementation in which the hardware components are temporarily configured (e.g., programmed), it is not necessary that each of the hardware components be configured or instantiated at any one time. For example, if the hardware components include a general-purpose processor 504 configured by software to be a special-purpose processor, the general-purpose processor 504 may be configured at different times as different special-purpose processors (e.g., including different hardware components).The software accordingly configures a particular processor 512, 514 or processor 504, for example, to configure a particular hardware component at one time and to configure a different hardware component at a different time.
[0312] Hardware components can provide information to and receive information from other hardware components. Thus, the described hardware components can be considered to be communicatively coupled. When multiple hardware components are present simultaneously, communication can be achieved by signal transmission (e.g., through appropriate circuits and buses) between two or more of the hardware components. In implementations in which multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storage and retrieval of information in memory structures accessible to the multiple hardware components. For example, an operation can be performed by one hardware component and the output of that operation can be stored in a communicatively coupled memory device. Then, at a later point in time, an additional hardware component accesses the memory device to retrieve and process the stored output.
[0313] Hardware components may also initiate communication with input or output devices and operate on resources (e.g., collections of information). Various operations of the example methods described herein may be performed, at least in part, by one or more processors 504 that are temporarily configured (e.g., by software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors 504 may constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors 504. Similarly, the methods described herein may be at least partially processor-implemented, with a particular processor 512, 514, or processor 504 being an example of hardware. For example, at least a portion of the operations of a method may be performed by one or more processors 504 or processor-implemented components. Additionally, one or more processors 504 may operate to support execution of the associated operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a group of computers (as an example of machine 500, including processor 504), which may be accessible via a network 534 (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). Certain performance of the operations may be distributed among processors and located across several machines, rather than just being present within a single machine 500. In some example implementations, processor 504 or components implemented by the processor may be located in a single geographic location (e.g., in a home environment, an office environment, or a server farm).In another example implementation, the processor 504 or components implemented by the processor may be distributed across several geographic locations.
[0314] FIG. 6 is a block diagram illustrating a system 600 including an example software architecture 602 that can be used in conjunction with various hardware architectures described herein. It will be appreciated that FIG. 6 is a non-limiting example of a software architecture, and that many other architectures can be implemented to facilitate the functionality described herein. The software architecture 602 may execute on hardware, such as the machine 500 of FIG. 5, which includes, among other things, a processor 504, memory / storage 506, and input / output (I / O) components 508. A representative hardware layer 604 is illustrated and may represent, for example, the machine 500 of FIG. 5. The representative hardware layer 604 includes a processing unit 606 having associated executable instructions 608. The executable instructions 608 represent the executable instructions of the software architecture 602, including implementations of methods, components, etc., described herein. The hardware layer 604 also includes at least one memory or storage module memory / storage 610, which also has the executable instructions 608. The hardware layer 604 may also include other hardware 612.
[0315] In the example architecture of FIG. 6 , software architecture 602 can be conceptualized as a stack of layers, with each layer providing specific functionality. For example, software architecture 602 may include layers such as operating system 614, libraries 616, framework / middleware 618, application 620, and presentation layer 622. In operation, application 620 or other components within a layer can invoke API calls 624 through the software stack and receive messages 626 in response to API calls 624. The illustrated layers are representative of the actual implementation, and not all software architectures have all layers. For example, some mobile or special-purpose operating systems may not provide framework / middleware 618, while others may provide such a layer. Other software architectures may include additional or different layers.
[0316] The operating system 614 may manage hardware resources and provide common services. The operating system 614 may include, for example, a kernel 628, services 630, and drivers 632. The kernel 628 may serve as an abstraction layer between the hardware and other software layers. For example, the kernel 628 may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security configuration, etc. The services 630 may provide other common services to the other software layers. The drivers 632 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 632 may include a display driver, a camera driver, a Bluetooth driver, a flash memory driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), a Wi-Fi driver, an audio driver, a power management driver, etc., depending on the hardware configuration.
[0317] Libraries 616 provide a common foundation used by applications 620 and / or other components or layers. Libraries 616 provide functions that allow other software components to perform tasks more easily than if they interfaced directly with underlying operating system 614 functions (e.g., kernel 628, services 630, drivers 632). Libraries 616 may include system libraries 634 (e.g., standard C libraries) that may provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, libraries 616 may include API libraries 636, such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.), graphics libraries (e.g., an OpenGL framework that can be used to render two-dimensional and three-dimensional graphical content on a display), database libraries (e.g., SQLite, which may provide various relational database functions), web libraries (e.g., WebKit, which may provide web browsing functions), etc. The library 616 may also include a wide variety of other libraries 638 to provide many other APIs to the application 620 and other software components / modules.
[0318] Frameworks / middleware 618 (sometimes also referred to as middleware) provide a higher-level common foundation that can be used by applications 620 or other software components / modules. For example, frameworks / middleware 618 may provide various graphical user interface functionality, high-level resource management, high-level location services, etc. Frameworks / middleware 618 may provide a wide range of other APIs that can be utilized by applications 620 or other software components / modules, some of which may be specific to a particular operating system 614 or platform.
[0319] The applications 620 include built-in applications 640 and third-party applications 642. Examples of representative built-in applications 640 may include, but are not limited to, a contacts application, a browser application, a book reader application, a location application, a media application, a messaging application, or a game application. The third-party applications 642 may include applications developed using the ANDROID® or IOS™ Software Development Kit (SDK) by an entity other than the vendor of a particular platform, and may be mobile software running on a mobile operating system, such as IOS™, ANDROID®, WINDOWS® Phone, or other mobile operating system. The third-party applications 642 can invoke API calls 624 provided by the mobile operating system (e.g., operating system 614) to facilitate the functionality described herein.
[0320] Applications 620 may create a UI using built-in operating system facilities (e.g., kernel 628, services 630, drivers 632), libraries 616, and frameworks / middleware 618 for interacting with a user of the system. Alternatively or additionally, in some systems, user interaction can be performed by a presentation layer, such as presentation layer 622. In these systems, application / component "logic" can be separated from the aspects of the application / component that interact with the user.
[0321] At least a portion of the processes described herein may be embodied in computer-readable instructions for execution by one or more processors, and thus the operations of the processes may be performed, in whole or in part, by the functional components of one or more computer systems. Accordingly, the computer-implemented processes described herein are, in some circumstances, referenced as examples. However, in other implementations, at least a portion of the operations of the computer-implemented processes described herein may be deployed in a variety of other hardware configurations. Accordingly, the computer-implemented processes described herein are not intended to be limited to the systems and configurations described with respect to FIGS. 5 and 6 , and may be implemented, in whole or in part, by one or more additional systems and / or components.
[0322] While the flowcharts described herein may depict operations as a sequential process, many of the operations may be performed in parallel or together. Moreover, the order of operations may be rearranged. A process terminates when its operations are completed. A process may correspond to a method, a procedure, an algorithm, etc. The operations of a method may be performed in whole or in part, may be performed in conjunction with some or all of the operations of other methods, and may be performed by any number of different systems, such as the systems described herein, or any portion thereof, e.g., a processor included in any of the systems. [Example]
[0323] Methylation data representing the count of methylated cytosines in the CG regions of the classification and control regions was determined for 143 samples. Samples with a molecular count of less than 50,000 were excluded. Samples were collected from subjects with several different types of cancer. For example, samples were collected from subjects with bladder cancer, breast cancer, colorectal cancer, gastric cancer, lung cancer, ovarian cancer, pancreatic cancer, and prostate cancer. A portion of the subjects with cancer were negative for homologous recombination deficiency, and a portion of the subjects with cancer were positive for homologous recombination deficiency. Subjects were considered positive for homologous recombination deficiency if germline or somatic deletions, including SNVs or indels, were detected in one of the following genes: ATM, BRCA1, BRCA2, CDK12, CHEK2, PALB2, or RAD51D.
[0324] The count of molecules included in the high distribution was determined for approximately 18,000 classification regions. The normalized molecular count for the classification region was determined based on the molecular count in the control region. A 10-fold cross-validation was performed using samples randomly divided so that one group contained 10% of the molecules used for testing and 90% of the other subset of molecules used for training.
[0325] The computational model training process involved selecting potential predictor regions by generating a logistic regression model for each classification region, where the response variable was the HRD status for a given sample and the explanatory variables were the normalized counts for each classification region and the cancer type for each sample. The top 300 regions based on the p-values generated using the normalized classification region counts were selected. A logistic regression model was generated using the normalized counts for the selected regions to predict the HRD status of the sample. Cross-validation was used to optimize the lambda value and fit the computational model to the training data. The computational model was generated using the glmnet-package in the R programming language, and the final prediction score was a probit value.
[0326] Figure 7 is a chart showing genomic regions identified as predictors of homologous recombination repair deficiency using the techniques described herein. At least some of these genomic regions, such as ST6GAL1, PAXX, BRCA1, and DNMT3A, are known to be involved in DNA repair mechanisms. The p2_adj column is an error rate adjustment of the p-value in the p2 column. Figure 8 includes several graphs showing probit values for determining whether a subject is homologous repair deficiency (HRD) positive or negative for several forms of cancer, including bladder cancer, breast cancer, colorectal cancer, gastric cancer, lung cancer, ovarian cancer, pancreatic cancer, and prostate cancer, using the techniques described herein. For several forms of cancer, a separation is observed between the probit values for determining positive and negative HRD status.
Claims
1. obtaining, with a computing system having one or more hardware processors and memory, training sequence data comprising training sequence representations from a plurality of samples, wherein each training sequence representation comprises a nucleotide sequence corresponding to a fragment of a nucleic acid contained in one sample of the plurality of samples, and wherein each sample of the plurality of samples corresponds to a subject classified as having a homology directed repair deficiency; determining, with the computing system, a subset of the training sequence representations that correspond to nucleic acids having at least a threshold amount of methylated cytosines within one or more regions of the nucleotide sequence; analyzing, with the computing system, a subset of training sequence representations to determine quantitative measures derived from the subset of training sequence representations, each quantitative measure corresponding to a classification region of a plurality of classification regions of a reference genome, each classification region of the plurality of classification regions having the threshold amount of methylated cytosines in a subject in which cancer is to be detected; analyzing, by said computing system and using one or more computational techniques, the quantitative measures of said plurality of classification regions to determine a subset of said plurality of classification regions having at least a threshold likelihood of being indicative of a homologous recombination repair deficiency; generating, by the computing system, a predictive model for determining the probability of a homology-directed repair deficiency being present in one or more additional subjects, the predictive model comprising a plurality of variables and a plurality of weights, each weight of the plurality of weights corresponding to a respective variable of the plurality of variables, each variable of the plurality of variables corresponding to a respective classification region of the subset of the plurality of classification regions, and each weight corresponding to the respective variable indicating the likelihood that the respective classification region exhibits a homology-directed repair deficiency; A method comprising:
2. analyzing, with the computing system, a subset of training sequence representations to determine additional quantitative measures derived from the subset of training sequence reads, each quantitative measure corresponding to a control region of a plurality of control regions of a reference genome, each control region of the plurality of control regions having the threshold amount of methylated cytosines in subjects in which cancer is detected and in additional subjects in which cancer is not detected; determining, by the computing system, normalized quantitative measures corresponding to the subset of the plurality of classification regions, each normalized quantitative measure being determined according to a quantitative measure corresponding to one classification region of the subset of the plurality of classification regions and the additional quantitative measure; The method of claim 1 , comprising:
3. determining the subset of the plurality of classified regions having at least a threshold likelihood of being indicative of a homologous recombination repair deficiency, determining, by the computing system, for each classification region of the plurality of classification regions, a difference between a first portion of the normalized quantitative measure derived from a sample corresponding to a subject in which a homologous recombination repair deficiency is present and a second portion of the normalized quantitative measure derived from a sample corresponding to an additional subject in which a homologous recombination repair deficiency is not present; determining, by the computing system, that an individual classification region be included in the subset of the plurality of classification regions based on the difference between the first portion of the normalized quantitative measure for the individual classification region and the second portion of the normalized quantitative measure for the individual classification region being at least a threshold difference; The method of claim 2 , comprising:
4. implementing, by the computing system, the predictive model to determine an individual probability of a homologous recombination repair deficiency in each sample of the plurality of samples based on the normalized quantitative measure corresponding to the individual sample; and determining, by the computing system, a threshold probability based on the individual probabilities to indicate the presence of a homologous recombination repair deficiency for a given subject.
5. determining, by the computing system, responsiveness to a treatment for a group of subjects, wherein cancer is detected in the group of subjects and the treatment is administered to treat the cancer; determining, by the computing system, the plurality of samples corresponding to subjects having a homologous recombination repair deficiency based on the responsiveness of a portion of the group of subjects to the treatment being at least a threshold level of responsiveness; The method according to any one of claims 1 to 4, comprising:
6. 6. The method of claim 5, wherein the treatment is a polyadenosine diphosphate (ADP) ribose polymerase (PARP) inhibitor.
7. analyzing, with the computing system, additional sequence reads from samples from a group of subjects in which cancer is detected to determine whether one or more genomic mutations are present with respect to one or more genomic regions, wherein the one or more genomic mutations correspond to a homologous recombination repair pathway; determining, with the computing system, the plurality of samples used to generate the training sequence representations by identifying a portion of the samples from the group of subjects in which the one or more genomic variations are present; The method according to any one of claims 1 to 6, comprising:
8. The method of any one of claims 1 to 7, wherein the one or more computational techniques include implementing one or more logistic regression models with elastic regularization.
9. 9. The method of any one of claims 1 to 8, comprising implementing, by the computing system, the predictive model to determine the probability of the presence of a homologous recombination repair deficiency in a plurality of additional samples, wherein the plurality of additional samples are from additional subjects, a first portion of the additional subjects having a first form of cancer detected, and a second portion of the additional subjects having a second form of cancer detected.
10. 10. The method of any one of claims 1 to 9, comprising implementing, by the computing system, the predictive model to determine the probability of the presence of a homologous recombination repair deficiency in a plurality of additional samples, wherein the plurality of additional samples are from additional subjects presenting with a single form of cancer.
11. analyzing, with the computing system, a subset of the training sequence reads to determine a group of training sequence reads corresponding to a plurality of genomic regions associated with a homology directed repair pathway; determining, by the computing system, one or more additional quantitative measures based on the number of groups of training sequence representations that correspond to at least a portion of the plurality of genomic regions; The method according to any one of claims 1 to 10, comprising:
12. determining, with the computing system, additional subsets of the training sequence representations corresponding to additional nucleic acids having methylation below an additional threshold amount; analyzing, with the computing system, the additional subset of training sequence reads to determine an additional set of training sequence representations corresponding to the plurality of genomic regions associated with the homology directed repair pathway; determining, by the computing system, one or more additional quantitative measures based on additional numbers of additional groups of the training sequence representations corresponding to at least a portion of the plurality of genomic regions; The method of claim 11 , comprising:
13. 13. The method of claim 12, further comprising analyzing, by the computing system, differences between the one or more additional quantitative measures and the one or more further quantitative measures to determine one or more additional variables for the predictive model.
14. The method of any one of claims 1 to 13, wherein the plurality of classification regions have a cytosine-guanine content of at least a threshold amount.
15. determining, by the computing system, tumor fraction estimates for a number of samples, the number of samples corresponding to subjects in which cancer is detected; analyzing, by the computing system, the tumor fraction estimate with respect to a threshold tumor fraction estimate; determining, by the computing system, the plurality of samples to use to obtain the training sequence reads based on identification of at least a portion of the number of samples having a tumor fraction estimate that corresponds to at least the threshold tumor fraction estimate; The method according to any one of claims 1 to 14, comprising:
16. obtaining, with the computing system, test sequence data from an additional subject not included in the plurality of subjects, the test sequence data comprising test sequencing representations from a sample of the additional subject, each test sequencing representation comprising a nucleotide sequence corresponding to a fragment of nucleic acid included in the additional sample, and each test sequencing read corresponding to a molecule having the threshold amount of methylated cytosines included within a region of the nucleotide; using said predictive model and said additional sequence data to determine the probability of a homologous recombination repair deficiency being present in said additional subject; The method according to any one of claims 1 to 15, comprising:
17. analyzing, by the computing system, the test sequencing reads to determine a first additional quantitative measure corresponding to each of the classification regions of the plurality of classification regions; analyzing, with the computing system, the test sequencing reads to determine a second, additional quantitative measure derived from the test sequencing reads corresponding to individual control regions of a plurality of control regions, the individual control regions of the plurality of control regions having the threshold amount of methylated cytosines in the subject in which cancer is detected and in additional subjects in which cancer is not detected; determining, by the computing system, additional normalized quantitative measures corresponding to the subset of the plurality of classification regions, each additional normalized quantitative measure being determined according to the first additional quantitative measure and the second additional quantitative measure; generating, by the computing system, an input vector comprising the normalized quantitative measure; Including, the predictive model uses the input vector to determine the probability that a homologous recombination repair deficiency is present in the additional subject.
17. The method of claim 16.
18. combining a plurality of nucleic acids from at least one of a subject's blood or tissue with a solution containing an amount of a methyl-binding domain (MBD) protein to produce a nucleic acid-MBD protein solution; performing multiple washes of the nucleic acid-MBD protein solution with a salt solution to produce several nucleic acid fractions, each nucleic acid fraction having a threshold number of methylated cytosines within a region of the plurality of nucleic acids having at least a threshold cytosine-guanine content; The method of any one of claims 1 to 17, comprising:
19. 20. The method of claim 18, wherein one of the plurality of washes is performed with a solution having one concentration of sodium chloride (NaCl) to produce one of the several nucleic acid fractions having a range of binding energies for MBD protein.
20. determining that the first nucleic acid fraction is associated with a first distribution of a plurality of distributions of nucleic acids, the first distribution corresponding to a first range of binding energies for MBD proteins; causing binding of a first molecular barcode to nucleic acids of the first nucleic acid fraction, the first molecular barcode associated with the first distribution; determining that a second nucleic acid fraction is associated with a second distribution of nucleic acids in the plurality of distributions, the second distribution corresponding to a second range of binding energies for MBD protein that is different from the first range of binding energies for MBD protein; causing binding of a second molecular barcode to nucleic acids of the second nucleic acid fraction, wherein the second molecular barcode is associated with the second distribution; 20. The method of claim 18, comprising:
21. combining at least a portion of said several nucleic acid fractions with a quantity of a restriction enzyme that cuts molecules having one or more unmethylated cytosines to produce at least a portion of said plurality of samples used to generate said training sequence representations; said threshold amount of methylated cytosines corresponds to a minimum frequency of methylated cytosines within a region having at least said threshold cytosine-guanine content.
20. The method of claim 18, comprising:
22. combining at least a portion of said several nucleic acid fractions with a quantity of a restriction enzyme that cuts molecules having one or more methylated cytosines to produce at least a portion of said plurality of samples used to generate said training sequence representations; said threshold amount of unmethylated cytosines corresponds to the highest frequency of uncleaved methylated cytosines within a region having at least said threshold cytosine-guanine content.
20. The method of claim 18, comprising:
23. 23. The method of any one of claims 1 to 22, comprising administering a treatment to the subject based on a determination that a homologous recombination repair deficiency is present in the subject.